Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2,848
datasets available to search
ShareScore release 0.9.0
Dataset results
2,848 results for “sequence data”
Data from: miR-122, small RNA annealing and sequence mutations alter the predicted structure of the Hepatitis C virus 5′ UTR RNA to stabilize and promote viral RNA accumulation
Annealing of the liver-specific microRNA, miR-122, to the Hepatitis C virus (HCV) 5′ UTR is required for efficient virus replication. By using siRNAs to pressure escape mutations, 30 replication-competent HCV genomes having nucleotide changes in the conserved 5′ untranslated region (UTR) were identified. In silico analysis predicted that miR-122 annealing induces canonical HCV genomic 5′ UTR RNA folding, and mutant 5′ UTR sequences that promoted miR-122-independent HCV replication favored the formation of the canonical RNA structure, even in the absence of miR-122. Additionally, some mutant viruses adapted to use the siRNA as a miR-122-mimic. We further demonstrate that small RNAs that anneal with perfect complementarity to the 5′ UTR stabilize and promote HCV genome accumulation. Thus, HCV genome stabilization and life-cycle promotion does not require the specific annealing pattern demonstrated for miR-122 nor 5′ end annealing or 3′ overhanging nucleotides. Replication promotion by perfect-match siRNAs was observed in Ago2 knockout cells revealing that other Ago isoforms can support HCV replication. At last, we present a model for miR-122 promotion of the HCV life cycle in which miRNA annealing to the 5′ UTR, in conjunction with any Ago isoform, modifies the 5′ UTR structure to stabilize the viral genome and promote HCV RNA accumulation.
Data from: Genome sequencing and comparative analysis of three Chlamydia pecorum strains associated with different pathogenic outcomes
Background: Chlamydia pecorum is the causative agent of a number of acute diseases, but most often causes persistent, subclinical infection in ruminants, swine and birds. In this study, the genome sequences of three C. pecorum strains isolated from the faeces of a sheep with inapparent enteric infection (strain W73), from the synovial fluid of a sheep with polyarthritis (strain P787) and from a cervical swab taken from a cow with metritis (strain PV3056/3) were determined using Illumina/Solexa and Roche 454 genome sequencing. Results: Gene order and synteny was almost identical between C. pecorum strains and C. psittaci. Differences between C. pecorum and other chlamydiae occurred at a number of loci, including the plasticity zone, which contained a MAC/perforin domain protein, two copies of a >3400 amino acid putative cytotoxin gene and four (PV3056/3) or five (P787 and W73) genes encoding phospholipase D. Chlamydia pecorum contains an almost intact tryptophan biosynthesis operon encoding trpABCDFR and has the ability to sequester kynurenine from its host, however it lacks the genes folA, folKP and folB required for folate metabolism found in other chlamydiae. A total of 15 polymorphic membrane proteins were identified, belonging to six pmp families. Strains possess an intact type III secretion system composed of 18 structural genes and accessory proteins, however a number of putative inc effector proteins widely distributed in chlamydiae are absent from C. pecorum. Two genes encoding the hypothetical protein ORF663 and IncA contain variable numbers of repeat sequences that could be associated with persistence of infection. Conclusions: Genome sequencing of three C. pecorum strains, originating from animals with different disease manifestations, has identified differences in ORF663 and pseudogene content between strains and has identified genes and metabolic traits that may influence intracellular survival, pathogenicity and evasion of the host immune system.
Data from: Diversity measures in environmental sequences are highly dependent on alignment quality—data from ITS and new LSU primers targeting basidiomycetes
The ribosomal DNA comprised of the ITS1-5.8S-ITS2 regions is widely used as a fungal marker in molecular ecology and systematics but cannot be aligned with confidence across genetically distant taxa. In order to study the diversity of Agaricomycotina in forest soils, we designed primers targeting the more alignable 28S (LSU) gene, which should be more useful for phylogenetic analyses of the detected taxa. This paper compares the performance of the established ITS1F/4B primer pair, which targets basidiomycetes, to that of two new pairs. Key factors in the comparison were the diversity covered, off-target amplification, rarefaction at different Operational Taxonomic Unit (OTU) cutoff levels, sensitivity of the method used to process the alignment to missing data and insecure positional homology, and the congruence of monophyletic clades with OTU assignments and BLAST-derived OTU names. The ITS primer pair yielded no off-target amplification but also exhibited the least fidelity to the expected phylogenetic groups. The LSU primers give complementary pictures of diversity, but were more sensitive to modifications of the alignment such as the removal of difficult-to align stretches. The LSU primers also yielded greater numbers of singletons but also had a greater tendency to produce OTUs containing sequences from a wider variety of species as judged by BLAST similarity. We introduced some new parameters to describe alignment heterogeneity based on Shannon entropy and the extent and contents of the OTUs in a phylogenetic tree space. Our results suggest that ITS should not be used when calculating phylogenetic trees from genetically distant sequences obtained from environmental DNA extractions and that it is inadvisable to define OTUs on the basis of very heterogeneous alignments.
Data from: Intraspecific trait variation and colonization sequence alter community assembly and disease epidemics
When individuals from multiple populations colonize a new habitat patch, intraspecific trait variation can make the arrival order of colonists an important factor for subsequent population and community dynamics. In particular, intraspecific priority effects (IPEs) allow early arrivers to limit the growth or establishment of later arrivers, even when competitively inferior on a per-capita basis. Through their effects on genes and traits, IPEs can alter short-term growth and long-term evolutionary change in single species metapopulations. Given their importance for intraspecific interactions, IPEs in a dominant species have the potential to affect the composition of entire communities. We conducted an experiment to determine whether and how arrival order and IPEs in the zooplankter Daphnia pulex affected its interactions with both competitors (the cladoceran Simocephalus vetulus) and parasites (the virulent fungus Metschnikowia bicuspidata). We found strong evidence for IPEs in Daphnia, as early arrivers inhibited late arrivers even when competitively inferior. These IPEs in Daphnia altered both the establishment success of interspecific competitors and the size of disease epidemics: early colonization by fast-growing D. pulex led to large Daphnia populations and low competitor establishment, but large disease epidemics. Early colonization by slow-growing D. pulex, on the other hand, resulted in small Daphnia populations with high competitor establishment, but smaller disease epidemics. Overall, our results demonstrate the importance of intraspecific variation and arrival order for community dynamics, and highlight IPEs as a general mechanism driving variation in natural communities.
Data from: Deconstruction of archaeal genome depict strategic consensus in core pathways coding sequence assembly
A comprehensive in silico analysis of 71 species representing the different taxonomic classes and physiological genre of the domain Archaea was performed. These organisms differed in their physiological attributes, particularly oxygen tolerance and energy metabolism. We explored the diversity and similarity in the codon usage pattern in the genes and genomes of these organisms, emphasizing on their core cellular pathways. Our thrust was to figure out whether there is any underlying similarity in the design of core pathways within these organisms. Analyses of codon utilization pattern, construction of hierarchical linear models of codon usage, expression pattern and codon pair preference pointed to the fact that, in the archaea there is a trend towards biased use of synonymous codons in the core cellular pathways and the Nc-plots appeared to display the physiological variations present within the different species. Our analyses revealed that aerobic species of archaea possessed a larger degree of freedom in regulating expression levels than could be accounted for by codon usage bias alone. This feature might be a consequence of their enhanced metabolic activities as a result of their adaptation to the relatively O2-rich environment. Species of archaea, which are related from the taxonomical viewpoint, were found to have striking similarities in their ORF structuring pattern. In the anaerobic species of archaea, codon bias was found to be a major determinant of gene expression. We have also detected a significant difference in the codon pair usage pattern between the whole genome and the genes related to vital cellular pathways, and it was not only species-specific but pathway specific too. This hints towards the structuring of ORFs with better decoding accuracy during translation. Finally, a codon-pathway interaction in shaping the codon design of pathways was observed where the transcription pathway exhibited a significantly different coding frequency signature.
Data from: 28 year temporal sequence of epidemic dynamics in a natural rust – host plant metapopulation
A long-term study of disease dynamics caused by the rust Uromyces valerianae in 31 discrete populations of Valeriana salina provides a rare opportunity to explore extended temporal patterns in the epidemiology of a natural host-pathogen metapopulation. Over a 28-year period, pathogen population dynamics varied across the metapopulation with disease incidence (presence/absence), prevalence (% plants infected) and severity (% leaf area covered by lesions) all showing strong population and year effects, indicative of heterogeneity among years and host populations in the suitability of conditions for the pathogen. Disease incidence within individual host populations was significantly affected by host population size, disease prevalence the previous year and the proximity of neighbouring populations infected in the current year. After accounting for these variables there was still a marked temporal component with winter sea level having a significant effect; as did summer rainfall in the second part of the study period (1997-2011). Disease prevalence was also effected by host population size and disease prevalence in the previous year. However, it was less affected by spatial aspects of disease spread than was disease incidence. Winter sea level and June rainfall significantly affected disease prevalence. Assessment of disease impact on plant performance found strong variation in disease severity associated with the aspect and positioning of host populations. Plants growing in lower disease environments produced significantly more seeds than those growing in high disease sites. Significant variation in reaction to infection by U. valerianae was detected among plants within four populations and between these different populations. Synthesis. The epidemiology of U. valerianae was highly influenced by host population size, previous disease and distance. After accounting for these factors, there was a clear temporal signal of change in disease incidence linked to winter sea level and summer rainfall. These patterns reinforce the importance of considering interactions in multiple populations over long periods of time in order to obtain a clear picture of the variability of disease-induced selection pressures across time and space. The behaviour of the pathogen fitted that predicted for a metapopulation with considerable asynchrony in epidemiological patterns among demes.
Data from: GHOST: Recovering Historical Signal from Heterotachously-evolved Sequence Alignments
<p><span>Molecular sequence data that have evolved under the influence of heterotachous evolutionary processes are known to mislead phylogenetic inference. We introduce the General Heterogeneous evolution On a Single Topology (GHOST) model of sequence evolution, implemented under a maximum-likelihood framework in the phylogenetic program IQ-TREE (</span><a class="link link-uri" href="http://www.iqtree.org/">http://www.iqtree.org</a><span>). Simulations show that using the GHOST model, IQ-TREE can accurately recover the tree topology, branch lengths, and substitution model parameters from heterotachously evolved sequences. We investigate the performance of the GHOST model on empirical data by sampling phylogenomic alignments of varying lengths from a plastome alignment. We then carry out inference under the GHOST model on a phylogenomic data set composed of 248 genes from 16 taxa, where we find the GHOST model concurs with the currently accepted view, placing turtles as a sister lineage of archosaurs, in contrast to results obtained using traditional variable rates-across-sites models. Finally, we apply the model to a data set composed of a sodium channel gene of 11 fish taxa, finding that the GHOST model is able to elucidate a subtle component of the historical signal, linked to the previously established convergent evolution of the electric organ in two geographically distinct lineages of electric fish. We compare inference under the GHOST model to partitioning by codon position and show that, owing to the minimization of model constraints, the GHOST model offers unique biological insights when applied to empirical data.</span></p>
Data from: Successful recovery of nuclear protein-coding genes from small insects in museums using illumina sequencing
In this paper we explore high-throughput Illumina sequencing of nuclear protein-coding, ribosomal, and mitochondrial genes in small, dried insects stored in natural history collections. We sequenced one tenebrionid beetle and 12 carabid beetles ranging in size from 3.7 to 9.7 mm in length that have been stored in various museums for 4 to 84 years. Although we chose a number of old, small specimens for which we expected low sequence recovery, we successfully recovered at least some low-copy nuclear protein-coding genes from all specimens. For example, in one 56-year-old beetle, 4.4 mm in length, our de novo assembly recovered about 63% of approximately 41,900 nucleotides in a target suite of 67 nuclear protein-coding gene fragments, and 70% using a reference-based assembly. Even in the least successfully sequenced carabid specimen, reference-based assembly yielded fragments that were at least 50% of the target length for 34 of 67 nuclear protein-coding gene fragments. Exploration of alternative references for reference-based assembly revealed few signs of bias created by the reference. For all specimens we recovered almost complete copies of ribosomal and mitochondrial genes. We verified the general accuracy of the sequences through comparisons with sequences obtained from PCR and Sanger sequencing, including of conspecific, fresh specimens, and through phylogenetic analysis that tested the placement of sequences in predicted regions. A few possible inaccuracies in the sequences were detected, but these rarely affected the phylogenetic placement of the samples. Although our sample sizes are low, an exploratory regression study suggests that the dominant factor in predicting success at recovering nuclear protein-coding genes is a high number of Illumina reads, with success at PCR of COI and killing by immersion in ethanol being secondary factors; in analyses of only high-read samples, the primary significant explanatory variable was body length, with small beetles being more successfully sequenced.
FIGURES 1‒3 in A revision of the genus Lepidobrya Womersley (Collembola: Entomobryidae) based on morphology and sequence data of the genotype
FIGURES 1‒3. Colour pattern of two individuals in Lepidobrya mawsoni. Scale bars: 1 mm.
FIGURE 13 in Taxonomy of the African army ant Dorylus gribodoi Emery, 1892 (Hymenoptera, Formicidae) — new insights from DNA sequence data and morphology
FIGURE 13. Dorsal view of the hypopygium of the D. gribodoi queen (HAD 2.53 mm).
FIGURE 7 in Taxonomy of the African army ant Dorylus gribodoi Emery, 1892 (Hymenoptera, Formicidae) — new insights from DNA sequence data and morphology
FIGURE 7. Lateral view of the head of a large D. gribodoi worker (HW 2.79 mm), Taï, Ivory Coast.
FIGURE 6 in Taxonomy of the African army ant Dorylus gribodoi Emery, 1892 (Hymenoptera, Formicidae) — new insights from DNA sequence data and morphology
FIGURE 6. Lateral overview of a large D. emeryi worker from Taï, Ivory Coast.
FIGURE 4 in Taxonomy of the African army ant Dorylus gribodoi Emery, 1892 (Hymenoptera, Formicidae) — new insights from DNA sequence data and morphology
FIGURE 4. Dorylus gribodoi male: genital capsule and subgenital plate.
FIGURE 3 in Taxonomy of the African army ant Dorylus gribodoi Emery, 1892 (Hymenoptera, Formicidae) — new insights from DNA sequence data and morphology
FIGURE 3. Frontal view of the head of a Dorylus gribodoi male from Taï, Ivory Coast.
FIGURE 2 in Taxonomy of the African army ant Dorylus gribodoi Emery, 1892 (Hymenoptera, Formicidae) — new insights from DNA sequence data and morphology
FIGURE 2. Lateral overview of a Dorylus gribodoi male from Taï, Ivory Coast.
FIGURE 5 in Taxonomy of the African army ant Dorylus gribodoi Emery, 1892 (Hymenoptera, Formicidae) — new insights from DNA sequence data and morphology
FIGURE 5. Lateral overview of a large Dorylus gribodoi worker from Taï, Ivory Coast.
FIGURE 1 in Taxonomy of the African army ant Dorylus gribodoi Emery, 1892 (Hymenoptera, Formicidae) — new insights from DNA sequence data and morphology
FIGURE 1. Definition of the subgenital plate measurements.
FIGURE 11 in Taxonomy of the African army ant Dorylus gribodoi Emery, 1892 (Hymenoptera, Formicidae) — new insights from DNA sequence data and morphology
FIGURE 11. Lateral overview of the D. gribodoi queen collected from a nest at Lamto, Ivory Coast.
FIGURE 8 in Taxonomy of the African army ant Dorylus gribodoi Emery, 1892 (Hymenoptera, Formicidae) — new insights from DNA sequence data and morphology
FIGURE 8. Lateral view of the head of a large D. emeryi worker (HW 3.62 mm) from Taï, Ivory Coast.
Raw sequencing data of Anaplamsa phagocytophilum loci (ankA, msp4, groEL) obtained from 454 and parameter files to clean these data using MOTHUR
<p>A compressed archive including: i) raw sequences in ssf file; ii) Mothur oligo files to sort out sequences among loci and individual samples.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.