Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
521
datasets available to search
ShareScore release 0.9.0
Dataset results
521 results for “DNA sequence data”
FIGURES 9–10. 9. Adelphydraenaspinosa n in Three new species and DNA sequence data of the rare South American water beetle genus Adelphydraena Perkins, 1989 (Coleoptera: Hydraenidae)
FIGURES 9–10. 9. Adelphydraenaspinosa n. sp., holotypeaedeagus. 10. Adelphydraenaorchymonti Perkins, non-type, aedeagus.
Data from: Changes in soil microbial communities in post mine ecological restoration: implications for monitoring using high throughput DNA sequencing
<p>The ecological restoration of ecosystem services and biodiversity is a key intervention used to reverse the impacts of anthropogenic activities such as mining. Assessment of the performance of restoration against completion criteria relies on biodiversity monitoring. However, monitoring usually overlooks soil microbial communities (SMC), despite increased awareness of their pivotal role in many ecological functions. Recent advances in cost, scalability and technology has led to DNA sequencing being considered as a cost-effective biological monitoring tool, particularly for otherwise difficult to survey groups such as microbes. However, such approaches for monitoring complex restoration sites such as post-mined landscapes have not yet been tested. Here we examine bacterial and fungal communities across chronosequences of mine site restoration at three locations in Western Australia to determine if there are consistent changes in SMC diversity, community composition and functional capacity. Although we detected directional changes in community composition indicative of microbial recovery, these were inconsistent between locations and microbial taxa (bacteria or fungi). Assessing functional diversity provided greater understanding of changes in site conditions and microbial recovery than could be determined through assessment of community composition alone. These results demonstrate that <span>high-throughput amplicon sequencing of environmental DNA (eDNA)</span> is an effective approach for monitoring the complex changes in SMC following restoration. Future monitoring of mine site restoration using eDNA should consider archiving samples to provide improved understanding of changes in communities over time. Expansion to include other biological groups (e.g. soil fauna) and substrates would also provide a more holistic understanding of biodiversity recovery. </p>
Data from: Phylogenetic relationships and timing of diversification in gonorynchiform fishes inferred using nuclear gene DNA sequences (Teleostei: Ostariophysi)
The Gonorynchiformes are the sister lineage of the species-rich Otophysi and provide important insights into the diversification of ostariophysan fishes. Phylogenies of gonorynchiforms inferred using morphological characters and mtDNA gene sequences provide differing resolutions with regard to the sister lineage of all other gonorynchiforms (Chanos vs. Gonorynchus) and support for monophyly of the two miniaturized lineages Cromeria and Grasseichthys. In this study the phylogeny and divergence times of gonorynchiforms are investigated with DNA sequences sampled from nine nuclear genes and a published morphological character matrix. Bayesian phylogenetic analyses reveal substantial congruence among individual gene trees with inferences from eight genes placing Gonorynchus as the sister lineage to all other gonorynchiforms. Seven gene trees resolve Cromeria and Grasseichthys as a clade, supporting previous inferences using morphological characters. Phylogenies resulting from either concatenating the nuclear genes, performing a multispecies coalescent species tree analysis, or combining the morphological and nuclear gene DNA sequences resolve Gonorynchus as the living sister lineage of all other gonorynchiforms, strongly support the monophyly of Cromeria and Grasseichthys, and resolve a clade containing Parakneria, Cromeria, and Grasseichthys. The morphological dataset, which includes 13 gonorynchiform fossil taxa that range in age from Early Cretaceous to Eocene, was analyzed in combination with DNA sequences from the nine nuclear genes and a relaxed molecular clock to estimate times of evolutionary divergence. This "tip dating" strategy accommodates uncertainty in the phylogenetic resolution of fossil taxa that provide calibration information in the relaxed molecular clock analysis. The estimated age of the most recent common ancestor (MRCA) of living gonorynchiforms is slightly older than estimates from previous node dating efforts, but the molecular tip dating estimated ages of Kneriinae (Kneria, Parakneria, Cromeria, and Grasseichthys) and the two paedomorphic lineages, Cromeria and Grasseichthys, are considerably younger.
Data from: Two new species of Limbodessus diving beetles from New Guinea - short verbal descriptions flanked by online content (digital photography, μCT scans, drawings and DNA sequence data)
Background: To date only one species of Limbodessus diving beetles has been reported from the Island of New Guinea, L. compactus (Clark, 1862), which is widerspread in the Australian region. New information: We describe two new species of microendemic New Guinea Limbodessus and use a compact descriptive format flanked by enriched online content in wiki powered species pages. Limbodessus baliem sp.n. is described from ca. 1,600 m altitude in the Baliem Valley of Papua and Limbodessus alexanderi sp.n. from >3,000 m altitude north of Sugapa, Papua. Based on our analysis, we also transfer three species from other genera to Limbodessus Guignot, 1939, with the following changes: Limbodessus deflectus (Ordish, 1966), new combination; Limbodessus leveri (J. Balfour-Browne, 1944), new combination; and Limbodessus plicatus (Sharp, 1882), new combination.
Data from: Phylogenetic systematics of subtribe Spiranthinae (Orchidaceae: Orchidoideae: Cranichideae) based on nuclear and plastid DNA sequences of a nearly complete generic sample
Subtribe Spiranthinae is the most species-rich lineage of terrestrial Neotropical orchids, encompassing > 500 species and 40 genera. We conducted maximum parsimony and maximum likelihood phylogenetic analyses of DNA sequence data of plastid matK-trnK and trnL-trnF and nuclear ribosomal ITS sequences for 36 genera and 182 species of Spiranthinae plus appropriate outgroups. The results strongly support monophyly of Spiranthinae (minus Discyphus, Discyphinae and Galeottiella, Galeottiellinae) and five major lineages, namely monospecific Cotylolabium (sister to the remaining Spiranthinae) and the Eurystyles, Pelexia, Spiranthes and Stenorrhynchos clades. Eighteen of the 27 genera of Spiranthinae for which more than one species was included in our analyses are monophyletic. Paraphyly of large genera, such as Cyclopogon and Sarcoglottis, resulted from segregation of particular species or groups of species exhibiting minor modifications of structures directly involved in pollination (e.g. nectary, rostellum and viscidium). Conversely, polyphyly has resulted from convergent evolution of floral attributes in distantly related species (e.g. Mesadenus). Some of the morphological characters used traditionally for generic delimitation and in non-molecular cladistic analyses of Spiranthinae are discussed against the evolutionary framework set by our molecular trees, emphasizing putative synapomorphies and problems derived from inappropriate character coding or incorrect homology assessments. Our ancestral area analysis indicates that Spiranthinae originated in eastern South America, with subsequent migrations and secondary radiations in Mesoamerica and North America, plus a derived migration from the latter region to the Old World (Spiranthes).
Data from: High-throughput sequencing of nematode communities from total soil DNA extractions
Background: Nematodes are extremely diverse and numbers of species are predicted to be more than a million. Studies on nematode diversity are difficult and laborious using standard methods such as identification based on morphology and therefore high-throughput sequencing is an attractive alternative. Generally, primers that have been used for generating amplicons for sequencing are not nematode specific and also amplify other groups such as fungi and plantae. Thus a nematode enrichment step must be included that may introduce biases. Results: An amplification strategy, including a new primer, which selectively amplifies nematodes and other metazoans was developed. When this strategy was tested on DNA templates from a set of 22 agricultural soils, we obtained 64.4 % sequences of nematode origin in total, whereas the remaining sequences were almost entirely metazoan. The nematode sequences were derived from a broad taxonomic range and most sequences were from nematode taxa that have previously been found to be abundant in soil such as Tylenchida, Rhabditida, Dorylaimida, Triplonchida and Araeolaimida. Conclusions: This amplification and sequencing strategy for assessing nematode diversity was demonstrated to be able to collect a broad taxonomy of nematodes without prior enrichment and thus the method will be highly valuable in ecological studies of nematodes. Keywords: nematode, community, next-generation sequencing, SSU, diversity, 18S, rDNA
Data from: Transatlantic secondary contact in Atlantic salmon, comparing microsatellites, a SNP array, and Restriction Associated DNA sequencing for the resolution of complex spatial structure
Identification of discrete and unique assemblages of individuals or populations is central to the management of exploited species. Advances in population genomics provide new opportunities for re-evaluating existing conservation units but comparisons among approaches remain rare. We compare the utility of RAD-seq, a single nucleotide polymorphism (SNP) array and a microsatellite panel to resolve spatial structuring under a scenario of possible trans-Atlantic secondary contact in a threatened Atlantic Salmon, Salmo salar, population in southern Newfoundland. Bayesian clustering indentified two large groups subdividing the existing conservation unit and multivariate analyses indicated significant similarity in spatial structuring among the three data sets. mtDNA alleles diagnostic for European ancestry displayed increased frequency in southeastern Newfoundland and were correlated with spatial structure in all marker types. Evidence consistent with introgression among these two groups was present in both SNP data sets but not the microsatellite data. Asymmetry in the degree of introgression was also apparent in SNP data sets with evidence of gene flow towards the east or European type. This work highlights the utility of RAD-seq based approaches for the resolution of complex spatial patterns, resolves a region of trans-Atlantic secondary contact in Atlantic Salmon in Newfoundland and demonstrates the utility of multiple marker comparisons in identifying dynamics of introgression.
Data from: Pleistocene climate change and phylogeographic structure of the Gymnocarpos przewalskii (Caryophyllaceae) in the northwest China: Evidence from plastid DNA, ITS sequences, and Microsatellite
Northwestern China has a wealth of endemic species, which has been hypothesized to be affected by the complex paleoclimatic and paleogeographic history during Quaternary. In this paper, we used Gymnocarpos przewalskii as a model to address the evolutionary history and current population genetic structure of species in northwestern China. We employed two chloroplast DNA fragments (rps16 and psbB‐psbI), one nuclear DNA fragment (ITS), and simple sequence repeat (SSRs) to investigate the spatial genetic pattern of G. przewalskii. High genetic diversity (cpDNA: hS = 0.330, hT = 0.866; ITS: hS = 0.458, hT = 0.872) was identified in almost all populations, and most of the population have private haplotypes. Moreover, multimodal mismatch distributions were observed and estimates of Tajima's D and Fu's FS tests did not identify significantly departures from neutrality, indicating that recent expansion of G. przewalskii was rejected. Thus, we inferred that G. przewalskii survived generally in northwestern China during the Pleistocene. All data together support the genotypes of G. przewalskii into three groups, consistent with their respective geographical distributions in the western regions—Tarim Basin, the central regions—Hami Basin and Hexi Corridor, and the eastern regions—Alxa Desert and Wulate Prairie. Divergence among most lineages of G. przewalskii occurred in the Pleistocene, and the range of potential distributions is associated with glacial cycles. We concluded that climate oscillation during Pleistocene significantly affected the distribution of the species.
Data from: DNA sequence variation among conspecific accessions of the legume Coursetia caribaea reveals geographically localized clades here ranked as species
Coursetia caribaea is geographically and morphologically the most variable species in the genus Coursetia and in the tribe Robinieae (Leguminosae, Papilionoideae). Because of potentially undetected species, we assessed the phylogenetic relationships among the eight taxonomic varieties of C. caribaea. Sampling included nuclear ribosomal internal transcribed spacer sequences from 489 Robinieae accessions representing all varieties of C. caribaea and 38 of the 40 species of Coursetia, in addition to chloroplast trnD-trnT sequences from 186 accessions. Separate and combined phylogenetic analyses resolved a clade of conspecific accessions of the Bolivian C. caribaea var. astragalina as sister to the central Andean Coursetia grandiflora clade. Also distantly related to Coursetia caribaea var. caribaea accessions were those of the coastal Oaxacan C. caribaea var. pacifica, which formed the sister clade to accessions of the central Andean C. caribaea var. ochroleuca. The estimated mean ages of the stem clades for these three lineages, 11, 7.7, and 7.7 Ma, respectively, contrasted to the estimated mean ages of the corresponding crown clades of 0, 0, and 1.5 Ma. The contrasting stem and crown ages suggest that these taxa, appropriately ranked as species, Coursetia astragalina, Coursetia diversifolia, and Coursetia ochroleuca, each have persisted over evolutionary time frames as distinct geographically localized populations in seasonally dry tropical forests and woodlands.
Data from: Estimation of a killer whale (Orcinus orca) population's diet using sequencing analysis of DNA from feces
Estimating diet composition is important for understanding interactions between predators and prey and thus illuminating ecosystem function. The diet of many species, however, is difficult to observe directly. Genetic analysis of fecal material collected in the field is therefore a useful tool for gaining insight into wild animal diets. In this study, we used high-throughput DNA sequencing to quantitatively estimate the diet composition of an endangered population of wild killer whales (Orcinus orca) in their summer range in the Salish Sea. We combined 175 fecal samples collected between May and September from five years between 2006 and 2011 into 13 sample groups. Two known DNA composition control groups were also created. Each group was sequenced at a ~330bp segment of the 16s gene in the mitochondrial genome using an Illumina MiSeq sequencing system. After several quality controls steps, 4,987,107 individual sequences were aligned to a custom sequence database containing 19 potential fish prey species and the most likely species of each fecal-derived sequence was determined. Based on these alignments, salmonids made up >98.6% of the total sequences and thus of the inferred diet. Of the six salmonid species, Chinook salmon made up 79.5% of the sequences, followed by coho salmon (15%). Over all years, a clear pattern emerged with Chinook salmon dominating the estimated diet early in the summer, and coho salmon contributing an average of >40% of the diet in late summer. Sockeye salmon appeared to be occasionally important, at >18% in some sample groups. Non-salmonids were rarely observed. Our results are consistent with earlier results based on surface prey remains, and confirm the importance of Chinook salmon in this population's summer diet.
Data from: Congruent species delimitation of two controversial gold-thread nanmu tree species based on morphological and restriction site-associated DNA sequencing data
Species delimitation is fundamental to conservation and sustainable use of economically important forest tree species. However, the delimitation of two highly valued gold-thread nanmu species (Phoebe bournei and P. zhennan) has been confusing and debated. To address this problem, we integrated morphology and restriction site-associated DNA sequencing (RADseq) to define their species boundaries. We obtained highly consistent results from both data sets, supporting two distinct lineages corresponding to P. bournei and P. zhennan. In Phoebe bournei, higher order leaf venation is more prominent, petioles are thicker and leaf apex angle is narrower, compared to P. zhennan. Both data sets also showed that putative P. bournei localities from north-eastern Guizhou were P. zhennan. The two species have different distributions and only overlap in the Wuling Mountains. Phoebe bournei occurs mainly in Central Fujian, southern Jiangxi, the Nanling Mountains and the Wuling Mountains, whereas P. zhennan is found in the adjoining eastern regions of the Qionglai Mountains, the Southern Sichuan Hills and the Wuling Mountains. The improved delimitation of P. bournei and P. zhennan and clarification of their ranges provide a better guidance for conservation and sustainable utilization of these tree species.
Data from: Developing nuclear DNA phylogenetic markers in the angiosperm genus Leucadendron (Proteaceae): a next-generation sequencing transcriptomic approach
Despite the recent advances in generating molecular data, reconstructing species-level phylogenies for non-models groups remains a challenge. The use of a number of independent genes is required to resolve phylogenetic relationships, especially for groups displaying low polymorphism. In such cases, low-copy nuclear exons and non-coding regions, such as 3′ untranslated regions (3′-UTRs) or introns, constitute a potentially interesting source of nuclear DNA variation. Here, we present a methodology meant to identify new nuclear orthologous markers using both public-nucleotide databases and transcriptomic data generated for the group of interest by using next generation sequencing technology. To identify PCR primers for a non-model group, the genus Leucadendron (Proteaceae), we adopted a framework aimed at minimizing the probability of paralogy and maximizing polymorphism. We anchored when possible the right-hand primer into the 3′-UTR and the left-hand primer into the coding region. Seven new nuclear markers emerged from this search strategy, three of those included 3′-UTRs. We further compared the phylogenetic potential between our new markers and the ribosomal internal transcribed spacer region (ITS). The sequenced 3′-UTRs yielded higher polymorphism rates than the ITS region did. We did not find strong incongruences with the phylogenetic signal contained in the ITS region and the seven new designed markers but they strongly improved the phylogeny of the genus Leucadendron. Overall, this methodology is efficient in isolating orthologous loci and is valid for any non-model group given the availability of transcriptomic data.
Data from: Restriction-site-associated DNA sequencing reveals a cryptic viburnum species on the North American coastal plain
Species are the starting point for most studies of ecology and evolution, but the proper circumscription of species can be extremely difficult in morphologically variable lineages, and there are still few convincing examples of molecularly-informed species delimitation in plants. We focus here on the Viburnum nudum complex, a highly variable clade that is widely distributed in eastern North America. Taxonomic treatments have mostly divided this complex into northern (V. nudum var. cassinoides) and southern (V. nudum var. nudum) entities, but additional names have been proposed. We used multiple lines of evidence, including RADseq, morphological, and geographic data, to test how many independently evolving lineages exist within the V. nudum complex. Genetic clustering and phylogenetic methods revealed three distinct groups—one lineage that is highly divergent, and two others that are recently diverged and morphologically similar. A combination of evidence that includes reciprocal monophyly, lack of introgression, and discrete rather than continuous patterns of variation supports the recognition of all three lineages as separate species. These results identify a surprising case of cryptic diversity in which two broadly sympatric species have consistently been lumped in taxonomic treatments. The clarity of our findings is directly related to the dense sampling and high quality genetic data in this study. We argue that there is a critical need for carefully sampled and integrative species delimitation studies to clarify species boundaries even in well-known plant lineages. Studies following the model that we have developed here are likely to identify many more cryptic lineages and will fundamentally improve our understanding of plant speciation and patterns of species richness.
Data from: Massively parallel multiplex DNA sequencing for specimen identification using an Illumina MiSeq platform
Genetic information is a valuable component of biosystematics, especially specimen identification through the use of species-specific DNA barcodes. Although many genomics applications have shifted to High-Throughput Sequencing (HTS) or Next-Generation Sequencing (NGS) technologies, sample identification (e.g., via DNA barcoding) is still most often done with Sanger sequencing. Here, we present a scalable double dual-indexing approach using an Illumina Miseq platform to sequence DNA barcode markers. We achieved 97.3% success by using half of an Illumina Miseq flowcell to obtain 658 base pairs of the cytochrome c oxidase I DNA barcode in 1,010 specimens from eleven orders of arthropods. Our approach recovers a greater proportion of DNA barcode sequences from individuals than does conventional Sanger sequencing, while at the same time reducing both per specimen costs and labor time by nearly 80%. In addition, the use of HTS allows the recovery of multiple sequences per specimen, for deeper analysis of genetic variation in target gene regions.
Data from: ITS all right mama: Investigating the formation of chimeric sequences in the ITS2 region by DNA metabarcoding analyses of fungal mock communities of different complexities
The formation of chimeric sequences can create significant methodological bias in PCR-based DNA metabarcoding analyses. During mixed-template amplification of barcoding regions, chimera formation is frequent and well documented. However, profiling of fungal communities typically uses the more variable rDNA region ITS. Due to a larger research community, tools for chimera detection have been developed mainly for the 16S/18S markers. However, these tools are widely applied to the ITS region without verification of their performance. We examined the rate of chimera formation during amplification and 454 sequencing of the ITS2 region from fungal mock communities of different complexities. We evaluated the chimera detecting ability of two common chimera-checking algorithms: Perseus and UCHIME. Large proportions of the chimeras reported were false positives. No false negatives were found in the dataset. Verified chimeras accounted for only 0.2% of the total ITS2 reads, which is considerably less than what is typically reported in 16S and 18S metabarcoding analyses. Verified chimeric "parent sequences" had significantly higher percent identity to one another than to random members of the mock communities. Community complexity increased the rate of chimera formation. GC content was higher around the verified chimeric break points, potentially facilitating chimera formation through base pair mismatching in the neighboring regions of high similarity in the chimeric region. We conclude that the hypervariable nature of the ITS region seem to buffer the rate of chimera formation in comparison to other, less variable barcoding regions, due to shorter regions of high sequence similarity.
FIGURES 14–22 in Revision of Australian jumping spider genus Servaea Simon 1887 (Aranaea: Salticidae) including use of DNA sequence data and predicted distributions
FIGURES 14–22. Servaea incana (cont.) 14 known and predicted distribution; 15–20 male palp (15–17 'light' male and 19–20 'dark' male; 21 anterior view of geniculate male chelicera; 22 anterior view of rounded female chelicera. Scale: 0.2 mm
FIGURES 39–46. Servaea spinibarbis. 39–40 in Revision of Australian jumping spider genus Servaea Simon 1887 (Aranaea: Salticidae) including use of DNA sequence data and predicted distributions
FIGURES 39–46. Servaea spinibarbis. 39–40 dorsal view (39 female, 40 male); 41–42 female genitalia (41 dorsal view of cleared specimen, 42 ventral view of external characteristics); 43–45 male palp (43 ventral view, 44 anterior lateral view, 45 posterior lateral view); 46 known and predicted distribution. Scale: total body 1 mm; remainder 0.2 mm.
FIGURES 23–30. Servaea melaina n in Revision of Australian jumping spider genus Servaea Simon 1887 (Aranaea: Salticidae) including use of DNA sequence data and predicted distributions
FIGURES 23–30. Servaea melaina n. sp. 23–24 dorsal view (23 female, 24 male); 25–26 female genitalia (25 dorsal view of cleared specimen, 26 ventral view of external characteristics); 27–29 male palp (27 ventral view, 28 anterior lateral view, 29 posterior lateral view); 30 known and predicted distribution. Scale: total body 1 mm; remainder 0.2 mm
FIGURES 47–54. Servaea villosa. 47–48 in Revision of Australian jumping spider genus Servaea Simon 1887 (Aranaea: Salticidae) including use of DNA sequence data and predicted distributions
FIGURES 47–54. Servaea villosa. 47–48 dorsal view (47 female, 48 male); 49–50 female genitalia (49 dorsal view of cleared specimen, 50 ventral view of external characteristics); 51–53 male palp (51 ventral view, 52 anterior lateral view, 53 posterior lateral view); 54 known and predicted distribution. Scale: total body 1 mm; remainder 0.2 mm.
FIGURE 5 in Revision of Australian jumping spider genus Servaea Simon 1887 (Aranaea: Salticidae) including use of DNA sequence data and predicted distributions
FIGURE 5. Phenetic tree resulting from Neighbour Joining analysis inferred from the COI data-set. Branch lengths are proportional to genetic differences (see scale bar).
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.