Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,696
datasets available to search
ShareScore release 0.9.0
Dataset results
1,696 results for “DNA sequence”
Data from: Estimation of a killer whale (Orcinus orca) population's diet using sequencing analysis of DNA from feces
Estimating diet composition is important for understanding interactions between predators and prey and thus illuminating ecosystem function. The diet of many species, however, is difficult to observe directly. Genetic analysis of fecal material collected in the field is therefore a useful tool for gaining insight into wild animal diets. In this study, we used high-throughput DNA sequencing to quantitatively estimate the diet composition of an endangered population of wild killer whales (Orcinus orca) in their summer range in the Salish Sea. We combined 175 fecal samples collected between May and September from five years between 2006 and 2011 into 13 sample groups. Two known DNA composition control groups were also created. Each group was sequenced at a ~330bp segment of the 16s gene in the mitochondrial genome using an Illumina MiSeq sequencing system. After several quality controls steps, 4,987,107 individual sequences were aligned to a custom sequence database containing 19 potential fish prey species and the most likely species of each fecal-derived sequence was determined. Based on these alignments, salmonids made up >98.6% of the total sequences and thus of the inferred diet. Of the six salmonid species, Chinook salmon made up 79.5% of the sequences, followed by coho salmon (15%). Over all years, a clear pattern emerged with Chinook salmon dominating the estimated diet early in the summer, and coho salmon contributing an average of >40% of the diet in late summer. Sockeye salmon appeared to be occasionally important, at >18% in some sample groups. Non-salmonids were rarely observed. Our results are consistent with earlier results based on surface prey remains, and confirm the importance of Chinook salmon in this population's summer diet.
Data from: Congruent species delimitation of two controversial gold-thread nanmu tree species based on morphological and restriction site-associated DNA sequencing data
Species delimitation is fundamental to conservation and sustainable use of economically important forest tree species. However, the delimitation of two highly valued gold-thread nanmu species (Phoebe bournei and P. zhennan) has been confusing and debated. To address this problem, we integrated morphology and restriction site-associated DNA sequencing (RADseq) to define their species boundaries. We obtained highly consistent results from both data sets, supporting two distinct lineages corresponding to P. bournei and P. zhennan. In Phoebe bournei, higher order leaf venation is more prominent, petioles are thicker and leaf apex angle is narrower, compared to P. zhennan. Both data sets also showed that putative P. bournei localities from north-eastern Guizhou were P. zhennan. The two species have different distributions and only overlap in the Wuling Mountains. Phoebe bournei occurs mainly in Central Fujian, southern Jiangxi, the Nanling Mountains and the Wuling Mountains, whereas P. zhennan is found in the adjoining eastern regions of the Qionglai Mountains, the Southern Sichuan Hills and the Wuling Mountains. The improved delimitation of P. bournei and P. zhennan and clarification of their ranges provide a better guidance for conservation and sustainable utilization of these tree species.
Data from: Developing nuclear DNA phylogenetic markers in the angiosperm genus Leucadendron (Proteaceae): a next-generation sequencing transcriptomic approach
Despite the recent advances in generating molecular data, reconstructing species-level phylogenies for non-models groups remains a challenge. The use of a number of independent genes is required to resolve phylogenetic relationships, especially for groups displaying low polymorphism. In such cases, low-copy nuclear exons and non-coding regions, such as 3′ untranslated regions (3′-UTRs) or introns, constitute a potentially interesting source of nuclear DNA variation. Here, we present a methodology meant to identify new nuclear orthologous markers using both public-nucleotide databases and transcriptomic data generated for the group of interest by using next generation sequencing technology. To identify PCR primers for a non-model group, the genus Leucadendron (Proteaceae), we adopted a framework aimed at minimizing the probability of paralogy and maximizing polymorphism. We anchored when possible the right-hand primer into the 3′-UTR and the left-hand primer into the coding region. Seven new nuclear markers emerged from this search strategy, three of those included 3′-UTRs. We further compared the phylogenetic potential between our new markers and the ribosomal internal transcribed spacer region (ITS). The sequenced 3′-UTRs yielded higher polymorphism rates than the ITS region did. We did not find strong incongruences with the phylogenetic signal contained in the ITS region and the seven new designed markers but they strongly improved the phylogeny of the genus Leucadendron. Overall, this methodology is efficient in isolating orthologous loci and is valid for any non-model group given the availability of transcriptomic data.
Data from: Restriction-site-associated DNA sequencing reveals a cryptic viburnum species on the North American coastal plain
Species are the starting point for most studies of ecology and evolution, but the proper circumscription of species can be extremely difficult in morphologically variable lineages, and there are still few convincing examples of molecularly-informed species delimitation in plants. We focus here on the Viburnum nudum complex, a highly variable clade that is widely distributed in eastern North America. Taxonomic treatments have mostly divided this complex into northern (V. nudum var. cassinoides) and southern (V. nudum var. nudum) entities, but additional names have been proposed. We used multiple lines of evidence, including RADseq, morphological, and geographic data, to test how many independently evolving lineages exist within the V. nudum complex. Genetic clustering and phylogenetic methods revealed three distinct groups—one lineage that is highly divergent, and two others that are recently diverged and morphologically similar. A combination of evidence that includes reciprocal monophyly, lack of introgression, and discrete rather than continuous patterns of variation supports the recognition of all three lineages as separate species. These results identify a surprising case of cryptic diversity in which two broadly sympatric species have consistently been lumped in taxonomic treatments. The clarity of our findings is directly related to the dense sampling and high quality genetic data in this study. We argue that there is a critical need for carefully sampled and integrative species delimitation studies to clarify species boundaries even in well-known plant lineages. Studies following the model that we have developed here are likely to identify many more cryptic lineages and will fundamentally improve our understanding of plant speciation and patterns of species richness.
Data from: Massively parallel multiplex DNA sequencing for specimen identification using an Illumina MiSeq platform
Genetic information is a valuable component of biosystematics, especially specimen identification through the use of species-specific DNA barcodes. Although many genomics applications have shifted to High-Throughput Sequencing (HTS) or Next-Generation Sequencing (NGS) technologies, sample identification (e.g., via DNA barcoding) is still most often done with Sanger sequencing. Here, we present a scalable double dual-indexing approach using an Illumina Miseq platform to sequence DNA barcode markers. We achieved 97.3% success by using half of an Illumina Miseq flowcell to obtain 658 base pairs of the cytochrome c oxidase I DNA barcode in 1,010 specimens from eleven orders of arthropods. Our approach recovers a greater proportion of DNA barcode sequences from individuals than does conventional Sanger sequencing, while at the same time reducing both per specimen costs and labor time by nearly 80%. In addition, the use of HTS allows the recovery of multiple sequences per specimen, for deeper analysis of genetic variation in target gene regions.
Data from: ITS all right mama: Investigating the formation of chimeric sequences in the ITS2 region by DNA metabarcoding analyses of fungal mock communities of different complexities
The formation of chimeric sequences can create significant methodological bias in PCR-based DNA metabarcoding analyses. During mixed-template amplification of barcoding regions, chimera formation is frequent and well documented. However, profiling of fungal communities typically uses the more variable rDNA region ITS. Due to a larger research community, tools for chimera detection have been developed mainly for the 16S/18S markers. However, these tools are widely applied to the ITS region without verification of their performance. We examined the rate of chimera formation during amplification and 454 sequencing of the ITS2 region from fungal mock communities of different complexities. We evaluated the chimera detecting ability of two common chimera-checking algorithms: Perseus and UCHIME. Large proportions of the chimeras reported were false positives. No false negatives were found in the dataset. Verified chimeras accounted for only 0.2% of the total ITS2 reads, which is considerably less than what is typically reported in 16S and 18S metabarcoding analyses. Verified chimeric "parent sequences" had significantly higher percent identity to one another than to random members of the mock communities. Community complexity increased the rate of chimera formation. GC content was higher around the verified chimeric break points, potentially facilitating chimera formation through base pair mismatching in the neighboring regions of high similarity in the chimeric region. We conclude that the hypervariable nature of the ITS region seem to buffer the rate of chimera formation in comparison to other, less variable barcoding regions, due to shorter regions of high sequence similarity.
Figure 10 in Target-enriched DNA sequencing from historical type material enables a partial revision of the Madagascar giant stream frogs (genus Mantidactylus)
Figure 10. Mantidactylus (Mantidactylus) radaka sp. nov. being prepared for human consumption. (a) Frogs and crabs are collected from broad streams. Then (b) the frogs are gutted and skinned, and the head, hands and feet removed. The frog is then rinsed in the stream, leaving (c) cleaned animals for cooking in a stew. Note the ovaries full with hundreds of eggs.
Figure 9 in Target-enriched DNA sequencing from historical type material enables a partial revision of the Madagascar giant stream frogs (genus Mantidactylus)
Figure 9. Preserved type specimens of the four nomina in the Mantidactylus subgenus Mantidactylus and one of the paralectotypes of Rana guttulata.
Figure 7 in Target-enriched DNA sequencing from historical type material enables a partial revision of the Madagascar giant stream frogs (genus Mantidactylus)
Figure 7. Photographs of living specimens of Mantidactylus (Mantidactylus) guttulatus, M. (M.) grandidieri, and of three candidate species. (a, b) M. (M.) guttulatus, female ZSM 1013/2003 (FGMV 2002.438) from Ranomafana. (c) Unidentified specimen from Ranomafana, assigned tentatively to M. (M.) guttulatus (no genetic evidence). (d, e) M. (M.) guttulatus, specimen KU 340853 (CRH729) from Ranomafana. (f) M. (M.) grandidieri, specimen ZSM 5077/2005 (ZCMV 2159) from Nosy Mangabe. (g) M. (M.) grandidieri, specimen ZSM 276/2005 (FGZC 2682) from Vohidrazana. (h) M. (M.) grandidieri, unidentified specimen (probably subadult) from Andranofotsy. (i, j) M. (M.) grandidieri, specimen KU
Figure 5. Per-base coverage plots for the 16S in Target-enriched DNA sequencing from historical type material enables a partial revision of the Madagascar giant stream frogs (genus Mantidactylus)
Figure 5. Per-base coverage plots for the 16S fragment in four Mantidactylus type specimens from the MNHN and BMNH collections. (a) BMNH 1947.2.25.48 (paralectotype of Rana guttulata); (b) BMNH 1947.2.25.51 (paralectotype of Rana guttulata); (c) MNHN 1895.255 (syntype of M. grandidieri); (d) MNHN 1883.520 (syntype of M. grandidieri).
Figure 3 in Target-enriched DNA sequencing from historical type material enables a partial revision of the Madagascar giant stream frogs (genus Mantidactylus)
Figure 3. Haplotype network of the subgenus Mantidactylus based on 1227 bp of the nuclear RAG-1 gene from 39 samples. Small black dots represent additional mutational steps.
Figure 2 in Target-enriched DNA sequencing from historical type material enables a partial revision of the Madagascar giant stream frogs (genus Mantidactylus)
Figure 2. Diagonal matrix visualising the mean uncorrected genetic distances (p-distances) in the mitochondrial 16S rRNA gene between the different lineages in the subgenus Mantidactylus, calculated from 514 bp of the 16S mitochondrial gene.
Figure 1. Maximum likelihood phylogenetic tree obtained from 514 in Target-enriched DNA sequencing from historical type material enables a partial revision of the Madagascar giant stream frogs (genus Mantidactylus)
Figure 1. Maximum likelihood phylogenetic tree obtained from 514 bp of the mitochondrial 16S rRNA gene. The values at the nodes are the bootstrap supports (not given for intra-lineage nodes for improved clarity). The type specimens of M. guttulatus and M. grandidieri from the London and Paris museum collections are highlighted in red and brown, respectively.
Figure 4 in Target-enriched DNA sequencing from historical type material enables a partial revision of the Madagascar giant stream frogs (genus Mantidactylus)
Figure 4. Stacked barplots showing the number of reads uniquely matching different reference sequences for the three targeted mitochondrial genes with a similarity threshold of 98%. The Rana pigra type was not included because the number of reads was too low.
Figure 8 in Target-enriched DNA sequencing from historical type material enables a partial revision of the Madagascar giant stream frogs (genus Mantidactylus)
Figure 8. Lateral views of the heads of preserved adult males of Mantidactylus (Mantidactylus) radaka sp. nov. in comparison with M. (M.) guttulatus and M. (M.) grandidieri. Note the more distinct and larger tympanum (indicated by yellow arrows) in the latter two species. Not to scale.
Figure 6 in Target-enriched DNA sequencing from historical type material enables a partial revision of the Madagascar giant stream frogs (genus Mantidactylus)
Figure 6. Photographs of living specimens of Mantidactylus radaka sp. nov. (a, b) Male holotype ZSM 644/2001 (field number FGMV 2001.132) from Manarikoba forest, Tsaratanana Massif. (c–f) Female paratype ZSM 1800/2010 (ZCMV 12345) from Camp 1 (Antevialambazaha), Tsaratanana Massif. (g, h) Female paratype ZSM 97/2016 (MSZC 0080) from Ampotsidy. (i, j) Male paratype MSZC 0120 (uncatalogued in UADBA) from Ampotsidy. (k) Unidentified specimen from Camp 0 (Ankijagna Lagnana), Tsaratanana Massif. (l) Paratype ZSM 582/2014 (DRV 6073) from Camp 0 (Ankijagna Lagnana). (m, n) Unidentified female specimen from Manongarivo (Camp 0), probably preserved in UADBA collection.
Mitochondrial DNA tree for COI sequences (DNA barcode) of the goby genus Trimma.
<p>Mitochondrial DNA tree for COI sequences (DNA barcode) of the goby genus Trimma</p>
FIGURE 4. Leiolopisma ceciliae n in A new Nactus gecko (Gekkonidae) and a new Leiolopisma skink (Scincidae) from La Réunion, Indian Ocean, based on recent fossil remains and ancient DNA sequence
FIGURE 4. Leiolopisma ceciliae n. sp., material from Grotte au Sable, St-Gilles, La Réunion. Holotype. Dentary (lateral and medial views). Scale in mm.
FIGURE 5. Leiolopisma ceciliae n in A new Nactus gecko (Gekkonidae) and a new Leiolopisma skink (Scincidae) from La Réunion, Indian Ocean, based on recent fossil remains and ancient DNA sequence
FIGURE 5. Leiolopisma ceciliae n. sp., material from Grotte au Sable, St-Gilles, La Réunion. Selected paratypes, from individuals of various sizes. Top row: left mandible (lateral view), right maxilla (lateral view), fused left postfrontal and postorbital bone (dorsal view), left quadrate (posterior view). Middle row: right dentary (medial view), posterior left mandible (lateral view). Bottom row: Left humerus, left pelvis (lateral view), sacrum (dorsal view), presacral vertebrae (lateral and ventral views), right femur, tibia. Scale in mm.
FIGURE 6. Leiolopisma ceciliae n in A new Nactus gecko (Gekkonidae) and a new Leiolopisma skink (Scincidae) from La Réunion, Indian Ocean, based on recent fossil remains and ancient DNA sequence
FIGURE 6. Leiolopisma ceciliae n. sp., material from Grotte au Sable, St-Gilles, La Réunion. Paratype. Frontal of juvenile animal (ventral and dorsal views, anterior upwards). Scale in mm.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.