Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,696

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,696 results for “DNA sequence”

Learn how ShareScore rates datasets ↗
zenodo28/100

Figure 1 in Comment on Psonis et al. (2015): 'Evaluation of the taxonomy of Helix cincta (Muller, 1774) and Helix nucula (Mousson, 1854); insights using mitochondrial DNA sequence data'

Figure 1. (A) Helix nucula Mousson, 1854, syntype ZMZ 506381a, shell diameter 28.8 mm. (B) Helix melanostoma Draparnaud, 1801, syntype of Helix uthicensis Pechaud, 1883, ruins of Uthica, Tunisia, MHNG 18185, shell diameter 36.2 mm. (C), Helix cincta Müller, 1774, syntype of Helix (Pomatia) cincta var. anatolica Kobelt, 1891, 'Hieronda' (southwest Turkey), SMF 9949, shell diameter 34.9 mm. (D) Helix pronuba Westerlund & Blanc, 1879, syntype of Helix thiesseana var. pronuba Westerlund & Blanc, 1879, GNM 1723, shell diameter 26.3 mm. (E) Helix borealis Mousson, 1859, possible syntype ZMZ 506307, Grecce, Argostoli, shell diameter 34.65 mm.

opencc-by-4.0Mar 2015View details →
zenodo28/100

Aligned DNA sequence matrix for phylogenetic analyses in the article "Three new species of frogs of the genus Pristimantis (Anura: Strabomantidae) with a redefinition of the P. lacrimosus species group"

<p>Aligned DNA sequence matrix for phylogenetic analyses of the article &quot;Three new species of frogs of the genus Pristimantis (Anura: Strabomantidae) with a redefinition of the P. lacrimosus species group&quot;. The matrix is in NEXUS format.</p> <p>Gene partitions are arranged as follows (tRNAs are included as part of larger adjacent genes):</p> <p>RAG1 codon position 1 = &nbsp;1-625\3;<br> RAG1 codon position 2&nbsp;= &nbsp;2-626\3;<br> RAG1 codon position 3&nbsp;= &nbsp;3-627\3;<br> 12S rRNA = 628-1440;<br> 16S rRNA = 1441-3071;<br> ND1 codon position 1 = &nbsp;3072-3972\3;<br> ND1 codon position 2&nbsp;= &nbsp;3073-3973\3;<br> ND1 codon position 3&nbsp;= &nbsp;3074-3974\3;</p>

opencc-by-4.0Dec 2019View details →
dryad28/100

Raw data associated with the article: "Single-molecule DNA sequencing of widely varying GC-content using nucleotide release, capture and detection in microdroplets.", NAR, Puchtler et.al.

<p>All data taken in the production of the corresponding paper: "Single-molecule DNA sequencing of widely varying GC-content using nucleotide release, capture and detection in microdroplets."</p> <p>The associated manuscript describes a method for DNA sequencing which involves the sequential release of nucleotides from a single, immobilised strand of DNA via pyrophosphorolysis (PPL). Released nucleotides, in the form of dNTPs, are captured in microdroplets which are manipulated using an optical-EWOD platform. A detection chemistry within each droplet releases a specific dye depending on which dNTPs are present, allowing the optical read-out of bases within each droplet. Hence, by capturing bases sequentially within droplets as they are cleaved from the strand of DNA, the sequence can be optically identified.</p>

opencc-zeroOct 2020View details →
zenodo28/100

FIGURE 13 in Three new species and DNA sequence data of the rare South American water beetle genus Adelphydraena Perkins, 1989 (Coleoptera: Hydraenidae)

FIGURE 13. Ventral habitus of Adelphydraena species.

opennotspecifiedSep 2020View details →
zenodo28/100

FIGURE 7. Adelphydraena spinosa n in Three new species and DNA sequence data of the rare South American water beetle genus Adelphydraena Perkins, 1989 (Coleoptera: Hydraenidae)

FIGURE 7. Adelphydraena spinosa n. sp., habitus of holotype.

opennotspecifiedSep 2020View details →
zenodo28/100

FIGURE 1 in Three new species and DNA sequence data of the rare South American water beetle genus Adelphydraena Perkins, 1989 (Coleoptera: Hydraenidae)

FIGURE 1. Dorsal habitus of Adelphydraena species.

opennotspecifiedSep 2020View details →
zenodo28/100

FIGURE 4 in Three new species and DNA sequence data of the rare South American water beetle genus Adelphydraena Perkins, 1989 (Coleoptera: Hydraenidae)

FIGURE 4. Adelphydraena spangleri Perkins, habitus of non-type.

opennotspecifiedSep 2020View details →
zenodo28/100

FIGURE 3 in Three new species and DNA sequence data of the rare South American water beetle genus Adelphydraena Perkins, 1989 (Coleoptera: Hydraenidae)

FIGURE 3. Adelphydraena orchymonti Perkins, habitus of non-type.

opennotspecifiedSep 2020View details →
zenodo28/100

FIGURE 8. Adelphydraena surinamensis n in Three new species and DNA sequence data of the rare South American water beetle genus Adelphydraena Perkins, 1989 (Coleoptera: Hydraenidae)

FIGURE 8. Adelphydraena surinamensis n. sp., habitus of holotype.

opennotspecifiedSep 2020View details →
zenodo28/100

FIGURE 2. Adelphydraena amazonica n in Three new species and DNA sequence data of the rare South American water beetle genus Adelphydraena Perkins, 1989 (Coleoptera: Hydraenidae)

FIGURE 2. Adelphydraena amazonica n. sp., habitus of holotype.

opennotspecifiedSep 2020View details →
dryad28/100

Data from: Robust DNA isolation and high-throughput sequencing library construction for herbarium specimens

Herbaria are an invaluable source of plant material that can be used in a variety of biological studies. The use of herbarium specimens is associated with a number of challenges including sample preservation quality, degraded DNA, and destructive sampling of rare specimens. In order to more effectively use herbarium material in large sequencing projects, a dependable and scalable method of DNA isolation and library preparation is needed. This paper demonstrates a robust, beginning-to-end protocol for DNA isolation and high-throughput library construction from herbarium specimens that does not require modification for individual samples. This protocol is tailored for low quality dried plant material and takes advantage of existing methods by optimizing tissue grinding, modifying library size selection, and introducing an optional reamplification step for low yield libraries. Reamplification of low yield DNA libraries can rescue samples derived from irreplaceable and potentially valuable herbarium specimens, negating the need for additional destructive sampling and without introducing discernible sequencing bias for common phylogenetic applications. The protocol has been tested on hundreds of grass species, but is expected to be adaptable for use in other plant lineages after verification. This protocol can be limited by extremely degraded DNA, where fragments do not exist in the desired size range, and by secondary metabolites present in some plant material that inhibit clean DNA isolation. Overall, this protocol introduces a fast and comprehensive method that allows for DNA isolation and library preparation of 24 samples in less than 13 hours, with only 8 hours of active hands-on time with minimal modifications.

opencc-zeroDec 2017View details →
dryad28/100

Data from: Quantifying sequence proportions in a DNA-based diet study using Ion Torrent amplicon sequencing: which counts count?

A goal of many environmental DNA barcoding studies is to infer quantitative information about relative abundances of different taxa based on sequence read proportions generated by high-throughput sequencing. However, potential biases associated with this approach are only beginning to be examined. We sequenced DNA amplified from faeces (scats) of captive harbour seals (Phoca vitulina) to investigate whether sequence counts could be used to quantify the seals' diet. Seals were fed fish in fixed proportions, a chordate-specific mitochondrial 16S marker was amplified from scat DNA and amplicons sequenced using an Ion Torrent PGM™. For a given set of bioinformatic parameters, there was generally low variability between scat samples in proportions of prey species sequences recovered. However, proportions varied substantially depending on sequencing direction, level of quality filtering (due to differences in sequence quality between species) and minimum read length considered. Short primer tags used to identify individual samples also influenced species proportions. In addition, there were complex interactions between factors; for example, the effect of quality filtering was influenced by the primer tag and sequencing direction. Resequencing of a subset of samples revealed some, but not all, biases were consistent between runs. Less stringent data filtering (based on quality scores or read length) generally produced more consistent proportional data, but overall proportions of sequences were very different than dietary mass proportions, indicating additional technical or biological biases are present. Our findings highlight that quantitative interpretations of sequence proportions generated via high-throughput sequencing will require careful experimental design and thoughtful data analysis.

opencc-zeroDec 2012View details →
dryad28/100

Data from: Universal and blocking primer mismatches limit the use of high-throughput DNA sequencing for the quantitative metabarcoding of arthropods

The quantification of the biological diversity in environmental samples using high-throughput DNA sequencing is hindered by the PCR bias caused by variable primer–template mismatches of the individual species. In some dietary studies, there is the added problem that samples are enriched with predator DNA, so often a predator-specific blocking oligonucleotide is used to alleviate the problem. However, specific blocking oligonucleotides could coblock nontarget species to some degree. Here, we accurately estimate the extent of the PCR biases induced by universal and blocking primers on a mock community prepared with DNA of twelve species of terrestrial arthropods. We also compare universal and blocking primer biases with those induced by variable annealing temperature and number of PCR cycles. The results show that reads of all species were recovered after PCR enrichment at our control conditions (no blocking oligonucleotide, 45 °C annealing temperature and 40 cycles) and high-throughput sequencing. They also show that the four factors considered biased the final proportions of the species to some degree. Among these factors, the number of primer–template mismatches of each species had a disproportionate effect (up to five orders of magnitude) on the amplification efficiency. In particular, the number of primer–template mismatches explained most of the variation (~3/4) in the amplification efficiency of the species. The effect of blocking oligonucleotide concentration on nontarget species relative abundance was also significant, but less important (below one order of magnitude). Considering the results reported here, the quantitative potential of the technique is limited, and only qualitative results (the species list) are reliable, at least when targeting the barcoding COI region.

opencc-zeroDec 2013View details →
dryad28/100

Data from: DNA barcoding meets molecular scatology: short mtDNA sequences for standardized species assignment of carnivore noninvasive samples

Although species assignment of scats is important to study carnivoran biology, there is still no standardized assay for the identification of carnivores worldwide, which would allow large-scale routine assessments and reliable cross-comparison of results. Here we evaluate the potential of two short mtDNA fragments (ATP6 [126 bp] and COI [187 bp]) to serve as standard markers for the Carnivora. Samples of 66 species were sequenced for one or both of these segments. Alignments were complemented with archival sequences, and analyzed with three approaches (tree-based, distance-based and character-based). Intraspecific genetic distances were generally lower than between-species distances, resulting in diagnosable clusters for 86% (ATP6) and and 85% (COI) of the species. Notable exceptions were recently diverged species, most of which could still be identified using diagnostic characters, uniqueness of haplotypes, or by reducing the geographic scope of the comparison. In silico comparative analyses were also performed with a 110-bp cytochrome b (cytb) segment, whose identification success was lower (70%), possibly due to the smaller number of informative sites and/or the influence of misidentified sequences obtained from GenBank. Finally, we performed case-studies with faecal samples, which supported the suitability of our two focal markers for poor-quality DNA, and allowed an assessment of prey-DNA co-amplification. No evidence of prey DNA contamination was found for ATP6, while some cases were observed for COI and subsequently eliminated by the design of more specific primers. Overall, our results indicate that these segments hold good potential as standard markers for accurate species-level identification in the Carnivora.

opencc-zeroDec 2010View details →
dryad28/100

Data from: DNA barcodes from century-old type specimens using next generation sequencing

Type specimens have high scientific importance because they provide the only certain connection between the application of a Linnean name and a physical specimen. Many other individuals may have been identified as a particular species, but their linkage to the taxon concept is inferential. Because type specimens are often more than a century old and have experienced conditions unfavorable for DNA preservation, success in sequence recovery has been uncertain. The present study addresses this challenge by employing next generation sequencing (NGS) to recover sequences for the barcode region of the cytochrome c oxidase 1 gene from small amounts of template DNA. DNA quality was first screened in more than 1800 century-old type specimens of Lepidoptera by attempting to recover 164bp and 94bp reads via Sanger sequencing. This analysis permitted the assignment of each specimen to one of three DNA quality categories – high (164bp sequence), medium (94bp sequence), or low (no sequence). Ten specimens from each category were subsequently analyzed via a PCR-based NGS protocol requiring very little template DNA. It recovered sequence information from all specimens with average read lengths ranging from 458bp to 610bp for the three DNA categories. By sequencing ten specimens in each NGS run, costs were similar to Sanger analysis. Future increases in the number of specimens processed in each run promise substantial reductions in cost, making it possible to anticipate a future where barcode sequences are available from most type specimens.

opencc-zeroDec 2014View details →
dryad28/100

Data from: Comparison of target-capture and restriction-site associated DNA sequencing for phylogenomics: a test in cardinalid tanagers (Aves, genus: Piranga)

Restriction-site associated DNA sequencing (RAD-seq) and target capture of specific genomic regions, such as ultraconserved elements (UCEs), are emerging as two of the most popular methods for phylogenomics using reduced-representation genomic datasets. These two methods were designed to target different evolutionary timescales: RAD-seq was designed for population-genomic level questions and UCEs for deeper phylogenetics. The utility of both datasets to infer phylogenies across a variety of taxonomic levels has not been adequately compared within the same taxonomic system. Additionally, the effects of uninformative gene trees on species tree analyses (for target capture data) have not been explored. Here, we utilize RAD-seq and UCE data to infer a phylogeny of the bird genus Piranga. The group has a range of divergence dates (0.5 my – 6 my), contains eleven recognized species, and lacks a resolved phylogeny. We compared two species tree methods for the RAD-seq data and six species tree methods for the UCE data. Additionally, in the UCE data, we analyzed a complete matrix as well as datasets with only highly informative loci. A complete matrix of 189 UCE loci with ten or more parsimony informative (PI) sites, and an ~80% complete matrix of 1128 PI SNPs (from RAD-seq) yield the same fully resolved phylogeny of Piranga. We inferred non-monophyletic relationships of P. lutea individuals, with all other a priori species identified as monophyletic. Finally, we found that species tree analyses that included predominantly uninformative gene trees provided strong support for different topologies, with consistent phylogenetic results when limiting species tree analyses to highly informative loci or only using less informative loci with concatenation or methods meant for SNPs alone.

opencc-zeroDec 2015View details →
dryad28/100

Data from: Phylogenetic relationships of Agaric fungi based on nuclear large subunit ribosomal DNA sequences

Phylogenetic relationships of mushrooms and their relatives within the order Agaricales were addressed using nuclear large subunit ribosomal DNA sequences. Approximately 900 bases of the 5' end of the nucleus-encoded large subunit RNA gene (nLSU-rDNA) were sequenced for 154 selected taxa representing most families within the Agaricales. Several phylogenetic methods were used, including weighted and equally weighted parsimony (MP), maximum likelihood (ML), and distance methods (NJ). The starting tree for branch swapping in the ML analyses was the tree with the highest ML score among previously produced MP and NJ trees. A high degree of consensus was observed between phylogenetic estimates obtained through MP and ML. NJ trees differed according to the distance model that was used, however, all NJ trees still supported most of the same terminal groupings as MP and ML trees. NJ trees were always significantly suboptimal when evaluated against the best MP and ML trees, using both parsimony and likelihood tests. Our analyses suggest that weighted parsimony and ML provide the best estimates of Agaricales phylogeny. Similar support was observed between bootstrapping and jackknifing methods for evaluation of tree robustness. Phylogenetic analyses revealed many groups of agaricoid fungi that are supported by moderate to high bootstrap or jackknife levels or are consistent with morphology-based classification schemes. Analyzes also support separate placement of the boletes and russules, which are basal to the main core group of gilled mushrooms (the Agaricineae of Singer). Examples of monophyletic groups include the families Amanitaceae, Coprinaceae (excluding Coprinus comatus and subfamily Panaeolideae), Agaricaceae (excluding the Cystodermateae), and Strophariaceae pro parte (Stropharia, Pholiota, and Hypholoma); the mycorrhizal species of Tricholoma (including Leucopaxillus, also mycorrhizal); Mycena and Resinomycena; Termitomyces, Podabrella, and Lyophyllum; and Pleurotus with Hohenbuehelia. Several nonmonophyletic groups revealed by these data include the families Tricholomataceae, Cortinariaceae, and Hygrophoraceae and the genera Clitocybe, Omphalina, and Marasmius. This study provides a framework for future systematics studies in the Agaricales and suggestions for analyzing large molecular data sets.

opencc-zeroDec 2008View details →
dryad28/100

Data from: Restriction site-associated DNA sequencing (RAD-seq) reveals an extraordinary number of transitions among gecko sex-determining systems

Sex chromosomes have evolved many times in animals and studying these replicate evolutionary "experiments" can help broaden our understanding of the general forces driving the origin and evolution of sex chromosomes. However this plan of study has been hindered by the inability to identify the sex chromosome systems in the large number of species with cryptic, homomorphic sex chromosomes. Restriction site-associated DNA sequencing (RAD-seq) is a critical enabling technology that can identify the sex chromosome systems in many species where traditional cytogenetic methods have failed. Using newly generated RAD-seq data from twelve gecko species, along with data from the literature, we reinterpret the evolution of sex-determining systems in lizards and snakes and test the hypothesis that sex chromosomes can routinely act as evolutionary traps. We uncovered between 17 and 25 transitions among gecko sex-determining systems. This is approximately ½ to ⅔ of the total number of transitions observed among all lizards and snakes. We find support for the hypothesis that sex chromosome systems can readily become trap-like and show that adding even a small number of species from understudied clades can greatly enhance hypothesis testing in a model-based phylogenetic framework. RAD-seq will undoubtedly prove useful in evaluating other species for male or female heterogamety, particularly the majority of fish, amphibian, and reptile species that lack visibly heteromorphic sex chromosomes, and will significantly accelerate the pace of biological discovery.

opencc-zeroDec 2014View details →
dryad28/100

Data from: Counting with DNA in metabarcoding studies: how should we convert sequence reads to dietary data?

Advances in DNA sequencing technology have revolutionised the field of molecular analysis of trophic interactions and it is now possible to recover counts of food DNA sequences from a wide range of dietary samples. But what do these counts mean? To obtain an accurate estimate of a consumer's diet should we work strictly with datasets summarising frequency of occurrence of different food taxa, or is it possible to use relative number of sequences? Both approaches are applied to obtain semi-quantitative diet summaries, but occurrence data is often promoted as a more conservative and reliable option due to taxa-specific biases in recovery of sequences. We explore representative dietary metabarcoding datasets and point out that diet summaries based on occurrence data often overestimate the importance of food consumed in small quantities (potentially including low-level contaminants) and are sensitive to the count threshold used to define an occurrence. Our simulations indicate that using relative read abundance (RRA) information often provide a more accurate view of population-level diet even with moderate recovery biases incorporated; however, RRA summaries are sensitive to recovery biases impacting common diet taxa. Both approaches are more accurate when the mean number of food taxa in samples is small. The ideas presented here highlight the need to consider all sources of bias and to justify the methods used to interpret count data in dietary metabarcoding studies. We encourage researchers to continue addressing methodological challenges, and acknowledge unanswered questions to help spur future investigations in this rapidly developing area of research.

opencc-zeroDec 2017View details →
dryad28/100

Data from: Genomic-scale capture and sequencing of endogenous DNA from feces

Genomic-level analyses of DNA from non-invasive sources would facilitate powerful conservation and evolutionary studies in natural populations of endangered and otherwise elusive species. However, the typical low quantity and poor quality of DNA that is extracted from non-invasive samples have generally precluded such work. Here we apply a modified DNA capture protocol that, when used in combination with massively-parallel sequencing technology, facilitates efficient and highly-accurate resequencing of megabases of specified nuclear genomic regions from fecal DNA samples. We validated our approach by comparing genetic variants identified from corresponding fecal and blood DNA samples of six western chimpanzees (Pan troglodytes verus) across more than 1.5 megabases of chromosome 21, chromosome X, and the complete mitochondrial genome. Our results suggest that it is now feasible to conduct genomic studies in natural populations for which constraints on invasive sampling have otherwise long been a barrier. The data we collected also provided an opportunity to examine western chimpanzee genetic diversity at unprecedented scale. Despite high mitochondrial genome diversity (pi = 0.585%), western chimpanzees have a low ratio (0.42) of X chromosomal (pi = 0.034%) to autosomal (chromosome 21 pi = 0.081%) sequence diversity, a pattern that may reflect an unusual demographic history of this subspecies.

opencc-zeroDec 2009View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record