Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

521

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

521 results for “DNA sequence data”

Learn how ShareScore rates datasets ↗
dryad36/100

DNA sequence data generated using non-invasive feather and eggshell samples from the Grenada Dove for two gene regions: Cyt b and ND2

<p>As an island endemic with a decreasing population, the Critically Endangered Grenada Dove <em>Leptotila wellsi</em> is threatened by accelerated loss of genetic diversity resulting from ongoing habitat fragmentation. Small, threatened populations are difficult to sample directly but advances in molecular methods mean that non-invasive samples can be used. We performed the first assessment of genetic diversity of populations of Grenada Dove by a) assessing mtDNA genetic diversity in the only two areas of occupancy on Grenada, b) defining the number of haplotypes present at each site and c) evaluating evidence of isolation between sites. We used non-invasively collected samples from two locations: Mt Hartman (n=18) and Perseverance (n=12). DNA extraction and PCR were used to amplify 1,751 bps of mtDNA from two mitochondrial markers: NADH dehydrogenase 2 (<em>ND2</em>) and Cytochrome b (<em>Cyt b</em>). Haplotype diversity (<em>h</em>) of 0.4, a nucleotide diversity (π) of 0.00023 and two unique haplotypes were identified within the <em>ND2</em> sequences; a single haplotype was identified within the <em>Cyt b </em>sequences. Of the two haplotypes identified; the most common haplotype (haplotype A = 73.9%) was observed at both sites and the other (haplotype B = 26.1%) was unique to Perseverance. Our results show low mitochondrial genetic diversity and clear evidence for genetically isolated populations. The Grenada Dove needs urgent conservation action, including habitat protection and potential augmentation of gene flow by translocation in order to increase genetic resilience and diversity with the ultimate aim of securing the long-term survival of this Critically Endangered species. </p>

opencc-zeroNov 2023View details →
dryad36/100

Data from: Benchmarking ultra-high molecular weight DNA preservation methods for long-read and long-range sequencing

<p>Studies in vertebrate genomics require sampling from a broad range of tissue types, taxa, and localities. Recent advancements in long-read and long-range genome sequencing have made it possible to produce high-quality chromosome-level genome assemblies for almost any organism. However, adequate tissue preservation for the requisite ultra-high molecular weight DNA (uHMW DNA) remains a major challenge. Here we present a comparative study of preservation methods for field and laboratory tissue sampling, across vertebrate classes and different tissue types. We find that no single method is best for all cases. Instead, the optimal storage and extraction methods vary by taxa, by tissue, and by down-stream application. Therefore, we provide sample preservation guidelines that ensure sufficient DNA integrity and amount required for use with long-read and long-range sequencing technologies across vertebrates. Our best practices generate the uHMW DNA needed for the high-quality reference genomes for Phase 1 of the Vertebrate Genomes Project (VGP), whose ultimate mission is to generate chromosome-level reference genome assemblies of all ~70,000 extant vertebrate species.</p>

opencc-zeroApr 2022View details →
dryad36/100

DNA metabarcoding sequence data for diet analysis of caribou

<p>Woodland caribou (<em>Rangifer tarandus caribou</em>) are threatened in Canada due to the drastic decline in population size caused primarily by human-induced landscape changes that decrease habitat and increase predation risk. Conservation efforts have largely focused on reducing predators and protecting critical habitat, whereas research on dietary niches and the role of potential food constraints in lichen-poor environments is limited. To improve our understanding of dietary niche variability, we used a next-generation sequencing approach with metabarcoding of DNA extracted from faecal pellets of woodland caribou located on Lake Superior in lichen-rich (mainland) and lichen-poor (island) environments. Amplicon sequencing of fungal ITS2 region revealed lichen-associated fungi as predominant in samples from both populations, but amplification at the chloroplast <em>trnL </em>region, which was only successful on island samples, revealed primary consumption of yew based on relative read abundance (<em>Taxus spp.</em>; 83.68%) with dogwood (<em>Cornus spp</em>.; 9.67%) and maple (<em>Acer spp.</em>; 4.10%) also prevalent. These results suggest that conservation efforts for caribou need to consider the availability of food resources beyond lichen to ensure successful outcomes.  More broadly, we provide a reliable methodology for assessing ungulate diet from archived faecal pellets that could reveal important dietary shifts over time in response to climate change.</p>

opencc-zeroJun 2022View details →
zenodo36/100

Data and scripts for the manuscript of svaRetro and svaNUMT: modular packages for annotating retrotransposed transcripts and nuclear integration of mitochondrial DNA in genome sequencing data

<p>This upload include data and scripts supporting&nbsp;the results described in the manuscript of&nbsp;<em>svaRetro and svaNUMT: modular packages for annotating retrotransposed transcripts and nuclear integration of mitochondrial DNA in genome sequencing data</em><em>.&nbsp;</em>Detailed description of the contents can be found in README.txt.</p>

opencc-by-4.0Feb 2022View details →
zenodo36/100

Raw DNA sequence data of an individual known as "whitequark" (part 1)

<p>Whole genome sequenced on NovaSeq 6000, paired-end 2x150bp with&nbsp;350bp insert.&nbsp;30-40&times;&nbsp;coverage.</p> <p>This dataset can be used by anyone, with attribution.</p>

opencc-by-nc-4.0Sep 2017View details →
zenodo36/100

Raw DNA sequence data of an indivdual known as "whitequark" (part 2)

<p>Whole genome sequenced on NovaSeq 6000, paired-end 2x150bp with 350bp insert. 30-40× coverage.</p> <p>This dataset can be used by anyone, with attribution.</p>

opencc-by-nc-4.0Sep 2017View details →
zenodo36/100

Experimental data for "An End-to-End Coding Scheme for DNA-Based Data Storage With Nanopore Sequenced Reads"

<p>The experimental dataset used in "An End-to-End Coding Scheme for DNA-Based Data Storage With Nanopore Sequenced Reads."</p> <p>A set of 91,766 150-nt oligos were synthesised with GenScript (oligos.fasta). Each oligo consists of a pseudo-random 110-nt payload flanked by 20-nt primers at each end. The strands are split in three roughly equal groups (two groups of 30,589 and one group of 30,588). Each group has a dedicated primer pair for targeted PCR amplification (the primer pairs used for amplification are provided in primers_synthesis.fasta). The pseudo-random payload was designed to avoid primer-payload collisions.</p> <p>For each file, a sample from the synthesised pool was PCR amplified using the corresponding primer pair and sequenced using Oxford Nanopore Technologies MinION sequencing device following the standard library preparation protocol for amplicon DNA. The raw reads were basecalled using guppy, either in fast- ("acc-false") or high-accuracy ("acc-true") regime. The basecaller generated two groups of reads&mdash;"passQ-true" for the reads that passed the quality-score threshold of 8 and "passQ-false" for those that did not. For each group of reads, a BLAST-based fuzzy search for primer sequences was performed and, based on the resulting alignments, the segments containing the correct primer pairs and located at a distance of 150+-15nt were extracted (separately for forward and reverse-complemented reads). The segments are then assigned to the closest synthesized strand based on Levenshtein distance. The resulting clusters are used to estimate the parameters of the end-to-end DNA storage channel model and to test the proposed error-correction scheme.</p> <p>The archive clustered_read_segments.tar.gz contains 12 sub-archives, for each file (0,1,2), accuracy ("acc-true" or "acc-false"), and Q-score ("passQ-true" or "passQ-false"). Within each sub-archive, there are two folders (one for forward read segments and one for backward read segments), and each folder contains two files: one for the reference synthesised (or "transmitted") sequences that correspond to the file in question ("TX__" &mdash; e.g., "TX__file=0_accBaCa=true_passQ=true_filter=true_forward_.txt") and another file for the sequenced (or "received") segment clusters ("RX__" &mdash; e.g., "RX__file=0_accBaCa=true_passQ=true_filter=true_forward_.txt"). The received clusters in the "RX__" file are ordered in correspondence with the synthesised sequences in the "TX__" file, and a line "===============================" is used as a separator.</p>

opencc-by-4.0May 2024View details →
dryad36/100

Data from: Affordable de novo generation of fish mitogenomes using amplification-free enrichment of mitochondrial DNA and deep sequencing of long fragments

<p>Biomonitoring surveys from environmental DNA make use of metabarcoding tools to describe the community composition. These studies match their sequencing results against public genomic databases to identify the species. However, mitochondrial genomic reference data are yet incomplete, only a few genes may be available, or the suitability of existing sequence data is suboptimal for species-level resolution. Here we present a dedicated and cost-effective workflow with no DNA amplification for generating complete fish mitogenomes for the purpose of strengthening fish mitochondrial databases. Two different long-fragment sequencing approaches using Oxford Nanopore sequencing coupled with mitochondrial DNA enrichment were used. One where the enrichment is achieved by preferential isolation of mitochondria followed by DNA extraction and nuclear DNA depletion ('mitoenrichment').  A second enrichment approach takes advantage of the CRISPR-Cas9 targeted scission on previously dephosphorylated DNA ('targeted mitosequencing'). The sequencing results varied between tissue, species, and integrity of the DNA. The mitoenrichment method yielded 0.17-12.33 % of sequences on target and a mean coverage ranging from 74.9 to 805-fold. The targeted mitosequencing experiment from native genomic DNA yielded 1.83-55 % of sequences on target and a 38 to 2123-fold mean coverage. This produced complete the mitogenome of species with homopolymeric regions, tandem repeats, and gene rearrangements. We demonstrate that deep sequencing of long fragments of native fish DNA is possible and can be achieved with low computational resources in a cost-effective manner, opening the discovery of mitogenomes of non-model or understudied fish taxa to a broad range of laboratories worldwide.</p>

opencc-zeroJun 2024View details →
zenodo36/100

Genus level DNA sequence data for three genes (matK, rbcL, trnH-psbA) for the paper: A comprehensive, genus-level time-calibrated phylogeny of the tree flora of Mediterranean Europe and an assessment of its vulnerability

<p>This data file contains the consensus DNA sequences in fasta format, of 64 tree genera found in Mediterranean Europe, following the checklist of M&eacute;dail et al. (2019).&nbsp;</p> <p>The data are used in a manuscript submitted for publication to Botany Letters and currently under revision. The manuscript is entitled: &quot;<em>A comprehensive, genus-level time-calibrated phylogeny of the tree flora of Mediterranean Europe and an assessment of its vulnerability</em>&quot;. Its authors are: Marwan Cheikh Albassatneh, Marcial Escudero, Loic Ponge<sup>*</sup>, Anne-Christine Monnet, Juan Arroyo, Toni Nikolic, Gianluigi Bacchetta, Francesca Bagnoli, Panayotis Dimopoulos, Agathe Leriche, Fr&eacute;d&eacute;ric M&eacute;dail, Anne Roig, Ilaria Spanu, Giovanni Giuseppe Vendramin, Arndt Hampe, Bruno Fady.</p>

opencc-by-4.0Sep 2019View details →
zenodo36/100

Data for paper: Analysis and control of untemplated DNA polymerase activity for guided synthesis of kilobase-scale DNA sequences

<p>This deposit contains data for the paper entitled: "<strong>Analysis and control of untemplated DNA polymerase activity for guided synthesis of kilobase-scale DNA sequences</strong>"</p> <p>Contents include:</p> <ul> <li><strong>00_sequencing-metadata.xlsx</strong> - Metadata for each sequencing sample.</li> <li><strong>01_raw-fastqs.zip</strong> - Raw basecalled FASTQ data.</li> <li><strong>02_clean-fastas.zip</strong> - Cleaned FASTA files (removal of adapters and barcode sequences).</li> <li><strong>03_read-statistics.zip</strong> - Read statistics for all samples.</li> <li><strong>04_sequence-analysis.zip</strong> - Sequence analysis output for all samples.</li> <li><strong>05_qpcr-data.xlsx</strong> - qPCR data for all the dNTP mixes studied.</li> <li><strong>06_analysis-scripts.zip</strong> - Analysis scripts.</li> </ul>

opencc-by-4.0Aug 2024View details →
dryad36/100

Data from: How "simple" methodological decisions affect interpretation of population structure based on reduced representation library DNA sequencing: a case study using the lake whitefish

Reduced representation (RRL) sequencing approaches (e.g., RADSeq, genotyping by sequencing) require decisions about how much to invest in genome coverage and sequencing depth (library quality), as well as choices of values for adjustable bioinformatics parameters. To empirically explore the importance of these "simple" decisions, we generated two independent sequencing libraries for the same 142 individual lake whitefish (Coregonus clupeaformis) using a nextRAD RRL approach: (1) A small number of loci and low sequencing depth (library A); and (2) more loci and higher sequencing depth (library B). The fish were selected from populations with different levels of expected genetic subdivision. Each library was analyzed using the STACKS pipeline followed by three types of population structure assessment (FST, DAPC and ADMIXTURE) with iterative increases in the stringency of sequencing depth and missing data requirements, as well as more specific a priori population maps. Library B was always able to resolve strong population differentiation in all three types of assessment regardless of the selected parameters. In contrast, library A produced more variable results; increasing the minimum sequencing depth threshold (-m) resulted in a reduced number of retained loci, and therefore lost resolution at high -m values for FST and ADMIXTURE, but not DAPC. FST and DAPC were robust to varying the population map and increasing the stringency of missing data requirements. In contrast, ADMIXTURE was unable to resolve strong population differentiation when increasing these same parameters in library A. Similarly, when examining fine scale population subdivision, library B was robust to changing parameters but library A lost resolution depending on the parameter set. We used library B to examine actual subdivision in our study populations. All three types of analysis found complete subdivision among populations in Lake Huron, ON and Dore Lake, SK, Canada using 10,640 SNP loci. Weak population subdivision was detected in Lake Huron with fish from sites in the north-west, Search Bay, North Point and Hammond Bay, showing slight differentiation. Overall, we show that apparently simple decisions about library quality and bioinformatics parameters can have potentially important impacts on the interpretation of population subdivision. Although costly, the early investment in a high-quality library and more conservative stringency settings on STACKS parameters lead to a final dataset that was more consistent and robust when examining both weak and strong population differentiation.

opencc-zeroMar 2020View details →
zenodo36/100

Fig. 5 in Copelatus sibelaemontis sp. nov. (Coleoptera: Dytiscidae) from the Moluccas with generic assignment based on morphology and DNA sequence data

Fig. 5. Distribution of Copelatus sibelaemontis sp. nov.

opencc-by-4.0Dec 2010View details →
dryad36/100

The topological nature of tag jumping in environmental DNA metabarcoding studies (sequencing raw data)

<p>Metabarcoding of environmental DNA constitutes a state-of-the-art tool for environmental studies. One fundamental principle implicit in most metabarcoding studies is that individual sample amplicons can still be identified after being pooled with others – based on their unique combinations of tags – during the so-called demultiplexing step that follows sequencing. Nevertheless, it has been recognized that tags can sometimes be changed (i.e. tag jumping), which ultimately leads to sample crosstalk. Here, using four DNA metabarcoding datasets derived from the analysis of soils and sediments, we show that tag jumping follows very specific and systematic patterns. Specifically, we find a strong correlation between the number of reads in blank samples and their topological position in the tag matrix (described by vertical and horizontal vectors). This observed spatial pattern of artefactual sequences could be explained by polymerase activity, which leads to the exchange of the 3' tag of single stranded tagged sequences through the formation of heteroduplexes with mixed barcodes. Importantly, tag jumping substantially distorted our datasets – despite our use of methods suggested to minimize this error. We developed a topologic model to estimate the noise based on the counts in our blanks, which suggested that 40-80% of the taxa in our soil and sedimentary samples were likely false positives introduced through tag jumping. We highlight that the amount of false positive detections caused by tag jumping strongly biased our community analyses. </p>

opencc-zeroNov 2022View details →
dryad36/100

The phylogeny and global biogeography of Primulaceae based on high-throughput DNA sequence data

<p>The angiosperm family Primulaceae is morphologically diverse and distributed nearly worldwide. However, phylogenetic uncertainty has limited the ability to identify major morphological and biogeographic transitions. We used target capture sequencing with the Angiosperms353 kit for over 300 species across Ericales, tree-based sequence curation, and multiple phylogenetic approaches to investigate the phylogenetics of the major clades of Primulaceae and their relationship to other Ericales. The study included 150 samples of Primulaceae comprising nearly all recognized genera of the family, with a particular focus on the most diverse subfamily, Myrsinoideae, for which previous phylogenetic knowledge was poor. We used fossil and secondary calibrations to generate dated phylogenetic trees and conducted broad-scale biogeographic analyses as well as ancestral state reconstructions of plant habit.</p>

opencc-zeroDec 2022View details →
dryad36/100

Perianth evolution and implications for generic delimitation in the Eucalypts (Myrtaceae): DNA sequences, morphological data

<p><em>Eucalyptus</em> was traditionally defined by the operculate perianth—hence the generic name (Latin, meaning "well-covered"). But after previous phylogenetic analysis placed <em>Angophora</em>, which has free sepals and petals, as sister to the bloodwood eucalypts, the latter were segregated into a new genus, <em>Corymbia</em>. We made a targeted capture of 101 low-copy nuclear exons from 392 samples representing 329 species-level taxa. The phylogeny was estimated using maximum likelihood (IQtree and RAxML) and the multi-species coalescent (Astral). We tested alternative relationships between four genera within Eucalypteae (<em>Arillastrum</em>, <em>Angophora</em>, <em>Eucalyptus</em>, <em>Corymbia</em>) at each of two nodes critical to generic delimitation using Shimodaira's Approximately Unbiased (AU) test. Monophyly of <em>Arillastrum</em> + (<em>Corymbia</em> + <em>Angophora</em>) relative to <em>Eucalyptus</em> sensu stricto was supported whereas monophyly of <em>Corymbia</em> relative to <em>Angophora</em> was decisively rejected. These results indicate that either <em>Eucalyptus</em> should be expanded to include all four genera or <em>Corymbia</em> should be split into two. All of the alternative relationships among the four currently recognised genera imply homoplasy in perianth evolution, specifically with respect to origins of the bud cap (operculum or calyptra), which has been traditionally used to define <em>Eucalyptus</em>. Inferred evolutionary transitions in perianth traits are generally congruent with divergences between major clades with a single exception: expression of separate sepals and petals in <em>Angophora</em>, which is nested within the operculate genus <em>Corymbia</em>, appears prima facie to be a reversal to the plesiomorphic perianth structure. Strictly, this is not a reversal because the petals of <em>Angophora</em> and <em>Corymbia</em> have a novel compound keel-and-limb structure that is absent in the outgroups. This structure is evident in early development, irrespective of whether the petals remain free or later become part of an operculum. Many of the currently recognised infrageneric taxa down to sectional level (and below in some cases) are well-supported by the sequence data and definable by morphological traits. Inclusion of <em>Angophora</em> within <em>Eucalyptus</em> was formally proposed two decades ago but did not gain acceptance. Here instead, we formally raise <em>Corymbia</em> subg. <em>Blakella</em> to genus rank and make the relevant new combinations.</p>

opencc-zeroFeb 2023View details →
dryad36/100

DNA sequence data for two Roscoea species, R. stenophylla and R. australis (Zingiberaceae)

<p><span>This dataset includes three genomic regions, nrITS (ITS1-5.8S–ITS2) and two chloroplast DNA (cpDNA) regions (psbA-trnH and trnL-F) for two <em>Roscoea</em> species <em>R. stenophylla</em> and <em>R. australis</em> (Zingiberaceae).</span></p>

opencc-zeroMay 2023View details →
zenodo36/100

Models and Data associated with: Single-cell gene expression prediction from DNA sequence at large contexts

<p>This archive holds trained models and associated data&nbsp;for the <a href="https://www.biorxiv.org/content/10.1101/2023.07.26.550634v1">manuscript</a>:<br> &quot;Single-cell gene expression prediction from DNA sequence at large contexts&quot;</p> <p>Structure:</p> <ul> <li>configs&nbsp;- example configs for the workflows to produce publication data&nbsp;</li> <li>data_* - pre-processed single cell data used for publication</li> <li>models_* - model checkpoints, hyperparameters and training progress in tensorboard logs</li> <li>preprocessing - additional data required to reproduce the pre-processing workflow</li> </ul> <p>&nbsp;</p> <p>&quot;Copyright 2023 GlaxoSmithKline Research &amp; Development Limited. All rights reserved.&quot;</p>

opencc-by-nc-nd-4.0Sep 2023View details →
dryad36/100

Data from: Benchmarking ultra-high molecular weight DNA preservation methods for long-read and long-range sequencing

Open the record for dataset details and reuse information.

publicApr 2022View details →
dryad36/100

Data from: Concealed by darkness: interactions between predatory bats and nocturnally migrating songbirds illuminated by DNA sequencing

Open the record for dataset details and reuse information.

publicAug 2016View details →
dryad36/100

Data from: Adaptive radiation of the Callicarpa genus in the Bonin Islands revealed through double-digest restriction site–associated DNA sequencing analysis

Open the record for dataset details and reuse information.

publicAug 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record