Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

2,848

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

2,848 results for “sequence data”

Learn how ShareScore rates datasets ↗
dryad32/100

Sequencing data and normalized counts for tripartite RNAseq of Drosophila, Wolbachia, and SINV virus

Open the record for dataset details and reuse information.

publicJan 2021View details →
dryad32/100

Data for morphometric analysis and DNA barcode sequence for the new fish species Polymixia hollisterae

Open the record for dataset details and reuse information.

publicJul 2021View details →
dryad32/100

Sequencing data from validation experiment for base editing of SCN2A

Open the record for dataset details and reuse information.

publicMay 2023View details →
dryad32/100

Sea otter sequence capture project data files

Open the record for dataset details and reuse information.

publicDec 2021View details →
dryad32/100

Data for: Longitudinal metatranscriptomic sequencing of Southern California wastewater representing 16 million people from August 2020-21 reveals widespread transcription of antibiotic resistance genes

Open the record for dataset details and reuse information.

publicJul 2022View details →
dryad32/100

Data for: Sequencing and variant detection of eight abundant plant-infecting tobamoviruses across Southern California wastewater

Open the record for dataset details and reuse information.

publicAug 2022View details →
edi32/100

FAB 1 Forests and Biodiversity Experiment - High density diversity experiment:Soil lipid (P/NLFA) and AMF (spore and sequence) data from selected plots

A forest biodiversity experiment (FAB) focused on trees of our region investigates the consequences of multiple dimensions of tree diversity for soil, food webs, plant communities and ecosystems. FAB is designed to unravel effects of three forms of biological diversity: species richness (SR), functional diversity (FD), and phylogenetic diversity (PD). We define FD as the representation of multiple traits of leaves, roots, seeds, and the whole organism that are correlated with species positions along gradients of resource supply, growth, and decomposition. PD is the representation of evolutionary lineages measured as the genetic distances between species. While PD and FD are often correlated, convergent evolution and adaptive differentiation can decouple them. When functional traits that drive specific ecosystem functions are not phylogenetically conserved, PD and FD may give contrasting predictions. SR, PD, and FD are not independent, and we posit that PD may help explain SR effects, and FD may help explain both PD and SR effects. Thus FAB is designed to examine the separate and combined effects of all three components of diversity for multiple ecosystem functions and to distinguish between ???sampling??? and ???complementarity??? effects of biodiversity. Due to the long lag between planting tree seedlings and determining effects of tree composition and diversity on ecosystem functioning, fewer experiments have been established to elucidate the role of biodiversity in the functioning of forest ecosystems than grassland experiments. FAB will contribute to this gap and is a member of the IDENT and TreeDiv network of forest biodiversity experiments (www.treedivnet.ugent.be). Hypotheses: 1. PD, FD, and SR will all contribute to increased productivity, stability, and diversity of other trophic levels (herbivores, predators, parasitoids, soil microbes, soil flora and fauna) as well as to greater soil C sequestration. 2. Because PD incorporates both the number of species a

openCC0Jun 2019View details →
zenodo28/100

Sequence Data for Templated Mutagenesis Analysis

<p>This repository contains the sequence data required for the analysis of templated mutagenesis.</p> <p>The sequences in the ebola&nbsp;folder are described in&nbsp; Bornholdt et al., &quot;Isolation of potent neutralizing antibodies from a survivor of the 2014 Ebola virus outbreak&quot; (dx.doi.org/10.1126/science.aad5788). The associated SRA accession numbers are listed in ebola/ebola_accessions_heavy.txt.</p> <p>The sequences in the yeap&nbsp;folder are described in Yeap et al., &quot;Sequence-Intrinsic Mechanisms that Target AID<br> Mutational Outcomes on Antibody Genes&quot; (http://dx.doi.org/10.1016/j.cell.2015.10.042). The associated SRA accession numbers are listed in yeap/SRR_Acc_List_gpt.txt.</p> <p>The sequences in the reference_sets&nbsp;folder are derived from IMGT (http://www.imgt.org/vquest/refseqh.html) and Retter et al., &quot;Sequence and Characterization of the Ig Heavy Chain Constant and Partial Variable Region of the Mouse Strain 129S1&quot; (https://doi.org/10.4049/jimmunol.179.4.2419). More information about how the sequences were accessed and processed is available in reference_sets/README.md.</p>

opencc-by-4.0Sep 2019View details →
zenodo28/100

Simulated data and results from "Accurate quantification of copy-number aberrations and whole-genome duplications in multi-sample tumor sequencing data"

<p>This dataset contains all the simulated data and the results of all the considered methods in the benchmark presented in &quot;Accurate quantification of copy-number aberrations and whole-genome duplications in multi-sample tumor sequencing data&quot; [Zaccaria &amp; Raphael, 2018]. All the data in this dataset and the corresponding formats are fully described at&nbsp;<a href="https://github.com/raphael-group/hatchet-paper">https://github.com/raphael-group/hatchet-paper</a>. The folder <em>simulations</em>&nbsp;which contains the entire dataset has been compressed with standard <em>zip</em>.</p>

opencc-by-4.0May 2020View details →
zenodo28/100

Simulated read data analysed in "Removing reference bias and improving indel calling in ancient DNA data analysis by mapping to a sequence variation graph"

<p>Simulated read data analyzed in &quot;Removing reference bias and improving indel calling in ancient DNA data analysis by mapping to a sequence variation graph&quot;.</p> <p><strong>1) Human sequence data</strong></p> <p><strong>HO_chr11_50bp_sliding_window*fq.gz:</strong><br> All possible 50 bp reads overlapping chromosome 11 SNPs in the Human Origins dataset. Files with the word &quot;alternate&quot; in their filename carry the alternate allele, otherwise, they carry the reference allele. Deamination has been added into these simulated reads using gargammel (Renaud 2016) based on empirically estimated post-mortem damage in a dataset of 102 ancient genomes (Allentoft et al., 2015).</p> <p><strong>2) microbial data</strong></p> <p><strong>simulation_*_s.fq.gz:</strong><br> Simulated microbial read data&nbsp;from a set of microbial reference genomes identified in the ancient Clovis genome (Rasmussen 2014), using gargammel.</p>

opencc-by-4.0Jul 2020View details →
dryad28/100

Data from: More affordable and effective noninvasive SNP genotyping using high-throughput amplicon sequencing

<p>Non-invasive genotyping methods have become key elements of wildlife research over the last two decades, but their widespread adoption is limited by high costs, low success rates, and high error rates. <span>The information lost when genotyping success is low may lead to decreased precision in animal population densities, which could misguide conservation and management actions.</span> <span>Single nucleotide polymorphisms (SNPs) provide a promising alternative to traditionally used microsatellites as SNPs allow amplification of shorter DNA fragments, are less prone to genotyping errors, and produce results that are easily shared among laboratories.</span> Here, we outline a detailed protocol for cost-effective and accurate noninvasive SNP genotyping using multiplexed amplicon sequencing optimized for degraded DNA. <span>We validated this method for individual identification by genotyping 216 scats, 18 hairs and 15 tissues from coyotes (<i>Canis latrans</i>) using 26 SNPs. </span><a name="_Hlk33181599">Our genotyping success rate for scat samples was 93%, and 100% for hair and tissue, representing a substantial increase compared to previous microsatellite-based studies while remaining at a low cost of under $5 per PCR replicate (excluding labor). </a>The accuracy of the genotypes was further corroborated in that genotypes from scats matching known, GPS-collared coyotes were always located within the territory of the known individual. We also show that different levels of multiplexing produced similar results, but that PCR product cleanup strategies can have substantial effects on genotyping success. By making noninvasive genotyping more affordable, accurate, and efficient, this research may allow for a substantial increase in the use of noninvasive methods to monitor and conserve free-ranging wildlife populations.</p>

opencc-zeroJun 2020View details →
zenodo28/100

Figure 1 in Comment on Psonis et al. (2015): 'Evaluation of the taxonomy of Helix cincta (Muller, 1774) and Helix nucula (Mousson, 1854); insights using mitochondrial DNA sequence data'

Figure 1. (A) Helix nucula Mousson, 1854, syntype ZMZ 506381a, shell diameter 28.8 mm. (B) Helix melanostoma Draparnaud, 1801, syntype of Helix uthicensis Pechaud, 1883, ruins of Uthica, Tunisia, MHNG 18185, shell diameter 36.2 mm. (C), Helix cincta Müller, 1774, syntype of Helix (Pomatia) cincta var. anatolica Kobelt, 1891, 'Hieronda' (southwest Turkey), SMF 9949, shell diameter 34.9 mm. (D) Helix pronuba Westerlund &amp; Blanc, 1879, syntype of Helix thiesseana var. pronuba Westerlund &amp; Blanc, 1879, GNM 1723, shell diameter 26.3 mm. (E) Helix borealis Mousson, 1859, possible syntype ZMZ 506307, Grecce, Argostoli, shell diameter 34.65 mm.

opencc-by-4.0Mar 2015View details →
dryad28/100

Pooled whole genome sequencing from year 2004 and Early-Late SNP data from year 2014

<p>Speciation underlies the generation of novel biodiversity. Yet, there is much to learn about how natural selection shapes genomes during speciation. Selection is assumed to act against gene flow at barrier loci, promoting reproductive isolation. However, evidence for gene flow and selection is often indirect and we know very little about the temporal stability of barrier loci. Here we utilize haplodiploidy to identify candidate male barrier loci in hybrids between two wood ant species. As ant males are haploid they are expected to reveal recessive barrier loci, which can be masked in diploid females if heterozygous. We then test for barrier stability in a sample collected ten years later and use survival analysis to provide a direct measure of natural selection acting on candidate male barrier loci. We find multiple candidate male barrier loci scattered throughout the genome. Surprisingly, a proportion of them are not stable after ten years, natural selection apparently switching from acting against to favoring introgression in the later sample. Instability of barrier effect and natural selection for introgressed alleles could be due to environment-dependent selection, emphasizing the need to consider temporal variation in the strength of natural selection and the stability of barrier effect at putative barrier loci in future speciation work.</p>

opencc-zeroAug 2020View details →
dryad28/100

Data from: Adapterama II: universal amplicon sequencing on Illumina platforms (TaggiMatrix)

Next-generation sequencing (NGS) of amplicons is used in a wide variety of contexts. In many cases, NGS amplicon sequencing remains overly expensive and inflexible, with library preparation strategies relying upon the fusion of locus-specific primers to full-length adapter sequences with a single identifying sequence or ligating adapters onto PCR products. In Adapterama I, we presented universal stubs and primers to produce thousands of unique index combinations and a modifiable system for incorporating them into Illumina libraries. Here, we describe multiple ways to use the Adapterama system and other approaches for amplicon sequencing on Illumina instruments. In the variant we use most frequently for large-scale projects, we fuse partial adapter sequences (TruSeq or Nextera) onto the 5' end of locus-specific PCR primers with variable-length tag sequences between the adapter and locus-specific sequences. These fusion primers can be used combinatorially to amplify samples within a 96-well plate (eight forward primers + 12 reverse primers yield 8 x 12 = 96 combinations), and the resulting amplicons can be pooled. The initial PCR products then serve as template for a second round of PCR with dual-indexed iTru or iNext primers (also used combinatorially) to make full-length libraries. The resulting quadruple-indexed amplicons have diversity at most base positions and can be pooled with any standard Illumina library for sequencing. The number of sequencing reads from the amplicon pools can be adjusted, facilitating deep sequencing when required or reducing sequencing costs per sample to an economically trivial amount when deep coverage is not needed. We demonstrate the utility and versatility of our approaches with results from six projects using different implementations of our protocols. Thus, we show that these methods facilitate amplicon library construction for Illumina instruments at reduced cost with increased flexibility. A simple web page to design fusion primers compatible with iTru primers is available at: http://baddna.uga.edu/tools-taggi.html. A fast and easy to use program to demultiplex amplicon pools with internal indexes is available at: https://github.com/lefeverde/Mr_Demuxy.

opencc-zeroSep 2020View details →
zenodo28/100

MALDI-TOF-MS reference spectra and sequence data for African bovid collagen for Zooarchaeology by Mass Spectrometry (ZooMS)

<p>MALDI-TOF-MS spectra of extracted collagen from modern African bovids used as reference spectra to develop markers for Zooarchaeology by Mass Spectrometry (ZooMS).&nbsp; Some of this material was also analyzed by LC-MS/MS.&nbsp; That data can be found at&nbsp;MassIVE MSV000084675&nbsp;(<a href="https://doi.org/doi:10.25345/C5239K">doi:10.25345/C5239K</a>).&nbsp; Information about the species of the samples can be found in Key for Labels.csv file.</p> <p>The sequence data contains annotated alignments of the proteins COL1A1 and COL1A2 and the alignments for the available bovid collagen protein sequences.&nbsp; More information on these files can be found in the corresponding manuscript to this dataset.</p>

opencc-by-4.0Jul 2020View details →
dryad28/100

Raw data associated with the article: "Single-molecule DNA sequencing of widely varying GC-content using nucleotide release, capture and detection in microdroplets.", NAR, Puchtler et.al.

<p>All data taken in the production of the corresponding paper: "Single-molecule DNA sequencing of widely varying GC-content using nucleotide release, capture and detection in microdroplets."</p> <p>The associated manuscript describes a method for DNA sequencing which involves the sequential release of nucleotides from a single, immobilised strand of DNA via pyrophosphorolysis (PPL). Released nucleotides, in the form of dNTPs, are captured in microdroplets which are manipulated using an optical-EWOD platform. A detection chemistry within each droplet releases a specific dye depending on which dNTPs are present, allowing the optical read-out of bases within each droplet. Hence, by capturing bases sequentially within droplets as they are cleaved from the strand of DNA, the sequence can be optically identified.</p>

opencc-zeroOct 2020View details →
dryad28/100

Data from: Spontaneous hybridization and introgression between walleye (Sander vitreus) and sauger (S. canadensis) in two large reservoirs: insights from genotyping-by-sequencing

<p>Anthropogenic activities may facilitate undesirable hybridization and genomic introgression between fish species. Walleye (<i>Sander vitreus</i>) and sauger (<i>Sander canadensis</i>) are economically valuable freshwater species that can spontaneously hybridize in areas of sympatry. Levels of genomic introgression between walleye and sauger may be increased by modifications to waterbodies (e.g., reservoir development) and inadvertent propagation of hybrids in stocking programs. We used genotyping by sequencing (GBS) to examine 217 fish from two large reservoirs with mixed populations of walleye and sauger in Saskatchewan, Canada (Lake Diefenbaker, Tobin Lake). Analyses with 20,038 (r90) and 478 (r100) SNPs clearly resolved walleye and sauger, and classified hybrids with high confidence. F<sub>1</sub>, F<sub>2</sub>, and multi-generation hybrids were detected in Lake Diefenbaker, indicating potentially high levels of genomic introgression. In contrast, only F<sub>1</sub> hybrids were detected in Tobin Lake. Field classification of fish was unreliable; 7% of fish were misidentified based on broad species categories. Important for activities such as brood stock selection, 12/173 (7%) fish field-identified as pure walleye, and 1/24 (4%) identified as pure sauger were actually hybrids. In addition, 2/15 (13%) field-identified hybrids were actually pure walleye or sauger. We conclude that hybridization and introgression are occurring in Saskatchewan reservoirs, and that caution is warranted when using these populations in stocking programs. GBS offers a powerful and flexible tool for examining hybridization without pre-identification of informative loci, eliminating some of the key challenges associated with other marker types.</p>

opencc-zeroNov 2020View details →
zenodo28/100

Supplementary material 3 from: Fryssouli V, Zervakis GI, Polemis E, Typas MA (2020) A global meta-analysis of ITS rDNA sequences from material belonging to the genus Ganoderma (Basidiomycota, Polyporales) including new data from selected taxa. MycoKeys 75: 71-143. https://doi.org/10.3897/mycokeys.75.59872

Figure S2a

opencc-zeroDec 2020View details →
zenodo28/100

Supplementary material 7 from: Fryssouli V, Zervakis GI, Polemis E, Typas MA (2020) A global meta-analysis of ITS rDNA sequences from material belonging to the genus Ganoderma (Basidiomycota, Polyporales) including new data from selected taxa. MycoKeys 75: 71-143. https://doi.org/10.3897/mycokeys.75.59872

Figure S2e

opencc-zeroDec 2020View details →
zenodo28/100

Supplementary material 8 from: Fryssouli V, Zervakis GI, Polemis E, Typas MA (2020) A global meta-analysis of ITS rDNA sequences from material belonging to the genus Ganoderma (Basidiomycota, Polyporales) including new data from selected taxa. MycoKeys 75: 71-143. https://doi.org/10.3897/mycokeys.75.59872

Figure S2f

opencc-zeroDec 2020View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record