Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

47

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

47 results for “sequence simulation”

Learn how ShareScore rates datasets ↗
dryad36/100

Data from Readsynth: short-read simulation for consideration of composition-biases in reduced metagenome sequencing approaches

Open the record for dataset details and reuse information.

publicApr 2024View details →
zenodo32/100

Fast Intratumor Heterogeneity Inference from Single-Cell Sequencing Data (simulated data - Extended Data Figures)

<p>This data repository contains simulated data used for benchmarking HUNTRESS against the existing alternative tools. Results of the benchmarking are shown in&nbsp;Extended Data Figures 1-10&nbsp;of the paper&nbsp;&quot;Fast Intratumor Heterogeneity Inference from Single-Cell Sequencing Data&quot; (to appear in Nature Computational Science).&nbsp;</p>

opencc-by-4.0Jul 2022View details →
zenodo32/100

Tree Sequence and Genealogical Forest Files for a Simulated Human Chromosome 20

<p>Dataset containing 640000 samples simulated using <a href="https://github.com/popsim-consortium/stdpopsim">stdpopsim</a> 0.2.0 and the <code>HapMapII_GRCh38</code> genetic map.<br>The tree sequence was converted to a genealogical forest files via <a href="https://github.com/lukashuebner/gfkit">gfkit</a> version <code>fbd2740</code>.</p>

opencc-by-4.0May 2024View details →
zenodo32/100

Sequencing reads simulated with ART from Homo sapiens CHM13 chr 21

<p>Sequencing reads simulated with ART from Homo sapiens CHM13 chr 21</p>

opencc-by-4.0Jul 2024View details →
zenodo32/100

Simulated hepatitis B virus (HBV) sequencing data and HBV sequence variation graph materials

<p>Simulated HBV sequencing data (InSilicoSeq, HiSeq error model) and a sequence variation graph constructed using HBV genome sequences described in <a href="https://doi.org/10.1099/jgv.0.001387">https://doi.org/10.1099/jgv.0.001387</a></p>

opencc-by-4.0Jun 2022View details →
dryad32/100

Simulated eukaryotic genomic sequencing, long and short reads

<p><span>As accuracy and throughput of nanopore sequencing improves, it is increasingly common to perform long-read-first </span><em>de novo</em> genome assemblies followed by polishing with accurate short reads (Kim et al. 2021). We briefly introduce FMLRC2, the successor to the original FM-index Long Read Corrector (FMLRC), and illustrate its performance as a fast and accurate <em>de novo</em> assembly polisher for both bacterial and eukaryotic genomes.</p>

opencc-zeroJan 2023View details →
dryad32/100

Simulated eukaryotic genomic sequencing, long and short reads

Open the record for dataset details and reuse information.

publicJan 2023View details →
zenodo28/100

Simulated data and results from "Accurate quantification of copy-number aberrations and whole-genome duplications in multi-sample tumor sequencing data"

<p>This dataset contains all the simulated data and the results of all the considered methods in the benchmark presented in &quot;Accurate quantification of copy-number aberrations and whole-genome duplications in multi-sample tumor sequencing data&quot; [Zaccaria &amp; Raphael, 2018]. All the data in this dataset and the corresponding formats are fully described at&nbsp;<a href="https://github.com/raphael-group/hatchet-paper">https://github.com/raphael-group/hatchet-paper</a>. The folder <em>simulations</em>&nbsp;which contains the entire dataset has been compressed with standard <em>zip</em>.</p>

opencc-by-4.0May 2020View details →
zenodo28/100

Simulated read data analysed in "Removing reference bias and improving indel calling in ancient DNA data analysis by mapping to a sequence variation graph"

<p>Simulated read data analyzed in &quot;Removing reference bias and improving indel calling in ancient DNA data analysis by mapping to a sequence variation graph&quot;.</p> <p><strong>1) Human sequence data</strong></p> <p><strong>HO_chr11_50bp_sliding_window*fq.gz:</strong><br> All possible 50 bp reads overlapping chromosome 11 SNPs in the Human Origins dataset. Files with the word &quot;alternate&quot; in their filename carry the alternate allele, otherwise, they carry the reference allele. Deamination has been added into these simulated reads using gargammel (Renaud 2016) based on empirically estimated post-mortem damage in a dataset of 102 ancient genomes (Allentoft et al., 2015).</p> <p><strong>2) microbial data</strong></p> <p><strong>simulation_*_s.fq.gz:</strong><br> Simulated microbial read data&nbsp;from a set of microbial reference genomes identified in the ancient Clovis genome (Rasmussen 2014), using gargammel.</p>

opencc-by-4.0Jul 2020View details →
zenodo28/100

Model trees and associated simulated nucleotide sequences for testing phylogenetic inference methods

<p>This repository contains 142 tar.gz archive files, each containing nucleotide sequence data that have been simulated using <a href="http://abacus.gene.ucl.ac.uk/software/indelible/"><em>INDELible</em></a> for testing alignment-free phylogenetic inference methods. These datasets were generated by using the results (trees and model parameters) of 142 phylogenomic analyses of real-case data as model (available <a href="https://zenodo.org/record/4034261">here</a>). Initial sequence length was 5 Mbs, and an indel rate of 0.01 was set with indel length drawn from [1, 50000] according to a Zipf distribution with parameter 1.5 (see <em>INDELible</em> <a href="http://abacus.gene.ucl.ac.uk/software/indelible/manual/model.shtml">manual</a>).</p> <p>Each archive contains the following files/directories:</p> <ul> <li><code>GTR.params.trees.tsv &nbsp; </code> &nbsp; a tab-delimited file summarizing the real-case GTR+&Gamma; model parameters and the phylogenetic tree used to simulate the sequence dataset (gathered from <a href="https://zenodo.org/record/4034261">https://zenodo.org/record/4034261</a>)</li> <li><code>tax.tsv &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp; </code> &nbsp; a tab-delimited file containing the initial (col 1) and simplified (col 2) taxon names</li> <li><code>model.nwk &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp; </code> &nbsp; a <a href="https://evolution.genetics.washington.edu/phylip/newicktree.html">Newick</a>-formatted file containing the initial model tree (gathered from <code>GTR.params.trees.tsv</code>) with simplified leaf names (following <code>tax.tsv</code>)</li> <li><code>control.txt &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp; </code> &nbsp; the <em>INDELible</em> input file used to simulate the evolution of a sequence along the tree in <code>model.nwk</code></li> <li><code>seq/ &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; </code> &nbsp; a directory containing the simulated sequences (one FASTA file per leaf in the tree in <code>model.nwk</code>)</li> </ul> <p>___</p> <p>Criscuolo A (2020) <em>On the transformation of MinHash-based uncorrected distances into proper evolutionary distances for phylogenetic inference</em>. F1000Research, 9:1309. <a href="https://doi.org/10.12688/f1000research.26930.1">doi:10.12688/f1000research.26930.1</a></p>

opencc-by-4.0Sep 2020View details →
zenodo28/100

Simulated pairs of nucleotide sequences for testing (alignment-free) genome distance estimate methods

<p>This repository contains 24,000 pairs of nucleotide sequences (and associated parameters) that have been simulated for testing alignment-free genome distance estimates. Given an evolutionary distance <em>d</em> varying from 0.05 to 1.00 nucleotide substitutions per character (step = 0.05), the program <a href="http://abacus.gene.ucl.ac.uk/software/indelible/"><em>INDELible</em></a> was used to simulate the evolution of 200 nucleotide sequence pairs with <em>d</em> substitution events per character under the models GTR and GTR+&Gamma;. Each model was adjusted with three different equilibrium frequencies:</p> <ul> <li><em>f</em><sub>1</sub>: equal frequencies, i.e. freq(A) = freq(C) = freq(G) = freq(T) = 0.25,</li> <li><em>f</em><sub>2</sub>: GC-rich, i.e. freq(A) = 0.1, freq(C) = 0.3, freq(G) = 0.4, freq(T) = 0.2,</li> <li><em>f</em><sub>3</sub>: AT-rich, i.e. freq(A) = freq(T) = 0.4, freq(C) = freq(G) = 0.1.</li> </ul> <p>For each simulated sequence pair, model parameters (i.e. GTR: six relative rates of nucleotide substitution; GTR+&Gamma;: six rates and one &Gamma; shape parameter) were randomly drawn from 142 sets of parameters derived from real-case data (see file <a href="https://zenodo.org/record/4034261/files/GTR.params.trees.tsv?download=1">GTR.params.trees.tsv</a> at <a href="https://zenodo.org/record/4034261">https://zenodo.org/record/4034261</a>). Initial sequence length was 5 Mbs, and an indel rate of 0.01 was set with indel length drawn from [1, 50000] according to a Zipf distribution with parameter 1.5 (see <em>INDELible</em> <a href="http://abacus.gene.ucl.ac.uk/software/indelible/manual/model.shtml">manual</a>).</p> <p>&nbsp;</p> <p>For each of the 20 evolutionary distances <em>d</em> = 0.05, 0.10, ..., 1.00, six XZ-compressed files containing 200 simulation data are available:</p> <ul> <li><code>data-d-f1-nogam.tsv.xz</code> &nbsp; data simulated under the model GTR with equilibrium frequencies <em>f</em><sub>1</sub></li> <li><code>data-d-f1-gamma.tsv.xz</code> &nbsp; data simulated under the model GTR+&Gamma; with equilibrium frequencies <em>f</em><sub>1</sub></li> <li><code>data-d-f2-nogam.tsv.xz</code> &nbsp; data simulated under the model GTR with equilibrium frequencies <em>f</em><sub>2</sub></li> <li><code>data-d-f2-gamma.tsv.xz</code> &nbsp; data simulated under the model GTR+&Gamma; with equilibrium frequencies <em>f</em><sub>2</sub></li> <li><code>data-d-f3-nogam.tsv.xz</code> &nbsp; data simulated under the model GTR with equilibrium frequencies <em>f</em><sub>3</sub></li> <li><code>data-d-f3-gamma.tsv.xz</code> &nbsp; data simulated under the model GTR+&Gamma; with equilibrium frequencies <em>f</em><sub>3</sub></li> </ul> <p>&nbsp;</p> <p>Each file is tab-delimited and contains the 18 following fields:</p> <ul> <li><code>[1]&nbsp; &nbsp;</code>&nbsp;&nbsp; integer <em>seed</em> value specified to <em>INDELible</em>,</li> <li><code>[2-5]&nbsp;</code> &nbsp; frequencies of T, C, A, G, respectively, specified to <em>INDELible</em>,</li> <li><code>[6-10]&nbsp;</code> C-T, A-T, G-T, A-C, C-G rate parameters, respectivly (normalized such that A-G rate = 1), specified to <em>INDELible</em>,</li> <li><code>[11] &nbsp; </code> &nbsp; &Gamma; shape parameter <em>alpha</em> (= 0 in the <code>nogam</code> files, i.e. GTR substitution model without &Gamma;) specified to <em>INDELible</em>,</li> <li><code>[12] &nbsp; </code> &nbsp; length <em>lgt1</em> of the first sequence <em>seq1</em> (i.e. no. A, C, G, T in <em>seq1</em>),</li> <li><code>[13] &nbsp; </code> &nbsp; length <em>lgt2</em> of the second sequence <em>seq2</em> (i.e. no. A, C, G, T in <em>seq2</em>),</li> <li><code>[14] &nbsp; </code> &nbsp; no. <em>sites</em> in aligned sequences <em>seq1</em> and <em>seq2</em> (i.e. no. A, C, G, T and gap character states in <em>seq1</em> or <em>seq2</em>),</li> <li><code>[15] &nbsp; </code> &nbsp; no. non-gapped sites (<em>core</em> sites) in aligned sequences <em>seq1</em> and <em>seq2</em>,</li> <li><code>[16] &nbsp; </code> &nbsp; observed <em>p-distance</em> between aligned sequences <em>seq1</em> and <em>seq2</em> (i.e. no. nucleotide mismatches divided by no. <em>core</em> sites),</li> <li><code>[17] &nbsp; </code> &nbsp; aligned <em>seq1</em> (containing indel gaps),</li> <li><code>[18] &nbsp; </code> &nbsp; aligned <em>seq2</em> (containing indel gaps).</li> </ul> <p>Of note, <em>seq1</em> and <em>seq2</em> (fields <code>[17-18]</code>) being aligned, these two entries are two strings with identical no. <em>sites</em> (field <code>[14]</code>). Gap character states (<code>-</code>) should be removed from <em>seq1</em> and <em>seq2</em> to obtain the unaligned sequences.</p> <p>_____</p> <p>Criscuolo A (2020) <em>On the transformation of MinHash-based uncorrected distances into proper evolutionary distances for phylogenetic inference</em>. F1000Research, 9:1309. <a href="https://doi.org/10.12688/f1000research.26930.1">doi:10.12688/f1000research.26930.1</a></p>

opencc-by-4.0Sep 2020View details →
dryad28/100

Data from: Simple chained guide trees give poorer multiple sequence alignments than inferred trees in simulation and phylogenetic benchmarks

Multiple sequence aligners typically work by progressively aligning the most closely related sequences or group of sequences according to guide trees. In PNAS, Boyce et al. report that alignments reconstructed using simple chained trees (i.e., comb-like topologies) with random leaf assignment performed better in protein structure-based benchmarks than those reconstructed using phylogenies estimated from the data as guide trees. The authors state that this result could turn decades of research in the field on its head. In light of this statement, it is important to check immediately whether their result holds under evolutionary criteria: recovery of homologous sequence residues and inference of phylogenetic trees from the alignments. We have done this and the results are entirely opposed to Boyce et al.'s findings.

opencc-zeroDec 2014View details →
dryad28/100

Data from: Unforeseen consequences of excluding missing data from next-generation sequences: simulation study of RAD sequences

There is a lack of consensus on how next-generation sequence data should be considered for phylogenetic and phylogeographic estimates, with some studies excluding loci with missing data, while others include them, even when sequences are missing from a large number of individuals. Here we use simulations, focusing specifically on RAD sequences, to highlight some of the unforeseen consequence of excluding missing data from next-generation sequencing. Specifically, we show that in addition to the obvious effects associated with reducing the amount of data used to make historical inferences, the decisions we make about missing data (such as the minimum number of individuals with a sequence for a locus to be included in the study) also impact the types of loci sampled for a study. In particular, as the tolerance for missing data becomes more stringent, the mutational spectrum represented in the sampled loci becomes truncated such that loci with the highest mutation rates are disproportionately excluded. This effect is exacerbated further by factors involved in the preparation of the genomic library (i.e., the use of reduced representation libraries, as well as the coverage) and the taxonomic diversity represented in the library (i.e., the level of divergence among the individuals). We demonstrate that the intuitive appeals about being conservative by removing loci may be misguided.

opencc-zeroDec 2013View details →
zenodo28/100

Datasets associated with the manuscript "Discovering SARS-CoV-2 neoepitopes and the associated TCR-pMHC recognition mechanisms by combining single-cell sequencing, deep learning, and molecular dynamics simulation techniques"

<p>meta_data_TCR-pMHC_from_STCRDab.tsv, TCR-pMHC structures used for contacts analysis.</p><p>tcr_gliph_input_sars2.tsv, input files (TCR sequences and related information) used for clustering TCRs targeting SARS-CoV-2 epitopes and epitope-unknown TCRs.</p><p>tcr_gliph_input_non-sars2.tsv, input files used for clustering TCRs targeting non-SARS-CoV-2 epitopes and epitope-unknown TCRs.</p><p>tcr_gliph_output*, output files from the GLIPH software, including the recognized TCR clusters by GLIPH (convergence-group.txt), the linkage information of TCR clusters (clone-network.txt), and the recognized motif in TCR clusters (kmer.txt).</p><p>md_trajs.tar, structures and MD simulation trajectories of TCR-614-pMHC and TCR-204-pMHC complexes.</p>

opencc-by-4.0Oct 2023View details →
zenodo28/100

research data supporting "Revealing the organization of catalytic sequence-defined oligomers via combined molecular dynamics simulations and network analysis"

<p>This repository contains all the data generated and analyzed including the starting structures, the input files, the trajectory files, the output data from cpptraj and network analyses, and in-house scripts used to prepare the network and module files shown in the paper&nbsp;<strong>&quot;Revealing the organization of catalytic sequence-defined oligomers via combined molecular dynamics simulations and network analysis&quot;</strong>&nbsp;published in&nbsp;<strong>Journal of Chemical Information and Modeling</strong>&nbsp;(DOI: 10.1021/acs.jcim.2c00101).&nbsp;</p>

openother-openMay 2022View details →
zenodo28/100

Squeegee: de novo identification of reagent and laboratory induced microbial contaminants in low biomass microbiomes, simulation dataset 0.5% spike-in contaminant sequences

<p>Computational analysis of host-associated microbiomes has opened the door to numerous discoveries relevant to human health and disease. However, contaminant sequences in metagenomic samples can potentially impact the interpretation of findings reported in microbiome studies, especially in low biomass environments. Our hypothesis is that contamination from DNA extraction kits or sampling lab environments will leave taxonomic &quot;bread crumbs&rdquo; across multiple distinct sample types, allowing for the detection of microbial contaminants when negative controls are unavailable. To test this hypothesis we implemented Squeegee, a de novo contamination detection tool. We tested Squeegee on simulated and real low biomass metagenomic datasets. On the low biomass samples, we compared Squeegee predictions to experimental negative control data and show that Squeegee accurately recovers known contaminants. We also analyzed 749 metagenomic datasets from the Human Microbiome Project and identified likely previously unreported kit contamination. Collectively, our results highlight that Squeegee can identify microbial contaminants with high precision.</p> <p>&nbsp;</p> <p>Simulation Dataset 0.5% contaminant spike-in.&nbsp;</p>

opencc-by-4.0Sep 2022View details →
dryad28/100

Data from: RAD sequencing and genomic simulations resolve hybrid origins within North American Canis

Top predators are disappearing worldwide, significantly changing ecosystems that depend on top-down regulation. Conflict with humans remains the primary roadblock for large carnivore conservation, but for the eastern wolf (Canis lycaon), disagreement over its evolutionary origins presents a significant barrier to conservation in Canada and has impeded protection for grey wolves (Canis lupus) in the USA. Here, we use 127 235 single-nucleotide polymorphisms (SNPs) identified from restriction-site associated DNA sequencing (RAD-seq) of wolves and coyotes, in combination with genomic simulations, to test hypotheses of hybrid origins of Canis types in eastern North America. A principal components analysis revealed no evidence to support eastern wolves, or any other Canis type, as the product of grey wolf × western coyote hybridization. In contrast, simulations that included eastern wolves as a distinct taxon clarified the hybrid origins of Great Lakes-boreal wolves and eastern coyotes. Our results support the eastern wolf as a distinct genomic cluster in North America and help resolve hybrid origins of Great Lakes wolves and eastern coyotes. The data provide timely information that will shed new light on the debate over wolf conservation in eastern North America.

opencc-zeroDec 2016View details →
dryad28/100

Data from: Simple chained guide trees give poorer multiple sequence alignments than inferred trees in simulation and phylogenetic benchmarks

Open the record for dataset details and reuse information.

publicDec 2015View details →
dryad28/100

Data from: Unforeseen consequences of excluding missing data from next-generation sequences: simulation study of RAD sequences

Open the record for dataset details and reuse information.

publicJun 2014View details →
dryad28/100

Data from: RAD sequencing and genomic simulations resolve hybrid origins within North American Canis

Open the record for dataset details and reuse information.

publicSep 2017View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record