Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
47
datasets available to search
ShareScore release 0.9.0
Dataset results
47 results for “sequence simulation”
Data from Readsynth: short-read simulation for consideration of composition-biases in reduced metagenome sequencing approaches
Open the record for dataset details and reuse information.
Fast Intratumor Heterogeneity Inference from Single-Cell Sequencing Data (simulated data - Extended Data Figures)
<p>This data repository contains simulated data used for benchmarking HUNTRESS against the existing alternative tools. Results of the benchmarking are shown in Extended Data Figures 1-10 of the paper "Fast Intratumor Heterogeneity Inference from Single-Cell Sequencing Data" (to appear in Nature Computational Science). </p>
Tree Sequence and Genealogical Forest Files for a Simulated Human Chromosome 20
<p>Dataset containing 640000 samples simulated using <a href="https://github.com/popsim-consortium/stdpopsim">stdpopsim</a> 0.2.0 and the <code>HapMapII_GRCh38</code> genetic map.<br>The tree sequence was converted to a genealogical forest files via <a href="https://github.com/lukashuebner/gfkit">gfkit</a> version <code>fbd2740</code>.</p>
Sequencing reads simulated with ART from Homo sapiens CHM13 chr 21
<p>Sequencing reads simulated with ART from Homo sapiens CHM13 chr 21</p>
Simulated hepatitis B virus (HBV) sequencing data and HBV sequence variation graph materials
<p>Simulated HBV sequencing data (InSilicoSeq, HiSeq error model) and a sequence variation graph constructed using HBV genome sequences described in <a href="https://doi.org/10.1099/jgv.0.001387">https://doi.org/10.1099/jgv.0.001387</a></p>
Simulated eukaryotic genomic sequencing, long and short reads
<p><span>As accuracy and throughput of nanopore sequencing improves, it is increasingly common to perform long-read-first </span><em>de novo</em> genome assemblies followed by polishing with accurate short reads (Kim et al. 2021). We briefly introduce FMLRC2, the successor to the original FM-index Long Read Corrector (FMLRC), and illustrate its performance as a fast and accurate <em>de novo</em> assembly polisher for both bacterial and eukaryotic genomes.</p>
Simulated eukaryotic genomic sequencing, long and short reads
Open the record for dataset details and reuse information.
Simulated data and results from "Accurate quantification of copy-number aberrations and whole-genome duplications in multi-sample tumor sequencing data"
<p>This dataset contains all the simulated data and the results of all the considered methods in the benchmark presented in "Accurate quantification of copy-number aberrations and whole-genome duplications in multi-sample tumor sequencing data" [Zaccaria & Raphael, 2018]. All the data in this dataset and the corresponding formats are fully described at <a href="https://github.com/raphael-group/hatchet-paper">https://github.com/raphael-group/hatchet-paper</a>. The folder <em>simulations</em> which contains the entire dataset has been compressed with standard <em>zip</em>.</p>
Simulated read data analysed in "Removing reference bias and improving indel calling in ancient DNA data analysis by mapping to a sequence variation graph"
<p>Simulated read data analyzed in "Removing reference bias and improving indel calling in ancient DNA data analysis by mapping to a sequence variation graph".</p> <p><strong>1) Human sequence data</strong></p> <p><strong>HO_chr11_50bp_sliding_window*fq.gz:</strong><br> All possible 50 bp reads overlapping chromosome 11 SNPs in the Human Origins dataset. Files with the word "alternate" in their filename carry the alternate allele, otherwise, they carry the reference allele. Deamination has been added into these simulated reads using gargammel (Renaud 2016) based on empirically estimated post-mortem damage in a dataset of 102 ancient genomes (Allentoft et al., 2015).</p> <p><strong>2) microbial data</strong></p> <p><strong>simulation_*_s.fq.gz:</strong><br> Simulated microbial read data from a set of microbial reference genomes identified in the ancient Clovis genome (Rasmussen 2014), using gargammel.</p>
Model trees and associated simulated nucleotide sequences for testing phylogenetic inference methods
<p>This repository contains 142 tar.gz archive files, each containing nucleotide sequence data that have been simulated using <a href="http://abacus.gene.ucl.ac.uk/software/indelible/"><em>INDELible</em></a> for testing alignment-free phylogenetic inference methods. These datasets were generated by using the results (trees and model parameters) of 142 phylogenomic analyses of real-case data as model (available <a href="https://zenodo.org/record/4034261">here</a>). Initial sequence length was 5 Mbs, and an indel rate of 0.01 was set with indel length drawn from [1, 50000] according to a Zipf distribution with parameter 1.5 (see <em>INDELible</em> <a href="http://abacus.gene.ucl.ac.uk/software/indelible/manual/model.shtml">manual</a>).</p> <p>Each archive contains the following files/directories:</p> <ul> <li><code>GTR.params.trees.tsv </code> a tab-delimited file summarizing the real-case GTR+Γ model parameters and the phylogenetic tree used to simulate the sequence dataset (gathered from <a href="https://zenodo.org/record/4034261">https://zenodo.org/record/4034261</a>)</li> <li><code>tax.tsv </code> a tab-delimited file containing the initial (col 1) and simplified (col 2) taxon names</li> <li><code>model.nwk </code> a <a href="https://evolution.genetics.washington.edu/phylip/newicktree.html">Newick</a>-formatted file containing the initial model tree (gathered from <code>GTR.params.trees.tsv</code>) with simplified leaf names (following <code>tax.tsv</code>)</li> <li><code>control.txt </code> the <em>INDELible</em> input file used to simulate the evolution of a sequence along the tree in <code>model.nwk</code></li> <li><code>seq/ </code> a directory containing the simulated sequences (one FASTA file per leaf in the tree in <code>model.nwk</code>)</li> </ul> <p>___</p> <p>Criscuolo A (2020) <em>On the transformation of MinHash-based uncorrected distances into proper evolutionary distances for phylogenetic inference</em>. F1000Research, 9:1309. <a href="https://doi.org/10.12688/f1000research.26930.1">doi:10.12688/f1000research.26930.1</a></p>
Simulated pairs of nucleotide sequences for testing (alignment-free) genome distance estimate methods
<p>This repository contains 24,000 pairs of nucleotide sequences (and associated parameters) that have been simulated for testing alignment-free genome distance estimates. Given an evolutionary distance <em>d</em> varying from 0.05 to 1.00 nucleotide substitutions per character (step = 0.05), the program <a href="http://abacus.gene.ucl.ac.uk/software/indelible/"><em>INDELible</em></a> was used to simulate the evolution of 200 nucleotide sequence pairs with <em>d</em> substitution events per character under the models GTR and GTR+Γ. Each model was adjusted with three different equilibrium frequencies:</p> <ul> <li><em>f</em><sub>1</sub>: equal frequencies, i.e. freq(A) = freq(C) = freq(G) = freq(T) = 0.25,</li> <li><em>f</em><sub>2</sub>: GC-rich, i.e. freq(A) = 0.1, freq(C) = 0.3, freq(G) = 0.4, freq(T) = 0.2,</li> <li><em>f</em><sub>3</sub>: AT-rich, i.e. freq(A) = freq(T) = 0.4, freq(C) = freq(G) = 0.1.</li> </ul> <p>For each simulated sequence pair, model parameters (i.e. GTR: six relative rates of nucleotide substitution; GTR+Γ: six rates and one Γ shape parameter) were randomly drawn from 142 sets of parameters derived from real-case data (see file <a href="https://zenodo.org/record/4034261/files/GTR.params.trees.tsv?download=1">GTR.params.trees.tsv</a> at <a href="https://zenodo.org/record/4034261">https://zenodo.org/record/4034261</a>). Initial sequence length was 5 Mbs, and an indel rate of 0.01 was set with indel length drawn from [1, 50000] according to a Zipf distribution with parameter 1.5 (see <em>INDELible</em> <a href="http://abacus.gene.ucl.ac.uk/software/indelible/manual/model.shtml">manual</a>).</p> <p> </p> <p>For each of the 20 evolutionary distances <em>d</em> = 0.05, 0.10, ..., 1.00, six XZ-compressed files containing 200 simulation data are available:</p> <ul> <li><code>data-d-f1-nogam.tsv.xz</code> data simulated under the model GTR with equilibrium frequencies <em>f</em><sub>1</sub></li> <li><code>data-d-f1-gamma.tsv.xz</code> data simulated under the model GTR+Γ with equilibrium frequencies <em>f</em><sub>1</sub></li> <li><code>data-d-f2-nogam.tsv.xz</code> data simulated under the model GTR with equilibrium frequencies <em>f</em><sub>2</sub></li> <li><code>data-d-f2-gamma.tsv.xz</code> data simulated under the model GTR+Γ with equilibrium frequencies <em>f</em><sub>2</sub></li> <li><code>data-d-f3-nogam.tsv.xz</code> data simulated under the model GTR with equilibrium frequencies <em>f</em><sub>3</sub></li> <li><code>data-d-f3-gamma.tsv.xz</code> data simulated under the model GTR+Γ with equilibrium frequencies <em>f</em><sub>3</sub></li> </ul> <p> </p> <p>Each file is tab-delimited and contains the 18 following fields:</p> <ul> <li><code>[1] </code> integer <em>seed</em> value specified to <em>INDELible</em>,</li> <li><code>[2-5] </code> frequencies of T, C, A, G, respectively, specified to <em>INDELible</em>,</li> <li><code>[6-10] </code> C-T, A-T, G-T, A-C, C-G rate parameters, respectivly (normalized such that A-G rate = 1), specified to <em>INDELible</em>,</li> <li><code>[11] </code> Γ shape parameter <em>alpha</em> (= 0 in the <code>nogam</code> files, i.e. GTR substitution model without Γ) specified to <em>INDELible</em>,</li> <li><code>[12] </code> length <em>lgt1</em> of the first sequence <em>seq1</em> (i.e. no. A, C, G, T in <em>seq1</em>),</li> <li><code>[13] </code> length <em>lgt2</em> of the second sequence <em>seq2</em> (i.e. no. A, C, G, T in <em>seq2</em>),</li> <li><code>[14] </code> no. <em>sites</em> in aligned sequences <em>seq1</em> and <em>seq2</em> (i.e. no. A, C, G, T and gap character states in <em>seq1</em> or <em>seq2</em>),</li> <li><code>[15] </code> no. non-gapped sites (<em>core</em> sites) in aligned sequences <em>seq1</em> and <em>seq2</em>,</li> <li><code>[16] </code> observed <em>p-distance</em> between aligned sequences <em>seq1</em> and <em>seq2</em> (i.e. no. nucleotide mismatches divided by no. <em>core</em> sites),</li> <li><code>[17] </code> aligned <em>seq1</em> (containing indel gaps),</li> <li><code>[18] </code> aligned <em>seq2</em> (containing indel gaps).</li> </ul> <p>Of note, <em>seq1</em> and <em>seq2</em> (fields <code>[17-18]</code>) being aligned, these two entries are two strings with identical no. <em>sites</em> (field <code>[14]</code>). Gap character states (<code>-</code>) should be removed from <em>seq1</em> and <em>seq2</em> to obtain the unaligned sequences.</p> <p>_____</p> <p>Criscuolo A (2020) <em>On the transformation of MinHash-based uncorrected distances into proper evolutionary distances for phylogenetic inference</em>. F1000Research, 9:1309. <a href="https://doi.org/10.12688/f1000research.26930.1">doi:10.12688/f1000research.26930.1</a></p>
Data from: Simple chained guide trees give poorer multiple sequence alignments than inferred trees in simulation and phylogenetic benchmarks
Multiple sequence aligners typically work by progressively aligning the most closely related sequences or group of sequences according to guide trees. In PNAS, Boyce et al. report that alignments reconstructed using simple chained trees (i.e., comb-like topologies) with random leaf assignment performed better in protein structure-based benchmarks than those reconstructed using phylogenies estimated from the data as guide trees. The authors state that this result could turn decades of research in the field on its head. In light of this statement, it is important to check immediately whether their result holds under evolutionary criteria: recovery of homologous sequence residues and inference of phylogenetic trees from the alignments. We have done this and the results are entirely opposed to Boyce et al.'s findings.
Data from: Unforeseen consequences of excluding missing data from next-generation sequences: simulation study of RAD sequences
There is a lack of consensus on how next-generation sequence data should be considered for phylogenetic and phylogeographic estimates, with some studies excluding loci with missing data, while others include them, even when sequences are missing from a large number of individuals. Here we use simulations, focusing specifically on RAD sequences, to highlight some of the unforeseen consequence of excluding missing data from next-generation sequencing. Specifically, we show that in addition to the obvious effects associated with reducing the amount of data used to make historical inferences, the decisions we make about missing data (such as the minimum number of individuals with a sequence for a locus to be included in the study) also impact the types of loci sampled for a study. In particular, as the tolerance for missing data becomes more stringent, the mutational spectrum represented in the sampled loci becomes truncated such that loci with the highest mutation rates are disproportionately excluded. This effect is exacerbated further by factors involved in the preparation of the genomic library (i.e., the use of reduced representation libraries, as well as the coverage) and the taxonomic diversity represented in the library (i.e., the level of divergence among the individuals). We demonstrate that the intuitive appeals about being conservative by removing loci may be misguided.
Datasets associated with the manuscript "Discovering SARS-CoV-2 neoepitopes and the associated TCR-pMHC recognition mechanisms by combining single-cell sequencing, deep learning, and molecular dynamics simulation techniques"
<p>meta_data_TCR-pMHC_from_STCRDab.tsv, TCR-pMHC structures used for contacts analysis.</p><p>tcr_gliph_input_sars2.tsv, input files (TCR sequences and related information) used for clustering TCRs targeting SARS-CoV-2 epitopes and epitope-unknown TCRs.</p><p>tcr_gliph_input_non-sars2.tsv, input files used for clustering TCRs targeting non-SARS-CoV-2 epitopes and epitope-unknown TCRs.</p><p>tcr_gliph_output*, output files from the GLIPH software, including the recognized TCR clusters by GLIPH (convergence-group.txt), the linkage information of TCR clusters (clone-network.txt), and the recognized motif in TCR clusters (kmer.txt).</p><p>md_trajs.tar, structures and MD simulation trajectories of TCR-614-pMHC and TCR-204-pMHC complexes.</p>
research data supporting "Revealing the organization of catalytic sequence-defined oligomers via combined molecular dynamics simulations and network analysis"
<p>This repository contains all the data generated and analyzed including the starting structures, the input files, the trajectory files, the output data from cpptraj and network analyses, and in-house scripts used to prepare the network and module files shown in the paper <strong>"Revealing the organization of catalytic sequence-defined oligomers via combined molecular dynamics simulations and network analysis"</strong> published in <strong>Journal of Chemical Information and Modeling</strong> (DOI: 10.1021/acs.jcim.2c00101). </p>
Squeegee: de novo identification of reagent and laboratory induced microbial contaminants in low biomass microbiomes, simulation dataset 0.5% spike-in contaminant sequences
<p>Computational analysis of host-associated microbiomes has opened the door to numerous discoveries relevant to human health and disease. However, contaminant sequences in metagenomic samples can potentially impact the interpretation of findings reported in microbiome studies, especially in low biomass environments. Our hypothesis is that contamination from DNA extraction kits or sampling lab environments will leave taxonomic "bread crumbs” across multiple distinct sample types, allowing for the detection of microbial contaminants when negative controls are unavailable. To test this hypothesis we implemented Squeegee, a de novo contamination detection tool. We tested Squeegee on simulated and real low biomass metagenomic datasets. On the low biomass samples, we compared Squeegee predictions to experimental negative control data and show that Squeegee accurately recovers known contaminants. We also analyzed 749 metagenomic datasets from the Human Microbiome Project and identified likely previously unreported kit contamination. Collectively, our results highlight that Squeegee can identify microbial contaminants with high precision.</p> <p> </p> <p>Simulation Dataset 0.5% contaminant spike-in. </p>
Data from: RAD sequencing and genomic simulations resolve hybrid origins within North American Canis
Top predators are disappearing worldwide, significantly changing ecosystems that depend on top-down regulation. Conflict with humans remains the primary roadblock for large carnivore conservation, but for the eastern wolf (Canis lycaon), disagreement over its evolutionary origins presents a significant barrier to conservation in Canada and has impeded protection for grey wolves (Canis lupus) in the USA. Here, we use 127 235 single-nucleotide polymorphisms (SNPs) identified from restriction-site associated DNA sequencing (RAD-seq) of wolves and coyotes, in combination with genomic simulations, to test hypotheses of hybrid origins of Canis types in eastern North America. A principal components analysis revealed no evidence to support eastern wolves, or any other Canis type, as the product of grey wolf × western coyote hybridization. In contrast, simulations that included eastern wolves as a distinct taxon clarified the hybrid origins of Great Lakes-boreal wolves and eastern coyotes. Our results support the eastern wolf as a distinct genomic cluster in North America and help resolve hybrid origins of Great Lakes wolves and eastern coyotes. The data provide timely information that will shed new light on the debate over wolf conservation in eastern North America.
Data from: Simple chained guide trees give poorer multiple sequence alignments than inferred trees in simulation and phylogenetic benchmarks
Open the record for dataset details and reuse information.
Data from: Unforeseen consequences of excluding missing data from next-generation sequences: simulation study of RAD sequences
Open the record for dataset details and reuse information.
Data from: RAD sequencing and genomic simulations resolve hybrid origins within North American Canis
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.