Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

47

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

47 results for “sequence simulation”

Learn how ShareScore rates datasets ↗
zenodo44/100

MACREL software benchmark data set: Simulated metagenomes with sequencing quality, errors profile and abundance distributions derived from real samples

<p>These metagenomes were used in the benchmarking of FACS pipeline, and were designed after NGLess benchmark dataset (doi.org/10.5281/zenodo.2560288).&nbsp; Metagenomes were simulated with <a href="https://www.niehs.nih.gov/research/resources/software/biostatistics/art/index.cfm">ART-bin-MountRainier-2016.06.05</a> using real abundance profiles (.abund files) available <a href="https://doi.org/10.5281/zenodo.2560288">elsewhere</a>, and <a href="http://progenomes1.embl.de/data/repGenomes/representatives.contigs.fasta.gz">proGenomes&#39; representative contigs</a> as reference genomes. There are available metagenomes with 40, 60 and 80 M (million of reads) based in the reference genomes and abundances of the following samples:</p> <pre><code>SAMEA2466916 SAMEA2466953 SAMEA2466965 SAMEA2621107 SAMEA2621229 SAMEA2621247</code></pre> <p>To convert them from the CRAM format back to fastq files:</p> <pre><code> ## 1. converting from cram to bam format: samtools view -b -T refgenome.fa -o file.bam file.cram ## 2. sorting the bam file: samtools sort -n file.bam -o input_sorted.bam # sort reads by identifier-name (-n) ## 3. converting from bam to fastq format: bedtools bamtofastq -i input_sorted.bam -fq output_r1.fastq -fq2 output_r2.fastq </code></pre> <p>&nbsp;</p>

opencc-by-4.0Nov 2019View details →
zenodo44/100

Simulation Data for "Community-Driven Code Comparisons for Three-Dimensional Dynamic Modeling of Sequences of Earthquakes and Aseismic Slip"

<p>Simulation data from Jiang et al. (2022), "Community-Driven Code Comparisons for Three-Dimensional Dynamic Modeling of Sequences of Earthquakes and Aseismic Slip," <em>Journal of Geophysical Research:&nbsp;Solid Earth</em><em>.</em></p> <p>The archive includes simulation data for 3D SEAS benchmarks BP4-QD and BP5-QD that are analyzed in our paper (descriptions in NOTES.txt)&nbsp;</p> <p><strong>BP4-QD Benchmark Simulations:</strong><br>1000 m: &nbsp;jiang.5, lambert.8, barbot.3, barbot.2, dliu.2, li.4<br>500 m:&nbsp; jiang.3, lambert.3, barbot.5, barbot.7, ozawa</p> <p><strong>BP5-QD Benchmark Simulations:</strong><br>2000 m: &nbsp;jiang.6, lambert.8, &nbsp;liu.4, cattania.5, &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;dli.7, barbot.3, dliu.10, li.3<br>1000 m:&nbsp; jiang.2, lambert.7, &nbsp;liu.5, cattania.3, ozawa, &nbsp; dli.5, barbot, &nbsp; dliu.6, &nbsp;li.2<br>500 m:&nbsp; jiang.4, lambert.9, &nbsp;liu.6, cattania.4, ozawa.2, dli.6, barbot.2, dliu.8<br>250 m:&nbsp; lambert.10, liu.7</p> <p><strong>BP5-QD with Off-Fault Data:</strong><br>1000 m: &nbsp;lambert.7, dli.5, barbot, &nbsp; dliu.6, li.2<br>500 m:&nbsp; lambert.9, dli.6, barbot.2, dliu.8</p> <p>Tables 2&ndash;4 in our paper summarizes details of numerical codes and selected simulations.</p> <p>The benchmark descriptions and the full suite of simulation data are available at SEAS online platform https://strike.scec.org/cvws/seas/.</p>

opencc-by-4.0Feb 2022View details →
zenodo44/100

A Simulated Heterozygous Diploid Genome for Third-gen Sequencing, Assembly, and Curation

<p>A simulated heterozygous diploid genome based on <em>Saccharomyces</em> <em>cerevisiae</em>, and <em>S. paradoxus</em> homologous chromosomes.</p> <p>Simulated PacBio subreads were generated from both parent haplomes and mixed together. A phased assembly was produced using FALCON assembler and FALCON Unzip (doi:10.1038/nmeth.4035). This dataset and assembly were then used to validate the Purge Haplotigs pipeline (https://bitbucket.org/mroachawri/purge_haplotigs). See workflow.sh for commands, comments and file descriptions.</p>

opencc-by-4.0Nov 2017View details →
zenodo44/100

bin3C - simulated community and associated sequencing datasets

<p>We simulated a human gut microbiome comprising 63 genomes from the GTDB annotated with an isolation source of faeces. No two genomes are more than 96% similar in terms of ANI.</p> <p>A Generalized Pareto distribution was used to model an abundance profile, which was assigned in random order to the references. There is a 50:1 difference between the most and least abundant member.</p> <p>Illumina shotgun and Hi-C reads were simulated using MetaART and sim3C (https://github.com/cerebis/sim3C).</p> <p>A sweep was performed over depth of coverage, by serially subsampling initial high depth readsets. Shotgun depth was parameterised by the most abundant at 250x, while Hi-C was parameterised by the number of pairs (200 million pairs).</p> <p>Shotgun was subsampled once, at half depth (125x), while Hi-C was subsampled 4 times (12.5, 25, 50, 100, 200 million pairs).</p> <p>The random seed used throughout was 12345.</p> <p>These simulated&nbsp;readsets were then analyzed using bin3C to retrieve metagenome-assembled genomes (MAGs). The resulting genome bins were validated using CheckM to estimate completeness and contamination.</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Aug 2018View details →
zenodo44/100

Simulated Arabidopsis thaliana sequencing datasets for chloroplast assembler benchmarking

<p><strong>Changes</strong></p> <ul> <li>Fixed non-circular sampling from chloroplast and mitochondrion in version 1.1.0</li> <li>Fixed off-by-one error in reverse read in version 1.0.0</li> </ul> <p><strong>Purpose and Documentation</strong></p> <p>See: <a href="https://github.com/chloroExtractorTeam/benchmark">github.com/chloroExtractorTeam/benchmark</a></p> <p><strong>Original data</strong><br> The original <em>Arabidopsis thaliana </em>sequences were downloaded from TAIR:&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &nbsp;</p> <p>The Arabidopsis Information Resource (<a>TAIR</a>) on www.arabidopsis.org, Mar 22, 2019 available under the <a href="http://www.arabidopsis.org/doc/about/tair_terms_of_use/417">TAIR Terms of Use</a>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &nbsp;</p> <p><em>Tanya Z. Berardini, Leonore Reiser, Donghui Li, Yarik Mezheritsky, Robert Muller, Emily Strait and Eva Huala. &quot;The Arabidopsis Information Resource: Making and mining the &quot;gold standard&quot; annotated reference plant genome.&quot;&nbsp;&nbsp;&nbsp; genesis 2015 <a href="https://doi.org/10.1002/dvg.22877">doi:10.1002/dvg.22877</a></em></p> <p><strong>Programs used to generate this data</strong><br> &nbsp;- <a href="https://github.com/shenwei356/seqkit">seqkit</a> (v0.10.1): Shen W, Le S, Li Y, Hu F (2016) &quot;SeqKit: A Cross-Platform and Ultrafast Toolkit for FASTA/Q File Manipulation.&quot; PLOS ONE 11(10): e0163962. <a href="https://doi.org/10.1371/journal.pone.0163962">doi:10.1371/journal.pone.0163962</a></p> <p>&nbsp;</p>

opencc-by-4.0Apr 2019View details →
zenodo40/100

Automatic message sequence chart creation from simulation run of the parametric colored Petri net model of the Chandy-Lamport algorithm with four processes

<p><span>The video shows the creation of the message sequence chart from a simulation run of the proposed parametric colored Petri net model of the Chandy-Lamport algorithm using the CPN tool with four constituting processes. The picture shows the resulting message sequence chart. </span></p> <p><strong><span>Message Sequence Chart of Parametric Model With 4 Processes via Automatic Simulation Run_SuppInfo.mp4</span></strong><span>: This video shows the automatic generation of a message sequence chart of the proposed parametric colored Petri net model of the Chandy-Lamport distributed global snapshot algorithm using the CPN tool version 4.0.0. The model's number of constituting processes is parametric and was set to four. The video was generated using the authors' updated CPN tool extension server. The automatic simulation run of the model has been used to create this video. The CPN tool randomly selects the enabled transition at each step in an automatic simulation run.</span></p> <p><strong><span>Picture of Message Sequence Chart of Parametric Model With 4 Processes_SuppInfo.png:</span></strong><span> This picture shows the automatically generated message sequence chart of the proposed parametric colored Petri net model of the Chandy-Lamport algorithm that is visible in the above video clip. The number of constituting processes was set to four. </span></p>

opencc-by-4.0Apr 2024View details →
zenodo40/100

Squeegee: de novo identification of reagent and laboratory induced microbial contaminants in low biomass microbiomes, simulation dataset 0.25% spike-in contaminant sequences

<p>Computational analysis of host-associated microbiomes has opened the door to numerous discoveries relevant to human health and disease. However, contaminant sequences in metagenomic samples can potentially impact the interpretation of findings reported in microbiome studies, especially in low biomass environments. Our hypothesis is that contamination from DNA extraction kits or sampling lab environments will leave taxonomic &quot;bread crumbs&rdquo; across multiple distinct sample types, allowing for the detection of microbial contaminants when negative controls are unavailable. To test this hypothesis we implemented Squeegee, a de novo contamination detection tool. We tested Squeegee on simulated and real low biomass metagenomic datasets. On the low biomass samples, we compared Squeegee predictions to experimental negative control data and show that Squeegee accurately recovers known contaminants. We also analyzed 749 metagenomic datasets from the Human Microbiome Project and identified likely previously unreported kit contamination. Collectively, our results highlight that Squeegee can identify microbial contaminants with high precision.</p> <p>&nbsp;</p> <p>Simulation Dataset 0.25% contaminant spike-in.</p>

opencc-by-4.0Sep 2022View details →
zenodo40/100

Squeegee: de novo identification of reagent and laboratory induced microbial contaminants in low biomass microbiomes, simulation dataset 1% spike-in contaminant sequences

<p>Computational analysis of host-associated microbiomes has opened the door to numerous discoveries relevant to human health and disease. However, contaminant sequences in metagenomic samples can potentially impact the interpretation of findings reported in microbiome studies, especially in low biomass environments. Our hypothesis is that contamination from DNA extraction kits or sampling lab environments will leave taxonomic &quot;bread crumbs&rdquo; across multiple distinct sample types, allowing for the detection of microbial contaminants when negative controls are unavailable. To test this hypothesis we implemented Squeegee, a de novo contamination detection tool. We tested Squeegee on simulated and real low biomass metagenomic datasets. On the low biomass samples, we compared Squeegee predictions to experimental negative control data and show that Squeegee accurately recovers known contaminants. We also analyzed 749 metagenomic datasets from the Human Microbiome Project and identified likely previously unreported kit contamination. Collectively, our results highlight that Squeegee can identify microbial contaminants with high precision.</p> <p>&nbsp;</p> <p>Simulation Dataset 1% contaminant spike-in.</p>

opencc-by-4.0Sep 2022View details →
zenodo40/100

Simulated exome-sequencing data for a family study of lymphoid cancer

<p>This repository contains all the data files for a simulated exome-sequencing study of 150 families, ascertained to contain at least four members affected with lymphoid cancer.&nbsp; Please note that previous versions of this repository omitted a key file linking the genotypes of individuals to their family and individual IDs; this file, geno_key.txt, is now included. All other files remain the same as in previous versions.</p> <p>The simulated data can be found in&nbsp;the&nbsp;files section below. The files are:</p> <ol> <li>SLiM_output.txt - contains the&nbsp;SLiM-simulated, exome-wide, SNV data generated&nbsp;under an American-admixture demographic model,&nbsp; for&nbsp;the&nbsp;American-admixed sub-population only.</li> <li>SLiM_output_chr8&amp;9.txt -&nbsp;contains the&nbsp;SLiM-simulated data above for all source populations as well as the American-admixed sub-population, but&nbsp;only for&nbsp;chromosomes 8 and 9.</li> <li>sample_info.txt - contains pedigree information of all the disease-affected individuals and individuals connecting them along a line of descent, for all 150 ascertained&nbsp;pedigrees.</li> <li>Genotypes.zip -&nbsp; a zipfile that&nbsp;contains 22 text files of&nbsp;genotypes for each chromosome. The genotypes are for simulated&nbsp;single-nucleotide variants on the exome and are&nbsp;in gene-dosage format.&nbsp;</li> <li>geno_key.txt &ndash; a plain-text file that links the genotyped individuals to their family and individual IDs.</li> <li>SNVmaps.zip -&nbsp; a zipfile that&nbsp;contains 22 text files giving&nbsp;the single-nucleotide&nbsp;variant information for each chromosome.&nbsp;</li> <li>familial_cRV.txt - contains the familial causal rare variants for all 150 ascertained&nbsp;pedigrees.</li> <li>study_peds.txt - contains the 150 pedigrees ascertained to contain four or more relatives affected with lymphoid cancer.</li> <li>PLINKfiles.zip -&nbsp; a zipfile that contains PLINK .fam, .bim and .bed files for all 22 of the chromosomes.</li> </ol> <p>All the scripts used to generate these data&nbsp;can be found in the GitHub repository archived at <a href="../records/12694914">https://zenodo.org/records/12694914</a></p> <p>We have also&nbsp;uploaded one intermediate .Rdata file,&nbsp;Chromwide.Rdata, to save the user substantial time when running&nbsp;the associated RMarkdown script for the simulation. We recommend loading Chromwide.Rdata into your R work-space rather than generating it from scratch.</p>

openagpl-3.0-or-laterDec 2021View details →
zenodo40/100

Automatic message sequence chart creation from simulation run of the Chandy-Lamport algorithm modeled by colored Petri net

<p>Videos of two message sequence chart creation from simulation runs of the Chandy-Lamport algorithm modeled by colored Petri net using the CPN tool.</p> <p><strong>Message Sequence Chart Of Model With Automatic Simulation Run_SuppInfo.mp4</strong>: This video shows the automatic generation of a message sequence chart of a proposed colored Petri net model of the Chandy-Lamport distributed global snapshot algorithm using the CPN tool version 4.0.1. The video has been generated by the authors&#39; updated extension server of the CPN tool. The automatic simulation run of the model has been used to create this video. The CPN tool randomly selects the enabled transition in an automatic simulation run.</p> <p><strong>Message Sequence Chart Of Model With Step-By-Step Simulation Run_SuppInfo.mp4:</strong> This video shows the automatic generation of a message sequence chart of the proposed colored Petri net model of the Chandy-Lamport algorithm in a step-by-step simulation run with our updated extension server of the CPN tools version 4.0.1. We fired our selected enabled transition of the model to create this video.</p>

opencc-by-4.0Sep 2023View details →
dryad36/100

Data from Readsynth: short-read simulation for consideration of composition-biases in reduced metagenome sequencing approaches

<p><strong>Background</strong></p> <p>The application of reduced metagenomic sequencing approaches holds promise as a middle ground between targeted amplicon sequencing and whole metagenome sequencing approaches but has not been widely adopted as a technique. A major barrier to adoption is the lack of read simulation software built to handle characteristic features of these novel approaches. Reduced metagenomic sequencing (RMS) produces unique patterns of fragmentation per genome that are sensitive to restriction enzyme choice, and the non-uniform size selection of these fragments may introduce novel challenges to taxonomic assignment as well as relative abundance estimates.</p> <p><strong>Results</strong></p> <p>Through the development and application of simulation software, readsynth, we compare simulated metagenomic sequencing libraries with existing RMS data to assess the influence of multiple library preparation and sequencing steps on downstream analytical results. Based on read depth per position, readsynth achieved 0.79 Pearson's correlation and 0.94 Spearman's correlation to these benchmarks. Application of a novel estimation approach, fixed length taxonomic ratios, improved quantification accuracy of simulated human gut microbial communities when compared to estimates of mean or median coverage.</p> <p><strong>Conclusions</strong></p> <p>We investigate the possible strengths and weaknesses of applying the RMS technique to profiling microbial communities via simulations with readsynth. The choice of restriction enzymes and size selection steps in library prep are non-trivial decisions that bias downstream profiling and quantification. The simulations investigated in this study illustrate the possible limits of preparing metagenomic libraries with a reduced representation sequencing approach, but also allow for the development of strategies for producing and handling the sequence data produced by this promising application.</p>

opencc-zeroApr 2024View details →
zenodo36/100

Automatic message sequence chart creation from simulation run of the proposed parametric colored Petri net model of the Chandy-Lamport algorithm

<p><span>These videos show the creation of two message sequence charts from simulation runs of the proposed parametric colored Petri net model of the Chandy-Lamport algorithm using the CPN tool with three constituting processes. </span></p> <p><strong><span>Message Sequence Chart of <span>&nbsp;</span>Parametric Model With 3 Processes via Automatic Simulation Run_SuppInfo.mp4</span></strong><span>: This video shows the automatic generation of a message sequence chart of the proposed parametric colored Petri net model of the Chandy-Lamport distributed global snapshot algorithm using the CPN tool version 4.0.0. The number of constituting processes is parametric in the model and was set to three. The video was generated using the authors' updated CPN tool extension server. The automatic simulation run of the model has been used to create this video. The CPN tool randomly selects the enabled transition at each step in an automatic simulation run.</span></p> <p><strong><span>Message Sequence Chart of Parametric Model With 3 Processes via Step-By-Step Simulation Run_SuppInfo.mp4:</span></strong><span> This video shows the automatic generation of a message sequence chart of the proposed parametric colored Petri net model of the Chandy-Lamport algorithm in a step-by-step simulation run with our updated extension server of the CPN tools version 4.0.0. The number of constituting processes is parametric in the model and was set to three. We manually fired our selected enabled transition of the model to create this video. </span></p>

opencc-by-4.0Apr 2024View details →
zenodo36/100

Simulated exome-sequencing data for a family study of lymphoid cancer

<p>This repository contains&nbsp;all the data files for a simulated exome-sequencing study of 150 families ascertained to contain at least four members affected with lymphoid cancer.</p> <p>The simulated data can be found in&nbsp;the&nbsp;files section below. The files are:</p> <ol> <li>SLiM_output.txt - contains the&nbsp;SLiM-simulated, exome-wide, SNV data generated&nbsp;under an American-admixture demographic model,&nbsp; for&nbsp;the&nbsp;American-admixed sub-population only.</li> <li>SLiM_output_chr8&amp;9.txt -&nbsp;contains the&nbsp;SLiM-simulated data above for all source populations as well as the American-admixed sub-population, but&nbsp;only for&nbsp;chromosomes 8 and 9.</li> <li>sample_info.txt - contains pedigree information of all the disease-affected individuals and individuals connecting them along a line of descent, for all 150 ascertained&nbsp;pedigrees.</li> <li>Genotypes.zip -&nbsp; a zipfile that&nbsp;contains 22 text files of&nbsp;genotypes for each chromosome. The genotypes are for simulated&nbsp;single-nucleotide variants on the exome and are&nbsp;in gene-dosage format.&nbsp;</li> <li>SNVmaps.zip -&nbsp; a zipfile that&nbsp;contains 22 text files giving&nbsp;the single-nucleotide&nbsp;variant information for each chromosome.&nbsp;</li> <li>familial_cRV.txt - contains the familial causal rare variants for all 150 ascertained&nbsp;pedigrees.</li> <li>study_peds.txt - contains the 150 pedigrees ascertained to contain four or more relatives affected with lymphoid cancer.</li> <li>PLINKfiles.zip -&nbsp; a zipfile that contains PLINK .fam, .bim and .bed files for all 22 of the chromosomes.</li> </ol> <p>All the scripts used to generate these data&nbsp;can be found in the GitHub repository archived at&nbsp;<a href="https://zenodo.org/record/6505385">https://zenodo.org/record/6505385</a>&nbsp;.</p> <p>We have also&nbsp;uploaded one intermediate .Rdata file,&nbsp;Chromwide.Rdata, to save the user substantial time when running&nbsp;the associated RMarkdown script for the simulation. We recommend loading Chromwide.Rdata into your R work-space rather than generating it from scratch.</p>

openagpl-3.0-or-laterDec 2021View details →
zenodo36/100

Supplementary Data for "The shaky foundations of simulating single-cell RNA sequencing data"

<p>Supplementary Data for &quot;The shaky foundations of simulating single-cell RNA sequencing data&quot;</p> <p>See description.txt, Supplementary Text, Methods and&nbsp;https://github.com/HelenaLC/simulation-comparison for further description of the files available here.</p>

opencc-by-4.0Aug 2022View details →
dryad36/100

Simulated data from: Reference-free assembly of long-read transcriptome sequencing data with RNA-Bloom2

<p><span>Long-read sequencing technologies have improved significantly since their emergence. Their read lengths, potentially spanning entire transcripts, is advantageous for reconstructing transcriptomes. Existing long-read transcriptome assembly methods are primarily reference-based and to date, there is little focus on reference-free transcriptome assembly. We introduce RNA-Bloom2, a reference-free assembly method for long-read transcriptome sequencing data. </span>RNA-Bloom2 is available on GitHub at: <a href="https://github.com/bcgsc/RNA-Bloom">https://github.com/bcgsc/RNA-Bloom</a>.</p> <p><span>We benchmarked the assembly quality and the computational performance of RNA-Bloom2 on simulated data. We prepared two mouse simulated datasets with Trans-NanoSim</span><span> for the cDNA and dRNA sequencing protocols model on experimental ONT data</span><span>. The datasets were simulated </span><span>based on the mouse ENSEMBL annotation for GRCm39.</span><span> To investigate the effect of sequencing depth, we subsampled each dataset to 2, 10, and 18 million reads, resulting in a total of six sets of reads for our benchmarking experiments. Using the simulated data, w</span><span>e showed that the transcriptome assembly quality of RNA-Bloom2 is competitive to those of reference-based methods.</span></p>

opencc-zeroSep 2022View details →
zenodo36/100

Strainy: phasing and assembly of strain haplotypes from long-read metagenome sequencing - Simulated datasets

<p>This repository contains the reads, assemblies, and references required to replicate the <strong>simulated</strong>&nbsp;results presented in the paper: https://doi.org/10.1101/2023.01.31.526521</p>

opencc-by-4.0May 2024View details →
zenodo36/100

Simulated wastewater sequencing data for benchmarking SARS-CoV-2 variant abundance estimation

<p>To evaluate the accuracy of variant abundance&nbsp;predictions from wastewater sequencing, we built a collection of benchmarking datasets that resemble real wastewater samples. For each variant (B.1.1.7, B.1.351, B.1.427, B.1.429, P.1) we created a series of 33 benchmarks by simulating sequencing reads from a variant genome, as well as a collection of background (non-variant of concern/interest) sequences, such that the variant abundance ranges from 0.05% to 100%. Analogously, we created a second series of benchmarks, simulating reads only from the Spike gene of each SARS-CoV-2 genome. We refer to the first set of benchmarks as &quot;whole genome&quot; (WG)&nbsp;and to the second set of benchmarks as &quot;S-only&quot;. We repeated these simulations at different sequencing depths: 100x and 1000x coverage for the whole genome benchmarks, and 100x, 1000x, and 10,000x coverage for the S-only benchmarks.</p>

opencc-by-4.0Aug 2021View details →
zenodo36/100

AliSim: Ultrafast and Realistic Sequence Alignment Simulator for Phylogenetics - Supplementary Data

<p>This supplementary data&nbsp;contains&nbsp;testing scripts, input/output data for validating and benchmarking AliSim.</p>

opencc-by-4.0Sep 2021View details →
ClinicalTrials.gov36/100

Impact of 2 Resuscitation Sequences on Management of Simulated Pediatric Cardiac Arrest

ClinicalTrials.gov study NCT05474170. IPD Sharing: YES. Countries: 1. Publications: 1.

controlledIPD-YESFeb 2026View details →
dryad36/100

Simulated data from: Reference-free assembly of long-read transcriptome sequencing data with RNA-Bloom2

Open the record for dataset details and reuse information.

publicSep 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record