Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

7

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

7 results for “RNA-seq data simulation”

Learn how ShareScore rates datasets ↗
zenodo44/100

Simulated RNA-seq data

<p>Simulated RNA-seq data shows that histograms from p value sets with around one hundred&nbsp;&nbsp;true effects out of 20,000 features can be classified as &#39;uniform&#39;.&nbsp;RNA-seq data was simulated with polyester R package <a href="https://doi.org/10.1093/bioinformatics/btv272">(Frazee, 2015)</a> on 20,000 transcripts from human transcriptome&nbsp;using grid of 3, 6, and 10 replicates and 100, 200, 400, and 800 effects for two groups.&nbsp;Fold changes were set to 0.5 and 2.&nbsp;Differential expression was assessed using DESeq2 R package <a href="https://doi.org/10.1186/s13059-014-0550-8">(Love, 2014)</a> using default settings&nbsp;and group 1 versus group 2 contrast.&nbsp;Effects denotes in facet labels the number of true effects and N denotes number of replicates.&nbsp;Red line denotes QC threshold used for dividing p histograms into discrete classes.&nbsp;Workflow and code used to run this simulation is available on <a href="https://github.com/rstats-tartu/simulate-rnaseq">rstats-tartu/simulate-rnaseq</a>.</p> <p>&nbsp;</p> <p>Files</p> <ul> <li>de_simulation_results.csv -- merged and processed DE analysis results of simulated data.</li> <li>simulate-reads-2021-01-25.tar.gz -- raw DE analysis results&nbsp;on 20,000 transcripts from human transcriptome&nbsp;using grid of 3, 6, and 10 replicates and 100, 200, 400, and 800 effects for two groups.&nbsp;Fold changes were set to 0.5, 1, and 2.&nbsp;Differential expression was assessed using DESeq2 with default settings.</li> <li>simulate-rnaseq.tar.gz -- snakemake workflow and input fasta file&nbsp;to simulate RNA-seq data with polyester and analyse results with DESeq2. Adjust settings in config.yaml to customise simulation. Includes software to run workflow on Linux, given that <a href="https://repo.anaconda.com/miniconda/Miniconda3-latest-Linux-x86_64.sh">Conda</a> and <a href="https://snakemake.readthedocs.io/en/stable/index.html">snakemake</a> are installed.</li> </ul> <p>The simulate-rnaseq.tar.gz&nbsp;archive can be re-executed on a vanilla machine that only has Conda and Snakemake installed via:</p> <pre><code class="language-bash">tar -xf simulate-rnaseq.tar.gz snakemake --use-conda -n</code></pre> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Jan 2021View details →
zenodo36/100

Isosceles paper simulated ovarian cell line ONT data (bulk RNA-Seq)

<div> <p>Simulated ovarian cell line ONT data (bulk RNA-Seq) for the Isosceles paper - more details can be found in the&nbsp;<a href="https://github.com/Genentech/Isosceles_Paper" target="_blank" rel="noopener">Isosceles_Paper</a> repository.</p> </div>

opencc-by-4.0Mar 2024View details →
zenodo36/100

Simulated RNA-seq data for differential splicing analysis with covariates

<p>The repository includes alignments of simulated RNA-seq data for evaluating differential splicing detection with covariates. Starting from an empirical transcript expression matrix trained on an RNA-seq data set from lung fibroblasts (GenBank A# SRR493366) and using GENCODE v.41 as reference, 11.5 million 100 bp long paired-end reads were generated per sample, from 2,000 genes with two or more expressed isoforms. RNA-seq data was simulated for one &lsquo;condition&rsquo;, with values &lsquo;control&rsquo;, &lsquo;disease&rsquo; and &lsquo;stage2&rsquo;, with one covariate, &lsquo;biological sex&rsquo;, with values &lsquo;M&rsquo; and &lsquo;F&rsquo;.&nbsp; 5 samples each were simulated for each (condition x sex) category. Changes were simulated in the expression (DE) and/or the splicing ratio (DS) of genes as follows. Changes in expression (DE) were simulated by either halving or doubling the expression level of the gene. Changes in splicing ratios (DS) were simulated by swapping the expression levels of the gene&rsquo;s top two transcript isoforms. All RNA-seq data was mapped to the hg38 genome with the spliced alignment tool STAR v2.7.10a.</p> <p>&nbsp;<em><u>Pairwise comparison alignment set</u></em>: Differences due to &lsquo;condition&rsquo; between two states, &lsquo;control&rsquo; and &lsquo;disease&rsquo;, were simulated at 600 genes, including 200 DE, 200 DS and 200 DE+DS genes. Differences in &lsquo;biological sex&rsquo; (covariate) were represented as changes in 300 genes, including 100 from each of the DS, DE and DE+DS categories. Hence, the target gene set for differential splicing ratio (DSR)<em> pairwise comparisons </em>consists of the pooled 200 DS and 200 DS+DE genes differentially spliced between the &lsquo;control&rsquo; and &lsquo;disease&rsquo; states, while for differential splicing abundance (DSA)<em> pairwise comparisons </em>the target gene set is the set of 600 modified genes, 200 in each of the DS, DE and DS+DE categories.</p> <p>&nbsp;<em><u>Multiway (3-way) comparison alignment set:</u></em> Additional changes between &lsquo;disease&rsquo; and &lsquo;stage2&rsquo; were made to 100 of the previously modified genes, as well as to a set of 200 additional genes not encountered previously, for each of the categories DE, DS and DE+DS. Therefore, for&nbsp;<em>DSR three-way comparisons</em>, the target gene set represents the 800 genes simulated as being DS or DE+DS between any of the &lsquo;control&rsquo;, &lsquo;disease&rsquo; and &rsquo;stage2&rsquo; categories, while for the <em>multi-way DSA comparisons</em> the target is the full set of 1,200 genes (400 DE, 400 DS and 400 DE+DS) simulated to have changed between any of the 'control', 'disease&rsquo; and &lsquo;stage2&rsquo; states.</p> <p>&nbsp;<em><u>Further details:</u></em> See the &lsquo;key&rsquo; directories in each package for the gene lists.</p>

opencc-by-4.0May 2024View details →
zenodo36/100

ESPRESSO: Robust discovery and quantification of transcript isoforms from error-prone long-read RNA-seq data (repository for simulated ONT RNA-seq data)

<p>Simulated ONT direct RNA and 1D cDNA sequencing data of varying sequencing depths (0.5 million, 1 million, 3 million, and 5 million simulated reads) used for benchmark evaluations of transcript discovery and quantification in our paper &quot;ESPRESSO: Robust discovery and quantification of transcript isoforms from error-prone long-read RNA-seq data&quot;. All details can be found in the <strong>Materials and Methods</strong> section of the paper.&nbsp;</p> <p><em>HEK293T_DirectRNA.transcriptome_quantification.tsv</em> and&nbsp;<em>HEK293T_DirectRNA.transcriptome_quantification.tsv </em>are tab-separated files containing estimated raw read counts and normalized abundance values (in TPM) of transcripts annotated in GENCODE v34lift37. Transcript quantification was done using NanoSim (version 3.1.0).&nbsp;</p> <p><em>HEK293T_DirectRNA.NanoSim_500k.fastq.gz</em>,<em>&nbsp;</em><em>HEK293T_DirectRNA.NanoSim_1M.fastq.gz</em>,&nbsp;<em>HEK293T_DirectRNA.NanoSim_3M.fastq.gz</em>, and<em>&nbsp;HEK293T_DirectRNA.NanoSim_5M.fastq.gz&nbsp;</em>are gzip compressed FASTQ files containing 0.5 million, 1 million, 3 million, and 5 million simulated ONT direct RNA sequencing&nbsp;reads respectively.&nbsp;</p> <p><em>HEK293T_1DcDNA.NanoSim_500k.fastq.gz</em>,<em>&nbsp;HEK293T_1DcDNA.NanoSim_1M.fastq.gz</em>,&nbsp;<em>HEK293T_1DcDNA.NanoSim_3M.fastq.gz</em>, and<em>&nbsp;HEK293T_1DcDNA.NanoSim_5M.fastq.gz&nbsp;</em>are gzip compressed FASTQ files containing 0.5 million, 1 million, 3 million, and 5 million simulated ONT 1D cDNA sequencing&nbsp;reads respectively.&nbsp;</p>

opencc-by-4.0Oct 2022View details →
zenodo32/100

SQUID: Transcriptomic Structural Variation Detection from RNA-seq -- simulation data part 3

<p>Simulation data part 3&nbsp;for SQUID software.</p>

openbsd-3-clauseMar 2018View details →
zenodo32/100

SQUID: Transcriptomic Structural Variation Detection from RNA-seq -- simulation data part 2

<p>Simulation data part 2&nbsp;for SQUID software.</p>

openbsd-3-clauseMar 2018View details →
zenodo32/100

SQUID: Transcriptomic Structural Variation Detection from RNA-seq -- simulation data part 1

<p>Simulation data part 1 for SQUID software.</p>

openbsd-3-clauseMar 2018View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record