Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

36,856

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

36,856 results for “rna”

Learn how ShareScore rates datasets ↗
zenodo44/100

Dataset for the article "Spatially coherent diffusion of human RNA Pol II depends on transcriptional state rather than chromatin motion" by Roman Barth and Haitham Shaban

<p>The data set comprises all raw microscopy images and DFCC analyses as presented in&nbsp;</p> <p><strong>Spatially coherent diffusion of human RNA Pol II depends on transcriptional state rather than chromatin motion</strong></p> <p>by Roman Barth and Haitham Shaban, published in Nucleus (https://doi.org/10.1080/19491034.2022.2088988)</p> <p>There are two folders for RNAPII and DNA each, one for the raw images and one for the processed DFCC data, supplied as .mat files.</p> <p>Every folder contains three sub-folders containing the data for the conditions: +Serum, -Serum, and +DRB.</p>

opencc-by-4.0Aug 2022View details →
zenodo44/100

Data files: Single-cell RNA profiling of Plasmodium vivax-infected hepatocytes reveals parasite- and host- specific transcriptomic signatures and therapeutic targets

<p>Scripts, preprocessed count matrices, and single-cell data objects generated&nbsp;in&nbsp;<strong>&ldquo;Single-cell RNA profiling of&nbsp;<em>Plasmodium vivax</em><em>-</em>infected hepatocytes reveals parasite- and host- specific transcriptomic signatures&nbsp;and therapeutic targets&rdquo;&nbsp;</strong></p>

opencc-by-4.0Jul 2022View details →
zenodo44/100

direct RNA seq data, triplicate of RSV - strain A2 in Calu-3 cells at 48 hours post infection

<p>Raw fastq data from Calue-3 cells infected with RSV strain A2</p>

opencc-by-4.0Sep 2022View details →
zenodo44/100

RNA-seq dataset for Integrative functional genomic analyses implicate specific molecular pathways and circuits in autism

<p>Data to be used along with <a href="https://github.com/neelroop/asd-development-coexpression-2013">code</a> from 2013 paper that was originally on a site hosted at UCLA, but may no longer be accessible.</p>

opencc-by-4.0May 2024View details →
zenodo44/100

Dataset of "Single-Cell RNA-Seq Reveals Transcriptional Heterogeneity in Latent and Reactivated HIV-infected Cells"

<p><strong>Detailed quantitative analysis of GFP expression in SAHA and TCR-treated cells &amp; Computational analysis of&nbsp;bulk and single-cell RNA-Seq data.</strong></p> <p>&nbsp;</p> <p><em><strong>Detailed quantitative analysis of GFP expression in SAHA and TCR-treated cells.</strong></em></p> <p>Cells were prepared for single cell analysis at the Genome Technology Facility (GTF) of the University of Lausanne. Cells were loaded on Fluidigm C1 IFC plates (5-10 &mu;m), with run ID smart33, smart34 and smart35, corresponding to untreated, SAHA- and TCR-treated conditions respectively. After single cell capture on the Fluidigm C1 IFC plate, each chamber was inspected visually by microscopy and pictures were captured with a Zeiss Axiovert 200 M fluorescence microscope equipped with a Roper Scientific CoolSnap HQ camera using a Plan-Neofluar 10X lens (smart34 run) or 20X lens (for smart35 run). For each capture chamber, pictures in bright field and FITC channel were taken with the MetaMorph 6.3 software. Picture analysis was then performed using ImageJ 1.50b software (open access software: website). Brightness and contrast were adjusted for qualitative assessment of the pictures.</p> <p><em><strong>Computational analysis of&nbsp;bulk and single-cell RNA-Seq data.</strong></em></p> <p>Upon bulk or single cell isolation, RNA extraction and library preparation was performed according to Illumina protocols. Bulk and single-cell RNA-Seq data analysis are detailed here.</p> <p>&nbsp;</p> <p>Linked to the paper published in Cell Reports (doi:10.1016/j.celrep.2018.03.102):&nbsp;</p> <p><strong>Single-Cell RNA-Seq Reveals Transcriptional Heterogeneity&nbsp;in Latent and Reactivated HIV-infected Cells</strong></p> <p>Despite effective treatment, HIV can persist in latent reservoirs, which represent a major obstacle towards HIV eradication. Targeting and reactivating latent cells is challenging due to the heterogeneous nature of HIV infected cells. Here, we used a primary model of HIV latency and single-cell RNA sequencing to characterize transcriptional heterogeneity during HIV latency and reactivation. Our analysis identified transcriptional programs leading to successful reactivation of HIV expression.</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Feb 2018View details →
zenodo44/100

Sample datasets for Galaxy RNA-seq tutorial

<p>This is downsampled dataset from&nbsp;http://dx.doi.org/10.1038/nprot.2016.095. It was prepared as follows:</p> <ol> <li>Mapping data from&nbsp;&nbsp;http://dx.doi.org/10.1038/nprot.2016.095 against hg38 using HISAT2</li> <li>Restricting resulting BAM datasets to chrX:70,000,000-80,000,000</li> <li>Extracting reads using picard SamToFastq tool</li> </ol> <p>File rnaseq_sex.tab contains mapping between accession numbers and sex of the sequences individuals.&nbsp;</p>

opencc-by-4.0Jun 2018View details →
zenodo44/100

Single-cell RNA-seq of breast cancer infiltrating T cells (case 1)

<p>Single cell suspensions were generated from two individual TNBC primary&nbsp;tumor samples (this&nbsp;entry contains case 2) and the viable cells were FACS sorted for CD3<sup>+</sup> T cells.&nbsp;Sorted cells were then counted and assessed for viability. Single cell library preparation was carried out as per the 10X Genomics Chromium Single cell protocol for the v2 reagent kit (10X Genomics, Pleasanton, CA, USA). Cell suspensions were loaded onto a Chromium Single Cell Chip along with the reverse transcription (RT) mastermix and single cell 3&rsquo; gel beads. Following generation of single cell gel bead-in-emulsions (GEMs), reverse transcription was performed using a C1000 Touch Thermal Cycler with a Deep Well Reaction Module (Bio-Rad Laboratories, Hercules, CA, USA). Amplified cDNA was purified using SPRIselect beads (Beckman Coulter, Lane Cove, NSW, Australia) and sheared to approximately 200bp with a Covaris S2 instrument (Covaris, Woburn, MA, USA) using the manufacturer&rsquo;s recommended parameters. Sequencing libraries were generated with unique sample indices (SI) for each sample. Libraries were sequenced on an Illumina HiSeq 2500 High Output Mode using V4 clustering and sequencing chemistry.</p> <p>This dataset&nbsp;contains&nbsp;the raw .bcl files.</p>

opencc-by-4.0Jun 2018View details →
zenodo44/100

Single-cell RNA-seq of breast cancer infiltrating T cells (case 2)

<p>Single cell suspensions were generated from two individual TNBC primary&nbsp;tumor samples (this&nbsp;entry contains case 1) and the viable cells were FACS sorted for CD3<sup>+</sup> T cells.&nbsp;Sorted cells were then counted and assessed for viability. Single cell library preparation was carried out as per the 10X Genomics Chromium Single cell protocol for the v2 reagent kit (10X Genomics, Pleasanton, CA, USA). Cell suspensions were loaded onto a Chromium Single Cell Chip along with the reverse transcription (RT) mastermix and single cell 3&rsquo; gel beads (this sample was divided into two channels). Following generation of single cell gel bead-in-emulsions (GEMs), reverse transcription was performed using a C1000 Touch Thermal Cycler with a Deep Well Reaction Module (Bio-Rad Laboratories, Hercules, CA, USA). Amplified cDNA was purified using SPRIselect beads (Beckman Coulter, Lane Cove, NSW, Australia) and sheared to approximately 200bp with a Covaris S2 instrument (Covaris, Woburn, MA, USA) using the manufacturer&rsquo;s recommended parameters. Sequencing libraries were generated with unique sample indices (SI) for each sample. Libraries were sequenced on an Illumina HiSeq 2500 High Output Mode using V4 clustering and sequencing chemistry.</p> <p>This dataset&nbsp;contains&nbsp;the raw .bcl files.</p>

opencc-by-4.0Jun 2018View details →
zenodo44/100

CellSIUS provides sensitive and specific detection of rare cell populations from complex single cell RNA-seq data: Codes and processed data

<p>Codes and processed data to reproduce the analysis discussed in:&nbsp;</p> <p>Wegmann <em>et Al.</em>,<strong> CellSIUS provides sensitive and specific detection of rare cell<br> populations from complex single cell RNA-seq data</strong>, Genome Biology 2019 (Accepted)<br> &nbsp;</p>

openapache2.0Jun 2019View details →
zenodo44/100

BaRTv1.0: an improved barley reference transcript dataset to determine accurate changes in the barley transcriptome using RNA-seq

<p>Background<br> Time consuming computational assembly and quantification of gene expression and splicing analysis from RNA-seq data vary considerably. Recent fast non-alignment tools such as Kallisto and Salmon overcome these problems, but these tools require a high quality, comprehensive reference transcripts dataset (RTD), which are rarely available in plants.</p> <p>Results<br> A high-quality, non-redundant barley gene RTD and database (Barley Reference Transcripts &ndash; BaRTv1.0) has been generated. BaRTv1.0, was constructed from a range of tissues, cultivars and abiotic treatments and transcripts assembled and aligned to the barley cv. Morex reference genome (Mascher et al., 2017). Full-length cDNAs from the barley variety Haruna nijo (Matsumoto et al., 2011) determined transcript coverage, and high-resolution RT-PCR validated alternatively spliced (AS) transcripts of 86 genes in five different organs and tissue. These methods were used as benchmarks to select an optimal barley RTD. BaRTv1.0-Quantification of Alternatively Spliced Isoforms (QUASI) was also made to overcome inaccurate quantification due to variation in 5&rsquo; and 3&rsquo; UTR ends of transcripts. BaRTv1.0-QUASI was used for accurate transcript quantification of RNA-seq data of five barley organs/tissues. This analysis identified 20,972 significant differentially expressed genes, 2,791 differentially alternatively spliced genes and 2,768 transcripts with differential transcript usage.</p> <p>Conclusion<br> A high confidence barley reference transcript dataset consisting of 60,444 genes with 177,240 transcripts has been generated. Compared to current barley transcripts, BaRTv1.0 transcripts are generally longer, have less fragmentation and improved gene models that are well supported by splice junction reads. Precise transcript quantification using BaRTv1.0 allows routine analysis of gene expression and AS.</p>

opencc-by-4.0Dec 2018View details →
zenodo44/100

Additional data for publication: Simple protocol for combined extraction of exocrine secretions and RNA in small arthropods.

<p>Additional data and results are given in this repository. It contains the trimmed reads (fastp; raw reads also on SRA accession numbers SRR29851544-SRR29851549, Bioproject PRJNA1136254), full busco reports for individual transcriptomes, assembly of all six RNAseqs together (transcriptome as base for differential expression analysis) and results of salmon.</p>

opencc-by-4.0Jul 2024View details →
zenodo44/100

Cameroonian blackflies (Diptera: Simuliidae) harbour a plethora of (RNA) viruses

<p>Fasta file (blackfly_viral_genomes.fasta) with identified virus sequences in Cameroonian blackflies and a corresponding table with their taxonomy when possible (taxonomy_data.tsv). Viruses were identified with geNomad v1.7.0 and Diamond blastx v2.0.11 with the NCBI nr database (accessed 23-03-2023).</p> <p>We also included a fata file with all assembled sequences across all blackfly samples (blackfly.all.fasta)</p>

opencc-by-4.0Sep 2024View details →
zenodo44/100

An example RNA-seq count table using nf-core/rnaseq

<p>An example RNA-seq count table generated from data deposited under GSE40419 (lung adenocarcinoma study) and nf-core/rnaseq for testing downstream RNA-seq analyses.</p>

opencc-by-4.0Oct 2024View details →
zenodo44/100

ZIRFs: zero-inflated random forests for estimating gene regulatory networks from single cell RNA-seq data (assessment of predictive accuracy and VIM stability)

<p>We developed a zero-inflated random forests (ZIRFs) algorithm to produce a metric of connection strength&nbsp;between regulator genes and target genes. This file contains SCENIC results for the aorta and diaphragm tissue data sets from the Tabula Muris Consortium results. SCENIC is a genetic regulatory network analysis published by Aibar et al. (2017). The purpose of the data sets and R source code are described by README files in each directory.</p>

opencc-by-3.0-usJul 2021View details →
zenodo44/100

xPore: Identification of differential RNA modifications from nanopore direct RNA sequencing - SGNEx data

<p>xPore is&nbsp;a Python package for identification and quantification of differential RNA modifications from direct RNA sequencing.</p> <p>The detailed usage&nbsp;is&nbsp;documented at&nbsp;<a href="https://xpore.readthedocs.io/en/latest/">https://xpore.readthedocs.io/en/latest</a>, while all&nbsp;scripts and source code are available at&nbsp;<a href="https://github.com/GoekeLab/xpore">https://github.com/GoekeLab/xpore</a>.</p> <p>All the preprocessed&nbsp;datasets&nbsp;used in the paper are provided here.&nbsp;</p> <p>Please cite our paper below&nbsp;when using these data.<br> Ploy N. Pratanwanich et al. &quot;Detection of differential RNA modifications from direct RNA sequencing of human cell lines.&quot; bioRxiv (2020).</p>

opencc-by-4.0Mar 2021View details →
zenodo44/100

Mammary single-cell RNA-seq analysis and prostate cancer survival as a function of H2AFJ expression for the paper entitled: The histone variant H2A.J is enriched in luminal epithelial cells

<p>H2A.J is a poorly studied mammalian-specific variant of histone H2A. We used immunohistochemistry to study its localization in various human and mouse tissues. H2A.J showed cell-type specific expression with a striking enrichment in luminal epithelial cells of multiple glands including those of breast, prostate, pancreas, thyroid, stomach, and salivary glands. H2A.J was also highly expressed in many carcinoma cell lines and in particular, those derived from luminal breast and prostate cancer. H2A.J thus appears to be a novel marker for luminal epithelial cancers. Knocking-out the H2AFJ gene in T47D luminal breast cancer cells reduced the expression of several estrogen-responsive genes which may explain its putative tumorigenic role in luminal-B breast cancer.</p>

opencc-by-4.0Sep 2021View details →
zenodo44/100

THOR RNA-sequencing results

<p>Summarized results of RNA-sequencing performed within THOR (targeting smooth muscle cells in atherosclerosis). THOR is a collaborative project of Aarhus University and Novo Nordisk A/S as a part&nbsp;of the Open Discovery Innovation Network (ODIN) initiative.</p> <p>See detailed data description&nbsp;in the file &quot;DESCRIPTION.md&quot;.</p>

opencc-zeroNov 2022View details →
zenodo44/100

D.melanogaster Genelab OSD Normalized RNA Seq Matrix

<p><em>D.melanogaster&nbsp;</em>normalized counts RNA seq data matrix developed from NASA Genelab&#39;s open science data repository. Created using R.</p>

opencc-by-4.0Dec 2022View details →
zenodo44/100

M.musculus Genelab OSD Unnormalized RNA seq Matrix

<p><em>M.musculus&nbsp;</em>unnormalized counts&nbsp;RNA seq data matrix from NASA Genelab&#39;s&nbsp;open science data repository. Created using R.</p>

opencc-by-4.0Dec 2022View details →
zenodo44/100

Whole-genome capture and sequencing of Mycobacterium tuberculosis directly from clinical samples - Design of RNA oligonucleotide baits for Agilent Technologies' SureSelect target enrichment

<p>This dataset comprises the sequence of <strong>44&nbsp;278&nbsp;RNA oligonucleotide &quot;baits&quot; (120 bp each) </strong>designed to perform&nbsp;<strong>whole-genome capture and sequencing of <em>Mycobacterium tuberculosis</em>&nbsp;directly from clinical samples</strong>&nbsp;(DNA)&nbsp;using Agilent Technologies&rsquo; SureSelect target enrichment system following the Illumina paired-end multiplexed sequencing library protocol.&nbsp;</p> <p>RNA oligonucleotide &ldquo;baits&rdquo; were designed to span the &sim;4.5 Mb of the <em>M. tuberculosis</em> genome. In brief, the reference genome sequence of the MTBC H37Rv strain (Genbank #AL123456) was <em>in silico</em> fragmented into 120 bp sequences twice, to ensure an overlap of 60 bp between sequences. Due to their rich GC content, which could interfere with DNA capture, all MTBC genes of the PE, PPE and PE-PGRS family were also independently fragmented into 120 bp sequences, in order to increase capture sensitivity. All resulting sequences were BLASTn searched against the Human Genomic + Transcript database to excluded homologous sequences to the human genome. Overall, a total of 42,278 RNA probes were generated and this custom bait library was then uploaded to the SureDesign software (https://earray.chem.agilent.com/suredesign) and synthesized by Agilent Technologies. During synthesis, the 2198 sequences complementary to the PE, PPE and PE-PGRS family were unbalanced 8:1 to potentiate capture.</p> <p>More details can be found in the following publication:</p> <p>- Macedo, R., Isidro, J., Ferreira, R., Pinto, M., Borges, V., Duarte, S., Vieira, L., &amp; Gomes, J. P. (2023). Molecular Capture of&nbsp;<em>Mycobacterium tuberculosis</em>&nbsp;Genomes Directly from Clinical Samples: A Potential Backup Approach for Epidemiological and Drug Susceptibility Inferences.&nbsp;<em>International journal of molecular sciences</em>,&nbsp;<em>24</em>(3), 2912. https://doi.org/10.3390/ijms24032912</p>

opencc-by-4.0Jan 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record