Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
6,609
datasets available to search
ShareScore release 0.7.1
Dataset results
6,609 results for “RNA sequencing”
xPore: Identification of differential RNA modifications from nanopore direct RNA sequencing
<p>xPore is a Python package for identification and quantification of differential RNA modifications from direct RNA sequencing.</p> <p>The detailed usage is documented at <a href="https://xpore.readthedocs.io/en/latest/">https://xpore.readthedocs.io/en/latest</a>, while all scripts and source code are available at <a href="https://github.com/GoekeLab/xpore">https://github.com/GoekeLab/xpore</a>.</p> <p>All the preprocessed datasets used in the paper are provided here. </p> <p>Please cite our paper below when using these data.<br> Ploy N. Pratanwanich et al. "Detection of differential RNA modifications from direct RNA sequencing of human cell lines." bioRxiv (2020).</p>
RNA sequencing data for bleomycin exposed THP-1 macrophages
<p>This dataset contains normalized counts matrices, from dds_deseq objects, from DeSeq2 analysis of RNA sequencing data, from THP-1 macrophages exposed to multiple doses of bleomycin in the range of 0-100µg/ml for 24H, 48H or 72H.</p>
RNA sequencing dataset for prediction of liver hepatocellular carcinoma using SIMON analysis
<p>The LIHC dataset was used for data mining and for the generation of machine learning model for the detection of liver hepatocellular carcinoma cells (LIHC) using the SIMON platform as described in the "SIMON: open-source knowledge discovery platform" publication (<a href="https://doi.org/10.1101/2020.08.16.252767">https://doi.org/10.1101/2020.08.16.252767</a>). The LIHC dataset was obtained from the <em>GSEABenchmarkeR</em> package ( <a href="https://doi.org/10.1093/bib/bbz158">https://doi.org/10.1093/bib/bbz158</a>) and it contains RNA expression data from 374 liver hepatocellular carcinoma (LIHC) cells and 50 adjacent normal cells.</p>
VirHunter: a deep learning-based method for detection of novel RNA viruses in plant sequencing data
<p>This storage contains 2 archives: toy datasets to test the training of the VirHunter and weights of the fully trained VirHunter models for 3 host species (peach, grapevine, sugar beet) and for fragment sizes 500 and 1000. .</p> <p>The toy dataset consists of 3 archived files: 'viruses.fasta', 'host.fasta', 'bacteria.fasta'.</p> <p>'viruses.fasta' contains 10000 randomly selected plant viruses from the virus dataset described in the paper.</p> <p>'host.fasta' consists of peach chromosome 2.</p> <p>'bacteria.fasta' consists of 10 bacterial genomes selected randomly: GCF_000284415, GCF_000590555, GCF_001548055, GCF_002795265, GCF_003330825, GCF_003957805, GCF_005845345, GCF_009176625, GCF_010748935, GCF_014681765</p> <p> </p>
xPore: Identification of differential RNA modifications from nanopore direct RNA sequencing - SGNEx data
<p>xPore is a Python package for identification and quantification of differential RNA modifications from direct RNA sequencing.</p> <p>The detailed usage is documented at <a href="https://xpore.readthedocs.io/en/latest/">https://xpore.readthedocs.io/en/latest</a>, while all scripts and source code are available at <a href="https://github.com/GoekeLab/xpore">https://github.com/GoekeLab/xpore</a>.</p> <p>All the preprocessed datasets used in the paper are provided here. </p> <p>Please cite our paper below when using these data.<br> Ploy N. Pratanwanich et al. "Detection of differential RNA modifications from direct RNA sequencing of human cell lines." bioRxiv (2020).</p>
THOR RNA-sequencing results
<p>Summarized results of RNA-sequencing performed within THOR (targeting smooth muscle cells in atherosclerosis). THOR is a collaborative project of Aarhus University and Novo Nordisk A/S as a part of the Open Discovery Innovation Network (ODIN) initiative.</p> <p>See detailed data description in the file "DESCRIPTION.md".</p>
Whole-genome capture and sequencing of Mycobacterium tuberculosis directly from clinical samples - Design of RNA oligonucleotide baits for Agilent Technologies' SureSelect target enrichment
<p>This dataset comprises the sequence of <strong>44 278 RNA oligonucleotide "baits" (120 bp each) </strong>designed to perform <strong>whole-genome capture and sequencing of <em>Mycobacterium tuberculosis</em> directly from clinical samples</strong> (DNA) using Agilent Technologies’ SureSelect target enrichment system following the Illumina paired-end multiplexed sequencing library protocol. </p> <p>RNA oligonucleotide “baits” were designed to span the ∼4.5 Mb of the <em>M. tuberculosis</em> genome. In brief, the reference genome sequence of the MTBC H37Rv strain (Genbank #AL123456) was <em>in silico</em> fragmented into 120 bp sequences twice, to ensure an overlap of 60 bp between sequences. Due to their rich GC content, which could interfere with DNA capture, all MTBC genes of the PE, PPE and PE-PGRS family were also independently fragmented into 120 bp sequences, in order to increase capture sensitivity. All resulting sequences were BLASTn searched against the Human Genomic + Transcript database to excluded homologous sequences to the human genome. Overall, a total of 42,278 RNA probes were generated and this custom bait library was then uploaded to the SureDesign software (https://earray.chem.agilent.com/suredesign) and synthesized by Agilent Technologies. During synthesis, the 2198 sequences complementary to the PE, PPE and PE-PGRS family were unbalanced 8:1 to potentiate capture.</p> <p>More details can be found in the following publication:</p> <p>- Macedo, R., Isidro, J., Ferreira, R., Pinto, M., Borges, V., Duarte, S., Vieira, L., & Gomes, J. P. (2023). Molecular Capture of <em>Mycobacterium tuberculosis</em> Genomes Directly from Clinical Samples: A Potential Backup Approach for Epidemiological and Drug Susceptibility Inferences. <em>International journal of molecular sciences</em>, <em>24</em>(3), 2912. https://doi.org/10.3390/ijms24032912</p>
Joint embedding of vertebrate brain single-cell RNA-Seq using sequence or structure
<p>Embeddings of single-cell RNA-Seq data from three adult vertebrate brain datasets into Orthogroup feature space or Structural cluster feature space. Orthogroups were generated using OrthoFinder v5.5.0; Structural clusters were assigned by using FoldSeek to cluster AlphaFold-v4 structural predictions.<br> <br> The three datasets used as the basis for these embeddings were:</p> <ul> <li>sample <a href="https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSM3768152">"Brain8"</a> from the <a href="https://www.frontiersin.org/articles/10.3389/fcell.2021.743421/full">Jiang et al. 2021</a> zebrafish cell atlas (files beginning with GSM3768152)</li> <li>sample <a href="https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSM2906405">"Brain1"</a> from the <a href="https://www.sciencedirect.com/science/article/pii/S0092867418301168#sec4">Han et al. 2018</a> mouse cell atlas (files beginning with GSM2906405)</li> <li>sample <a href="https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSM6214268">"Xenopus_brain_COL65"</a> from the <a href="https://www.nature.com/articles/s41467-022-31949-2">Liao et al. 2022</a> Xenopus laevis adult cell atlas (files beginning with GSM6214268)</li> </ul> <p>For each dataset, we also generated a standardized cell type annotation file based on the author's originally provided cell type annotation data. The first column is the cell barcode for that species and the second column is the original study's cell type annotation for that cell.</p> <p>For the Xenopus brain data, we removed around ~18k cells that were not annotated in the original data to simplify data analyses - these are reflected in the files with the "subsampled" suffix. Subsampled versions of the data are also available for the joint embedding space (prefixed with "DrerMmusXlae").</p> <p>For the final datasets used in our analyses, we also provide features x cell matrices as .h5ad files for smaller file sizes and faster loading using Scanpy. </p> <p>For visualizing our UMAP plots of our top200 embedding space, we provide ".tsv" files with a variety of metrics and the x and y positions of each cell in the UMAP. See "DrerMmusXlae_adultbrain_FoldSeek_plotlydata.tsv" and "DrerMmusXlae_adultbrain_OrthoFinder_plotlydata.tsv"</p> <p>These data are part of the Arcadia Science Pub titled <a href="https://doi.org/10.57844/arcadia-vw5e-2670">"Comparing gene expression across species based on protein structure instead of sequence"</a>.</p>
Reads-per-UMI tables across single-cell RNA sequencing protocols
<p>Data analyzed in <a href="https://www.biorxiv.org/content/10.1101/2023.08.02.551637v1">Lause, Ziegenhain et al. (2023)</a>.</p> <p>Code to obtain these tables from public data sources is available on <a href="https://github.com/berenslab/read-normalization">github</a>.</p> <p> </p> <p>Each row in the table is a UMI-tag detected in a certain cell (column RG) attached to a molecule from a specific gene (column GE) with a certain barcode (column UB). Column N gives the number of times the UMI was detected for that gene and cell.</p> <p>Data sources and protocols are given with the respective file names below.</p> <p><strong>Johnsson2022_Smartseq3_PE.hd1.txt.gz</strong>: Mouse fibroblasts profiled with <strong>Smart-seq3</strong> paired-end; accession E-MTAB-10148, sample plate2,<br> <a href="https://doi.org/10.1038/s41588-022-01014-1">Paper</a><br> <br> <strong>Hagemann-Jensen2020_Smartseq3_SE.hd1.txt.gz: </strong>Mouse fibroblasts profiled with <strong>Smart-seq3</strong> single-end; accession E-MTAB-8735, sample Smartseq3.Fibroblasts.smFISH<br> <a href="https://doi.org/10.1038/s41587-020-0497-0">Paper</a><br> <br> <strong>Hagemann-Jensen2022_Smartseq3xpress.hd1.txt.gz: </strong>HEK293 cells profiled with <strong>Smart-seq3Xpress</strong>; accession E-MTAB-11467.<br> <a href="https://www.biorxiv.org/content/10.1101/2021.07.10.451889v1">Paper</a><br> <br> <strong>Ziegenhain2017.hd1.txt.gz: </strong>Mouse embryonic stem cells profiled by <strong>CEL-seq2, Drop-seq, MARS-seq, </strong>and<strong> SCRB-seq</strong>; GEO accession GSE75790<br> <a href="https://doi.org/10.1016/j.molcel.2017.01.023">Paper</a></p>
Single-cell RNA sequencing identifies shared differentiation paths of mouse thymic innate T cells
<p>scRNA sequencing datasets used in the paper titled 'Single-cell RNA sequencing identifies shared differentiation paths of mouse thymic innate T cells' published in Nature Communications<br> <br> https://www.nature.com/articles/s41467-020-18155-8</p>
Full-length, single-cell RNA-sequencing of human bone marrow subpopulations reveals hidden complexity
<p><a href="http://www.biorxiv.org/content/10.1101/2021.07.28.454226v2">Full-length, single-cell RNA-sequencing of human bone marrow subpopulations reveals hidden complexity</a></p> <p>Bone marrow progenitor cell differentiation has frequently been used as a model for studying cellular plasticity and cell-fate decisions. Recent analysis at the level of single-cells has expanded knowledge of the transcriptional landscape of human hematopoietic cell lineages. Using single-molecule real-time (SMRT) full-length RNA sequencing, we have previously shown that human bone marrow lineage-negative (Lin-neg) cell populations contain a surprisingly diverse set of mRNA isoforms. Here, we report from single cell, full-length RNA sequencing that this diversity is also reflected at the single-cell level. From fresh human bone marrow unselected and lineage-negative progenitor cells were isolated by droplet-based single-cell selection (10xGenomics). The single cell-derived mRNAs were analyzed by full-length SMRT and short-read sequencing. In both samples we detected an average of 8000 different genes using short-read sequencing. Differential expression analysis arranged the single-cells of the total bone marrow into only four clusters whereas the Lin-neg population was much more diverse with nine clusters. mRNA isoform analysis of the single-cell populations using full-length sequencing revealed that Lin-neg cells contain on average 24% more novel splice variants than the total bone marrow cells. Interestingly, among the most frequent genes expressing novel isoforms were members of the spliceosome, e.g. HNRNPs, DEAD box helicases and SRSFs. Mapping the isoforms from all genes to the cell type clusters revealed that total bone marrow cells express novel isoforms only in a small subset of clusters. On the other hand, lineage-negative progenitor cells expressing novel isoforms were present in nearly all subpopulations. In conclusion, on a single-cell level lineage-negative cells express a higher diversity of genes and more alternatively spliced novel isoforms suggesting that cells in this subpopulation are poised for different fates. </p> <p> </p>
Deep sequencing datasets from: RNA-catalyzed evolution of catalytic RNA
<p>This dataset includes raw and processed sequencing data from evolving RNA populations described in Nikos Papastavrou, David P Horning, Gerald F Joyce. "RNA-Catalyzed Evolution of Catalytic RNA" (submitted). Briefly, directed evolution of a hammerhead ribozyme sequence was carried out over eight rounds of three steps each: 1) templated synthesis of the reverse-complement hammerhead RNA by a polymerase ribozyme; 2) templated synthesis of a new copy of the hammerhead RNA from the reverse-complement by a polymerase ribozyme; 3) selective recovery of hammerhead RNA that cleaved an attached RNA substrate. Cleaved RNA was reverse transcribed, PCR amplified, and archived for sequencing, while a portion was in vitro transcribed with T7 RNA polymerase to initiate the next round of evolution. Two distinct branches of evolution were carried out for 8 rounds, using the '52-2' or '71-89' polymerase ribozymes to replicate RNA, respectively. Sequenced RNA populations were analyzed to determine polymerase ribozyme fidelity and study the evolution of hammerhead sequences replicated by polymerases with low or high RNA copying fidelities. The dataset includes raw sequence files, processed tables of mutations by position along the sequence for each polymerase, processed tables of the sequence frequency distribution from each round in the evolving populations, and spreadsheets containing final processed data used directly in manuscript figures and tables.</p>
RNA sequencing of macrophages co-cultured with MSCs and RNA sequencing of alveolar macrophages from mice with lung injury treated with MSCs
<p>RNA sequencing of macrophages co-cultured with MSCS Table 5</p> <p>RNA sequencing of alveolar macrophages from mice with lung injury treated with MSCS Table 8</p>
Harnessing single cell RNA sequencing to identify dendritic cell types, characterize their biological states and infer their activation trajectory
<p><strong>Summary: </strong>Dendritic cells (DCs) orchestrate innate and adaptive immunity, by translating the sensing of distinct danger signals into the induction of different effector lymphocyte responses, to induce different defense mechanisms suited to face distinct types of threats. Hence, DCs are very plastic, which results from two key characteristics. First, DCs encompass distinct cell types specialized in different functions. Second, each DC type can undergo different activation states, fine-tuning its functions depending on its tissue microenvironment and the pathophysiological context, by adapting the output signals it delivers to the input signals it receives. Hence, to better understand DC biology and harness it in the clinic, we must determine which combinations of DC types and activation states mediate which functions, and how.<br> To decipher the nature, functions and regulation of DC types and their physiological activation states, one of the methods that can be harnessed most successfully is ex vivo single cell RNA sequencing (scRNAseq). However, for new users of this approach, determining which analytics strategy and computational tools to choose can be quite challenging, considering the rapid evolution and broad burgeoning of the field. In addition, awareness must be raised on the need for specific, robust and tractable strategies to annotate cells for cell type identity and activation states. It is also important to emphasize the necessity of examining whether similar cell activation trajectories are inferred by using different, complementary methods. In this chapter, we take these issues into account for providing a pipeline for scRNAseq analysis and illustrating it with a tutorial reanalyzing a public dataset of mononuclear phagocytes isolated from the lungs of naïve or tumor-bearing mice. We describe this pipeline step-by-step, including data quality controls, dimensionality reduction, cell clustering, cell cluster annotation, inference of the cell activation trajectories and investigation of the underpinning molecular regulation. It is accompanied with a more complete tutorial on Github. We anticipate that this method will be helpful for both wet lab and bioinformatics researchers interested in harnessing scRNAseq data for deciphering the biology of DCs or other cell types, and that it will contribute to establishing high standards in the field.</p> <p> </p> <p><strong>Data:</strong></p> <p>1. negative_cDC1_relative_signatures.csv : Negative signatures for performing Connectivity Map (cMAP) Analysis</p> <p>2. positive_cDC1_relative_signatures.csv : Positive signatures for performing Connectivity Map (cMAP) Analysis</p>
Integrated Data of Single cell RNA sequencing for Human Pancreatic Adenocarcinoma
<p>These data are collected and integrated from five available deposit data and one original data of single cell RNA sequencing from human pancreatic adenocarcinoma. Further analyses data for bulk transcriptomics (such as TCGA )using scRNAseq data and re-clustering for ductal epithelial cells and fibroblasts are also stored in step by step. Moreover, all R code is uploaded.</p>
Sequencing data of RNA editing, RNA modifications, and transcriptional units in Listeria monocytogenes
<p>Sequencing data for "RNA editing, RNA modifications, and transcriptional units in <em>Listeria monocytogenes</em>" manuscript, which is submitted to BMC genomics.</p>
Direct RNA targeted in situ sequencing for transcriptomic profiling in tissue
<p>You can find here the Direct RNA In Situ Sequencing (HybISS-based) maps generated using the Hight Sensitivity kit from CARTANA AB. They include half a mouse brain coronal section, targeting 50 genes. Genes were targeted in a sequential manner. Both reads, DAPI staining and segmented cells are included. The analysis of the same cells, but using 10X magnification are also provided in an anndata object.</p>
Data files: Single-cell RNA sequencing of Plasmodium vivax sporozoites reveals stage- and species-specific transcriptomic signatures
<p>Scripts, preprocessed count matrices, single-cell data objects, and generated data (tables and .rds files) from the scRNA-seq analyses performed in <strong>“Single-cell RNA sequencing of Plasmodium vivax sporozoites reveals stage- and species-specific transcriptomic signatures".</strong></p> <p> </p>
Single-cell and single-nucleus RNA-sequencing from paired normal-adenocarcinoma lung samples provides both common and discordant biological insights
<p>The datasets generated by <em>Cellranger </em>for all 24 samples (.h5 format).<br><br></p>
FedscGen: privacy-aware federated batch effect correction of single-cell RNA sequencing data -- Preprocessed datasets
<div> <div> <div> <div> <p>This dataset accompanies the publication "FedscGen: Privacy-Aware Federated Batch Effect Correction of Single-Cell RNA Sequencing Data" and includes eight single-cell RNA sequencing (scRNA-seq) datasets used to benchmark the FedscGen and scGen methods. The datasets are provided in <code>.h5ad</code> format and include comprehensive metadata necessary for replication and further analysis.</p> <h3>Datasets</h3> <p>We analyze various datasets to compare FedscGen against scGen (centralized) in terms of batch correction. For simplicity, we refer to the dataset by abbreviations:</p> <ol> <li> <p><strong>Cell Line (CL)</strong>:</p> <ul> <li>Derived from the 293t_jurkat experiment with three batches: Zheng et al., 2017.</li> </ul> </li> <li> <p><strong>Human Dendritic Cells (HDC)</strong>:</p> <ul> <li>scRNA-seq data of human dendritic cells across two batches: Villani et al., 2017.</li> </ul> </li> <li> <p><strong>Human Pancreas (HP)</strong>:</p> <ul> <li>Consolidated data from five sources with 14,767 cells each: Baron et al., 2016; Muraro et al., 2016; Segerstolpe et al., 2016; Wang et al., 2016; Xin et al., 2016.</li> </ul> </li> <li> <p><strong>Mouse Brain (MB)</strong>:</p> <ul> <li>Merged datasets with 691,600 and 141,606 cells: Saunders et al., 2018; Rosenberg et al., 2018.</li> </ul> </li> <li> <p><strong>Mouse Cell Atlas (MCA)</strong>:</p> <ul> <li>Data focusing on 11 cell types from various organs: Han et al., 2018; The Tabula Muris Consortium, 2018.</li> </ul> </li> <li> <p><strong>Mouse Hematopoietic Stem and Progenitor Cells (MHSPC)</strong>:</p> <ul> <li>Data from SMART-seq2 and MARS-seq protocols: Nestorowa et al., 2016; Paul et al., 2015.</li> </ul> </li> <li> <p><strong>Mouse Retina (MR)</strong>:</p> <ul> <li>Data from two unassociated laboratories with 26,830 and 44,808 cells: Macosko et al., 2015; Shekhar et al., 2016.</li> </ul> </li> <li> <p><strong>PBMC (human Peripheral Blood Mononuclear Cell)</strong>:</p> <ul> <li>scRNA-seq data with two batches: Zheng et al., 2017.</li> </ul> </li> </ol> <p><strong>Usage Notes</strong>: Each dataset is provided in <code>.h5ad</code> format, compatible with common single-cell analysis tools such as Scanpy. Detailed metadata is included within each file.</p> <p><strong>Keywords</strong>: Single-cell RNA sequencing, scRNA-seq, Batch effect correction, Privacy-aware, Federated learning, scGen, FedscGen, Clinical multi-center studies, Genomics, Bioinformatics</p> <p><strong>Contact</strong>: For questions or further information, please contact Mohammad Bakhtiari at <a href="mailto:mohammad.bakhtiari@uni-hamburg.de.">mohammad.bakhtiari@uni-hamburg.de.</a></p> <p><strong>License</strong>: Creative Commons Attribution 4.0 International (CC BY 4.0)</p> </div> </div> </div> </div> <div> <div> <div> </div> </div> </div>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.