Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,401
datasets available to search
ShareScore release 0.9.0
Dataset results
1,401 results for “single-cell Seq”
Teaching Datasets for Single-Cell RNA-seq Analysis Course
<p>This repository contains teaching datasets used in the Single-Cell RNA-seq Analysis Course, which is part of the SeuratExtend project (<a href="https://github.com/huayc09/SeuratExtend">https://github.com/huayc09/SeuratExtend</a>). The course materials are available at <a href="https://huayc09.github.io/SeuratExtend/#single-cell-rna-seq-analysis-course-new-in-v110-1">https://huayc09.github.io/SeuratExtend/#single-cell-rna-seq-analysis-course-new-in-v110-1</a>, with code and scripts hosted at <a href="https://github.com/huayc09/single-cell-course">https://github.com/huayc09/single-cell-course</a>.</p> <p>The datasets include:</p> <ol> <li>3k Peripheral Blood Mononuclear Cells (PBMCs)</li> <li>Paired PBMC samples processed with 10x Genomics' 3' kit</li> <li>Paired PBMC samples processed with 10x Genomics' 5' kit</li> </ol> <p>All original data were obtained from 10x Genomics (<a href="https://www.10xgenomics.com/resources/datasets">https://www.10xgenomics.com/resources/datasets</a>) and processed for educational purposes.</p>
Accuracy, robustness and scalability of dimensionality reduction methods for single-cell RNA-seq analysis
<p>A detailed list of the selected scRNA-seq datasets used in the paper, also provided in Additional file <a href="https://genomebiology.biomedcentral.com/articles/10.1186/s13059-019-1898-6#MOESM1">1</a>: Table S1-S2.</p>
single-cell RNA-seq of EHT process
<p>This file 'atfr.rds' contains onject including the raw and log-normalizd expression counts, and TF activity matrix generated from metaRegulon. In addition, the file 'cell_weight_pseudotime.rds' contains the pseudotime information.</p>
Single-cell RNA-seq supplementary data
<p>These files are supplementary data for the article Laczik M, Erdős E, Ozgyin L, Hevessy Zs, Csősz É, Kalló G, Nagy T, Barta E, Póliska Sz, Szatmári I and Bálint BL: Extensive proteome and functional genomic profiling of variability between genetically identical human B-lymphoblastoid cells. The files are outputs from the 10x Genomics software Cell Ranger and Loupe Browser, they contain QC data for the single cell sequencing described in the article, and also comparative analyses between 3 datasets (GM22648old, GM22648new, GM22649). For further details please see the article.</p>
Single-cell RNA-seq datasets derived from human parasitic nematode Brugia malayi microfilariae
<p>Single-cell RNA-seq data of the parasitic nematode <em>Brugia malayi</em> in the microfilariae development stage. Data includes untreated cellular transcriptional states and the transcriptional response to ivermectin (1 µM). </p> <p>The unfiltered gene expression matrix is provided as both a Seurat and AnnData object. Filtered datasets are provided as .csv files for direct import into R. </p>
AsaruSim: a single-cell and spatial RNA-Seq Nanopore long-reads simulation workflow
Open the record for dataset details and reuse information.
Single-cell landscape of innate and acquired drug resistance in acute myeloid leukemia: scRNA-seq and CyTOF processed datasets
<p><strong>This data was generated as part of the Tumor Profiler study. If you use it in your research, please cite:</strong></p> <p>Wegmann, R., Bonilla, X., Casanova, R. <em>et al.</em> Single-cell landscape of innate and acquired drug resistance in acute myeloid leukemia. <em>Nat Commun</em> 15, 9402 (2024). https://doi.org/10.1038/s41467-024-53535-4</p> <p><strong>Derived data - scRNA-seq</strong></p> <p>This is an R data set (.RDS) containing a SingleCellExperiment object with the following slots:</p> <div> <ul> <li>Assays: <ul> <li>counts: raw counts</li> </ul> </li> </ul> </div> <div> <ul> <li>colData: Cell-level metadata <ul> <li> barcodes: The cell barcode</li> <li> fractionMT: Fraction mitochondrial genes per cell</li> <li> n_umi: Total number of UMIs per cell</li> <li> n_gene: Total number of genes per cell</li> <li> log_umi: log10 total number of UMIs per cell</li> <li> g2m_score: Cell cycle phase score for G2M</li> <li>s_score: Cell cycle phase score for S</li> <li>cycle_phase: predicted cell cycle phase</li> <li>celltype_major_full_ct_name: Major cell type full name</li> <li>celltype_major: Major cell type short name</li> <li>celltype_final_full_ct_name: Cell subtype full name</li> <li>celltype_final: Cell subtype short name </li> <li>sample_id </li> </ul> </li> </ul> </div> <div> <ul> <li>rowData: Gene-level metadata <ul> <li>gene_ids</li> <li>gene_names</li> </ul> </li> </ul> </div> <p><strong>Derived data - CyTOF</strong></p> <p>This is an R data set (.RDS) containing a SingleCellExperiment object with the following slots:</p> <ul> <li>Assays:<br> <ul> <li>counts_raw: signal intensity based on CyTOF dual counts</li> <li>exprs_raw: arcsinh transformed raw counts (cofactor 5)</li> <li>counts: batch corrected raw counts (linear scaling based on a quantile)</li> <li>exprs: arcsin transformed counts (cofactor 5)</li> <li>scaled: 0-1 normalized exprs (clipped to the 99.95th percentile)</li> </ul> </li> <li>colData (cell metadata) <ul> <li>bc_id: barcode of the sample during staining </li> <li>run: CyTOF experiment batch, named after the first sample of the batch</li> <li>type: Sample type (blood or bone marrow)</li> <li>sample_id: TuPro sample ID</li> <li>pred_id: Predicted cell type [char]</li> <li>pred_n: Predicted cell type [integer]</li> </ul> </li> <li>rowData (marker metadata) <ul> <li>channel_name: Name and isotopic mass of the metal ion corresponding to this marker</li> <li>marker_name: Protein name</li> <li>channel_group, channel_group_integer: Biological processes the channel identifies, e.g. specific cell type, signalling, cell death</li> <li>tsne_channel: Logical - use this channel for dimensionality reduction?</li> <li>channel_order: Define the order of channels for plotting</li> <li>cluster_channel: Logical - use this channel for clustering?</li> </ul> </li> </ul>
scCTS: identifying the cell type-specific marker genes from population-level single-cell RNA-seq
<p>Single cell RNA-sequencing (scRNA-seq) provides gene expression profiles of individual cells from complex samples, facilitating the detection of cell type-specific marker genes. In scRNA-seq experiments with multiple donors, the population level variation brings an extra layer of complexity in cell type-specific gene detection, for example, they may not appear in all donors. Motivated by this observation, we develop a statistical model named scCTS to identify cell type-specific genes from population-level scRNA-seq data. Extensive data analyses demonstrate that the proposed method identifies more biologically meaningful cell type-specific genes compared to traditional methods.</p>
Processed Single-cell RNA-seq Data for Exercise Rejuvenation Intervention on Mouse Subventricular Zone
<p>Exercise single-cell RNA-seq data of mouse subventricular zone (in processed Seurat object) used in the publication "Cell type-specific aging clocks to quantify aging and rejuvenation in regenerative regions of the brain" (preprint available at <a href="https://doi.org/10.1101/2022.01.10.475747">https://doi.org/10.1101/2022.01.10.475747</a>).</p>
Training material for the mapping and quantification of single-cell ATAC-seq 10X Datasets
<p>The data provided here is part of the Galaxy Training Network tutorial that analyses 10x genomics single-cell ATAC-seq data from the 10x platform. </p> <p>Due to time constraints during training, the datasets were subsampled to reads that map to chrY.</p> <p>The 10x Genomics Datasets follow the <a href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution</a> license.</p>
Single-Cell RNA-Seq Identifies Pathways and Genes Contributing to the Hyperandrogenemia Associated with Polycystic Ovary Syndrome
<p>Polycystic ovary syndrome (PCOS) is a common endocrine disorder characterized by hyperandrogenemia of ovarian thecal cell origin, resulting in anovulation/oligo-ovulation and infertility. Our previous studies established that ovarian theca cells isolated and propagated from ovaries of normal ovulatory women and women with PCOS, have distinctive molecular and cellular signatures that underlie the increased androgen biosynthesis in PCOS. To evaluate differences between gene expression in single cells from passaged cultures of theca cells from ovaries of normal ovulatory women and women with PCOS, we performed single-cell RNA sequencing (scRNA-seq). Results from these studies revealed differentially expressed pathways and genes involved in the acquisition of cholesterol, the precursor of steroid hormones, and steroidogenesis. Bulk RNA-seq and microarray studies confirmed the theca cell differential gene expression profiles. The expression profiles appear to be directed largely by increased levels or activity of the transcription factors SREBF1, which regulates genes involved in cholesterol acquisition (<em>LDLR, LIPA, NPC1, CYP11A1, FDX1, FDXR)</em> and GATA6, which regulates expression of genes encoding steroidogenic enzymes (<em>CYP17A1) </em>in concert with other differentially expressed transcription factors (<em>SP1</em>, <em>NR5A2</em>). This study provides insights into the molecular mechanisms underlying the hyperandrogenemia associated with PCOS, and highlights potential targets for molecular diagnosis and therapeutic intervention.</p> <p><strong>scRNA-seq data in the form of 10x Cell Ranger files are available for the following samples:</strong></p> <p>PCOS affected - Mc03, Mc10, Mc16, Mc26, Mc27</p> <p>Normal cycling women - Mc02, Mc06, Mc31, Mc40, Mc50</p> <p>A F in the sample name indicates forskolin treatment and a C indicates untreated control samples.</p> <p> </p>
Immune single-cell RNA-seq data from PyMT-M tumor and its peripheral blood samples
<p><em>Immune single-cell RNA-seq data collection and preprocessing</em></p> <p>Blood and tumor samples were harvested from PyMT-M tumor-bearing mice. Blood samples (n=3) are collected retro-orbitally using caliper tubes and processed with red blood cell lysis buffer (Tonbo Biosciences) before library preparation. A tumor sample are dissociated with Tumor Dissociation Kit following the manufacturer’s instructions (Miltenyi Biotec). After isolation and filtering through a 70µm filter, CD45+DAPI- cells were sorted using FACSAria cell sorter (BD Biosciences) at the Cytometry and Cell Sorting Core. The single-cell libraries were prepared using Chromium Controller (10X Genomics) at the Single Cell Genomics Core and sequenced using NovaSeq 6000 at the Genomics and RNA Profiling Core of Baylor College of Medicine. The FASTQ files were processed using Cell Ranger pipelines (10X Genomics) to generate feature-barcode matrices.</p> <p><em>Integrating immune single-cell RNA-seq data from the blood and tumor of PyMT-M mouse</em></p> <p>We followed the Seurat tutorial on single-cell RNA integration from <a href="https://satijalab.org/seurat/articles/integration_introduction.html">https://satijalab.org/seurat/articles/integration_introduction.html</a>. Specifically, both datasets were library-size normalized and log-scaled. Then, variable genes from both datasets were extracted, and overlapped variable genes were used as anchors to integrate the two datasets to generate a combined immune single-cell RNA-seq dataset.</p> <p> </p> <p><em>Immune cell type annotation using SingleR</em></p> <p>After having the integrated immune single-cell RNA-sequencing data, we performed the standard pipeline for clustering, including scaling the expression data, performing dimension reduction using PCA and UMAP, and finding clusters by a shared nearest neighbor (SNN) modularity optimization (<a href="https://satijalab.org/seurat/articles/integration_introduction.html">https://satijalab.org/seurat/articles/integration_introduction.html</a>). Then, we used the R package <em>SingleR </em>to assign cell-type labels to each identified cluster using the <em>ImmGen</em> reference data from the Immunological Genome Project.</p> <p> </p>
Improving cell type identification with Gaussian noise-augmented single-cell RNA-seq contrastive learning
<p>The benchmark datasets used to evaluate Gaussian noise augmentation-based scRNA-seq contrastive learning (GsRCL) against scRNA-seq cell-type identification tasks.</p>
Single-cell RNA-seq dataset to determine cell-type specific response to fungal pathogen infection in plant leaves
<p>Single-cell RNA-seq dataset to determine cell-type specific response to fungal pathogen infection in plant leaves</p>
Raw and normalized count data for "Probabilistic cell-type assignment of single-cell RNA-seq for tumor microenvironment profiling"
<p>SingleCellExperiment objects containing raw and normalized counts, as well as reduced dimension representations and cell type annotations for both the follicular lymphoma samples (sce_follicular_annotated_final.rds) and high grade serous ovarian cancer samples (sce_hgsc_annotated_final.rds) as detailed in the paper, as well as reactive lymph node data (sce_RLN_normalized.rds).</p>
Data Repository: Single-cell mapper (scMappR): using scRNA-seq to infer cell-type specificities of differentially expressed genes
<p>Data repository for the scMappR manuscript:</p> <p>Abstract from biorXiv (https://www.biorxiv.org/content/10.1101/2020.08.24.265298v1.full).</p> <p>RNA sequencing (RNA-seq) is widely used to identify differentially expressed genes (DEGs) and reveal biological mechanisms underlying complex biological processes. RNA-seq is often performed on heterogeneous samples and the resulting DEGs do not necessarily indicate the cell types where the differential expression occurred. While single-cell RNA-seq (scRNA-seq) methods solve this problem, technical and cost constraints currently limit its widespread use. Here we present single cell Mapper (scMappR), a method that assigns cell-type specificity scores to DEGs obtained from bulk RNA-seq by integrating cell-type expression data generated by scRNA-seq and existing deconvolution methods. After benchmarking scMappR using RNA-seq data obtained from sorted blood cells, we asked if scMappR could reveal known cell-type specific changes that occur during kidney regeneration. We found that scMappR appropriately assigned DEGs to cell-types involved in kidney regeneration, including a relatively small proportion of immune cells. While scMappR can work with any user supplied scRNA-seq data, we curated scRNA-seq expression matrices for ∼100 human and mouse tissues to facilitate its use with bulk RNA-seq data alone. Overall, scMappR is a user-friendly R package that complements traditional differential expression analysis available at CRAN.</p>
Supplementary Material for the paper entitled "Scalable nonparametric clustering with unified marker gene selection for single-cell RNA-seq data"
<p>This repo contain supplementary tables from the manuscript entitled: "Scalable nonparametric clustering with unified marker gene selection for single-cell RNA-seq data". Clustering is a common way to identify cell types in single-cell RNA-sequencing (scRNA-seq) data. Unfortunately, current methods (i) require users to make human-in-the-loop decisions, which adds significant runtime to bioinformatic analyses, and (ii) reuse the same data twice when testing for differentially expressed genes, which can lead to an increased number of false discoveries. In this work, we overcome these limitations with NCLUSION: a Bayesian nonparametric method that simultaneously clusters cells and selects marker genes. NCLUSION operates without user-defined heuristics to set model parameters and leverages variational expectation-maximization (EM) for posterior inference which allows it to scale well up to 1 million cells. By analyzing publicly available datasets, we illustrate that NCLUSION matches the state-of-the-art clustering performance of competing approaches, achieves improved computational efficiency, and directly enables identification of biologically relevant gene sets driving cluster definitions.</p>
Single-cell RNA-seq dataset and code for human colorectal cancer liver metastasis study
Open the record for dataset details and reuse information.
Supplementary information to "ScRNA-IMM: Single-cell RNA-Seq Imputation method using Mean/Median Imputation"
Open the record for dataset details and reuse information.
An integrated single-cell RNA-seq atlas of the mouse hypothalamic paraventricular nucleus links transcriptional and function types
<p>The hypothalamic paraventricular nucleus (PVN) is a highly complex brain region that is crucial for homeostatic<br> regulation through neuroendocrine signalling, outflow of the autonomic nervous system (ANS), and projections<br> to other brain areas. The past years, single-cell datasets of the hypothalamus have contributed immensely<br> to the current understanding of the diverse hypothalamic cellular composition. While the PVN has been<br> adequately classified functionally, its molecular classification is currently still insufficient.</p> <p>To address this, we created a detailed atlas of PVN transcriptional cell types by integrating various PVN<br> single-cell datasets into a recently published hypothalamus single-cell transcriptome atlas. Furthermore, we<br> functionally profiled transcriptional cell types, based on relevant literature, existing retrograde tracing data<br> and existing single-cell data of a PVN-projection target region.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.