Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2,750
datasets available to search
ShareScore release 0.9.0
Dataset results
2,750 results for “scRNA seq”
SpaGE Spatial Gene Enhancement using scRNA-seq
<p>Spatial transcriptomics and scRNA-seq datasets using for integration and prediction of spatially non-measured genes using SpaGE</p>
Post-processed datasets for scRNA-seq clustering analysis in PPML-Omics
Open the record for dataset details and reuse information.
Isosceles paper simulated ovarian cell line ONT data (scRNA-Seq)
<div> <div> <p>Simulated ovarian cell line ONT data (scRNA-Seq) for the Isosceles paper - more details can be found in the <a href="https://github.com/Genentech/Isosceles_Paper" target="_blank" rel="noopener">Isosceles_Paper</a> repository.</p> </div> </div>
MTN deficient scRNA-seq
<p>"3KR" refers to scRNA-seq data from the intestinal tissues of mice fed a methionine-, tryptophan-, and niacin-deficient diet followed by recovery on a regular diet.<br>"3K" refers to scRNA-seq data from the intestines of mice continuously fed a methionine-, tryptophan-, and niacin-deficient diet.<br>"REG" represents scRNA-seq data from the intestinal tissues of mice maintained on a regular diet.<br>The file <em>mouse_combined.rds</em> contains the integrated dataset of these three conditions, processed using Seurat.</p>
processed scRNA-seq data for Neuwirth & Malzl et al. 2024
<p>This dataset contains the analysed scRNA-seq data used in Neuwirth & Malzl et al. 2024 as AnnData objects. You can find the code used to produce and analyse this data on <a href="https://github.com/menchelab/Neuwirth_Malzl_et_al_2024">GitHub</a>. All data was preprocessed using cellranger v6.0.1 with GRCh38 3.0.0 as reference.</p> <h3><strong>Psoriasis and Sarcoidosis PBMC data</strong></h3> <p><strong><em>pbmc.scps.integrated.clustered.h5ad</em></strong>: contains scVI-integrated and Leiden-clustered data of Psoriasis and Sarcoidosis patient blood as well as healthy controls (Psoriasis data was generated within this study; Sarcoidosis data was reprocessed from 10.1016/j.immuni.2023.01.014)<br><strong><em>tcells.pbmc.scps.integrated.clustered.h5ad</em></strong>: contains scVI-integrated and Leiden-clustered T cell subset of Psoriasis and Sarcoidosis PBMC data</p> <h3><strong>Atopic dermatitis skin data</strong></h3> <p><strong><em>tissue.ad.integrated.clustered.h5ad</em></strong>: contains scVI-integrated and Leiden-clustered data of healthy and atopic dermatitis patient skin (reprocessed from 10.1126/science.aba6500)<br><strong><em>tcells.tissue.ad.integrated.clustered.h5ad</em></strong>: contains scVI-integrated and Leiden-clustered T cell subset of atopic dermatitis data</p> <h3><strong>Psoriasis and Sarcoidosis skin data</strong></h3> <p><strong><em>tissue.scps.integrated.annotated.h5ad</em></strong>: contains scVI-integrated and celltypist-annotated data from Psoriasis and Sarcoidosis patient skin as well as healthy controls (Psoriasis data was reprocessed from 10.1126/science.aba6500; Sarcoidosis data was reprocessed from 10.1016/j.immuni.2023.01.014)<br><strong><em>tcells.tissue.scps.integrated.clustered.h5ad</em></strong>: contains scVI-integrated and Leiden-clustered T cell subset of Psoriasis and Sarcoidosis skin data<br><strong><em>tregs.tissue.scps.integrated.annotated.h5ad</em></strong>: contains scVI-integrated and SAT1 status annotated regulatory T cell subset of Psoriasis and Sarcoidosis skin data<br><em><strong>tregs.tissue.scps.integrated.milo.h5ad</strong></em>: basically same as above but with cell neighborhood overrepresentation analysis on top</p> <h3><strong>IBD colon data</strong></h3> <p><strong><em>tissue.uc.integrated.clustered.h5ad</em></strong>: contains scVI-integrated and Leiden-clustered data of Crohn's disease and ulcerative colitis patients as well as healthy controls (reprocessed from 10.1126/sciimmunol.abb4432)<br><strong><em>tcells.tissue.uc.integrated.clustered.h5ad</em></strong>: contains scVI-integrated and Leiden-clustered T cell subset of IBD data</p> <h3><strong>Raw and unfiltered data</strong></h3> <p><strong><em>inflammatory_disease.h5ad</em></strong>: contains the raw, unfiltered and unprocessed data of all the files above (i.e. combined cellranger output) and is the source data file of all analyses in this study. If you just want the untouched data this is what you want to use.</p>
PBMC scRNA-seq datasets measured using different 10X Chromium chemistries
<p>PBMC scRNA-seq datasets measured using different 10X Chromium chemistries<br> Obtained from: https://www.10xgenomics.com/resources/datasets</p>
scRNA-seq dataset of iPSC-derived pancreatic islet cells
Open the record for dataset details and reuse information.
Processed Seurat objects from scRNA-seq data of the aging subventricular zone (SVZ) neurogenic niche with partial reprogramming
<p>This repository contains the processed Seurat objects from the publication "Restoration of neuronal progenitors by partial reprogramming in the aged neurogenic niche" (https://doi.org/10.1038/s43587-024-00594-3).</p> <p>Raw sequencing data is available at the Gene Expression Omnibus (GEO) under accession number GSE224438. Code used to process and analyze the data is available on GitHub (https://github.com/gitlucyxu/SVZreprogramming). </p> <p>These Seurat objects are filtered to high-quality singlets for samples included in the publication. Descriptions and notable metadata:</p> <ul> <li>svz_iOSKM_cohort1_toshare.rds - SVZ after whole-body partial reprogramming, cohort 1 <ul> <li>Celltype - cell type annotation</li> <li>Treatment - condition <ul> <li>untr: old control</li> <li>2Dox0: old+OSKM</li> </ul> </li> <li>hash.ID - mouse ID (biological replicate)</li> </ul> </li> <li>svz_iOSKM_cohort2_toshare.rds - SVZ after whole-body partial reprogramming, cohort 2 <ul> <li>Celltype - cell type annotation</li> <li>Age_Treatment - condition <ul> <li>young_untr: young control</li> <li>old_untr: old control</li> <li>old_2Dox0: old+OSKM</li> </ul> </li> <li>hash.ID - mouse ID (biological replicate)</li> </ul> </li> <li>svz_ciOSKM_toshare.rds - SVZ after SVZ-targeted partial reprogramming <ul> <li>Celltype - cell type annotation</li> <li>Age_Treatment - condition <ul> <li>young_untr: young control</li> <li>old_untr: old control</li> <li>old_Dox: old+OSKM(SVZ)</li> </ul> </li> <li>MULTI_classification_rescued - mouse ID (biological replicate)</li> </ul> </li> </ul> <p> </p> <p><em>Updated 2024/07/01 (v2): replaced a corrupted file. </em></p>
DynToy simulated scRNA-seq datasets
Open the record for dataset details and reuse information.
scRNA-seq of murine tendon from sham and injured group
<p>ScRNA-seq was performed on murine single cell suspension collected from sham and injured group. <span>Tendon</span><span> tissues were collected and dissected into pieces. Then they were digested in </span><span>a mix of collagenase II (Roche, 3 mg/ml) and dispase II (Sigma-Aldrich, 4 mg/ml), prepared in DMEM for 1 hour at 37 °C. Digestions were subsequently quenched with 10% FBS DMEM and filtered through 40μm sterile strainers. Cells were then washed in PBS with 0.04% BSA, counted and resuspended at a concentration of ~1000 cells/μl. Cell viability was assessed with Trypan blue exclusion on a Countess II (Thermo Fisher Scientific) automated counter and only samples with >85% viability were processed for further sequencing.</span><span> </span></p>
Identification of Kidney Cell Types in scRNA-seq and snRNA-seq Data Using Machine Learning Algorithms
<p>Metadata files for our final analysis as part of the "<span><span>Identification of Kidney Cell Types in scRNA-seq and snRNA-seq Data Using Machine Learning Algorithms" study.</span></span></p>
Single-cell landscape of innate and acquired drug resistance in acute myeloid leukemia: scRNA-seq and CyTOF processed datasets
<p><strong>This data was generated as part of the Tumor Profiler study. If you use it in your research, please cite:</strong></p> <p>Wegmann, R., Bonilla, X., Casanova, R. <em>et al.</em> Single-cell landscape of innate and acquired drug resistance in acute myeloid leukemia. <em>Nat Commun</em> 15, 9402 (2024). https://doi.org/10.1038/s41467-024-53535-4</p> <p><strong>Derived data - scRNA-seq</strong></p> <p>This is an R data set (.RDS) containing a SingleCellExperiment object with the following slots:</p> <div> <ul> <li>Assays: <ul> <li>counts: raw counts</li> </ul> </li> </ul> </div> <div> <ul> <li>colData: Cell-level metadata <ul> <li> barcodes: The cell barcode</li> <li> fractionMT: Fraction mitochondrial genes per cell</li> <li> n_umi: Total number of UMIs per cell</li> <li> n_gene: Total number of genes per cell</li> <li> log_umi: log10 total number of UMIs per cell</li> <li> g2m_score: Cell cycle phase score for G2M</li> <li>s_score: Cell cycle phase score for S</li> <li>cycle_phase: predicted cell cycle phase</li> <li>celltype_major_full_ct_name: Major cell type full name</li> <li>celltype_major: Major cell type short name</li> <li>celltype_final_full_ct_name: Cell subtype full name</li> <li>celltype_final: Cell subtype short name </li> <li>sample_id </li> </ul> </li> </ul> </div> <div> <ul> <li>rowData: Gene-level metadata <ul> <li>gene_ids</li> <li>gene_names</li> </ul> </li> </ul> </div> <p><strong>Derived data - CyTOF</strong></p> <p>This is an R data set (.RDS) containing a SingleCellExperiment object with the following slots:</p> <ul> <li>Assays:<br> <ul> <li>counts_raw: signal intensity based on CyTOF dual counts</li> <li>exprs_raw: arcsinh transformed raw counts (cofactor 5)</li> <li>counts: batch corrected raw counts (linear scaling based on a quantile)</li> <li>exprs: arcsin transformed counts (cofactor 5)</li> <li>scaled: 0-1 normalized exprs (clipped to the 99.95th percentile)</li> </ul> </li> <li>colData (cell metadata) <ul> <li>bc_id: barcode of the sample during staining </li> <li>run: CyTOF experiment batch, named after the first sample of the batch</li> <li>type: Sample type (blood or bone marrow)</li> <li>sample_id: TuPro sample ID</li> <li>pred_id: Predicted cell type [char]</li> <li>pred_n: Predicted cell type [integer]</li> </ul> </li> <li>rowData (marker metadata) <ul> <li>channel_name: Name and isotopic mass of the metal ion corresponding to this marker</li> <li>marker_name: Protein name</li> <li>channel_group, channel_group_integer: Biological processes the channel identifies, e.g. specific cell type, signalling, cell death</li> <li>tsne_channel: Logical - use this channel for dimensionality reduction?</li> <li>channel_order: Define the order of channels for plotting</li> <li>cluster_channel: Logical - use this channel for clustering?</li> </ul> </li> </ul>
Pan-Cancer T cell atlas from "The combined use of scRNA-seq and network propagation highlights key features of pan-cancer Tumor-Infiltrating T cells" (https://doi.org/10.1371/journal.pone.0315980)
<p>The scRNA-seq data were collected from previously published datasets (GSE140228, GSE139555, GSE155698, GSE121636, and GSE139324), adhering to the following selection criteria: 1) presence of T cells, 2) treatment-naïve patients, 3) solid tumors, and 4) inclusion of at least tumor and blood samples.<br>Each scRNA-seq dataset underwent separate preprocessing in R (v4.0.2). We filtered out cells from the original count matrices that had fewer than 200 genes detected or more than 10% mitochondrial UMI counts and we only kept genes detected in at least 3 cells. Then, we applied Seurat (v4.0.5) with default parameters for count data normalization and scaling. Each cell was assigned a cell cycle score using the CellCycleScoring function and we computed the difference between the G2M and S phase scores. This approach allows for the separation of non-cycling from cycling cells while minimizing the differences in cell cycle phase among proliferating cells. The SelectIntegrationFeatures function was ran with the nfeatures parameter set to 3,000 before merging all samples from each dataset. These integration features were then used for Principal Component Analysis (PCA) and Uniform Manifold Approximation and Projection (UMAP). Clustering was performed using the Louvain algorithm with the resolution parameter set to 2.0 for all datasets. Finally, T cells were isolated based on CD3D and CD3G genes expression (CD3D or CD3G expression level > 0).</p> <p>To integrate heterogeneous data from different sources, a two-step procedure was applied. We first concatenated all datasets together and ran the scaling and PCA steps based on the top 3,000 highly variable genes identified by the FindVariableFeatures function with the “vst” method. Harmony was applied for batch effect correction then UMAP and clustering using the Louvain algorithm with the resolution parameter set to 2.0 were performed on the harmony reduction. Examining the result from the first clustering run, we identified contamination clusters and clusters that arose from unwanted factors: we removed the contamination clusters including low quality cells highly expressing marker genes associated with apoptosis and tissue dissociation operation, pancreatic acinar cells (expressing PRSS1, CLPS, PNLIP and CTRB1 among others), myeloid cells (expressing CD68) and B cells (expressing CD79A). Then, we performed the second run of integration and clustering excluding immunoglobulin, ribosome-protein-coding, and T cell receptor (TCR) genes (gene symbol with string pattern "^IGK|^IGH|^IGL|^IGJ|^IGS|^IGD|IGFN1", "^RP([0–9]+-|[LS])", and "^TRA|^TRB|^TRG" respectively) from the top 3,000 highly variable genes and regressing out the cell cycle difference effect as well as the percentage of mitochondrial UMI counts. Harmony (v0.1.0) was applied again for batch effect correction and UMAP was performed on the harmony reduction.<br>T cell subtypes identification and annotation was performed by clustering cells using the Louvain algorithm with the resolution parameter set to 4.1 after iterative testing from 3.5 to 5.0 by 0.1 (more granular than default), computing clusters signatures based on differential gene expression using the FindAllMarkers function with the “MAST” method and interrogating known gene markers expression. A resolution value of 4.1 was notably found to be the lowest resolution value enabling the correct separation of proliferating CD4+ T cells from proliferating CD8+ T cells.</p>
scRNA-seq dataset from "Glucose deprivation and identification of TXNIP as an immunometabolic modulator of T cell activation in cancer"
<p>Gene expression profiling analysis of single cell RNA-seq data from MLR, anti-CD3/anti-CD28 treated and paired untreated CD4+ T cells samples under high glucose (11 mM) and low glucose (1 mM) conditions.</p> <p>For each sample, cells suspensions in culture medium were recovered, washed once with 0.04% BSA in 1X PBS and processed through 10x Cell Multiplexing Oligo Labeling protocol (10x Genomics, USA) according to the manufacturer’s instructions. ~1,600 cells/µl pooled cell suspensions were prepared with equal number of cells per sample: one for MLR samples, one for non-stimulated T cells samples, and one for anti-CD3/anti-CD28-stimulated T cells samples. Libraries were prepared using the 10x Chromium Single-Cell 3’ v3.1 protocol with Feature Barcode (10x Genomics, USA), according to the manufacturer’s instructions. Sequencing was performed on a NovaSeq 6000 sequencer (Illumina, USA).</p> <p>Cell Ranger (v6.0.1, 10x Genomics Inc) was applied for demultiplexing, reads mapping against the GRCh38 human reference genome, and UMI counting. Seurat package (v4.4.0) was used to generate Seurat objects. Only genes detected in at least 3 cells were kept. Cells with fewer than 200 genes detected or >15% mitochondrial UMI counts were filtered out. Samples were merged in a unique Seurat object then count data normalization and scaling was performed using Seurat with default parameters. The 2000 most highly variable genes were used for Principal Component Analysis (PCA). Harmony (v0.1.1) was applied for batch effect correction then Uniform Manifold Approximation and Projection (UMAP) and clustering using the Louvain algorithm were performed on the harmony reduction. Non-T or -MoDC clusters were removed for further analysis.</p>
Meta-analysis of scRNA-seq Co-expression in Human Neural Organoids Reveals High Variability in Recapitulating Primary Tissue
<p>Contains all code and data for Werner and Gillis, Meta-analysis of scRNA-seq Co-expression in Human Neural Organoids Reveals High Variability in Recapitulating Primary Tissue, 2024. </p> <p>Additionally, the code and data for this paper can be found at https://github.com/JonathanMWerner/meta_organoid_analysis with an easy to view github markdown file containing all the code used to generate all figure panel plots at https://github.com/JonathanMWerner/meta_organoid_analysis/blob/main/figure_plots_with_data_code.md.</p> <p>Due to file size limits on github, there are several data files not available on github, but are available here on zenodo in the data_for_plots.zip file, see below:</p> <pre>umap_embeddings_Fig2A.Rdata<br>cross_dataset_aggregated_exp_metaMarker_all_fetal_SuppFig1B_Fig2E.Rdata<br>organoid_egad_results_ranked_6_26_24_Fig3D.Rdata<br>fetal_egad_results_ranked_6_26_24_Fig3D.Rdata<br>org_eigenvec_matrices_SuppFig3CD.Rdata</pre> <p><br>The R package developed for this paper is available at https://github.com/JonathanMWerner/preservedCoexp</p>
Lung ECs scRNA-seq: Gene Expression and Metadata
<p>Normalized gene expression and cell metadata derived from a Seurat object.</p>
Human embryonic head scRNA-seq data
<p>These data contain scRNA-seq data of 50,059 single cells from the heads of five human embryos aged four to six weeks in Seurat object format. The data are the result of a re-analysis of primary scRNA-seq data from whole embryos (Xu et al., 2023 Nat. Cell Biol.; GEO accession number: GSE157329).</p> <p>The re-analysis is described in:</p> <p>Siewert, A., Hoeland, S., Mangold, E. <em>et al.</em> Combining genetic and single-cell expression data reveals cell types and novel candidate genes for orofacial clefting. <em>Sci Rep</em> <strong>14</strong>, 26492 (2024). https://doi.org/10.1038/s41598-024-77724-9</p> <div> <div> </div> </div>
Isosceles paper mouse E18 brain scRNA-Seq (Lebrigand et al.) data analysis input files
<p>Mouse E18 brain scRNA-Seq (Lebrigand et al.) data analysis input files for the Isosceles paper - more details can be found in the <a href="https://github.com/Genentech/Isosceles_Paper" target="_blank" rel="noopener">Isosceles_Paper</a> repository.</p>
Integrated PF-ILD scRNA-seq atlas
<p>Processed counts object of the integrated PF-ILD atlas from publicly available scRNA-seq datasets.</p>
scRNA-seq reveals an immune microenvironment and JUN-mediated NK cell exhaustion in relapsed T-ALL
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.