Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2,917
datasets available to search
ShareScore release 0.7.1
Dataset results
2,917 results for “scRNA”
moFluMemB - Dataset : scRNA-seq from Lymph node, Spleen and Lung - 10x_190712_m_moFluMemB
<p><strong>Title</strong></p> <p>Viral infection engenders bona fide and bystander subsets of lung-resident memory B cells through a permissive mechanism<br><br><strong>Authors</strong><br>Claude Gregoire,1 Lionel Spinelli,1 Sergio Villazala-Merino,1 Laurine Gil,1 María Pía Holgado,1 Myriam Moussa,1 Chuang Dong,1 Ana Zarubica,2 Mathieu Fallet,1 Jean-Marc Navarro,1 Bernard Malissen,1,2 Pierre Milpied,1,* and Mauro Gaya1,*<br><br><strong>Affiliations</strong><br>1 Centre d'Immunologie de Marseille-Luminy (CIML), Aix Marseille Université, INSERM, CNRS, Marseille, France<br>2 Centre d'Immunophénomique (CIPHE), Aix Marseille Université, INSERM, CNRS, Marseille, France<br>* Correspondence: milpied@ciml.univ-mrs.fr (P.M.), gaya@ciml.univ-mrs.fr (M.G.)<br><br><strong>Summary</strong><br>Lung-resident memory B cells (MBCs) provide localized protection against reinfection in the respiratory airways. Currently, the biology of these cells remains largely unexplored. Here, we combined influenza and SARS-CoV-2 infection with fluorescent-reporter mice to identify MBCs regardless of antigen specificity. We found that two main transcriptionally distinct subsets of MBCs colonized the lung peribronchial niche after infection. These subsets arose from different progenitors and were both class-switched, somatically mutated and intrinsically biased in their differentiation fate towards plasma cells. Combined analysis of antigen-specificity and B cell receptor repertoire segregated these subsets into “bona fide” virus-specific MBCs and “bystander” MBCs with no apparent specificity for eliciting viruses and generated through an alternative permissive mechanism. Thus, diverse transcriptional programs in MBCs are not linked to specific effector fates but rather to divergent strategies of the immune system to simultaneously provide rapid protection from reinfection while diversifying the initial B cell repertoire.</p> <p><strong>Data: 1</strong>0x_190712_m_moFluMemB_processedData.tar.gz : pre-processed data of 10x 5’ scRNA-Seq on memory B cells sorted from single-cell suspensions of spleen, lymph nodes and lungs with enzymatic digestion of lung tissue at 37°C.</p> <p>See the three other Zenodo deposit for the rest of the data:</p> <p><strong>10.5281/zenodo.5566674</strong></p> <p><strong>10.5281/zenodo.5565863</strong></p> <p><strong>10.5281/zenodo.10559312</strong></p>
Data from: Clustering Deviation Index (CDI): A robust and accurate internal measure for evaluating scRNA-seq data clustering
<div> <div> <p>The clustering of cells has been widely used to explore the heterogeneity of cell populations in single-cell RNA-sequencing (scRNA-seq). We proposed a parametric model for monoclonal and polyclonal scRNA-seq data to evaluate clustering results. Based on the parametric model, we proposed a metric (CDI) to quantify the goodness-of-fit of cell clustering to the data. Here we presented CT26.WT and T-CELL as two datasets to examine the performance of our model and metric. CT26.WT contains wild-type CT26 cells from the murine colorectal carcinoma cell line, and cells in CT26.WT are highly homogeneous. T-CELL contains T-cells from tumor tissue of mice three weeks after 4T1 tumor injection. From these datasets and public datasets, we validated our model and benchmarked our metric.</p> </div> </div>
Integrated and annotated human placenta scRNA-seq matrices
Open the record for dataset details and reuse information.
Automated cell annotation in scRNA-seq data using unique marker gene sets
<p>Single-cell RNA sequencing has revolutionized the study of cellular heterogeneity, yet accurate cell type annotation remains a significant challenge. Inconsistent labels, technological variability, and limitations in transferring annotations from reference datasets hinder precise annotation. This study presents a novel approach for accurate cell type annotation in scRNA-seq data using unique marker gene sets. By manually curating cell type names and markers from 280 publications, we verified marker expression profiles across these datasets and unified nomenclatures to consistently identify 166 cell types and subtypes. Our customized algorithm, which builds on the AUCell method, achieves accurate cell labeling at single-cell resolution and surpasses the performance of reference-based tools like Azimuth, especially in distinguishing closely related subtypes. To enhance accessibility and practical utility for researchers, we have also developed a user-friendly application that automates the cell typing process, enabling efficient verification and supporting comprehensive downstream analyses. The desktop application can be accessed at <a href="https://omnibusx.com/">https://omnibusx.com</a>.</p>
Processed scRNA and scATAC data in DRCTdb
<p><span>Understanding the molecular mechanisms underlying </span><span>genetic</span><span> diseases is challenging due to the involvement of both environmental and genetic factors. Genome-wide association studies (GWAS) have identified numerous genetic loci, but their functional implications remain largely unknown. Single-cell multiomics sequencing has emerged as a powerful tool to study disease-specific cell types and their relationship with genetic variants. However, there is a lack of comprehensive databases for exploring genetic disease-related cell types and their mechanisms across different human tissues. In this study, we present the disease-related cell type database (DRCTdb), a database that integrates GWAS data and single-cell multiomics data to identify disease-related cell types and elucidate their regulatory mechanisms. DRCTdb contains well-processed single-cell multiomics data in 16 studies, encompassing transcriptome and epigenetic information overall 4 million cells within 28 tissues. Through DRCTdb, user can easily browse relationships and regulatory mechanisms between SNPs of 42 genetic disease and cell type in different human tissue based on GWAS and single cell multiomics data. Moreover, DRCTdb also provides data download,</span><span> which</span><span> allowing users to download well-processed <a name="_Int_HEP27DcN"></a>single-cell multiomics data and analysis result from DRCTdb</span></p>
The Hitchhiker's Guide to scRNA-seq course
<p>This repository comprises the intermediate results, data, and the respective script to create it, to be use on the second day of the course <a href="https://www.medicina.ulisboa.pt/en/hitchhikers-guide-scrna-seq" target="_blank" rel="noopener">The Hitchhiker's Guide to scRNA-seq</a> (08-12/07/2024, iMM, Lisbon, Portugal), focused on integration. </p> <p>File description: </p> <ul> <li>data: <ul> <li><strong>pbmcref.rds</strong>: a Seurat R object of a reference of PBMCs retrieved from the R package SeuratData (v.0.2.2.9001)</li> <li><strong>pbmc3k_panc8.rds</strong>: a Seurat R object of two data sets - 3k human PBMCs from 10X Genomics and pancreatic islets from indrop1 - retrieved from SeuratData package (v.0.2.2.9001) </li> <li><strong>jurkat.rds</strong>: a Seurat R object comprising three data sets - Jurkat, HEK293T and Jurkat:HEK293T (50:50) - retrieved from 10X genomics and published by <a href="https://doi.org/10.1038/ncomms14049" target="_blank" rel="noopener">Zheng et al., 2017</a></li> <li><strong>ifnb.rds</strong>: a Seurat R object of two human PBMCs data sets - resting/control and interferon-stimulated - retrieved from the R package SeuratData (v.0.2.2.9001)</li> <li><strong>covid.rds</strong>: a Seurat R object of a COVID-19 PBMCs data set from <a href="https://doi.org/10.1038/s41467-020-17834-w" target="_blank" rel="noopener">Guo et al., 2020</a> retrieved from <a href="https://cellxgene.cziscience.com/e/ae5341b8-60fb-4fac-86db-86e49ee66287.cxg" target="_blank" rel="noopener">cziscience</a></li> </ul> </li> <li><strong>01_create_datasets.R</strong>: R script used to retrieve and parse all the Seurat R objects mentioned above</li> <li><strong>results.zip</strong>: compressed folder with intermediate results used for the hands-on exercises</li> </ul>
Sampling time-dependent artifacts in single-cell genomics studies: scRNA-seq data
<p>Robust protocols and automation now enable large-scale single-cell RNA and ATAC sequencing experiments and their application on biobank and clinical cohorts. However, technical biases introduced during sample acquisition can hinder solid, reproducible results, and a systematic benchmarking is required before entering large-scale data production. Here, we report the existence and extent of gene expression and chromatin accessibility artifacts introduced during sampling and identify experimental and computational solutions for their prevention.</p> <p>This repository contains the expression matrices and Seurat objects associated with the scRNA-seq data of the manuscript: "Sampling time-dependent artifacts in single-cell genomics studies" published in Genome Biology in 2020. The purpose of this repo is to share processed files and metadata for immediate access and reproducibility. The code to analyze it is thoroughly documented at the associated Github repository (https://github.com/massonix/sampling_artifacts).</p>
scNCL transfers labels from scRNA-seq to scATAC-seq data with neighborhood contrastive regularization
<p>Data used in our manuscript.</p> <p>'scNCL_data' folder contains gene expression data (with protein) and gene activitiy data (with protein).</p> <p>'GLUE_data' folder contains gene expression data and raw chromatin accessibility data. </p> <p>'HFA-resampling' folder contains multiple independent sampling of HFA-subset dataset.</p> <p>'Accuray on HFA-subsets.xlsx' file contains accuray results of all compared methods on HFA-subsets datasets.</p>
CAKE: clustering scRNA-seq data via combining contrastive learning with knowledge distillation
<p>All datasets used in our paper "<strong>CAKE: clustering scRNA-seq data via combining contrastive learning with knowledge distillation"</strong></p>
Multi-subject simulated scRNA-seq trajectory data
<p>- Simulated scRNA-seq counts with a trajectory structure</p> <p>- Simulation was performed using <a href="https://github.com/rhondabacher/scaffold/tree/master">the scaffold R package</a></p> <p>- Based on pancreas reference dataset from <a href="https://doi.org/10.1016/j.cels.2016.08.011">Baron et al (2017)</a></p> <p>- Data are comprised of 3 subjects with 400 cells apiece for a total of 1200 cells</p> <p>- Counts have been processed using a SingleCellExperiment-based workflow (normalization, dimension reduction, clustering)</p> <p>- Ground-truth pseudotime for each cell is stored in the colData slot in the variable <strong>cell_time_normed</strong></p>
Data from: Clustering Deviation Index (CDI): A robust and accurate internal measure for evaluating scRNA-seq data clustering
Open the record for dataset details and reuse information.
Colorectal cancer interleukin-10 blockade scRNA-seq
Open the record for dataset details and reuse information.
Colorectal cancer scRNA-seq 10xG-format data matrix
Open the record for dataset details and reuse information.
SpaGE Spatial Gene Enhancement using scRNA-seq
<p>Spatial transcriptomics and scRNA-seq datasets using for integration and prediction of spatially non-measured genes using SpaGE</p>
Post-processed datasets for scRNA-seq clustering analysis in PPML-Omics
Open the record for dataset details and reuse information.
Isosceles paper simulated ovarian cell line ONT data (scRNA-Seq)
<div> <div> <p>Simulated ovarian cell line ONT data (scRNA-Seq) for the Isosceles paper - more details can be found in the <a href="https://github.com/Genentech/Isosceles_Paper" target="_blank" rel="noopener">Isosceles_Paper</a> repository.</p> </div> </div>
MTN deficient scRNA-seq
<p>"3KR" refers to scRNA-seq data from the intestinal tissues of mice fed a methionine-, tryptophan-, and niacin-deficient diet followed by recovery on a regular diet.<br>"3K" refers to scRNA-seq data from the intestines of mice continuously fed a methionine-, tryptophan-, and niacin-deficient diet.<br>"REG" represents scRNA-seq data from the intestinal tissues of mice maintained on a regular diet.<br>The file <em>mouse_combined.rds</em> contains the integrated dataset of these three conditions, processed using Seurat.</p>
processed scRNA-seq data for Neuwirth & Malzl et al. 2024
<p>This dataset contains the analysed scRNA-seq data used in Neuwirth & Malzl et al. 2024 as AnnData objects. You can find the code used to produce and analyse this data on <a href="https://github.com/menchelab/Neuwirth_Malzl_et_al_2024">GitHub</a>. All data was preprocessed using cellranger v6.0.1 with GRCh38 3.0.0 as reference.</p> <h3><strong>Psoriasis and Sarcoidosis PBMC data</strong></h3> <p><strong><em>pbmc.scps.integrated.clustered.h5ad</em></strong>: contains scVI-integrated and Leiden-clustered data of Psoriasis and Sarcoidosis patient blood as well as healthy controls (Psoriasis data was generated within this study; Sarcoidosis data was reprocessed from 10.1016/j.immuni.2023.01.014)<br><strong><em>tcells.pbmc.scps.integrated.clustered.h5ad</em></strong>: contains scVI-integrated and Leiden-clustered T cell subset of Psoriasis and Sarcoidosis PBMC data</p> <h3><strong>Atopic dermatitis skin data</strong></h3> <p><strong><em>tissue.ad.integrated.clustered.h5ad</em></strong>: contains scVI-integrated and Leiden-clustered data of healthy and atopic dermatitis patient skin (reprocessed from 10.1126/science.aba6500)<br><strong><em>tcells.tissue.ad.integrated.clustered.h5ad</em></strong>: contains scVI-integrated and Leiden-clustered T cell subset of atopic dermatitis data</p> <h3><strong>Psoriasis and Sarcoidosis skin data</strong></h3> <p><strong><em>tissue.scps.integrated.annotated.h5ad</em></strong>: contains scVI-integrated and celltypist-annotated data from Psoriasis and Sarcoidosis patient skin as well as healthy controls (Psoriasis data was reprocessed from 10.1126/science.aba6500; Sarcoidosis data was reprocessed from 10.1016/j.immuni.2023.01.014)<br><strong><em>tcells.tissue.scps.integrated.clustered.h5ad</em></strong>: contains scVI-integrated and Leiden-clustered T cell subset of Psoriasis and Sarcoidosis skin data<br><strong><em>tregs.tissue.scps.integrated.annotated.h5ad</em></strong>: contains scVI-integrated and SAT1 status annotated regulatory T cell subset of Psoriasis and Sarcoidosis skin data<br><em><strong>tregs.tissue.scps.integrated.milo.h5ad</strong></em>: basically same as above but with cell neighborhood overrepresentation analysis on top</p> <h3><strong>IBD colon data</strong></h3> <p><strong><em>tissue.uc.integrated.clustered.h5ad</em></strong>: contains scVI-integrated and Leiden-clustered data of Crohn's disease and ulcerative colitis patients as well as healthy controls (reprocessed from 10.1126/sciimmunol.abb4432)<br><strong><em>tcells.tissue.uc.integrated.clustered.h5ad</em></strong>: contains scVI-integrated and Leiden-clustered T cell subset of IBD data</p> <h3><strong>Raw and unfiltered data</strong></h3> <p><strong><em>inflammatory_disease.h5ad</em></strong>: contains the raw, unfiltered and unprocessed data of all the files above (i.e. combined cellranger output) and is the source data file of all analyses in this study. If you just want the untouched data this is what you want to use.</p>
PBMC scRNA-seq datasets measured using different 10X Chromium chemistries
<p>PBMC scRNA-seq datasets measured using different 10X Chromium chemistries<br> Obtained from: https://www.10xgenomics.com/resources/datasets</p>
scRNA-seq dataset of iPSC-derived pancreatic islet cells
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.