Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2,750
datasets available to search
ShareScore release 0.9.0
Dataset results
2,750 results for “scRNA seq”
Colorectal cancer scRNA-seq 10xG-format data matrix
<p>Metastatic colorectal cancer (CRC) is a major cause of cancer-related death and incidence is rising in the younger population (<50 years). Current chemotherapies can achieve response rates above 50%, but immunotherapies have limited value for patients with microsatellite-stable (MSS) cancers. The present study investigates the impact of chemotherapy on the tumor immune microenvironment. We treat human liver metastases slices with 5-Fluorouracil (5FU) plus either irinotecan or oxaliplatin, then perform single-cell transcriptome analyses. Results from eight cases reveal two cellular subtypes with divergent responses to chemotherapy. Susceptible tumors are characterized by a stemness signature, an activated interferon pathway, and suppression of PD-1 ligands in response to 5FU+irinotecan. Conversely, immune checkpoint TIM-3 ligands are maintained or up-regulated by chemotherapy in CRC with an enterocyte-like signature, and combining chemotherapy with TIM-3 blockade leads to synergistic tumor killing. Together, our analyses highlight chemo-modulation of the immune microenvironment and provide a framework for combined chemo-immunotherapies. </p>
scRNA-seq revealed the rules for CDR3 length pairing in TCR beta and alpha chains and BCR heavy and light chains
<p>The scRNAseq datasets of CDR3 length pairing in TCR beta and alpha chains which come from human cental and peripheral samples and mouse peripheral samples.</p> <p>The scRNAseq datasets of CDR3 length pairing in BCR heavy and light chainsCDR3 length pairing in TCR beta and alpha chains and BCR heavy and light chains human cental and peripheral samples and mouse cental and peripheral samples.</p> <p> </p>
scRNA-seq dataset "A novel in vitro tubular model to recapitulate features of distal airways: the bronchioid"
<p>We provide a .Rds file of an annotated Seurat Object of scRNA-seq data of two bronchioid models derived from distinct donors after 21days of culture using 10x genomics 3' v3 chemistry. Raw data was processed using CellRanger v7.1.0. Cells were filtered based on detected UMIs (>2000) and fraction of mitochondrial counts (<10%).<br>Metadata annotations contain:<br>- Patient -> patient information for every cell (patient1 or patient2)<br>- nCount_RNA -> UMI counts per cell<br>- nFeature_RNA -> genes detected per cell<br>- percent.mt -> mitochondrial count fraction per cell<br>- seurat_clusters -> unsupervised clustering results using Louvain algorithm with resolution = 0.5<br>- Manual.Annotation -> Cell types annotated based on marker gene expression<br>- Celltypist.prediction -> Cell types predicted with CellTypist Python package<br>- Celltypist.prediction.ari -> Cell types predicted with CellTypist Python package, with harmonized names for comparison with manual annotation</p>
LungMAP Azimuth Reference - Human Adult Lung scRNA-Seq
<p>The LungMAP scRNA-Seq reference associated with the Lung CellCards resource. The initial reference integrated 259k cells from 72 donors from five published (PMIDs: 32726565, 32427931, 30554520, 32832599, 32832598) and one unpublished single cell RNA-seq cohort. Non-diseased adult and pediatric healthy lung single-cell 10x Genomics captures (3’ v2 and v3). Cells from different donors were integrated using Batchlor. Preliminary cell types were called based on Leiden clustering analysis and expression patterns of LungMAP cell card markers. UMAP embeddings were generated using monocle3. This reference, along with a corresponding single-nucleus specific version of this atlas, is under active construction. We expect to release the beta version of the reference in November 2021. Conforms to<strong> </strong>Azimuth reference data structure described at <a href="https://github.com/satijalab/azimuth/wiki/Azimuth-Reference-Format">https://github.com/satijalab/azimuth/wiki/Azimuth-Reference-Format</a>.</p>
moFluMemB - Dataset : scRNA-seq from Lymph node, Spleen and Lung
<p><strong>Title</strong></p> <p>Viral infection engenders bona fide and bystander subsets of lung-resident memory B cells through a permissive mechanism<br><br><strong>Authors</strong><br>Claude Gregoire,1 Lionel Spinelli,1 Sergio Villazala-Merino,1 Laurine Gil,1 María Pía Holgado,1 Myriam Moussa,1 Chuang Dong,1 Ana Zarubica,2 Mathieu Fallet,1 Jean-Marc Navarro,1 Bernard Malissen,1,2 Pierre Milpied,1,* and Mauro Gaya1,*<br><br><strong>Affiliations</strong><br>1 Centre d'Immunologie de Marseille-Luminy (CIML), Aix Marseille Université, INSERM, CNRS, Marseille, France<br>2 Centre d'Immunophénomique (CIPHE), Aix Marseille Université, INSERM, CNRS, Marseille, France<br>* Correspondence: milpied@ciml.univ-mrs.fr (P.M.), gaya@ciml.univ-mrs.fr (M.G.)<br><br><strong>Summary</strong><br>Lung-resident memory B cells (MBCs) provide localized protection against reinfection in the respiratory airways. Currently, the biology of these cells remains largely unexplored. Here, we combined influenza and SARS-CoV-2 infection with fluorescent-reporter mice to identify MBCs regardless of antigen specificity. We found that two main transcriptionally distinct subsets of MBCs colonized the lung peribronchial niche after infection. These subsets arose from different progenitors and were both class-switched, somatically mutated and intrinsically biased in their differentiation fate towards plasma cells. Combined analysis of antigen-specificity and B cell receptor repertoire segregated these subsets into “bona fide” virus-specific MBCs and “bystander” MBCs with no apparent specificity for eliciting viruses and generated through an alternative permissive mechanism. Thus, diverse transcriptional programs in MBCs are not linked to specific effector fates but rather to divergent strategies of the immune system to simultaneously provide rapid protection from reinfection while diversifying the initial B cell repertoire.</p> <p><strong>Data</strong></p> <ul> <li>custom_201216_m_moFluMemB_processedData.tar.gz : pre-processed data of FB5P-seq protocol (Attaf et al., 2020) on memory B cells sorted from single-cell suspensions of lungs with enzymatic digestion of lung tissue at 37°C, with index sorting information for a panel of antibodies identifying subsets of memory B cells.</li> <li>moFluMemB_DockerImages.tar.gz: Docker images used by the analysis</li> <li>moFluMemB_SingularityImages.tar.gz: Singularity images used by the analysis (conversion of the docker images)<br> </li> </ul> <p>See the three other Zenodo deposit for the rest of the data:</p> <p><strong>10.5281/zenodo.5565863</strong></p> <p><strong>10.5281/zenodo.5564624</strong></p> <p><strong>10.5281/zenodo.10559312</strong></p>
Colorectal cancer interleukin-10 blockade scRNA-seq
<p><em>Objective:</em> PD-1 checkpoint inhibition and adoptive cellular therapy have limited success in patients with microsatellite stable colorectal cancer liver metastases (CRLM). We demonstrate that interleukin-10 (IL-10) blockade enhances endogenous T cell and chimeric antigen receptor T (CAR-T) cell anti-tumor function in CRLM slice cultures.<br><br><em>Design:</em> We created organotypic slice cultures from human CRLM (n = 38) and tested the anti-tumor effects of a neutralizing antibody against IL-10 (αIL-10). We evaluated slice cultures with single and multiplex immunohistochemistry, in situ hybridization, single cell RNA sequencing, and time-lapse fluorescent microscopy. In addition, we studied the effects of αIL-10 on carcinoembryonic antigen (CEA)-specific CAR-T cells exogenously administered to both human CRLM slice cultures and a CRLM murine model. <br><br><em>Results: </em>There was little effect of PD-1 blockade in CRLM slice cultures. In contrast, αIL-10 generated 1.8-fold increase in T cell-mediated carcinoma cell death, and increased proportion of CD8+ T cells and inflammatory polarization of macrophages. In addition to effects on endogenous immune cells in human CRLM, αIL-10 also rescued murine CAR-T cell proliferation and cytotoxicity from myeloid cell-mediated immunosuppression. In human CRLM slices, αIL-10 dramatically improved CEA-specific CAR-T cell cytotoxicity, generating nearly 70% carcinoma apoptosis across multiple human tumors. We saw a less dramatic, but similar effect of pretreatment of CAR-T cells with an IL-10 receptor blocking antibody, demonstrating that IL-10 inhibits CAR-T function in the CRLM tumor microenvironment.</p> <p><em>Conclusion:</em> Neutralizing the effects of IL-10 in human CRLM has therapeutic potential as a stand-alone treatment and to augment the function of adoptively transferred CAR-T cells.</p>
Supplementary table of PRIDE datasets analyzed for "FAVA: High-quality functional association networks inferred from scRNA-seq and proteomics data"
<p>Our proteomics dataset comes from The PRoteomics IDEntifications (PRIDE) database, the world’s largest data repository of mass spectrometry-based proteomics data. Specifically, we used 633 human proteomics project experiments with a total of 32,546 runs and reanalyzed them using ionbot with an FDR threshold of 0.01 [16], resulting in a total of 154,885,151 peptide spectrum matches for 18,846 proteins. Here is the full list of projects, runs, and general statistics.</p>
moFluMemB - Dataset : scRNA-seq from Lymph node, Spleen and Lung - 10x_191105_m_moFluMemB
<p><strong>Title</strong></p> <p>Viral infection engenders bona fide and bystander subsets of lung-resident memory B cells through a permissive mechanism<br><br><strong>Authors</strong><br>Claude Gregoire,1 Lionel Spinelli,1 Sergio Villazala-Merino,1 Laurine Gil,1 María Pía Holgado,1 Myriam Moussa,1 Chuang Dong,1 Ana Zarubica,2 Mathieu Fallet,1 Jean-Marc Navarro,1 Bernard Malissen,1,2 Pierre Milpied,1,* and Mauro Gaya1,*<br><br><strong>Affiliations</strong><br>1 Centre d'Immunologie de Marseille-Luminy (CIML), Aix Marseille Université, INSERM, CNRS, Marseille, France<br>2 Centre d'Immunophénomique (CIPHE), Aix Marseille Université, INSERM, CNRS, Marseille, France<br>* Correspondence: milpied@ciml.univ-mrs.fr (P.M.), gaya@ciml.univ-mrs.fr (M.G.)<br><br><strong>Summary</strong><br>Lung-resident memory B cells (MBCs) provide localized protection against reinfection in the respiratory airways. Currently, the biology of these cells remains largely unexplored. Here, we combined influenza and SARS-CoV-2 infection with fluorescent-reporter mice to identify MBCs regardless of antigen specificity. We found that two main transcriptionally distinct subsets of MBCs colonized the lung peribronchial niche after infection. These subsets arose from different progenitors and were both class-switched, somatically mutated and intrinsically biased in their differentiation fate towards plasma cells. Combined analysis of antigen-specificity and B cell receptor repertoire segregated these subsets into “bona fide” virus-specific MBCs and “bystander” MBCs with no apparent specificity for eliciting viruses and generated through an alternative permissive mechanism. Thus, diverse transcriptional programs in MBCs are not linked to specific effector fates but rather to divergent strategies of the immune system to simultaneously provide rapid protection from reinfection while diversifying the initial B cell repertoire.</p> <p><strong>Data:</strong> 10x_191105_m_moFluMemB_processedData.tar.gz : pre-processed data of 10x 5’ scRNA-Seq on memory B cells sorted from single-cell suspensions of spleen, lymph nodes and lungs with mechanical dissociation of lung tissue at 4°C.</p> <p>See the three other Zenodo deposit for the rest of the data:</p> <p><strong>10.5281/zenodo.5566674</strong></p> <p><strong>10.5281/zenodo.5564624</strong></p> <p><strong>10.5281/zenodo.10559312</strong></p>
moFluMemB - Dataset : scRNA-seq from Lymph node, Spleen and Lung - 10x_190712_m_moFluMemB
<p><strong>Title</strong></p> <p>Viral infection engenders bona fide and bystander subsets of lung-resident memory B cells through a permissive mechanism<br><br><strong>Authors</strong><br>Claude Gregoire,1 Lionel Spinelli,1 Sergio Villazala-Merino,1 Laurine Gil,1 María Pía Holgado,1 Myriam Moussa,1 Chuang Dong,1 Ana Zarubica,2 Mathieu Fallet,1 Jean-Marc Navarro,1 Bernard Malissen,1,2 Pierre Milpied,1,* and Mauro Gaya1,*<br><br><strong>Affiliations</strong><br>1 Centre d'Immunologie de Marseille-Luminy (CIML), Aix Marseille Université, INSERM, CNRS, Marseille, France<br>2 Centre d'Immunophénomique (CIPHE), Aix Marseille Université, INSERM, CNRS, Marseille, France<br>* Correspondence: milpied@ciml.univ-mrs.fr (P.M.), gaya@ciml.univ-mrs.fr (M.G.)<br><br><strong>Summary</strong><br>Lung-resident memory B cells (MBCs) provide localized protection against reinfection in the respiratory airways. Currently, the biology of these cells remains largely unexplored. Here, we combined influenza and SARS-CoV-2 infection with fluorescent-reporter mice to identify MBCs regardless of antigen specificity. We found that two main transcriptionally distinct subsets of MBCs colonized the lung peribronchial niche after infection. These subsets arose from different progenitors and were both class-switched, somatically mutated and intrinsically biased in their differentiation fate towards plasma cells. Combined analysis of antigen-specificity and B cell receptor repertoire segregated these subsets into “bona fide” virus-specific MBCs and “bystander” MBCs with no apparent specificity for eliciting viruses and generated through an alternative permissive mechanism. Thus, diverse transcriptional programs in MBCs are not linked to specific effector fates but rather to divergent strategies of the immune system to simultaneously provide rapid protection from reinfection while diversifying the initial B cell repertoire.</p> <p><strong>Data: 1</strong>0x_190712_m_moFluMemB_processedData.tar.gz : pre-processed data of 10x 5’ scRNA-Seq on memory B cells sorted from single-cell suspensions of spleen, lymph nodes and lungs with enzymatic digestion of lung tissue at 37°C.</p> <p>See the three other Zenodo deposit for the rest of the data:</p> <p><strong>10.5281/zenodo.5566674</strong></p> <p><strong>10.5281/zenodo.5565863</strong></p> <p><strong>10.5281/zenodo.10559312</strong></p>
Data from: Clustering Deviation Index (CDI): A robust and accurate internal measure for evaluating scRNA-seq data clustering
<div> <div> <p>The clustering of cells has been widely used to explore the heterogeneity of cell populations in single-cell RNA-sequencing (scRNA-seq). We proposed a parametric model for monoclonal and polyclonal scRNA-seq data to evaluate clustering results. Based on the parametric model, we proposed a metric (CDI) to quantify the goodness-of-fit of cell clustering to the data. Here we presented CT26.WT and T-CELL as two datasets to examine the performance of our model and metric. CT26.WT contains wild-type CT26 cells from the murine colorectal carcinoma cell line, and cells in CT26.WT are highly homogeneous. T-CELL contains T-cells from tumor tissue of mice three weeks after 4T1 tumor injection. From these datasets and public datasets, we validated our model and benchmarked our metric.</p> </div> </div>
Integrated and annotated human placenta scRNA-seq matrices
Open the record for dataset details and reuse information.
Automated cell annotation in scRNA-seq data using unique marker gene sets
<p>Single-cell RNA sequencing has revolutionized the study of cellular heterogeneity, yet accurate cell type annotation remains a significant challenge. Inconsistent labels, technological variability, and limitations in transferring annotations from reference datasets hinder precise annotation. This study presents a novel approach for accurate cell type annotation in scRNA-seq data using unique marker gene sets. By manually curating cell type names and markers from 280 publications, we verified marker expression profiles across these datasets and unified nomenclatures to consistently identify 166 cell types and subtypes. Our customized algorithm, which builds on the AUCell method, achieves accurate cell labeling at single-cell resolution and surpasses the performance of reference-based tools like Azimuth, especially in distinguishing closely related subtypes. To enhance accessibility and practical utility for researchers, we have also developed a user-friendly application that automates the cell typing process, enabling efficient verification and supporting comprehensive downstream analyses. The desktop application can be accessed at <a href="https://omnibusx.com/">https://omnibusx.com</a>.</p>
The Hitchhiker's Guide to scRNA-seq course
<p>This repository comprises the intermediate results, data, and the respective script to create it, to be use on the second day of the course <a href="https://www.medicina.ulisboa.pt/en/hitchhikers-guide-scrna-seq" target="_blank" rel="noopener">The Hitchhiker's Guide to scRNA-seq</a> (08-12/07/2024, iMM, Lisbon, Portugal), focused on integration. </p> <p>File description: </p> <ul> <li>data: <ul> <li><strong>pbmcref.rds</strong>: a Seurat R object of a reference of PBMCs retrieved from the R package SeuratData (v.0.2.2.9001)</li> <li><strong>pbmc3k_panc8.rds</strong>: a Seurat R object of two data sets - 3k human PBMCs from 10X Genomics and pancreatic islets from indrop1 - retrieved from SeuratData package (v.0.2.2.9001) </li> <li><strong>jurkat.rds</strong>: a Seurat R object comprising three data sets - Jurkat, HEK293T and Jurkat:HEK293T (50:50) - retrieved from 10X genomics and published by <a href="https://doi.org/10.1038/ncomms14049" target="_blank" rel="noopener">Zheng et al., 2017</a></li> <li><strong>ifnb.rds</strong>: a Seurat R object of two human PBMCs data sets - resting/control and interferon-stimulated - retrieved from the R package SeuratData (v.0.2.2.9001)</li> <li><strong>covid.rds</strong>: a Seurat R object of a COVID-19 PBMCs data set from <a href="https://doi.org/10.1038/s41467-020-17834-w" target="_blank" rel="noopener">Guo et al., 2020</a> retrieved from <a href="https://cellxgene.cziscience.com/e/ae5341b8-60fb-4fac-86db-86e49ee66287.cxg" target="_blank" rel="noopener">cziscience</a></li> </ul> </li> <li><strong>01_create_datasets.R</strong>: R script used to retrieve and parse all the Seurat R objects mentioned above</li> <li><strong>results.zip</strong>: compressed folder with intermediate results used for the hands-on exercises</li> </ul>
Sampling time-dependent artifacts in single-cell genomics studies: scRNA-seq data
<p>Robust protocols and automation now enable large-scale single-cell RNA and ATAC sequencing experiments and their application on biobank and clinical cohorts. However, technical biases introduced during sample acquisition can hinder solid, reproducible results, and a systematic benchmarking is required before entering large-scale data production. Here, we report the existence and extent of gene expression and chromatin accessibility artifacts introduced during sampling and identify experimental and computational solutions for their prevention.</p> <p>This repository contains the expression matrices and Seurat objects associated with the scRNA-seq data of the manuscript: "Sampling time-dependent artifacts in single-cell genomics studies" published in Genome Biology in 2020. The purpose of this repo is to share processed files and metadata for immediate access and reproducibility. The code to analyze it is thoroughly documented at the associated Github repository (https://github.com/massonix/sampling_artifacts).</p>
scNCL transfers labels from scRNA-seq to scATAC-seq data with neighborhood contrastive regularization
<p>Data used in our manuscript.</p> <p>'scNCL_data' folder contains gene expression data (with protein) and gene activitiy data (with protein).</p> <p>'GLUE_data' folder contains gene expression data and raw chromatin accessibility data. </p> <p>'HFA-resampling' folder contains multiple independent sampling of HFA-subset dataset.</p> <p>'Accuray on HFA-subsets.xlsx' file contains accuray results of all compared methods on HFA-subsets datasets.</p>
CAKE: clustering scRNA-seq data via combining contrastive learning with knowledge distillation
<p>All datasets used in our paper "<strong>CAKE: clustering scRNA-seq data via combining contrastive learning with knowledge distillation"</strong></p>
Multi-subject simulated scRNA-seq trajectory data
<p>- Simulated scRNA-seq counts with a trajectory structure</p> <p>- Simulation was performed using <a href="https://github.com/rhondabacher/scaffold/tree/master">the scaffold R package</a></p> <p>- Based on pancreas reference dataset from <a href="https://doi.org/10.1016/j.cels.2016.08.011">Baron et al (2017)</a></p> <p>- Data are comprised of 3 subjects with 400 cells apiece for a total of 1200 cells</p> <p>- Counts have been processed using a SingleCellExperiment-based workflow (normalization, dimension reduction, clustering)</p> <p>- Ground-truth pseudotime for each cell is stored in the colData slot in the variable <strong>cell_time_normed</strong></p>
Data from: Clustering Deviation Index (CDI): A robust and accurate internal measure for evaluating scRNA-seq data clustering
Open the record for dataset details and reuse information.
Colorectal cancer interleukin-10 blockade scRNA-seq
Open the record for dataset details and reuse information.
Colorectal cancer scRNA-seq 10xG-format data matrix
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.