Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
894
datasets available to search
ShareScore release 0.7.1
Dataset results
894 results for “single-cell RNA-seq”
Recovery and analysis of transcriptome subsets from pooled single-cell RNA-seq libraries
<p>Processed data files for manuscript: "Recovery and analysis of transcriptome subsets from pooled single-cell RNA-seq libraries" <a href="https://doi.org/10.1093/nar/gky1204">https://doi.org/10.1093/nar/gky1204</a> . Scripts for generating figures are found here: https://github.com/rnabioco/scrna-subsets</p>
Time-resolved single-cell RNA-seq using scSLAM-seq and GRAND-SLAM
<p>This is the example data for the protocol "Time-resolved single-cell RNA-seq using scSLAM-seq and GRAND-SLAM".</p> <p>The test data comprises sequencing data from 12 different cells. Five of them in mock condition with 4sU, five infected with MCMV with 4sU and 2 uninfected cells without 4sU. For each sample, paired-end sequencing was conducted and therefore two files exist for each sample. The files can be found at data/fastq.</p> <p>The reads should be mapped against the mouse genome and the MCMV genome, which are located in data/genome. The sequence is in fasta-fomat and the features in gtf-format.The scripts conducting quality control, read mapping and GRAND-SLAM analysis can be found in the processing folder. For each computational step in this protocol there is a shell script for running the necessary steps. Executing all scripts will reproduce the analysis of our test data</p>
Single-cell RNA-Seq and TCR-Seq analysis of PD-1+ CD8+ T-cells responding to anti-PD-1 and anti-PD-1/CTLA-4 immunotherapy in melanoma
<p><strong>This dataset details the scRNASeq and TCR-Seq analysis of sorted PD-1+ CD8+ T cells from patients with melanoma treated with checkpoint therapy (anti-PD-1 monotherapy and anti-PD-1 & anti-CTLA-4 combination therapy) at baseline and after the first cycle of therapy. A major publication using this dataset is accessible here: (reference) </strong></p> <p> </p> <p><strong>*experimental design</strong></p> <p> Single-cell RNA sequencing was performed using 10x Genomics with feature barcoding technology to multiplex cell samples from different patients undergoing mono or dual therapy so that they can be loaded on one well to reduce costs and minimize technical variability. Hashtag oligomers (oligos) were obtained as purified and already oligo-conjugated in TotalSeq-C format from BioLegend. Cells were thawed, counted and 20 million cells per patient and time point were used for staining. Cells were stained with barcoded antibodies together with a staining solution containing antibodies against CD3, CD4, CD8, PD-1/IgG4 and fixable viability dye (eBioscience) prior to FACS sorting. Barcoded antibody concentrations used were 0.5 µg per million cells, as recommended by the manufacturer (BioLegend) for flow cytometry applications. After staining, cells were washed twice in PBS containing 2% BSA and 0.01% Tween 20, followed by centrifugation (300 xg 5 min at 4 °C) and supernatant exchange. After the final wash, cells were resuspended in PBS and filtered through 40 µm cell strainers and proceeded for sorting. Sorted cells were counted and approximately 75,000 cells were processed through 10x Genomics single-cell V(D)J workflow according to the manufacturer’s instructions. Gene expression, hashing and TCR libraries were pooled to desired quantities to obtain the sequencing depths of 15,000 reads per cell for gene expression libraries and 5,000 reads per cell for hashing and TCR libraries. Libraries were sequenced on a NovaSeq 6000 flow cell in a 2X100 paired-end format.</p> <p> </p> <p><strong>*extract protocol</strong></p> <p> PBMCs were thawed, counted and 20 million cells per patient and time point were used for staining. Cells were stained with barcoded antibodies together with a staining solution containing antibodies against CD3, CD4, CD8, PD-1/IgG4 and fixable viability dye (eBioscience) prior to FACS sorting. Barcoded antibody concentrations used were 0.5 µg per million cells, as recommended by the manufacturer (BioLegend) for flow cytometry applications. After staining, cells were washed twice in PBS containing 2% BSA and 0.01% Tween 20, followed by centrifugation (300 xg 5 min at 4 °C) and supernatant exchange. After the final wash, cells were resuspended in PBS and filtered through 40 µm cell strainers and proceeded for sorting. Sorted cells were counted and approximately 75,000 cells were processed through 10x Genomics single-cell V(D)J workflow according to the manufacturer’s instructions.</p> <p> </p> <p><strong>*library construction protocol</strong></p> <p> Sorted cells were counted and approximately 75,000 cells were processed through 10x Genomics single-cell V(D)J workflow according to the manufacturer’s instructions. Gene expression, hashing and TCR libraries were pooled to desired quantities to obtain the sequencing depths of 15,000 reads per cell for gene expression libraries and 5,000 reads per cell for hashing and TCR libraries. Libraries were sequenced on a NovaSeq 6000 flow cell in a 2X100 paired-end format.</p> <p> </p> <p><strong>*library strategy</strong></p> <p> scRNA-seq and scTCR-seq</p> <p> </p> <p><strong>*data processing step</strong></p> <p> Pre-processing of sequencing results to generate count matrices (gene expression and HTO barcode counts) was performed using the 10x genomics Cell Ranger pipeline.</p> <p> Further processing was done with Seurat (cell and gene filtering, hashtag identification, clustering, differential gene expression analysis based on gene expression).</p> <p> </p> <p> <strong>*genome build/assembly</strong></p> <p> Alignment was performed using prebuilt Cell Ranger human reference GRCh38.</p> <p> </p> <p><strong>*processed data files format and content</strong></p> <p> RNA counts and HTO counts are in sparse matrix format and TCR clonotypes are in csv format.</p> <p>Datasets were merged and analyzed by Seurat and the analyzed objects are in rds format.</p> <p> </p> <table> <tbody> <tr> <td> <p><strong>file name</strong></p> </td> <td> <p><strong>file checksum</strong></p> </td> </tr> <tr> <td> <p>PD1CD8_160421_filtered_feature_bc_matrix.zip</p> </td> <td> <p>da2e006d2b39485fd8cf8701742c6d77</p> </td> </tr> <tr> <td> <p>PD1CD8_190421_filtered_feature_bc_matrix.zip</p> </td> <td> <p>e125fc5031899bba71e1171888d78205</p> </td> </tr> <tr> <td> <p>PD1CD8_160421_filtered_contig_annotations.csv</p> </td> <td> <p>927241805d507204fbe9ef7045d0ccf4</p> </td> </tr> <tr> <td> <p>PD1CD8_190421_filtered_contig_annotations.csv</p> </td> <td> <p>8ca544d27f06e66592b567d3ab86551e</p> </td> </tr> </tbody> </table> <p> </p> <table> <tbody> <tr> <td> <p><strong>*processed data file </strong></p> </td> <td> <p><strong>antibodies/tags</strong></p> </td> </tr> <tr> <td> <p>PD1CD8_160421_filtered_feature_bc_matrix.zip</p> </td> <td> <p>none</p> </td> </tr> <tr> <td> <p>PD1CD8_160421_filtered_feature_bc_matrix.zip</p> </td> <td> <p>TotalSeq™-C0251 anti-human Hashtag 1 Antibody - (HASH_1) - M1_base_monotherapy<br>TotalSeq™-C0252 anti-human Hashtag 2 Antibody - (HASH_2) - M1_post_monotherapy<br>TotalSeq™-C0253 anti-human Hashtag 3 Antibody - (HASH_3) - C1_base_combined_therapy<br>TotalSeq™-C0254 anti-human Hashtag 4 Antibody - (HASH_4) - C1_post_combined_therapy<br>TotalSeq™-C0255 anti-human Hashtag 5 Antibody - (HASH_5) - C2_base_combined_therapy<br>TotalSeq™-C0256 anti-human Hashtag 6 Antibody - (HASH_6) - C2_post_combined_therapy</p> </td> </tr> <tr> <td> <p>PD1CD8_160421_filtered_contig_annotations.csv</p> </td> <td> <p>none</p> </td> </tr> <tr> <td> <p>PD1CD8_190421_filtered_feature_bc_matrix.zip</p> </td> <td> <p>none</p> </td> </tr> <tr> <td> <p>PD1CD8_190421_filtered_feature_bc_matrix.zip</p> </td> <td> <p>TotalSeq™-C0251 anti-human Hashtag 1 Antibody - (HASH_1) - M2_base_monotherapy<br>TotalSeq™-C0252 anti-human Hashtag 2 Antibody - (HASH_2) - M2_post_monotherapy<br>TotalSeq™-C0253 anti-human Hashtag 3 Antibody - (HASH_3) - M3_base_monotherapy<br>TotalSeq™-C0254 anti-human Hashtag 4 Antibody - (HASH_4) - M3_post_monotherapy<br>TotalSeq™-C0255 anti-human Hashtag 5 Antibody - (HASH_5) - C3_base_combined_therapy<br>TotalSeq™-C0256 anti-human Hashtag 6 Antibody - (HASH_6) - C3_post_combined_therapy</p> </td> </tr> <tr> <td> <p>PD1CD8_190421_filtered_contig_annotations.csv</p> </td> <td> <p>none</p> </td> </tr> </tbody> </table> <p> </p>
Encompassing view of spatial and single-cell RNA-seq renews the role of the microvasculature in human atherosclerosis
<p>Here we generate an integrative high-resolution map of human atherosclerotic plaques combining single-cell RNA-seq from multiple studies and spatial transcriptomics data from 12 human specimens, with different stages of atherosclerosis. We show cell-type and atherosclerosis-specific expression changes and spatially constrained alterations in cell-cell communication.</p>
Single-cell RNA-seq count data used in differential expression benchmark study
<p>Count matrices and meta data tables from simulated and real world immune cell single-cell RNA-seq experiments.</p> <p>All files are in Rds format and can be read by R using "readRDS()". </p> <ul> <li>10k_*: These files contain a filtered version of the 10k Human PBMCs, 3' v3.1 data <a href="https://www.10xgenomics.com/resources/datasets/10k-human-pbmcs-3-v3-1-chromium-controller-3-1-high">published</a> by 10x Genomics</li> <li>blueprint_data.Rds: This file contains the bulk RNA-seq data downloaded from <a href="http://dcc.blueprint-epigenome.eu">BLUEPRINT</a></li> <li>blueprint_immune_comparisons.Rds: Results from running three bulk RNA-seq differential expression methods</li> <li>sim_data*: These files contain the count matrices and meta data tables for the simulated data. Every file contains a list of 13 replicates.</li> </ul> <p> </p>
Bulk and single-cell RNA-seq of human fetal pancreatic organoids
Open the record for dataset details and reuse information.
Single-cell RNA-seq of the rare virosphere reveals the native hosts of giant viruses in the marine environment
Open the record for dataset details and reuse information.
Single-cell RNA-seq of the embryonic zebrafish heads from wild-type siblings and betaPix CRISPR mutants at 1 dpf and 2 dpf
Open the record for dataset details and reuse information.
How to perform PCA on single-cell RNA-Seq data in three simple steps
<p>Video on YouTube: <a href="https://www.youtube.com/watch?v=IOe0X-q7FtE">https://www.youtube.com/watch?v=IOe0X-q7FtE</a></p> <p>My Twitter: <a href="https://twitter.com/flo_compbio">https://twitter.com/flo_compbio</a></p> <p>Savannah Bertrand’s fundraiser: <a href="https://www.gofundme.com/f/help-a-black-lesbian-academic-get-out?utm_source=twitter&utm_medium=social&utm_campaign=p_cf+share-flow-1">https://www.gofundme.com/f/help-a-black-lesbian-academic-get-out?utm_source=twitter&utm_medium=social&utm_campaign=p_cf+share-flow-1</a></p> <p>-------------------<br> References:</p> <p>Batson, Joshua, Loïc Royer, and James Webber. “Molecular Cross-Validation for Single-Cell RNA-Seq.” BioRxiv, September 30, 2019, 786269. <a href="https://doi.org/10.1101/786269">https://doi.org/10.1101/786269</a>.</p> <p>Grün, Dominic, Lennart Kester, and Alexander van Oudenaarden. “Validation of Noise Models for Single-Cell Transcriptomics.” Nature Methods 11, no. 6 (June 2014): 637–40.<a href="https://doi.org/10.1038/nmeth.2930"> https://doi.org/10.1038/nmeth.2930</a>.</p> <p>Hafemeister, Christoph, and Rahul Satija. “Normalization and Variance Stabilization of Single-Cell RNA-Seq Data Using Regularized Negative Binomial Regression.” Genome Biology 20, no. 1 (23 2019): 296. <a href="https://doi.org/10.1186/s13059-019-1874-1">https://doi.org/10.1186/s13059-019-1874-1</a>.</p> <p>Hsu, Lauren L., and Aedin C. Culhane. “Impact of Data Preprocessing on Integrative Matrix Factorization of Single Cell Data.” Frontiers in Oncology 10 (2020). <a href="https://doi.org/10.3389/fonc.2020.00973">https://doi.org/10.3389/fonc.2020.00973</a>.</p> <p>Sun, Shiquan, Jiaqiang Zhu, Ying Ma, and Xiang Zhou. “Accuracy, Robustness and Scalability of Dimensionality Reduction Methods for Single-Cell RNA-Seq Analysis.” Genome Biology 20, no. 1 (10 2019): 269. <a href="https://doi.org/10.1186/s13059-019-1898-6">https://doi.org/10.1186/s13059-019-1898-6</a>.</p> <p>Townes, F. William, Stephanie C. Hicks, Martin J. Aryee, and Rafael A. Irizarry. “Feature Selection and Dimension Reduction for Single-Cell RNA-Seq Based on a Multinomial Model.” Genome Biology 20, no. 1 (23 2019): 295. <a href="https://doi.org/10.1186/s13059-019-1861-6">https://doi.org/10.1186/s13059-019-1861-6</a>.</p> <p>Tsuyuzaki, Koki, Hiroyuki Sato, Kenta Sato, and Itoshi Nikaido. “Benchmarking Principal Component Analysis for Large-Scale Single-Cell RNA-Sequencing.” Genome Biology 21, no. 1 (20 2020): 9. <a href="https://doi.org/10.1186/s13059-019-1900-3">https://doi.org/10.1186/s13059-019-1900-3</a>.</p> <p>Wagner, Florian, Dalia Barkley, and Itai Yanai. “Accurate Denoising of Single-Cell RNA-Seq Data Using Unbiased Principal Component Analysis.” BioRxiv, June 17, 2019, 655365.<a href="https://doi.org/10.1101/655365"> https://doi.org/10.1101/655365</a>.</p> <p>Wagner, Florian. “Monet: An Open-Source Python Package for Analyzing and Integrating ScRNA-Seq Data Using PCA-Based Latent Spaces.” BioRxiv, 2020. <a href="https://doi.org/10.1101/2020.06.08.140673">https://doi.org/10.1101/2020.06.08.140673</a>.</p>
BecomingLTi - Dataset : RNA-seq from Embryo Periphery and Fetal Liver at stage 13.5 and Single-cell RNA-seq for Embryo Periphery at stage 13.5 and 14.5
<p><strong>Dataset from article</strong> : Distinct waves from the hemogenic endothelium give rise to layered Lymphoid Tissue Inducer cell ontogeny</p> <p><strong>Summary:</strong> During embryogenesis Lymphoid Tissue Inducer (LTi) cells are essential for lymph node organogenesis. These cells are part of the Innate Lymphoid Cell (ILC) family. Although their earliest embryonic hematopoietic origin is unclear, other innate immune cells were shown to be derived from both early hemogenic endothelium in the yolk-sac as well as the aorta-gonad-mesonephros. A proper model to discriminate between these locations was unavailable. In this study, using a new Cxcr4-CreERT2 lineage tracing model, we identify a major contribution from embryonic hemogenic endothelium, but not yolk-sac, towards the LTi progenitors. Conversely, embryonic LTi cells are replaced by hematopoietic stem cell derived cells in adult. We further show that within the fetal liver common lymphoid progenitors differentiate into highly dynamic alpha-lymphoid precursor cells, which at this embryonic stage preferentially mature into LTi precursors and establish their functional LTi cell identity only after reaching the periphery.</p> <p><strong>Data </strong>:</p> <p>1. SPlab_BecomingLTi_Bulk_Stage13.5_2tissues_00_RawData : Bulk RNA-seq data for Mouse Embryo Periphery and Fetal Liver at Stage 13.5. It contains matrix of gene expression counts for each sample in the dataset</p> <p>2. SPlab_BecomingLTi_Stage13.5_Periphery_CellRangerV3_00_RawData : Single-cell RNA-seq data for Mouse Embryo Periphery at stage 13.5. It contains the full output of CellRanger count (v3) analysis.</p> <p>3. SPlab_BecomingLTi_Stage14.5_Periphery_CellRangerV3_00_RawData : Single-cell RNA-seq data for Mouse Embryo Periphery at stage 14.5. It contains the full output of CellRanger count (v3) analysis.</p>
Becoming LTi - Dataset : Single-cell RNA-seq of Embryo Fetal Liver tissue at stage 13.5 days
<p><strong>Dataset from article</strong> : Distinct waves from the hemogenic endothelium give rise to layered Lymphoid Tissue Inducer cell ontogeny</p> <p><strong>Summary:</strong> During embryogenesis Lymphoid Tissue Inducer (LTi) cells are essential for lymph node organogenesis. These cells are part of the Innate Lymphoid Cell (ILC) family. Although their earliest embryonic hematopoietic origin is unclear, other innate immune cells were shown to be derived from both early hemogenic endothelium in the yolk-sac as well as the aorta-gonad-mesonephros. A proper model to discriminate between these locations was unavailable. In this study, using a new Cxcr4-CreERT2 lineage tracing model, we identify a major contribution from embryonic hemogenic endothelium, but not yolk-sac, towards the LTi progenitors. Conversely, embryonic LTi cells are replaced by hematopoietic stem cell derived cells in adult. We further show that within the fetal liver common lymphoid progenitors differentiate into highly dynamic alpha-lymphoid precursor cells, which at this embryonic stage preferentially mature into LTi precursors and establish their functional LTi cell identity only after reaching the periphery.</p> <p><strong>Data </strong>:</p> <p>1. SPlab_BecomingLTi_Stage13.5_FetalLiver_00_RawData : Single-cell RNA-seq data for Mouse Embryo Fetal Liver at stage 13.5. It contains the full output of CellRanger count (v3) analysis.</p>
Clustering-independent estimation of cell abundances in bulk tissues using single-cell RNA-seq data
<p>ConDecon is a clustering-independent method for inferring the likelihood for each cell in a single-cell dataset to be present in a bulk tissue. This repository contains the raw data of the benchmarking analyses presented in the original publication using the pipeline of Avila-Cobos et al. (10.1038/s41467-020-19015-1). We used this pipeline to evaluate the ability of ConDecon and 17 other deconvolution methods to infer discrete cell type abundances in bulk tissues. The compressed file in this repository contains the synthetic bulk data, ground truth cell type proportions, and the predicted cell type proportions for each method and dataset associated with these analyses. Additional details can be found in the Methods section of the ConDecon publication.</p>
Teaching Datasets for Single-Cell RNA-seq Analysis Course
<p>This repository contains teaching datasets used in the Single-Cell RNA-seq Analysis Course, which is part of the SeuratExtend project (<a href="https://github.com/huayc09/SeuratExtend">https://github.com/huayc09/SeuratExtend</a>). The course materials are available at <a href="https://huayc09.github.io/SeuratExtend/#single-cell-rna-seq-analysis-course-new-in-v110-1">https://huayc09.github.io/SeuratExtend/#single-cell-rna-seq-analysis-course-new-in-v110-1</a>, with code and scripts hosted at <a href="https://github.com/huayc09/single-cell-course">https://github.com/huayc09/single-cell-course</a>.</p> <p>The datasets include:</p> <ol> <li>3k Peripheral Blood Mononuclear Cells (PBMCs)</li> <li>Paired PBMC samples processed with 10x Genomics' 3' kit</li> <li>Paired PBMC samples processed with 10x Genomics' 5' kit</li> </ol> <p>All original data were obtained from 10x Genomics (<a href="https://www.10xgenomics.com/resources/datasets">https://www.10xgenomics.com/resources/datasets</a>) and processed for educational purposes.</p>
Accuracy, robustness and scalability of dimensionality reduction methods for single-cell RNA-seq analysis
<p>A detailed list of the selected scRNA-seq datasets used in the paper, also provided in Additional file <a href="https://genomebiology.biomedcentral.com/articles/10.1186/s13059-019-1898-6#MOESM1">1</a>: Table S1-S2.</p>
single-cell RNA-seq of EHT process
<p>This file 'atfr.rds' contains onject including the raw and log-normalizd expression counts, and TF activity matrix generated from metaRegulon. In addition, the file 'cell_weight_pseudotime.rds' contains the pseudotime information.</p>
Single-cell RNA-seq supplementary data
<p>These files are supplementary data for the article Laczik M, Erdős E, Ozgyin L, Hevessy Zs, Csősz É, Kalló G, Nagy T, Barta E, Póliska Sz, Szatmári I and Bálint BL: Extensive proteome and functional genomic profiling of variability between genetically identical human B-lymphoblastoid cells. The files are outputs from the 10x Genomics software Cell Ranger and Loupe Browser, they contain QC data for the single cell sequencing described in the article, and also comparative analyses between 3 datasets (GM22648old, GM22648new, GM22649). For further details please see the article.</p>
Single-cell RNA-seq datasets derived from human parasitic nematode Brugia malayi microfilariae
<p>Single-cell RNA-seq data of the parasitic nematode <em>Brugia malayi</em> in the microfilariae development stage. Data includes untreated cellular transcriptional states and the transcriptional response to ivermectin (1 µM). </p> <p>The unfiltered gene expression matrix is provided as both a Seurat and AnnData object. Filtered datasets are provided as .csv files for direct import into R. </p>
AsaruSim: a single-cell and spatial RNA-Seq Nanopore long-reads simulation workflow
Open the record for dataset details and reuse information.
scCTS: identifying the cell type-specific marker genes from population-level single-cell RNA-seq
<p>Single cell RNA-sequencing (scRNA-seq) provides gene expression profiles of individual cells from complex samples, facilitating the detection of cell type-specific marker genes. In scRNA-seq experiments with multiple donors, the population level variation brings an extra layer of complexity in cell type-specific gene detection, for example, they may not appear in all donors. Motivated by this observation, we develop a statistical model named scCTS to identify cell type-specific genes from population-level scRNA-seq data. Extensive data analyses demonstrate that the proposed method identifies more biologically meaningful cell type-specific genes compared to traditional methods.</p>
Processed Single-cell RNA-seq Data for Exercise Rejuvenation Intervention on Mouse Subventricular Zone
<p>Exercise single-cell RNA-seq data of mouse subventricular zone (in processed Seurat object) used in the publication "Cell type-specific aging clocks to quantify aging and rejuvenation in regenerative regions of the brain" (preprint available at <a href="https://doi.org/10.1101/2022.01.10.475747">https://doi.org/10.1101/2022.01.10.475747</a>).</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.