Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
341
datasets available to search
ShareScore release 0.7.1
Dataset results
341 results for “PBMC”
Kotliarov 2020 Vaccine Responsiveness PBMC dataset for Besca
<p>Kotliarov, Y., Sparks, R., Martins, A.J. <em>et al.</em> Broad immune activation underlies shared set point signatures for vaccine responsiveness in healthy individuals and disease activity in patients with lupus. <em>Nat Med</em> <strong>26, </strong>618–629 (2020). https://doi.org/10.1038/s41591-020-0769-8. We reprocessed the dataset using the Besca package (<a href="https://github.com/bedapub/besca">https://github.com/bedapub/besca</a>). The original gene expression data are available from <a href="https://doi.org/10.35092/yhjc.c.4753772">https://doi.org/10.35092/yhjc.c.4753772</a>.</p>
Suco - PBMC
<p>Suco (Single cell universal classification omnibus) is a large standardized reference dataset for cell type classification in single cell RNA sequencing data. Suco seeks to tackle the lack of standardized datasets for classification tasks in the single cell genomics field. In other fields of artificial intelligence, like computer vision, standardized datasets such as MNIST or ImageNet have transformed the development of powerful new machine learning methods.</p> <p>Suco features manual uniform standardized hierarchical cell type annotations in independently analyzed datasets. This collection of independent datasets ensures that machine learning classifiers can be tested using statistically independent data & labels. </p> <p>Here, we present the peripheral blood mononuclear cell dataset within Suco which includes > 1200 independent manual cell type cluster labels from 12 datasets totalling >500 individuals and >5 millions cells. </p> <p>______________________________________________________________________________________________________</p> <p><br><strong><em>Structure of the dataset:</em></strong></p> <p>.zip compressed folder containing datasets from individual studies which have been reprocessed, clustered and annotated independently by two different expert raters (human immunology)</p> <ul> <li>Filenames: DATASET_ID.h5ad <ul> <li>The dataset ID has the following format (each line followed by ‘-X-‘ separator <ul> <li>Tissue/cell type: here PBMC</li> <li>Disease context</li> <li>Publication year</li> <li>First author (optional: followed by _BATCHNAME)</li> <li>DOI (/ in DOI is replace by _ for compatibility with file systems</li> </ul> </li> <li>the .h5ad files have the following structure <ul> <li>load using the scanpy python package adata = sc.read(FILE_PATH)</li> <li><em>cell barcode (adata.obs_names)</em></li> <li>Study-ID + '-X-' + internal barcode</li> <li>adata.obs[‘sample_id’] <ul> <li>the sample ID should be the patient ID + '-X-' separator + internal sample ID <ul> <li>e.g. <em>TIL-X-BRCA-X-scRNAseq-X-Bassez-X-2021-X-10.1038_s41591-021-01323-8-X-2-X-Pre</em> <ul> <li>with <em>TIL-X-BRCA-X-scRNAseq-X-Bassez-X-2021-X-10.1038_s41591-021-01323-8-X-2 </em>being the patient ID</li> <li><em>-X-</em> the separator</li> </ul> </li> </ul> </li> </ul> </li> <li>adata.obs['patient_id'] <ul> <li>dataset id followed by an '-X-' separator and the internal patient id</li> <li>e.g. <em>TIL-X-BRCA-X-scRNAseq-X-Bassez-X-2021-X-10.1038_s41591-021-01323-8-X-35</em></li> <li><em>-X-</em> is the separator</li> <li><em>35</em> is the internal patient</li> </ul> </li> <li>adata.obs[cluster_final'] <ul> <li>final clustering used for the cell type annotation</li> <li>granularity can differ between subsets --> e.g. clustering from myeloid cells can originate from myeloid subset, clustering from TNK from TNK subset and epithelial from all leukocyte subset</li> <li>should be preceeded by the prefix used for subtyping e.g. 'TNK' for TNK cells followed by a '_' seperator and the cluster number:· </li> <li>e.g. <ul> <li>cluster 0 in TNK would be 'TNK_0'</li> <li>cluster 1 in M would be 'M_1'</li> </ul> </li> </ul> </li> <li>adata.obs[cluster_all'] <ul> <li>containing coarse clustering format 'all_CLUSTERNUMBER' </li> </ul> </li> <li>adata.obs[‘annotation’] <ul> <li>Most granular annotation based on adata.obs[‘cluster_final’]</li> </ul> </li> <li>adata.obs[‘annotation_all’] <ul> <li> annotation based on adata.obs[‘cluster_all’]</li> </ul> </li> </ul> </li> </ul> </li> </ul> <p>______________________________________________________________________________________________________</p> <p>This datasets contains individually reprocessed and reannotated data from the following studies:</p> <p> </p> <ol> <li>Zhang, Y.<em> et al.</em> Single-cell analyses reveal key immune cell subsets associated with response to PD-L1 blockade in triple-negative breast cancer. <em>Cancer Cell</em> <strong>39</strong>, 1578-1593.e1578 (2021). <a href="https://doi.org:10.1016/j.ccell.2021.09.010">https://doi.org:10.1016/j.ccell.2021.09.010</a></li> <li>Keenan, B. P.<em> et al.</em> Circulating monocytes associated with anti-PD-1 resistance in human biliary cancer induce T&#xa0;cell paralysis. <em>Cell Reports</em> <strong>40</strong> (2022). <a href="https://doi.org:10.1016/j.celrep.2022.111384">https://doi.org:10.1016/j.celrep.2022.111384</a></li> <li><span> C</span>he, L.-H.<em> et al.</em> A single-cell atlas of liver metastases of colorectal cancer reveals reprogramming of the tumor microenvironment in response to preoperative chemotherapy. <em>Cell Discovery</em> <strong>7</strong>, 80 (2021). <a href="https://doi.org:10.1038/s41421-021-00312-y">https://doi.org:10.1038/s41421-021-00312-y</a></li> <li> Wang, F.<em> et al.</em> Single-cell and spatial transcriptome analysis reveals the cellular heterogeneity of liver metastatic colorectal cancer. <em>Science Advances</em> <strong>9</strong>, eadf5464 (2023). <a href="https://doi.org:doi:10.1126/sciadv.adf5464">https://doi.org:doi:10.1126/sciadv.adf5464</a></li> <li> Liu, C.<em> et al.</em> Time-resolved systems immunology reveals a late juncture linked to fatal COVID-19. <em>Cell</em> <strong>184</strong>, 1836-1857.e1822 (2021). <a href="https://doi.org:10.1016/j.cell.2021.02.018">https://doi.org:10.1016/j.cell.2021.02.018</a></li> <li> Ren, X.<em> et al.</em> COVID-19 immune features revealed by a large-scale single-cell transcriptome atlas. <em>Cell</em> <strong>184</strong>, 1895-1913.e1819 (2021). <a href="https://doi.org:10.1016/j.cell.2021.01.053">https://doi.org:10.1016/j.cell.2021.01.053</a></li> <li> Hao, Y.<em> et al.</em> Integrated analysis of multimodal single-cell data. <em>bioRxiv</em>, 2020.2010.2012.335331 (2020). <a href="https://doi.org:10.1101/2020.10.12.335331">https://doi.org:10.1101/2020.10.12.335331</a></li> <li> Terekhova, M.<em> et al.</em> Single-cell atlas of healthy human blood unveils age-related loss of NKG2C+GZMB-CD8+ memory T cells and accumulation of type 2 memory T cells. <em>Immunity</em> <strong>56</strong>, 2836-2854.e2839 (2023). <a href="https://doi.org:10.1016/j.immuni.2023.10.013">https://doi.org:10.1016/j.immuni.2023.10.013</a></li> <li>Oelen, R.<em> et al.</em> Single-cell RNA-sequencing of peripheral blood mononuclear cells reveals widespread, context-specific gene expression regulation upon pathogenic exposure. <em>Nature Communications</em> <strong>13</strong>, 3267 (2022). <a href="https://doi.org:10.1038/s41467-022-30893-5">https://doi.org:10.1038/s41467-022-30893-5</a></li> <li>Steele, N. G.<em> et al.</em> Multimodal mapping of the tumor and peripheral blood immune landscape in human pancreatic cancer. <em>Nature Cancer</em> <strong>1</strong>, 1097-1112 (2020). <a href="https://doi.org:10.1038/s43018-020-00121-4">https://doi.org:10.1038/s43018-020-00121-4</a></li> </ol>
scGeneAI pbmc multimodal dataset
<p>The input pbmc_multimodal_2023 dataset used in the full-size examples in scGenAI is uploaded here</p>
PBMC CITE-seq, 10x Multiome, and TEA-seq multiomic datasets from Swanson, et al. eLife (2021).
<p>Assembled multiomic datasets from Swanson, et al. <em>Simultaneous trimodal single-cell measurement of transcripts, epitopes, and chromatin accessibility using TEA-seq</em>. eLife 2021;10:e63632 DOI: <a href="https://doi.org/10.7554/eLife.63632">10.7554/eLife.63632</a> </p> <p>Datasets were assembled from the GEO repository: <a href="https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE158013">GSE158013</a> </p> <p>Data are provided as SeuratObjects stored in separate .rds files using the saveRDS() function in R.</p> <p>MuData files for use with Python tools (muon, scanpy, scvitools, et al.) will be added soon.</p> <p>Contact Lucas Graybuck (lucasg at alleninstitute dot org) if there are problems with these datasets.</p>
CyTOF data of PBMC samples of patients with metastatic pancreatic ductal adenocarcinoma
<p>These two CyTOF datasets are a part of the manuscript by M. Baretti "E<span>ntinostat in combination with nivolumab in metastatic pancreatic ductal adenocarcinoma: a phase 2 clinical trial" accepted in Nature Communications. The datasets contain FCS files of PBMCs samples of patients with metastatic pancreatic ductal adenocarcinoma treated with entinostat and nivolumab. PBMC samples were run with myeloid- and lymphoid-oriented panels.<br></span></p>
Bulk RNA-Seq PBMC data of SLE patients and healthy volunteers/ profiling of 29 individual immune cell types as well as PBMCs of healthy donors
<p>This Zenodo project contains processed gene expression data from two publicly available data sets. It includes the gene expression data of peripheral blood mononuclear cells (PBMCs) of systemic lupus erythematosus (SLE) patients as well as healthy volunteers (GSE122459). The project also comprises the bulk RNA-Seq profiling of 29 immune cell types as well as PBMCs of healthy individuals (GSE107011). In both cases, the raw RNA-Seq data was downloaded, aligned and processed. The gene expression data is available in form of a count matrix (GSE107011) or count matrix and transcript-per-million (TPM) values (GSE122459). For the latter, an annotation file is attached. Further details are provided in the information file. </p>
Azimuth ATAC Reference - Human PBMC
<p>Here we provide the reference data files used to run the Azimuth ATAC Human PBMC reference web application. For a full description of the reference data structure, please see the wiki at <a href="https://github.com/satijalab/azimuth/wiki/Azimuth-Reference-Format">https://github.com/satijalab/azimuth/wiki/Azimuth-Reference-Format</a>.</p>
High frequency of X4/DM-tropic viruses in PBMC samples from HIV-1 recently infected blood donors by massively parallel sequencing: the REDS II Study
<p>Here is a sub-library of the <em>env</em> V3 massively parallel sequencing proviral data generated (by Illumina MiSeq platform) during the early phase of HIV-1 infection in a group of first-time blood donors. Only paired-end reads that encompass the complete V3 region from each dataset were extracted, uploaded and considered for the analysis to avoid artificial generation of <em>in silico</em> chimeras through assembly and to evade inflating the diversity estimates of the V3 region</p>
10K-Cell Subset of PBMC CITE-Seq Dataset for CITEViz
<p>This repository contains an example CITE-Seq data (10K peripheral blood mononuclear cells) to test the CITEViz program. The CITEViz preprint is available <a href="https://www.biorxiv.org/content/10.1101/2022.05.15.491411v1">here</a>, and the documentation website is located <a href="https://maxsonbraunlab.github.io/CITEViz/">here</a>. The original data underlying this article are available in GEO (Gene Expression Omnibus) at <a href="https://www.ncbi.nlm.nih.gov/geo/">https://www.ncbi.nlm.nih.gov/geo/</a>, and can be accessed with GSE164378. </p>
T-cell receptor Vβ (TCRVB) deep sequencing on human T cells isolated from humanized mice engrafted with human PBMC and treated or not with PTCy post-transplantation
<p>Nucleotide sequence data of TCRVB sequencing performed on T cells isolated from the pre-transplantation hPBMC (donor T cells) or from mice organs at day 21 post-transplantation (injected or not with PTCy at day 3) to determine the impact of PTCy on the T cell V beta (TCRVB) receptor repertoire diversity.</p> <p>NSG mice were engrafted with human PBMC to develop xeno-GVHD and treated or not with 100 mg/kg PTCy. Spleens and lungs from 10 NSG mice per group were pooled and stained to sorted human CD4+ and CD8+ . DNA of human CD4+ and CD8+ T cells sorted from each organ and from the PBMC donor was extracted. One hundred fifty µg from each sample were used for T-cell receptor Vβ (TCRVB) deep sequencing performed by Adaptive Biotechnologies.<br> </p>
Mass Cytometry (CyTOF) FCS files from Priest et al. 2024. Human PBMC from longitudinal analysis of COVID-19, Bacterial Sepsis, mRNA vaccination cohorts.
<p>Mass Cytometry (CyTOF) FCS files from Priest et al. "Non-classical CD45RB<sup>lo</sup> memory B-cells are the majority of circulating antigen-specific B-cells following mRNA vaccination and COVID-19 infection." Research Square 2024. </p> <p>Files are already normalised, debarcoded, gated, batch corrected and compensated as described in Priest et al. </p> <p>Data is from Human PBMCs of londitudanal cohorts of Severe COVID-19, Sepsis and mRNA vaccine recipients. </p> <p>Samples were barcoded, mixed and then split magnetically before staining with seperate antibody panels for CD3+ (CD4, Treg, Tfh, CD8, gdT) or CD3- (B cells, DC, NK, Monocytes) to give approximatly 1280 FCS files from 218 individuals. </p> <p>A follow up experiment with a B-cell specific panel and Tetramers is included. </p> <p>Patient level metadata and antibody panel details are included. </p> <p> </p>
Processing of 10X 500 PBMC single cell ATAC-seq data with SnapATAC2
<p>Input files for SnapATAC2 tutorial on Galaxy</p>
PBMC CITE-seq reference
<p>This PBMC CITE-seq reference object was constructed using Seurat v5.</p>
COVID-19 PBMC sample information and the VCF file of variants around OAS1 gene.
<p>We investigated our recently published PBMC scRNA-seq data (Stephenson et al. 2021 Nat Med) obtained from 112 donors, including 84 COVID-19 positive individuals, and profiled using the CITE-seq approach, as an independent in vivo validation of OAS1 eQTL colocalisation with GWAS locus for COVID-19 susceptibility. The dataset includes both sample information and genotype information of variants around OAS1 gene (chr12:111906777-113906777) as in VCF format. The whole variant data of the COVID-19 PBMC samples is available upon request.</p>
Haploidentical PBMC Transplant for Severe Congenital Anemias
ClinicalTrials.gov study NCT00977691. IPD Sharing: Not stated. Countries: 1. Publications: 3.
PBMC 3k test datasets for besca
<p>This is a single cell transcriptomics dataset containing roughly 3,000 PBMCs. The original data was downloaded from the Seurat 3k PBMC tutorial: https://satijalab.org/seurat/v3.0/pbmc3k_tutorial.html. We reprocessed the dataset using the Besca package (https://github.com/bedapub/besca).</p> <p> </p>
SCARlink test data (PBMC)
<p>Seurat and ArchR test files for running SCARlink <a href="https://github.com/snehamitra/SCARlink/blob/main/notebooks/tutorial.ipynb">tutorial notebook</a>. Related manuscript: <a href="https://www.nature.com/articles/s41588-024-01689-8">https://www.nature.com/articles/s41588-024-01689-8</a></p>
PBMC scRNA-seq datasets measured using different 10X Chromium chemistries
<p>PBMC scRNA-seq datasets measured using different 10X Chromium chemistries<br> Obtained from: https://www.10xgenomics.com/resources/datasets</p>
AIDA PBMC sQTL
<p>This is AIDA PBMC cis-sQTL dataset including 19 cell types.</p>
Genetically Engineered PBMC and PBSC Expressing NY-ESO-1 TCR After a Myeloablative Conditioning Regimen to Treat Patients With Advanced Cancer
ClinicalTrials.gov study NCT03240861. IPD Sharing: Not stated. Countries: 1. Publications: 1.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.