Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
207
datasets available to search
ShareScore release 0.9.0
Dataset results
207 results for “Single-Cell Genomics”
The evolution of genomic, transcriptomic, and single-cell protein markers of metastatic upper tract urothelial carcinoma
<p>The molecular characteristics of metastatic upper tract urothelial carcinoma (UTUC) are unknown. The genomic and transcriptomic differences between primary and metastatic UTUC is not well described either. We combined whole-exome sequencing, RNA-sequencing, and Imaging Mass Cytometry<sup>TM</sup> (IMC<sup>TM</sup>) of 44 tumor samples from 28 patients with high-grade primary and metastatic UTUC. IMC enables spatially resolved single-cell analyses to examine the evolution of cancer cell, immune cell, and stromal cell markers using mass cytometry with lanthanide metal-conjugated antibodies. We discovered that actionable genomic alterations are frequently discordant between primary and metastatic UTUC tumors in the same patient. In contrast, molecular subtype membership and immune depletion signature were stable across primary and matched metastatic UTUC. Molecular and immune subtypes were consistent between bulk RNA-sequencing and mass cytometry of protein markers from 340,798 single-cells. Molecular subtyping at the single cell level was highly conserved between primary and metastatic UTUC tumors within the same patient.</p>
Joint Trajectory Inference for Single-cell Genomics Using Deep Learning with a Mixture Prior
<p>The datasets used in the paper "Joint Trajectory Inference for Single-cell Genomics Using Deep Learning with a Mixture Prior". A detailed description of these datasets is available at https://github.com/jaydu1/VITAE/tree/master/data.</p>
GWAS to single cell: Intersecting single-cell transcriptomics and genome wide association studies identifies crucial cell-populations and candidate genes for atherosclerosis.
<p><strong>Background</strong></p> <p>Genome-wide association studies (GWAS) have discovered hundreds of common genetic variants for atherosclerotic disease and cardiovascular risk factors. The translation of susceptibility loci into biological mechanisms and targets for drug discovery remains challenging. Intersecting genetic and gene expression data has led to identification of candidate genes. However, the assayed tissues are often non-diseased and heterogeneous in cell composition confounding the candidate prioritization. We collected single-cell transcriptomics (scRNA-seq) from atherosclerotic plaques and aimed to identify cell-type-specific expression of disease-associated genes. </p> <p> </p> <p><strong>Methods and Results</strong></p> <p>To identify disease-associated candidate genes, we applied gene-based analyses using GWAS summary statistics from 46 atherosclerotic, cardiometabolic, and other traits. Next we intersected these candidates with single-cell transcriptomics (scRNA-seq) to identify those genes that are specifically expressed in individual cell (sub)populations of atherosclerotic plaques. We derive an enrichment score and show that loci that associated with coronary artery disease demonstrated a prominent substrate in plaque smooth muscle cells (<em>SKI</em>, <em>KANK2</em>, <em>SORT1</em>), endothelial cells (<em>SLC44A1</em>, <em>ATP2B1</em>), and macrophages (<em>APOE</em>, <em>HNRNPUL1</em>). Further sub clustering of SMC-subtypes revealed genes in risk loci for coronary calcification specifically enriched in a synthetic cluster of SMCs. To verify the robustness of our approach, we used liver-derived scRNAseq-data and showed enrichment of circulating lipids-associated loci in hepatocytes.</p> <p><br> <strong>Conclusion</strong></p> <p>We confirm known gene-cell pairs relevant for atherosclerotic disease, and discovered novel pairs pointing to new biological mechanisms amenable for therapy. We present an intuitive single-cell transcriptomics driven workflow rooted in human large-scale genetic studies to identify putative candidate genes and affected cells associated with cardiovascular traits.</p> <p> </p>
Single-cell mouse and PC9 data for "TP53 loss with whole genome doubling mediates heterogeneous intra-patient therapy response through Chromosomal Instability"
<p>This repository includes the processed data (including copy number profiles and related analysis) for the E/EP mouse tumors and for the PC9 resistance cell lines for all the analyses of the manuscript "TP53 loss with whole genome doubling mediates heterogeneous intra-patient therapy response through Chromosomal Instability".</p><p>The code for the related analyses is available in GitHub at https://github.com/zaccaria-lab/TP53loss_WGD</p>
Linking regulatory variants to target genes by integrating single-cell multiome methods and genomic distance
<p>The below data are associated with our paper entitled "Linking regulatory variants to target genes by integrating single-cell multiome methods and genomic distance."</p> <p>1) SNP-gene link predictions generated by pgBoost and existing methods SCENT (Sakaue et al. 2024 <em>Nat Genet</em>), Signac (Stuart et al. 2021 <em>Nat Methods</em>), ArchR (Granja et al. 2021 <em>Nat Genet</em>), and Cicero (Pliner et al. 2018 <em>Mol Cell</em>).</p> <p><strong>pgBoost_scores.tsv.gz </strong>contains linking predictions made by pgBoost.</p> <p><strong>constituent_method_scores.tsv.gz</strong> contains linking predictions made by constituent methods.</p> <p><em><span>**NOTE: promoters (+/- 1kb from TSS) and candidate links >500kb are excluded from linking predictions (see manuscript)**</span></em></p> <p>Linking scores and percentiles are reported for each method (pgBoost score, SCENT FDR, Signac correlation, ArchR correlation, Cicero co-accessibility). Rank percentiles are computed as: 1 - (rank / n). When multiple links receive the same score, they are assigned the percentile of the top rank. Links unscored by each method (denoted by zeros* in the linking score column) are assigned a percentile equivalent to the percent of links unscored by the focal method. See the Methods section of the paper for further details on computing linking scores and summarizing scores across cell types and data sets.</p> <p>*Candidate links tested and assigned a co-accessibility of zero by the Cicero method are given a score of 1e-100 in the "Cicero" column to distinguish between unscored candidate links and candidate links assigned a partial correlation of zero (see Pliner et al. 2018 <em>Mol Cell</em>).</p> <p><em>NOTE: The predictions associated with this release (version 2) were generated using an expanded set of data sets, an expanded training set, and corrected TSS coordinates.</em></p> <p>2) GWAS-derived evaluation SNP-gene link evaluation set.</p> <p><strong>gwas_evaluation.tsv</strong>: GWAS-derived evaluation SNP-gene link evaluation set. Column 1 provides SNP coordinates in the format <chr-start-end>. This evaluation framework was proposed by Weeks et al. 2024 <em>Nature Genetics</em> based on fine-mapping results from Kanai et al. <em>medrxiv</em> (see Methods: <em>Evaluation data sets</em> of Dorans et al.). "True" links (gold = 1) are non-coding variants fine-mapped to a focal trait (PIP > 0.1) with a coding variant for exactly one candidate gene within 1 Mb attaining PIP > 0.5 for the same trait. "False" links (gold = 0) are candidate SNP-gene pairs involving a SNP with a "true" link. This file of SNP-gene links was adapted from credible set-gene links <a href="https://github.com/Deylab999/GWAS_benchmark_IGVF/blob/bb91d08cc02d59cdd829eb1430569057ac26c5fe/V2G/ENCODE_E2G_2023/UKBiobank.ABCGene.anyabc.tsv">here</a> (the "truth" column defines true/false links) by identifying SNPs with PIP > 0.1 within each credible set-gene link.</p>
SEACells: Inference of transcriptional and epigenomic cellular states from single-cell genomics data
<p>Processed data for the manuscript "" available on bioRxiv at ""</p> <p>Data is available for the following samples</p> <ol> <li>CD34+ Multiome data : 2 replicates </li> <li>T-cell depleted bone marrow Multiome data: 2 replicates </li> </ol> <p> </p> <p>The following counts and fragments files are available for each replicate </p> <ol> <li><sample>_filtered_feature_bc_matrix.h5: Feature counts from CellRanger ARC</li> <li><sample>_atac_fragments.tsv.gz: ATAC fragments file from Cellranger ARC</li> <li><sample>_atac_fragments.tsv.gz.tbi: Index files for ATAC fragments file from Cellranger ARC</li> </ol> <p> </p> <p>The following scanpy anndata objects are also available</p> <ol> <li>cd34_multiome_rna.h5ad: Anndata object with normalized data, cell type annotations and clusters for the RNA modality of CD34+ hematopoietic stem and progenitor cells </li> <li>cd34_multiome_atac.h5ad: Anndata object with peak counts, cell type annotations and clusters for the ATAC modality of of CD34+ hematopoietic stem and progenitor cells.</li> <li>cd34_multiome_rna.h5ad: Anndata object with normalized data, cell type annotations and clusters for the RNA modality of T-cell depleted bone marrow dataset.</li> <li>bm_multiome_atac.h5ad: Anndata object with peak counts, cell type annotations and clusters for the ATAC modality of T-cell depleted bone marrow dataset.</li> </ol>
STING OPS: HeLa Genome-wide Screen Single-Cell Features (Part 2/5)
<div> <div> <p><strong>Classification and functional characterization of regulators of intracellular STING trafficking identified by genome-wide optical pooled screening</strong></p> <p>Single-cell features and coordinates for HeLa genome-wide screen, Zenodo dataset part 2/5.</p> <p>Images available at gs://opspublic-east1/STINGOpticalPooledScreen/Images/GW*.</p> <p>README for additional image information is at gs://opspublic-east1/STINGOpticalPooledScreen/STING_README.</p> </div> </div>
STING OPS: HeLa Genome-wide Screen Single-Cell Features (Part 5/5)
<div> <div> <p><strong>Classification and functional characterization of regulators of intracellular STING trafficking identified by genome-wide optical pooled screening</strong></p> <p>Single-cell features and coordinates for HeLa genome-wide screen, Zenodo dataset part 5/5.</p> <p>Images available at gs://opspublic-east1/STINGOpticalPooledScreen/Images/GW*.</p> <p>README for additional image information is at gs://opspublic-east1/STINGOpticalPooledScreen/STING_README.</p> </div> </div>
STING OPS: HeLa Genome-wide Screen Single-Cell Features (Part 1/5)
<div> <div> <p><strong>Classification and functional characterization of regulators of intracellular STING trafficking identified by genome-wide optical pooled screening</strong></p> <p>Single-cell features and coordinates for HeLa genome-wide screen, Zenodo dataset part 1/5.</p> <p>Images available at gs://opspublic-east1/STINGOpticalPooledScreen/Images/GW*.</p> <p>README for additional image information is at gs://opspublic-east1/STINGOpticalPooledScreen/STING_README.</p> </div> </div>
STING OPS: HeLa Genome-wide Screen Single-Cell Features (Part 3/5)
<div> <div> <p><strong>Classification and functional characterization of regulators of intracellular STING trafficking identified by genome-wide optical pooled screening</strong></p> <p>Single-cell features and coordinates for HeLa genome-wide screen, Zenodo dataset part 3/5.</p> <p>Images available at gs://opspublic-east1/STINGOpticalPooledScreen/Images/GW*.</p> <p>README for additional image information is at gs://opspublic-east1/STINGOpticalPooledScreen/STING_README.</p> </div> </div>
STING OPS: HeLa Genome-wide Screen Single-Cell Features (Part 4/5)
<div> <div> <p><strong>Classification and functional characterization of regulators of intracellular STING trafficking identified by genome-wide optical pooled screening</strong></p> <p>Single-cell features and coordinates for HeLa genome-wide screen, Zenodo dataset part 4/5.</p> <p>Images available at gs://opspublic-east1/STINGOpticalPooledScreen/Images/GW*.</p> <p>README for additional image information is at gs://opspublic-east1/STINGOpticalPooledScreen/STING_README.</p> </div> </div>
Resolving Organoid Brain Region Identities by Mapping Single-Cell Genomic Data to Reference Atlases
<p>Data underlying the figures in the publication “Resolving organoid brain region identities by mapping single-cell genomic data to reference atlases”, published in <em>Cell Stem Cell, </em><strong>2021</strong><em>, </em>28, 1148–1159.</p> <p><a href="https://www.sciencedirect.com/science/article/pii/S1934590921000655">https://www.sciencedirect.com/science/article/pii/S1934590921000655</a></p> <p>Table of contents:</p> <p><strong>1. patscreen_srt.rds</strong>; Numerical data for <em>Figure 7</em>: RNA-seq data of a patterning screen in organoids with an array of small molecules. The dataset is in the rds data format, which can be opened in the R programming language using the function `readRDS()`. Once opened, the dataset is a Seurat object (https://satijalab.org/seurat/) and contains both the transcript counts and the metadata for all samples in the screen. The raw data used in figure 7 was also deposited in ArrayExpress (<a href="https://www.ebi.ac.uk/arrayexpress/experiments/E-MTAB-10037/">https://www.ebi.ac.uk/arrayexpress/experiments/E-MTAB-10037/</a>)</p>
Sampling time-dependent artifacts in single-cell genomics studies: scRNA-seq data
<p>Robust protocols and automation now enable large-scale single-cell RNA and ATAC sequencing experiments and their application on biobank and clinical cohorts. However, technical biases introduced during sample acquisition can hinder solid, reproducible results, and a systematic benchmarking is required before entering large-scale data production. Here, we report the existence and extent of gene expression and chromatin accessibility artifacts introduced during sampling and identify experimental and computational solutions for their prevention.</p> <p>This repository contains the expression matrices and Seurat objects associated with the scRNA-seq data of the manuscript: "Sampling time-dependent artifacts in single-cell genomics studies" published in Genome Biology in 2020. The purpose of this repo is to share processed files and metadata for immediate access and reproducibility. The code to analyze it is thoroughly documented at the associated Github repository (https://github.com/massonix/sampling_artifacts).</p>
Data from: Morphological identification and single-cell genomics of marine diplonemids
Recent global surveys of marine biodiversity have revealed that a group of organisms known as "marine diplonemids" constitutes one of the most abundant and diverse planktonic lineages [1]. Though discovered over a decade ago [2 and 3], their potential importance was unrecognized, and our knowledge remains restricted to a single gene amplified from environmental DNA, the 18S rRNA gene (small subunit [SSU]). Here, we use single-cell genomics (SCG) and microscopy to characterize ten marine diplonemids, isolated from a range of depths in the eastern North Pacific Ocean. Phylogenetic analysis confirms that the isolates reflect the entire range of marine diplonemid diversity, and comparisons to environmental SSU surveys show that sequences from the isolates range from rare to superabundant, including the single most common marine diplonemid known. SCG generated a total of ∼915 Mbp of assembled sequence across all ten cells and ∼4,000 protein-coding genes with homologs in the Kyoto Encyclopedia of Genes and Genomes (KEGG) orthology database, distributed across categories expected for heterotrophic protists. Models of highly conserved genes indicate a high density of non-canonical introns, lacking conventional GT-AG splice sites. Mapping metagenomic datasets [4] to SCG assemblies reveals virtually no overlap, suggesting that nuclear genomic diversity is too great for representative SCG data to provide meaningful phylogenetic context to metagenomic datasets. This work provides an entry point to the future identification, isolation, and cultivation of these elusive yet ecologically important cells. The high density of nonconventional introns, however, also portends difficulty in generating accurate gene models and highlights the need for the establishment of stable cultures and transcriptomic analyses.
scooby: Modeling multi-modal genomic profiles from DNA sequence at single-cell resolution - Supplementary data and code
<p>Data and code to reproduce the analyses from the study: "scooby: Modeling multi-modal genomic profiles from DNA sequence at single-cell resolution". </p>
Single-cell RNA sequencing of CNS-infiltrating HSC-derived phagocytes of Ms4a3Ai14, BM chimeric mice (CD45.2 Csf2rb-/-: CD45.1 Csf2rb+/+ and CD45.2 Ifngr1-/-: CD45.1 Ifngr1+/+) using 10X Genomics platform. IFN-γ and GM-CSF control complementary differentiation programs in the monocyte to phagocyte transition during neuroinflammation.
<p><strong>Single-cell RNA sequencing of CNS-infiltrating HSC-derived phagocytes of <em>Ms4a3</em><sup>Ai14</sup> at onset and peak EAE, BM chimeric mice (CD45.2 <em>Csf2rb</em><sup>-/-</sup>: CD45.1 <em>Csf2rb</em><sup>+/+</sup> and CD45.2 <em>Ifngr1<sup>-/-</sup></em>: CD45.1 <em>Ifngr1<sup>+/+</sup></em>) using 10X Genomics platform.</strong></p> <p>The sorted cells were loaded into 10x Genomics Chromium in parallel. Libraries were prepared as per the manufacturer's protocol (Chromium Next GEM Single Cell 3ʹ Reagent Kits v3.1 protocol) and sequenced on an Illumina NovaSeq sequencer according to 10X Genomics recommendations (paired-end reads, R1=28, i7=8, R2=91) to a depth of around 50,000 reads per cell.</p> <p>Initial processing was done using Cell Ranger (v3.1.0) mkfastq and count (reads were aligned to GENCODE reference build GRCm38.p6 Release M23 with added tdTomato sequence for the dataset from <em>Ms4a3</em><sup>Ai14</sup> mouse and collapse UMIs). Starting from the filtered gene-cell count matrix produced by CellRranger's in-built cell calling algorithms, we proceeded with Seurat v4 workflow.</p>
Single-cell RNA sequencing of Lymph node-infiltrating HSC-derived phagocytes of Ms4a3Ai14 using 10X Genomics platform. IFN-γ and GM-CSF control complementary differentiation programs in the monocyte to phagocyte transition during neuroinflammation.
<p><strong>Single-cell RNA sequencing of Lymph node-infiltrating HSC-derived phagocytes of <em>Ms4a3</em><sup>Ai14</sup> at onset and peak EAE using 10X Genomics platform.</strong></p> <p>The sorted cells were loaded into 10x Genomics Chromium in parallel. Libraries were prepared as per the manufacturer's protocol (Chromium Next GEM Single Cell 3ʹ Reagent Kits v3.1 protocol) and sequenced on an Illumina NovaSeq sequencer according to 10X Genomics recommendations (paired-end reads, R1=28, i7=8, R2=91) to a depth of around 50,000 reads per cell.</p> <p>Initial processing was done using Cell Ranger (v3.1.0) mkfastq and count (reads were aligned to GENCODE reference build GRCm38.p6 Release M23 with added tdTomato sequence for the dataset from <em>Ms4a3</em><sup>Ai14</sup> mouse and collapse UMIs). Starting from the filtered gene-cell count matrix produced by CellRranger's in-built cell calling algorithms, we proceeded with Seurat v4 workflow.</p>
Single-cell RNA sequencing of Bone Marrow-infiltrating HSC-derived phagocytes of Ms4a3Ai14 using 10X Genomics platform. IFN-γ and GM-CSF control complementary differentiation programs in the monocyte to phagocyte transition during neuroinflammation.
<p><strong>Single-cell RNA sequencing of Bone Marrow-infiltrating HSC-derived phagocytes of <em>Ms4a3</em><sup>Ai14</sup> at onset and peak EAE using 10X Genomics platform.</strong></p> <p>The sorted cells were loaded into 10x Genomics Chromium in parallel. Libraries were prepared as per the manufacturer's protocol (Chromium Next GEM Single Cell 3ʹ Reagent Kits v3.1 protocol) and sequenced on an Illumina NovaSeq sequencer according to 10X Genomics recommendations (paired-end reads, R1=28, i7=8, R2=91) to a depth of around 50,000 reads per cell.</p> <p>Initial processing was done using Cell Ranger (v3.1.0) mkfastq and count (reads were aligned to GENCODE reference build GRCm38.p6 Release M23 with added tdTomato sequence for the dataset from <em>Ms4a3</em><sup>Ai14</sup> mouse and collapse UMIs). Starting from the filtered gene-cell count matrix produced by CellRranger's in-built cell calling algorithms, we proceeded with Seurat v4 workflow.</p>
Single-cell RNA sequencing of Blood-infiltrating HSC-derived phagocytes of Ms4a3Ai14 using 10X Genomics platform. IFN-γ and GM-CSF control complementary differentiation programs in the monocyte to phagocyte transition during neuroinflammation.
<p><strong>Single-cell RNA sequencing of Blood-infiltrating HSC-derived phagocytes of <em>Ms4a3</em><sup>Ai14</sup> at onset and peak EAE using 10X Genomics platform.</strong></p> <p>The sorted cells were loaded into 10x Genomics Chromium in parallel. Libraries were prepared as per the manufacturer's protocol (Chromium Next GEM Single Cell 3ʹ Reagent Kits v3.1 protocol) and sequenced on an Illumina NovaSeq sequencer according to 10X Genomics recommendations (paired-end reads, R1=28, i7=8, R2=91) to a depth of around 50,000 reads per cell.</p> <p>Initial processing was done using Cell Ranger (v3.1.0) mkfastq and count (reads were aligned to GENCODE reference build GRCm38.p6 Release M23 with added tdTomato sequence for the dataset from <em>Ms4a3</em><sup>Ai14</sup> mouse and collapse UMIs). Starting from the filtered gene-cell count matrix produced by CellRranger's in-built cell calling algorithms, we proceeded with Seurat v4 workflow.</p>
Systematic dissection of transcriptional regulatory networks by genome-scale and single-cell CRISPR screens
Millions of putative transcriptional regulatory elements (TREs) have been cataloged in the human genome, yet their functional relevance in specific pathophysiological settings remains to be determined. This is critical to understand how oncogenic transcription factors (TFs) engage specific TREs to impose transcriptional programs underlying malignant phenotypes. Here, we combine cutting edge CRISPR screens and epigenomic profiling to functionally survey ≈15,000 TREs engaged by estrogen receptor (ER). We show that ER exerts its oncogenic role in breast cancer by engaging TREs enriched in GATA3, TFAP2C, and H3K27Ac signal. These TREs control critical downstream TFs, among which TFAP2C plays an essential role in ER-driven cell proliferation. Together, our work reveals novel insights into a critical oncogenic transcription program and provides a framework to map regulatory networks, enabling to dissect the function of the noncoding genome of cancer cells.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.