Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
371
datasets available to search
ShareScore release 0.9.0
Dataset results
371 results for “single cell genomics”
Improving the efficiency of single cell genome sequencing based on overlapping pooling strategy
Single cell genome sequencing has become a useful tool in medicine and biology studies. However, an independent library is required for each cell in single cell genome sequencing, so that the cost grows in step with the number of cells. In this study, we report a study on efficient single-cell copy number variation (CNV) analysis based on overlapping pooling strategy together with branch and bound (B&B) algorithm. Single cells are overlapped pooled before sequencing, and later are assorted into specific types by estimating their CNV patterns by B&B algorithm. Instead of constructing libraries for each cell, a library is required only for each pool. As long as the number of pools is smaller than the cells, fewer libraries are needed, and a lower cost is spent. Through computer simulations, we overlapping pooled 80 cells into 40 and 27 pools and classified them into cell types based on CNV pattern. The results showed that 84% cells in 40 pools and 76.5% cells in 27 pools were correctly classified on average, while only half or one-third of the sequencing libraries are required. Combining with traditional approaches, our method is expected to significantly improve the efficiency of single cell genome sequencing.
Sampling time-dependent artifacts in single-cell genomics studies: scRNA-seq data
<p>Robust protocols and automation now enable large-scale single-cell RNA and ATAC sequencing experiments and their application on biobank and clinical cohorts. However, technical biases introduced during sample acquisition can hinder solid, reproducible results, and a systematic benchmarking is required before entering large-scale data production. Here, we report the existence and extent of gene expression and chromatin accessibility artifacts introduced during sampling and identify experimental and computational solutions for their prevention.</p> <p>This repository contains the expression matrices and Seurat objects associated with the scRNA-seq data of the manuscript: "Sampling time-dependent artifacts in single-cell genomics studies" published in Genome Biology in 2020. The purpose of this repo is to share processed files and metadata for immediate access and reproducibility. The code to analyze it is thoroughly documented at the associated Github repository (https://github.com/massonix/sampling_artifacts).</p>
High quality genomes produced from single MinION flow cells clarify polyploid and demographic histories of critically endangered Fraxinus (ash) species
Open the record for dataset details and reuse information.
Improving the efficiency of single cell genome sequencing based on overlapping pooling strategy
Open the record for dataset details and reuse information.
Accurate SNV detection in single cells by transposon-based whole-genome amplification of complementary strands
<p>Common SNPs from gnomAD</p>
Data from: Morphological identification and single-cell genomics of marine diplonemids
Recent global surveys of marine biodiversity have revealed that a group of organisms known as "marine diplonemids" constitutes one of the most abundant and diverse planktonic lineages [1]. Though discovered over a decade ago [2 and 3], their potential importance was unrecognized, and our knowledge remains restricted to a single gene amplified from environmental DNA, the 18S rRNA gene (small subunit [SSU]). Here, we use single-cell genomics (SCG) and microscopy to characterize ten marine diplonemids, isolated from a range of depths in the eastern North Pacific Ocean. Phylogenetic analysis confirms that the isolates reflect the entire range of marine diplonemid diversity, and comparisons to environmental SSU surveys show that sequences from the isolates range from rare to superabundant, including the single most common marine diplonemid known. SCG generated a total of ∼915 Mbp of assembled sequence across all ten cells and ∼4,000 protein-coding genes with homologs in the Kyoto Encyclopedia of Genes and Genomes (KEGG) orthology database, distributed across categories expected for heterotrophic protists. Models of highly conserved genes indicate a high density of non-canonical introns, lacking conventional GT-AG splice sites. Mapping metagenomic datasets [4] to SCG assemblies reveals virtually no overlap, suggesting that nuclear genomic diversity is too great for representative SCG data to provide meaningful phylogenetic context to metagenomic datasets. This work provides an entry point to the future identification, isolation, and cultivation of these elusive yet ecologically important cells. The high density of nonconventional introns, however, also portends difficulty in generating accurate gene models and highlights the need for the establishment of stable cultures and transcriptomic analyses.
scooby: Modeling multi-modal genomic profiles from DNA sequence at single-cell resolution - Supplementary data and code
<p>Data and code to reproduce the analyses from the study: "scooby: Modeling multi-modal genomic profiles from DNA sequence at single-cell resolution". </p>
Single-cell RNA sequencing of CNS-infiltrating HSC-derived phagocytes of Ms4a3Ai14, BM chimeric mice (CD45.2 Csf2rb-/-: CD45.1 Csf2rb+/+ and CD45.2 Ifngr1-/-: CD45.1 Ifngr1+/+) using 10X Genomics platform. IFN-γ and GM-CSF control complementary differentiation programs in the monocyte to phagocyte transition during neuroinflammation.
<p><strong>Single-cell RNA sequencing of CNS-infiltrating HSC-derived phagocytes of <em>Ms4a3</em><sup>Ai14</sup> at onset and peak EAE, BM chimeric mice (CD45.2 <em>Csf2rb</em><sup>-/-</sup>: CD45.1 <em>Csf2rb</em><sup>+/+</sup> and CD45.2 <em>Ifngr1<sup>-/-</sup></em>: CD45.1 <em>Ifngr1<sup>+/+</sup></em>) using 10X Genomics platform.</strong></p> <p>The sorted cells were loaded into 10x Genomics Chromium in parallel. Libraries were prepared as per the manufacturer's protocol (Chromium Next GEM Single Cell 3ʹ Reagent Kits v3.1 protocol) and sequenced on an Illumina NovaSeq sequencer according to 10X Genomics recommendations (paired-end reads, R1=28, i7=8, R2=91) to a depth of around 50,000 reads per cell.</p> <p>Initial processing was done using Cell Ranger (v3.1.0) mkfastq and count (reads were aligned to GENCODE reference build GRCm38.p6 Release M23 with added tdTomato sequence for the dataset from <em>Ms4a3</em><sup>Ai14</sup> mouse and collapse UMIs). Starting from the filtered gene-cell count matrix produced by CellRranger's in-built cell calling algorithms, we proceeded with Seurat v4 workflow.</p>
Single-cell RNA sequencing of Lymph node-infiltrating HSC-derived phagocytes of Ms4a3Ai14 using 10X Genomics platform. IFN-γ and GM-CSF control complementary differentiation programs in the monocyte to phagocyte transition during neuroinflammation.
<p><strong>Single-cell RNA sequencing of Lymph node-infiltrating HSC-derived phagocytes of <em>Ms4a3</em><sup>Ai14</sup> at onset and peak EAE using 10X Genomics platform.</strong></p> <p>The sorted cells were loaded into 10x Genomics Chromium in parallel. Libraries were prepared as per the manufacturer's protocol (Chromium Next GEM Single Cell 3ʹ Reagent Kits v3.1 protocol) and sequenced on an Illumina NovaSeq sequencer according to 10X Genomics recommendations (paired-end reads, R1=28, i7=8, R2=91) to a depth of around 50,000 reads per cell.</p> <p>Initial processing was done using Cell Ranger (v3.1.0) mkfastq and count (reads were aligned to GENCODE reference build GRCm38.p6 Release M23 with added tdTomato sequence for the dataset from <em>Ms4a3</em><sup>Ai14</sup> mouse and collapse UMIs). Starting from the filtered gene-cell count matrix produced by CellRranger's in-built cell calling algorithms, we proceeded with Seurat v4 workflow.</p>
Single-cell RNA sequencing of Bone Marrow-infiltrating HSC-derived phagocytes of Ms4a3Ai14 using 10X Genomics platform. IFN-γ and GM-CSF control complementary differentiation programs in the monocyte to phagocyte transition during neuroinflammation.
<p><strong>Single-cell RNA sequencing of Bone Marrow-infiltrating HSC-derived phagocytes of <em>Ms4a3</em><sup>Ai14</sup> at onset and peak EAE using 10X Genomics platform.</strong></p> <p>The sorted cells were loaded into 10x Genomics Chromium in parallel. Libraries were prepared as per the manufacturer's protocol (Chromium Next GEM Single Cell 3ʹ Reagent Kits v3.1 protocol) and sequenced on an Illumina NovaSeq sequencer according to 10X Genomics recommendations (paired-end reads, R1=28, i7=8, R2=91) to a depth of around 50,000 reads per cell.</p> <p>Initial processing was done using Cell Ranger (v3.1.0) mkfastq and count (reads were aligned to GENCODE reference build GRCm38.p6 Release M23 with added tdTomato sequence for the dataset from <em>Ms4a3</em><sup>Ai14</sup> mouse and collapse UMIs). Starting from the filtered gene-cell count matrix produced by CellRranger's in-built cell calling algorithms, we proceeded with Seurat v4 workflow.</p>
Single-cell RNA sequencing of Blood-infiltrating HSC-derived phagocytes of Ms4a3Ai14 using 10X Genomics platform. IFN-γ and GM-CSF control complementary differentiation programs in the monocyte to phagocyte transition during neuroinflammation.
<p><strong>Single-cell RNA sequencing of Blood-infiltrating HSC-derived phagocytes of <em>Ms4a3</em><sup>Ai14</sup> at onset and peak EAE using 10X Genomics platform.</strong></p> <p>The sorted cells were loaded into 10x Genomics Chromium in parallel. Libraries were prepared as per the manufacturer's protocol (Chromium Next GEM Single Cell 3ʹ Reagent Kits v3.1 protocol) and sequenced on an Illumina NovaSeq sequencer according to 10X Genomics recommendations (paired-end reads, R1=28, i7=8, R2=91) to a depth of around 50,000 reads per cell.</p> <p>Initial processing was done using Cell Ranger (v3.1.0) mkfastq and count (reads were aligned to GENCODE reference build GRCm38.p6 Release M23 with added tdTomato sequence for the dataset from <em>Ms4a3</em><sup>Ai14</sup> mouse and collapse UMIs). Starting from the filtered gene-cell count matrix produced by CellRranger's in-built cell calling algorithms, we proceeded with Seurat v4 workflow.</p>
Single cell Iso-Sequencing enables rapid genome annotation for scRNAseq analysis
<p>Single <span>cell RNA sequencing (scRNAseq) is a powerful technique that continues to expand across various biological applications. However, incomplete 3' UTR annotations can impede single cell analysis resulting in genes that are partially or completely uncounted. Performing scRNAseq with incomplete 3' UTR annotations can hinder the identification of cell identities and gene expression patterns and lead to erroneous biological inferences. We demonstrate that performing single cell isoform sequencing (ScISOr-Seq) in tandem with scRNAseq can rapidly improve 3' UTR annotations. Using threespine stickleback fish (</span><em>Gasterosteus aculeatus</em><span>), we show that gene models resulting from a minimal embryonic ScISOr-Seq dataset retained 26.1% greater scRNAseq reads than gene models from Ensembl alone. Furthermore, pooling our ScISOr-Seq isoforms with a previously published adult bulk Iso-Seq dataset from stickleback, and merging the annotation with the Ensembl gene models, resulted in a marginal improvement (+0.8%) over the ScISOr-Seq only dataset. In addition, isoforms identified by ScISOr-Seq included thousands of new splicing variants. The improved gene models obtained using ScISOr-Seq lead to successful identification of cell types and increased the reads identified of many genes in our scRNAseq stickleback dataset. Our work illuminates ScISOr-Seq as a cost-effective and efficient mechanism to rapidly annotate genomes for scRNAseq.</span></p>
Systematic dissection of transcriptional regulatory networks by genome-scale and single-cell CRISPR screens
Millions of putative transcriptional regulatory elements (TREs) have been cataloged in the human genome, yet their functional relevance in specific pathophysiological settings remains to be determined. This is critical to understand how oncogenic transcription factors (TFs) engage specific TREs to impose transcriptional programs underlying malignant phenotypes. Here, we combine cutting edge CRISPR screens and epigenomic profiling to functionally survey ≈15,000 TREs engaged by estrogen receptor (ER). We show that ER exerts its oncogenic role in breast cancer by engaging TREs enriched in GATA3, TFAP2C, and H3K27Ac signal. These TREs control critical downstream TFs, among which TFAP2C plays an essential role in ER-driven cell proliferation. Together, our work reveals novel insights into a critical oncogenic transcription program and provides a framework to map regulatory networks, enabling to dissect the function of the noncoding genome of cancer cells.
Supporting data for "Studying stochastic systems biology of the cell with single-cell genomics data"
<p>The dataset tar.gz files contain the raw unspliced and spliced count matrices generated by <em>kallisto</em>|<em>bustools</em> 0.26.0 from six datasets.</p> <p>The GVP_2023...zip file is a mirror of the related GitHub repository.</p>
Single-cell somatic copy number variants in brain using different amplification methods and reference genomes
<p>Variable and constant sized bins for GRCh38 and T2T-Chm13 were generated using the buildGenome scripts provided with Ginkgo (<a href="https://github.com/robertaboukhalil/ginkgo/tree/master/genomes/scripts">https://github.com/robertaboukhalil/ginkgo/tree/master/genomes/scripts</a>).</p>
Single Cell Genomics of Psoriatic Skin
ClinicalTrials.gov study NCT02929745. IPD Sharing: Not stated. Countries: 1. Publications: 1.
Genome-wide Single Cell Haplotyping as a Generic Method for Preimplantation Genetic Diagnosis
ClinicalTrials.gov study NCT01336400. IPD Sharing: Not stated. Countries: 1. Publications: 2.
Single cell Iso-Sequencing enables rapid genome annotation for scRNAseq analysis
Open the record for dataset details and reuse information.
Data from: Morphological identification and single-cell genomics of marine diplonemids
Open the record for dataset details and reuse information.
Systematic dissection of transcriptional regulatory networks by genome-scale and single-cell CRISPR screens
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.