Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

2,598

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

2,598 results for “single cell sequencing”

Learn how ShareScore rates datasets ↗
zenodo48/100

Clonal decomposition and DNA replication states defined by scaled single cell genome sequencing

<p><strong>OV2295&nbsp;Tables</strong></p> <p>ov2295_breakpoint_counts.csv.gz: Table of breakpoint counts per cell</p> <ul> <li>prediction_id: identifier for the breakpoint</li> <li>cell_id: identifier for the cell</li> <li>read_count: number of reads</li> <li>library_id: identifier for the DNA library</li> <li>sample_id: identifier for the sequenced sample</li> <li>chromosome_1: chromosome of breakend 1</li> <li>strand_1: orientation of break end 1</li> <li>position_1: position of break end 1</li> <li>chromosome_2: chromosome of breakend 2</li> <li>strand_2: orientation of break end 2</li> <li>position_2: position of break end 2</li> </ul> <p>ov2295_cell_cn.csv.gz: Table of cell specific copy number</p> <ul> <li>cell_id: identifier for the cell</li> <li>sample_id: identifier for the sequenced sample</li> <li>library_id: identifier for the DNA library</li> <li>chr: chromosome of bin</li> <li>start: start of bin</li> <li>end: end of bin</li> <li>reads: number of reads</li> <li>copy: raw normalized copy number</li> <li>state: copy number state</li> <li>gc: percent gc of the bin</li> <li>map: average mappability of the bin</li> </ul> <p>ov2295_cell_metrics.csv.gz: Table of cell metrics</p> <ul> <li>cell_id: identifier of the cell</li> <li>unpaired_mapped_reads: number of unpaired mapped reads</li> <li>paired_mapped_reads: number of mapped reads that were properly paired</li> <li>unpaired_duplicate_reads: number of unpaired duplicated reads</li> <li>paired_duplicate_reads: number of paired reads that were also marked as duplicate</li> <li>unmapped_reads: number of unmapped reads</li> <li>percent_duplicate_reads: percentage of duplicate reads</li> <li>estimated_library_size: scaled total number of mapped reads</li> <li>total_reads: total number of reads, regardless of mapping status</li> <li>total_mapped_reads: total number of mapped reads</li> <li>total_duplicate_reads: number of duplicate reads</li> <li>total_properly_paired: number of properly paired reads</li> <li>coverage_breadth: percentage of genome covered by some read</li> <li>coverage_depth: average reads per nucleotide position in the genome</li> <li>median_insert_size: median insert size between paired reads</li> <li>mean_insert_size: mean insert size between paired reads</li> <li>standard_deviation_insert_size: standard deviation of the insert size between paired reads</li> <li>index_sequence: index sequence of the adaptor sequence</li> <li>column: column of the cell on the nanowell chip</li> <li>img_col: column of the cell from the perspective of the microscope</li> <li>index_i5: id of the i5 index adapter sequence</li> <li>sample_type: type of the sample</li> <li>primer_i7: id of the i5 index primer sequence</li> <li>experimental_condition: experimental treatment of the cell, includes controls</li> <li>index_i7: id of the i7 index adapter sequence</li> <li>cell_call: living/dead classification of the cell based on staining usually, C1 == living, C2 == dead</li> <li>sample_id: name of the sample</li> <li>primer_i5: id of the i5 index primer sequence</li> <li>row: row of the cell on the nanowell chip</li> <li>library_id: identifier for the DNA library</li> <li>index: ignored</li> <li>multiplier: during parameter searching, the set [1..6] that was chosen</li> <li>MSRSI_non_integerness: median of segment residuals from segment integer copy number states</li> <li>MBRSI_dispersion_non_integerness: median of bin residuals from segment integer copy number states</li> <li>MBRSM_dispersion: median of bin residuals from segment median copy number values</li> <li>autocorrelation_hmmcopy: hmmcopy copy autocorrelation</li> <li>cv_hmmcopy: ignored</li> <li>empty_bins_hmmcopy: number of empty bins in hmmcopy</li> <li>mad_hmmcopy: median absolute deviation of hmmcopy copy</li> <li>mean_hmmcopy_reads_per_bin: mean reads per hmmcopy bin</li> <li>median_hmmcopy_reads_per_bin: median reads per hmmcopy bin</li> <li>std_hmmcopy_reads_per_bin: standard deviation value of reads in hmmcopy bins</li> <li>total_halfiness: summed halfiness penality score of the cell</li> <li>total_mapped_reads_hmmcopy: total mapped reads in all hmmcopy bins</li> <li>scaled_halfiness: summed scaled halfiness penalty score of the cell</li> <li>mean_state_mads: mean value for all median absolute deviation scores for each state</li> <li>mean_state_vars: variance value for all median absolute deviation scores for each state</li> <li>mad_neutral_state: median absolute deviation score of the neutral 2 copy state</li> <li>breakpoints: number of breakpoints, as indicated by state changes not at the ends of chromosomes</li> <li>mean_copy: mean hmmcopy copy value</li> <li>state_mode: the most commonly occuring state</li> <li>log_likelihood: hmmcopy log likelihood for the cell</li> <li>true_multiplier: the exact decimal value used to scale the copy number for segmentation</li> <li>order: order of the cell in the hierarchical clustering tree</li> <li>quality: random forest classifier proability score that cell is good</li> </ul> <p>ov2295_clone_alleles.csv.gz: Table of clone specific allele data</p> <ul> <li>chr: chromosome of bin</li> <li>start: start of bin</li> <li>end: end of bin</li> <li>hap_label: haplotype block identifier</li> <li>clone_id: clone identifier</li> <li>allele_1_sum: number of reads for allele 1 of the haplotype block</li> <li>allele_2_sum: number of reads for allele 2 of the haplotype block</li> <li>total_counts_sum: total reads for the haplotype block</li> </ul> <p>ov2295_clone_breakpoints.csv.gz:&nbsp;Table of breakpoints per clone for OV2295 samples. Columns:</p> <ul> <li>prediction_id: identifier for the breakpoint</li> <li>chromosome_1: chromosome of breakend 1</li> <li>strand_1: orientation of break end 1</li> <li>position_1: position of break end 1</li> <li>chromosome_2: chromosome of breakend 2</li> <li>strand_2: orientation of break end 2</li> <li>position_2: position of break end 2</li> <li>clone_id: clone identifier</li> <li>read_count: number of reads</li> <li>is_present: presence=1, absent=0</li> </ul> <p>ov2295_clone_clusters.csv.gz: Table of cell clusters as putative clones</p> <ul> <li>cell_id: identifier for the cell</li> <li>clone_id: clone identifier</li> </ul> <p>ov2295_clone_cn.csv.gz:&nbsp;Table of allele specific copy number per clone for OV2295 samples. Columns:</p> <ul> <li>chr: chromosome of bin</li> <li>start: start of bin</li> <li>end: end of bin</li> <li>total_cn: HMMCopy predicted total copy number&nbsp;</li> <li>minor_cn: HMM predicted minor copy number&nbsp;</li> <li>major_cn: HMM predicted major copy number&nbsp;</li> <li>clone_id: clone identifier</li> </ul> <p>ov2295_clone_snvs.csv.gz:&nbsp;Table of SNVs per clone for OV2295 samples.&nbsp; Columns:</p> <ul> <li>chrom: chromosome</li> <li>coord: genome position</li> <li>ref: reference nucleotide</li> <li>alt: alternate nucleotide</li> <li>clone_id: clone identifier</li> <li>ref_counts: number of reads at this position matching the reference nucleotide</li> <li>alt_counts: number of reads at this position matching the alternate nucleotide</li> <li>total_counts: total number of reads at this position</li> <li>is_present: presence=0, absent=1</li> <li>is_het:&nbsp;is heterozygous</li> <li>is_hom: is homozygous for the alternate</li> </ul> <p>ov2295_nodes.csv.gz: Table of phylogenetic information for SNV evolution</p> <ul> <li>variant_id: identifier for the SNV as chrom:coord:ref:alt</li> <li>node: node in the phylogenetic tree</li> <li>loss: probability the SNV was lost at this node</li> <li>origin: probability the SNV originated at this node</li> <li>presence: probability the SNV is present at this node</li> <li>ml_origin: binary indicator the SNV originated at this node</li> <li>ml_presence: binary indicator the SNV is present at this node</li> <li>ml_loss: binary indicator the SNV was lost at this node</li> </ul> <p>ov2295_snv_counts.csv.gz: Table of SNV counts</p> <ul> <li>chrom: chromosome</li> <li>coord: genome position</li> <li>ref: reference nucleotide</li> <li>alt: alternate nucleotide</li> <li>ref_counts: number of reads at this position matching the reference nucleotide</li> <li>alt_counts: number of reads at this position matching the alternate nucleotide</li> <li>cell_id: identifier for the cell</li> <li>total_counts: total number of reads at this position</li> <li>sample_id: identifier for the sequenced sample</li> </ul> <p>ov2295_tree.pickle: Phylogenetic tree in python pickle format.&nbsp; Requires installation of the stochastic dollo code at: https://bitbucket.org/dranew/dollo, version 0.4.2.</p> <p>Note the following sample mapping: &lsquo;SA922&rsquo;: &lsquo;OV2295(R2)&rsquo;, &lsquo;SA921&rsquo;: &lsquo;TOV2295(R)&rsquo;, &lsquo;SA1090&rsquo;: &lsquo;OV2295&rsquo;,</p> <p><strong>Plots</strong></p> <p>ov_supp_clone_allele_cn.png: Clone allele ratios for each OV2295 sample.</p> <p>ov_supp_clone_total_cn.png: Clone copy number for each OV2295 sample.</p> <p>ov_supp_sample_total_cn.png: Bulk copy number for each OV2295 sample.</p> <p>ov_supp_sample_allele_cn.png: Bulk allele ratios for each OV2295 sample.</p>

opencc-by-4.0Sep 2019View details →
zenodo44/100

Raw read counts and phased SNP counts for every single cell in the sequencing datasets of the breast cancer patient S1 from "Characterizing allele- and haplotype-specific copy numbers in single cells with CHISEL"

<p>This dataset contains the raw read counts and phased SNP counts&nbsp;for every single cell in the sequencing datasets of breast cancer patient S1 from &ldquo;Characterizing allele- and haplotype-specific copy numbers in single cells with CHISEL&rdquo; [Zaccaria &amp; Raphael, 2020]. These data enable the full reproduction of all the results in the related manuscript for breast cancer patient S1. Specifically, the data are provided in two files for every dataset <em>DAT</em> of patient S1 with the following format:</p> <ol> <li><em>DAT.raw</em>_<em>read</em>_<em>counts.bed.gz&nbsp;</em>is a multi-cell BED file containing the raw read counts in the following fields: <ul> <li>CHROMOSOME: the name of a human chromosome</li> <li>START: the starting genomic position of a genomic bin in the chromosome</li> <li>END: the ending genomic position of the genomic bin in the chromosome</li> <li>CELL: the cell barcode that uniquely identifies a cell</li> <li>NORMAL: the raw read count&nbsp;for the specified bin from a matched-normal sample</li> <li>COUNT: the raw read count&nbsp;for the specified bin in the specified cell</li> <li>RDR: the estimated read-depth ratio for the specified bin in the specified cell</li> </ul> </li> <li><em>DAT.phased</em>_<em>snps</em>_<em>counts.pos.gz&nbsp;</em>is a multi-cell POS file containing the phased SNP counts in the following fields: <ul> <li>CHROMOSOME: the name of a human chromosome</li> <li>POS: the genomic position in the chromosome of a germline SNP</li> <li>CELL: the cell barcode that uniquely identifies a cell</li> <li>COUNT_HAPLOTYPE_A: the count of reads that cover&nbsp;the SNP and that belong to haplotype A in the specified cell</li> <li>COUNT_HAPLOTYPE_B: the count of reads that cover&nbsp;the SNP and that belong to haplotype B in the specified cell</li> </ul> </li> </ol> <p>All the files have been compressed using standard <em>gzip</em>.</p>

opencc-by-4.0May 2020View details →
zenodo44/100

Raw read counts and phased SNP counts for every single cell in the sequencing datasets of the breast cancer patient S0 from "Characterizing allele- and haplotype-specific copy numbers in single cells with CHISEL"

<p>This dataset contains the raw read counts and phased SNP counts&nbsp;for every single cell in the sequencing datasets of breast cancer patient S0 from &ldquo;Characterizing allele- and haplotype-specific copy numbers in single cells with CHISEL&rdquo; [Zaccaria &amp; Raphael, 2020]. These data enable the full reproduction of all the results in the related manuscript for breast cancer patient S0. Specifically, the data are provided in two files for every dataset <em>DAT</em> of patient S0 with the following format:</p> <ol> <li><em>DAT.raw</em>_<em>read</em>_<em>counts.bed.gz&nbsp;</em>is a multi-cell BED file containing the raw read counts in the following fields: <ul> <li>CHROMOSOME: the name of a human chromosome</li> <li>START: the starting genomic position of a genomic bin in the chromosome</li> <li>END: the ending genomic position of the genomic bin in the chromosome</li> <li>CELL: the cell barcode that uniquely identifies a cell</li> <li>NORMAL: the raw read count&nbsp;for the specified bin from a matched-normal sample</li> <li>COUNT: the raw read count&nbsp;for the specified bin in the specified cell</li> <li>RDR: the estimated read-depth ratio for the specified bin in the specified cell</li> </ul> </li> <li><em>DAT.phased</em>_<em>snps</em>_<em>counts.pos.gz&nbsp;</em>is a multi-cell POS file containing the phased SNP counts in the following fields: <ul> <li>CHROMOSOME: the name of a human chromosome</li> <li>POS: the genomic position in the chromosome of a germline SNP</li> <li>CELL: the cell barcode that uniquely identifies a cell</li> <li>COUNT_HAPLOTYPE_A: the count of reads that cover&nbsp;the SNP and that belong to haplotype A in the specified cell</li> <li>COUNT_HAPLOTYPE_B: the count of reads that cover&nbsp;the SNP and that belong to haplotype B in the specified cell</li> </ul> </li> </ol> <p>All the files have been compressed using standard <em>gzip</em>.</p>

opencc-by-4.0May 2020View details →
zenodo44/100

Development of Multiomics in situ Pairwise Sequencing (MiP-Seq) for Single-cell Resolution Multidimensional Spatial Omics

<p>The original data used in the article:&nbsp;Development of Multiomics in situ Pairwise Sequencing (MiP-Seq) for Single-cell Resolution Multidimensional Spatial Omics</p> <p>Delineating the spatial multiomics landscape will pave the way to understanding the molecular basis of physiology and pathology. However, current spatial omics technology development is still in its infancy. Here, we developed a high-throughput targeted in situ sequencing strategy, multiomics in situ pairwise sequencing (MiP-Seq), to efficiently decipher multiplexed DNAs, RNAs, proteins, and small biomolecules at subcellular resolution. MiP-Seq simultaneously sequenced the dual barcode base of padlock probes, dramatically increasing the detection capacity to 10N by N rounds of sequencing. We delineated spatial gene profiles in the hypothalamus using MiP-Seq. Moreover, MiP-Seq was unitized to detect tumor gene mutations and allele-specific expression of parental genes and to differentiate sites with and without the m6A RNA modification at specific sites. MiP-Seq was combined with in vivo Ca2+ imaging and Raman imaging to obtain a spatial multiomics atlas correlated to neuronal activity and cellular biochemical fingerprints. Importantly, we proposed a &ldquo;signal dilution strategy&rdquo; to resolve the crowded signals that challenge the applicability of in situ sequencing. Together, our method improves spatial multiomics and precision diagnostics, and facilitates analyzing cell function in connection with gene profiles.</p>

opencc-by-4.0Jan 2023View details →
zenodo44/100

Joint embedding of vertebrate brain single-cell RNA-Seq using sequence or structure

<p>Embeddings of single-cell RNA-Seq data from three adult vertebrate brain datasets into Orthogroup feature space or Structural cluster feature space. Orthogroups were generated using OrthoFinder v5.5.0; Structural clusters were assigned by using FoldSeek to cluster AlphaFold-v4 structural predictions.<br> <br> The three datasets used as the basis for these embeddings were:</p> <ul> <li>sample&nbsp;<a href="https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSM3768152">&quot;Brain8&quot;</a>&nbsp;from the&nbsp;<a href="https://www.frontiersin.org/articles/10.3389/fcell.2021.743421/full">Jiang et al. 2021</a>&nbsp;zebrafish cell atlas (files beginning with&nbsp;GSM3768152)</li> <li>sample&nbsp;<a href="https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSM2906405">&quot;Brain1&quot;</a>&nbsp;from the&nbsp;<a href="https://www.sciencedirect.com/science/article/pii/S0092867418301168#sec4">Han et al. 2018</a>&nbsp;mouse cell atlas (files beginning with&nbsp;GSM2906405)</li> <li>sample&nbsp;<a href="https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSM6214268">&quot;Xenopus_brain_COL65&quot;</a>&nbsp;from the&nbsp;<a href="https://www.nature.com/articles/s41467-022-31949-2">Liao et al. 2022</a>&nbsp;Xenopus laevis adult cell atlas (files beginning with GSM6214268)</li> </ul> <p>For each dataset, we also generated a standardized cell type annotation file based on the author&#39;s originally provided cell type annotation data. The first column is the cell barcode for that species and the second column is the original study&#39;s cell type annotation for that cell.</p> <p>For the Xenopus brain data, we removed around ~18k cells that were not annotated in the original data to simplify data analyses - these are reflected in the files with the &quot;subsampled&quot; suffix. Subsampled versions of the data are also available for the joint embedding space (prefixed with &quot;DrerMmusXlae&quot;).</p> <p>For the final datasets used in our analyses, we also provide features x cell matrices as .h5ad files for smaller file sizes and faster loading using Scanpy.&nbsp;</p> <p>For visualizing our UMAP plots of our top200 embedding space, we provide &quot;.tsv&quot; files with a variety of metrics and the x and y positions of each cell in the UMAP. See &quot;DrerMmusXlae_adultbrain_FoldSeek_plotlydata.tsv&quot; and &quot;DrerMmusXlae_adultbrain_OrthoFinder_plotlydata.tsv&quot;</p> <p>These data are part of the Arcadia Science Pub titled <a href="https://doi.org/10.57844/arcadia-vw5e-2670">&quot;Comparing gene expression across species based on protein structure instead of sequence&quot;</a>.</p>

opencc-by-4.0Apr 2023View details →
zenodo44/100

Reads-per-UMI tables across single-cell RNA sequencing protocols

<p>Data analyzed in <a href="https://www.biorxiv.org/content/10.1101/2023.08.02.551637v1">Lause, Ziegenhain et al. (2023)</a>.</p> <p>Code to obtain these tables from public data sources is available on <a href="https://github.com/berenslab/read-normalization">github</a>.</p> <p>&nbsp;</p> <p>Each row in the table is a UMI-tag detected in a certain cell (column RG) attached to a molecule from a specific gene (column GE) with a certain barcode (column UB). Column N gives the number of times the UMI was detected for that gene and cell.</p> <p>Data sources and protocols are given with the respective file names below.</p> <p><strong>Johnsson2022_Smartseq3_PE.hd1.txt.gz</strong>: Mouse fibroblasts profiled with <strong>Smart-seq3</strong> paired-end; accession E-MTAB-10148, sample plate2,<br> <a href="https://doi.org/10.1038/s41588-022-01014-1">Paper</a><br> <br> <strong>Hagemann-Jensen2020_Smartseq3_SE.hd1.txt.gz: </strong>Mouse fibroblasts profiled with <strong>Smart-seq3</strong> single-end; accession E-MTAB-8735, sample Smartseq3.Fibroblasts.smFISH<br> <a href="https://doi.org/10.1038/s41587-020-0497-0">Paper</a><br> <br> <strong>Hagemann-Jensen2022_Smartseq3xpress.hd1.txt.gz: </strong>HEK293 cells profiled with <strong>Smart-seq3Xpress</strong>; accession E-MTAB-11467.<br> <a href="https://www.biorxiv.org/content/10.1101/2021.07.10.451889v1">Paper</a><br> <br> <strong>Ziegenhain2017.hd1.txt.gz: </strong>Mouse embryonic stem cells profiled by <strong>CEL-seq2, Drop-seq, MARS-seq, </strong>and<strong> SCRB-seq</strong>; GEO accession GSE75790<br> <a href="https://doi.org/10.1016/j.molcel.2017.01.023">Paper</a></p>

opencc-by-4.0Jul 2023View details →
zenodo40/100

Direct chromosome-length haplotyping by single-cell sequencing.

<p>Selected Strand-seq libraries from PMID:27646535 study. Data were originally shared on the European Nucleotide Archive (http://www.ebi.ac.uk/ena) under the accession number: PRJEB14185</p>

opencc-by-4.0Nov 2016View details →
zenodo40/100

Single-cell RNA sequencing identifies shared differentiation paths of mouse thymic innate T cells

<p>scRNA sequencing&nbsp;datasets&nbsp;used in the paper titled &#39;Single-cell RNA sequencing identifies shared differentiation paths of mouse thymic innate T cells&#39; published in Nature Communications<br> <br> https://www.nature.com/articles/s41467-020-18155-8</p>

opencc-by-4.0Jul 2023View details →
zenodo40/100

Full-length, single-cell RNA-sequencing of human bone marrow subpopulations reveals hidden complexity

<p><a href="http://www.biorxiv.org/content/10.1101/2021.07.28.454226v2">Full-length, single-cell RNA-sequencing of human bone marrow subpopulations reveals hidden complexity</a></p> <p>Bone marrow progenitor cell differentiation has frequently been used as a model for studying cellular plasticity and cell-fate decisions. Recent analysis at the level of single-cells has expanded knowledge of the transcriptional landscape of human hematopoietic cell lineages. Using single-molecule real-time (SMRT) full-length RNA sequencing, we have previously shown that human bone marrow lineage-negative (Lin-neg) cell populations contain a surprisingly diverse set of mRNA isoforms. Here, we report from single cell, full-length RNA sequencing that this diversity is also reflected at the single-cell level. From fresh human bone marrow unselected and lineage-negative progenitor cells were isolated by droplet-based single-cell selection (10xGenomics). The single cell-derived mRNAs were analyzed by full-length SMRT and short-read sequencing. In both samples we detected an average of 8000 different genes using short-read sequencing. Differential expression analysis arranged the single-cells of the total bone marrow into only four clusters whereas the Lin-neg population was much more diverse with nine clusters. mRNA isoform analysis of the single-cell populations using full-length sequencing revealed that Lin-neg cells contain on average 24% more novel splice variants than the total bone marrow cells. Interestingly, among the most frequent genes expressing novel isoforms were members of the spliceosome, e.g. HNRNPs, DEAD box helicases and SRSFs. Mapping the isoforms from all genes to the cell type clusters revealed that total bone marrow cells express novel isoforms only in a small subset of clusters. On the other hand, lineage-negative progenitor cells expressing novel isoforms were present in nearly all subpopulations. In conclusion, on a single-cell level lineage-negative cells express a higher diversity of genes and more alternatively spliced novel isoforms suggesting that cells in this subpopulation are poised for different fates.&nbsp;&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Sep 2022View details →
zenodo40/100

Single-cell sequencing data of human umbilical cord and placental mesenchymal stem cells

<p>Expression matrix of umbilical cord and placenta single-cell sequencing data from the same donor.Table1 is the umbilical cord and Table2 is the placenta.</p>

opencc-by-4.0Apr 2024View details →
zenodo40/100

Harnessing single cell RNA sequencing to identify dendritic cell types, characterize their biological states and infer their activation trajectory

<p><strong>Summary: </strong>Dendritic cells (DCs) orchestrate innate and adaptive immunity, by translating the sensing of distinct danger signals into the induction of different effector lymphocyte responses, to induce different defense mechanisms suited to face distinct types of threats. Hence, DCs are very plastic, which results from two key characteristics. First, DCs encompass distinct cell types specialized in different functions. Second, each DC type can undergo different activation states, fine-tuning its functions depending on its tissue microenvironment and the pathophysiological context, by adapting the output signals it delivers to the input signals it receives. Hence, to better understand DC biology and harness it in the clinic, we must determine which combinations of DC types and activation states mediate which functions, and how.<br> To decipher the nature, functions and regulation of DC types and their physiological activation states, one of the methods that can be harnessed most successfully is ex vivo single cell RNA sequencing (scRNAseq). However, for new users of this approach, determining which analytics strategy and computational tools to choose can be quite challenging, considering the rapid evolution and broad burgeoning of the field. In addition, awareness must be raised on the need for specific, robust and tractable strategies to annotate cells for cell type identity and activation states. It is also important to emphasize the necessity of examining whether similar cell activation trajectories are inferred by using different, complementary methods. In this chapter, we take these issues into account for providing a pipeline for scRNAseq analysis and illustrating it with a tutorial reanalyzing a public dataset of mononuclear phagocytes isolated from the lungs of na&iuml;ve or tumor-bearing mice. We describe this pipeline step-by-step, including data quality controls, dimensionality reduction, cell clustering, cell cluster annotation, inference of the cell activation trajectories and investigation of the underpinning molecular regulation. It is accompanied with a more complete tutorial on Github. We anticipate that this method will be helpful for both wet lab and bioinformatics researchers interested in harnessing scRNAseq data for deciphering the biology of DCs or other cell types, and that it will contribute to establishing high standards in the field.</p> <p>&nbsp;</p> <p><strong>Data:</strong></p> <p>1. negative_cDC1_relative_signatures.csv : Negative signatures for performing Connectivity Map (cMAP) Analysis</p> <p>2. positive_cDC1_relative_signatures.csv : Positive signatures for performing Connectivity Map (cMAP) Analysis</p>

opencc-by-4.0Dec 2021View details →
zenodo40/100

Integrated Data of Single cell RNA sequencing for Human Pancreatic Adenocarcinoma

<p>These data are collected and integrated from five available deposit data and one original data of single cell RNA sequencing from human pancreatic adenocarcinoma. Further analyses data for bulk transcriptomics (such as TCGA )using scRNAseq data and re-clustering for ductal epithelial cells and fibroblasts are also stored in step by step. Moreover, all R code is uploaded.</p>

opencc-by-4.0Feb 2022View details →
zenodo40/100

Data files: Single-cell RNA sequencing of Plasmodium vivax sporozoites reveals stage- and species-specific transcriptomic signatures

<p>Scripts, preprocessed count matrices, single-cell data objects, and generated data (tables and .rds files)&nbsp;from the scRNA-seq analyses performed in&nbsp;<strong>&ldquo;Single-cell RNA sequencing of Plasmodium vivax sporozoites reveals stage- and species-specific transcriptomic signatures&quot;.</strong></p> <p>&nbsp;</p>

opencc-by-4.0Apr 2022View details →
zenodo40/100

Phertilizer: growing a clonal tree from single-cell DNA sequencing data of tumors

<p>The is the supplementary data repository for the simulation input data for&nbsp;Phertilizer: growing a clonal tree from single-cell DNA sequencing data of tumors.</p>

opencc-by-4.0Apr 2022View details →
zenodo40/100

Single-cell and single-nucleus RNA-sequencing from paired normal-adenocarcinoma lung samples provides both common and discordant biological insights

<p>The datasets generated by&nbsp;<em>Cellranger </em>for all 24 samples (.h5 format).<br><br></p>

opencc-by-4.0May 2024View details →
zenodo40/100

FedscGen: privacy-aware federated batch effect correction of single-cell RNA sequencing data -- Preprocessed datasets

<div> <div> <div> <div> <p>This dataset accompanies the publication "FedscGen: Privacy-Aware Federated Batch Effect Correction of Single-Cell RNA Sequencing Data" and includes eight single-cell RNA sequencing (scRNA-seq) datasets used to benchmark the FedscGen and scGen methods. The datasets are provided in <code>.h5ad</code> format and include comprehensive metadata necessary for replication and further analysis.</p> <h3>Datasets</h3> <p>We analyze various datasets to compare FedscGen against scGen (centralized) in terms of batch correction. For simplicity, we refer to the dataset by abbreviations:</p> <ol> <li> <p><strong>Cell Line (CL)</strong>:</p> <ul> <li>Derived from the 293t_jurkat experiment with three batches: Zheng et al., 2017.</li> </ul> </li> <li> <p><strong>Human Dendritic Cells (HDC)</strong>:</p> <ul> <li>scRNA-seq data of human dendritic cells across two batches: Villani et al., 2017.</li> </ul> </li> <li> <p><strong>Human Pancreas (HP)</strong>:</p> <ul> <li>Consolidated data from five sources with 14,767 cells each: Baron et al., 2016; Muraro et al., 2016; Segerstolpe et al., 2016; Wang et al., 2016; Xin et al., 2016.</li> </ul> </li> <li> <p><strong>Mouse Brain (MB)</strong>:</p> <ul> <li>Merged datasets with 691,600 and 141,606 cells: Saunders et al., 2018; Rosenberg et al., 2018.</li> </ul> </li> <li> <p><strong>Mouse Cell Atlas (MCA)</strong>:</p> <ul> <li>Data focusing on 11 cell types from various organs: Han et al., 2018; The Tabula Muris Consortium, 2018.</li> </ul> </li> <li> <p><strong>Mouse Hematopoietic Stem and Progenitor Cells (MHSPC)</strong>:</p> <ul> <li>Data from SMART-seq2 and MARS-seq protocols: Nestorowa et al., 2016; Paul et al., 2015.</li> </ul> </li> <li> <p><strong>Mouse Retina (MR)</strong>:</p> <ul> <li>Data from two unassociated laboratories with 26,830 and 44,808 cells: Macosko et al., 2015; Shekhar et al., 2016.</li> </ul> </li> <li> <p><strong>PBMC (human Peripheral Blood Mononuclear Cell)</strong>:</p> <ul> <li>scRNA-seq data with two batches: Zheng et al., 2017.</li> </ul> </li> </ol> <p><strong>Usage Notes</strong>: Each dataset is provided in <code>.h5ad</code> format, compatible with common single-cell analysis tools such as Scanpy. Detailed metadata is included within each file.</p> <p><strong>Keywords</strong>: Single-cell RNA sequencing, scRNA-seq, Batch effect correction, Privacy-aware, Federated learning, scGen, FedscGen, Clinical multi-center studies, Genomics, Bioinformatics</p> <p><strong>Contact</strong>: For questions or further information, please contact Mohammad Bakhtiari at <a href="mailto:mohammad.bakhtiari@uni-hamburg.de.">mohammad.bakhtiari@uni-hamburg.de.</a></p> <p><strong>License</strong>: Creative Commons Attribution 4.0 International (CC BY 4.0)</p> </div> </div> </div> </div> <div> <div> <div>&nbsp;</div> </div> </div>

opencc-by-4.0Jun 2024View details →
zenodo40/100

Supplementary data for: Detection of expressed mutations in acute myeloid leukemia cells using single cell RNA-sequencing

<p>Supplemental data for the publication:<br> Detection of expressed mutations in acute myeloid leukemia cells using single cell RNA-sequencing&nbsp;</p> <p>Contents:&nbsp;<br> - expression_matrices.tar&nbsp; -&nbsp;Gene/Barcode expression matrices from `cellranger count`<br> - *.seurat.rds&nbsp; - R object files&nbsp;with Seurat analyses and data structures for each sample<br> - scrna_mutations.tar.gz&nbsp; -&nbsp;copy of a git repository&nbsp;containing additional scripts and data - also hosted at&nbsp;<a href="https://github.com/genome/scrna_mutations">https://github.com/genome/scrna_mutations</a>&nbsp;(snapshot as&nbsp;of May&nbsp;20, 2019)</p>

opencc-by-4.0May 2019View details →
zenodo40/100

A comparison of automatic cell identification methods for single-cell RNA-sequencing data

<p>Benchmark datasets used to evaluate the performance of 22 classifiers for cell type classification for scRNA-seq data</p>

opencc-by-4.0Dec 2018View details →
zenodo40/100

Nonlinear methods for dimensionality reduction and clustering of bacterial single-cell sequencing data - intermediate data and figures (MSc thesis)

<p>Data, intermediate results and figures for analyses of my master's thesis in biostatistics at LMU Munich. I took a look on how to use Nonlinear Matrix Decomposition (NMD) (<a href="https://doi.org/10.1137/21M1405769">Saul, L., 2022</a>) in the context of bacterial scRNA-seq analysis (Heumos, L., et. al. 2023), replacing Principal Component Analysis in the optimized workflow, as outlined in Ostner, J. (2024).</p> <p>My thesis was structured along the following objectives:</p> <ul> <li>implement the algorithms from <a href="https://arxiv.org/abs/2305.08687">Seraghiti, G., et. al. (2023)</a> in the Python module <a href="https://github.com/flatironinstitute/nomad/">nomad</a> in cooperation with <a href="https://www.simonsfoundation.org/flatiron/" rel="nofollow">Flatiron Institute</a></li> <li>code for the simulation study of the algorithms in <a href="https://arxiv.org/abs/2305.08687">Seraghiti, G., et. al. (2023)</a> with varying sparsity can be found in <code>/simulation</code></li> <li>apply NMD in the context of the BacSC workflow (<a href="https://www.biorxiv.org/content/10.1101/2024.06.22.600071v1">Ostner, J., et. al. (2024)</a>) on raw and normalized counts (found in <code>/application/analysis</code>), also for manually set number of latent dimensions</li> <li>explore NMD's potential for imputation of <a href="https://www.nature.com/articles/s41467-021-27729-z" rel="nofollow">sampling zeros</a> (check <code>/application/NMD_zero_imputation /</code>)</li> <li>potential of Poisson-Hurdle model-based clustering (<a href="https://academic.oup.com/bioinformatics/article/39/1/btac782/6873739">Qiao, Z., et. al. (2023)</a>) for scRNA-seq (<code>/application/poisson_hurdle</code>).</li> </ul>

opencc-by-4.0Oct 2024View details →
zenodo40/100

Additional files of scTensor paper "scTensor detects many-to-many cell-cell interactions from single cell RNA-sequencing data"

<p>Complex biological systems are described as a multitude of cell-cell interactions (CCIs). Recent single-cell RNA-sequencing studies focus on CCIs based on ligand-receptor (L-R) gene co-expression. However, the analytical methods are still not mature; such methods cannot detect CCIs and the related L-R pairs simultaneously or also are not appropriate to detect many-to-many CCIs.</p> <p>In this work, we propose scTensor, a novel method for extracting representative triadic relationships (or hypergraphs), which include ligand-expression, receptor-expression, and related L-R pairs. Through extensive studies with simulated and empirical datasets, we have shown that scTensor could detect some hypergraphs, which cannot be detected by conventional methods, especially when those CCIs are many-to-many relationships.</p>

opencc-by-4.0Sep 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record