Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

11,090

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

11,090 results for “single cell”

Learn how ShareScore rates datasets ↗
zenodo52/100

Data for "Profiling the transcriptomic age of single-cells in humans"

<p>This is a supplementary data for the article titled "Profiling transcriptomic age of human single-cells". Data created in this project is shared here for the scientific community.&nbsp;</p> <p>Here we used available scRNA-seq data of 1,058,909 blood cells of 508 healthy, human donors, for developing cell-type-specific single-cell transcriptomic clocks and predicting the age of human blood cells. &nbsp;We also applied our clocks to different external datasets and evaluated the age of single cells originated from COVID-19 patients and human embryos.</p> <p>For the description of the content of the dataset see the ReadMe file.</p>

opencc-by-4.0Jun 2024View details →
zenodo48/100

Bulk and single-cell gene expression profiling of SARS-CoV-2 infected human cell lines identifies molecular targets for therapeutic intervention

<p>Single cell RNA seq datasets used for analysis in the&nbsp;Bulk and single-cell gene expression profiling of SARS-CoV-2 infected human cell lines identifies molecular targets for therapeutic intervention</p>

opencc-by-4.0Sep 2020View details →
zenodo48/100

The evolution of genomic, transcriptomic, and single-cell protein markers of metastatic upper tract urothelial carcinoma

<p>The molecular characteristics of metastatic upper tract urothelial carcinoma (UTUC) are unknown. The genomic and transcriptomic differences between primary and metastatic UTUC is not well described either. We combined whole-exome sequencing, RNA-sequencing, and Imaging Mass Cytometry<sup>TM</sup>&nbsp;(IMC<sup>TM</sup>) of 44 tumor samples from 28 patients with high-grade primary and metastatic UTUC. IMC enables spatially resolved single-cell analyses to examine the evolution of cancer cell, immune cell, and stromal cell markers using mass cytometry with lanthanide metal-conjugated antibodies. We discovered that actionable genomic alterations are frequently discordant between primary and metastatic UTUC tumors in the same patient. In contrast, molecular subtype membership and immune depletion signature were stable across primary and matched metastatic UTUC. Molecular and immune subtypes were consistent between bulk RNA-sequencing and mass cytometry of protein markers from 340,798 single-cells. Molecular subtyping at the single cell level was highly conserved between primary and metastatic UTUC tumors within the same patient.</p>

opencc-by-4.0Nov 2021View details →
zenodo48/100

CAbiNet: Joint clustering and visualization of cells and genes for single-cell transcriptomics

<p>We here provide the data sets to reproduce the results in our manuscript "CAbiNet: Joint clustering and visualization of cells and genes for single-cell transcriptomics". Our package "CAbiNet" can be downloaded from https://github.com/VingronLab/CAbiNet. The scripts to reproduce the results in our manuscript can be found from https://github.com/VingronLab/CAbiNet_paper.</p><p>You can find the description of folders in 'Data.zip' in the README.md file.</p>

opencc-by-4.0Dec 2023View details →
zenodo48/100

Joint Trajectory Inference for Single-cell Genomics Using Deep Learning with a Mixture Prior

<p>The datasets used in the paper "Joint Trajectory Inference for Single-cell Genomics Using Deep Learning with a Mixture Prior". A detailed description of these datasets is available at https://github.com/jaydu1/VITAE/tree/master/data.</p>

opencc-by-4.0Dec 2020View details →
zenodo48/100

Clonal decomposition and DNA replication states defined by scaled single cell genome sequencing

<p><strong>OV2295&nbsp;Tables</strong></p> <p>ov2295_breakpoint_counts.csv.gz: Table of breakpoint counts per cell</p> <ul> <li>prediction_id: identifier for the breakpoint</li> <li>cell_id: identifier for the cell</li> <li>read_count: number of reads</li> <li>library_id: identifier for the DNA library</li> <li>sample_id: identifier for the sequenced sample</li> <li>chromosome_1: chromosome of breakend 1</li> <li>strand_1: orientation of break end 1</li> <li>position_1: position of break end 1</li> <li>chromosome_2: chromosome of breakend 2</li> <li>strand_2: orientation of break end 2</li> <li>position_2: position of break end 2</li> </ul> <p>ov2295_cell_cn.csv.gz: Table of cell specific copy number</p> <ul> <li>cell_id: identifier for the cell</li> <li>sample_id: identifier for the sequenced sample</li> <li>library_id: identifier for the DNA library</li> <li>chr: chromosome of bin</li> <li>start: start of bin</li> <li>end: end of bin</li> <li>reads: number of reads</li> <li>copy: raw normalized copy number</li> <li>state: copy number state</li> <li>gc: percent gc of the bin</li> <li>map: average mappability of the bin</li> </ul> <p>ov2295_cell_metrics.csv.gz: Table of cell metrics</p> <ul> <li>cell_id: identifier of the cell</li> <li>unpaired_mapped_reads: number of unpaired mapped reads</li> <li>paired_mapped_reads: number of mapped reads that were properly paired</li> <li>unpaired_duplicate_reads: number of unpaired duplicated reads</li> <li>paired_duplicate_reads: number of paired reads that were also marked as duplicate</li> <li>unmapped_reads: number of unmapped reads</li> <li>percent_duplicate_reads: percentage of duplicate reads</li> <li>estimated_library_size: scaled total number of mapped reads</li> <li>total_reads: total number of reads, regardless of mapping status</li> <li>total_mapped_reads: total number of mapped reads</li> <li>total_duplicate_reads: number of duplicate reads</li> <li>total_properly_paired: number of properly paired reads</li> <li>coverage_breadth: percentage of genome covered by some read</li> <li>coverage_depth: average reads per nucleotide position in the genome</li> <li>median_insert_size: median insert size between paired reads</li> <li>mean_insert_size: mean insert size between paired reads</li> <li>standard_deviation_insert_size: standard deviation of the insert size between paired reads</li> <li>index_sequence: index sequence of the adaptor sequence</li> <li>column: column of the cell on the nanowell chip</li> <li>img_col: column of the cell from the perspective of the microscope</li> <li>index_i5: id of the i5 index adapter sequence</li> <li>sample_type: type of the sample</li> <li>primer_i7: id of the i5 index primer sequence</li> <li>experimental_condition: experimental treatment of the cell, includes controls</li> <li>index_i7: id of the i7 index adapter sequence</li> <li>cell_call: living/dead classification of the cell based on staining usually, C1 == living, C2 == dead</li> <li>sample_id: name of the sample</li> <li>primer_i5: id of the i5 index primer sequence</li> <li>row: row of the cell on the nanowell chip</li> <li>library_id: identifier for the DNA library</li> <li>index: ignored</li> <li>multiplier: during parameter searching, the set [1..6] that was chosen</li> <li>MSRSI_non_integerness: median of segment residuals from segment integer copy number states</li> <li>MBRSI_dispersion_non_integerness: median of bin residuals from segment integer copy number states</li> <li>MBRSM_dispersion: median of bin residuals from segment median copy number values</li> <li>autocorrelation_hmmcopy: hmmcopy copy autocorrelation</li> <li>cv_hmmcopy: ignored</li> <li>empty_bins_hmmcopy: number of empty bins in hmmcopy</li> <li>mad_hmmcopy: median absolute deviation of hmmcopy copy</li> <li>mean_hmmcopy_reads_per_bin: mean reads per hmmcopy bin</li> <li>median_hmmcopy_reads_per_bin: median reads per hmmcopy bin</li> <li>std_hmmcopy_reads_per_bin: standard deviation value of reads in hmmcopy bins</li> <li>total_halfiness: summed halfiness penality score of the cell</li> <li>total_mapped_reads_hmmcopy: total mapped reads in all hmmcopy bins</li> <li>scaled_halfiness: summed scaled halfiness penalty score of the cell</li> <li>mean_state_mads: mean value for all median absolute deviation scores for each state</li> <li>mean_state_vars: variance value for all median absolute deviation scores for each state</li> <li>mad_neutral_state: median absolute deviation score of the neutral 2 copy state</li> <li>breakpoints: number of breakpoints, as indicated by state changes not at the ends of chromosomes</li> <li>mean_copy: mean hmmcopy copy value</li> <li>state_mode: the most commonly occuring state</li> <li>log_likelihood: hmmcopy log likelihood for the cell</li> <li>true_multiplier: the exact decimal value used to scale the copy number for segmentation</li> <li>order: order of the cell in the hierarchical clustering tree</li> <li>quality: random forest classifier proability score that cell is good</li> </ul> <p>ov2295_clone_alleles.csv.gz: Table of clone specific allele data</p> <ul> <li>chr: chromosome of bin</li> <li>start: start of bin</li> <li>end: end of bin</li> <li>hap_label: haplotype block identifier</li> <li>clone_id: clone identifier</li> <li>allele_1_sum: number of reads for allele 1 of the haplotype block</li> <li>allele_2_sum: number of reads for allele 2 of the haplotype block</li> <li>total_counts_sum: total reads for the haplotype block</li> </ul> <p>ov2295_clone_breakpoints.csv.gz:&nbsp;Table of breakpoints per clone for OV2295 samples. Columns:</p> <ul> <li>prediction_id: identifier for the breakpoint</li> <li>chromosome_1: chromosome of breakend 1</li> <li>strand_1: orientation of break end 1</li> <li>position_1: position of break end 1</li> <li>chromosome_2: chromosome of breakend 2</li> <li>strand_2: orientation of break end 2</li> <li>position_2: position of break end 2</li> <li>clone_id: clone identifier</li> <li>read_count: number of reads</li> <li>is_present: presence=1, absent=0</li> </ul> <p>ov2295_clone_clusters.csv.gz: Table of cell clusters as putative clones</p> <ul> <li>cell_id: identifier for the cell</li> <li>clone_id: clone identifier</li> </ul> <p>ov2295_clone_cn.csv.gz:&nbsp;Table of allele specific copy number per clone for OV2295 samples. Columns:</p> <ul> <li>chr: chromosome of bin</li> <li>start: start of bin</li> <li>end: end of bin</li> <li>total_cn: HMMCopy predicted total copy number&nbsp;</li> <li>minor_cn: HMM predicted minor copy number&nbsp;</li> <li>major_cn: HMM predicted major copy number&nbsp;</li> <li>clone_id: clone identifier</li> </ul> <p>ov2295_clone_snvs.csv.gz:&nbsp;Table of SNVs per clone for OV2295 samples.&nbsp; Columns:</p> <ul> <li>chrom: chromosome</li> <li>coord: genome position</li> <li>ref: reference nucleotide</li> <li>alt: alternate nucleotide</li> <li>clone_id: clone identifier</li> <li>ref_counts: number of reads at this position matching the reference nucleotide</li> <li>alt_counts: number of reads at this position matching the alternate nucleotide</li> <li>total_counts: total number of reads at this position</li> <li>is_present: presence=0, absent=1</li> <li>is_het:&nbsp;is heterozygous</li> <li>is_hom: is homozygous for the alternate</li> </ul> <p>ov2295_nodes.csv.gz: Table of phylogenetic information for SNV evolution</p> <ul> <li>variant_id: identifier for the SNV as chrom:coord:ref:alt</li> <li>node: node in the phylogenetic tree</li> <li>loss: probability the SNV was lost at this node</li> <li>origin: probability the SNV originated at this node</li> <li>presence: probability the SNV is present at this node</li> <li>ml_origin: binary indicator the SNV originated at this node</li> <li>ml_presence: binary indicator the SNV is present at this node</li> <li>ml_loss: binary indicator the SNV was lost at this node</li> </ul> <p>ov2295_snv_counts.csv.gz: Table of SNV counts</p> <ul> <li>chrom: chromosome</li> <li>coord: genome position</li> <li>ref: reference nucleotide</li> <li>alt: alternate nucleotide</li> <li>ref_counts: number of reads at this position matching the reference nucleotide</li> <li>alt_counts: number of reads at this position matching the alternate nucleotide</li> <li>cell_id: identifier for the cell</li> <li>total_counts: total number of reads at this position</li> <li>sample_id: identifier for the sequenced sample</li> </ul> <p>ov2295_tree.pickle: Phylogenetic tree in python pickle format.&nbsp; Requires installation of the stochastic dollo code at: https://bitbucket.org/dranew/dollo, version 0.4.2.</p> <p>Note the following sample mapping: &lsquo;SA922&rsquo;: &lsquo;OV2295(R2)&rsquo;, &lsquo;SA921&rsquo;: &lsquo;TOV2295(R)&rsquo;, &lsquo;SA1090&rsquo;: &lsquo;OV2295&rsquo;,</p> <p><strong>Plots</strong></p> <p>ov_supp_clone_allele_cn.png: Clone allele ratios for each OV2295 sample.</p> <p>ov_supp_clone_total_cn.png: Clone copy number for each OV2295 sample.</p> <p>ov_supp_sample_total_cn.png: Bulk copy number for each OV2295 sample.</p> <p>ov_supp_sample_allele_cn.png: Bulk allele ratios for each OV2295 sample.</p>

opencc-by-4.0Sep 2019View details →
zenodo48/100

bollito: a flexible pipeline for comprehensive single-cell RNA-seq analyses - Melanoma tutorial

<p>Downsampled version of the melanoma dataset originally published by&nbsp;<em><a href="https://genome.cshlp.org/content/28/9/1353">Ho et al </a>(1)</em>. The&nbsp;dataset is composed by cells from the 451Lu cell line. There&nbsp;are two samples available:</p> <table> <tbody> <tr> <td><strong>Name</strong></td> <td><strong>Description</strong></td> <td><strong>R1/R2</strong></td> </tr> <tr> <td>451LU</td> <td>Parental cell line</td> <td>2500K_451LU_L003_R*_001.fastq.gz</td> </tr> <tr> <td>451LUBR3</td> <td>Vemurafenib-resistant sample treated with targeted BRAF inhibitors</td> <td>500K_451LUBR3_L004_R*_001.fastq.gz</td> </tr> </tbody> </table> <p><br> (1)&nbsp;Ho YJ, Anaparthy N, Molik D, et al. Single-cell RNA-seq analysis identifies markers of resistance to targeted BRAF inhibitors in melanoma cell populations.&nbsp;<em>Genome Res</em>. 2018;28(9):1353-1363. doi:10.1101/gr.234062.117</p>

opencc-by-4.0Feb 2022View details →
zenodo48/100

Cardelino: Integrating whole exomes and single-cell transcriptomes to reveal phenotypic impact of somatic variants

<p>This dataset&nbsp;consists of the reference data files, metadata and processed results files for the paper &quot;Cardelino: Integrating whole exomes and single-cell transcriptomes to reveal phenotypic impact of somatic variants,&quot; which&nbsp;investigates clonality in normal human dermal fibroblast cell populations in 32 cell lines from distinct donors, using bulk whole-exome sequencing and single-cell RNA-sequencing data.</p> <p>This dataset contains everything required to reproduce the results presented in the paper from&nbsp;processed data and results of our data processing workflows. Our analyses can be reproduced using the <a href="https://github.com/davismcc/fibroblast-clonality">source code</a>&nbsp;and instructions available at our <a href="https://davismcc.github.io/fibroblast-clonality/">project website</a>.</p> <p>The <em>entire</em> analysis workflow from raw data to final results is also reproducible but&nbsp;is substantially more complicated and computationally intensive.&nbsp;It also requires large datasets to be obtained from other repositories. Specifically, single-cell RNA-seq data have been deposited in the ArrayExpress database at EMBL-EBI under accession number E-MTAB-7167. Whole-exome sequencing data is available through the HipSci portal (www.hipsci.org). Combined with the dataset in this repository and following the instructions on the project website, it is possible to run our entire analysis pipeline.</p> <p>&nbsp;</p>

opencc-by-4.0Aug 2018View details →
zenodo48/100

Single-molecule DNA methylation patterns of full-length human-specific LINE-1 (L1HS) retrotransposons in a panel of cell lines.

<p>We used bs-ATLAS-seq to comprehensively map the genomic location and assess the DNA methylation status of&nbsp;full-length human-specific LINE-1 elements (L1HS). The approach capture region 1-210 of L1HS elements, which corresponds to the most 5&#39; end of its promoter sequence. This was performed in a panel of 12 human primary or transformed cell lines (BJ, IMR90, MRC5, H1, K562, HCT116, HeLa S3, HepG2, MCF7, HEK-293, HEK-293T, 2102Ep), many being shared with the encode project.</p> <p>These datasets provide a visualization for DNA methylation patterns at the single molecule level for each L1HS loci.</p>

opencc-by-4.0Sep 2022View details →
zenodo48/100

Cell metadata for "The emergent landscape of the mouse gut endoderm at single-cell resolution"

<p>Cell metadata for the data published in&nbsp;&quot;The emergent landscape of the mouse gut endoderm at single-cell resolution&quot;</p> <p>&nbsp;</p>

opencc-by-4.0May 2019View details →
zenodo44/100

Raw read counts and phased SNP counts for every single cell in the sequencing datasets of the breast cancer patient S1 from "Characterizing allele- and haplotype-specific copy numbers in single cells with CHISEL"

<p>This dataset contains the raw read counts and phased SNP counts&nbsp;for every single cell in the sequencing datasets of breast cancer patient S1 from &ldquo;Characterizing allele- and haplotype-specific copy numbers in single cells with CHISEL&rdquo; [Zaccaria &amp; Raphael, 2020]. These data enable the full reproduction of all the results in the related manuscript for breast cancer patient S1. Specifically, the data are provided in two files for every dataset <em>DAT</em> of patient S1 with the following format:</p> <ol> <li><em>DAT.raw</em>_<em>read</em>_<em>counts.bed.gz&nbsp;</em>is a multi-cell BED file containing the raw read counts in the following fields: <ul> <li>CHROMOSOME: the name of a human chromosome</li> <li>START: the starting genomic position of a genomic bin in the chromosome</li> <li>END: the ending genomic position of the genomic bin in the chromosome</li> <li>CELL: the cell barcode that uniquely identifies a cell</li> <li>NORMAL: the raw read count&nbsp;for the specified bin from a matched-normal sample</li> <li>COUNT: the raw read count&nbsp;for the specified bin in the specified cell</li> <li>RDR: the estimated read-depth ratio for the specified bin in the specified cell</li> </ul> </li> <li><em>DAT.phased</em>_<em>snps</em>_<em>counts.pos.gz&nbsp;</em>is a multi-cell POS file containing the phased SNP counts in the following fields: <ul> <li>CHROMOSOME: the name of a human chromosome</li> <li>POS: the genomic position in the chromosome of a germline SNP</li> <li>CELL: the cell barcode that uniquely identifies a cell</li> <li>COUNT_HAPLOTYPE_A: the count of reads that cover&nbsp;the SNP and that belong to haplotype A in the specified cell</li> <li>COUNT_HAPLOTYPE_B: the count of reads that cover&nbsp;the SNP and that belong to haplotype B in the specified cell</li> </ul> </li> </ol> <p>All the files have been compressed using standard <em>gzip</em>.</p>

opencc-by-4.0May 2020View details →
zenodo44/100

Raw read counts and phased SNP counts for every single cell in the sequencing datasets of the breast cancer patient S0 from "Characterizing allele- and haplotype-specific copy numbers in single cells with CHISEL"

<p>This dataset contains the raw read counts and phased SNP counts&nbsp;for every single cell in the sequencing datasets of breast cancer patient S0 from &ldquo;Characterizing allele- and haplotype-specific copy numbers in single cells with CHISEL&rdquo; [Zaccaria &amp; Raphael, 2020]. These data enable the full reproduction of all the results in the related manuscript for breast cancer patient S0. Specifically, the data are provided in two files for every dataset <em>DAT</em> of patient S0 with the following format:</p> <ol> <li><em>DAT.raw</em>_<em>read</em>_<em>counts.bed.gz&nbsp;</em>is a multi-cell BED file containing the raw read counts in the following fields: <ul> <li>CHROMOSOME: the name of a human chromosome</li> <li>START: the starting genomic position of a genomic bin in the chromosome</li> <li>END: the ending genomic position of the genomic bin in the chromosome</li> <li>CELL: the cell barcode that uniquely identifies a cell</li> <li>NORMAL: the raw read count&nbsp;for the specified bin from a matched-normal sample</li> <li>COUNT: the raw read count&nbsp;for the specified bin in the specified cell</li> <li>RDR: the estimated read-depth ratio for the specified bin in the specified cell</li> </ul> </li> <li><em>DAT.phased</em>_<em>snps</em>_<em>counts.pos.gz&nbsp;</em>is a multi-cell POS file containing the phased SNP counts in the following fields: <ul> <li>CHROMOSOME: the name of a human chromosome</li> <li>POS: the genomic position in the chromosome of a germline SNP</li> <li>CELL: the cell barcode that uniquely identifies a cell</li> <li>COUNT_HAPLOTYPE_A: the count of reads that cover&nbsp;the SNP and that belong to haplotype A in the specified cell</li> <li>COUNT_HAPLOTYPE_B: the count of reads that cover&nbsp;the SNP and that belong to haplotype B in the specified cell</li> </ul> </li> </ol> <p>All the files have been compressed using standard <em>gzip</em>.</p>

opencc-by-4.0May 2020View details →
zenodo44/100

Data archive: CICT for single cell RNA-seq network inference

<p>This archive contains benchmarking input data and results for using single cell gene expression data to infer gene regulatory networks (GRN) by the Causal Inference with Composition of Transactions (CICT) method and a selected set of published methods. This accompanies the manuscript "Robust discovery of gene regulatory networks from single-cell gene expression data by Causal Inference Using Composition of Transactions" (Shojaee and Huang, Brief in Bioinform 2023. DOI: 10.1093/bib/bbad370). The CICT code is available at the GitHub repo (https://github.com/hlab1/scRNAseqWithCICT/).</p><p>The original CICT algorithm was described in Shojaee et al. (arXiv:1608.02658, 2016). The benchmarked methods were included in the BEELINE benchmarking pipeline (Pratapa et al., Nat Methods 2020), to which we added DEEPDRIM (Chen et al., Brief Bioinform 2021), SCENIC (Aibar et al., Nat Methods 2017), Inferelator 3.0 (Gibbs et al., Bioinformatics 2022), and CellOracle (Kamimoto et al., Nature 2023). The output directory names are (subdirectories within each dataset):</p><p>* CICT_ewMIshrink_RFmaxdepth10_RFntrees20/: CICT for simulated data<br>* CICT_v2/: CICT for experimental data<br>* CELLORACLEDB/: CellOracle for experimental data<br>* DEEPDRIM72_ewMIshrink_RFmaxdepth10_RFntrees20/: DEEPDRIM for simulated data<br>* DEEPDRIM72_v2/: DEEPDRIM for experimental data<br>* INFERELATOR38_ewMIshrink_RFmaxdepth10_RFntrees20/: Inferelator-Prior for simulated data<br>* INFERELATOR38_v2/: Inferelator-Prior for experimental data<br>* INFERELATOR34_ewMIshrink_RFmaxdepth10_RFntrees20/: Inferelator-NoPrior for experimental data<br>* INFERELATOR34_v2/: Inferelator-NoPrior for experimental data<br>* GENIE3/: GENIE3<br>* GRNBOOST2/: GRNBOST2<br>* LEAP/: LEAP<br>* PIDC/: PIDC<br>* PPCOR/: PPCOR<br>* SCENICDB/: SCENIC for experimental data<br>* SCNS/: SCNS<br>* SCODE/: SCODE<br>* SCRIBE/: SCRIBE<br>* SINCERITIES/: SINCERITIES<br>* SINGE/: SINGE<br>* RANDOM/: RANDOM</p><p>The methods were benchmarked against two kinds of scRNA-seq datasets:<br>* Simulated datasets produced by the SERGIO simulator from a synthetic network (Dibaeinia et al., Cell Systems 2020), including complete datasets and datasets with dropouts with shape parameter k=6.5 and rate parameter q=10, 30, 50, 70, 80.&nbsp;<br>* Experimental datasets compiled by the BEELINE pipeline, evaluated at three different levels L0, L1 and L2, with three types of ground truth networks.<br>&nbsp; &nbsp; * Evaluation levels:<br>&nbsp;&nbsp; &nbsp; &nbsp; &nbsp;* L0: 500 highly varying genes plus TFs<br>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;* L1: 1000 highly varying genes plus TFs<br>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;* L2: 500 highly varying genes, TFs and 500 genes randomly selected that excluded the 1000 highly varying genes from L1.<br>&nbsp; &nbsp; * Types of ground truths:<br>&nbsp;&nbsp; &nbsp; &nbsp; &nbsp;* Cell-type-specific ChIP-seq ground truth (L0, L1, L2)<br>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;* Non-specific ChIP-seq ground truth (L0_ns, L1_ns, L2_ns)<br>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;* Loss-of-function/gain-of-function ground truth (L0_lofgof, L1_lofgof, L2_lofgof)</p><p>The directory structure is organized in accordance with the BEELINE benchmarking pipeline. For complete details please please see the BEELINE documentation (https://murali-group.github.io/Beeline/) and Github repo (https://github.com/Murali-group/Beeline).</p><p>&nbsp;</p>

opencc-by-nc-sa-4.0Jun 2023View details →
zenodo44/100

Advanced Non-Clear Cell Renal Cell Carcinoma Treatments and Survival: A Real-World Single-Centre Experience

<p>Dataset of the paper "Advanced Non-Clear Cell Renal Cell Carcinoma Treatments and Survival: A Real-World Single-Centre Experience"</p>

opencc-by-4.0Nov 2023View details →
zenodo44/100

Development of Multiomics in situ Pairwise Sequencing (MiP-Seq) for Single-cell Resolution Multidimensional Spatial Omics

<p>The original data used in the article:&nbsp;Development of Multiomics in situ Pairwise Sequencing (MiP-Seq) for Single-cell Resolution Multidimensional Spatial Omics</p> <p>Delineating the spatial multiomics landscape will pave the way to understanding the molecular basis of physiology and pathology. However, current spatial omics technology development is still in its infancy. Here, we developed a high-throughput targeted in situ sequencing strategy, multiomics in situ pairwise sequencing (MiP-Seq), to efficiently decipher multiplexed DNAs, RNAs, proteins, and small biomolecules at subcellular resolution. MiP-Seq simultaneously sequenced the dual barcode base of padlock probes, dramatically increasing the detection capacity to 10N by N rounds of sequencing. We delineated spatial gene profiles in the hypothalamus using MiP-Seq. Moreover, MiP-Seq was unitized to detect tumor gene mutations and allele-specific expression of parental genes and to differentiate sites with and without the m6A RNA modification at specific sites. MiP-Seq was combined with in vivo Ca2+ imaging and Raman imaging to obtain a spatial multiomics atlas correlated to neuronal activity and cellular biochemical fingerprints. Importantly, we proposed a &ldquo;signal dilution strategy&rdquo; to resolve the crowded signals that challenge the applicability of in situ sequencing. Together, our method improves spatial multiomics and precision diagnostics, and facilitates analyzing cell function in connection with gene profiles.</p>

opencc-by-4.0Jan 2023View details →
zenodo44/100

scGraph2Vec: a deep generative model for gene embedding augmented by Graph Neural Network and single-cell omics data

<p>This repository contains the training data and source code to reproduce the results of our paper:<br>scGraph2Vec: a deep generative model for gene embedding augmented by Graph Neural Network and single-cell omics data</p> <p>More description can be also found in GitHub (https://github.com/LPH-BIG/scGraph2Vec).</p>

opencc-zeroJun 2024View details →
zenodo44/100

Single-cell RNA-seq profiles of tumor-bearing mice treated with PAGln with or without anti-PD-1

<p>single-cell RNA sequencing (scRNA-seq) profiles of&nbsp; tumor-bearing mice treated using Phenylacetylglutamine (PAGln) with or without anti-PD-1 were performed to compare the alterations of immune microenvironment affected by PAGln under the condition of anti-PD-1 treatment.</p>

opencc-by-4.0Nov 2024View details →
zenodo44/100

Additional data: Longitudinal single-cell multiomic atlas of high-risk neuroblastoma reveals chemotherapy-induced tumor microenvironment rewiring

<p>This repository provides additional data for the manuscript titled "Longitudinal single-cell multiomic atlas of high-risk neuroblastoma reveals chemotherapy-induced tumor microenvironment rewiring", currently under revision at Nature Genetics. The primary data cohort has been deposited in the HTAN data portal. This repository includes processed 10x Xenium spatial transcriptomic data for six TH-MYCN mice (three chemotherapy-treated and three treatment-naive) as well as processed scRNA-seq data for CHLA15 and CHLA20 neuroblastoma (NBL) cells. The scRNA-seq data includes mono-cultured, co-cultured cells with THP-1 macrophages, and co-culture cells treated with Afatinib/CRM197.&nbsp;&nbsp;</p>

opencc-by-4.0Dec 2024View details →
zenodo44/100

Systematic reconstruction of molecular pathway signatures using scalable single-cell perturbation screens

<p>This repo contains Seurat objects, differential expression analysis results, and pathway gene lists for the manuscript "Systematic reconstruction of molecular pathway signatures using scalable single-cell perturbation screens"<br>List of files:</p> <p>1. Seurat_object_IFNB_Perturb_seq.rds: &nbsp; &nbsp; Seurat object of the Perturb-seq data for Interferon-beta pathway<br>2. Seurat_object_IFNG_Perturb_seq.rds: &nbsp; &nbsp;Seurat object of the Perturb-seq data for Interferon-gamma pathway<br>3. Seurat_object_TNFA_Perturb_seq.rds: &nbsp; Seurat object of the Perturb-seq data for TNF-alpha pathway<br>4. Seurat_object_TGFB1_Perturb_seq.rds: Seurat object of the Perturb-seq data for TGF-beta1 pathway<br>5. Seurat_object_INS_Perturb_seq.rds: &nbsp; &nbsp; &nbsp;Seurat object of the Perturb-seq data for insulin pathway<br>6. Pathway_genelist.rds: &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; The pathway gene lists from MultiCCA analysis<br>7. Pathway_Exclusive_genelist.rds: &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;The pathway exclusive gene lists generated from Pathway_genelist.rds<br>8. HClust_Pathway_celltype_specific_genelist.rds: &nbsp; &nbsp; The cell-line specific pathway gene lists from hierarchical clustering analysis independently done on each cell line<br>9. DE_results_all_pathway.zip: &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; The DE test results for all the regulators, cell lines, and pathways (from Mixscale weighted DE test.)<br>10. Bulk_RNAseq_Seurat_object_IFNG_and_TGFB_stim.rds: &nbsp; &nbsp; &nbsp; Seurat object for the bulk RNA-seq data for interferon-gamma and TGF-beta stimulation experiments<br>11. Parse_Guide_Capture_Protocol.pdf: &nbsp; &nbsp; &nbsp;The guide RNA capture protocol developed for Parse Evercode Whole Transcriptome kit</p>

opencc-by-4.0Jan 2024View details →
zenodo44/100

Single-cell analyses of axolotl forebrain organization, neurogenesis, and regeneration

<p>Preprint:&nbsp;https://doi.org/10.1101/2022.03.21.485045</p> <p>Abstract:</p> <p>Salamanders are important tetrapod models to study brain organization and regeneration, however the identity and evolutionary conservation of brain cell types is largely unknown. Here, we delineate cell populations in the axolotl telencephalon during homeostasis and regeneration, representing the first single-cell genomic and spatial profiling of an anamniote tetrapod brain. We identify glutamatergic neurons with similarities to amniote neurons of hippocampus, dorsal and lateral cortex, and conserved GABAergic neuron classes. We infer transcriptional dynamics and gene regulatory relationships of postembryonic, region-specific direct and indirect neurogenesis, and unravel conserved signatures. Following brain injury, ependymoglia activate an injury-specific state before reestablishing lost neuron populations and axonal connections. Together, our analyses yield key insights into the organization, evolution, and regeneration of a tetrapod nervous system.</p> <p>&nbsp;</p> <p>File description:</p> <p>all_nuclei_clustered_highlevel_anno.rds - Seurat object including all snRNA-seq data from uninjured pallium, both from microdissections and whole pallium multiome.</p> <p>pallium_metadata_simp.csv - csv file containing a simplified version of the metadata for the uninjured pallium</p> <p>Edu_1_2_4_6_8_12_fil_highvarfeat.rds - Seurat object containing all Div-seq data for the pallium injury time course</p> <p>divseq_predicted_metadata.csv - csv file containing a simplified version of the metadata for the pallium injury time course</p> <p>ep_wpi_srat.rds - Seurat object containing an integrated version of ependymoglia cells from uninjured and injured pallium (see Fig 6 in the preprint).</p> <p>D1_113_sub_b.rds - Seurat object containing a Visium data for the axolotl pallium</p> <p>multiome_integATAC_SCT.rds - Signac object containing the data used for multiome analysis of the uninjured whole pallium</p> <p>predictions_cell2loc.csv - csv file containing cell2location scores for the uninjured pallium cell types in the Visium dataset</p>

opencc-by-4.0Mar 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record