Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,133

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

1,133 results for “Single-cell analysis”

Learn how ShareScore rates datasets ↗
zenodo44/100

Single-cell analysis of megakaryopoiesis in peripheral CD34+ cells: insights into ETV6-related thrombocytopenia

<p>This repository contains necessary files for reproducing the analysis in Bigot et al, 2023. The instructions for reproducing the analysis are given in github (https://github.com/poggiteam/ETV6_2020).</p>

opencc-by-4.0Sep 2022View details →
zenodo44/100

Mammary single-cell RNA-seq analysis and prostate cancer survival as a function of H2AFJ expression for the paper entitled: The histone variant H2A.J is enriched in luminal epithelial cells

<p>H2A.J is a poorly studied mammalian-specific variant of histone H2A. We used immunohistochemistry to study its localization in various human and mouse tissues. H2A.J showed cell-type specific expression with a striking enrichment in luminal epithelial cells of multiple glands including those of breast, prostate, pancreas, thyroid, stomach, and salivary glands. H2A.J was also highly expressed in many carcinoma cell lines and in particular, those derived from luminal breast and prostate cancer. H2A.J thus appears to be a novel marker for luminal epithelial cancers. Knocking-out the H2AFJ gene in T47D luminal breast cancer cells reduced the expression of several estrogen-responsive genes which may explain its putative tumorigenic role in luminal-B breast cancer.</p>

opencc-by-4.0Sep 2021View details →
zenodo40/100

Robustness and applicability of transcription factor and pathway analysis tools on single-cell RNA-seq data

<p>Data used to test the robustness and applicability of transcription factor and pathway analysis tools on single-cell RNA-seq data, described in <a href="https://doi.org/10.1186/s13059-020-1949-z">Holland et al. 2020</a>.</p> <p>The folder&nbsp;<em>data </em>contains<em>&nbsp;</em>raw data and the folder <em>output</em> contains intermediate and final results of all analyses.&nbsp;</p> <p>The associated analyses code and more information are available on&nbsp;<a href="https://github.com/saezlab/FootprintMethods_on_scRNAseq">GitHub</a>.</p> <p>&nbsp;</p> <p><strong>Abstract</strong></p> <p><strong>Background</strong></p> <p>Many functional analysis tools have been developed to extract functional and mechanistic insight from bulk transcriptome data. With the advent of single-cell RNA sequencing (scRNA-seq), it is in principle possible to do such an analysis for single cells. However, scRNA-seq data has characteristics such as drop-out events and low library sizes. It is thus not clear if functional TF and pathway analysis tools established for bulk sequencing can be applied to scRNA-seq in a meaningful way.</p> <p><strong>Results</strong></p> <p>To address this question, we perform benchmark studies on simulated and real scRNA-seq data. We include the bulk-RNA tools PROGENy, GO enrichment, and DoRothEA that estimate pathway and transcription factor (TF) activities, respectively, and compare them against the tools SCENIC/AUCell and metaVIPER, designed for scRNA-seq. For the in silico study, we simulate single cells from TF/pathway perturbation bulk RNA-seq experiments. We complement the simulated data with real scRNA-seq data upon CRISPR-mediated knock-out. Our benchmarks on simulated and real data reveal comparable performance to the original bulk data. Additionally, we show that the TF and pathway activities preserve cell type-specific variability by analyzing a mixture sample sequenced with 13 scRNA-seq protocols. We also provide the benchmark data for further use by the community.</p> <p><strong>Conclusions</strong></p> <p>Our analyses suggest that bulk-based functional analysis tools that use manually curated footprint gene sets can be applied to scRNA-seq data, partially outperforming dedicated single-cell tools. Furthermore, we find that the performance of functional analysis tools is more sensitive to the gene sets than to the statistic used.</p> <p>&nbsp;</p> <p>For questions related to the data please write an email to christian.holland@bioquant.uni-heidelberg.de or use the <a href="https://github.com/saezlab/FootprintMethods_on_scRNAseq/issues">GitHub issue system</a>.</p>

opencc-by-4.0Dec 2019View details →
zenodo40/100

Data for "Tuning parameters of dimensionality reduction methods for single-cell RNA-seq analysis"

<p>The files named&nbsp;<code>df_scran.csv</code>,&nbsp;<code>df_seurat.csv</code>,&nbsp;<code>df_zinbwave.csv</code>,&nbsp;<code>df_dca.csv</code>, and&nbsp;<code>df_scvi.csv</code>&nbsp;contain one row per configuration that we ran successfully.</p> <p>The files named&nbsp;<code>DATASET.METHOD.h5ad</code>&nbsp;are encoded with anndata&nbsp;<code>v0.7.0</code>&nbsp;(be careful as they are not readable with previous versions) and contain 100 embeddings each. The embeddings are in the&nbsp;<code>obsm</code>&nbsp;attribute of the object. All the embeddings can be listed with the&nbsp;<code>obsm_keys()</code>&nbsp;method. The name of the embedding contains the parameters used to generate that embedding and are written like that&nbsp;<code>method=zinbwave.dims=10.epsilon=1000.features=300.gene_covariate=0</code>.</p> <p>&nbsp;</p> <p>For questions on this dataset please contact fraimundo@google.com</p>

openapache2.0Jul 2020View details →
zenodo40/100

EpiScanpy: integrated single-cell epigenomic analysis

<p>All files below are in &quot;h5ad&quot; format, which can be opened by the Python package AnnData (see https://anndata.readthedocs.io/en/latest). The files are organized as:</p> <p><br> 1. Single-nucleus methylcytosine sequencing (snmC-seq) of neurons from the frontal cortex of young adult mouse brains from Luo et al., 2017. The files are the preprocessed input for EpiScanpy for multiple CG imputed feature spaces, 100k base pair windows, enhancers, gene bodies, promoters and CH gene bodies.<br> - processed_enhancers_CG_luo_et_al_nov2020_paper_resubmission.h5ad<br> - processed_genebodies_CG_luo_et_al_nov2020_paper_resubmission.h5ad<br> - processed_genebodies_CH_luo_et_al_nov2020_paper_resubmission.h5ad<br> - processed_promoter_CG_luo_et_al_nov2020_paper_resubmission.h5ad<br> - processed_windows_CG_luo_et_al_nov2020_paper_resubmission.h5ad<br> &nbsp;</p> <p>2. Single cell ATAC sequencing (scATAC-seq) from 10X Genomics, preprocessed peak count matrix of Next GEM v1.1 10k Peripheral blood mononuclear cells (PBMCs) from a healthy donor. Peak matrices of Next GEM v1.1 10k PBMCs and whole blood fresh data (GEO:GSE129785) from Satpathy et al. 2019 (Greenleaf&#39;s lab). Concatenated 5000bp matrix of Fresh cortex from adult mouse brain (P50) from 10X Genomics and CEMBA180312_3B mouse brain sample from Fang et al. 2019.<br> - atac_pbmc_10k_nextgem_fragments_macs2_peaks_outter_all_chrom.h5ad<br> - atac_pbmc_10k_nextgem_fragments_merged_peaks_for_integration_greenleaf_outter_all_chrom.h5ad<br> - raw_greenleaf_pre_integration_oct_2020.h5ad (Satpathy et al. 2019)<br> - preprocessed_10x_genomics_5k_adulte_mouse_brain_Fang_et_al_2019_CEMBA180312_3B_5kb_windows.h5ad</p>

opencc-by-4.0Nov 2020View details →
zenodo40/100

Computational Analysis of Two-dimensional High-throughput Data from Large-scale RNAi Screens and Single-cell Transcriptomics

<p>This publication&nbsp;provides&nbsp;a singularity definition file to reproduce the computational environment along with the scripts to reproduce every figure or table in the revised manuscript using ZetaSuite Perl module and R package.</p> <p>First, generate a new folder and then download all the files into the folder.</p> <p>Then, uncompressed the files DataSets_part1.tar.gz,DataSets_part2.tar.gz,DataSets_part3.tar.gz,DataSets_part4.tar.gz, and scripts.tar.gz. within the folder.</p> <p>Next, move all the files in DataSets_part1 folder,&nbsp;DataSets_part2&nbsp;folder,DataSets_part3&nbsp;folder and&nbsp;DataSets_part4&nbsp;folder to a new folder called DataSets.</p> <p>Finally, run the following scripts to generate the&nbsp;figures and tables in our manuscript.</p> <p>Regeneration of Figure2 and S2: singularity exec ZetaSuite.sif sh Figure2andS2.sh&nbsp;&nbsp;</p> <p>Regeneration of Figure3 and S3: singularity exec ZetaSuite.sif sh Figure3andS3.sh&nbsp;&nbsp;</p> <p>Regeneration of Figure4 and S4: singularity exec ZetaSuite.sif sh Figure4andS4.sh&nbsp;&nbsp;</p> <p>Regeneration of Figure5 and S5: singularity exec ZetaSuite.sif sh Figure5andS5.sh&nbsp;&nbsp;</p> <p>Regeneration of Figure6 and S6: singularity exec ZetaSuite.sif sh Figure6andS6.sh&nbsp;&nbsp;</p> <p>Regeneration of Figure7 and S7: singularity exec ZetaSuite.sif sh Figure7andS7.sh&nbsp;&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Mar 2022View details →
zenodo40/100

Image data for bioRxiv article named: mtFociCounter - Reproducible, open source and quantitative single-cell analysis of mitochondrial nucleoids and other foci

<p>Raw imaging data to reproduce and test the findings of the bioRxiv article: <strong>mtFociCounter </strong>- Reproducible, open source and quantitative single-cell analysis of mitochondrial nucleoids and other foci. It contains data from three imaging days and 2 or three technical replicates on each day.</p> <p>&nbsp;</p>

opencc-by-4.0Aug 2022View details →
zenodo40/100

Optimized summary-statistic-based single-cell meta-analysis. Input files

<p>This dataset contains information about the input files used in the Optimized summary-statistic-based single-cell meta-analysis research project.&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0May 2024View details →
zenodo40/100

Data for: "A high-throughput microscopy method for single-cell analysis of event-time correlations in nanoparticle-induced cell death"

<p>Data related to the&nbsp;publication Murschhauser <em>et al.</em>: <a href="https://doi.org/10.1038/s42003-019-0282-0">A high-throughput microscopy method for single-cell analysis of event-time correlations in nanoparticle-induced cell death</a>. It contains fluorescence time traces of single cells marked with cell-event markers and observed by time-lapse microscopy. The cells were treated with nanoparticles at different doses (NP25 and NP100), with staurosporine (sts) or were left untreated for control (ctrl). See the above-mentioned publication for more details.</p> <p>The format of the data is described below.</p> <p>The file <code>Data_A549.zip</code> contains data measured with A549 cells, and the file <code>Data_Huh7.zip</code> contains data measured with Huh7 cells. Both files have the same structure. Each file contains the directories <code>Raw</code> and <code>Fitted</code> as well as a checksum file. The <code>Raw</code> directory contains single-cell fluorescence time courses as obtained by time-lapse microscopy. The <code>Fitted</code> directory contains the results of fitting model functions as well as properties of identified events, such as event times. The checksum file contains SHA256 checksums of all files within these directories and can be used to check file integrity.</p> <p>Both directories contain measurement directories. Each measurement directory contains the data corresponding to&nbsp;one experiment. The name of the measurement directory is the measurement identifier. Each measurement directory contains condition directories. Each condition directory contains data corresponding to one condition measured in the measurement and is named after the condition. Each condition directory contains marker directories. They are named after the fluorescence markers measured and contain&nbsp;files with single-cell data corresponding to the respective markers.</p> <p>The names of those files consist of multiple parts separated by underscores. The first two parts identify a position of the microscope. Since pairs of markers were measured, each position is present in two marker directories. The third part is the measurement identifier. The other parts will be described below.</p> <p>The <code>Raw</code> directory contains only CSV files with the raw fluorescence time courses. The filenames contain no other parts and have the suffix &ldquo;.txt&rdquo;. The first row of each CSV file is the time (in units of 10 minutes), and the other rows are the fluorescence time courses of the cells observed at the corresponding position (in arbitrary units). Each file in the <code>Raw</code> directory corresponds to a group of files in the <code>Fitted</code> directory.</p> <p>The <code>Fitted</code> directory contains three types of CSV files. Their names have &ldquo;ALL&rdquo; as fourth part,&nbsp;a session identifier as sixth part and the suffix &ldquo;.csv&rdquo;. The fifth part indicates the type of file and is one of the following:</p> <ul> <li>&ldquo;PARAMS&rdquo; indicates the estimated values for the model parameters. Each row stands for one cell and each column for a parameter of the model function fitted to the data. The model functions are published with the&nbsp;<a href="https://doi.org/10.5281/zenodo.1418465">fitting software</a>.</li> <li>&ldquo;SIMULATED&rdquo; indicates&nbsp;the fitted traces. The traces are calculated using the model functions and the estimated parameters. The format is the same as for the raw traces, but the time is in units of hours and has a higher resolution.</li> <li>&ldquo;STATE&rdquo; indicates additional information extracted from the fitted traces. Each row stands for a cell and each column for a property. The first column is the number of the cell. The second column is the event time&nbsp;found (in hours); non-finite values indicate that no event time was found. The third and fourth columns contain the absolute and relative amplitude of the trace, respectively. The fifth column is the logarithmic likelihood of the best fit. The sixth column indicates an algorithm used for postprocessing, and the seventh column indicates the trace slope at the event. See the fitting software for details.</li> </ul> <p>&nbsp;</p>

opencc-by-4.0Sep 2018View details →
zenodo40/100

Microscopy data for the paper: Analysis and design of single-cell experiments to harvest fluctuation information while rejecting measurement noise.

<p>Microscopy data for the paper: Analysis and design of single-cell experiments to harvest fluctuation information while rejecting measurement noise.</p> <p>&nbsp;</p> <p>List of files used for each dataset.</p> <p>&nbsp;</p> <p>Dataset 0 : MS2-CY5_Cyto543_560_woStim</p> <p>&nbsp;&nbsp;&nbsp;&nbsp;Images in the dataset :</p> <p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;ROI001_XY1657814108_Z00_T0_merged.tif &nbsp;&nbsp;- Image Id Number: &nbsp;0</p> <p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;ROI002_XY1657815441_Z00_T0_merged.tif &nbsp;&nbsp;- Image Id Number: &nbsp;1</p> <p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;ROI003_XY1657814110_Z00_T0_merged.tif &nbsp;&nbsp;- Image Id Number: &nbsp;2</p> <p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;ROI004_XY1657814111_Z00_T0_merged.tif &nbsp;&nbsp;- Image Id Number: &nbsp;3</p> <p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;ROI005_XY1657814112_Z00_T0_merged.tif &nbsp;&nbsp;- Image Id Number: &nbsp;4</p> <p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;ROI006_XY1657814113_Z00_T0_merged.tif &nbsp;&nbsp;- Image Id Number: &nbsp;5</p> <p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;ROI007_XY1657814114_Z00_T0_merged.tif &nbsp;&nbsp;- Image Id Number: &nbsp;6</p> <p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;ROI008_XY1657814115_Z00_T0_merged.tif &nbsp;&nbsp;- Image Id Number: &nbsp;7</p> <p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;ROI009_XY1657814116_Z00_T0_merged.tif &nbsp;&nbsp;- Image Id Number: &nbsp;8</p> <p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;ROI010_XY1657814117_Z00_T0_merged.tif &nbsp;&nbsp;- Image Id Number: &nbsp;9</p> <p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;ROI011_XY1657814118_Z00_T0_merged.tif &nbsp;&nbsp;- Image Id Number: &nbsp;10</p> <p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;ROI012_XY1657814119_Z00_T0_merged.tif &nbsp;&nbsp;- Image Id Number: &nbsp;11</p> <p>&nbsp;</p> <p>Datset 1 : MS2-CY5_Cyto543_560_18minTPL_5uM</p> <p>&nbsp;&nbsp;&nbsp;&nbsp;Images in the dataset :</p> <p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;ROI001 - Position 1_XY1657818948_Z00_T0_merged.tif &nbsp;&nbsp;- Image Id Number: &nbsp;0</p> <p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;ROI001 - Position 2_XY1657818949_Z00_T0_merged.tif &nbsp;&nbsp;- Image Id Number: &nbsp;1</p> <p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;ROI001 - Position 4_XY1657818951_Z00_T0_merged.tif &nbsp;&nbsp;- Image Id Number: &nbsp;2</p> <p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;ROI001 - Position 5_XY1657818952_Z00_T0_merged.tif &nbsp;&nbsp;- Image Id Number: &nbsp;3</p> <p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;ROI001 - Position 6_XY1657818953_Z00_T0_merged.tif &nbsp;&nbsp;- Image Id Number: &nbsp;4</p> <p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;ROI001 - Position 7_XY1657818954_Z00_T0_merged.tif &nbsp;&nbsp;- Image Id Number: &nbsp;5</p> <p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;ROI001 - Position 8_XY1657818955_Z00_T0_merged.tif &nbsp;&nbsp;- Image Id Number: &nbsp;6</p> <p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;ROI001 - Position 9_XY1657818956_Z00_T0_merged.tif &nbsp;&nbsp;- Image Id Number: &nbsp;7</p> <p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;ROI001 - Position 10_XY1657818957_Z00_T0_merged.tif &nbsp;&nbsp;- Image Id Number: &nbsp;8</p> <p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;ROI001 - Position 11_XY1657818958_Z00_T0_merged.tif &nbsp;&nbsp;- Image Id Number: &nbsp;9</p> <p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;ROI001 - Position 12_XY1657818959_Z00_T0_merged.tif &nbsp;&nbsp;- Image Id Number: &nbsp;10</p> <p>&nbsp;</p> <p>Dataset 2: MS2-CY5_Cyto543_560_5hTPL_5uM</p> <p>&nbsp;&nbsp;&nbsp;&nbsp;Images in the datset :</p> <p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;ROI001_XY1657822809_Z00_T0_merged.tif &nbsp;&nbsp;- Image Id Number: &nbsp;0</p> <p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;ROI002_XY1657822933_Z00_T0_merged.tif &nbsp;&nbsp;- Image Id Number: &nbsp;1</p> <p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;ROI003_XY1657822934_Z00_T0_merged.tif &nbsp;&nbsp;- Image Id Number: &nbsp;2</p> <p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;ROI005_XY1657822936_Z00_T0_merged.tif &nbsp;&nbsp;- Image Id Number: &nbsp;3</p> <p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;ROI006_XY1657822937_Z00_T0_merged.tif &nbsp;&nbsp;- Image Id Number: &nbsp;4</p> <p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;ROI007_XY1657822938_Z00_T0_merged.tif &nbsp;&nbsp;- Image Id Number: &nbsp;5</p> <p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;ROI008_XY1657822939_Z00_T0_merged.tif &nbsp;&nbsp;- Image Id Number: &nbsp;6</p> <p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;ROI010_XY1657822941_Z00_T0_merged.tif &nbsp;&nbsp;- Image Id Number: &nbsp;7</p> <p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;ROI013_XY1657822944_Z00_T0_merged.tif &nbsp;&nbsp;- Image Id Number: &nbsp;8</p> <p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;ROI014_XY1657822945_Z00_T0_merged.tif &nbsp;&nbsp;- Image Id Number: &nbsp;9</p> <p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;ROI015_XY1657822946_Z00_T0_merged.tif &nbsp;&nbsp;- Image Id Number: &nbsp;10</p> <p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;ROI016_XY1657822947_Z00_T0_merged.tif &nbsp;&nbsp;- Image Id Number: &nbsp;11</p> <p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;ROI017_XY1657822948_Z00_T0_merged.tif &nbsp;&nbsp;- Image Id Number: &nbsp;12</p> <p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;ROI018_XY1657822949_Z00_T0_merged.tif &nbsp;&nbsp;- Image Id Number: &nbsp;13</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Apr 2023View details →
zenodo40/100

Comparative Analysis of Droplet- vs. Microwell-based Whole Transcriptome Single-Cell Sequencing Technologies in Complex Human Tissues

<p>In the past decade, high-dimensional single-cell omics tools have enabled scientists to study the tumor microenvironment (TME) in unprecedented detail. However, recent investigations suggest that each technique has its unique strengths but also technology-inherent limitations. Here we directly compared two commercially available high-throughput single-cell RNA sequencing (scRNA-seq) technologies - droplet-based 10X&nbsp;Chromium <em>vs.</em> microwell-based BD&nbsp;Rhapsody - using paired samples from patients with localized prostate cancer (PCa) undergoing a radical prostatectomy.</p> <p>Although high technical consistency was observed in unraveling the whole transcriptome, the relative abundance of detectable cell populations differed. This could in part be ascribed to differences in the performance to recover cells with low-mRNA content. Hence, immune cells such as neutrophils are underrepresented in data generated with the widely used droplet-based scRNA-seq protocol, highlighting the importance of considering platform limitations in low mRNA content cell recovery. In contrast, droplet-based scRNA-seq demonstrated superiority in terms of recovering cells of epithelial origin. Moreover, we discovered platform-dependent variabilities in mRNA quantification and cell-type marker annotation, affecting the composition of identified tissue profiles and the exploratory value of the generated datasets. Overall, our study emphasizes the importance of carefully selecting the appropriate scRNA-seq platform to improve cell type representation and obtain a more comprehensive and accurate understanding of the TME.</p>

opencc-by-4.0Jun 2023View details →
zenodo40/100

Model-based analysis of sample index hopping reveals its widespread artifacts in multiplexed single-cell RNA-sequencing

<p>Supplementary data&nbsp;that are needed to rerun&nbsp;the reproducible notebooks from the first steps using Alevin output and configuration files.</p> <p>Intermediate R data object that can be used to rerun the reproducible notebooks after the filtering steps.</p> <p>Validation data for inferring the sample index hopping rate. The <em>hiseq4000_joined_datatable_plexed_nonplexed.zip file contains read counts for four samples (two non-multiplexed and two multiplexed)&nbsp; joined by&nbsp; a cell-barcode, UMI, and gene-ID (CUG) key combination. The hiseq4000_inner_joined_with_labels.zip file contains only those CUGs that are observed in both the non-multiplexed and multiplexed samples.</em><em> </em></p>

opencc-by-4.0Jul 2019View details →
dryad40/100

Vizgen MERFISH files for Single-cell analysis reveals M. tuberculosis ESX-1-mediated accumulation of anti-inflammatory macrophages in infected mouse lungs

Open the record for dataset details and reuse information.

publicDec 2024View details →
dryad40/100

Single-cell analysis reveals M. tuberculosis ESX-1-mediated accumulation of anti-inflammatory macrophages in infected mouse lungs

Open the record for dataset details and reuse information.

publicDec 2024View details →
dryad36/100

Affected cell types for hundreds of Mendelian diseases revealed by analysis of human and mouse single-cell data

<p>Hereditary diseases manifest clinically in certain tissues, however their affected cell types typically remain elusive. Single-cell expression studies showed that overexpression of disease-associated genes may point to the affected cell types. Here, we developed a method that infers disease-affected cell types from the preferential expression of disease-associated genes in cell types (PrEDiCT). We applied PrEDiCT to single-cell expression data of six human tissues, to infer the cell types affected in 1,459 hereditary diseases. Overall, we identified 114 cell types affected by 1,140 diseases. We corroborated our findings by literature text-mining and recapitulation in mouse corresponding tissues. Based on these findings, we explored features of disease-affected cell types and cell classes, highlighted cell types affected by mitochondrial diseases and heritable cancers, and identified diseases that perturb intercellular communication. This study expands our understanding of disease mechanisms and cellular vulnerability.</p>

opencc-zeroJan 2024View details →
zenodo36/100

Efficient statistical method for single-cell QTL analysis

<p>This upload contains data objects associated with our paper "Efficient statistical method for single-cell QTL analysis" introducing the SAIGE-QTL method (preprint available soon!).</p>

opencc-by-4.0Mar 2024View details →
zenodo36/100

SeuratExtend Tutorial: Curated Example Datasets for Single-Cell Analysis

<p>This repository contains example datasets specifically curated for the SeuratExtend tutorial, aimed at facilitating advanced analyses and visualization techniques in single-cell genomics. The datasets have been derived from publicly available data obtained from the 10X Genomics website and have undergone careful preprocessing to serve specific tutorial goals.</p> <p>The collection includes the following datasets:</p> <ol> <li> <p><strong>Myeloid Subset from PBMC 10k Dataset:</strong> This subset focuses on myeloid cells extracted from the larger PBMC 10k dataset, showcasing a preprocessed SeuratObject stored as an RDS file. The data serve as a primary example for demonstrating the capabilities of SeuratExtend differentiation trajectory analysis.</p> </li> <li> <p><strong>Velocyto LOOM File of Myeloid Subset from PBMC 10k Dataset:</strong> Accompanying the first dataset, this Velocyto-generated LOOM file represents a subset of the same myeloid cells, focusing on RNA velocity analyses. It provides a dynamic perspective on gene expression changes over time, enriching the tutorial with advanced single-cell transcriptomics insights.</p> </li> <li> <p><strong>SCENIC-Processed PBMC 3k Dataset:</strong> An outcome of running the SCENIC workflow on the PBMC 3k dataset, this LOOM file represents a refined dataset highlighting gene regulation networks. It serves as an advanced example for users interested in exploring gene regulatory mechanisms using SeuratExtend.</p> </li> </ol> <p>Each dataset has been subsetted and processed, making them ideal for users ranging from beginners to advanced researchers in the field of single-cell genomics. The provided data are intended for educational and tutorial purposes, allowing users to gain hands-on experience with real-world single-cell analysis scenarios.</p> <p>&nbsp;</p>

opencc-zeroApr 2024View details →
zenodo36/100

Deep cross-omics cycle attention model for joint analysis of single-cell multi-omics data

<p>We proposed DCCA for accurately dissecting the cellular heterogeneity on joint-profiling multi-omics data from the same individual cell by transferring representation between each other.</p>

opencc-by-4.0May 2021View details →
zenodo36/100

Single-cell transcriptome analysis of the in vivo response to viral infection in the cave nectar bat Eonycteris spelaea

<p>Bats are reservoir hosts of many zoonotic viruses with pandemic potential in humans. Here, we<br> utilized single-cell transcriptome sequencing (scRNA-seq) to provide detailed comparative<br> analyses of the immune repertoire and the transcriptional responses in the bat lungs upon in<br> vivo infection with a double-stranded RNA virus, Pteropine orthoreovirus PRV3M. Neutrophils<br> were observed to have basally high IDO1 expression, uniquely amongst mammals currently<br> profiled by scRNA-seq. NK/T cells were the most abundant immune cell type in lung tissue, and<br> included three distinct CD8 + effector T cell populations delineated by the differential expression<br> of KLRB1, GFRA2 and DPP4. We identified NK/T clusters which up-regulated genes involved in<br> T-cell activation and effector function early after viral infection. Alveolar macrophages and<br> classical monocytes were key drivers of antiviral interferon signaling. Infection also resulted in<br> the expansion of a CSF1R + population expressing collagen-like genes, which became the<br> predominant myeloid cell type after infection. This work uncovers novel features relevant to viral<br> disease tolerance in bats, lays a foundation for future in vivo and in vitro experimental<br> investigations, and serves as a key resource for comparative immunology studies across bats<br> and other mammals.</p> <p>&nbsp;</p> <p>This upload is the transcriptome fasta file used for alignment for the dataset.</p>

opencc-by-4.0Aug 2022View details →
zenodo36/100

Analysis of public single-cell sequencing database of COVID lung samples

<p>Lung endothelial cells from three published scRNA-seq datasets (GSE122960, GSE149878, GSE171668) of healthy subjects and COVID-19 patients were collected for further integrative analyses. The endothelial cells were classified into three sub-groups according to their distinguished expression of IL7R, DKK2, and EDNRB. For differential analysis of gene expression, counts per million of aggregated UMIs in each group were adopted in Wilcoxon rank-sum test.</p>

opencc-by-4.0Jun 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record