Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
6,040
datasets available to search
ShareScore release 0.9.0
Dataset results
6,040 results for “Single-cell”
Robustness and applicability of transcription factor and pathway analysis tools on single-cell RNA-seq data
<p>Data used to test the robustness and applicability of transcription factor and pathway analysis tools on single-cell RNA-seq data, described in <a href="https://doi.org/10.1186/s13059-020-1949-z">Holland et al. 2020</a>.</p> <p>The folder <em>data </em>contains<em> </em>raw data and the folder <em>output</em> contains intermediate and final results of all analyses. </p> <p>The associated analyses code and more information are available on <a href="https://github.com/saezlab/FootprintMethods_on_scRNAseq">GitHub</a>.</p> <p> </p> <p><strong>Abstract</strong></p> <p><strong>Background</strong></p> <p>Many functional analysis tools have been developed to extract functional and mechanistic insight from bulk transcriptome data. With the advent of single-cell RNA sequencing (scRNA-seq), it is in principle possible to do such an analysis for single cells. However, scRNA-seq data has characteristics such as drop-out events and low library sizes. It is thus not clear if functional TF and pathway analysis tools established for bulk sequencing can be applied to scRNA-seq in a meaningful way.</p> <p><strong>Results</strong></p> <p>To address this question, we perform benchmark studies on simulated and real scRNA-seq data. We include the bulk-RNA tools PROGENy, GO enrichment, and DoRothEA that estimate pathway and transcription factor (TF) activities, respectively, and compare them against the tools SCENIC/AUCell and metaVIPER, designed for scRNA-seq. For the in silico study, we simulate single cells from TF/pathway perturbation bulk RNA-seq experiments. We complement the simulated data with real scRNA-seq data upon CRISPR-mediated knock-out. Our benchmarks on simulated and real data reveal comparable performance to the original bulk data. Additionally, we show that the TF and pathway activities preserve cell type-specific variability by analyzing a mixture sample sequenced with 13 scRNA-seq protocols. We also provide the benchmark data for further use by the community.</p> <p><strong>Conclusions</strong></p> <p>Our analyses suggest that bulk-based functional analysis tools that use manually curated footprint gene sets can be applied to scRNA-seq data, partially outperforming dedicated single-cell tools. Furthermore, we find that the performance of functional analysis tools is more sensitive to the gene sets than to the statistic used.</p> <p> </p> <p>For questions related to the data please write an email to christian.holland@bioquant.uni-heidelberg.de or use the <a href="https://github.com/saezlab/FootprintMethods_on_scRNAseq/issues">GitHub issue system</a>.</p>
Data for "Tuning parameters of dimensionality reduction methods for single-cell RNA-seq analysis"
<p>The files named <code>df_scran.csv</code>, <code>df_seurat.csv</code>, <code>df_zinbwave.csv</code>, <code>df_dca.csv</code>, and <code>df_scvi.csv</code> contain one row per configuration that we ran successfully.</p> <p>The files named <code>DATASET.METHOD.h5ad</code> are encoded with anndata <code>v0.7.0</code> (be careful as they are not readable with previous versions) and contain 100 embeddings each. The embeddings are in the <code>obsm</code> attribute of the object. All the embeddings can be listed with the <code>obsm_keys()</code> method. The name of the embedding contains the parameters used to generate that embedding and are written like that <code>method=zinbwave.dims=10.epsilon=1000.features=300.gene_covariate=0</code>.</p> <p> </p> <p>For questions on this dataset please contact fraimundo@google.com</p>
EI Single-Cell RNA-Seq Workshop 2020
<p>Datasets to be used for the "Single-Cell RNA-Seq Workshop 2020" at the Earlham Institute, Norwich, UK.</p>
EpiScanpy: integrated single-cell epigenomic analysis
<p>All files below are in "h5ad" format, which can be opened by the Python package AnnData (see https://anndata.readthedocs.io/en/latest). The files are organized as:</p> <p><br> 1. Single-nucleus methylcytosine sequencing (snmC-seq) of neurons from the frontal cortex of young adult mouse brains from Luo et al., 2017. The files are the preprocessed input for EpiScanpy for multiple CG imputed feature spaces, 100k base pair windows, enhancers, gene bodies, promoters and CH gene bodies.<br> - processed_enhancers_CG_luo_et_al_nov2020_paper_resubmission.h5ad<br> - processed_genebodies_CG_luo_et_al_nov2020_paper_resubmission.h5ad<br> - processed_genebodies_CH_luo_et_al_nov2020_paper_resubmission.h5ad<br> - processed_promoter_CG_luo_et_al_nov2020_paper_resubmission.h5ad<br> - processed_windows_CG_luo_et_al_nov2020_paper_resubmission.h5ad<br> </p> <p>2. Single cell ATAC sequencing (scATAC-seq) from 10X Genomics, preprocessed peak count matrix of Next GEM v1.1 10k Peripheral blood mononuclear cells (PBMCs) from a healthy donor. Peak matrices of Next GEM v1.1 10k PBMCs and whole blood fresh data (GEO:GSE129785) from Satpathy et al. 2019 (Greenleaf's lab). Concatenated 5000bp matrix of Fresh cortex from adult mouse brain (P50) from 10X Genomics and CEMBA180312_3B mouse brain sample from Fang et al. 2019.<br> - atac_pbmc_10k_nextgem_fragments_macs2_peaks_outter_all_chrom.h5ad<br> - atac_pbmc_10k_nextgem_fragments_merged_peaks_for_integration_greenleaf_outter_all_chrom.h5ad<br> - raw_greenleaf_pre_integration_oct_2020.h5ad (Satpathy et al. 2019)<br> - preprocessed_10x_genomics_5k_adulte_mouse_brain_Fang_et_al_2019_CEMBA180312_3B_5kb_windows.h5ad</p>
Beyondcell: targeting cancer therapeutic heterogeneity in single-cell RNA-seq
<p><strong><a href="https://gitlab.com/bu_cnio/Beyondcell">Beyondcell</a> </strong>is a methodology for the identification of drug vulnerabilities in single cell RNA-seq data. To this end, <strong>Beyondcell</strong> focuses on the analysis of drug-related commonalities between cells by classifying them into distinct therapeutic clusters. We have validated the tool in a population of MCF7-AA cells exposed to 500nM of bortezomib and collected at different time points: t0 (before treatment), t12, t48 and t96 (72h treatment followed by drug wash and 24h of recovery) obtained from <a href="https://www.nature.com/articles/s41586-018-0409-3"><strong><em>Ben-David U, et al., Nature, 2018</em></strong></a>. Here, you can find the integrated Seurat object obtained from this analysis. This object is meant to help users follow <strong>Beyondcell's</strong> <a href="https://gitlab.com/bu_cnio/Beyondcell/-/tree/master/tutorial/analysis_workflow">analysis workflow</a>.</p> <p> </p>
Direct chromosome-length haplotyping by single-cell sequencing.
<p>Selected Strand-seq libraries from PMID:27646535 study. Data were originally shared on the European Nucleotide Archive (http://www.ebi.ac.uk/ena) under the accession number: PRJEB14185</p>
Single-cell RNA sequencing identifies shared differentiation paths of mouse thymic innate T cells
<p>scRNA sequencing datasets used in the paper titled 'Single-cell RNA sequencing identifies shared differentiation paths of mouse thymic innate T cells' published in Nature Communications<br> <br> https://www.nature.com/articles/s41467-020-18155-8</p>
Single-cell datasets for cell cycle plasticity underlies fractional resistance to palbociclib in ER+/HER2- breast tumor cells
<p>There are 7 files uploaded in the data.</p><p>tumor_preprocessed.h5ad: Full primary tumor dataset post-feature selection and standardization across three treatment conditions (0, 10, and 100 nM palbociclib). AnnData object format. 14 cell cycle features, phase labels and other cell metadata, and two PHATE dimensions for manifold visualization.</p><p>T47D_preprocssed.h5ad: Full dataset of main text T47D dataset post-feature selection and standardization across three treatment conditions (0, 10, and 100 nM palbociclib). AnnData object format. 14 cell cycle features, phase labels and other cell metadata, and two PHATE dimensions for manifold visualization.</p><p>sketched_integrated.h5ad: After downsample 6,000 (2,000 per condition) from T47D_preprocessed and tumor_preprocessed, we integrate the two datasets into one joint latent space using TRANSACT. Now included in the data are the consensus component columns ('0',..,'13'). AnnData object.</p><p>sketched_integrated_df.csv: sketched_integrated.h5ad in .csv format.</p><p>T47D_replicate_preprocessed: Replicate experimental dataset of T47D for supplementary analysis post-feature selection and standardization across three treatment conditions (0, 10, and 100 nM palbociclib). 15 cell cycle features (same 14 but with CDK6).</p><p>sketched_rep.h5ad: Representative downsample of the T47D_replicate_preprocessed. Selecting 6,000 cells (2,000 for each of the three treatment conditions) using kernel herding sketching. AnnData object.</p><p>sketched_rep_df.csv: Same data as sketched_rep.h5ad in csv format.</p><p>T47D_triplicate_preprocessed.h5ad: T47D biological replicate sample collected in triplicate form (three wells for 0, 10, and 100 nM of palbociclib). Wells were joined and the data were sketched down to 20,000 per condition.</p><p>T47D_triplicate_preprocessed.h5ad: T47D triplicate in .csv form.</p><p>tumor_2_preprocessed.h5ad: An additional tumor sample from a new patient with the same treatment conditions of palbociclib. Sketched down to 2,000 cells per condition.</p><p>tumor_2_preprocessed.csv: The additional tumor sample in .csv form.</p><p> </p><p>Further description of sketched_integrated: This is the joint dataset between the T47D and primary tumor, after subsampling using kernel herding sketching. This is a dataset consisting of T47D and primary tumor cells resected from a consented patient. The samples were imaged using iterative indirect immunofluorescent imaging (4i) to get proteomic measurements on a single-cell level. The T47D and tumor samples were gathered, cultured, and imaged separately. Each sample was treated with three conditions of CDK4/6 inhibitor palbociclib (control, 10 nM, and 100 nM). Then, we used kernel sketching to representatively downsample each dataset, selecting 2,000 from each of the three treatment conditions (6,000 cells from each of the two sources). We used an integration method called TRANSACT to integrate the two datasets into one shared, latent space. The dataset here is consisting of these 12,000 cells. The columns ('0','1',...'13') are the principal vectors of the joint latent space. After that, there are the columns of the standardized proteomic measurements of different cell cycle effectors, and biological annotations of interest. The standardization is done for each data source separately. Well refers to the treatment condition. 'prb_ratio' is a marker of if a cell is still proliferating or arrested, found by selecting the upper modality of pRB/RB values. 'phase' are cell cycle phase labels found by unsupervised clustering done on a handful of known cell cycle markers.</p>
Single-cell mouse and PC9 data for "TP53 loss with whole genome doubling mediates heterogeneous intra-patient therapy response through Chromosomal Instability"
<p>This repository includes the processed data (including copy number profiles and related analysis) for the E/EP mouse tumors and for the PC9 resistance cell lines for all the analyses of the manuscript "TP53 loss with whole genome doubling mediates heterogeneous intra-patient therapy response through Chromosomal Instability".</p><p>The code for the related analyses is available in GitHub at https://github.com/zaccaria-lab/TP53loss_WGD</p>
Anabaena circadian clock behavior under nitrogen-poor conditions from single-cell measurements of fluorescence intensity
<p>Circadian clock arrays in multicellular filaments of the heterocyst-forming cyanobacterium Anabaena sp. strain PCC 7120 display remarkable spatio-temporal coherence under nitrogen-replete conditions. To shed light on the interplay between circadian clocks and the formation of developmental patterns, we followed the expression of a clock-controlled gene under nitrogen deprivation, at the level of individual cells. Our experiments showed that differentiation into heterocysts took place preferentially within a limited interval of the circadian clock cycle, that gene expression in different vegetative intervals along a developed filament was discoordinated, and that the circadian clock was active in individual heterocysts. Furthermore, Anabaena mutants lacking the kaiABC genes encoding the circadian clock core components produced heterocysts but failed in diazotrophy. Therefore, genes related to some aspect of nitrogen fixation, rather than early or mid-heterocyst differentiation genes, are likely affected by the absence of the clock. A bioinformatics analysis supports the notion that RpaA may play a role as master regulator of clock outputs in Anabaena, the temporal control of differentiation by the circadian clock and the involvement of the clock in proper diazotrophic growth. Together, these results suggest that under nitrogen-deficient conditions, the clock coherent unit in Anabaena is reduced from a full filament under nitrogen-rich conditions to the vegetative cell interval between heterocysts.</p>
Molecular features of luminal breast cancer defined through spatial and single-cell transcriptomics (codes and data files)
<p>This dataset includes all the relevant codes and data files associated with the paper ("Molecular features of luminal breast cancer defined through spatial and single-cell transcriptomics") in Clinical and Translational Medicine journal.</p>
Full-length, single-cell RNA-sequencing of human bone marrow subpopulations reveals hidden complexity
<p><a href="http://www.biorxiv.org/content/10.1101/2021.07.28.454226v2">Full-length, single-cell RNA-sequencing of human bone marrow subpopulations reveals hidden complexity</a></p> <p>Bone marrow progenitor cell differentiation has frequently been used as a model for studying cellular plasticity and cell-fate decisions. Recent analysis at the level of single-cells has expanded knowledge of the transcriptional landscape of human hematopoietic cell lineages. Using single-molecule real-time (SMRT) full-length RNA sequencing, we have previously shown that human bone marrow lineage-negative (Lin-neg) cell populations contain a surprisingly diverse set of mRNA isoforms. Here, we report from single cell, full-length RNA sequencing that this diversity is also reflected at the single-cell level. From fresh human bone marrow unselected and lineage-negative progenitor cells were isolated by droplet-based single-cell selection (10xGenomics). The single cell-derived mRNAs were analyzed by full-length SMRT and short-read sequencing. In both samples we detected an average of 8000 different genes using short-read sequencing. Differential expression analysis arranged the single-cells of the total bone marrow into only four clusters whereas the Lin-neg population was much more diverse with nine clusters. mRNA isoform analysis of the single-cell populations using full-length sequencing revealed that Lin-neg cells contain on average 24% more novel splice variants than the total bone marrow cells. Interestingly, among the most frequent genes expressing novel isoforms were members of the spliceosome, e.g. HNRNPs, DEAD box helicases and SRSFs. Mapping the isoforms from all genes to the cell type clusters revealed that total bone marrow cells express novel isoforms only in a small subset of clusters. On the other hand, lineage-negative progenitor cells expressing novel isoforms were present in nearly all subpopulations. In conclusion, on a single-cell level lineage-negative cells express a higher diversity of genes and more alternatively spliced novel isoforms suggesting that cells in this subpopulation are poised for different fates. </p> <p> </p>
Single-cell RNA-seq profiles of lung adenocarcinoma patients and tumor-bearing mice
<p>single-cell RNA sequencing (scRNA-seq) profiles from eight patients with lung adenocarcinoma (LUAD) and four samples of tumor tissues from tumor bearing mice were performed. By integrating other scRNA-seq data and clinical information, we identified activated adaptive immune responses in older patients, reflected by enriched dysfunctional T cell signature scores and immune checkpoint molecules. Our study shows increased efficacy of immune checkpoint blockade therapy in older patients, addressing the prominent role of age when considering immunotherapy.</p>
Automatic monitoring of neural activity with single-cell resolution in behaving Hydra
<p>The ability to record every spike from every neuron in a behaving animal is one of the holy grails of neuroscience. Here, we report coming one step closer towards this goal with the development of an end-to-end pipeline that automatically tracks and extracts calcium signals from individual neurons in the cnidarian <em>Hydra vulgaris</em>. We imaged dually labeled (nuclear tdTomato and cytoplasmic GCaMP7s) transgenic <em>Hydra </em>and developed an open-source Python platform (TraSE-IN) for the Tracking and Spike Estimation of Individual Neurons in the animal during behavior. The TraSE-IN platform comprises a series of modules that segments and tracks each nucleus over time and extracts the corresponding calcium activity in the GCaMP channel. Another series of signal processing modules allows robust prediction of individual spikes from each neuron's calcium signal. This complete pipeline will facilitate the automatic generation and analysis of large-scale datasets of single-cell resolution neural activity in <em>Hydra</em>, and potentially other model organisms, paving the way towards deciphering the neural code of an entire animal.</p>
Single-cell sequencing data of human umbilical cord and placental mesenchymal stem cells
<p>Expression matrix of umbilical cord and placenta single-cell sequencing data from the same donor.Table1 is the umbilical cord and Table2 is the placenta.</p>
Linking regulatory variants to target genes by integrating single-cell multiome methods and genomic distance
<p>The below data are associated with our paper entitled "Linking regulatory variants to target genes by integrating single-cell multiome methods and genomic distance."</p> <p>1) SNP-gene link predictions generated by pgBoost and existing methods SCENT (Sakaue et al. 2024 <em>Nat Genet</em>), Signac (Stuart et al. 2021 <em>Nat Methods</em>), ArchR (Granja et al. 2021 <em>Nat Genet</em>), and Cicero (Pliner et al. 2018 <em>Mol Cell</em>).</p> <p><strong>pgBoost_scores.tsv.gz </strong>contains linking predictions made by pgBoost.</p> <p><strong>constituent_method_scores.tsv.gz</strong> contains linking predictions made by constituent methods.</p> <p><em><span>**NOTE: promoters (+/- 1kb from TSS) and candidate links >500kb are excluded from linking predictions (see manuscript)**</span></em></p> <p>Linking scores and percentiles are reported for each method (pgBoost score, SCENT FDR, Signac correlation, ArchR correlation, Cicero co-accessibility). Rank percentiles are computed as: 1 - (rank / n). When multiple links receive the same score, they are assigned the percentile of the top rank. Links unscored by each method (denoted by zeros* in the linking score column) are assigned a percentile equivalent to the percent of links unscored by the focal method. See the Methods section of the paper for further details on computing linking scores and summarizing scores across cell types and data sets.</p> <p>*Candidate links tested and assigned a co-accessibility of zero by the Cicero method are given a score of 1e-100 in the "Cicero" column to distinguish between unscored candidate links and candidate links assigned a partial correlation of zero (see Pliner et al. 2018 <em>Mol Cell</em>).</p> <p><em>NOTE: The predictions associated with this release (version 2) were generated using an expanded set of data sets, an expanded training set, and corrected TSS coordinates.</em></p> <p>2) GWAS-derived evaluation SNP-gene link evaluation set.</p> <p><strong>gwas_evaluation.tsv</strong>: GWAS-derived evaluation SNP-gene link evaluation set. Column 1 provides SNP coordinates in the format <chr-start-end>. This evaluation framework was proposed by Weeks et al. 2024 <em>Nature Genetics</em> based on fine-mapping results from Kanai et al. <em>medrxiv</em> (see Methods: <em>Evaluation data sets</em> of Dorans et al.). "True" links (gold = 1) are non-coding variants fine-mapped to a focal trait (PIP > 0.1) with a coding variant for exactly one candidate gene within 1 Mb attaining PIP > 0.5 for the same trait. "False" links (gold = 0) are candidate SNP-gene pairs involving a SNP with a "true" link. This file of SNP-gene links was adapted from credible set-gene links <a href="https://github.com/Deylab999/GWAS_benchmark_IGVF/blob/bb91d08cc02d59cdd829eb1430569057ac26c5fe/V2G/ENCODE_E2G_2023/UKBiobank.ABCGene.anyabc.tsv">here</a> (the "truth" column defines true/false links) by identifying SNPs with PIP > 0.1 within each credible set-gene link.</p>
# Single-cell network biology characterizes cell type gene regulation for drug repurposing and phenotype prediction in Alzheimer's disease
<p>Dysregulation of gene expression in Alzheimer’s disease (AD) remains elusive, especially at the cell type level. Gene regulatory network, a key molecular mechanism linking transcription factors (TFs) and regulatory elements to govern target gene expression, can change across cell types in the human brain and thus serve as a model for studying gene dysregulation in AD. However, it is still challenging to understand how cell type networks work abnormally under AD. To address this, we integrated single-cell multi-omics data and predicted the gene regulatory networks in AD and control for four major cell types, excitatory and inhibitory neurons, microglia and oligodendrocytes. Importantly, we applied network biology approaches to analyze the changes of network characteristics across these cell types, and between AD and control. For instance, many hub TFs target different genes between AD and control (rewiring). Also, these networks show strong hierarchical structures in which top TFs (master regulators) are largely common across cell types, whereas different TFs operate at the middle levels in some cell types (e.g., microglia). The regulatory logics of enriched network motifs (e.g., feed-forward loops) further uncover cell type-specific TF-TF cooperativities in gene regulation. The cell type networks are highly modular and several network modules with cell-type-specific expression changes in AD pathology are enriched with AD-risk genes and putative targets of approved and pending AD drugs, suggesting possible cell-type genomic medicine in AD. Finally, using the cell type gene regulatory networks, we developed machine learning models to classify and prioritize additional AD genes. We found that top prioritized genes predict clinical phenotypes (e.g., cognitive impairment) with reasonable accuracy. Overall, this single-cell network biology analysis provides a comprehensive map linking genes, regulatory networks, cell types and drug targets and reveals dysregulated cell type gene dysregulatory mechanisms in AD.</p>
SEACells: Inference of transcriptional and epigenomic cellular states from single-cell genomics data
<p>Processed data for the manuscript "" available on bioRxiv at ""</p> <p>Data is available for the following samples</p> <ol> <li>CD34+ Multiome data : 2 replicates </li> <li>T-cell depleted bone marrow Multiome data: 2 replicates </li> </ol> <p> </p> <p>The following counts and fragments files are available for each replicate </p> <ol> <li><sample>_filtered_feature_bc_matrix.h5: Feature counts from CellRanger ARC</li> <li><sample>_atac_fragments.tsv.gz: ATAC fragments file from Cellranger ARC</li> <li><sample>_atac_fragments.tsv.gz.tbi: Index files for ATAC fragments file from Cellranger ARC</li> </ol> <p> </p> <p>The following scanpy anndata objects are also available</p> <ol> <li>cd34_multiome_rna.h5ad: Anndata object with normalized data, cell type annotations and clusters for the RNA modality of CD34+ hematopoietic stem and progenitor cells </li> <li>cd34_multiome_atac.h5ad: Anndata object with peak counts, cell type annotations and clusters for the ATAC modality of of CD34+ hematopoietic stem and progenitor cells.</li> <li>cd34_multiome_rna.h5ad: Anndata object with normalized data, cell type annotations and clusters for the RNA modality of T-cell depleted bone marrow dataset.</li> <li>bm_multiome_atac.h5ad: Anndata object with peak counts, cell type annotations and clusters for the ATAC modality of T-cell depleted bone marrow dataset.</li> </ol>
Computational Analysis of Two-dimensional High-throughput Data from Large-scale RNAi Screens and Single-cell Transcriptomics
<p>This publication provides a singularity definition file to reproduce the computational environment along with the scripts to reproduce every figure or table in the revised manuscript using ZetaSuite Perl module and R package.</p> <p>First, generate a new folder and then download all the files into the folder.</p> <p>Then, uncompressed the files DataSets_part1.tar.gz,DataSets_part2.tar.gz,DataSets_part3.tar.gz,DataSets_part4.tar.gz, and scripts.tar.gz. within the folder.</p> <p>Next, move all the files in DataSets_part1 folder, DataSets_part2 folder,DataSets_part3 folder and DataSets_part4 folder to a new folder called DataSets.</p> <p>Finally, run the following scripts to generate the figures and tables in our manuscript.</p> <p>Regeneration of Figure2 and S2: singularity exec ZetaSuite.sif sh Figure2andS2.sh </p> <p>Regeneration of Figure3 and S3: singularity exec ZetaSuite.sif sh Figure3andS3.sh </p> <p>Regeneration of Figure4 and S4: singularity exec ZetaSuite.sif sh Figure4andS4.sh </p> <p>Regeneration of Figure5 and S5: singularity exec ZetaSuite.sif sh Figure5andS5.sh </p> <p>Regeneration of Figure6 and S6: singularity exec ZetaSuite.sif sh Figure6andS6.sh </p> <p>Regeneration of Figure7 and S7: singularity exec ZetaSuite.sif sh Figure7andS7.sh </p> <p> </p>
Human single-cell TCR data from irradiated sporozoite trial
<p>This is the MiAIRR-compliant single-cell AIRR-seq data of the study "Clonal evolution and TCR specificity of the human<br> TFH cell response to Plasmodium falciparum CSP" by Wahl <em>et al.,</em> <em>Sci Immunol</em> 7:eabm9644 (2022). The study defines the clonal evolution and epitope specificity of the human cTFH cell response to Plasmodium falciparum CSP.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.