Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

25,372

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

25,372 results for “Transcriptomics”

Learn how ShareScore rates datasets ↗
zenodo40/100

SMART: Spatial transcriptomics deconvolution using marker-gene-assisted topic model

<p>Source code and simulated datasets used in manuscript "SMART: Spatial transcriptomics deconvolution using marker-gene-assisted topic model"</p>

opengpl-3.0-or-laterDec 2023View details →
zenodo40/100

Corallorhiza maculata genomic and transcriptomic analysis

<p>Novoplasty assemblies of plastid genomes from two different Corallorhiza maculata plants (Circularized_assembly_1_CM_1A.fasta and&nbsp;Circularized_assembly_1_CM_2A.fasta).</p> <p>Spades assembly of total cellular genomic DNA from one Corallorhiza maculata plant (scaffolds.fasta.gz). This assembly includes scaffolds of mitochondrial origin (as well as plastid and nuclear).</p> <p>Trinity assembly of rRNA-depleted RNA-seq reads from one Corallorhiza maculata plant (trinity_out_dir.Trinity.fasta.gz).</p>

opencc-by-4.0Nov 2024View details →
zenodo40/100

An integrated transcriptomic cell atlas of human neural organoids: Full Dataset

<p>This deposition includes the full HNOCA dataset for the following pre-print:&nbsp;</p> <blockquote> <p>He, Z., Dony, L., Fleck, J.S. <em>et al.</em> An integrated transcriptomic cell atlas of human neural organoids. <em>Nature</em> <strong>635</strong>, 690&ndash;698 (2024). https://doi.org/10.1038/s41586-024-08172-8</p> </blockquote> <p><strong>This file contains additional data representations and metadata from intermediate processing steps of the HNOCA. For day-to-day use of the HNOCA as a resource, we recommend using the cleaned-up HNOCA object that can be found together with the disease atlas and the extended version of HNOCA in the&nbsp;<a href="https://doi.org/10.5281/zenodo.11203684" target="_blank" rel="noopener">original Zenodo deposition</a>.</strong></p> <p>&nbsp;</p> <p>Abstract:</p> <p>Neural tissues generated from human pluripotent stem cells in vitro (known as neural organoids) are becoming useful tools to study human brain development, evolution and disease. The characterization of neural organoids using single-cell genomic methods has revealed a large diversity of neural cell types with molecular signatures similar to those observed in primary human brain tissue. However, it is unclear which domains of the human nervous system are covered by existing protocols. It is also difficult to quantitatively assess variation between protocols and the specific cell states in organoids as compared to primary counterparts. Single-cell transcriptome data from primary tissue and neural organoids derived with guided or unguided approaches and under diverse conditions combined with large-scale integrative analyses make it now possible to address these challenges. Recent advances in computational methodology enable the generation of integrated atlases across many data sets. Here, we integrated 36 single-cell transcriptomics data sets spanning 26 protocols into one integrated human neural organoid cell atlas (HNOCA) totaling over 1.7 million cells. We harmonize cell type annotations by incorporating reference data sets from the developing human brain. By mapping to the developing human brain reference, we reveal which primary cell states have been generated in vitro, and which are under-represented. We further compare transcriptomic profiles of neuronal populations in organoids to their counterparts in the developing human brain. To support rapid organoid phenotyping and quantitative assessment of new protocols, we provide a programmatic interface to browse the atlas and query new data sets, and showcase the power of the atlas to annotate new query data sets and evaluate new organoid protocols. Taken together, the HNOCA will be useful to assess the fidelity of organoids, characterize perturbed and diseased states and facilitate protocol development in the future.</p>

opencc-by-4.0Oct 2023View details →
zenodo40/100

An integrated transcriptomic cell atlas of human neural organoids: Cleaned datasets

<p>This deposition includes the datasets in h5ad format for the following publication:&nbsp;</p> <blockquote> <p>He, Z., Dony, L., Fleck, J.S. <em>et al.</em> An integrated transcriptomic cell atlas of human neural organoids. <em>Nature</em> <strong>635</strong>, 690&ndash;698 (2024). https://doi.org/10.1038/s41586-024-08172-8</p> </blockquote> <p><strong>The file `hnoca_cleanedmeta.h5ad` is a cleaned up version of the HNOCA dataset. While it contains all cells, some metadata and representations originating from intermediate processing steps are removed. You can find the full (and significantly larger) object in a <a href="https://doi.org/10.5281/zenodo.12536006" target="_blank" rel="noopener">second Zenodo deposition</a>.</strong></p> <p><strong>You can find the minimal HNOCA and primary reference h5ad files for query-to-reference-mapping in a <a href="https://doi.org/10.5281/zenodo.15004817">third Zenodo deposition</a>.</strong></p> <p>&nbsp;</p> <p>Abstract:</p> <p>Neural tissues generated from human pluripotent stem cells in vitro (known as neural organoids) are becoming useful tools to study human brain development, evolution and disease. The characterization of neural organoids using single-cell genomic methods has revealed a large diversity of neural cell types with molecular signatures similar to those observed in primary human brain tissue. However, it is unclear which domains of the human nervous system are covered by existing protocols. It is also difficult to quantitatively assess variation between protocols and the specific cell states in organoids as compared to primary counterparts. Single-cell transcriptome data from primary tissue and neural organoids derived with guided or unguided approaches and under diverse conditions combined with large-scale integrative analyses make it now possible to address these challenges. Recent advances in computational methodology enable the generation of integrated atlases across many data sets. Here, we integrated 36 single-cell transcriptomics data sets spanning 26 protocols into one integrated human neural organoid cell atlas (HNOCA) totaling over 1.7 million cells. We harmonize cell type annotations by incorporating reference data sets from the developing human brain. By mapping to the developing human brain reference, we reveal which primary cell states have been generated in vitro, and which are under-represented. We further compare transcriptomic profiles of neuronal populations in organoids to their counterparts in the developing human brain. To support rapid organoid phenotyping and quantitative assessment of new protocols, we provide a programmatic interface to browse the atlas and query new data sets, and showcase the power of the atlas to annotate new query data sets and evaluate new organoid protocols. Taken together, the HNOCA will be useful to assess the fidelity of organoids, characterize perturbed and diseased states and facilitate protocol development in the future.</p>

opencc-by-4.0Oct 2023View details →
zenodo40/100

Platynereis dumerilii full-length transcriptome of developmental stages

<p>To generate a high-quality full-length transcriptome for the annelid&nbsp;<em>Platynereis dumerilii</em>, we collected samples from representative developmental stages, from unfertilized eggs to 5 days post-fertilization. Each sample consisted of a bulk mix from 1 to 5 batches of embryos fertilized by different parents. We incubated all batches at 18 degrees Celsius until the desired time point, then collected the embryos into a clean tube and snap-froze them in liquid nitrogen with as little seawater as possible. The samples were stored at -80 degrees Celsius until RNA extraction. We extracted total RNA from the samples using a Trizol protocol. After measuring the RNA concentration with NanoDrop, we created a bulk RNA mix by combining 1 &micro;L from each sample into a new tube. We gave the sample to the Sequencing and Genotyping facility of the Max Planck Institute of Molecular Cell Biology and Genetics, who ran aliquots of this bulk mix through a Bioanalyzer and gel electrophoresis. They found no evidence of RNA degradation. From this sample, they prepared PacBio Iso-Seq libraries using the Express Template Prep Kit 2.0 and sequenced full-length transcripts on a SMRT 8M Cell for 30 hours using a PacBio Sequel II System. They processed the raw movie subreads with SMRT Analysis software, following the Iso-Seq v3 workflow to generate representative circular consensus sequences, demultiplex and remove primers, trim poly(A) tails, and remove concatemers. After transcript clustering and merging, the resulting dataset contained 176,122 polished high-quality isoforms. Using Cogent, we removed redundant isoforms and obtained a dataset with 117,524 transcripts. From this, we generated a dataset containing only the longest isoform for each gene, with 70,003 sequences in total. We calculated descriptive metrics using Transrate. To estimate their completeness, we used BUSCO for metazoa and obtained a score of 85%. Finally, we annotated the longest-isoform dataset using EnTAP. About 85% of the transcripts have a coding sequence. We obtained annotations for 67% of the sequences, while 33% have remained unannotated.</p> <h2>Datasets</h2> <table> <tbody> <tr> <td><strong>file name</strong></td> <td><strong>file size (zipped)</strong></td> <td><strong>sequences</strong></td> <td><strong>description</strong></td> </tr> <tr> <td><a href="https://zenodo.org/records/14250773/files/0-Pdum_workflow.zip?download=1">0-Pdum_workflow.zip</a> (folder)</td> <td>3.40 GB</td> <td>-</td> <td>entire pipeline with notebook entries and analyses</td> </tr> <tr> <td><a href="https://zenodo.org/records/14250773/files/1-Pdum_hq_isoforms.zip?download=1">1-Pdum_hq_isoforms.zip</a> (fasta)</td> <td>180.30 MB</td> <td>176,122</td> <td>polished high-quality isoforms from CCS</td> </tr> <tr> <td><a href="https://zenodo.org/records/14250773/files/2-Pdum_co_isoforms.zip?download=1">2-Pdum_co_isoforms.zip</a> (fasta)</td> <td>70.68 MB</td> <td>117,524</td> <td>non-redundant polished high-quality isoforms</td> </tr> <tr> <td><a href="https://zenodo.org/records/14250773/files/3-Pdum_co_longest.zip?download=1">3-Pdum_co_longest.zip</a> (fasta)</td> <td>54.85 MB</td> <td>70,003</td> <td>longest of non-redundant polished high-quality isoforms</td> </tr> <tr> <td><a href="https://zenodo.org/records/14250773/files/4-Pdum_co_longest_annotations.zip?download=1">4-Pdum_co_longest_annotations.zip</a> (tsv)</td> <td>34.37 MB</td> <td>70,003 (46,635 annotated)</td> <td>annotations for longest-isoform dataset</td> </tr> </tbody> </table> <p>&nbsp;</p>

opencc-by-4.0Nov 2024View details →
zenodo40/100

Innateness Transcriptome Gradients Characterize Mouse T Lymphocyte Populations

<p>Whole blot image repository for :</p> <p><strong>Innateness Transcriptome Gradients Characterize Mouse T Lymphocyte Populations</strong></p> <p>Gabriel Ascui<sup>1,2,3 </sup>*, Viankail Cedillo-Castelan<sup>1 </sup>*, Alba Mendis<sup>1</sup>, Eleni Phung<sup>1</sup>, Hsin-Yu Liu<sup>1</sup>, Greet Verstichel<sup>1</sup>, Shilpi Chandra<sup>1</sup>, Mallory P. Murray<sup>1,3</sup>, Cindy Luna<sup>1</sup>, Hilde Cheroutre<sup>1</sup>, Mitchell Kronenberg<sup>1,2,3</sup>.</p> <p><sup>1</sup> La Jolla Institute for Immunology, La Jolla, California, US; <sup>2</sup> Department of Molecular Biology, University of California San Diego, La Jolla, California, US, US; <sup>3</sup> Immunological Genome Project Consortium.</p> <p>*: Equal contribution.</p> <p><strong>Corresponding author</strong>: <a href="mailto:mitch@lji.org">mitch@lji.org</a></p> <p>&nbsp;</p> <p>Image Description:&nbsp;</p> <p>Blots were revealed and later stripped for next primary and secondary antibody staining.&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Jan 2024View details →
zenodo40/100

Memory B Cells and Their Transcriptomic Profiles Associated with Belimumab Resistance in Systemic Lupus Erythematosus in the Maintenance Phase

<p>Normalized count of RNA sequences from systemic lupus erythematosus patients treated with belimumab.</p> <p>We collected blood samples from patients with systemic lupus erythematosus (n=44) before and approximately 3 and 6 months after treatment by belimumab, respectively. We also included healthy individuals (n=17). There were no missing samples in the before-treatment specimen for 44 patients; however, five samples were missing at the three-month time, and four were missing at the six-month time after treatment. Whole blood samples were stored in PAXgene tubes (QUIAGEN). RNA was extracted by PAXgene Blood RNA Kit (QUIAGEN). Library preparation was performed using TruSeq stranded Total RNA Library PrepKit with Ribo-Zero Globin Human. Sequencing was conducted by NovaSeq6000 in the 100-bp paired-end mode. Sequencing reads were trimmed using Trimmomatic ver. 0.36 (leading: 20, trailing: 20, slidingwindow: 4:15, minlen: 36) and aligned to hg38 reference genome using STAR (ver. 2.7.3a). Gene counts were generated by RSEM (ver. 1.3.1) using Homo_sapiens.GRCh38.95.gtf from the Ensembl database. Gene counts were normalized by size factor implemented in DESeq2&nbsp;and converted to count per million (CPM), and log<sub>2</sub>(CPM+1) was calculated.&nbsp;</p> <p>Sample names starting with 'R_' indicate responders, while those ending with '_NR' indicate non-responders. '0M' represents before treatment, '3M' represents three months after treatment, and '6M' represents six months after treatment. The numeric ID, excluding the last two characters, represents individual patients.</p> <p>&nbsp;</p> <p>If you use this data, please cite Iwasaki et al. Front. Immunol., 04 February 2025 <a href="https://doi.org/10.3389/fimmu.2025.1506298">https://doi.org/10.3389/fimmu.2025.1506298</a></p>

opencc-by-4.0Oct 2024View details →
zenodo40/100

Simulation data for benchmarking de novo long read transcriptome assembly software

<p>Method of simulation of differentially expressed biological replicates</p> <p>We first obtained a subset of transcripts that are widely expressed in the GTEx v9 dataset (92 samples) using Gencode comprehensive annotation (v44). We kept transcripts with more than 5 reads in at least 15 samples after Salmon quantification (18145 genes, 40509 transcripts), and stored their mean count per million (CPM) values as the control group&rsquo;s baseline expression. We then generated a perturbed set of CPM values where transcript expression was changed by: (1) randomly selecting 1000 genes and changing all transcripts belonging to that gene concordantly (500 genes 2 fold up and 500 genes 2 fold down), (2) selected another 1000 genes randomly, and then select 2 random transcripts from the gene and swap their expression, (3) selected another 1000 genes randomly, and then select 1 random transcript to change its expression (500 transcripts 2 fold up and 500 transcripts 2 fold down). The updated CPM were stored as the perturbed group baseline expression. We then generated a count matrix and CPM matrix for 3 control replicates and 3 perturbed replicates with gamma distribution, followed by a Poisson distribution <a href="https://www.zotero.org/google-docs/?cUP4ui">(Baldoni et al., 2024)</a>. Both long-read and short-read FASTQ files were simulated using SQANTI-SIM with default settings and ONT R9.4 cDNA error profile (v 0.2.1) <a href="https://www.zotero.org/google-docs/?Qyopst">(Mestre-Tom&aacute;s et al., 2023)</a>. The long read data contained 6 million reads in total, and an average read length of 1085 bp, and short read data was 100 bp paired-end. We then subsampled the short-read data to match the total number of base pairs in the long read data (6.5 billion bases). The simulated data was non-stranded, and contains 2000 DE genes, 2000 genes with DTU, 5927 transcripts with DTU and 6933 DE transcripts.</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2024View details →
zenodo40/100

Phenotype Driven Data Augmentation Methods for Transcriptomic Data

<p>This repository contains the data and associated results of all experiments conducted in our work "<em>Phenotype Driven Data Augmentation Methods for Transcriptomic Data</em>". In this work, we introduce two classes of phenotype driven data augmentation approaches &ndash; signature-dependent and signature-independent. The signature-dependent methods assume the existence of distinct gene signatures describing some phenotype and are simple, non-parametric, and novel data augmentation methods. The signature-independent methods are a modification of the established Gamma-Poisson and Poisson sampling methods for gene expression data. We benchmark our proposed methods against random oversampling, SMOTE, unmodified versions of Gamma-Poisson and Poisson sampling, and unaugmented data.&nbsp;<br>&nbsp;</p> <p>This repository contains data used for all our experiments. This includes the original data based off which augmentation was performed, the cross validation split indices as a json file, the training and validation data augmented by the various augmentation methods mentioned in our study, a test set (containing only real samples) and an external test set standardised accordingly with respect to each augmentation method and training data per CV split.&nbsp;</p> <p>The compressed files&nbsp;<code>5x5stratified_{x}percent.zip</code>&nbsp;contains data that were augmented on <code>x%</code> of the available real data.&nbsp;<code>brca_public.zip</code> contains data used for the breast cancer experiments. <code>distribution_size_effect.zip</code> contains data used for hyperparameter tuning the reference set size for the modified Poisson and Gamma-Poisson methods.&nbsp;</p> <p>The compressed file&nbsp;<code>results.zip</code>&nbsp;contains all the results from all the experiments. This includes the parameter files used to train the various models, the metrics (balanced accuracy and auc-roc) computed including p-values, as well as the latent space of train, validation and test (for the (N)VAE) for all 25 (5x5) CV splits.</p> <p><strong>PLEASE NOTE:</strong>&nbsp;If any part of this repository is used in any form for your work, please&nbsp;<strong>attribute</strong>&nbsp;the following, in addition to attributing the original data source &nbsp;- TCGA, CPTAC, GSE20713 and METABRIC, accordingly:</p> <pre>@article{janakarajan2025phenotype,<br>&nbsp; title={Phenotype driven data augmentation methods for transcriptomic data},<br>&nbsp; author={Janakarajan, Nikita and Graziani, Mara and Rodr{\'\i}guez Mart{\'\i}nez, Mar{\'\i}a},<br>&nbsp; journal={Bioinformatics Advances},<br>&nbsp; volume={5},<br>&nbsp; number={1},<br>&nbsp; pages={vbaf124},<br>&nbsp; year={2025},<br>&nbsp; publisher={Oxford University Press}<br>}</pre> <p>&nbsp;</p>

opencc-by-4.0Sep 2023View details →
zenodo40/100

A collection of draft gene regulatory networks and perturbation transcriptomics data

<p>These&nbsp;are&nbsp;collections of previously published gene regulatory networks and perturbation transcriptomics data&nbsp;analyzed in our manuscript "A systematic comparison of computational methods for expression forecasting". For more information and related code, see&nbsp;https://github.com/ekernf01/perturbation_benchmarking .&nbsp;</p>

opencc-by-4.0Dec 2023View details →
zenodo40/100

Datasets for CASSL: A cell-type annotation method for single cell transcriptomics data using semi-supervised learning

<p>This repository contains datasets used in the project CASSL:&nbsp;A cell-type annotation method for single cell transcriptomics data using semi-supervised learning. This project aims at learning cell annotations for missing cell labels via NMF and recursive k-Means clustering.</p>

opencc-by-4.0Nov 2021View details →
zenodo40/100

Transcriptomic profiling of clobetasol propionate-induced immunosuppression in challenged zebrafish embryos

<p>We have conducted an immune challenge experiment on 48h zebrafish embryos, which were previously treated with the immunosuppressive drug clobetasol propionate (CP). The embryos' immune system was challenged by injection of a mix of different pathogen associated molecular patterns (PAMPs). RNA sequencing was performed in order to detect altered molcular expression levels induced by CP, PAMPs and a combination of both. The data was published in <a href="https://doi.org/10.1016/j.ecoenv.2022.113346" target="_blank" rel="noopener">Essfeld <em>et al.</em> 2022</a>.</p> <p>The uploaded data archive (<a href="https://www.7-zip.org/">7-zip</a> compressed) consists of three major data types:<br>1. MultiQC reports from raw RNA-Seq read processing and QC (50bp SR)<br>2. Result tables from differential gene expression analysis (DGEA) with DESEq2<br>3. Result tables from Overrespresentation Analysis (ORA) with clusterProfiler</p> <p>Gene count normalization and DGEA was conducted with DESeq2 (<a href="https://genomebiology.biomedcentral.com/articles/10.1186/s13059-014-0550-8">Love et al., 2014</a>, DOI 10.1186/s13059-014-0550-8) . Three biological replicates per condition, exposure treatments were compared with respect to the control group in a pairwise fashion, applying Wald&rsquo;s t-test. P values were corrected for multiple testing with independent hypothesis weighting (IHW) (<a href="https://www.nature.com/articles/nmeth.3885">Ignatiadis et al., 2016</a>, DOI 10.1038/nmeth.3885 ) after Benjamini-Hochberg (BH). To improve the signal to statistical noise ratio, the obtained log<sub>2</sub>-fold change (lfc) values were shrunk with the apeglm method described by Zhu and colleagues (<a href="https://academic.oup.com/bioinformatics/article/35/12/2084/5159452?login=true">2019</a>, DOI 10.1093/bioinformatics/bty895 ) before DGEA result tables were subjected to ORA via clusterProfiler (<a href="https://www.liebertpub.com/doi/10.1089/omi.2011.0118">Yu et al., 2012</a>, DOI 10.1089/omi.2011.0118).</p> <p>The ArrayExpress accession number E-MTAB-11092, provides access to the raw and DESeq2 normalized gene count matrices upon which these analysis were performed. Genes were annotated through the biomaRt package (<a href="https://www.nature.com/articles/nprot.2009.97.pdf?origin=ppub">Durinck et al., 2009</a>, DOI 10.1038/nprot.2009.97 ) in R (<a href="https://www.r-project.org/">R Core Team 2021</a>).</p>

opencc-by-4.0Dec 2021View details →
zenodo40/100

raw data for Optogenetic Stimulation of Prelimbic Pyramidal Neurons Maintains Fear Memories and Modulates Amygdala Pyramidal Neuron Transcriptome

<p>Figure Legend &nbsp;</p> <p>Modulation of cellular excitability of Prelimbic (PrL) pyramidal neurons by optogenetic stimulation. (<strong>A</strong>) Representative traces in current-clamp configuration reporting evoked firing activity triggered by a series of depolarizing current steps (0 to 400 pA) applied to PrL pyramidal neurons of SHAM FEAR (black,&nbsp;<em>n</em>&nbsp;= 8 neurons from 5 mice), OPTO FEAR (red,&nbsp;<em>n</em>= 8 neurons from 5 mice), and No-EX (green,&nbsp;<em>n</em>&nbsp;= 5 neurons from 3 mice) groups. The cumulative plot shows the changes in firing activity. (<strong>B</strong>) Representative traces of PrL pyramidal neurons of SHAM FEAR (black,&nbsp;<em>n</em>&nbsp;= 8 neurons from 5 mice), OPTO FEAR (red,&nbsp;<em>n</em>= 8 neurons from 5 mice), and No-EX (green,&nbsp;<em>n</em>&nbsp;= 5 neurons from 3 mice) groups showing the firing activity triggered by linear depolarization from 0 to 800 pA. Graph (on the right) reports the effects of optogenetic stimulation on rheobase value. Namely, PrL pyramidal neurons of OPTO FEAR and No-EX groups recorded after optogenetic stimulation showed a clear reduction in the rheobase value in comparison to neurons of SHAM FEAR group (* at least&nbsp;<em>p</em>= 0.01). (<strong>C</strong>) Representative traces of Excitatory Post-Synaptic Currents (EPSC) of PrL pyramidal neurons of SHAM FEAR (black,&nbsp;<em>n</em>&nbsp;= 8 neurons from 5 mice), OPTO FEAR (red,&nbsp;<em>n</em>= 8 neurons from 5 mice), and No-EX (green,&nbsp;<em>n</em>&nbsp;= 5 neurons from 3 mice) groups. Graph plot (in the middle) and cumulative curve (on the right) depict the clear increase in firing frequency in PrL pyramidal neurons of OPTO FEAR and No-EX groups (* at least&nbsp;<em>p</em>&nbsp;= 0.01). (<strong>D</strong>) Graphs and cumulative curves report no significant differences in cellular excitability in PrL pyramidal neurons of SHAM NOT FEAR (black) and OPTO NOT FEAR (blue) groups. Data are reported as median with interquartile range.</p>

opencc-by-4.0Jan 2021View details →
zenodo40/100

Raw Data for the article: Changes in the Transcriptome Profiles of Human Amnion-Derived Mesenchymal Stromal/Stem Cells Induced by Three-Dimensional Culture: A Potential Priming Strategy to Improve Their Properties

<p>Mesenchymal stromal/stem cells (MSCs) are believed to function in vivo as a homeostatic tool that shows therapeutic properties for tissue repair/regeneration. Conventionally, these cells are expanded in two-dimensional (2D) cultures, and, in that case, MSCs undergo genotypic/phenotypic changes resulting in a loss of their therapeutic capabilities. Moreover, several clinical trials using MSCs have shown controversial results with moderate/insufficient therapeutic responses. Different priming methods were tested to improve MSC effects, and three-dimensional (3D) culturing techniques were also examined. MSC spheroids display increased therapeutic properties, and, in this context, it is crucial to understand molecular changes underlying spheroid generation. To address these limitations, we performed RNA-seq on human amnion-derived MSCs (hAMSCs) cultured in both 2D and 3D conditions and examined the transcriptome changes associated with hAMSC spheroid formation. We found a large number of 3D culture-sensitive genes and identified selected genes related to 3D hAMSC therapeutic effects. In particular, we observed that these genes can regulate proliferation/differentiation, as well as immunomodulatory and angiogenic processes. We validated RNA-seq results by qRT-PCR and methylome analysis and investigation of secreted factors. Overall, our results showed that hAMSC spheroid culture represents a promising approach to cell-based therapy that could significantly impact hAMSC application in the field of regenerative medicine.</p>

opencc-by-4.0Mar 2022View details →
zenodo40/100

Natural Killer cells demonstrate distinct eQTL and transcriptome-wide disease associations, highlighting their role in autoimmunity.

<p><strong>Abstract</strong>&nbsp;</p> <p>Natural Killer (NK) cells are innate lymphocytes with central roles in immunosurveillance and are implicated in autoimmune pathogenesis. The degree to which regulatory variants affect NK gene expression is poorly understood. We performed expression quantitative trait locus (eQTL) mapping of negatively selected NK cells from a population of healthy Europeans (n=245). We find a significant subset of genes demonstrate eQTL specific to NK cells and these are highly informative of human disease, in particular autoimmunity. An NK cell transcriptome-wide association study (TWAS) across five common autoimmune diseases identified further novel associations at 27 genes. In addition to these <em>cis</em> observations, we find novel master-regulatory regions impacting expression of <em>trans</em> gene networks at regions including 19q13.4, the Killer cell Immunoglobulin-like Receptor (KIR) Region, <em>GNLY</em>,&nbsp;<em>MC1R</em> and<em> UVSSA</em>. Our findings provide new insights into the unique biology of NK cells, demonstrating markedly different eQTL from other immune cells, with implications for disease mechanisms.</p> <p><strong>Preprint</strong></p> <p>https://www.biorxiv.org/content/10.1101/2021.05.10.443088v1</p> <p><strong>Dataset</strong></p> <p>nk_raw_for_zenodo.txt: Matrix of raw gene expression at 47,209 probes in primary human NK cells from 245 healthy individuals of European ancestry. Gene expression is quantified using the&nbsp;Illumina HumanHT-12 v4 BeadChip gene expression array platform. Column names represent Array Address ID for each probe, and row names represent pseudonymised sample identifiers, which can be matched to sample genotypes.&nbsp;Sample genotypes are available at the European Genome-Phenome Archive with accession ID EGAS00000000109).</p> <p>probes_passing_QC.txt: List of probes passing quality control; probe sequences mapping to a unique genomic locus, and probe sequences not containing common genomic variation (minor allele frequency &gt;1%), n=29,002. Column names are Array Address ID (probeID), Ensembl ID (ensembl), Gene ID (gene), and Illumina probe ID (ilmn).&nbsp;</p>

opencc-by-4.0Mar 2022View details →
zenodo40/100

Comparative host transcriptomics as a tool to identify candidate biomarkers for immune reactions in leprosy: A meta-analysis study

<p>The&nbsp;dataset consists of R&nbsp;source code for the individual dataset analysis of the studies and their meta-analysis. It also contains supplementary tables and figure.</p>

opencc-by-4.0Mar 2022View details →
zenodo40/100

Computational Analysis of Two-dimensional High-throughput Data from Large-scale RNAi Screens and Single-cell Transcriptomics

<p>This publication&nbsp;provides&nbsp;a singularity definition file to reproduce the computational environment along with the scripts to reproduce every figure or table in the revised manuscript using ZetaSuite Perl module and R package.</p> <p>First, generate a new folder and then download all the files into the folder.</p> <p>Then, uncompressed the files DataSets_part1.tar.gz,DataSets_part2.tar.gz,DataSets_part3.tar.gz,DataSets_part4.tar.gz, and scripts.tar.gz. within the folder.</p> <p>Next, move all the files in DataSets_part1 folder,&nbsp;DataSets_part2&nbsp;folder,DataSets_part3&nbsp;folder and&nbsp;DataSets_part4&nbsp;folder to a new folder called DataSets.</p> <p>Finally, run the following scripts to generate the&nbsp;figures and tables in our manuscript.</p> <p>Regeneration of Figure2 and S2: singularity exec ZetaSuite.sif sh Figure2andS2.sh&nbsp;&nbsp;</p> <p>Regeneration of Figure3 and S3: singularity exec ZetaSuite.sif sh Figure3andS3.sh&nbsp;&nbsp;</p> <p>Regeneration of Figure4 and S4: singularity exec ZetaSuite.sif sh Figure4andS4.sh&nbsp;&nbsp;</p> <p>Regeneration of Figure5 and S5: singularity exec ZetaSuite.sif sh Figure5andS5.sh&nbsp;&nbsp;</p> <p>Regeneration of Figure6 and S6: singularity exec ZetaSuite.sif sh Figure6andS6.sh&nbsp;&nbsp;</p> <p>Regeneration of Figure7 and S7: singularity exec ZetaSuite.sif sh Figure7andS7.sh&nbsp;&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Mar 2022View details →
zenodo40/100

de novo transcriptome assemblies for Peperomia dahlstedtii and Peperomia pellucida

<p><em>De novo</em> transcriptome assemblies generated using Trinity using young leaf tissue from <em>Peperomia dahlstedtii</em> and <em>Peperomia pellucida</em>.&nbsp;</p>

opencc-by-3.0-usDec 2021View details →
zenodo40/100

Development of a machine learning model to predict non- durable response to anti-TNF therapy in Crohn's disease using transcriptome imputed from genotypes

<p>This is the expression value predicted using PrediXcan version 7 to find a gene feature that can distinguish between patients with and without effect on infliximab.</p> <p>Among the various tissue models provided by PrediXcan v7, three models were selected and used: whole blood, Colon&nbsp;transverse, and terminal ileum of small intestine, and the predicted gene counts of each model were 6,294, 5,612 and 3,107.</p> <p>For each of the three models, predicted gene expression values and phenotype information per sample were submitted.</p>

opencc-by-4.0Apr 2022View details →
zenodo40/100

Supplementary files for Transcriptomic alterations in roots of two contrasting Coffea arabica cultivars after hexanoic acid priming

<p>Supplementar files for Transcriptomic alterations in roots of two contrasting Coffea arabica cultivars after hexanoic acid priming.</p> <p>Supplementary Table 1. Differentially expressed genes (DEGs) identified in Obat&atilde; and Catua&iacute; cultivars<br> Supplementary Table 2. Gene ontology (GO) enrichment analysis of differentially expressed genes (DEGs) identified in Obat&atilde; and Catua&iacute; cultivars<br> Supplementary Table 3. MapMan pathway analysis of differentially genes expressed (DEGs) identified in Obat&atilde; and Catua&iacute; cultivars<br> Supplementary Table 4. FPKM values for Catua&iacute; cultivar genes and Roots DEGs<br> Supplementary Table 5. FPKM values for Obat&atilde; cultivar genes and Roots DEGs</p>

opencc-by-4.0Apr 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record