Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
25,372
datasets available to search
ShareScore release 0.7.1
Dataset results
25,372 results for “transcriptomics”
De novo transcriptome assembly of hyperaccumulating Noccaea praecox
<p>Trinity de novo assembly for hyperaccumulating plant species Noccaea praecox (syn. Thlaspi praecox), Brassicaceae. The dataset includes annotations from SwissProt, Pfam, Rfam databases and information on transmembrane regions and signal peptide cleavage sites (annotated using BLAST, HMMER, infernal, tmHMM and signalP through Trinotate). Detailed information on the preprocessing, assembly, post-processing and annotations are described in the ReadMe file.</p> <p>Supplementary material for Bočaj, V., Pongrac, P., Fischer, S. <em>et al.</em> <em>De novo</em> transcriptome assembly of hyperaccumulating <em>Noccaea praecox</em> for gene discovery. <em>Sci Data</em> <strong>10</strong>, 856 (2023). <a href="https://doi.org/10.1038/s41597-023-02776-x">https://doi.org/10.1038/s41597-023-02776-x</a></p>
Library size confounds biology in spatial transcriptomics data
<p>This dataset contains annotated sub-cellular localised spatial measurements from the Visium, Xenium and CosMx platforms. Specifically, it includes datasets analysed in the publication Bhuva et. al, 2023 titled "Library size confounds biology in spatial transcriptomics data". Raw transcript detections are presented. Data is best accessed through the accompanying <em>SubcellularSpatialData</em> R/Bioconductor package. Region files used to annotate individual transcript detections are presented in the form of <a href="https://geojson.org/">GeoJSON</a> files. </p>
SpatialMETA: A Novel Framework for Integrating Spatial Transcriptomics and Metabolomics Data
<p>Multimodal analysis of spatial transcriptomics (ST) and spatial metabolomics (SM) has rapidly advanced for characterizing tissue microenvironments. However, integrating ST and SM data remains challenging due to differing morphologies, resolutions, and batch effects. We developed SpatialMETA (Spatial Metabolomics and Transcriptomics Analysis), a novel method for integrating spatial multi-omics data, which aligns ST and SM to a unified resolution, enables both cross-modal and cross-sample integration to identify ST-SM associated spatial patterns, and provides extensive visualization and analysis capabilities. The datasets for SpatialMETA is avaiable. </p>
Spatial transcriptome data from coronal mouse brain sections after striatal injection of heme and heme-hemopexin
<p>This dataset and the associated Python notebooks are related to the publication "Spatial transcriptome data from coronal mouse brain sections after striatal injection of heme and heme-hemopexin".</p>
Ramonda serbica de novo transcriptome database
<p>Ramonda serbica de novo transcriptome database obtained from desiccated and hydrated leaves.</p>
Manually curated transcriptomics data collection for toxicogenomic assessment of engineered nanomaterials
<p>Toxicogenomics (TGx) approaches are increasingly applied to gain insight into the possible toxicity mechanisms of engineered nanomaterials (ENMs). Omics data can be valuable to elucidate the mechanism of action of chemicals and develop predictive models in toxicology. While vast amounts of transcriptomics data from ENM exposures have already been accumulated, a unified, easily accessible and reusable collection of transcriptomics data for ENMs is currently lacking. In an attempt to improve the FAIRness of already existing transcriptomics data for nanomaterials, we curated a collection of homogenized transcriptomics data from human, mouse and rat ENM exposures <em>in vitro</em> and <em>in vivo</em>.</p>
Data files: Single-cell RNA profiling of Plasmodium vivax-infected hepatocytes reveals parasite- and host- specific transcriptomic signatures and therapeutic targets
<p>Scripts, preprocessed count matrices, and single-cell data objects generated in <strong>“Single-cell RNA profiling of <em>Plasmodium vivax</em><em>-</em>infected hepatocytes reveals parasite- and host- specific transcriptomic signatures and therapeutic targets” </strong></p>
microSPLiT single-cell and bulk transcriptomes analysed with STAR - Pseudomonas putida KT2440/pKJK5
<h3>Description of the data and file structure</h3> <p>Data are displayed as 2 files</p> <p><strong>1. Bulk transcriptomics results (Bulk_STAR.csv)</strong></p> <p>STAR processed data combined in a gene x sample table</p> <p><strong>2. microSPLiT single-cell results (microSPLiT_STARsolo.xlsx)</strong></p> <p>STARsolo processed data combined as sublibraries’ gene associated transcript numbers (UMIs) per cell for the control (E1) and experiment (E2) sublibraries (F1-8) - (1 sublibrary per table).</p> <div> <p> </p> </div>
Single-cell transcriptomic profiling unveils dysregulation of cardiac progenitor cells and cardiomyocytes in a mouse model of maternal hyperglycemia
<p>Congenital heart disease (CHD) is the most prevalent structural malformations of the heart affecting ∼1% of live births. To date, both damaging genetic variations and adverse environmental exposure such as maternal diabetes have been found to cause CHD. Clinical studies show ∼fivefold higher risk of CHD in the offspring of mothers with pregestational diabetes. Maternal pregestational diabetes affects the gene regulatory networks key to proper cardiac development in the fetus. However, the cell-type specificity of these gene regulatory responses to maternal diabetes and their association with the observed cardiac defects in the fetuses remains unknown. To uncover the transcriptional responses to maternal diabetes in the early embryonic heart, we used an established murine model of pregestational diabetes. In this model, we have previously demonstrated an increased incidence of CHD. Here, we show maternal hyperglycemia (matHG) elicits diverse cellular responses during heart development by single-cell RNA-sequencing in embryonic hearts exposed to control and matHG environment. Through differential gene-expression and pseudotime trajectory analyses of this data, we identified changes in lineage specifying transcription factors, predominantly affecting Isl1+ second heart field progenitors and Tnnt2+cardiomyocytes with matHG. Using in vivo cell-lineage tracing studies, we confirmed that matHG exposure leads to impaired second heart field-derived cardiomyocyte differentiation. Finally, this work identifies matHG-mediated transcriptional determinants in cardiac cell lineages elevate CHD risk and show perturbations in Isl1-dependent gene-regulatory network (Isl1-GRN) affect cardiomyocyte differentiation. Functional analysis of this GRN in cardiac progenitor cells will provide further mechanistic insights into matHG-induced severity of CHD associated with diabetic pregnancies.</p>
laparocerus tessellatus adult full-body transcriptome
<p>Pooled transcriptome of two males and two females, Anaga peninsula, Tenerife, 2015.</p> <p>TruSeq Stranded mRNA kit, HiSeq paired-end 100bp.</p> <p>Trinity v2.0.6</p> <p>Longest isoforms only</p> <p> </p> <p> </p>
Applications and raw data for SciPipe genomics and transcriptomics case studies
<p>Accompanying applications and raw data for the genomics and transcriptomics (RNA-Seq) case studies for SciPipe [1] available at https://github.com/pharmbio/scipipe-demo </p> <p>[1] http://scipipe.org</p>
Data from A functional transcriptomics analysis in the relict marsupial Dromiciops gliroides reveals adaptive regulation of protective functions during hibernation
<p>This dataset contains files with the differentially expressed genes, raw counts, DESeq2 analyses and assembled transcriptome of D. gliroides. This information is linked to the manuscript published in Molecular Ecology.</p>
Genome alignments for the project "Whole transcriptome analysis of thousands of FACS-sorted single cells with the single cell nanoCAGE protocol" - Protocol optimization
<p>Genome alignments for data generated in the project "<em>Whole transcriptome analysis of thousands of FACS-sorted single cells with the single cell nanoCAGE protocol – Optimization of the protocol.</em>" Files names indicate unique identifiers of MOIRAI workflow runs, with the following structure: library name, dot, workflow ID (OP-WORKFLOW-CAGEscan-short-reads-v2.0.), dot, timestamp. The raw (FASTQ) data of each library is also deposited in Zenodo (<a href="https://doi.org/10.5281/zenodo.250156">10.5281/zenodo.250156</a>). Library names correspond to the following runs:</p> <ul> <li> NC33: 151007_M00528_0161_000000000-AEBDC</li> <li> NC37: 151204_M00528_0173_000000000-AEBEF</li> <li> NC38: 151211_M00528_0175_000000000-AE9PJ</li> <li> NC39: 160122_M00528_0185_000000000-AEB18</li> <li> NC42: 160302_M00528_0192_000000000-AELYK</li> </ul> <p>This data can be analysed using the "CAGEr" software package available from Bioconductor. The "multiplex_files.zip" file contains tables indicating which samples are biological replicates of each other or negative controls.</p>
BaRTv1.0: an improved barley reference transcript dataset to determine accurate changes in the barley transcriptome using RNA-seq
<p>Background<br> Time consuming computational assembly and quantification of gene expression and splicing analysis from RNA-seq data vary considerably. Recent fast non-alignment tools such as Kallisto and Salmon overcome these problems, but these tools require a high quality, comprehensive reference transcripts dataset (RTD), which are rarely available in plants.</p> <p>Results<br> A high-quality, non-redundant barley gene RTD and database (Barley Reference Transcripts – BaRTv1.0) has been generated. BaRTv1.0, was constructed from a range of tissues, cultivars and abiotic treatments and transcripts assembled and aligned to the barley cv. Morex reference genome (Mascher et al., 2017). Full-length cDNAs from the barley variety Haruna nijo (Matsumoto et al., 2011) determined transcript coverage, and high-resolution RT-PCR validated alternatively spliced (AS) transcripts of 86 genes in five different organs and tissue. These methods were used as benchmarks to select an optimal barley RTD. BaRTv1.0-Quantification of Alternatively Spliced Isoforms (QUASI) was also made to overcome inaccurate quantification due to variation in 5’ and 3’ UTR ends of transcripts. BaRTv1.0-QUASI was used for accurate transcript quantification of RNA-seq data of five barley organs/tissues. This analysis identified 20,972 significant differentially expressed genes, 2,791 differentially alternatively spliced genes and 2,768 transcripts with differential transcript usage.</p> <p>Conclusion<br> A high confidence barley reference transcript dataset consisting of 60,444 genes with 177,240 transcripts has been generated. Compared to current barley transcripts, BaRTv1.0 transcripts are generally longer, have less fragmentation and improved gene models that are well supported by splice junction reads. Precise transcript quantification using BaRTv1.0 allows routine analysis of gene expression and AS.</p>
Supporting transcriptomic data and code for: "Rapid and dose-dependent Natural Killer (NK) cell modulation and cytokine correlations after human rVSV-ZEBOV Ebolavirus vaccination"
<p>The counts_NK.csv file contains gene expression data (counts) for genes in the Ion Ampliseq human Gene expression kit panel. Data were obtained from whole blood RNA. Subjects were vaccinated with a high dose of the rVSV-ZEBOV vaccine against Ebola virus disease in the Geneva clinical trial.</p> <p>The Descriptive_Table_NK_2.csv contains descriptive data of the subjects for differential expression analysis.</p> <p>The NK_analysis_code.R file contains the code used for analysis.</p>
RDF version of the data from Saarimaki et al. Manually curated transcriptomics data collection for toxicogenomic assessment of engineered nanomaterials (Version 1.0.0) [Zenodo Dataset] (2020)
<p>This is an RDFied version of the dataset published by Saarimaki et al. Manually curated transcriptomics data collection for toxicogenomic assessment of engineered nanomaterials (Version 1.0.0) [Zebodo Dataset] (2020)</p> <p>The original dataset publication DOI: <a href="http://doi.org/10.5281/zenodo.4146981">http://doi.org/10.5281/zenodo.4146981</a></p> <p>The Original publication authors: Saarimaki, Laura Aliisa, Federico, Antonio, Lynch, Iseult, Papadiamantis, Anastasios G., Tsoumanis, Andreas, Melagraki, Georgia, Afantitis, Antreas, Serra, Angela, & Greco, Dario</p>
GWAS to single cell: Intersecting single-cell transcriptomics and genome wide association studies identifies crucial cell-populations and candidate genes for atherosclerosis.
<p><strong>Background</strong></p> <p>Genome-wide association studies (GWAS) have discovered hundreds of common genetic variants for atherosclerotic disease and cardiovascular risk factors. The translation of susceptibility loci into biological mechanisms and targets for drug discovery remains challenging. Intersecting genetic and gene expression data has led to identification of candidate genes. However, the assayed tissues are often non-diseased and heterogeneous in cell composition confounding the candidate prioritization. We collected single-cell transcriptomics (scRNA-seq) from atherosclerotic plaques and aimed to identify cell-type-specific expression of disease-associated genes. </p> <p> </p> <p><strong>Methods and Results</strong></p> <p>To identify disease-associated candidate genes, we applied gene-based analyses using GWAS summary statistics from 46 atherosclerotic, cardiometabolic, and other traits. Next we intersected these candidates with single-cell transcriptomics (scRNA-seq) to identify those genes that are specifically expressed in individual cell (sub)populations of atherosclerotic plaques. We derive an enrichment score and show that loci that associated with coronary artery disease demonstrated a prominent substrate in plaque smooth muscle cells (<em>SKI</em>, <em>KANK2</em>, <em>SORT1</em>), endothelial cells (<em>SLC44A1</em>, <em>ATP2B1</em>), and macrophages (<em>APOE</em>, <em>HNRNPUL1</em>). Further sub clustering of SMC-subtypes revealed genes in risk loci for coronary calcification specifically enriched in a synthetic cluster of SMCs. To verify the robustness of our approach, we used liver-derived scRNAseq-data and showed enrichment of circulating lipids-associated loci in hepatocytes.</p> <p><br> <strong>Conclusion</strong></p> <p>We confirm known gene-cell pairs relevant for atherosclerotic disease, and discovered novel pairs pointing to new biological mechanisms amenable for therapy. We present an intuitive single-cell transcriptomics driven workflow rooted in human large-scale genetic studies to identify putative candidate genes and affected cells associated with cardiovascular traits.</p> <p> </p>
Processed data used in transcriptome- metabolome-wide association study
<p>Processed_RNASeq_RPKM.txt contains RPKM levels for 45484 genes quantified in 555 individuals from RNA-Seq of lymphoblastoid cell lines (LCLs).</p> <p>Processed_NMRpeaks_baseline.txt contains binned, normalized and standardised (z-scored) NMR peak intensities for 1276 bins quantified in 555 individuals from urine samples taken at baseline. NMR spectra were acquired at 300 K on a Bruker 16.4 T Avance II 700 MHz NMR spectrometer (Bruker Biospin, Rheinstetten, Germany) using a standard 1H detection pulse sequence with water suppression. The spectra were referenced to the TSP signal and phase and baseline corrected.</p> <p>Processed_NMRpeaks_followup.txt contains binned, normalized and standardised (z-scored) NMR peak intensities for 1289 bins quantified in 315 individuals from urine samples taken during follow-up. NMR spectra were acquired with an Avance III HD 600 NMR spectrometer. Spectra were referenced to the TSP signal and phase and baseline corrected.</p> <p>More details on the data set can be found in Sönmez Flitman et al. (doi: https://doi.org/10.1101/2020.05.22.110197).</p>
Assembled transcriptomes of ovary, testis, and brain (male and female) of Amphibolurus muricatus (jacky dragon) generated using Trinity v2.11.0
<p><strong><em>A. muricatus</em> transcriptome assemblies generated using Trinity v2.11.0 (Haas et al. 2013; Grabherr et al. 2011; Henschel et al. 2012)</strong><br> • Amphibolurus-muricatus_brain.fa.tar.gz: Combined Trinity assembly of <em>A. muricatus</em> brain (male and female).<br> • Amphibolurus-muricatus_combined.fa.tar.gz: Combined Trinity assembly of <em>A. muricatus</em> ovary, testis, and brain (male and female).<br> • Amphibolurus-muricatus_female_brain.fa.tar.gz: Trinity assembly of female <em>A. muricatus</em> brain.<br> • Amphibolurus-muricatus_male_brain.fa.tar.gz: Trinity assembly of male <em>A. muricatus</em> brain.<br> • Amphibolurus-muricatus_ovary.fa.tar.gz: Trinity assembly of <em>A. muricatus</em> ovary.<br> • Amphibolurus-muricatus_testis.fa.tar.gz: Trinity assembly of <em>A. muricatus</em> testis.</p> <p> </p> <p><strong>References</strong></p> <ul> <li>Grabherr, M.G., B.J. Haas, M. Yassour, J.Z. Levin, D.A. Thompson et al., 2011 Full-length transcriptome assembly from RNA-Seq data without a reference genome. Nat Biotechnol 29 (7):644-652.</li> <li>Haas, B.J., A. Papanicolaou, M. Yassour, M. Grabherr, P.D. Blood et al., 2013 De novo transcript sequence reconstruction from RNA-seq using the Trinity platform for reference generation and analysis. Nat Protoc 8 (8):1494-1512.</li> <li>Henschel, R., M. Lieber, L.-S. Wu, P.M. Nista, B.J. Haas et al., 2012 Trinity RNA-Seq assembler performance optimization, pp. 45 in Proceedings of the 1st Conference of the Extreme Science and Engineering Discovery Environment: Bridging from the eXtreme to the campus and beyond. Association for Computing Machinery, Chicag, IL, USA.</li> </ul> <p> </p>
Data from: Transcriptomic meta-analysis reveals unannotated long non-coding RNAs related to the immune response in sheep
<p>This dataset contains additional files from the manuscript: "Transcriptomic meta-analysis reveals unannotated long non-coding RNAs related to the immune response in sheep".</p> <p>The files included are:</p> <p>- All novel lncRNA transcript annotation GTF file ( lncrnas.gtf )</p> <p>- High-confidence lncRNA gene annotation GTF file ( lncrnas_evidence.gtf )</p> <p>- All novel lncRNA transcript annotation GTF file remapped to the ARS-UI_Ramb_v2.0 genome ( lncrnas_remapped_v2.gtf )</p> <p>- Raw count estimates of the extended annotation ( rawcounts.csv )</p> <p>- TPM values of the extended annotation ( tpmcounts.csv )</p> <p>- Supplementary data to the published article (.xlsx, .pdf)</p> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.