Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
25,372
datasets available to search
ShareScore release 0.7.1
Dataset results
25,372 results for “Transcriptomics”
Assemblies and annotation products of a transcriptome of regenerating and non-regenerating Lumbriculus variegatus worms
<p>These are the assemblies and annotation products of a transcriptome assembled from regenerating and non-regenerating tissues of the California blackworm <em>Lumbriculus variegatus</em>. The assembly and annotation strategies are described in the article <strong>Transcriptome analysis during early regeneration of <em>Lumbriculus variegatus</em></strong>, published in Gene Reports (https://doi.org/10.1016/j.genrep.2021.101050).</p> <p>blastp.outfmt6: homology results from BLASTp (v2.10.0) of the predicted proteins</p> <p>blastx.outfmt6: homology results from BLASTx (v2.10.0) of the assembled transcripts</p> <p>LV_transcriptome.fasta: assembled transcriptome after duplicated sequences were filtered out with CD-HIT-EST (v4.7)</p> <p>LV_transcriptome.fasta.transdecoder.cds: coding sequences identified with TransDecoder (v5.5.0)</p> <p>LV_transcriptome.fasta.transdecoder.gff3: positional annotation of the ORFs identified with TransDecoder (v5.5.0)</p> <p>LV_transcriptome.fasta.transdecoder.pep: predicted proteins identified with TransDecoder (v5.5.0)</p> <p>LV_transcriptome_non_filtered.fasta: transcriptome assembled with Trinity (Galaxy v0.0.1)</p> <p>signalp.out: signal peptide predictions from signalP (v4.1)</p> <p>tmhmm.out: transmembrane domains predictions from tmHMM (v2.0)</p> <p>TrinotatePFAM.out: protein domains predictions from HMMER (v3.3)</p>
Transcriptomes for Ranitomeya imitator and R. variabilis: Evidence for a Parabasalian Gut Symbiote in Egg-Feeding Poison Frog Tadpoles in Peru
<p>This dataset contains the assembled transcriptomes for our paper. The three assemblies are for <em>Ranitomeya imitator, R. variabilis, </em>and a merged assembly of the two species. For methodological details, see the published manuscript.</p>
Spatiotemporally resolved transcriptomics reveals subcellular RNA kinetic landscape
<p>Spatiotemporal regulation of the cellular transcriptome is crucial for proper protein expression and cellular function. However, the intricate subcellular dynamics of RNA synthesis, decay, export, and translocation remain obscured due to the limitations of existing transcriptomics methods Here, we report a spatiotemporally resolved RNA mapping method (TEMPOmap) to uncover subcellular RNA profiles across time and space at the single-cell level in heterogeneous cell populations. TEMPOmap integrates pulse-chase metabolic labeling of the transcriptome with highly multiplexed three-dimensional (3D) in situ sequencing to simultaneously profile the age and location of individual RNA molecules. Using TEMPOmap, we constructed the subcellular RNA kinetic landscape of 991 genes in human HeLa cells from upstream transcription to downstream subcellular translocation. Clustering analysis of critical RNA kinetic parameters across single cells revealed kinetic gene clusters whose expression patterns were shaped by multistep kinetic sculpting. Importantly, these kinetic gene clusters are functionally segregated, suggesting that subcellular RNA kinetics are differentially regulated to serve molecular and cellular functions in a cell-cycle-dependent manner. We further demonstrated that functionally segregated RNA kinetics could be seen in heterogeneous human primary cell cultures, revealing cell-type-dependent RNA dynamic regulation. Together, these single-cell spatiotemporally resolved transcriptomics measurements provide us the gateway to uncovering new gene regulation principles and understanding how kinetic strategies enable precise RNA expression in time and space.</p> <p>Please use the most recent version of the dataset.</p>
Multivariate adaptive shrinkage improves cross-population transcriptome prediction and association studies in underrepresented populations
<p>This Zenodo file collection contains transcriptome prediction models built for PrediXcan, as well as concatenated raw results from S-PrediXcan. Files belong to the research paper "Multivariate adaptive shrinkage improves cross-population transcriptome prediction and association studies in underrepresented populations" published at HGG Advances (2023, doi: 10.1016/j.xhgg.2023.100216). Please refer to the README file for a more detailed description.</p>
Processed data and scripts supporting the manuscript "Single-cell transcriptomics reveals immune suppression and cell states predictive of patient outcomes in rhabdomyosarcoma"
<p>This submission contains the compiled count table, processed R objects and various scripts and output files accompanying our manuscript "Single-cell transcriptomics reveals immune suppression and cell states predictive of patient outcomes in rhabdomyosarcoma" (Nature Communications, 2023, https://doi.org/10.1038/s41467-023-38886-8)</p>
Supporting data for "Dissecting the cellular architecture of neuroblastoma bone marrow metastasis using single-cell transcriptomics and epigenomics unravels the role of monocytes at the metastatic niche"
<p>This data repository contains several datasets supplementing the paper “Dissecting the cellular architecture of neuroblastoma bone marrow metastasis using single-cell transcriptomics and epigenomics unravels the role of monocytes at the metastatic niche” by Fetahu, Esser-Skala, Dnyansagar et al. (2023).</p> <ul> <li>HOMER_Results.zip: detailed results of the HOMER analysis</li> <li>nblast_scopen_gene_activity_normalized_motifs_added.rds: Seurat object with scATAC-seq data</li> <li>snp_array.tgz: SNP array data</li> <li>R_data_generated.tgz: Files generated by the scRNA-seq analysis scripts in the GitHub repository associated with the publication.</li> </ul>
Comparative Analysis of Droplet- vs. Microwell-based Whole Transcriptome Single-Cell Sequencing Technologies in Complex Human Tissues
<p>In the past decade, high-dimensional single-cell omics tools have enabled scientists to study the tumor microenvironment (TME) in unprecedented detail. However, recent investigations suggest that each technique has its unique strengths but also technology-inherent limitations. Here we directly compared two commercially available high-throughput single-cell RNA sequencing (scRNA-seq) technologies - droplet-based 10X Chromium <em>vs.</em> microwell-based BD Rhapsody - using paired samples from patients with localized prostate cancer (PCa) undergoing a radical prostatectomy.</p> <p>Although high technical consistency was observed in unraveling the whole transcriptome, the relative abundance of detectable cell populations differed. This could in part be ascribed to differences in the performance to recover cells with low-mRNA content. Hence, immune cells such as neutrophils are underrepresented in data generated with the widely used droplet-based scRNA-seq protocol, highlighting the importance of considering platform limitations in low mRNA content cell recovery. In contrast, droplet-based scRNA-seq demonstrated superiority in terms of recovering cells of epithelial origin. Moreover, we discovered platform-dependent variabilities in mRNA quantification and cell-type marker annotation, affecting the composition of identified tissue profiles and the exploratory value of the generated datasets. Overall, our study emphasizes the importance of carefully selecting the appropriate scRNA-seq platform to improve cell type representation and obtain a more comprehensive and accurate understanding of the TME.</p>
Data from: A transcriptome-based phylogeny of Scarabaeoidea confirms the sister group relationship of dung beetles and phytophagous pleurostict scarabs (Coleoptera)
<p><span>Scarab beetles (Scarabaeidae) are a diverse and ecologically important group of angiosperm-associated insects. As conventionally understood, scarab beetles comprise two major lineages: dung beetles and the phytophagous Pleurosticti. However, previous phylogenetic analyses have not been able to convincingly answer the question whether or not the two lineages form a monophyletic group. Here we report our results from phylogenetic analyses of more than 4,000 genes mined from transcriptomes of more than 50 species of Scarabaeidae and non-scarabaeid Scarabaeoidea. Our results provide convincing support for the monophyly of Scarabaeidae, confirming the debated sister group relationship of dung beetles and phytophagous pleurostict scarabs. Supermatrix-based maximum likelihood and multispecies coalescent phylogenetic analyses strongly imply the subfamily Melolonthinae as currently understood being paraphyletic. We consequently suggest various changes in the systematics of Melolonthinae: Sericinae Kirby, 1837 stat. rest. and sensu n. to include the tribes Sericini, Ablaberini and Diphucephalini, and Sericoidinae Erichson, 1847 stat. rest. and sensu n. to include the tribes </span><span>Automoliini, Heteronychini, Liparetrini, Maechidiini, Scitalini, Sericoidini, and Phyllotocini. Both subfamilies appear to consistently form a monophyletic sister group to all remaining subfamilies so far included within pleurostict scarabs except Orphninae. Our results represent a major step towards understanding the diversification history of one of the largest angiosperm-associated radiations of beetles.</span></p>
STGMVA: clustering, imputation, and integration for spatial resolved transcriptomics using spatiotemporal gaussian mixture variational autoencoder
<p> In this study, we present STGMVA, a comprehensive analysis toolkit employs a spatiotemporal gaussian mixture variational autoencoder to tackle these tasks effectively. STGMVA consists of two stages: pretraining the gene expression and spatial location using a gaussian mixture model, and learning the embedding vectors through a variational graph autoencoder. Results demonstrate STGMVA surpasses state-of-the-art approaches on various spatial transcriptomics datasets, exhibiting superior performance across different scales and resolutions. Notably, STGMVA achieves the highest clustering accuracy in human brain, mouse hippocampus, and mouse olfactory bulb tissues. Furthermore, STGMVA enhances and denoises gene expression patterns for gene imputation task. Additionally, STGMVA has the capability to correct batch effects and achieve joint analysis when integrating multiple tissue slices.</p>
Reliable imputation of spatial transcriptome with uncertainty estimation and spatial regularization
<p>Imputation of missing features in spatial transcriptomics is urgently demanded due to technology limitations, while most existing computational methods suffer from moderate accuracy and cannot estimate the reliability of the imputation. <br> To fill the research gaps, we introduce a computational model, TransImp, that imputes the missing feature modality in spatial transcriptomics by mapping it from single-cell reference. Uniquely, we derived a set of attributes that can accurately predict imputation uncertainty, hence enabling us to select reliably imputed genes. Also, we introduced a spatial auto-correlation metric as a regularization to avoid overestimating spatial patterns. Multiple datasets from various platforms have demonstrated that our approach significantly improves the reliability of downstream analyses in detecting spatial variable genes and interacting ligand-receptor pairs. Therefore, TransImp offers a way towards a reliable spatial analysis of missing features for both matched and unseen modalities, e.g., nascent RNAs.</p>
Context-dependent perturbations in chromatin folding and the transcriptome by cohesin and related factors
<p>Cohesin plays vital roles in chromatin folding and gene expression regulation, cooperating with such factors as cohesin loaders, unloaders, and the insulation factor CTCF. Although models of regulation have been proposed (e.g., loop extrusion), how cohesin and related factors collectively or individually regulate the hierarchical chromatin structure and gene expression remains unclear. We have depleted cohesin and related factors and then conducted a comprehensive evaluation of the resulting 3D genome, transcriptome and epigenome data. We observed substantial variation in depletion effects among factors at topologically associating domain (TAD) boundaries and on interTAD interactions, which were related to epigenomic status. Gene expression changes were highly correlated with direct cohesin binding and gain of TAD boundaries than with the loss of boundaries. Moreover, cohesin was broadly enriched in active compartment A chromosomes, which were retained after CTCF depletion. Our results demonstrate context-specific roles of cohesin for gene expression and chromatin folding.</p>
Transcriptome analysis of anuran breeding glands reveals a surprisingly high expression and diversity of NNMT-like genes
<p><strong>Abstract</strong></p> <p>In many amphibians, males have sexually dimorphic breeding glands, which can produce proteinaceous or volatile pheromones, used for intraspecific communication. In this study we analyse two types of glands in the Mexican treefrog species <em>Ptychohyla macrotympanum </em>(Hylidae) – large ventrolateral glands and small nuptial pads on their fingers – using histology, whole-transcriptome sequencing and phylogenetic analyses. We found strong differences in glandular tissue composition and gene expression patterns between the two breeding gland types. In both glands we only found low expression of protein pheromone candidates. Instead, in the ventrolateral glands, gene expression was strikingly dominated by nicotinamide N-methyltransferase (NNMT)-like genes. Diversity of these genes was remarkably high, with at least 68 distinct NNMT-like genes. Our phylogenetic comparative analysis of the diversity of NNMT-like genes across vertebrates indicates that the extreme diversity of this gene is largely a frog-specific phenomenon and can be traced to large numbers of relatively recent gene duplications occurring independently in many lineages. The strong dominance and astonishing diversity of NNMT-like genes found in anurans in general, and in their sexually dimorphic breeding glands specifically, suggests an important function of NNMT-like proteins for anuran reproduction, possibly being related to volatile pheromone production.In many amphibians, males have sexually dimorphic breeding glands, which can produce proteinaceous or volatile pheromones, used for intraspecific communication. In this study we analyse two types of glands in the Mexican treefrog species <em>Ptychohyla macrotympanum </em>(Hylidae) – large ventrolateral glands and small nuptial pads on their fingers – using histology, whole-transcriptome sequencing and phylogenetic analyses. We found strong differences in glandular tissue composition and gene expression patterns between the two breeding gland types. In both glands we only found low expression of protein pheromone candidates. Instead, in the ventrolateral glands, gene expression was strikingly dominated by nicotinamide N-methyltransferase (NNMT)-like genes. Diversity of these genes was remarkably high, with at least 68 distinct NNMT-like genes. Our phylogenetic comparative analysis of the diversity of NNMT-like genes across vertebrates indicates that the extreme diversity of this gene is largely a frog-specific phenomenon and can be traced to large numbers of relatively recent gene duplications occurring independently in many lineages. The strong dominance and astonishing diversity of NNMT-like genes found in anurans in general, and in their sexually dimorphic breeding glands specifically, suggests an important function of NNMT-like proteins for anuran reproduction, possibly being related to volatile pheromone production.</p> <p> </p> <p><strong>Supplementary datasets accompanying the paper:</strong></p> <p>- final RNAseq assemblies of the ventrolateral glands and the nuptial pads of <em>Ptychohyla macrotympanum</em><br> - fasta-file of all <em>Ptychohyla</em>-NNMT-like genes found in this study</p>
Transcriptomic analysis of light-induced genes in Nasonia vitripennis: possible implications for circadian light entrainment pathways
<p class="MDPI17abstract"><span>Circadian entrainment to the environmental day-night cycle is essential for the optimal use of environmental resources. In insects, opsin-based photoreception in the compound eye and ocelli, and CRYPTOCHROME1 (CRY1) in circadian clock neurons are thought to be involved in sensing photic information, but genetic regulation of circadian light entrainment in species without light-sensitive CRY1 remains unclear. To elucidate a possible CRY1-independent light transduction cascade, we analysed light-induced gene expression through RNA-sequencing in <em>Nasonia vitripennis</em>. Entrained wasps were subjected to a light pulse in the subjective night to reset the circadian clock and light-induced changes in gene expression were characterized at four different time points in wasp heads. We used co-expression, functional annotation, and transcription factor binding motif analyses to gain insight into the molecular pathways in response to acute light stimulus and form a hypothesis about the circadian light resetting pathway. Maximal gene induction was found after 2h of light stimulation (1432 genes), including the opsin <em>opblue</em> and the core clock genes <em>cry2</em> and <em>npas2</em>. Pathway and cluster analyses revealed light activation of glutamatergic and GABA-ergic neurotransmission, including CREB and AP-1 transcription pathway signalling. This suggests that circadian photic entrainment in <em>Nasonia</em> may require pathways that are similar to mammals. We propose a model for hymenopteran circadian light resetting that involves opsin-based photoreception, glutamatergic neurotransmission, and gene induction of <em>cry2</em> and <em>npas2</em> to reset the circadian clock.</span></p>
Transcriptome-wide meta-analysis of codon usage in Escherichia coli
<p>Data generated by the CUBseq pipeline on Escherichia coli RNA-seq data.</p>
Transcriptome Analysis of Cisplatin, Cannabidiol, and Intermittent Serum Starvation Alone and in Various Combinations on Colorectal Cancer Cells
<p>* See README file for the description of data files available in this repository</p> <p>1. Study Description:</p> <p>Platinum-derived chemotherapy medications are often combined with other conventional therapies for treating different tumours, including colorectal cancer. However, the development of drug resistance and multiple adverse effects remain common in clinical settings. Thus, there is a necessity to find novel treatments and drug combinations that could effectively target colorectal cancer cells and lower the probability of disease relapse. To find potential synergistic interaction, we designed multiple different combinations between cisplatin, cannabidiol, and intermittent serum starvation on colorectal cancer cell lines. Based on the cell viability assay, we found that combinations between cannabidiol and intermittent serum starvation, cisplatin, and intermittent serum starvation, as well as cisplatin, cannabidiol and intermittent serum starvation can work in a synergistic fashion on different colorectal cancer cell lines. Furthermore, we analyzed differentially expressed genes and affected pathways in colorectal cancer cell lines to understand further the potential molecular mechanisms behind the treatments and their interactions. We found that synergistic interaction between cannabidiol and intermittent serum starvation can be related to changes in the transcription of genes responsible for cell metabolism and cancer’s stress pathways. Moreover, when we added cisplatin to the treatments, there was a strong enrichment of genes taking part in G2/M cell cycle arrest and apoptosis.</p> <p> </p> <p>2. Bioinformatics workflow:</p> <p>Initial quality control was conducted using FastQC v0.11.9 https://www.bioinformatics.babraham.ac.uk/projects/fastqc/. Sequencing reads were trimmed of adapter sequences and low-quality bases using Trimmomatic. Trimmed sequence files were examined with FastQC to verify the trimming results. Trimmed sequencing reads were mapped to Human genome (GRCh37, Ensembl) downloaded from Illumina iGenome website (<a href="https://support.illumina.com/sequencing/sequencing_software/igenome.html">https://support.illumina.com/sequencing/sequencing_software/igenome.html</a>). Mapping was done using splice aware aligner HISAT2 2.1.0. Alignment files in SAM format were converted to BAM, sorted and indexed with samtools v.1.3.1. Mapping quality and statistics were collected with QualiMap software package v.2.2.2 <a href="http://qualimap.conesalab.org/">http://qualimap.conesalab.org/</a> The counts if reads mapping to features (genes) were counted using FeatureCounts v.2.0.1 software.</p> <p>Data exploration, visualization and statistical comparisons were conducted using R language version 4.2.2. Pair-wise comparisons between experimental groups were done with DESeq2 v.2.1.36 as described in the package manual. To decrease computational time, only the genes with at least 5 reads across 3 samples were kept in the analysis. In addition to hard threshold filtering mentioned above, DESeq2 implements independent filtering based on mean of normalized count as a filter statistic.</p> <p>We used hierarchical clustering (HC) and principal components analysis (PCA) to investigate the relationship between samples and detect potential outliers. Prior to HC and PCA analysis, DESeq2 normalized values underwent variance stabilizing transformation with using vst() function from DESeq2. HC was done using hclust() function implemented in R, with the clustering method set as “complete” for the matrices of sample-to-sample distances, and “Ward.D2” in case of the sample and gene clustering based on top 500 most variable genes. The distance measure in HC analysis was set to “euclidean”. Principal components analysis (PCA), applied to top 500 highly variable genes, was conducted using prcomp() function implemented in R with default options.</p> <p>Differentially expressed genes (DEGs) were detected with DESeq2 function results() with default options. DESeq2 uses Wald test to determine significantly changed genes between groups. The independent filtering option was set to TRUE with alpha threshold (adjusted p-value) kept at 0.1. Multiple comparison adjustment was done using Bejamini-Hochberg procedure.</p> <p> </p> <p> </p> <p> </p> <p> </p> <p> </p>
Thesis: Transcriptome analysis of insecticide resistant Drosophila suzukii
Open the record for dataset details and reuse information.
Supplementary data for: Transcriptomics of mosaic brain differentiation underlying complex division of labor in a social insect
Open the record for dataset details and reuse information.
The cacao gene atlas: A transcriptome developmental atlas reveals highly tissue-specific and dynamically-regulated gene networks in Theobroma cacao L
Open the record for dataset details and reuse information.
Wound image and transcriptome datasets of swine acute wounds
Open the record for dataset details and reuse information.
Reproducible data, example subsets, and analysis pipeline for the extended TAaCGH study of breast cancer genomic and transcriptomic profiles
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.