Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
25,372
datasets available to search
ShareScore release 0.7.1
Dataset results
25,372 results for “Transcriptomics”
A transcriptome for the early-branching fern Botrychium lunaria enables fine-grained resolution of population structure
<p>File S1: Alignment of Botrychium CRY2cA sequences used to infer the genus-level phylogeny.</p> <p>File S2: Peptide sequences used to infer the orthogroups named according species names or identifiers (see Table S1).</p> <p>File S3: Alignments of the orthogroup sequences subset used to infer the phylogenomy.</p> <p>File S4: Output files from modeltest-ng named by orthogroup names.</p> <p>File S5: Output files from raxml-ng named by orthogroup names.</p>
PacBio IsoSeq reference transcriptomes for Pinus taeda L.
<p>Fusiform rust disease, caused by the endemic fungus <i>Cronartium quercuum</i> f. sp. <i>fusiforme</i>, is the most damaging disease affecting economically important pine species in the southeast United States. In this report, we detail the genomic localization and sequence-level discovery of candidate race-nonspecific broad-spectrum fusiform rust resistance genes in <i>Pinus taeda </i>L. Two full-sib families, each with ~1000 progeny, were challenged with a complex inoculum consisting of over 150 pathogen isolates. High-density linkage mapping revealed three QTL distributed on two linkage groups. The two QTL on linkage group 2 were additive with respect to their effects on the probability of disease outcome. All three QTL were validated using a population of 2057 cloned pine genotypes in a six-year-old multi-environmental field trial. As a complement to the QTL mapping approach, bulked segregant RNAseq analysis revealed a small number of candidate nucleotide binding leucine rich repeat genes harboring SNP significantly associated with disease resistance. The results of this study demonstrate that single qualitative resistance genes can confer effective resistance against genetically diverse mixtures of an endemic pathogen.</p>
Error, noise and bias in de novo transcriptome assemblies
<p><i>De novo</i> transcriptome assembly is a powerful tool, widely used over the last decade for making evolutionary inferences. However, it relies on two implicit assumptions: that the assembled transcriptome is an unbiased representation of the underlying expressed transcriptome, and that expression estimates from the assembly are good, if noisy approximations of the relative abundance of expressed transcripts. Using publicly available data for model organisms, we demonstrate that, across assembly algorithms and data sets, these assumptions are consistently violated. Bias exists at the nucleotide level, with genotyping error rates ranging from 30-83%. As a result, diversity is underestimated in transcriptome assemblies, with consistent under-estimation of heterozygosity in all but the most inbred samples. Even at the gene level, expression estimates show wide deviations from map-to-reference estimates, and positive bias at lower expression levels. Standard filtering of transcriptome assemblies improves the robustness of gene expression estimates but leads to the loss of a meaningful number of protein-coding genes, including many that are highly expressed. We demonstrate a computational method, length-rescaled CPM, to partly alleviate noise and bias in expression estimates. Researchers should consider ways to minimize the impact of bias in transcriptome assemblies.</p>
RNASeq fastq files associated with the manuscript entitled 'Interspecies transcriptome analyses identify genes that control the development and evolution of limb skeletal proportion'
<p>This next-generation sequencing dataset is associated with the research manuscript entitled ‘<em>Interspecies transcriptome analyses identify genes that control the development and evolution of limb skeletal proportion</em>’ (https://www.biorxiv.org/content/10.1101/754002v2).</p> <p>The folder ‘<strong>Zenodo_Saxena_etal_2021_RNASeq_FastqFiles</strong>’ contains raw/unprocessed RNASeq Fatsq read files for postnatal day 5 (P5) mouse (Mus) and jerboa (Jac) cartilage samples (Metatarsal = MT; Radius/Ulna = RU).</p> <p>> The <strong>Jac_P5</strong> subfolder contains single-end reads (R1) for five jerboa metatarsals (MT1-5) and radius/ulna (RU1-5) biological replicates. Jac_MT1-3 and Jac_RU1-3 were used in the primary differential expression analysis (n=3). Jac_MT4-5 and Jac_RU4-5 were used for independent validation (n=2) of the the primary analysis. </p> <p>> The <strong>Mus_P5</strong> subfolder contains single-end reads (R1) for five mouse metatarsals (MT1-5) and radius/ulna (RU1-5) biological replicates. Mus_MT1-3 and Mus_RU1,3 & 4 were used in the primary differential expression analysis (n=3). Mus_MT4-5 and Mus_RU4 & 5 were used for independent validation (n=2) of the the primary analysis. </p>
Differential transcriptomic regulation in sweet orange fruit (Citrus sinensis L. Osbeck) following dehydration and rehydration conditions leading to peel damage
<p>Water stress is the most important environmental agent causing crop productivity and quality losses globally. In citrus, water stress is a main driver of fruit peel disorders that impact quality and marketability. An increasingly present postharvest peel disorder is non-chilling peel pitting (NCPP). NCPP is manifested as collapsed areas of flavedo randomly scattered on the fruit and its incidence increases due to abrupt increases in environmental relative humidity (RH) during postharvest fruit manipulation. In this work we have used a custom-made cDNA microarray containing 44k unigenes from <em>Citrus sinensis</em> (L. Osbeck) covering for the first time the whole genome from this species, to study transcriptomic responses of mature citrus fruit to water stress. We have compared global gene expression profiles of flavedo from Navelate oranges subjected to severe water stress with those of fruit subjected to rehydration stress provoked by changes in RH during postharvest, which enhances the development of NCPP. Our results show that NCPP is a complex physiological process that shares molecular responses with those from prolonged dehydration in fruit, but the damage associated to NCPP may be explained by unique features of rehydration stress at the molecular level, including membrane disorganization, cell wall modification, and proteolysis.</p>
Neonatal exposure to BPA, BDE-99, and PCB produces persistent changes in hepatic transcriptome associated with gut dysbiosis in adult mouse livers
<p class="western"><span><span><span><b>Background</b>. Recent evidence suggests that multigenic and complex environmentally modulated diseases result from early life exposure to toxicants at least partly via gut microbial influences. Environmental toxicants, polybrominated diphenyl ethers (PBDEs), and polychlorinated biphenyls (PCBs) are breast milk-enriched persistent organic pollutants (POPs) and thus remain a continuing threat to human health despite being banned from production. Recent findings focused on the liver developmental reprogramming capabilities from neonatal BPA exposure; however, little is known on how PBDEs and PCBs regulate the liver transcriptome with respect to the gut microbiome. </span></span></span></p> <p class="western"><span><span><span><b>Objectives</b>. We investigated whether the gut microbiome can be persistently reprogrammed with the liver following neonatal exposure to POPs, and whether microbial biomarkers associated with disease-prone changes in the hepatic epigenetic and transcriptomic landscape in adulthood. </span></span></span></p> <p class="western"><span><span><span><b>Methods. </b>C57BL/6 male and female mouse pups were orally administered vehicle, bisphenol A (BPA), BDE-99 (a breast milk-enriched PBDE congener), or the Fox River PCB mixture (an environmentally relevant PCB mixture), between postnatal day (PND) 2 to 4, once daily for three consecutive days. Tissues were collected at PND5 and PND60 for 16S rDNA sequencing and targeted metabolomics.</span></span></span></p> <p class="western"><span><span><span><b>Results</b>. Neonatal exposure to BDE-99, followed by BPA and PCB, produced the greatest persistent changes in the adult hepatic transcriptome, including an inflammation and cancer-prone transcriptomic signature. BDE-99 exposure resulted in a persistent increase in <i>Akkermansia muciniphila</i> throughout the intestinal sections and feces. We observed persistent increases in acetate and succinate, metabolites <i>A. muciniphila</i><span> is able to produce.</span> Correspondingly, liver H3K4me1 and H3K27 acetylation were enriched around the loci encoding liver cancer-related genes following neonatal BDE-99 exposure. </span></span></span></p> <p class="western"><span><span><span><b>Conclusion. </b><span>Similar to BPA, early life exposure to BDE-99 also produced a cancer-prone hepatic transcriptomic signature corresponding to an increase in permissive epigenetic signatures around cancer-related genes in adulthood. This positively associates with BDE-99 mediated increase in </span><i><span>A. muciniphila</span></i><span> and its metabolites which are established epigenetic modifiers. </span></span></span></span></p>
The EGFRvIII Transcriptome in glioblastoma - public data repository
<p>Compiled dataset from a large omics EGFRvIII study.</p>
Transcriptome analysis of T47D cells and H2A.J-KO derivatives for the paper entitled: The histone variant H2A.J is enriched in luminal epithelial gland cells
<p>H2A.J is a poorly studied mammalian-specific variant of histone H2A. We used immunohistochemistry to study its localization in various human and mouse tissues. H2A.J showed cell-type specific expression with a striking enrichment in luminal epithelial cells of multiple glands including those of breast, prostate, pancreas, thyroid, stomach, and salivary glands. H2A.J was also highly expressed in many carcinoma cell lines and in particular, those derived from luminal breast and prostate cancer. H2A.J thus appears to be a novel marker for luminal epithelial cancers. Knocking-out the H2AFJ gene in T47D luminal breast cancer cells reduced the expression of several estrogen-responsive genes which may explain its putative tumorigenic role in luminal-B breast cancer.</p>
Combined tumor and immune signals from genomes or transcriptomes predict outcomes of checkpoint inhibition in melanoma
<p>Code and data for manuscript "Combined tumor and immune signals from genomes or transcriptomes predict outcomes of checkpoint inhibition in melanoma"</p>
Dynamic prostate cancer transcriptome analysis delineates the trajectory to disease progression.
<p>This file contains vst-normalized gene expression data along with annotations which can be used to reproduce our findings.</p>
Alternative developmental and transcriptomic responses to host plant water limitation in a butterfly metapopulation
<p>The dataset is from a study examining the effects of host plant water stress on the developmental and transcriptomic responses of its specialist Lepidopteran herbivore. The study combines host plant metabolic profiling with development assays and full-transcriptome sequencing of herbivore larvae. First, we profiled metabolic differences between well-watered and water-limited ribwort plantain (<em>Plantago lanceolata</em>) using proton nuclear magnetic resonance spectroscopy (<sup>1</sup>H-NMR). Second, we tested how performance of developing Glanville fritillary (<em>Melitaea cinxia</em>) larvae was affected by host plant water limitation experienced at different larval developmental stages. Third, we examined larval gene regulatory responses to water limited host plants by sequencing full transcriptomes of 77 female larvae (RNA seq). Finally, to examine intrapopulation variation in the responses of the larvae, we compared the phenotypic and transcriptomic responses across full-sib families originating from different parts of the metapopulation. In this dataset, we provide data for the <em>P. lanceolata</em> metabolite responses to water limitation and developmental responses of the <em>M. cinxia</em> larvae to feeding on water limited <em>P. lanceolata</em>. The transcriptomic data are available from NCBI's Gene Expression Omnibus, with the accession number GSE159376.</p>
Cross-species transcriptomics uncovers genes underlying genetic accommodation of developmental plasticity in spadefoot toads
<p>That hardcoded genomes can manifest as plastic phenotypes responding to environmental perturbations is a fascinating feature of living organisms. How such developmental plasticity is regulated at the molecular level is beginning to be uncovered aided by the development of -omic techniques. Here, we compare the transcriptome-wide responses of two species of spadefoot toads with differing capacity for developmental acceleration of their larvae in the face of a shared environmental risk: pond drying. By comparing gene expression profiles over time and performing cross-species network analyses, we identified orthologues and functional gene pathways whose environmental sensitivity in expression have diverged between species. Genes related to lipid, cholesterol and steroid biosynthesis and metabolism make up most of a module of genes environmentally responsive in one species, but canalized in the other. The evolutionary changes in the regulation of the genes identified through these analyses may have been key in the genetic accommodation of developmental plasticity in this system.</p>
Large-scale integration of single-cell transcriptomic data captures transitional progenitor states in mouse skeletal muscle regeneration
<p>Skeletal muscle repair is driven by the coordinated self-renewal and fusion of myogenic stem and progenitor cells. Single-cell gene expression analyses of myogenesis have been hampered by the poor sampling of rare and transient cell states that are critical for muscle repair, and do not inform the spatial context that is important for myogenic differentiation. Here, we demonstrate how large-scale integration of single-cell and spatial transcriptomic data can overcome these limitations. We created a single-cell transcriptomic dataset of mouse skeletal muscle by integration, consensus annotation, and analysis of 23 newly collected scRNAseq datasets and 88 publicly available single-cell (scRNAseq) and single-nucleus (snRNAseq) RNA-sequencing datasets. The resulting dataset includes more than 365,000 cells and spans a wide range of ages, injury, and repair conditions. Together, these data enabled identification of the predominant cell types in skeletal muscle, and resolved cell subtypes, including endothelial subtypes distinguished by vessel-type of origin, fibro/adipogenic progenitors defined by functional roles, and many distinct immune populations. The representation of different experimental conditions and the depth of transcriptome coverage enabled robust profiling of sparsely expressed genes. We built a densely sampled transcriptomic model of myogenesis, from stem cell quiescence to myofiber maturation and identified rare, transitional states of progenitor commitment and fusion that are poorly represented in individual datasets. We performed spatial RNA sequencing of mouse muscle at three time points after injury and used the integrated dataset as a reference to achieve a high-resolution, local deconvolution of cell subtypes. We also used the integrated dataset to explore ligand-receptor co-expression patterns and identify dynamic cell-cell interactions in muscle injury response. We provide a public web tool to enable interactive exploration and visualization of the data. Our work supports the utility of large-scale integration of single-cell transcriptomic data as a tool for biological discovery.</p>
Ephestia kuehniella transcriptome
<p>This dataset contains:</p> <p>1) Ephestia_kuehniella_transcriptome_initial.fa: the initial <em>de novo</em> assembly of the silk gland-specific transcriptome of <em>Ephestia kuehniella</em>.</p> <p>2) Ephestia_kuehniella_transcriptome_improved.fa: the improved transcriptome by incorporating long-read genomic data.</p> <p>For details please see:</p> <p>Wu BC-H, Šauman I, Maaroufi HO, Žaloudíková A, Žurovcová M, Kludkiewicz B, Hradilová M and Žurovec M (2022), Characterization of silk genes in <em>Ephestia kuehniella</em> and <em>Galleria mellonella</em> revealed duplication of sericin genes and highly divergent sequences encoding fibroin heavy chains. Front. Mol. Biosci. 9:1023381. doi: 10.3389/fmolb.2022.1023381</p>
Functional annotation of the reference transcriptome of Mesodinium rubrum strain JAMR
<p>Raw reads were pre-processed by removing the adaptors and low-quality reads using BBMap. The filtered reads were normalized for depth based on kmer counts using BBNorm function. De novo transcriptomes were generated using both Trinity and velvet-oases. CD-HIT-EST was used to merge the two de novo transcriptomes and reduce the transcript redundancy to 98% similarity and generate unique genes. Transcriptome assembly completeness was evaluated with BUSCO (Benchmarking Universal Single Copy Orthologs) database. Functional annotation was done using blastp function of ncbi-blast using the nr database with evalue 1E-20 and num_alignments 3.</p>
Functional annotation of the reference transcriptome of Dinophysis acuminata strain DAVA01
<p>Raw reads were pre-processed by removing the adaptors and low-quality reads using BBMap. The filtered reads were normalized for depth based on kmer counts using BBNorm function. De novo transcriptomes were generated using both Trinity and velvet-oases. CD-HIT-EST was used to merge the two de novo transcriptomes and reduce the transcript redundancy to 98% similarity and generate unique genes. Transcriptome assembly completeness was evaluated with BUSCO (Benchmarking Universal Single Copy Orthologs) database. Functional annotation was done using blastp function of ncbi-blast using the nr database with evalue 1E-20 and num_alignments 3.</p>
Functional annotation of the reference transcriptome of Mesodinium rubrum strain MBL-DK2009
<p>Raw reads were pre-processed by removing the adaptors and low-quality reads using BBMap. The filtered reads were normalized for depth based on kmer counts using BBNorm function. De novo transcriptomes were generated using both Trinity and velvet-oases. CD-HIT-EST was used to merge the two de novo transcriptomes and reduce the transcript redundancy to 98% similarity and generate unique genes. Transcriptome assembly completeness was evaluated with BUSCO (Benchmarking Universal Single Copy Orthologs) database. Functional annotation was done using blastp function of ncbi-blast using the nr database with evalue 1E-20 and num_alignments 3.</p>
Functional annotation of the reference transcriptome of Teleaulax amphioxeia
<p>Raw reads were pre-processed by removing the adaptors and low-quality reads using BBMap. The filtered reads were normalized for depth based on kmer counts using BBNorm function. De novo transcriptomes were generated using both Trinity and velvet-oases. CD-HIT-EST was used to merge the two de novo transcriptomes and reduce the transcript redundancy to 98% similarity and generate unique genes. Transcriptome assembly completeness was evaluated with BUSCO (Benchmarking Universal Single Copy Orthologs) database. Functional annotation was done using blastp function of ncbi-blast using the nr database with evalue 1E-20 and num_alignments 3.</p>
Functional annotation of the reference transcriptome of Dinophysis ovum strain DoSS3195
<p>Raw reads were pre-processed by removing the adaptors and low-quality reads using BBMap. The filtered reads were normalized for depth based on kmer counts using BBNorm function. De novo transcriptomes were generated using both Trinity and velvet-oases. CD-HIT-EST was used to merge the two de novo transcriptomes and reduce the transcript redundancy to 98% similarity and generate unique genes. Transcriptome assembly completeness was evaluated with BUSCO (Benchmarking Universal Single Copy Orthologs) database. Functional annotation was done using blastp function of ncbi-blast using the nr database with evalue 1E-20 and num_alignments 3.</p>
datset of "Networks and genes modulated by posterior hypothalamic stimulation in patients with aggressive behaviours: Analysis of probabilistic mapping, normative connectomics, and atlas-derived transcriptomics of the largest international multi-centre dataset"
<p>This dataset accompanies the manuscript:<br> "Networks and genes modulated by posterior hypothalamic stimulation in patients with aggressive behaviours: Analysis of probabilistic mapping, normative connectomics, and atlas-derived transcriptomics of the largest international multi-centre dataset."<br> DOI: (https://doi.org/10.1101/2022.10.29.22281666)</p> <p>by</p> <p>Flavia Venetucci Gouveia1,2,3*†,Jürgen Germann4,5†, Gavin JB Elias4,5, Alexandre Boutet4,6, Aaron Loh4,5, Adriana Lucia Lopez Rios7,8, Cristina V Torres Diaz9, William Omar Contreras Lopez10,11, Raquel CR Martinez3,12, Erich T Fonoff13, Juan C Benedetti-Isaac14, Peter Giacobbe 2,15,16, Pablo M Arango Pava17, Han Yan5,18, George M Ibrahim5, 18,19,20, Nir Lipsman2,5,15, Andres M Lozano4,5, Clement Hamani2,5,15*</p> <p>1. Neuroscience and Mental Health, Hospital for Sick Children Research Institute; Toronto, Canada <br> 2. Sunnybrook Research Institute; Toronto, Canada<br> 3. Division of Neuroscience, Sírio-Libanês Hospital; São Paulo, Brazil<br> 4. Division of Neurosurgery, Department of Surgery, University Health Network, Toronto, Canada<br> 5. Division of Neurosurgery, Department of Surgery, University of Toronto; Toronto, Canada<br> 6. Joint Department of Medical Imaging, University of Toronto; Toronto, Canada<br> 7. Department of Functional and Stereotactic Neurosurgery, University Hospital San Vicente Fundación,<br> Medellín, Colombia<br> 8. Department of Functional and Stereotactic Neurosurgery, San Vicente Fundación, Rionegro, Colombia<br> 9. Department of Neurosurgery, University Hospital La Princesa; Madrid, Spain<br> 10. Nemod Research Group, Universidad Autónoma de Bucaramanga; Bucaramanga, Colombia<br> 11. Division of Functional Neurosurgery, Department of Neurosurgery, FOSCAL Clinic; Bucaramanga,<br> Colombia<br> 12. LIM 23, Institute of Psychiatry, School of Medicine, University of São Paulo; São Paulo, Brazil<br> 13. Department of Neurology, Integrated Clinic of Neuroscience, School of Medicine, University of São Paulo;<br> São Paulo, Brazil.<br> 14. Stereotactic and Functional Neurosurgery Division of the International Misericordia Clinic; Barranquilla,<br> Colombia<br> 15. Harquail Centre for Neuromodulation, Sunnybrook Health Sciences Centre; Toronto, Canada<br> 16. Department of Psychiatry, University of Toronto; Toronto, Canada<br> 17. Servicio de Neuocirugia Funcional y Esterotaxia, Clinica Comuneros Bucaramanga, Clinica Desa y Clinica<br> Dime Neurocardiovascular de Cali; Clinica Nueva del Lago, Bogota, Colombia.<br> 18. Division of Neurosurgery, The Hospital for Sick Children; Toronto, Canada<br> 19. Institute of Biomedical Engineering, University of Toronto; Toronto, Canada<br> 20. Institute of Medical Science, University of Toronto; Toronto, Canada<br> † Flavia Venetucci Gouveia and Jürgen Germann contributed equally to this work and share first authorship.</p> <p>* Corresponding Author: Dr. Flavia Venetucci Gouveia. Neuroscience and Mental Health, Hospital for Sick Children Research Institute. 686, Bay Street, Toronto, ON, M5G 0A4, Canada. flavia.venetuccigouveia@sickkids.ca<br> * Corresponding Author: Dr. Clement Hamani. Sunnybrook Research Institute. 2075 Bayview Ave, S126. Toronto, ON, M4N3M5, Canada. clement.hamani@sunnybrook.ca</p> <p>It contains a zip folder ("estimated_binary_Volume_of_Tissue_Activated.zip") with one file (in nii.gz format) per patient estimating the Volume of Activated Tissue for that patient (the estimated 'reach' of the active DBS stimulation) and a demographics file.<br> The case numbers are identical to Table 1 in the manuscript.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.