Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
25,372
datasets available to search
ShareScore release 0.7.1
Dataset results
25,372 results for “Transcriptomics”
Data from: De novo assembly and characterization of the Hucho taimen transcriptome
Taimen (Hucho taimen) is an important ecological and economic species that is classified as vulnerable by the IUCN Red List of Threatened Species; however, limited genomic information is available on this species. RNA-Seq is a useful tool for obtaining genetic information and developing genetic markers for non-model species in addition to its application in gene expression profiling. In this study, we performed a comprehensive RNA-Seq analysis of taimen. We obtained 157 M clean reads (14.7 Gb) and used them to de novo assemble a high-quality transcriptome with a N50 size of 1060 bp. In the assembly, 82% of the transcripts were annotated using several databases, and 14,666 of the transcripts contained a full open reading frame. The assembly covered 75% of the transcripts of Atlantic salmon and 57.3% of the protein-coding genes of rainbow trout. To learn about the genome evolution, we performed a systematic comparative analysis across 11 teleosts including 8 salmonids, and found 313 unique gene families in taimen. Using Atlantic salmon and rainbow trout transcriptomes as the background, we identified 250 positive selection transcripts. The pathway enrichment analysis revealed a unique characteristic of taimen: it possesses more immune-related genes than Atlantic salmon and rainbow trout; moreover, some genes have undergone strong positive selection. We also developed a pipeline for identifying microsatellite marker genotypes in samples, and successfully identified 24 polymorphic microsatellite markers for taimen. These data and tools are useful for studying conservation genetics, phylogenetics, evolution among salmonids and selective breeding for threatened taimen.
Data from: Illuminating the base of the annelid tree using transcriptomics
Annelida is one of three animal groups possessing segmentation and is central in considerations about the evolution of different character traits. It has even been proposed that the bilaterian ancestor resembled an annelid. However, a robust phylogeny of Annelida, especially with respect to the basal relationships, has been lacking. Our study based on transcriptomic data comprising 68,750 – 170,497 amino acid sites from 305 – 622 proteins resolves annelid relationships, including Chaetopteridae, Amphinomidae, Sipuncula, Oweniidae, Magelonidae in the basal part of the tree. Myzostomida, which have been indicated to belong to the basal radiation as well, are now found deeply nested within Annelida as sister group to Errantia in most analyses. Based on our reconstruction of a robust annelid phylogeny, we show that the basal branching taxa include a huge variety of life-styles such as tube-dwelling and deposit-feeding, endobenthic and burrowing, tubicolous and filter-feeding, as well as errant and carnivorous forms. Ancestral character state reconstruction suggests that the ancestral annelid possessed a pair of either sensory or grooved palps, bicellular eyes, biramous parapodia bearing simple chaeta and lacked nuchal organs. Since the oldest fossil of Annelida is reported for Sipuncula (520 Mya), we infer that the early diversification of annelids took place at least in the Lower Cambrian.
Data from: Parasite infection of public databases: a data mining approach to identify apicomplexan contaminations in animal genome and transcriptome assemblies
Background: Contaminations from various exogenous sources are a common problem in next-generation sequencing. Another possible source of contaminating DNA are endogenous parasites. On the one hand, undiscovered contaminations of animal sequence assemblies may lead to erroneous interpretation of data; on the other hand, when identified, parasite-derived sequences may provide a valuable source of information. Results: Here we show that sequences deriving from apicomplexan parasites can be found in many animal genome and transcriptome projects, which in most cases derived from an infection of the sequenced host specimen. The apicomplexan sequences were extracted from the sequence assemblies using a newly developed bioinformatic pipeline (ContamFinder) and tentatively assigned to distinct taxa employing phylogenetic methods. We analysed 920 assemblies and found 20,907 contigs of apicomplexan origin in 51 of the datasets. The contaminating species were identified as members of the apicomplexan taxa Gregarinasina, Coccidia, Piroplasmida, and Haemosporida. For example, in the platypus genome assembly, we found a high number of contigs derived from a piroplasmid parasite (presumably Theileria ornithorhynchi). For most of the infecting parasite species, no molecular data had been available previously, and some of the datasets contain sequences representing large amounts of the parasite's gene repertoire. Conclusion: Our study suggests that parasite-derived contaminations represent a valuable source of information that can help to discover and identify new parasites, and provide information on previously unknown host-parasite interactions. We, therefore, argue that uncurated assembly data should routinely be made available in addition to the final assemblies.
Data from: Thermal fluctuations affect the transcriptome through mechanisms independent of average temperature
Terrestrial ectotherms are challenged by variation in both mean and variance of temperature. Phenotypic plasticity (thermal acclimation) might mitigate adverse effects, however, we lack a fundamental understanding of the molecular mechanisms of thermal acclimation and how they are affected by fluctuating temperature. Here we investigated the effect of thermal acclimation in Drosophila melanogaster on critical thermal maxima (CTmax) and associated global gene expression profiles as induced by two constant and two ecologically relevant (non-stressful) diurnally fluctuating temperature regimes. Both mean and fluctuation of temperature contributed to thermal acclimation and affected the transcriptome. The transcriptomic response to mean temperatures comprised modification of a major part of the transcriptome, while the response to fluctuations affected a much smaller set of genes, which was highly independent of both the response to a change in mean temperature and to the classic heat shock response. Although the independent transcriptional effects caused by fluctuations were relatively small, they are likely to contribute to our understanding of thermal adaptation. We provide evidence that environmental sensing, particularly phototransduction, is a central mechanism underlying the regulation of thermal acclimation to fluctuating temperatures. Thus, genes and pathways involved in phototransduction are likely of importance in fluctuating climates.
Data from: "Transcriptome resources for the marmot flea and plague vector, Oropsylla silantiewi" in Genomic Resources Notes accepted 1 February 2014 to 31 March 2014
This article documents the public availability of raw transcriptome sequence data, 47,882 assembled unigenes as well as their functional annotations, and 13,052 deduced ORFs (open reading frame) of a plague vector Oropsylla silantiewi.
Data from: A non-lethal sampling method to obtain, generate and assemble whole-blood transcriptomes from small, wild mammals
The acquisition of tissue samples from wild populations is a constant challenge in conservation biology, especially for endangered species and protected species where nonlethal sampling is the only option. Whole blood has been suggested as a nonlethal sample type that contains a high percentage of bodywide and genomewide transcripts and therefore can be used to assess the transcriptional status of an individual, and to infer a high percentage of the genome. However, only limited quantities of blood can be nonlethally sampled from small species and it is not known if enough genetic material is contained in only a few drops of blood, which represents the upper limit of sample collection for some small species. In this study, we developed a nonlethal sampling method, the laboratory protocols and a bioinformatic pipeline to sequence and assemble the whole blood transcriptome, using Illumina RNA-Seq, from wild greater mouse-eared bats (Myotis myotis). For optimal results, both ribosomal and globin RNAs must be removed before library construction. Treatment of DNase is recommended but not required enabling the use of smaller amounts of starting RNA. A large proportion of protein-coding genes (61%) in the genome were expressed in the blood transcriptome, comparable to brain (65%), kidney (63%) and liver (58%) transcriptomes, and up to 99% of the mitogenome (excluding D-loop) was recovered in the RNA-Seq data. In conclusion, this nonlethal blood sampling method provides an opportunity for a genomewide transcriptomic study of small, endangered or critically protected species, without sacrificing any individuals.
Data from: Identification and qualification of 500 nuclear, single-copy, orthologous genes for the Eupulmonata (Gastropoda) using transcriptome sequencing and exon capture
The qualification of orthology is a significant challenge when developing large, multiloci phylogenetic data sets from assembled transcripts. Transcriptome assemblies have various attributes, such as fragmentation, frameshifts and mis-indexing, which pose problems to automated methods of orthology assessment. Here, we identify a set of orthologous single-copy genes from transcriptome assemblies for the land snails and slugs (Eupulmonata) using a thorough approach to orthology determination involving manual alignment curation, gene tree assessment and sequencing from genomic DNA. We qualified the orthology of 500 nuclear, protein-coding genes from the transcriptome assemblies of 21 eupulmonate species to produce the most complete phylogenetic data matrix for a major molluscan lineage to date, both in terms of taxon and character completeness. Exon capture targeting 490 of the 500 genes (those with at least one exon >120 bp) from 22 species of Australian Camaenidae successfully captured sequences of 2825 exons (representing all targeted genes), with only a 3.7% reduction in the data matrix due to the presence of putative paralogs or pseudogenes. The automated pipeline Agalma retrieved the majority of the manually qualified 500 single-copy gene set and identified a further 375 putative single-copy genes, although it failed to account for fragmented transcripts resulting in lower data matrix completeness when considering the original 500 genes. This could potentially explain the minor inconsistencies we observed in the supported topologies for the 21 eupulmonate species between the manually curated and 'Agalma-equivalent' data set (sharing 458 genes). Overall, our study confirms the utility of the 500 gene set to resolve phylogenetic relationships at a range of evolutionary depths and highlights the importance of addressing fragmentation at the homolog alignment stage for probe design.
Data from: De novo transcriptome assembly databases in the butterfly orchid Phalaenopsis equestris
Orchids are renowned for their spectacular flowers and ecological adaptations. After the sequencing of the genome of the tropical epiphytic orchid Phalaenopsis equestris, we combined Illumina HiSeq2000 for RNA-Seq and Trinity for de novo assembly to characterize the transcriptomes for 11 diverse P. equestris tissues representing the root, stem, leaf, flower buds, column, lip, petal, sepal and three developmental stages of seeds. Our aims were to contribute to a better understanding of the molecular mechanisms driving the analysed tissue characteristics and to enrich the available data for P. equestris. Here, we present three databases. The first dataset is the RNA-Seq raw reads, which can be used to execute new experiments with different analysis approaches. The other two datasets allow different types of searches for candidate homologues. The second dataset includes the sets of assembled unigenes and predicted coding sequences and proteins, enabling a sequence-based search. The third dataset consists of the annotation results of the aligned unigenes versus the Nonredundant (Nr) protein database, Kyoto Encyclopaedia of Genes and Genomes (KEGG) and Clusters of Orthologous Groups (COG) databases with low e-values, enabling a name-based search.
Data from: Characterization of the genome and transcriptome of the blue tit Cyanistes caeruleus: polymorphisms, sex-biased expression and selection signals
Decoding genomic sequences and determining their variation within populations has potential to reveal adaptive processes and unravel the genetic basis of ecologically relevant trait variation within a species. The blue tit Cyanistes caeruleus – a long-time ecological model species – has been used to investigate fitness consequences of variation in mating and reproductive behaviour. However, very little is known about the underlying genetic changes due to natural and sexual selection in the genome of this songbird. As a step to bridge this gap, we assembled the first draft genome of a single blue tit, mapped the transcriptome of five females and five males to this reference, identified genomewide variants and performed sex-differential expression analysis in the gonads, brain and other tissues. In the gonads, we found a high number of sex-biased genes, and of those, a similar proportion were sex-limited (genes only expressed in one sex) in males and females. However, in the brain, the proportion of female-limited genes within the female-biased gene category (82%) was substantially higher than the proportion of male-limited genes within the male-biased category (6%). This suggests a predominant on-off switching mechanism for the female-limited genes. In addition, most male-biased genes were located on the Z-chromosome, indicating incomplete dosage compensation for the male-biased genes. We called more than 500 000 SNPs from the RNA-seq data. Heterozygote detection in the single reference individual was highly congruent between DNA-seq and RNA-seq calling. Using information from these polymorphisms, we identified potential selection signals in the genome. We list candidate genes which can be used for further sequencing and detailed selection studies, including genes potentially related to meiotic drive evolution. A public genome browser of the blue tit with the described information is available at http://public-genomes-ngs.molgen.mpg.de.
Data from: New Zealand tree and giant wētā (Orthoptera) transcriptomics reveal divergent selection patterns in metabolic loci
Exposure to low temperatures requires an organism to overcome physiological challenges. New Zealand wētā belonging to the genera Hemideina and Deinacrida are found across a wide range of thermal environments and therefore subject to varying selective pressures. Here we assess the selection pressures across the wētā phylogeny, with a particular emphasis on identifying genes under positive or diversifying selection. We used RNA-seq to generate transcriptomes for all 18 Deinacrida and Hemideina species. A total of 755 orthologous genes were identified using a bidirectional best hit approach, with the resulting gene set encompassing a diverse range of functional classes. Analysis of orthologue ratios of synonymous to non-synonymous amino acid changes found 83 genes that are under positive selection for at least one codon. A wide variety of Gene Ontology terms, enzymes and KEGG pathways are represented among these genes. In particular, enzymes involved in oxidative phosphorylation, melanin synthesis and free-radical scavenging are represented, consistent with physiological and metabolic changes that are associated with adaptation to alpine environments. Structural alignment of the transcripts with the most codons under positive selection revealed that the majority of sites are surface residues, and therefore have the potential to influence the thermostability of the enzyme, with the exception of prophenoloxidase where two residues near the active site are under selection. These proteins provide interesting candidates for further analysis of protein evolution
Data from: Transcriptome and proteome dynamics of a light-dark synchronized bacterial cell cycle
BACKGROUND: Growth of the ocean's most abundant primary producer, the cyanobacterium Prochlorococcus, is tightly synchronized to the natural 24-hour light-dark cycle. We sought to quantify the relationship between transcriptome and proteome dynamics that underlie this obligate photoautotroph's highly choreographed response to the daily oscillation in energy supply. METHODOLOGY/PRINCIPAL FINDINGS: Using Illumina RNA-sequencing transcriptomics and mass spectrometry-based quantitative proteomics, we measured timecourses of paired mRNA-protein abundances for 312 genes every 2 hours over a light-dark cycle. These temporal expression patterns reveal strong oscillations in transcript abundance that are broadly damped at the protein level, with mRNA levels varying on average 2.3 times more than the corresponding protein. The single strongest observed protein-level oscillation is in a ribonucleotide reductase, which may reflect a defense strategy against phage infection. The peak in abundance of most proteins also lags that of their transcript by 2-8 hours, and the two are completely antiphase for some genes. While abundant antisense RNA was detected, it apparently does not account for the observed divergences between expression levels. The redirection of flux through central carbon metabolism from daytime carbon fixation to nighttime respiration is associated with quite small changes in relative enzyme abundances. CONCLUSIONS/SIGNIFICANCE: Our results indicate that expression responses to periodic stimuli that are common in natural ecosystems (such as the diel cycle) can diverge significantly between the mRNA and protein levels. Protein expression patterns that are distinct from those of cognate mRNA have implications for the interpretation of transcriptome and metatranscriptome data in terms of cellular metabolism and its biogeochemical impact.
Single cell RNA-seq transcriptomic profile of circulating immune and progenitor cells in a mouse model of neonatal hypoxic/ischemic (HI) brain damage.
<p>Hematopoietic cells play a pivotal role in regulating the inflammatory and reparative immune responses triggered after ischemic tissue damage. The response initiated within the injured tissue leads to compositional and transcriptional changes in circulating hematopoietic and progenitor cells, which have been utilized as biomarkers. While the importance of different immune and progenitor cell subtypes in the development of ischemic damage has been extensively researched in adult tissue injuries, there has been limited investigation in neonates. This is a critical developmental stage where ischemic damage can result in severe and irreversible health consequences if not promptly treated. To determine how ischemic damage could affect circulating cells in neonates, we have induced hypoxic-ischemic (HI) brain damage in seven-day-old mice, characterized by focal white and gray-matter injury (Rice-Vannucci model). Brain damage and circulating cells were analyzed at 48h post-HI, the intermediate reparative/inflammatory response phase post-injury. We applied scRNAseq to dissect the transcriptional and cellular composition changes in the peripheral blood of HI-treated and SHAM control neonates. This study provides the first scRNAseq dataset for immune and progenitor circulating cells in newborns with cerebral HI damage. It may help to identify biomarkers and selective therapeutic approaches aimed at modulating inflammatory and reparative pathways. </p><p>CD1 postnatal day 7 (P7) mice were subjected to hypoxic/ischemic (HI) brain injury by permanent ligation of the left common carotid artery followed, after 2h recover, by relocation to a hypoxia chamber for 90 minutes. Sham control mice (SHAM) underwent a skin incision and wound closure followed by hypoxia exposure. Circulating blood cells were collected from SHAM and HI mice at P9. After red blood cells (RBC) lysis, 7AAD-Ter119- cells were FACS sorted and analysed using sc RNAseq. Other samples were FACS sorted for CD45+CD11+ and CD45-CD31+ cells and mixed.</p>
PROTEOMIC AND TRANSCRIPTOMIC RESPONSE OF HUMAN SKELETAL MUSCLE TO 12-WEEK RESISTANCE TRAINING
Open the record for dataset details and reuse information.
Immuno-transcriptomic profiling of blood and tumor tissue identifies gene signatures associated with immunotherapy response in metastatic bladder cancer
<p>The dataset contains RNA-sequencing data (raw fastq files, not trimmed) of whole blood and whole bladder tissue samples and metadata file with sample annotation. The samples have been collected in the context of the study "Immuno-transcriptomic profiling of blood and tumor tissue identifies gene signatures associated with immunotherapy response in metastatic bladder cancer". In this study, we investigated which are the local and systemic immune changes in an experimental model of muscle-invasive bladder cancer upon immunotherapy. The RNA-sequencing data have been used to produce the results presented in the study. Please refer to the study publication material and methods and the metadata file for details.</p>
Expansion of the RNAStructuromeDB to include secondary structural data spanning the human protein-coding transcriptome
<p>This dataset includes the -2, -1, and no filter z-score dot bracket files from ScanFold for all protein coding transcript isoforms.</p>
Spatially Resolved Transcriptomics Atlas of Matched Primary and Metastatic Pancreatic Cancer Reveal Principles of Ecological Adaptation
Open the record for dataset details and reuse information.
STCGAN: a novel Cycle-Consistent Generative Adversarial Network for Spatial Transcriptomics Cellular Deconvolution
Open the record for dataset details and reuse information.
IF and SCRINSHOT image data of probe set selection for targeted spatial transcriptomics
Open the record for dataset details and reuse information.
Accurate Spatial Heterogeneity Dissection and Gene Regulation Interpretation for Spatial Transcriptomics using Dual Graph Contrastive Learning
Open the record for dataset details and reuse information.
Transcriptomic analysis and epigenetic regulators in human oocytes at different stages of oocyte meiotic maturation
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.