Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

25,372

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

25,372 results for “Transcriptomics”

Learn how ShareScore rates datasets ↗
dryad28/100

Data from: Cross-species transferability of SSR loci developed from transcriptome sequencing in lodgepole pine

With the advent of next generation sequencing technologies, transcriptome level sequence collections are arising as prominent resources for the discovery of gene-based molecular markers. In a previous study more than 15 000 simple sequence repeats (SSRs) in expressed sequence tag (EST) sequences resulting from 454 pyrosequencing of Pinus contorta cDNA were identified. From these we developed PCR primers for approximately 4000 candidate SSRs. Here, we tested 184 of these SSRs for successful amplification across P. contorta and eight other pine species and examined patterns of polymorphism and allelic variability for a subset of these SSRs. Cross-species transferability was high, with high percentages of loci producing PCR products in all species tested. In addition, 50% of the loci we screened across panels of individuals from three of these species were polymorphic and allelically diverse. We examined levels of diversity in a subset of these SSRs by collecting genotypic data across several populations of Pinus ponderosa in northern Wyoming. Our results indicate the utility of mining pyrosequenced EST collections for gene-based SSRs and provide a source of molecular markers that should bolster evolutionary genetic investigations across the genus Pinus.

opencc-zeroDec 2010View details →
dryad28/100

Data from: Transcriptome-wide mining, characterization, and development of microsatellite markers in Lychnis kiusiana (Caryophyllaceae)

Background: Lychnis kiusiana Makino is an endangered perennial herb native to wetland areas in Korea and Japan. Despite its conservational and evolutionary significance, population genetic resources are lacking for this species. Next-generation sequencing has been accepted as a rapid and cost-effective solution for the identification of microsatellite markers in nonmodel plants. Results: Using Illumina HiSeq 2000 sequencing technology, we assembled 67,498,600 reads into 91,900 contigs and identified 11,403 microsatellite repeat motifs in 9,563 contigs. A total of 4,510 microsatellite-containing transcripts had Gene Ontology (GO) annotations, and Kyoto Encyclopedia of Genes and Genomes (KEGG) analysis identified 124 pathways with significant scores. Many microsatellites in the L. kiusiana leaf transcriptome were linked to genes involved in the plant response to light intensity, salt stress, temperature stimulus, and nutrient and water deprivation. A total of 12,486 single-nucleotide polymorphisms (SNPs) were identified on transcripts harboring microsatellites. The analysis of nucleotide substitution rates for 2,389 unigenes indicated that 39 genes were under strong positive selection. The primers of 6,911 microsatellites were designed, and 40 of 50 selected primer pairs were consistently and successfully amplified from 51 individuals. Twenty-five of these were polymorphic, and the average number of alleles per SSR locus was 6.96, with a range from 2 to 15. The observed and expected heterozygosities ranged from 0.137 to 0.902 and 0.131 to 0.827, respectively, and locus-specific FIS estimates ranged from -0.116 to 0.290. Eleven of the 25 primer pairs were successfully amplified in three additional species of Lychnis: 56% in L. wilfordii, 64% in L. cognata and 80% in L. fulgens. Conclusions: The transcriptomic SSR markers of Lychnis kiusiana provide a valuable resource for understanding the population genetics, evolutionary history, and effective conservation management of this species. Furthermore, the identified microsatellite loci linked to the annotated genes should be useful for developing functional markers of L. kiusiana. The developed markers represent a potentially valuable source of transcriptomic SSR markers for population genetic analyses with moderate levels of cross-taxon portability.

opencc-zeroDec 2017View details →
dryad28/100

Data from: Transcriptome sequencing and microarray development for the woodrat (Neotoma spp.): custom genetic tools for exploring herbivore ecology

Massively parallel sequencing has enabled the creation of novel, in-depth genetic tools for nonmodel, ecologically important organisms. We present the de novo transcriptome sequencing, analysis and microarray development for a vertebrate herbivore, the woodrat (Neotoma spp.). This genus is of ecological and evolutionary interest, especially with respect to ingestion and hepatic metabolism of potentially toxic plant secondary compounds. We generated a liver transcriptome of the desert woodrat (Neotoma lepida) using the Roche 454 platform. The assembled contigs were well annotated using rodent references (99.7% annotation), and biotransformation function was reflected in the gene ontology. The transcriptome was used to develop a custom microarray (eArray, Agilent). We tested the microarray with three experiments: one across species with similar habitat (thus, dietary) niches, one across species with different habitat niches and one across populations within a species. The resulting one-colour arrays had high technical and biological quality. Probes designed from the woodrat transcriptome performed significantly better than functionally similar probes from the Norway rat (Rattus norvegicus). There were a multitude of expression differences across the woodrat treatments, many of which related to biotransformation processes and activities. The pattern and function of the differences indicate shared ecological pressures, and not merely phylogenetic distance, play an important role in shaping gene expression profiles of woodrat species and populations. The quality and functionality of the woodrat transcriptome and custom microarray suggest these tools will be valuable for expanding the scope of herbivore biology, as well as the exploration of conceptual topics in ecology.

opencc-zeroDec 2012View details →
dryad28/100

Data from: De novo transcriptome characterization and development of genomic tools for Scabiosa columbaria L. using next-generation sequencing techniques.

Next-generation sequencing (NGS) technologies are increasingly applied in many organisms, including non-model organisms that are important for ecological and conservation purposes. Illumina and 454 sequencing are among the most used NGS technologies and have been shown to produce optimal results at reasonable costs when used together. Here, we describe the combined application of these two NGS technologies to characterize the transcriptome of a plant species of ecological and conservation relevance for which no genomic resource is available, Scabiosa columbaria. We obtained 528,557 reads from a 454 GS-FLX run and a total of 28,993,627 reads from two lanes of an Illumina GAII single run. After reads trimming, the de novo assembly of both types of reads produced 109,630 contigs. Both the contigs and the >75 bp remaining singletons were blasted against Uniprot/Swissprot database, resulting in 29,676 and 10,515 significant hits, respectively. Based on sequence similarity with known gene products, these sequences represent at least 12,516 unique genes, most of which are well covered by contig sequences. In addition, we identified 4,320 microsatellite loci, of which 856 had flanking sequences suitable for PCR primer design. We also identified 75,054 putative SNPs. This annotated sequence collection and the relative molecular markers represent a main genomic resource for S. columbaria which should contribute to future research in conservation and population biology studies. Our results demonstrate the utility of NGS technologies as starting point for the development of genomic tools in nonmodel but ecologically important species.

opencc-zeroDec 2009View details →
dryad28/100

Data from: A comprehensive transcriptome of early development in yellowtail kingfish (Seriola lalandi)

Seriola lalandi is an ecologically and economically important species that is globally distributed in temperate and subtropical marine waters. The aim of this study was to identify large numbers of genic single nucleotide polymorphisms (SNPs) and differential gene expression (DGE) related to the early development of normal and deformed S. lalandi larvae using high-throughput RNA-seq data. A de novo assembly of reads generated 40 066 genes ranging from 300 bases to 64 799 bases with an N90 of 788 bases. Homology search and protein signature recognition assigned gene ontology (GO) terms to a total of 15 744 (39.34%) genes. A search against the Kyoto Encyclopedia of Genes and Genomes Pathway database (KEGG) retrieved 6808 KEGG orthology (KO) identifiers for 10 520 genes (26.25%), and mapping of KO identifiers generated 337 KEGG pathways. Comparisons of annotated genes revealed that 1262 genes were downregulated and 1047 genes were upregulated in the deformed larvae group compared to the normal group of larvae. Additionally, we identified 6989 high-quality SNPs from the assembled transcriptome. These putative SNPs contain 4415 transitions and 2574 transversions, which will be useful for further ecological studies of S. lalandi. This is the first study to use a global transcriptomic approach in S. lalandi, and the resources generated can be used further for investigation of gene expression of marine teleosts to investigate larval developmental biology. The results of the GO enrichment analysis highlight the crucial role of the extracellular matrix in normal skeleton development, which could be important for future studies of skeletal deformities in S. lalandi and other marine species.

opencc-zeroAug 2015View details →
dryad28/100

Data from: Phylogenomic analysis of transcriptome data elucidates co-occurrence of a paleopolyploid event and the origin of bimodal karyotypes in Agavoideae (Asparagaceae)

PREMISE OF THE STUDY: The stability of the bimodal karyotype found in Agave and closely related species has long interested botanists. The origin of the bimodal karyotype has been attributed to allopolyploidy, but this hypothesis has not been tested. Next Generation transcriptome sequence data were used to test whether a paleopolyploid event occurred on the same branch of the Agavoideae phylogenetic tree as the origin of the Yucca-Agave bimodal karyotype. METHODS: Illumina RNAseq data were generated for phylogenetically strategic species in Agavoideae. Paleopolyploidy was inferred in analyses of frequency plots for synonymous substitutions per synonymous site (Ks) between Hosta, Agave and Chlorophytum paralogous and orthologous gene pairs. Phylogenies of gene families including paralogous genes for these species and outgroup species were estimated in order to place inferred paleopolyploid events on a species tree. KEY RESULTS: Ks frequency plots suggested paleopolyploid events in the history of the genera Agave, Hosta and Chlorophytum. Phylogenetic analyses of gene families estimated from transcriptome data revealed two polyploid events: one predating the last common ancestor of Agave and Hosta and one within the lineage leading to Chlorophytum. CONCLUSIONS: We found that allopolyoidy and the origin of the Yucca-Agave bimodal karyotype co-occur on the same lineage consistent with the hypothesis that the bimodal karyotype is a consequence of allopolyploidy. We discuss this and alternative mechanisms for the formation of the Yucca-Agave bimodal karyotype. More generally, we illustrate how the use of next generation sequencing technology is a cost-efficient means for assessing genome evolution in non-model species.

opencc-zeroDec 2011View details →
dryad28/100

Data from: Microsatellite analyses across three diverse vertebrate transcriptomes (Acipenser fulvescens, Ambystoma tigrinum, and Dipodomys spectabilis)

Historically, many population genetics studies have utilized microsatellite markers sampled at random from the genome and presumed to be selectively neutral. Recent studies, however, have shown that microsatellites can occur in transcribed regions, where they are more likely to be under selection. In this study, we mined microsatellites from transcriptomes generated by 454-pyrosequencing for three vertebrate species: lake sturgeon (Acipenser fulvescens), tiger salamander (Ambystoma tigrinum), and kangaroo rat (Dipodomys spectabilis). We evaluated (i) the occurrence of microsatellites across species; (ii) whether particular gene ontology terms were over-represented in genes that contained microsatellites; (iii) whether repeat motifs were located in untranslated regions or coding sequences of genes; and (iv) in silico polymorphism. Microsatellites were less common in tiger salamanders than in either lake sturgeon or kangaroo rats. Across libraries, trinucleotides were found more frequently than any other motif type, presumably because they do not cause frameshift mutations. By evaluating variation across reads assembled to a given contig, we were able to identify repeat motifs likely to be polymorphic. Our study represents one of the first comparative data sets on the distribution of vertebrate microsatellites within expressed genes. Our results reinforce the idea that microsatellites do not always occur in noncoding DNA, but commonly occur in expressed genes.

opencc-zeroDec 2013View details →
dryad28/100

Data from: NEMBASE4: the nematode transcriptome resource

Nematode parasites are of major importance in human health and agriculture, and free-living species deliver essential ecosystem services. The genomics revolution has resulted in the production of many datasets of expressed sequence tags (ESTs) from a phylogenetically wide range of nematode species, but these are not easily compared. NEMBASE4 presents a single portal onto extensively functionally annotated, EST-derived transcriptomes from over sixty species of nematodes, including plant and animal parasites and free-living taxa. Using the PartiGene suite of tools, we have assembled the ESTs publicly available for each species into a high-quality set of putative transcripts. These transcripts have been translated to produce a protein sequence resource, and each annotated with functional information derived from comparison to well-studied nematode species such as Caenorhabditis elegans and also other non-nematode resources. By cross-comparing the sequences within NEMBASE4, we have also generated a protein family assignment for each translation. The data are presented in an openly-accessible, interactive database. To demonstrate the utility of NEMBASE4, we have used the database to examine the uniqueness of the transcriptomes of major clades of parasitic nematodes, identifying lineage-restricted genes that may underpin particular parasitic phenotypes, possible viral pathogens of nematodes, and nematode-unique protein families that may be developed as drug targets.

opencc-zeroDec 2010View details →
dryad28/100

Data from: Transcriptome resources for the perennial sunflower Helianthus maximiliani obtained from ecologically divergent populations

Next generation sequencing (NGS) technologies provide a rapid means to generate genomic resources for species exhibiting interesting ecological and evolutionary variation but for which such resources are scant or nonexistent. In the current report, we utilize 454 pyrosequencing to obtain transcriptome information for multiple individuals and tissue types from geographically disparate and ecologically differentiated populations of the perennial sunflower species Helianthus maximiliani. A total of 850,275 raw reads were obtained averaging 355 bp in length. Reads were assembled, post processing, into 16,681 unique contigs with an N50 of 898 bp and a total length of 13.6 Mb. A majority (67%) of these contigs were annotated based on comparison to the Arabidopsis thaliana genome (TAIR10). Contigs were identified that exhibit high similarity to genes associated with natural variation in flowering time and freezing tolerance in other plant species and will facilitate future studies aimed at elucidating the molecular basis of clinal life history variation and adaptive differentiation in H. maximiliani. Large numbers of gene-associated simple sequence repeats (SSRs) and single nucleotide polymorphisms (SNPs) also were identified that can be deployed in mapping and population genomic analyses.

opencc-zeroDec 2013View details →
dryad28/100

Data from: Insights into the human mesenchymal stromal/stem cell identity through integrative transcriptomic profiling

Background: Mesenchymal Stromal/Stem Cells (MSCs), isolated under the criteria established by the ISCT, still have a poorly characterized phenotype that is difficult to distinguish from similar cell populations. Although the field of transcriptomics and functional genomics has quickly grown in the last decade, a deep comparative analysis of human MSCs expression profiles in a meaningful cellular context has not been yet performed. There is also a need to find a well-defined MSCs gene-signature because many recent biomedical studies show that key cellular interaction processes (i.e. inmuno-modulation, cellular cross-talk, cellular maintenance, differentiation, epithelial-mesenchymal transition) are dependent on the mesenchymal stem cells within the stromal niche. Results: In this work we define a core mesenchymal lineage signature of 489 genes based on a deep comparative analysis of multiple transcriptomic expression data series that comprise: (i) MSCs of different tissue origins; (ii) MSCs in different states of commitment; (iii) other related non-mesenchymal human cell types. The work integrates several public datasets, as well as de-novo produced microarray and RNA-Seq datasets. The results present tissue-specific signatures for adipose tissue, chorionic placenta, and bone marrow MSCs, as well as for dermal fibroblasts; providing a better definition of the relationship between fibroblasts and MSCs. Finally, novel CD marker patterns and cytokine-receptor profiles are unravelled, especially for BM-MSCs; with MCAM (CD146) revealed as a prevalent marker in this subtype of MSCs. Conclusions: The improved biomolecular characterization and the released genome-wide expression signatures of human MSCs provide a comprehensive new resource that can drive further functional studies and redesigned cell therapy applications.

opencc-zeroDec 2015View details →
dryad28/100

Data from: Transcriptome profile analysis from different sex types of Ginkgo biloba L.

In plants, sex determination is a comprehensive process of correlated events, which involves genes that are differentially and/or specifically expressed in distinct developmental phases. Exploring gene expression profiles from different sex types will contribute to fully understanding sex determination in plants. In this study, we conducted RNA-sequencing of female and male buds (FB and MB) as well as ovulate strobilus and staminate strobilus (OS and SS) of Ginkgo biloba to gain insights into the genes potentially related to sex determination in this species. Approximately 60 Gb of clean reads were obtained from eight cDNA libraries. De novo assembly of the clean reads generated 108,307 unigenes with an average length of 796 bp. Among these unigenes, 51,953 (47.97%) had at least one significant match with a gene sequence in the public databases searched. A total of 4,709 and 9,802 differentially expressed genes (DEGs) were identified in MB vs. FB and SS vs. OS, respectively. Genes involved in plant hormone signal and transduction as well as those encoding DNA methyltransferase were found to be differentially expressed between different sex types. Their potential roles in sex determination of G. biloba were discussed. Pistil-related genes were expressed in male buds while anther-specific genes were identified in female buds, suggesting that dioecism in G. biloba was resulted from the selective arrest of reproductive primordia. High correlation of expression level was found between the RNA-Seq and quantitative real-time PCR results. The transcriptome resources that we generated allowed us to characterize gene expression profiles and examine differential expression profiles, which provided foundations for identifying functional genes associated with sex determination in G. biloba.

opencc-zeroDec 2015View details →
dryad28/100

Data from: An evaluation of transcriptome-based exon capture for frog phylogenomics across multiple scales of divergence (Class: Amphibia, Order: Anura)

Custom sequence capture experiments are becoming an efficient approach for gathering large sets of orthologous markers in nonmodel organisms. Transcriptome-based exon capture utilizes transcript sequences to design capture probes, typically using a reference genome to identify intron–exon boundaries to exclude shorter exons (<200 bp). Here, we test directly using transcript sequences for probe design, which are often composed of multiple exons of varying lengths. Using 1260 orthologous transcripts, we conducted sequence captures across multiple phylogenetic scales for frogs, including outgroups ~100 Myr divergent from the ingroup. We recovered a large phylogenomic data set consisting of sequence alignments for 1047 of the 1260 transcriptome-based loci (~561 000 bp) and a large quantity of highly variable regions flanking the exons in transcripts (~70 000 bp), the latter improving substantially by only including ingroup species (~797 000 bp). We recovered both shorter (<100 bp) and longer exons (>200 bp), with no major reduction in coverage towards the ends of exons. We observed significant differences in the performance of blocking oligos for target enrichment and nontarget depletion during captures, and differences in PCR duplication rates resulting from the number of individuals pooled for capture reactions. We explicitly tested the effects of phylogenetic distance on capture sensitivity, specificity, and missing data, and provide a baseline estimate of expectations for these metrics based on a priori knowledge of nuclear pairwise differences among samples. We provide recommendations for transcriptome-based exon capture design based on our results, cost estimates and offer multiple pipelines for data assembly and analysis.

opencc-zeroDec 2015View details →
dryad28/100

Data from: Animal tracking meets migration genomics: transcriptomic analysis of a partially migratory bird species

Seasonal migration is a widespread phenomenon, which is found in many different lineages of animals. This spectacular behaviour allows animals to avoid seasonally adverse environmental conditions to exploit more favourable habitats. Migration has been intensively studied in birds, which display astonishing variation in migration strategies, thus providing a powerful system for studying the ecological and evolutionary processes that shape migratory behaviour. Despite intensive research, the genetic basis of migration remains largely unknown. Here we used state-of-the-art radio-tracking technology to characterize the migratory behaviour of a partially migratory population of European blackbirds (Turdus merula) in southern Germany. We compared gene expression of resident and migrant individuals using high-throughput transcriptomics in blood samples. Analyses of sequence variation revealed a non-significant genetic structure between blackbirds differing by their migratory phenotype. We detected only four differentially expressed genes between migrants and residents, which might be associated with hyperphagia, moulting, and enhanced DNA replication and transcription. The most pronounced changes in gene expression occurred between migratory birds depending on when, in relation to their date of departure, blood was collected. Overall, the differentially expressed genes detected in this analysis may play crucial roles in determining the decision to migrate, or in controlling the physiological processes required for the onset of migration. These results provide new insights into, and testable hypotheses for, the molecular mechanisms controlling the migratory phenotype and its underlying physiological mechanisms in blackbirds and other migratory bird species.

opencc-zeroDec 2016View details →
dryad28/100

Data from: Insights into the maize pan-genome and pan-transcriptome

Genomes at the species level are dynamic, with genes present in every individual (core) and genes in a subset of individuals (dispensable) that collectively constitute the pan-genome. Using transcriptome sequencing of seedling RNA from 503 maize (Zea mays) inbred lines to characterize the maize pan-genome, we identified 8681 representative transcript assemblies (RTAs) with 16.4% expressed in all lines and 82.7% expressed in subsets of the lines. Interestingly, with linkage disequilibrium mapping, 76.7% of the RTAs with at least one single nucleotide polymorphism (SNP) could be mapped to a single genetic position, distributed primarily throughout the nonpericentromeric portion of the genome. Stepwise iterative clustering of RTAs suggests, within the context of the genotypes used in this study, that the maize genome is restricted and further sampling of seedling RNA within this germplasm base will result in minimal discovery. Genome-wide association studies based on SNPs and transcript abundance in the pan-genome revealed loci associated with the timing of the juvenile-to-adult vegetative and vegetative-to-reproductive developmental transitions, two traits important for fitness and adaptation. This study revealed the dynamic nature of the maize pan-genome and demonstrated that a substantial portion of variation may lie outside the single reference genome for a species.

opencc-zeroDec 2013View details →
dryad28/100

Data from: Genome reannotation of the lizard Anolis carolinensis based on 14 adult and embryonic deep transcriptomes

Background: The green anole lizard, Anolis carolinensis, is a key species for both laboratory and field-based studies of evolutionary genetics, development, neurobiology, physiology, behavior, and ecology. As the first non-avian reptilian genome sequenced, A. carolinesis is also a prime reptilian model for comparison with other vertebrate genomes. The public databases of Ensembl and NCBI have provided a first generation gene annotation of the anole genome that relies primarily on sequence conservation with related species. A second generation annotation based on tissue-specific transcriptomes would provide a valuable resource for molecular studies. Results: Here we provide an annotation of the A. carolinensis genome based on de novo assembly of deep transcriptomes of 14 adult and embryonic tissues. This revised annotation describes 59,373 transcripts, compared to 16,533 and 18,939 currently for Ensembl and NCBI, and 22,962 predicted protein-coding genes. A key improvement in this revised annotation is coverage of untranslated region (UTR) sequences, with 79% and 59% of transcripts containing 5' and 3' UTRs, respectively. Gaps in genome sequence from the current A. carolinensis build (Anocar2.0) are highlighted by our identification of 16,542 unmapped transcripts, representing 6,695 orthologues, with less than 70% genomic coverage. Conclusions: Incorporation of tissue-specific transcriptome sequence into the A. carolinensis genome annotation has markedly improved its utility for comparative and functional studies. Increased UTR coverage allows for more accurate predicted protein sequence and regulatory analysis. This revised annotation also provides an atlas of gene expression specific to adult and embryonic tissues.

opencc-zeroDec 2012View details →
dryad28/100

Data from: Combining animal personalities with transcriptomics resolves individual variation within a wild-type zebrafish population and identifies underpinning molecular differences in brain function

Resolving phenotype variation within a population in response to environmental perturbation is central to understanding biological adaptation. Relating meaningful adaptive changes at the level of the transcriptome requires the identification of processes that have a functional significance for the individual. This remains a major objective towards understanding the complex interactions between environmental demand and an individual's capacity to respond to such demands. The interpretation of such interactions and the significance of biological variation between individuals from the same or different populations remain a difficult and under-addressed question. Here, we provide evidence that variation in gene expression between individuals in a zebrafish population can be partially resolved by a priori screening for animal personality and accounts for >9% of observed variation in the brain transcriptome. Proactive and reactive individuals within a wild-type population exhibit consistent behavioural responses over time and context that relates to underlying differences in regulated gene networks and predicted protein–protein interactions. These differences can be mapped to distinct regions of the brain and provide a foundation towards understanding the coordination of underpinning adaptive molecular events within populations.

opencc-zeroDec 2013View details →
dryad28/100

Data from: Transcriptome analysis indicates considerable divergence in alternative splicing between duplicated genes in Arabidopsis thaliana

Gene and genome duplication events have created a large number of new genes in plants that can diverge by evolving new expression profiles and functions (neofunctionalization) or dividing extant ones (subfunctionalization). Alternative splicing (AS) generates multiple types of mRNA from a single type of pre-mRNA by differential intron splicing. It can result in new protein isoforms or down-regulation of gene expression by transcript decay. Using RNA-seq we investigated the degree to which alternative splicing patterns are conserved between duplicated genes in Arabidopsis thaliana. Our results revealed that 30% of AS events in alpha whole genome duplicates, and 33% of AS events in tandem duplicates, are qualitatively conserved within leaf tissue. Loss of ancestral splice forms, as well as asymmetric gain of new splice forms, may account for this divergence. Conserved events had different frequencies, as only 31% of shared AS events in alpha whole genome duplicates and 41% of shared AS events in tandem duplicates had similar frequencies in both paralogs, indicating considerable quantitative divergence. Analysis of published RNA-seq data from nonsense mediated decay (NMD) mutants indicated that 85% of alpha whole genome duplicates and 89% of tandem duplicates have diverged in their AS-induced NMD. Our results indicate that alternative splicing shows a high degree of divergence between paralogs such that qualitatively conserved alternative splicing events tend to have quantitative divergence. Divergence in AS patterns between duplicates may be a mechanism of regulating expression level divergence.

opencc-zeroDec 2013View details →
dryad28/100

Data from: De novo transcriptomic analyses for non-model organisms: an evaluation of methods across a multi-species data set

High-throughput sequencing (HTS) is revolutionizing biological research by enabling scientists to quickly and cheaply query variation at a genomic scale. Despite the increasing ease of obtaining such data, using these data effectively still poses notable challenges, especially for those working with organisms without a high-quality reference genome. For every stage of analysis – from assembly to annotation to variant discovery – researchers have to distinguish technical artefacts from the biological realities of their data before they can make inference. In this work, I explore these challenges by generating a large de novo comparative transcriptomic data set data for a clade of lizards and constructing a pipeline to analyse these data. Then, using a combination of novel metrics and an externally validated variant data set, I test the efficacy of my approach, identify areas of improvement, and propose ways to minimize these errors. I find that with careful data curation, HTS can be a powerful tool for generating genomic data for non-model organisms.

opencc-zeroDec 2012View details →
dryad28/100

Data from: Evaluation of the impact of RNA preservation methods of spiders for de novo transcriptome assembly

With advances in high-throughput sequencing technologies, de novo transcriptome sequencing and assembly has become a cost-effective method to obtain comprehensive genetic information of a species of interest, especially in non-model species with large genomes such as spiders. However, high-quality RNA is essential for successful sequencing and sample preservation conditions require careful consideration for the effective storage of field-collected samples. To this end, we report a streamlined feasibility study of various storage conditions and their effects on de novo transcriptome assembly results. The storage parameters considered include temperatures ranging from room temperature to -80°C; preservatives, including ethanol, RNAlater, TRIzol, and RNAlater-ICE; and sample submersion states. As a result, intact RNA was extracted and assembly was successful when samples were preserved at low temperatures regardless of the type of preservative used. The assemblies as well as the gene expression profiles were shown to be robust to RNA degradation, when 30 million 150 bp paired-end reads are obtained. The parameters for sample storage, RNA extraction, library preparation, sequencing, and in silico assembly considered in this work provide a guideline for the study of field-collected samples of spiders.

opencc-zeroDec 2014View details →
dryad28/100

Data from: Adaptation of a polyphagous herbivore to a novel host plant extensively shapes the transcriptome of herbivore and host

Generalist arthropod herbivores rapidly adapt to a broad range of host plants. However, the extent of transcriptional reprogramming in the herbivore and its hosts associated with adaptation remains poorly understood. Using the spider mite Tetranychus urticae and tomato as models with available genomic resources, we investigated the reciprocal genomewide transcriptional changes in both spider mite and tomato as a consequence of mite's adaptation to tomato. We transferred a genetically diverse mite population from bean to tomato where triplicated populations were allowed to propagate for 30 generations. Evolving populations greatly increased their reproductive performance on tomato relative to their progenitors when reared under identical conditions, indicative of genetic adaptation. Analysis of transcriptional changes associated with mite adaptation to tomato revealed two main components. First, adaptation resulted in a set of mite genes that were constitutively downregulated, independently of the host. These genes were mostly of an unknown function. Second, adapted mites mounted an altered transcriptional response that had greater amplitude of changes when re-exposed to tomato, relative to nonadapted mites. This gene set was enriched in genes encoding detoxifying enzymes and xenobiotic transporters. Besides the direct effects on mite gene expression, adaptation also indirectly affected the tomato transcriptional responses, which were attenuated upon feeding of adapted mites, relative to the induced responses by nonadapted mite feeding. Thus, constitutive downregulation and increased transcriptional plasticity of genes in a herbivore may play a central role in adaptation to host plants, leading to both a higher detoxification potential and reduced production of plant defence compounds.

opencc-zeroDec 2014View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record