Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
25,372
datasets available to search
ShareScore release 0.7.1
Dataset results
25,372 results for “Transcriptomics”
Data from: Integrated network analysis identifies fight-club nodes as a class of hubs encompassing key putative switch genes that induce major transcriptome reprogramming during grapevine development
We developed an approach that integrates different network-based methods to analyze the correlation network arising from large-scale gene expression data. By studying grapevine (Vitis vinifera) and tomato (Solanum lycopersicum) gene expression atlases and a grapevine berry transcriptomic data set during the transition from immature to mature growth, we identified a category named "fight-club hubs" characterized by a marked negative correlation with the expression profiles of neighboring genes in the network. A special subset named "switch genes" was identified, with the additional property of many significant negative correlations outside their own group in the network. Switch genes are involved in multiple processes and include transcription factors that may be considered master regulators of the previously reported transcriptome remodeling that marks the developmental shift from immature to mature growth. All switch genes, expressed at low levels in vegetative/green tissues, showed a significant increase in mature/woody organs, suggesting a potential regulatory role during the developmental transition. Finally, our analysis of tomato gene expression data sets showed that wild-type switch genes are downregulated in ripening-deficient mutants. The identification of known master regulators of tomato fruit maturation suggests our method is suitable for the detection of key regulators of organ development in different fleshy fruit crops.
Data from: Transcriptomics in the wild: hibernation physiology in free‐ranging dwarf lemurs
Hibernation is an adaptive strategy some mammals use to survive highly seasonal or unpredictable environments. We present the first investigation on the transcriptomics of hibernation in a natural population of primate hibernators: Crossley's dwarf lemurs (Cheirogaleus crossleyi). Using capture–mark–recapture techniques to track the same animals over a period of 7 months in Madagascar, we used RNA‐seq to compare gene expression profiles in white adipose tissue (WAT) during three distinct physiological states. We focus on pathway analysis to assess the biological significance of transcriptional changes in dwarf lemur WAT and, by comparing and contrasting what is known in other model hibernating species, contribute to a broader understanding of genomic contributions of hibernation across Mammalia. The hibernation signature is characterized by a suppression of lipid biosynthesis, pyruvate metabolism and mitochondrial‐associated functions, and an accumulation of transcripts encoding ribosomal components and iron‐storage proteins. The data support a key role of pyruvate dehydrogenase kinase isoenzyme 4 (PDK4) in regulating the shift in fuel economy during periods of severe food deprivation. This pattern of PDK4 holds true across representative hibernating species from disparate mammalian groups, suggesting that the genetic underpinnings of hibernation may be ancestral to mammals.
Data from: Genomics of Compositae crops: reference transcriptome assemblies, and evidence of hybridization with wild relatives
Although the Compositae harbours only two major food crops, sunflower and lettuce, many other species in this family are utilized by humans and have experienced various levels of domestication. Here we have used next generation sequencing technology to develop 15 reference transcriptome assemblies for Compositae crops or their wild relatives. These data allow us to gain insight into the evolutionary and genomic consequences of plant domestication. Specifically, we performed Illumina sequencing of Cichorium endivia, Cichorium intybus, Echinacea angustifolia, Iva annua, Helianthus tuberosus, Dahlia hybrida, Leontodon taraxacoides and Glebionis segetum, as well 454 sequencing of Guizotia scabra, Stevia rebaudiana, Parthenium argentatum and Smallanthus sonchifolius. Illumina reads were assembled using Trinity, and 454 reads were assembled using MIRA and CAP3. We evaluated the coverage of the transcriptomes using BLASTX analysis of a set of ultra-conserved orthologs (UCOs) and recovered most of these genes (88-98%). We found a correlation between contig length and read length for the 454 assemblies, and greater contig lengths for the 454 compared to the Illumina assemblies. This suggests that longer reads can aid in the assembly of more complete transcripts. Finally, we compared the divergence of orthologs at synonymous sites (Ks) between Compositae crops and their wild relatives and found greater divergence when the progenitors were self-incompatible. We also found greater divergence between pairs of taxa that had some evidence of post-zygotic isolation. For several more distantly related congeners, such as chicory and endive, we identified a signature of introgression in the distribution of Ks values.
Data from: "White-tailed deer (Odocoileus virginianus) transcriptome assembly and SNP discovery" in Genomic Resources Notes accepted 1 June 2013-31 July 2013
White-tailed deer (Odocoileus virginianus) are among the most abundant and widespread large mammals in the Americas, comprising up to 38 subspecies ranging from Northern Canada to Peru. Although believed to have high genetic diversity, surprisingly few genomic resources are currently available, despite the species' ecological and economic importance. White-tailed deer and other cervids throughout central North America are currently being afflicted by chronic wasting disease (CWD), one of the degenerative prion diseases collectively known as transmissible spongiform encephalopathies. Although CWD is of major importance to white-tailed deer management, little is currently known about innate resistance or susceptibility to CWD outside of polymorphisms in the prion protein gene, Prnp, though a recent study using microsatellites suggests that the disease may have additional underlying genetic components. Further association analysis is hindered by low marker density. In this study, we used high-throughput SOLiD sequencing to create novel sequence data for white-tailed deer and identify single-nucleotide polymorphisms, using the pooled blood transcriptomes of six individuals. In total, we generated 14,010 contigs of length ≥ 200 nt, representing 4,104,760 nt of unique sequence data, and we identified 66,596 SNPs. This data represents one of the largest genetic resources currently available for any cervid. We hope it will facilitate future research for population genomics and assist with the identification of genetic factors that underlie disease resistance and other traits relevant for conservation and management.
Data from: Comparative transcriptomics reveals domestication-associated features of Atlantic salmon lipid metabolism
<p>Domestication of animals imposes strong targeted selection for desired traits but can also result in unintended selection due to new domestic environments. Atlantic salmon was domesticated in the 1970s and has subsequently been selected for faster growth in systematic breeding programmes. More recently, salmon aquaculture has replaced fish oils (FO) with vegetable oils (VO) in feed, radically changing the levels of essential long-chain polyunsaturated fatty acids (LC-PUFA). Our aim was to study the impact of domestication on metabolism and explore the hypothesis that the shift to VO-diets has unintentionally selected for a domestication-specific lipid metabolism. We conducted a 96-day feeding trial of domesticated and wild salmon fed diets based on FO, VO or phospholipids (PL), and compared transcriptomes and fatty acids in tissues involved in lipid absorption (pyloric caeca) and lipid turnover and synthesis (liver). Domesticated salmon had faster growth and higher gene expression in glucose and lipid metabolism compared to wild fish, possibly linked to differences in regulation of circadian rhythm pathways. Only the domesticated salmon increased expression of LC-PUFA synthesis genes when given VO. This transcriptome response difference was mirrored at the physiological level, with domesticated salmon having higher LC-PUFA but lower 18:3n-3 and 18:2n-6 levels. In line with this, the VO diet decreased growth rate in wild but not domesticated salmon. Our study revealed a clear impact of domestication on transcriptomic regulation linked to metabolism and suggests that unintentional selection in the domestic-environment has resulted in evolution of stronger compensatory mechanisms to a diet low in LC-PUFA. </p>
Data from: De novo transcriptome analysis of the common New Zealand stick insect Clitarchus hookeri (Phasmatodea) reveals genes involved in olfaction, digestion and sexual reproduction
Phasmatodea, more commonly known as stick insects, have been poorly studied at the molecular level for several key traits, such as components of the sensory system and regulators of reproduction and development, impeding a deeper understanding of their functional biology. Here, we employ de novo transcriptome analysis to identify genes with primary functions related to female odour reception, digestion, and male sexual traits in the New Zealand common stick insect Clitarchus hookeri (White). The female olfactory gene repertoire revealed ten odorant binding proteins with three recently duplicated, 12 chemosensory proteins, 16 odorant receptors, and 17 ionotropic receptors. The majority of these olfactory genes were over-expressed in female antennae and have the inferred function of odorant reception. Others that were predominantly expressed in male terminalia (n = 3) and female midgut (n = 1) suggest they have a role in sexual reproduction and digestion, respectively. Over-represented transcripts in the midgut were enriched with digestive enzyme gene families. Clitarchus hookeri is likely to harbour nine members of an endogenous cellulase family (glycoside hydrolase family 9), two of which appear to be specific to the C. hookeri lineage. All of these cellulase sequences fall into four main phasmid clades and show gene duplication events occurred early in the diversification of Phasmatodea. In addition, C. hookeri genome is likely to express γ-proteobacteria pectinase transcripts that have recently been shown to be the result of horizontal transfer. We also predicted 711 male terminalia-enriched transcripts that are candidate accessory gland proteins, 28 of which were annotated to have molecular functions of peptidase activity and peptidase inhibitor activity, two groups being widely reported to regulate female reproduction through proteolytic cascades. Our study has yielded new insights into the genetic basis of odour detection, nutrient digestion, and male sexual traits in stick insects. The C. hookeri reference transcriptome, together with identified gene families, provides a comprehensive resource for studying the evolution of sensory perception, digestive systems, and reproductive success in phasmids.
Data from: Comparative ecological transcriptomics and the contribution of gene expression to the evolutionary potential of a threatened fish
Understanding whether small populations with low genetic diversity can respond to rapid environmental change via phenotypic plasticity is an outstanding research question in biology. RNA sequencing (RNA-seq) has recently provided the opportunity to examine variation in gene expression, a surrogate for phenotypic variation, in non-model species. We used a comparative RNA-seq approach to assess expression variation within and among adaptively divergent populations of a threatened freshwater fish, Nannoperca australis, found across a steep hydroclimatic gradient in the Murray-Darling Basin, Australia. These populations evolved under contrasting selective environments (e.g. dry/hot lowland; wet/cold upland) and represent opposite ends of the species' spectrum of genetic diversity and population size. We tested the hypothesis that environmental variation among isolated populations has driven the evolution of divergent expression at ecologically important genes using differential expression (DE) analysis and an ANOVA-based comparative phylogenetic expression variance and evolution model framework based on 27,425 de novo assembled transcripts. Additionally, we tested whether gene expression variance within-populations was correlated with levels of standing genetic diversity. We identified 290 DE candidate transcripts, 33 transcripts with evidence for high expression plasticity, and 50 candidates for divergent selection on gene expression after accounting for phylogenetic structure. Variance in gene expression appeared unrelated to levels of genetic diversity. Functional annotation of the candidate transcripts revealed variation in water quality is an important factor influencing expression variation for N. australis. Our findings suggest that gene expression variation can contribute to the evolutionary potential of small populations.
Data from: Transcriptomics and in vivo tests reveal novel mechanisms underlying endocrine disruption in an ecological sentinel, Nucella lapillus
Anthropogenic endocrine disruptors now contaminate all environments globally, with concomitant deleterious effects across diverse taxa. While most studies on endocrine disruption (ED) have focused on vertebrates, the superimposition of male sexual characteristics in the female dogwhelk, Nucella lapillus (imposex), caused by organotins, provides one of the most clearcut ecological examples of anthropogenically induced ED in aquatic ecosystems. To identify the underpinning mechanisms of imposex for this 'nonmodel' species, we combined Roche 454 pyrosequencing with custom oligoarray fabrication inexpensively to both generate gene models and identify those responding to chronic tributyltin (TBT) treatment. The results supported the involvement of steroid, neuroendocrine peptide hormone dysfunction and retinoid mechanisms, but suggested additionally the involvement of putative peroxisome proliferator–activated receptor (PPAR) pathways. Application of rosiglitazone, a well-known vertebrate PPARγ ligand, to dogwhelks induced imposex in the absence of TBT. Thus, while TBT-induced imposex is linked to the induction of many genes and has a complex phenotype, it is likely also to be driven by PPAR-responsive pathways, hitherto not described in invertebrates. Our findings provide further evidence for a common signalling pathway between invertebrate and vertebrate species that has previously been overlooked in the study of endocrine disruption.
Data from: Scrimer: designing primers from transcriptome data
With the rise of next-generation sequencing methods, it has become increasingly possible to obtain genomewide sequence data even for nonmodel species. Such data are often used for the development of single nucleotide polymorphism (SNP) markers, which can subsequently be screened in a larger population sample using a variety of genotyping techniques. Many of these techniques require appropriate locus-specific PCR and genotyping primers. Currently, there is no publicly available software for the automated design of suitable PCR and genotyping primers from next-generation sequence data. Here we present a pipeline called Scrimer that automates multiple steps, including adaptor removal, read mapping, selection of SNPs and multiple primer design from transcriptome data. The designed primers can be used in conjunction with several widely used genotyping methods such as SNaPshot or MALDI-TOF genotyping. Scrimer is composed of several reusable modules and an interactive bash workflow that connects these modules. Even the basic steps are presented, so the workflow can be executed in a step-by-step manner. The use of standard formats throughout the pipeline allows data from various sources to be plugged in, as well as easy inspection of intermediate results with visualization tools of the user's choice.
Data from: A NGS approach to the encrusting Mediterranean sponge Crella elegans (Porifera, Demospongiae, Poecilosclerida): transcriptome sequencing, characterization and overview of the gene expression along three life cycle stages
Sponges can be dominant organisms in many marine and freshwater habitats where they play essential ecological roles. They also represent a key group to address important questions in early metazoan evolution. Recent approaches for improving knowledge on sponge biological and ecological functions as well as on animal evolution have focused on the genetic toolkits involved in ecological responses to environmental changes (biotic and abiotic), development and reproduction. These approaches are possible thanks to newly available, massive sequencing technologies–such as the Illumina platform, which facilitate genome and transcriptome sequencing in a cost-effective manner. Here we present the first NGS (next-generation sequencing) approach to understanding the life cycle of an encrusting marine sponge. For this we sequenced libraries of three different life cycle stages of the Mediterranean sponge Crella elegans and generated de novo transcriptome assemblies. Three assemblies were based on sponge tissue of a particular life cycle stage, including non-reproductive tissue, tissue with sperm cysts and tissue with larvae. The fourth assembly pooled the data from all three stages. By aggregating data from all the different life cycle stages we obtained a higher total number of contigs, contigs with blast hit and annotated contigs than from one stage-based assemblies. In that multi-stage assembly we obtained a larger number of the developmental regulatory genes known for metazoans than in any other assembly. We also advance the differential expression of selected genes in the three life cycle stages to explore the potential of RNA-seq for improving knowledge on functional processes along the sponge life cycle.
Data from: Temporal transcriptomics suggest that twin-peaking genes reset the clock
The mammalian suprachiasmatic nucleus (SCN) drives daily rhythmic behavior and physiology, yet a detailed understanding of its coordinated transcriptional programmes is lacking. To reveal the finer details of circadian variation in the mammalian SCN transcriptome we combined laser-capture microdissection and RNA-seq over a 24-hour light/dark cycle. We show that 7-times more genes exhibited a classic sinusoidal expression signature than previously observed in the SCN. Another group of 766 genes unexpectedly peaked twice, near both the start and end of the dark phase; this twin-peaking group is significantly enriched for synaptic transmission genes that are crucial for light-induced phase shifting of the circadian clock. 341 intergenic non-coding RNAs, together with novel exons of annotated protein-coding genes, including Cry1, also show specific circadian expression variation. Overall, our data provide an important chronobiological resource (www.wgpembroke.com/shiny/SCNseq/) and allow us to propose that transcriptional timing in the SCN is gating clock resetting mechanisms.
Data from: Transcriptomic signatures of social experience during early development in a highly social cichlid fish
<p><span>The social environment encountered early during development can temporarily or permanently influence life history decisions and behaviour of individuals and correspondingly shape molecular pathways. In the highly social cichlid fish <i>Neolamprologus pulcher,</i> deprivation of brood care permanently affects social behaviour, and alters the expression of stress axis genes in juveniles and adults. It is unclear when gene expression patterns change during early life depending on social experience, and which genes are involved. We compared brain gene expression of <i>N. pulcher </i>at two time points during the social experience phase when juveniles were reared either with or without brood care, and one time point shortly afterwards. We compared (i) whole transcriptomes and (ii) expression of 79 genes related to stress regulation, in order to define a neurogenomic state of stress for each fish. At developmental day 75, that is, after the social experience phase, 43 genes were down-regulated in fish having experienced social deprivation, while two genes involved in learning and memory and in post translational modifications of proteins (PTM), respectively, were up-regulated. Down-regulated genes were mainly associated with immunity, PTM and brain function. In contrast, during the experience phase no genes were differentially expressed when assessing the whole transcriptome. When focusing on the neurogenomic state associated with the stress response, we found that individuals from the two social treatments differed in how their brain gene expression profiles changed over developmental stages. Our results indicate that the early social environment influences the transcriptional activation in fish brains, both during and after an early social experience, possibly affecting plasticity, immune system function and stress axis regulation. </span></p>
Data from: Characterization of the transcriptome, nucleotide sequence polymorphism, and natural selection in the desert adapted mouse Peromyscus eremicus
As a direct result of intense heat and aridity, deserts are thought to be among the most harsh of environments, particularly for their mammalian inhabitants. Given that osmoregulation can be challenging for these animals, with failure resulting in death, strong selection should be observed on genes related to the maintenance of water and solute balance. One such animal, Peromyscus eremicus, is native to the desert regions of the southwest United States and may live its entire life without oral fluid intake. As a first step toward understanding the genetics that underlie this phenotype, we present a characterization of the P. eremicus transcriptome. We assay four tissues (kidney, liver, brain, testes) from a single individual and supplement this with population level renal transcriptome sequencing from 15 additional animals. We identified a set of transcripts undergoing both purifying and balancing selection based on estimates of Tajima's D. In addition, we used the branch-site test to identify a transcript—Slc2a9, likely related to desert osmoregulation—undergoing enhanced selection in P. eremicus relative to a set of related non-desert rodents.
Data from: A priori and a posteriori approaches for finding genes of evolutionary interest in non-model species: osmoregulatory genes in the kidney transcriptome of the desert rodent Dipodomys spectabilis (banner-tailed kangaroo rat)
One common goal in evolutionary biology is the identification of genes underlying adaptive traits of evolutionary interest. Recently next-generation sequencing techniques have greatly facilitated such evolutionary studies in species otherwise depauperate of genomic resources. Kangaroo rats (Dipodomys sp.) serve as exemplars of adaptation in that they inhabit extremely arid environments, yet require no drinking water because of ultra-efficient kidney function and osmoregulation. As a basis for identifying water conservation genes in kangaroo rats, we conducted a priori bioinformatics searches in model rodents (Mus musculus and Rattus norvegicus) to identify candidate genes with known or suspected osmoregulatory function. We then obtained 446,758 reads via 454 pyrosequencing to characterize genes expressed in the kidney of banner-tailed kangaroo rats (Dipodomys spectabilis). We also determined candidates a posteriori by identifying genes that were overexpressed in the kidney. The kangaroo rat sequences revealed nine different a priori candidate genes predicted from our Mus and Rattus searches, as well as 32 a posteriori candidate genes that were overexpressed in kidney. Mutations in two of these genes, Slc12a1 and Slc12a3, cause human renal diseases that result in the inability to concentrate urine. These genes are likely key determinants of physiological water conservation in desert rodents.
Data from: Developmental and transcriptomal responses to seasonal dietary shifts in the cactophilic Drosophila mojavensis of North America
Drosophila mojavensis normally breeds in necrotic columnar cactus, but they also feed and breed in Opuntia fruit (prickly pear) which serves as a seasonal resource. The prickly pear fruits are much different chemically from cacti, mainly in their free sugars and lipid content, raising the question of the effects of this seasonal shift on fitness and on gene expression. Here we reared three isofemale strains of D. mojavensis collected from different parts of the species' range on semi-natural medium of either cactus or prickly pear fruit and measured the development time, survival, body weights and desiccation resistance. All these parameters were affected by diet and by interaction with strain and or sex. Interestingly, however, there appear to be tradeoffs: flies developed faster in prickly pear and the emerging adults were heavier, but those having grown in cactus were more resistant to desiccation. We also evaluated gene expression of emerging male and female adult flies using RNA-Seq. While more genes were down-regulated in prickly pear fruit than up-regulated in both sexes, the sexes did differ in expression patterns. The majority of the genes that were preferentially expressed comparing prickly pear fruit vs cactus underlie metabolism. Genes involved with carbohydrate and lipid metabolism, as well as with the amino acid serine, and their relationship to growth and development reflect the ways in which these dietary differences affect the flies.
Data from: De novo assembly of the transcriptome of an invasive snail and its multiple ecological applications
Studying how invasive species respond to environmental stress at the molecular level can help us assess their impact and predict their range expansion. Development of markers of genetic polymorphism can help us reconstruct their invasive route. However, to conduct such studies requires the presence of substantial amount of genomic resources. This study aimed to generate and characterize genomic resources using high throughput transcriptome sequencing for Pomacea canaliculata, a non-model gastropod indigenous to Argentina that has invaded Asia, Hawaii and southern United States. De novo assembly of the transcriptome resulted in 128,436 unigenes with an average length of 419 bp (range: 150 to 8,556 bp). Many of the unigenes (2,439) contained transposable elements, showing the existence of a source of genetic variability in response to stressful conditions. A total of 3,196 microsatellites were detected in the transcriptome; among 20 of the randomly tested microsatellites, 10 were validated to exhibit polymorphism. A total of 15,412 single-nucleotide polymorphisms (SNPs) were detected in the ORFs. LC-MS/MS analysis of the proteome of juveniles revealed 878 proteins, of which many are stress-related. This study has demonstrated the great potential of high throughput DNA sequencing for rapid development of genomic resources for a non-model organism. Such resources can facilitate various molecular ecological studies, such as stress physiology and range expansion.
Data from: Long read reference genome-free reconstruction of a full-length transcriptome from Astragalus membranaceus reveals transcript variants involved in bioactive compound biosynthesis
Astragalus membranaceus, also known as Huangqi in China, is one of the most widely used medicinal herbs in Traditional Chinese Medicine. Traditional Chinese Medicine formulations from Astragalus membranaceus have been used to treat a wide range of illnesses, such as cardiovascular disease, type 2 diabetes, nephritis and cancers. Pharmacological studies have shown that immunomodulating, anti-hyperglycemic, anti-inflammatory, antioxidant and antiviral activities exist in the extract of Astragalus membranaceus. Therefore, characterising the biosynthesis of bioactive compounds in Astragalus membranaceus, such as Astragalosides, Calycosin and Calycosin-7-O-β-D-glucoside, is of particular importance for further genetic studies of Astragalus membranaceus. In this study, we reconstructed the Astragalus membranaceus full-length transcriptomes from leaf and root tissues using PacBio Iso-Seq long reads. We identified 27 975 and 22 343 full-length unique transcript models in each tissue respectively. Compared with previous studies that used short read sequencing, our reconstructed transcripts are longer, and are more likely to be full-length and include numerous transcript variants. Moreover, we also re-characterised and identified potential transcript variants of genes involved in Astragalosides, Calycosin and Calycosin-7-O-β-D-glucoside biosynthesis. In conclusion, our study provides a practical pipeline to characterise the full-length transcriptome for species without a reference genome and a useful genomic resource for exploring the biosynthesis of active compounds in Astragalus membranaceus.
Data from: Signatures of rapid evolution in urban and rural transcriptomes of white-footed mice (Peromyscus leucopus) in the New York metropolitan area
Urbanization is a major cause of ecological degradation around the world, and human settlement in large cities is accelerating. New York City (NYC) is one of the oldest and most urbanized cities in North America, but still maintains 20% vegetation cover and substantial populations of some native wildlife. The white-footed mouse, Peromyscus leucopus, is a common resident of NYC's forest fragments and an emerging model system for examining the evolutionary consequences of urbanization. In this study, we developed transcriptomic resources for urban P. leucopus to examine evolutionary changes in protein-coding regions for an exemplar 'urban adapter'. We used Roche 454 GS FLX+ high throughput sequencing to derive transcriptomes from multiple tissues from individuals across both urban and rural populations. From these data, we identified 31,015 SNPs and several candidate genes potentially experiencing positive selection in urban populations of P. leucopus. These candidate genes are involved in xenobiotic metabolism, innate immune response, demethylation activity, and other important biological phenomena in novel urban environments. This study is one of the first to report candidate genes exhibiting signatures of directional selection in divergent urban ecosystems.
Data from: Dissecting molecular evolution in the highly diverse plant clade Caryophyllales using transcriptome sequencing
Many phylogenomic studies based on transcriptomes have been limited to "single-copy" genes due to methodological challenges in homology and orthology inferences. Only a relatively small number of studies have explored analyses beyond reconstructing species relationships. We sampled 69 transcriptomes in the hyperdiverse plant clade Caryophyllales and 27 outgroups from annotated genomes across eudicots. Using a combined similarity- and phylogenetic tree-based approach, we recovered 10,960 homolog groups, where each was represented by at least eight ingroup taxa. By decomposing these homolog trees, and taking gene duplications into account, we obtained 17,273 ortholog groups, where each was represented by at least ten ingroup taxa. We reconstructed the species phylogeny using a 1,122-gene data set with a gene occupancy of 92.1%. From the homolog trees, we found that both synonymous and nonsynonymous substitution rates in herbaceous lineages are up to three times as fast as in their woody relatives. This is the first time such a pattern has been shown across thousands of nuclear genes with dense taxon sampling. We also pinpointed regions of the Caryophyllales tree that were characterized by relatively high frequencies of gene duplication, including three previously unrecognized whole-genome duplications. By further combining information from homolog tree topology and synonymous distance between paralog pairs, phylogenetic locations for 13 putative genome duplication events were identified. Genes that experienced the greatest gene family expansion were concentrated among those involved in signal transduction and oxidoreduction, including a cytochrome P450 gene that encodes a key enzyme in the betalain synthesis pathway. Our approach demonstrates a new approach for functional phylogenomic analysis in nonmodel species that is based on homolog groups in addition to inferred ortholog groups.
Data from: Blood transcriptomes and de novo identification of candidate loci for mating success in lekking great snipe (Gallinago media)
We assembled the great snipe blood transcriptome using data from fourteen lekking males, in order to de novo identify candidate genes related to sexual selection, and determined the expression profiles in relation to mating success. The three most highly transcribed genes were encoding different haemoglobin subunits. All tended to be overexpressed in males with high mating success. We also called Single Nucleotide Polymorphisms (SNPs) from the transcriptome data and found considerable genetic variation for many genes expressed during lekking. Among these we identified 14 polymorphic candidate SNPs that had a significant genotypic association with mating success (number of females mated with) and/or mating status (mated or not). Four of the candidate SNPs were found in HBAA (encoding the haemoglobin α-chain). Heterozygotes for one of these and one SNP in the gene PABPC1 appeared to enjoy higher mating success compared to males homozygous for either of the alleles. In a larger dataset of individuals we genotyped 38 of the identified SNPs but found low support for consistent selection since only one of the zygosities of previously identified candidate SNPs and none of their genotypes were associated with mating status. However, candidate SNPs generally showed lower levels of spatial genetic structure compared to non-candidate markers. We also scored the prevalence of avian malaria in a sub-sample of birds. Males infected with avian malaria parasites had lower mating success in the year of sampling than non-infected males. Parasite infection and its interaction with specific genes may thus affect performance on the lek.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.