Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,715
datasets available to search
ShareScore release 0.7.1
Dataset results
1,715 results for “Arabidopsis thaliana; Arabidopsis”
Data from: The roles of genetic drift and natural selection in quantitative trait divergence along an altitudinal gradient in Arabidopsis thaliana
Open the record for dataset details and reuse information.
Data from: Fitness of Arabidopsis thaliana mutation accumulation lines whose spontaneous mutations are known
Open the record for dataset details and reuse information.
Phenotype pictures of Arabidopsis thaliana in high light and low light conditions
Open the record for dataset details and reuse information.
cDNA sequence of E2 gene family in Arabidopsis thaliana and data of statistical analysis
Open the record for dataset details and reuse information.
Data from: Heterosis and outbreeding depression in crosses between natural populations of Arabidopsis thaliana
Open the record for dataset details and reuse information.
The cis-regulatory codes of response to combined heat and drought stress in Arabidopsis thaliana
<p>Datasets used to train and test random forest and convolutional neural networks to predict transcriptional response patterns to single and combined heat and drought stress in Arabidopsis. Rows correspond to genes. The first column denotes the class with 1 indicating response group and 0 indicating non-responsive group (e.g. "NNU_merged_df.txt": 1 = NNU, 0 = NNN). The remaining columns are the pCRE and pCRE-omic overlap features, where "1" denotes the pCRE is present in the promoter region of that gene (or present and overlaping with the omic-feature) and "0" denotes the pCRE is not present (or present but not overlapping with the omic-feature). Feature names indicate the pCRE and omic-feature: "pCRE_OmicFeature" </p> <p><strong>For more information on how these datasets were generated and code used to implement and interpret the machine learning models see the manuscript and associated GitHub repository.</strong></p> <p>GitHub: <a href="https://github.com/ShiuLab/Manuscript_Code/tree/master/2019_CRC_HeatDrought">https://github.com/ShiuLab/Manuscript_Code/tree/master/2019_CRC_HeatDrought</a></p> <p>Abstract: Plants respond to their environment by dynamically modulating gene expression. A powerful approach for understanding how these responses are regulated is to integrate information about <em>cis-</em>regulatory elements (CREs) into models called <em>cis-</em>regulatory codes. Transcriptional response to combined stress is typically not the sum of the responses to the individual stresses. However, <em>cis-</em>regulatory codes underlying combined stress response have not been established. Here we modeled transcriptional response to single and combined heat and drought stress in <em>Arabidopsis thaliana.</em> We grouped genes by their pattern of response (independent, antagonistic, synergistic) and trained machine learning models to predict their response using putative CREs (pCREs) as features (median F-measure = 0.64). We then developed a deep learning approach to integrate additional omics information (sequence conservation, chromatin accessibility, histone modification) into our models, improving performance by 6.2%. While pCREs important for predicting independent and antagonistic responses tended to resemble binding motifs of transcription factors associated with heat and/or drought stress, important synergistic pCREs resembled binding motifs of transcription factors not known to be associated with stress. These findings demonstrate how <em>in silico</em> approaches can improve our understanding of the complex codes regulating response to combined stress and help us identify prime targets for future characterization.</p>
Data from: Adaptive reduction of male gamete number in the selfing plant Arabidopsis thaliana
<p class="AbstractSummary">The number of male gametes is critical for reproductive success and varies between and within species. The evolutionary reduction of the number of pollen grains encompassing the male gametes is widespread in selfing plants. Here, we employ genome-wide association study (GWAS) to identify underlying loci and to assess the molecular signatures of selection on pollen number-associated loci in the predominantly selfing plant <i>Arabidopsis thaliana</i>. Regions of strong association with pollen number are enriched for signatures of selection, indicating polygenic selection. We isolate the gene <i>REDUCED POLLEN NUMBER1 </i>(<i>RDP1</i>) at the locus with the strongest association. We validate its effect using a quantitative complementation test with CRISPR/Cas9-generated null mutants in nonstandard wild accessions. In contrast to pleiotropic null mutants, only pollen numbers are significantly affected by natural allelic variants. These data support theoretical predictions that reduced investment in male gametes is advantageous in predominantly selfing species.</p>
The influence of experimentally induced polyploidy on the relationships between endopolyploidy and plant function in Arabidopsis thaliana
<p>Whole genome duplication, leading to polyploidy and endopolyploidy, occurs in all domains and kingdoms and is especially prevalent in vascular plants. Both polyploidy and endopolyploidy increase cell size, but it is uncertain whether both processes have similar effects on plant morphology and function, or whether polyploidy influences the magnitude of endopolyploidy. To address these gaps in knowledge, fifty-five geographically separated diploid genotypes (i.e., accessions) of <i>Arabidopsis thaliana</i> that span a gradient of endopolyploidy were experimentally manipulated to induce polyploidy. Both the diploids and artificially induced tetraploids were grown in a common greenhouse environment and evaluated with respect to nine reproductive and vegetative characteristics. Induced polyploidy decreased leaf endopolyploidy and stem endopolyploidy along with specific leaf area, stem height, but increased days to bolting, leaf size, leaf dry mass and leaf water content. Phenotypic responses to induced polyploidy varied significantly among genotypes but this did not affect the relationship between phenotypic traits and endopolyploidy. Our results provide the experimental support for a trade-off between induced polyploidy and endopolyploidy, which caused induced polyploids to have lower endopolyploidy than diploids. Though polyploidy did not influence the relationship between endopolyploidy and plant traits, phenotypic responses to experimental genome duplication could not be easily predicted because of strong cytotype by genotype interactions.</p>
Evolution of conserved noncoding sequences in Arabidopsis thaliana
<p>Recent pangenome studies have revealed that a large fraction (>20%) of the gene content within a species exhibits presence-absence variation (PAV). However, coding regions alone provide an incomplete assessment of functional genomic sequence variation at the species level. Little to no attention has been paid to noncoding regulatory regions in pangenome studies, though these sequences directly modulate gene expression and phenotype. To uncover regulatory genetic variation, we generated chromosome-scale genome assemblies for thirty Arabidopsis thaliana accessions from multiple distinct habitats and characterized species level variation in Conserved Noncoding Sequences (CNS). Our analyses uncovered not only evidence for PAV and Positional Variation (PosV) but that diversity in CNS is non-random, with variants shared across different accessions. Using evolutionary analyses and chromatin accessibility data, we provide further evidence supporting conserved and variable CNS roles in gene regulation. Characterizing species-level diversity in all functional genomic sequences may later uncover previously unknown mechanistic links between genotype and phenotype.<span> </span></p>
Data from: Meta-analysis of Arabidopsis thaliana phospho-proteomics data reveals compartmentalization of phosphorylation motifs
Protein (de)phosphorylation plays an important role in plants. To provide a robust foundation for subcellular phosphorylation signaling network analysis and kinase-substrate relationships, we performed a meta-analysis of 27 published and unpublished in-house mass spectrometry–based phospho-proteome data sets for Arabidopsis thaliana covering a range of processes, (non)photosynthetic tissue types, and cell cultures. This resulted in an assembly of 60,366 phospho-peptides matching to 8141 nonredundant proteins. Filtering the data for quality and consistency generated a set of medium and a set of high confidence phospho-proteins and their assigned phospho-sites. The relation between single and multiphosphorylated peptides is discussed. The distribution of p-proteins across cellular functions and subcellular compartments was determined and showed overrepresentation of protein kinases. Extensive differences in frequency of pY were found between individual studies due to proteomics and mass spectrometry workflows. Interestingly, pY was underrepresented in peroxisomes but overrepresented in mitochondria. Using motif-finding algorithms motif-x and MMFPh at high stringency, we identified compartmentalization of phosphorylation motifs likely reflecting localized kinase activity. The filtering of the data assembly improved signal/noise ratio for such motifs. Identified motifs were linked to kinases through (bioinformatic) enrichment analysis. This study also provides insight into the challenges/pitfalls of using large-scale phospho-proteomic data sets to nonexperts.
Data from: Different gene families in Arabidopsis thaliana transposed in different epochs and at different frequencies throughout the rosids
Certain types of gene families, such as those encoding most families of transcription factors, maintain their chromosomal syntenic positions throughout Angiosperm evolutionary time. Other, non-syntenic gene families are prone to deletion, tandem duplication, and transposition. Here we describe the chromosomal positional history of all genes in Arabidopsis thaliana (A. thaliana) throughout the rosid superorder. We introduce a public database where researchers can look up the positional history of their favorite A. thaliana gene or gene family. Finally, we show that specific gene families transposed at specific points in evolutionary time, particularly after whole-genome duplication events in the Brassicales, and suggest that genes in mobile gene families are under different selection pressure than syntenic genes.
Data from: Genomic and phenotypic differentiation of Arabidopsis thaliana along altitudinal gradients in the North Italian Alps
Altitudinal gradients in mountain regions are short-range clines of different environmental parameters such as temperature or radiation. We investigated genomic and phenotypic signatures of adaptation to such gradients in five Arabidopsis thaliana populations from the North Italian Alps that originated from 580 to 2350 m altitude by resequencing pools of 19–29 individuals from each population. The sample includes two pairs of low- and high-altitude populations from two different valleys. High-altitude populations showed a lower nucleotide diversity and negative Tajima's D values and were more closely related to each other than to low-altitude populations from the same valley. Despite their close geographic proximity, demographic analysis revealed that low- and high-altitude populations split between 260 000 and 15 000 years before present. Single nucleotide polymorphisms whose allele frequencies were highly differentiated between low- and high-altitude populations identified genomic regions of up to 50 kb length where patterns of genetic diversity are consistent with signatures of local selective sweeps. These regions harbour multiple genes involved in stress response. Variation among populations in two putative adaptive phenotypic traits, frost tolerance and response to light/UV stress was not correlated with altitude. Taken together, the spatial distribution of genetic diversity reflects a potentially adaptive differentiation between low- and high-altitude populations, whereas the phenotypic differentiation in the two traits investigated does not. It may resemble an interaction between adaptation to the local microhabitat and demographic history influenced by historical glaciation cycles, recent seed dispersal and genetic drift in local populations.
Data from: Genome-wide association study in Arabidopsis thaliana of natural variation in seed oil melting point, a widespread adaptive trait in plants
Seed oil melting point is an adaptive, quantitative trait determined by the relative proportions of the fatty acids that compose the oil. Micro- and macro-evolutionary evidence suggests selection has changed the melting point of seed oils to covary with germination temperatures because of a trade-off between total energy stores and the rate of energy acquisition during germination under competition. The seed oil compositions of 391 natural accessions of Arabidopsis thaliana, grown under common-garden conditions, were used to assess whether seed oil melting point within a species varied with germination temperature. In support of the adaptive explanation, long-term monthly spring and fall field temperatures of the accession collection sites significantly predicted their seed oil melting points. In addition, a genome-wide association study (GWAS) was performed to determine which genes were most likely responsible for the natural variation in seed oil melting point. The GWAS found a single highly significant association within the coding region of FAD2, which encodes a fatty acid desaturase central to the oil biosynthesis pathway. In a separate analysis of fifteen a priori oil synthesis candidate genes, two (FAD2 and FATB) were located near significant SNPs associated with seed oil melting point. These results comport with others' molecular work showing that lines with alterations in these genes affect seed oil melting point as expected. Our results suggest natural selection has acted on a small number of loci to alter a quantitative trait in response to local environmental conditions.
Data from: Effects of multi-generational stress exposure and offspring environment on the expression and persistence of transgenerational effects in Arabidopsis thaliana
Plant phenotypes can be affected by environments experienced by their parents. Parental environmental effects are reported for the first offspring generation and some studies showed persisting environmental effects in second and further offspring generations. However, the expression of these transgenerational effects proved context-dependent and their reproducibility can be low. Here we study the context-dependency of transgenerational effects by evaluating parental and transgenerational effects under a range of parental induction and offspring evaluation conditions. We systematically evaluated two factors that can influence the expression of transgenerational effects: single- versus multiple-generation exposure and offspring environment. For this purpose, we exposed a single homozygous Arabidopsis thaliana Col-0 line to salt stress for up to three generations and evaluated offspring performance under control and salt conditions in a climate chamber and in a natural environment. Parental as well as transgenerational effects were observed in almost all traits and all environments and traced back as far as great-grandparental environments. The length of exposure exerted strong effects; multiple-generation exposure often reduced the expression of the parental effect compared to single-generation exposure. Furthermore, the expression of transgenerational effects strongly depended on offspring environment for rosette diameter and flowering time, with opposite effects observed in field and greenhouse evaluation environments. Our results provide important new insights into the occurrence of transgenerational effects and contribute to a better understanding of the context-dependency of these effects.
Data from: Transmission ratio distortion is frequent in Arabidopsis thaliana controlled crosses
The equal probability of transmission of alleles from either parent during sexual reproduction is a central tenet of genetics and evolutionary biology. Yet, there are many cases where this rule is violated. The preferential transmission of alleles or genotypes is termed transmission ratio distortion (TRD). Examples of TRD have been identified in many species, implying that they are universal, but the resolution of species-wide studies of TRD are limited. We have performed a species-wide screen for TRD in over 500 segregating F2 populations of Arabidopsis thaliana using pooled reduced-representation genome sequencing. TRD was evident in up to a quarter of surveyed populations. Most populations exhibited distortion at only one genomic region, with some regions being repeatedly affected in multiple populations. Our results begin to elucidate the species-level architecture of biased transmission of genetic material in A. thaliana, and serve as a springboard for future studies into the biological basis of TRD in this species.
Data from: Transcriptome analysis indicates considerable divergence in alternative splicing between duplicated genes in Arabidopsis thaliana
Gene and genome duplication events have created a large number of new genes in plants that can diverge by evolving new expression profiles and functions (neofunctionalization) or dividing extant ones (subfunctionalization). Alternative splicing (AS) generates multiple types of mRNA from a single type of pre-mRNA by differential intron splicing. It can result in new protein isoforms or down-regulation of gene expression by transcript decay. Using RNA-seq we investigated the degree to which alternative splicing patterns are conserved between duplicated genes in Arabidopsis thaliana. Our results revealed that 30% of AS events in alpha whole genome duplicates, and 33% of AS events in tandem duplicates, are qualitatively conserved within leaf tissue. Loss of ancestral splice forms, as well as asymmetric gain of new splice forms, may account for this divergence. Conserved events had different frequencies, as only 31% of shared AS events in alpha whole genome duplicates and 41% of shared AS events in tandem duplicates had similar frequencies in both paralogs, indicating considerable quantitative divergence. Analysis of published RNA-seq data from nonsense mediated decay (NMD) mutants indicated that 85% of alpha whole genome duplicates and 89% of tandem duplicates have diverged in their AS-induced NMD. Our results indicate that alternative splicing shows a high degree of divergence between paralogs such that qualitatively conserved alternative splicing events tend to have quantitative divergence. Divergence in AS patterns between duplicates may be a mechanism of regulating expression level divergence.
Data from: The recombination landscape in Arabidopsis thaliana F2 populations
Recombination during meiosis shapes the complement of alleles segregating in the progeny of hybrids, and has important consequences for phenotypic variation. We examined allele frequencies as well as crossover locations and frequencies in over 7000 plants from 17 F2 populations derived from crosses between 18 Arabidopsis thaliana accessions. We observe segregation distortion between parental alleles in over half of our populations. The potential causes of distortion include variation in seed dormancy and lethal epistatic interactions. Such a high occurrence of distortion was only detected here because of the large sample size of each population, and the number of populations characterized. Most plants carry only one or two crossovers per chromosome pair, and therefore inherit very large, non-recombined genomic fragments from each parent. Recombination frequencies vary between populations but consistently increase adjacent to the centromeres. Importantly, recombination rates do not correlate with whole-genome sequence differences between parental accessions, suggesting that sequence diversity within A. thaliana does not normally reach levels that are high enough to exert a major influence on the formation of crossovers. A global knowledge of the patterns of recombination in F2 populations is crucial to better understand the segregation of phenotypic traits in hybrids, in the laboratory or in the wild.
Data from: Bacillus cereus AR156 activates defense responses to Pseudomonas syringae pv. tomato in Arabidopsis thaliana similarly to flg22
Bacillus cereus AR156 (AR156) is a plant growth promoting rhizobacterium capable of inducing systemic resistance to Pseudomonas syringae pv. tomato (Pst) in Arabidopsis thaliana (Arabidopsis). Here we show that when applied to Arabidopsis leaves, AR156 acted similarly to flg22, a typical pathogen-associated molecular pattern (PAMP), in initiating PAMP-triggered immunity (PTI). AR156-elicited PTI responses included phosphorylation of MPK3 and MPK6, induction of the expression of defense-related genes PR1, FRK1, WRKY22, and WRKY29; production of reactive oxygen species; and callose deposition. Pretreatment with AR156 still significantly reduced Pst multiplication and disease severity in NahG transgenic plants and mutants sid2-2, jar1, etr1, ein2, npr1, and fls2. This suggests that AR156-induced PTI responses require neither salicylic acid, jasmonic acid, and ethylene signaling; nor flagella receptor kinase FLS2, the receptor of flg22. On the other hand, AR156 and flg22 acted in concert to differentially regulate a number of AGO1-bound miRNAs that function to mediate PTI. A full-genome transcriptional profiling analysis indicated that AR156 and flg22 activated similar transcriptional programs, co-regulating the expression of 117 genes; their concerted regulation of 16 genes was confirmed by real-time quantitative PCR analysis. These results suggest that AR156 activates basal defense responses to Pst in Arabidopsis similarly to flg22.
Data from: Adaptation to warmer climates by parallel functional evolution of CBF genes in Arabidopsis thaliana
The evolutionary processes and genetics underlying local adaptation at a specieswide level are largely unknown. Recent work has indicated that a frameshift mutation in a member of a family of transcription factors, C-repeat binding factors or CBFs, underlies local adaptation and freezing tolerance divergence between two European populations of Arabidopsis thaliana. To ask whether the specieswide evolution of CBF genes in Arabidopsis is consistent with local adaptation, we surveyed CBF variation from 477 wild accessions collected across the species' range. We found that CBF sequence variation is strongly associated with winter temperature variables. Looking specifically at the minimum temperature experienced during the coldest month, we found that Arabidopsis from warmer climates exhibit a significant excess of nonsynonymous polymorphisms in CBF genes and revealed a CBF haplotype network whose structure points to multiple independent transitions to warmer climates. We also identified a number of newly described mutations of significant functional effect in CBF genes, similar to the frameshift mutation previously indicated to be locally adaptive in Italy, and find that they are significantly associated with warm winters. Lastly, we uncover relationships between climate and the position of significant functional effect mutations between and within CBF paralogs, suggesting variation in adaptive function of different mutations. Cumulatively, these findings support the hypothesis that disruption of CBF gene function is adaptive in warmer climates, and illustrate how parallel evolution in a transcription factor can underlie adaptation to climate.
Data from: Fitness benefits and costs of cold acclimation in Arabidopsis thaliana
When resources are limited, there is a tradeoff between growth/reproduction and stress defense in plants. Most temperate plant species, including Arabidopsis thaliana, can enhance freezing tolerance through cold acclimation at low but non-freezing temperatures. Induction of the cold acclimation pathway should be beneficial in environments where plants frequently encounter freezing stress, but might represent a cost in environments where freezing events are rare. In A. thaliana, induction of the cold acclimation pathway critically involves a small subfamily of genes known as the CBFs. Here, we test for a cost of cold acclimation by utilizing (1) natural accessions of A. thaliana that originate from different regions of the species native range and that have experienced different patterns of historical selection on their CBF genes, and (2) transgenic CBF over-expression and T-DNA insertion (knockdown/knockout) lines. While benefits of cold acclimation in the presence of freezing stress were confirmed, no cost of cold acclimation was detected in the absence of freezing stress. These findings suggest that cold acclimation is unlikely to be selected against in warmer environments, and that naturally occurring mutations disrupting CBF function in the southern part of the species range are likely to be selectively neutral. An unanticipated finding was that cold acclimation, in the absence of a subsequent freezing stress, resulted in increased fruit production, i.e., fitness.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.