Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
229
datasets available to search
ShareScore release 0.7.1
Dataset results
229 results for “plant genomics”
Data from: Genome-wide search for quantitative trait loci controlling important plant and flower traits in petunia using an interspecific recombinant inbred population of Petunia axillaris and Petunia exserta
Open the record for dataset details and reuse information.
Data from: Population genomic signatures of founding events in autonomously self-fertilizing plants: A test with <em>Impatiens capensis</em>
Open the record for dataset details and reuse information.
Data from: Population genomics and demographic sampling of the ant-plant Vachellia drepanolobium and its symbiotic ants from sites across its range in East Africa.
Open the record for dataset details and reuse information.
Data From: Evolution of woody plants to the land‐sea interface: The atypical genomic features of mangroves with atypical phenotypic adaptation
Open the record for dataset details and reuse information.
Data for: Evolution and genomic basis of the plant-penetrating ovipositor: a key morphological trait in herbivorous Drosophilidae
Open the record for dataset details and reuse information.
Supplementary tables S5, S7, S9, S10, original protein models fasta files used for alignments, aligned and manually curated protein modes files used for phylogenies (PHYLIP format), and phylogenetic trees of plant cell wall decomposition gene families from 44 basidiomycete genomes (.tre files)
<p><span><span><span><span><span><span><span><span><span><span><span>Litter-decomposing Agaricales play key role in terrestrial carbon cycling, but little is known about their decomposition mechanisms. We assembled datasets of 42 gene families involved in plant-cell-wall decomposition from seven newly sequenced litter decomposers and 35 other Agaricomycotina members, mostly white-rot and brown-rot species. Using sequence similarity and phylogenetics, we split the families into phylogroups and compared their gene composition across nutritional strategies. Subsequently, we used Raman spectroscopy to examine the ability of litter decomposers, white-rot fungi, and brown-rot fungi to decompose crystalline cellulose. Both litter decomposers and white-rot fungi share the enzymatic cellulose decomposition, whereas brown-rot fungi possess a distinct mechanism that disrupts cellulose crystallinity. However, litter decomposers and white-rot fungi differ with respect to hemicellulose and lignin degradation phylogroups, suggesting adaptation of the former group to the litter environment. Litter decomposers show high phylogroup diversity, which is indicative of high functional versatility within the group, whereas a set of white-rot species shows adaptation to bulk-wood decomposition. In both groups, we detected species that have unique characteristics associated with hitherto unknown adaptations to diverse wood and litter substrates. Our results suggest that the terms white-rot fungi and litter decomposers mask a much larger functional diversity.</span></span></span></span></span></span></span></span></span></span></span></p>
Data from: From algae to angiosperms–inferring the phylogeny of green plants (Viridiplantae) from 360 plastid genomes
Background: Next-generation sequencing has provided a wealth of plastid genome sequence data from an increasingly diverse set of green plants (Viridiplantae). Although these data have been useful for reconstructing the phylogeny of numerous clades of photosynthetic organisms (e.g., green algae, angiosperms, and gymnosperms), their utility for inferring relationships across all green plants is uncertain. Viridiplantae originated 700-1500 million years ago and may comprise as many as 500,000 species. This clade represents a major source of photosynthetic carbon and contains an immense diversity of life forms, including some of the smallest and largest eukaryotes. Here we explore the limits and challenges of inferring a comprehensive green plant phylogeny from available complete or nearly complete plastid genome data. Results: We assembled protein-coding sequence data for 78 genes from 360 diverse green plant taxa with complete or nearly complete plastid genome sequences available from GenBank. Phylogenetic analyses of the plastid data recovered well-supported backbone relationships and strong support for relationships that were not observed in previous analyses of major subclades within Viridiplantae. However, there also is evidence of systematic error in some analyses. In several instances we obtained strongly supported but conflicting topologies from analyses of nucleotides versus amino acid characters, and the considerable variation in GC content among lineages and within single genomes affected the phylogenetic placement of several taxa. Conclusions: Analyses of the plastid data recovered a strongly supported framework of relationships for green plants. This includes the placement of Zygnematophyceace as sister to land plants (Embryophyta) and a clade of extant gymnosperms (Acrogymnospermae) with cycads + Ginkgo sister to remaining members and with gnetophytes (Gnetophyta) sister to non-Pinaceae conifers (Gnecup trees); within the monilophyte clade (Monilophyta), relationships are strongly supported with Equisetales + Psilotales sister to Marattiales + leptosporangiate ferns. We also highlight the challenges of using plastid genome sequences in deep-level phylogenomic analyses and provide suggestions for future analyses that will likely incorporate plastid genome data for thousands of species. We particularly emphasize the importance of exploring the effects of different partitioning and character coding protocols for the entire data set as well as subsets of the data.
Data from: How do cold-adapted plants respond to climatic cycles? interglacial expansion explains current distribution and genomic diversity in Primula farinosa L.
Understanding the effects of past climatic fluctuations on the distribution and population-size dynamics of cold-adapted species is essential for predicting their responses to ongoing global climate change. In spite of the heterogeneity of cold-adapted species, two main contrasting hypotheses have been proposed to explain their responses to Late Quaternary glacial cycles, namely, the interglacial contraction versus the interglacial expansion hypotheses. Here, we use the cold-adapted plant Primula farinosa to test two demographic models under each of the two alternative hypotheses and a fifth, null model. We first approximate the time and extent of demographic contractions and expansions during the Late Quaternary by projecting species distribution models across the last 72 ka. We also generate genome-wide sequence data using a Reduced Representation Library approach to reconstruct the spatial structure, genetic diversity, and phylogenetic relationships of lineages within P. farinosa. Finally, by integrating the results of climatic and genomic analyses in an Approximate Bayesian Computation framework, we propose the most likely model for the extent and direction of population-size changes in P. farinosa through the Late Quaternary. Our results support the interglacial expansion of P. farinosa, differing from the prevailing paradigm that the observed distribution of cold-adapted species currently fragmented in high altitude and latitude regions reflects the consequences of postglacial contraction processes.
Data from: Genome of the pitcher plant Cephalotus reveals genetic changes associated with carnivory
Carnivorous plants exploit animals as a nutritional source and have inspired long-standing questions about the origin and evolution of carnivory-related traits. To investigate the molecular bases of carnivory, we sequenced the genome of the heterophyllous pitcher plant Cephalotus follicularis, in which we succeeded in regulating the developmental switch between carnivorous and non-carnivorous leaves. Transcriptome comparison of the two leaf types and gene repertoire analysis identified genetic changes associated with prey attraction, capture, digestion and nutrient absorption. Analysis of digestive fluid proteins from C. follicularis and three other carnivorous plants with independent carnivorous origins revealed repeated co-options of stress-responsive protein lineages coupled with convergent amino acid substitutions to acquire digestive physiology. These results imply constraints on the available routes to evolve plant carnivory.
Data from: Horizontal gene acquisitions, mobile element proliferation, and genome decay in the host - restricted plant pathogen Erwinia tracheiphila
Modern industrial agriculture depends on high density cultivation of genetically similar crop plants, creating favorable conditions for the emergence of novel pathogens with increased fitness in managed compared to ecologically intact settings. Here, we present the genome sequence of six strains of the cucurbit bacterial wilt pathogen Erwinia tracheiphila (Enterobacteriaceae) isolated from infected squash plants in New York, Pennsylvania, Kentucky, and Michigan. These genomes exhibit a high proportion of recent horizontal gene acquisitions, invasion and remarkable amplification of mobile genetic elements, and pseudogenization of ~20% of the coding sequences. These genome attributes indicate that E. tracheiphila recently emerged as a host-restricted pathogen. Furthermore, chromosomal rearrangements associated with phage and transposable element proliferation contributes to substantial differences in gene content and genetic architecture between the six E. tracheiphila strains and other Erwinia species. Together, these data lead us to hypothesize that E. tracheiphila has undergone recent evolution via both genome decay (pseudogenization) and genome expansion (horizontal gene transfer and mobile element amplification). Despite evidence of dramatic genomic changes, the six strains are genetically monomorphic, suggesting a recent population bottleneck and emergence into E. tracheiphila's current ecological niche.
Data from: Genomic regions repeatedly involved in divergence among plant-specialized pea aphid biotypes
Understanding the genetic bases of biological diversification is a long-standing goal in evolutionary biology. Here we investigate whether replicated cases of adaptive divergence involve the same genomic regions in the pea aphid, Acyrthosiphon pisum, a large complex of genetically differentiated biotypes, each specialized on different species of legumes. A previous study identified genomic regions putatively involved in host-plant adaptation and/or reproductive isolation by performing a hierarchical genome scan in three biotypes. This led to the identification of 11 FST outliers among 390 polymorphic microsatellite markers. In this study, the outlier status of these 11 loci was assessed in eight biotypes specialized on other host plants. Four of the 11 previously identified outliers showed greater genetic differentiation among these additional biotypes than expected under the null hypothesis of neutral evolution (α<0.01). Whether these hotspots of genomic divergence result from adaptive events, intrinsic barriers or reduced recombination is discussed.
Input and output data and code for PPS2 on 333 vertebrate and111 plant genome assemblies, and for DDS2+ on fungal genome assemblies
<p>The datasets contain PPS2 and DDS2+ source and binary code, other scripts, and some input and output data. Please see the README file after unpacking it. The absolute paths in the scripts need to be modified in order to duplicate the results in this dataset. Most of the input and output data have to be removed from the datasets for quick uploading and downloading; otherwise, the datasets would exceed the 50-Gb size limit.</p>
NLRome dataset from 124 genomes of plants in the Solanaceae family
<p>We report a dataset of 66,665 NLR immune receptor sequences extracted using NLRtracker from the proteomes of 124 genomes of plants in the Solanaceae family. These include samples from the following genera: <i>Solanum</i> (108 genomes, 30 species), <i>Physalis</i> (1, 1), <i>Capsicum</i> (7, 3), and <i>Nicotiana</i> (8, 5). This dataset includes NLR protein sequences and their domain annotations. The dataset is made available to the community as an open access resource as part of the OpenPlantNLR initiative.</p>
Supplementary Information: CHAPTER 3 - Classification of genomic features of plant-associated bacteria using machine learning
<p>Appendix A- List of all bacterial genomes used in orthologous genes clustering in the feature extraction step and in the further steps to build and test classifiers’ models. The list includes the isolation source information and the related category for the genome classification and features selection purposes.</p> <p>Appendix B - Distribution of genomes by phylum, family, and genus among the categories defined according to bacteria lifestyle association.</p> <p>Appendix C - Enriched orthogroups by genus according to each enrichment test (Material and Methods). Values for each test are "Y" (enriched), "N" (not enriched), or "Untested" (clusters were untested when there was insufficient phylogenetic signal, they were too small or were found in all genomes).</p> <p>Appendix D - Classification performance of random forest and logistic regression techniques applied to genus-specific datasets of genomic features (orthogroups) using both matrices from gene count number and presence/absence values. Sensitivity is a measure of how well a test identifies true positives; Specificity: is a measure how well a test or model avoids false positives; Positive Predictive Value (Pos. Pred. Value): The probability that a positive prediction is correct; Negative Predictive Value (Neg. Pred. Value): The probability that a negative prediction is correct; Precision: The accuracy of positive predictions; Recall (Sensitivity): The ability to find all relevant cases; F1 Score: A combined measure of precision and recall; Prevalence: The proportion of positive cases in the total; Detection Rate: The proportion of true positive cases identified; Detection Prevalence: The proportion of positive predictions; Balanced Accuracy: An average of sensitivity and specificity; Area Under the Curve (AUC): The overall performance of the model in distinguishing between positive and negative cases.</p> <p>Appendix E - Orthogroups assigned with predicted COGs as an important feature for classifying plant-associated genomes. COG categories: A - RNA processing and modification; B - Chromatin structure and dynamics; C - Energy production and conversion; D - Cell cycle control, cell division, chromosome partitioning; E - Amino acid transport and metabolism; F - Nucleotide transport and metabolism; G - Carbohydrate transport and metabolism; H - Coenzyme transport and metabolism; I - Lipid transport and metabolism; J - Translation, ribosomal structure and biogenesis; K - Transcription; L - Replication, recombination and repair; M - Cell wall/membrane/envelope biogenesis; N - Cell motility; O - Posttranslational modification, protein turnover, chaperones; P - Inorganic ion transport and metabolism; Q - Secondary metabolites biosynthesis, transport and catabolism; R - General function prediction only; S - Function unknown; T - Signal transduction mechanisms; U - Intracellular trafficking, secretion, and vesicular transport; V - Defense mechanisms; W - Extracellular structures; X - Mobilome: prophages, transposons; Y - Nuclear structure; Z - Cytoskeleton.</p>
Sequence and functional analyses of native plasmids from plant pathogenic Gammaproteobacteria: comparative genomics, conjugative mobilization and fitness effects
<p>These data tables are part of the Supplementary Material for Chapter I of the thesis titled <em>"Sequence and Functional Analyses of Native Plasmids from Plant-Pathogenic Gammaproteobacteria: Comparative Genomics, Conjugative Mobilization, and Fitness Effects."</em></p>
Genome-wide analysis of natural and restored eastern oyster populations reveals local adaptation and positive impacts of planting frequency and broodstock number
<p>The release of captive-bred plants and animals has increased worldwide to augment declining species. However, insufficient attention has been given to understanding how neutral and adaptive genetic variation are partitioned within and among proximal natural populations, and the patterns and drivers of gene flow over small spatial scales, which can be important for restoration success. A seascape genomics approach was used to investigate population structure, local adaptation, and the extent to which environmental gradients influence genetic variation among natural and restored populations of Chesapeake Bay eastern oysters <i>Crassostrea virginica</i>. We also investigated the impact of hatchery practices on neutral genetic diversity of restored reefs and quantified the broader genetic impacts of large-scale hatchery-based bivalve restoration. Restored reefs showed similar levels of diversity as natural reefs, and striking relationships were found between planting frequency and broodstock numbers and genetic diversity metrics (effective population size and relatedness), suggesting that hatchery practices can have a major impact on diversity. Despite long-term restoration activities, haphazard historical translocations, and high dispersal potential of larvae that could homogenize allele frequencies among populations, moderate neutral population genetic structure was uncovered. Moreover, environmental factors, namely salinity, pH, and temperature, play a major role in the distribution of neutral and adaptive genetic variation. For marine invertebrates in heterogeneous seascapes, collecting broodstock from large populations experiencing similar environments to candidate sites may provide the most appropriate sources for restoration and ensure population resilience in the face of rapid environmental change. This is one of a few studies to demonstrate empirically that hatchery practices have a major impact on the retention of genetic diversity. Overall, these results contribute to the growing body of evidence for fine-scale genetic structure and local adaptation in broadcast-spawning marine species and provide novel information for the management of an important fisheries resource.</p>
A genome for Bidens hawaiensis: a member of a hexaploid Hawaiian plant adaptive radiation
Abstract <p></p><p>The plant genus <em>Bidens</em> (Asteraceae or Compositae; Coreopsidae) is a species-rich and circumglobally distributed taxon. The 19 hexaploid species endemic to the Hawaiian Islands are considered an iconic example of adaptive radiation, of which many are imperiled and of high conservation concern. Until now, no genomic resources were available for this genus, which may serve as a model system for understanding the evolutionary genomics of explosive plant diversification. Here, we present a high-quality reference genome for the Hawai'i Island endemic species <em>B. hawaiensis</em> A. Gray reconstructed from long-read, high-fidelity sequences generated on a Pacific Biosciences Sequel II System. The haplotype-aware, draft genome assembly consisted of ~6.67 Giga bases (Gb), close to the holoploid genome size estimate of 7.56 Gb (± 0.44 SD) determined by flow cytometry. After removal of alternate haplotigs and contaminant filtering, the consensus haploid reference genome was comprised of 15,904 contigs containing ~3.48 Gb, with a contig N50 value of 422,594. The high interspersed repeat content of the genome, approximately 74%, along with hexaploid status, contributed to assembly fragmentation. Both the haplotype-aware and consensus haploid assemblies recovered >96% of Benchmarking Universal Single-copy Orthologs. Yet, the removal of alternate haplotigs did not substantially reduce the proportion of duplicated benchmarking genes (~79% versus ~68%). This reference genome will support future work on the speciation process during adaptive radiation, including resolving evolutionary relationships, determining the genomic basis of trait evolution, and supporting ongoing conservation efforts.</p><p></p>
The effect of methodological considerations on the construction of gene-based plant pan-genomes
<p>Pan-genomics is an emerging approach for studying the genetic diversity within plant populations. In contrast to common resequencing studies that compare whole genome sequencing data to a single reference genome, the construction of a pan-genome involves the direct comparison of multiple genomes to one another, thereby enabling the detection of genomic sequences and genes not present in the reference, as well as the analysis of gene content diversity. While multiple studies describing pan-genomes of various plant species have been published in recent years, our understanding regarding the effect of the computational procedures used for pan-genome construction is still limited.</p> <p>Here we examine the effect of several key methodological factors on the obtained gene pool and on gene presence-absence detections by constructing and comparing multiple pan-genomes of Arabidopsis thaliana and cultivated soybean, as well as conducting a meta-analysis on published pan-genomes. These factors include the construction method, the sequencing depth, and the extent of input data used for gene annotation. We observe substantial differences between pan-genomes constructed using three common procedures (De novo assembly and annotation, Map-to-pan, and Iterative assembly), and that results are dependent on the extent of the input data. Specifically, we report low agreement between the gene content inferred using different procedures and input data. Our results should increase the awareness of the community to the consequences of methodological decisions made during the process of pan-genome construction and emphasize the need for further investigation of commonly applied methodologies.</p>
Supplementary Information: CHAPTER 2 - Unveiling genomic features linked to traits of plant-growth-promoting bacterial communities from sugarcane
<p>Appendix A. Summary of counts of subreads and circular consensus sequencing (CCS) sequences obtained for PacBio sequencing of SMRT libraries. (EMS_1.xlsx)</p> <p>Appendix B. Taxonomy assignment of MAGs at the higher taxonomic rank obtained from GTDB-tk and Kraken tools. (EMS_2.xlsx)</p> <p>Appendix C. Report of the classification workflow using GTDB-tk. (EMS_3.xlsx)</p> <p>Appendix D. Matrix of the KEGG Orthology (KOs) frequencies annotated by the EnrichM tool. (EMS_4.xlsx)</p> <p>Appendix E. Reconstruction and completeness of KEGG modules annotated by EnrichM. The asterisks (*) in the header represent additional values obtained by the script ‘classKEGGModules.pl’ (https://github.com/dgpinheiro/bioinfoutilities) to estimate PGPTs in KEGG modules. (EMS_5.xlsx)</p> <p>Appendix F. The secondary metabolite biosynthesis gene clusters (BGCs) identified with AntiSMASH. (EMS_6.xlsx)</p> <p>Appendix G. The raw count of plant growth-promoting traits (PGPTs) annotations, according to KEGG Orthology (KO) predictions for MAGs. (EMS_7.xlsx)</p> <p>Appendix H. The raw count of plant growth-promoting traits (PGPTs) that comprises the 39 classes (level 5 hierarchy) identified as enriched according to the results of Pearson's Chi-square test (qvalue ≤ 0.1). (EMS_8.xlsx)</p>
Supplementary material 1 from: Zúñiga JD, Gostel MR, Mulcahy DG, Barker K, Hill A, Sedaghatpour M, Vo SQ, Funk VA, Coddington JA (2017) Data Release: DNA barcodes of plant species collected for the Global Genome Initiative for Gardens Program, National Museum of Natural History, Smithsonian Institution. PhytoKeys 88: 119-122. https://doi.org/10.3897/phytokeys.88.14607
List of samples collected for the Global Genome Initiative for Gardens project selected for DNA barcoding, with GenBank accession numbers and genetic sample identification numbers. All the sequences are included in the GGI-Gardens BioProject. : Explanation note: List of samples collected for the Global Genome Initiative for Gardens project selected for DNA barcoding, with GenBank accession numbers and genetic sample identification numbers.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.