Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
479
datasets available to search
ShareScore release 0.7.1
Dataset results
479 results for “genome evolution”
Data from: Environmental variation influences genome evolution in Hispaniolan trunk anoles (<i>Anolis distichus</i>)
Open the record for dataset details and reuse information.
Data for: The 3-dimensional genome drives the evolution of asymmetric gene duplicates via enhancer capture-divergence
Open the record for dataset details and reuse information.
Dataset for: The redlegged earth mite draft genome provides new insights into pesticide resistance evolution and demography in its invasive Australian range
Open the record for dataset details and reuse information.
Comparative genomics sheds new light on the convergent evolution of infrared vision in snakes
Open the record for dataset details and reuse information.
The Evolution and Genomic Basis of Beetle Diversity
<p>Datasets S1-S4. (each is contained in a separate .zip file)</p> <p><strong>Dataset S1.</strong></p> <p>Gene trees for plant cell wall degrading enzyme phylogenetic analyses in the directory ‘Pfam Candidate Genes trees and phy’</p> <p>1. Phylip formatted files for each gene used in ML analyses.</p> <p>2. ML tree files for each gene studied showing TBE bootstrap support (100 replicates) (corresponding to Figs. S15-S26).</p> <p>3. ML tree files for each gene studied showing ML bootstrap support (100 replicates) from IQtree (corresponding to Figs. S15-S26).</p> <p><strong>Dataset S2.</strong></p> <p>Directory (Blast_10best_hits_Pfam_Candidate_Genes) including: Blast results (10 best hits) for all sequences extracted from the transcriptome and genome assemblies for the plant cell wall degrading enzyme analysis in the directory ‘Pfam_Candidate_Genes_Fas’ (before filtering).</p> <p><strong>Dataset S3.</strong></p> <p>Directory (Pfam_Candidate_Genes_Fas) including: Candidate genes/transcripts encoding plant cell wall degrading enzymes extracted from the transcriptome and genome assemblies (before filtering).</p> <p><strong>Dataset S4.</strong></p> <p>Directory (Supermatrices_partitions) including:</p> <ul> <li>Supermatrix for Fig. 1 (amino acid and nucleotide level, PHYLIP formats) and Supermatrix for Fig. S10 (amino acid level, PHYLIP format).</li> <li>Partition schemes for supermatrix for Fig. 1 (Fig_1_Partition_finder_best_scheme).</li> <li>Partition schemes for supermatrix for Fig. 1 prior to Partitionfinder (AA_partitions).</li> <li>Partition schemes for supermatrix for NT prior to Partitionfinder (NT_partitions).</li> <li>Partition schemes for supermatrix for Fig. S10 (Fig_S10_Partition_finder_best_scheme)</li> </ul> <p><strong>Data Use Statement.</strong> Data on genetic material contained in this paper are published for non-commercial use only. Utilization by third parties for purposes other than non-commercial scientific research may infringe the conditions under which the genetic resources were originally accessed, and should not be undertaken without obtaining consent from the corresponding author of the paper and/or obtaining permission from the original provider of the genetic material.</p>
Phylogenomics and contrasting modes of genome evolution in Ascomycota
<p>332 Saccharomycotina assemblies <br> 761 Pezizomycotina assemblies<br> 14 Taphrinomycotina assemblies<br> 6 Basidiomycota (outgroup) assemblies</p> <p>Proteomes of 1,107 Ascomycota</p>
Investigating evolution at the catalytic site of the main SARS-CoV-2 protease using over 15,000 genomes
<p>We investigated evolution and genomic variation of SARS-CoV-2 within the current pandemic at the catalytic site of the main SARS-CoV-2 protease (see https://zenodo.org/record/3834875#.Xs1IHsZ7nyk and <a href="https://openlabnotebooks.org/mapping-the-genetic-variations-of-sars-cov-2-onto-its-proteins-crystal-structures-post-1/">https://openlabnotebooks.org/mapping-the-genetic-variations-of-sars-cov-2-onto-its-proteins-crystal-structures-post-1/ </a>).<br> We used more than 15,000 genomic sequences from GISAID (<a href="https://www.epicov.org/">https://www.epicov.org/</a>) available on the 17th of May 2020.<br> We use a new approach based on phylogenetic inference of homoplasy, clustering of mutations, and ambiguous consensus sequence characters, to identify sites that are likely affected by sequencing artefacts.<br> We find that these sites are mostly conserved, and the amino acid variants observed are only M49I, P52S, N142S, and P168S, all of which appear only at extremely low frequencies (maximum of two samples each).</p>
Data from: Genome assembly of the ragweed leaf beetle, a step forward to better predict rapid evolution of a weed biocontrol agent to environmental novelties
<p><span>Rapid evolution of weed biological control agents (BCAs) to new biotic and abiotic conditions is poorly understood and so far, only little considered both in pre-release and post-release studies, despite potential major negative or positive implications for risks of non-targeted attacks or for colonizing yet unsuitable habitats, respectively. Provision of genetic resources, such as assembled and annotated genomes, is essential to assess potential adaptive processes by identifying underlying genetic mechanisms. Here, we provide the first sequenced genome of a phytophagous insect used as a BCA, <i>i.e.</i> the leaf beetle <i>Ophraella communa</i>, a promising BCA of common ragweed, recently and accidentally introduced into Europe. A total 33.98 Gb of raw DNA sequences, representing c. 43-fold coverage, were obtained using the PacBio SMRT-Cell sequencing approach. Among the five different assemblers tested, the SMARTdenovo assembly displaying the best scores was then corrected with Illumina short reads. A final genome of 774 Mb containing 7,003 scaffolds was obtained. The reliability of the final assembly was then assessed by benchmarking universal single-copy orthologous genes (> 96.0% of the 1,658 expected insect genes) and by remapping tests of Illumina short reads (average of 98.6% ± 0.7% without filtering). The number of protein-coding genes of 75,642, representing 82% of the published antennal transcriptome, and the phylogenetic analyses based on 825 orthologous genes placing <i>O. communa </i>in the monophyletic group of Chrysomelidae, confirm the relevance of our genome assembly. Overall, the genome provides a valuable resource for studying potential risks and benefits of this BCA facing environmental novelties.</span></p>
Data from: Chromosome-level genome of the melon thrips yields insights into evolution of a sap-sucking lifestyle and pesticide resistance
<p>Thrips are tiny insects from the order Thysanoptera (Hexapoda: Condylognatha), including many sap-sucking pests that are causing increasing damage to crops worldwide. In contrast to their closest relatives of Hemiptera (Hexapoda: Condylognatha) including numerous sap-sucking species, there are few genomic resources available for thrips. In this study, we assembled the first thrips genome at the chromosome level from the melon thrips, <i>Thrips palmi</i>, a notorious pest in agriculture, using PacBio long-read and Illumina short-read sequences. The assembled genome was 270.43 Mb in size with 4,120 contigs and a contig N50 of 426 kb. All contigs were assembled into 16 linkage groups assisted by the Hi-C technique. In total, 16,333 protein-coding genes were predicted, of which 88.13% were functionally annotated. Among sap-sucking insects, polyphagous species usually possess more detoxification genes than oligophagous species. The polyphagous thrips genomes characterized so far have relatively more detoxification genes in the GST and CCE families than polyphagous aphids, but they have fewer UGTs. HSP genes, especially from the Hsp70s group, have expanded in thrips compared to other hemipteran insects. These differences point to different genetic mechanisms associated with detoxification and stress responses in these two groups of sap-sucking insects. The expansion of these gene families may contribute to the rapid development of pesticide resistance in thrips, as supported by a transcriptome comparison of resistant and sensitive populations of <i>T. palmi</i>. The high-quality genome developed here provides an invaluable resource for understanding the ecology, genetics and evolution of thrips as well as their relatives more generally.</p>
Data from: Phylogenomic analysis of Wolbachia strains reveals patterns of genome evolution and recombination
<p><i>Wolbachia</i> are widespread intracellular bacteria that mediate many important biological processes in arthropod species. In this study, we identified 210 conserved single-copy genes in 33 genome-sequenced <i>Wolbachia</i> strains in the A, B, C, D, E and F supergroups. Phylogenomic analysis with these core genes indicate that all 33 <i>Wolbachia</i> strains maintain the supergroup relationship classified previously based on the multilocus sequence typing (MLST) genes. Using an interclade recombination screening method, 14 inter-supergroup recombination events were discovered in six genes (2.9%) among 210 single copy orthologs. This finding suggests a relatively low frequency of intergroup recombination. Interestingly, they have occurred not only between A and B supergroups (9 events), but also between A and E supergroups (5 events). Maintenance of such transfers suggests possible roles in <i>Wolbachia</i> infection related functions. Comparisons of strain divergence using the five genes of the MLST system show a high correlation (Pearson correlation coefficient r = 0.98) between MLST and whole genome divergences, indicating that MLST is a reliable method for identifying related strains when whole genome data are not available. The phylogenomic analysis and the identified core gene set in our study will serve as a valuable foundation for strain identification and the investigation of recombination and genome evolution in <i>Wolbachia</i>.</p>
Genomic and phenotypic evolution of Escherichia coli in a novel citrate-only resource environment
Evolutionary innovations allow populations to colonize new ecological niches. We previously reported that aerobic growth on citrate (Cit+) evolved in an Escherichia coli population during adaptation to a minimal glucose medium containing citrate (DM25). Cit+ variants can also grow in citrate-only medium (DM0), a novel environment for E. coli. To study adaptation to this niche, we founded two sets of Cit+ populations and evolved them for 2500 generations in DM0 or DM25. The evolved lineages acquired numerous parallel mutations, many mediated by transposable elements. Several also evolved amplifications of regions containing the maeA gene. Unexpectedly, some evolved populations and clones show apparent declines in fitness. We also found evidence of substantial cell death in Cit+ clones. Our results thus demonstrate rapid trait refinement and adaptation to the new citrate niche, while also suggesting a recalcitrant mismatch between E. coli physiology and growth on citrate.
Data from: Distinct genomic signals of lifespan and life history evolution in response to postponed reproduction and larval diet in Drosophila
Reproduction and diet are two major factors controlling the physiology of aging and life history, but how they interact to affect the evolution of longevity is unknown. Moreover, while studies of large-effect mutants suggest an important role of nutrient sensing pathways in regulating aging, the genetic basis of evolutionary changes in lifespan remains poorly understood. To address these questions, we analyzed the genomes of experimentally evolved Drosophila melanogaster populations subjected to a factorial combination of two selection regimes: reproductive age (early versus postponed), and diet during the larval stage ('low', 'control', 'high'), resulting in six treatment combinations with four replicate populations each. Selection on reproductive age consistently affected lifespan, with flies from the postponed reproduction regime having evolved a longer lifespan. In contrast, larval diet affected lifespan only in early-reproducing populations: flies adapted to the 'low' diet lived longer than those adapted to control diet. Here we find genomic evidence for strong independent evolutionary responses to either selection regime, as well as loci that diverged in response to both regimes, thus representing genomic interactions between the two. Overall, we find that the genomic basis of longevity is largely independent of dietary adaptation. Differentiated loci were not enriched for 'canonical' longevity genes, suggesting that naturally occurring genic targets of selection for longevity differ qualitatively from variants found in mutant screens. Comparing our candidate loci to those from other 'evolve-and-resequence' studies of longevity demonstrated significant overlap among independent experiments. This suggests that the evolution of longevity, despite its presumed complex and polygenic nature, might be to some extent convergent and predictable.
Supplementary datasets for: Large-scale genome sequencing reveals the driving forces of viruses in microalgal evolution
<p>Microalgae are integral primary producers for global ecosystems whose genomes can be mined for ecological insights, but representative genome sequences are lacking for many phyla. We cultured and sequenced 107 microalgae species from 11 different phyla indigenous to varied geographies and climates. This genome collection was used to resolve genomic differences between saltwater and freshwater microalgae. Freshwater species showed domain-centric ontology enrichment for nuclear and nuclear membrane functions, while saltwater species were enriched in organellar and cellular membrane functions. Marine species contained significantly more viral families in their genomes (<span>p-value = 8 x 10(-4))</span>. Viral sequences were identified from Chlorovirus, Coccolithovirus, Pandoravirus, Marseillevirus, Tupanvirus, and others integrated into algal genomes. Algal, viral-origin sequences were found to be expressed and to code for a wide variety of functions. Our results clarify the poorly characterized occurrences of viral elements in algal genomes and define a unified adaptive strategy for algal halotolerance.</p>
Potential causes and consequences of rapid mitochondrial genome evolution in thermoacidophilic Galdieria (Rhodophyta)
<p>The Cyanidiophyceae is an early-diverged red algal class that thrives in extreme conditions around acidic hot springs. Although this lineage has been highlighted as a model for understanding the biology of extremophilic eukaryotes, little is known about the molecular evolution of their mitochondrial genomes (mitogenomes).</p> <p>To fill this knowledge gap, we sequenced five mitogenomes from representative clades of Cyanidiophyceae and identified two major groups, here referred to as Galdieria-type (G-type) and Cyanidium-type (C-type). G-type mitogenomes exhibit the following three features: (i) reduction in genome size and gene inventory, (ii) evolution of unique protein properties including charge, hydropathy, stability, amino acid composition, and protein size, and (iii) distinctive GC-content and skewness of nucleotides. Based on GC-skew-associated characteristics, we postulate that unidirectional DNA replication may have resulted in the rapid evolution of G-type mitogenomes.</p> <p>The high divergence of G-type mitogenomes was likely driven by natural selection in the multiple extreme environments that Galdieria species inhabit combined with their highly flexible heterotrophic metabolism. We speculate that the interplay between mitogenome divergence and adaptation may help explain the dominance of Galdieria species in diverse extreme habitats.</p>
Data from: Reticulate evolution within a spruce (Picea) species complex revealed by population genomic analysis
The role of reticulation in the rapid diversification of organisms is attracting greater attention in evolutionary biology. Here, we report a population genomics approach to test the role of hybridization and introgression in the evolution of the Picea likiangensis species complex. Based on 84,793 SNPs detected in transcriptomes of 82 trees collected from 35 localities, we identified 18 hybrids (including backcrosses) distributed within the range boundaries of the four taxa. Coalescent simulations, for each pair of taxa and for all taxa taken together, rejected several tree-like divergence models and supported instead a reticulate evolution model with secondary contacts occurring during Pleistocene glacial cycles after initial divergence in the late Pliocene. Significant gene flow occurred among some taxa after secondary contact according to an analysis based on modified ABBA-BABA statistics that accommodated a rapid diversification scenario. A novel finding was that introgression between certain taxa can contribute to increasing divergence (and possibly reproductive isolation) between those taxa and other taxa within a complex at some loci. These results illuminate the reticulate nature of evolution within the P. likiangensis complex and highlight the value of population genomic data in detecting the effects of introgression in the rapid diversification of related taxa.
Data from: Evolution of the plastid genomes in diatoms
Diatoms are a monophyletic group of eukaryotic, single-celled heterokont algae. Despite years of phylogenetic research, relationships among major groups of diatoms remain uncertain. Here we assess diatom phylogenetic relationships using the plastid genome (plastome). The 22 previously published diatom plastomes showed variable genome size, gene content and extensive rearrangement. We report another 18 diatom plastome sequences ranging in size from 119,120 to 201,816 bp. Plagiogramma staurophorum had the largest plastome sequenced so far due to large inverted repeats and a 2971 bp group II intron insertion in petD. The previously reported loss of psaE, psaI and psaM genes in Rhizosolenia imbricata also occurred in the closely related species Rhizosolenia fallax. In the largest genome-scale phylogeny yet published for diatoms based on 103 shared plastid-coding genes from 40 diatoms and Triparma laevis as the outgroup, Leptocylindrus was recovered as sister to the remaining diatoms and the clade of Attheya plus Biddulphia was recovered as sister to pennate diatoms, strongly rejecting monophyly of two of the three proposed classes of diatoms. Our study also revealed extensive gene loss and a strong positive correlation between sequence divergence and gene order change in diatom plastomes.
Data from: Combining experimental evolution and genomics to understand how seed beetles adapt to a marginal host plant
<p>Genes that affect adaptive traits have been identified, but our knowledge of the genetic basis of adaptation in a more general sense (across multiple traits) remains limited. We combined population-genomic analyses of evolve and resequence experiments, genome-wide association mapping of performance traits, and analyses of gene expression to fill this knowledge gap, and shed light on the genomics of adaptation to a marginal host (lentil) by the seed beetle <em>Callosobruchus maculatus</em>. Using population-genomic approaches, we detected modest parallelism in allele frequency change across replicate lines during adaptation to lentil. Mapping populations derived from each lentil-adapted line revealed a polygenic basis for two host-specific performance traits (weight and development time), which had low to modest heritabilities. We found less evidence of parallelism in genotype-phenotype associations across these lines than in allele frequency changes during the experiments. Differential gene expression caused by differences in recent evolutionary history exceeded that caused by immediate rearing host. Together, the three genomic data sets suggest that genes affecting traits other than weight and development time are likely to be the main causes of parallel evolution, and that detoxification genes (especially cytochrome P450s and beta-glucosidase) could be especially important for colonization of lentil by <em>C. maculatus</em>.</p>
Data from: Clonal evolution and genome stability in a 2,500-year-old fungal individual
Individuals of the basidiomycete fungus Armillaria are well-known for their ability to spread from woody substrate to substrate on the forest floor through the growth of rhizomoprhs. Here we made 248 collections of A. gallica in one locality in Michigan's Upper Peninsula. To identify individuals, we genotyped collections with molecular markers and somatic compatibility testing. We found several different individuals in proximity to one another, but one genetic individual stood out as exceptionally large, covering hundreds of tree root systems over approximate 75 hectares of forest floor. Based on observed growth rates of the fungus, we estimate the minimum age of the large individual as 2,500 years. With whole-genome sequencing and variant discovery, we also found that mutation had occurred within the somatic cells of the individual, reflecting its historical pattern of growth from a single point. The overall rate of mutation over the 90 mb genome, however, was extremely low. This same individual was first discovered in the late 1980s, but its full spatial extent and internal mutation dynamic was unkown at that time. The large individual of A. gallica has been remarkably resistant to genomic change as it has persisted in place.
Data from: Population genomics of rapid evolution in natural populations: polygenic selection in response to power station thermal effluents
Background: Examples of rapid evolution are common in nature but difficult to account for with the standard population genetic model of adaptation. Instead, selection from the standing genetic variation permits rapid adaptation via soft sweeps or polygenic adaptation. Empirical evidence of this process in nature is currently limited but accumulating. Results: We provide genome-wide analyses of rapid evolution in two Fundulus heteroclitus populations subjected to recently elevated temperatures due to coastal power station thermal effluents. Bayesian and multivariate analyses of population genomic structure reveal a substantial portion of genetic variation that is most parsimoniously explained by selection at the site of thermal effluents. An FST outlier approach in conjunction with additional conservative requirements identify significant allele frequency differentiation that exceeds neutral expectations among exposed and closely related reference populations. Genomic variation patterns near these candidate loci reveal that individuals living near thermal effluents have rapidly evolved from the standing genetic variation through small allele frequency changes at many loci in a pattern consistent with polygenic selection on the standing genetic variation. Conclusions: While the ultimate trajectory of selection in these populations is unknown, our findings suggest that polygenic models of adaptation may play important roles in large, natural populations experiencing recent selection due to environmental changes that cause broad physiological impacts.
Deciphering genomic evolution of metastatic organotropism with 535 paired primary lung cancers and metastases
<p>Scripts and processed data used in the analyses of the manuscript:</p> <p><em>Deciphering genomic evolution of metastatic organotropism with 535 paired primary lung cancers and metastases.</em></p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.