Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
479
datasets available to search
ShareScore release 0.9.0
Dataset results
479 results for “genomic evolution”
Data from: Adaptive evolution and segregating load contribute to the genomic landscape of divergence in two tree species connected by episodic gene flow
Speciation often involves repeated episodes of genetic contact between divergent populations before reproductive isolation (RI) is complete. Whole-genome sequencing (WGS) holds great promise for unravelling the genomic bases of speciation. We have studied two ecologically divergent, hybridizing species of the 'model tree' genus Populus (poplars, aspens, cottonwoods), Populus alba and P. tremula, using >8.6 million single nucleotide polymorphisms (SNPs) from WGS of population pools. We used the genomic data to (i) scan these species' genomes for regions of elevated and reduced divergence, (ii) assess key aspects of their joint demographic history based on genomewide site frequency spectra (SFS) and (iii) infer the potential roles of adaptive and deleterious coding mutations in shaping the genomic landscape of divergence. We identified numerous small, unevenly distributed genome regions without fixed polymorphisms despite high overall genomic differentiation. The joint SFS was best explained by ancient and repeated gene flow and allowed pinpointing candidate interspecific migrant tracts. The direction of selection (DoS) differed between genes in putative migrant tracts and the remainder of the genome, thus indicating the potential roles of adaptive divergence and segregating deleterious mutations on the evolution and breakdown of RI. Genes affected by positive selection during divergence were enriched for several functionally interesting groups, including well-known candidate 'speciation genes' involved in plant innate immunity. Our results suggest that adaptive divergence affects RI in these hybridizing species mainly through intrinsic and demographic processes. Integrating genomic with molecular data holds great promise for revealing the effects of particular genetic pathways on speciation.
Data from: Recurrent selection explains parallel evolution of genomic regions of high relative but low absolute differentiation in a ring species
Recent technological developments allow investigation of the repeatability of evolution at the genomic level. Such investigation is particularly powerful when applied to a ring species, in which spatial variation represents changes during the evolution of two species from one. We examined genomic variation among three subspecies of the greenish warbler ring species, using genotypes at 13 013 950 nucleotide sites along a new greenish warbler consensus genome assembly. Genomic regions of low within-group variation are remarkably consistent between the three populations. These regions show high relative differentiation but low absolute differentiation between populations. Comparisons with outgroup species show the locations of these peaks of relative differentiation are not well explained by phylogenetically conserved variation in recombination rates or selection. These patterns are consistent with a model in which selection in an ancestral form has reduced variation at some parts of the genome, and those same regions experience recurrent selection that subsequently reduces variation within each subspecies. The degree of heterogeneity in nucleotide diversity is greater than explained by models of background selection, but is consistent with selective sweeps. Given the evidence that greenish warblers have had both population differentiation for a long period of time and periods of gene flow between those populations, we propose that some genomic regions underwent selective sweeps over a broad geographic area followed by within-population selection-induced reductions in variation. An important implication of this 'sweep-before-differentiation' model is that genomic regions of high relative differentiation may have moved among populations more recently than other genomic regions.
Data from: A single interacting species leads to widespread parallel evolution of the stickleback genome
Biotic interactions are potent, widespread causes of natural selection and divergent phenotypic evolution, and can lead to genetic differentiation with gene flow among wild populations ("isolation by ecology") [1-4]. Biotic selection has been predicted to act on more genes than abiotic selection thereby driving greater adaptation [5]. However, difficulties in isolating the genome-wide effect of single biotic agents of selection have limited our ability to identify and quantify the number and type of specific genetic regions responding to biotic selection [6-9]. We identified geographically interspersed lakes in which threespine stickleback fish (Gasterosteus aculeatus) have repeatedly adapted to the presence/absence of a single member of the ecological community, prickly sculpin (Cottus asper), a fish species that is both competitor and predator of stickleback [10]. Whole genome sequencing revealed that sculpin presence/absence accounted for the majority of genetic divergence among populations, more so than geography. The major axis of stickleback genomic variation within and between the two lake types was correlated with multiple traits, indicating parallel natural selection across a gradient of biotic environments. A large proportion of the genome - about 1.8%, encompassing more than 600 genes – differentiated stickleback from the two biotic environments. Divergence occurred in 141 discrete genomic clumps located mainly in regions of low recombination within the stickleback genome, suggesting that genes brought to lakes by the colonizing ancestral population often evolved together in linked blocks. Strong selection and a wealth of standing genetic variation explain how a single member of the biotic community can have such a rapid and profound evolutionary impact.
Data from: Major improvements to the Heliconius melpomene genome assembly used to confirm 10 chromosome fusion events in 6 million years of butterfly evolution
The Heliconius butterflies are a widely studied adaptive radiation of 46 species spread across Central and South America, several of which are known to hybridize in the wild. Here, we present a substantially improved assembly of the Heliconius melpomene genome, developed using novel methods that should be applicable to improving other genome assemblies produced using short read sequencing. First, we whole-genome-sequenced a pedigree to produce a linkage map incorporating 99% of the genome. Second, we incorporated haplotype scaffolds extensively to produce a more complete haploid version of the draft genome. Third, we incorporated ∼20x coverage of Pacific Biosciences sequencing, and scaffolded the haploid genome using an assembly of this long-read sequence. These improvements result in a genome of 795 scaffolds, 275 Mb in length, with an N50 length of 2.1 Mb, an N50 number of 34, and with 99% of the genome placed, and 84% anchored on chromosomes. We use the new genome assembly to confirm that the Heliconius genome underwent 10 chromosome fusions since the split with its sister genus Eueides, over a period of about 6 million yr.
Data from: The population genomics of sunflowers and genomic determinants of protein evolution revealed by RNAseq
Few studies have investigated the causes of evolutionary rate variation among plant nuclear genes, especially in recently diverged species still capable of hybridizing in the wild. The recent advent of Next Generation Sequencing (NGS) permits investigation of genome wide rates of protein evolution and the role of selection in generating and maintaining divergence. Here, we use individual whole-transcriptome sequencing (RNAseq) to refine our understanding of the population genomics of a wild species of sunflowers (Helianthus spp.) and the factors that affect rates of protein evolution. We aligned 35 GB of transcriptome sequencing data and identified 433,257 polymorphic sites (SNPs) in a reference transcriptome comprising 16,312 genes. Using SNP markers, we identified strong population clustering largely corresponding to the three species analyzed here (Helianthus annuus, H. petiolaris, H. debilis), with one distinct early generation hybrid. Then, we calculated the proportions of adaptive substitution fixed by selection (alpha) and identified gene ontology categories with elevated values of alpha. The "response to biotic stimulus" category had the highest mean alpha across the three interspecific comparisons, implying that natural selection imposed by other organisms plays an important role in driving protein evolution in wild sunflowers. Finally, we examined the relationship between protein evolution (dN/dS ratio) and several genomic factors predicted to co-vary with protein evolution (gene expression level, divergence and specificity, genetic divergence [FST], and nucleotide diversity [pi]). We find that variation in rates of protein divergence was correlated with gene expression level and specificity, consistent with results from a broad range of taxa and timescales. This would in turn imply that these factors govern protein evolution both at a microevolutionary and macroevolutionary timescale. Our results contribute to a general understanding of the determinants of rates of protein evolution and the impact of selection on patterns of polymorphism and divergence.
Duck pan-genome reveals two transposon-derived structural variations caused bodyweight enlarging and white plumage phenotype formation during evolution
<p><span>Structural variations (SVs) are a major source of domestication and improvement traits. We present the first duck pan-genome constructed using five genome assemblies capturing ~40.98 Mb new sequences. This pan-genome together with high-depth sequencing data (>46.5X) identified 101,041 SVs, of which substantial proportions were derived from transposable element (TE) activity. Many TE-derived SVs anchored in a gene body or regulatory region are linked to domestication and improvement. By combining quantitative genetics with molecular experiments, we dissect how TE-derived SVs change gene expression of <em>IGF2BP1</em> and generate novel transcripts of <em>MITF</em>, shaping body weight and plumage color. In the <em>IGF2BP1</em> locus, the TE-derived SV explains the largest effect on body weight among avian species (27.61% of phenotypic variation). Our findings highlight the </span><span>importance of using a pan-genome as a reference in genomics studies</span><span> and explore the roles of TE-derived SVs in trait formation and in livestock breeding.</span></p>
Figure 2 in The tight genome size of ants: diversity and evolution under ancestral state reconstruction and base composition
Figure 2. Bayesian consensus tree resulting from the LW-Rh and Wg gene alignments (871 bp). Coloured dots on the branches indicate the values of posterior probability (PP): green dots represent values between 1.00 and 0.95, yellow dots between 0.94 and 0.90, and red dots ≤ 0.89. The nodes are indicated with numbers. Values above and below the branches represent the ancestral genome size (GS; 1C-values, in picograms) at particular nodes: in blue is the value generated by the maximum likelihood (ML) [asterisks are related to confidence interval (CI) values shown in Supporting Information, Table S4]; orange is the value generated by maximum parsimony (MP); and black, given below the branches, is the value generated by Bayesian inference (BI). Genome size data (1C-values) were obtained in the present work (pink dots) or taken from the literature (grey dots).
Figure 1 in The tight genome size of ants: diversity and evolution under ancestral state reconstruction and base composition
Figure 1. Fluorescence intensity histograms obtained from three different species, with Drosophila melanogaster as internal standard, stained with propidium iodide (PI; A–C) or 4,6-diamidino-2-phenylindole (DAPI; D–F). The x-axis corresponds to the scale of fluorescence intensity, and the y-axis represents the number of nuclei with that fluorescence intensity.
Figure 3 in The tight genome size of ants: diversity and evolution under ancestral state reconstruction and base composition
Figure 3. Mean genome size (in picograms and megabase pairs) estimated for Formicidae subfamilies. The phylogenetic tree generated in the present study was redrawn, with collapsed branches corresponding to species of the same subfamily.
Plastid genome evolution in subtribe Gentianinae (Gentianaceae)
<p>We investigated plastome evolution in Subtribe Gentianinae of the Gentianaceae, which encompasses ca. 450 species distributed around the world, particularly in alpine and subalpine environments. We sequenced, assembled and annotated the plastomes of 41 species, representing all six genera in subtribe Gentianinae as well as all 14 sections of the species-rich genus <i>Gentiana</i>. Here, we upload (1) original alignments, Gblock alignments and tree topology file in phylogenetic analysis, and (2) results about a 5 kb insertion in <em>Gentiana</em> <i>cuneibarba</i>, including Sanger sequencing results which verified its two boundaries as well as the middle, and annotation results.</p>
Predictability and parallelism in the contemporary evolution of hybrid genomes
<p>Hybridization between species is widespread across the tree of life. As a result, many species, including our own, harbor regions of their genome derived from hybridization. Despite the recognition that this process is widespread, we understand little about how the genome stabilizes following hybridization, and whether the mechanisms driving this stabilization tend to be shared across species. Here, we dissect the drivers of variation in local ancestry across the genome in replicated hybridization events between two species pairs of swordtail fish: <em>Xiphophorus birchmanni </em>× <em>X. cortezi</em> and <em>X. birchmanni </em>× <em>X. malinche</em> . We find surprisingly high levels of repeatability in local ancestry across the two types of hybrid populations. This repeatability is attributable in part to the fact that the recombination landscape and locations of functionally important elements play a major role in driving variation in local ancestry in both types of hybrid populations. Beyond these broad scale patterns, we identify dozens of regions of the genome where minor parent ancestry is unusually low or high across species pairs. Analysis of these regions points to shared sites under selection across species pairs, and in some cases, shared mechanisms of selection. We show that one such region is a previously unknown hybrid incompatibility that is shared across <em>X. birchmanni</em> × <em>X. cortezi</em> and <em>X. birchmanni</em> × <em>X. malinche</em> hybrid populations. </p>
Supplementary File to "Both binding strength and evolutionary accessibility affect the population frequency of transcription factor binding sequences in Arabidopsis thaliana" (Genome Biology and Evolution)
<p>This data is supplementary file 1 of the following publication:</p> <p>Schweizer G, Wagner A. "Both binding strength and evolutionary accessibility affect the population frequency of transcription factor binding sequences in Arabidopsis thaliana" (Genome Biology and Evolution)</p>
Novel genomic insights into body size evolution in cetaceans and a resolution of Peto's Paradox
<p>Cetaceans (whales, dolphins, and porpoises) have undergone a radical transformation from the typical terrestrial mammalian body plan to a streamlined one while exhibited dramatic inter-specific size ranges. However, the molecular mechanisms underlying the diversifying evolution of cetacean body size are largely unknown. Here, by using genome and phenotypic data from 22 cetaceans, we seek to investigate the genome-wide gene-phenotype correlation and to explore the genetic basis under the high diversity of body size in cetaceans. Results of the functional enrichment showed that body size-related genes in cetaceans were enriched in pathways associated with immunity, cell growth, and metabolism, suggesting their potential roles in the diversifying evolution of body size in cetaceans. A series of genes was also found coevolution with body size that are mainly involved in immune surveillance, tumor suppression function, and development of 'cheater' tumors. This in turn suggests that the genes play a role in tumor control and thus resolve Peto's paradox, a finding that the expansion in body size and thereby cell number does not correlate with increases in cancer incidence in larger whales. The present study could provide novel insights into the evolution of great body size variation in cetaceans.</p>
Data from: Correlated evolution of larval development, egg size, and genome size across two genera of snapping shrimp
<p>Across plants and animals, genome size is often correlated with life history traits: large genomes are correlated with larger seeds, slower development, larger body size, and slower cell division. Among decapod crustaceans, caridean shrimps are among the most variable both in terms of genome size variation and life history characteristics such as larval development mode and egg size, but the extent to which these traits are associated in a phylogenetic context is largely unknown. In this study, we examine correlations among egg size, larval development, and genome size in two different genera of snapping shrimp, <em>Alpheus </em>and <em>Synalpheus, </em>using phylogenetically informed analyses<em>. </em>In both <em>Alpheus </em>and <em>Synalpheus, </em>egg size is strongly linked to larval development mode: species with abbreviated development had significantly larger eggs than species with extended larval development. We produced the first comprehensive dataset of genome size in <em>Alpheus </em>(n = 37 species), and demonstrated that genome size was strongly and positively correlated with egg size in both <em>Alpheus </em>and <em>Synalpheus. </em>Correlated trait evolution analyses showed that in <em>Alpheus</em>, changes in genome size were clearly dependent on egg size. In <em>Synalpheus, </em>evolutionary path analyses suggest that changes in development mode (from extended to abbreviated) drove increases in egg volume; and larger eggs, in turn, resulted in larger genomes. These data suggest that variation in reproductive traits may underpin the high degree of variation in genome size seen in a wide variety of caridean shrimp groups more generally.</p>
FIGURE 2 in Parental origin and genome evolution of several Eurasian hexaploid species of Chenopodium (Chenopodiaceae)
FIGURE 2. Neighbour net analysis of nrITS sequences. The main clusters are shaded in color; (grey) A-genome diploids and Asian tetraploid C. sosnowskyi; (green) Eurasian diploids (B-genome), American and Asian tetraploids and Taiwan hexaploid C. formosanum; (yellow) two Asian diploids; (red) Asian diploid (C. acuminatum) and most of the Asian and Eurasian polyploids. Bootstrap values>80% are given only for the major groups of the analyzed accessions.
FIGURE 3 in Parental origin and genome evolution of several Eurasian hexaploid species of Chenopodium (Chenopodiaceae)
FIGURE 3. Neighbour net analysis of 5S rDNA NTS sequences. The main clusters are shaded in color; (grey) sequences from A-genome diploids and American allotetraploids; (green) sequences from Eurasian diploids (B-genome), American allotetraploids and all analyzed hexaploid; (red) sequences from Eurasian tetraploids (C. betaceum, C. striatiforme) and two Eurasian hexaploids; (orange) sequences of Eurasian tetraploids (C. betaceum, C. striatiforme) and three Eurasian hexaploids. Bootstrap values>80% are given only for the major groups of the analyzed accessions.
FIGURE 1 in Parental origin and genome evolution of several Eurasian hexaploid species of Chenopodium (Chenopodiaceae)
FIGURE 1. Neighbour net analysis of cpDNA sequences. The main clusters are shaded in color; (grey) A-genome diploids and American allotetraploids with AABB genome composition; (red) Eurasian tetraploids (C. betaceum and C. striatiforme) and Eurasian hexaploid species (C. pedunculare, C. album and C. giganteum); (green) Eurasian diploids (B-genome) and Taiwan hexaploid C. formosanum. Bootstrap values>80% are given only for the major groups of the analyzed accessions.
FIGURE 6 in Parental origin and genome evolution of several Eurasian hexaploid species of Chenopodium (Chenopodiaceae)
FIGURE 6. Evolution of rDNA loci and genome size in allohexaploids C. album, C. giganteum, C. pedunculare, C. formosanum and C. opulifolium. Putative diploid and tetraploid ancestors are indicated by green and grey arrows, respectively. Genome size (1C), subgenome B (green), chromosome carrying 35S (yellow), or 5S (magenta) loci and their putative provenance (♂ for paternal and ♀ for maternal parent), as well as a type of ITS and 5S rDNA non-transcribed spacer (NTS) sequences (♂ for paternal and ♀ for maternal parent) are shown.
FIGURE 4 in Parental origin and genome evolution of several Eurasian hexaploid species of Chenopodium (Chenopodiaceae)
FIGURE 4. Somatic metaphases of (a) C. album (BGG 4356), (b) C. pedunculare, (c) C. giganteum (PI 596372), (d) C. formosanum and (e) C. opulifolium after double GISH with gDNA isolated from C. ficifolium (green) and tetraploid C. betaceum (red). The scale bar=5 μm.
FIGURE 5 in Parental origin and genome evolution of several Eurasian hexaploid species of Chenopodium (Chenopodiaceae)
FIGURE 5. Distribution of 35S (yellow) and 5S rDNA (magenta) loci in chromosomes of Chenopodium species; (a) C. betaceum; (b) C. striatiforme; (c) C. giganteum (IPK Chen61), (d) C. pedunculare. Somatic metaphases of (e) C. album (UHBG 36), (f) C. pedunculare, (g) C. formosanum, (h–i) C. giganteum (PI 596372) and (j) C. opulifolium after GISH/FISH with gDNA isolated from C. ficifolium (green), 35S rDNA (yellow) and 5S rDNA (magenta). The scale bar=5 μm.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.