Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
899
datasets available to search
ShareScore release 0.9.0
Dataset results
899 results for “allele”
Data from: Epistasis and allele specificity in the emergence of a stable polymorphism in Escherichia coli
Ecological opportunities promote population divergence into coexisting lineages. However, the genetic mechanisms that enable new lineages to exploit these opportunities are poorly understood except in cases of single mutations. We examined how two Escherichia coli lineages diverged from their common ancestor at the outset of a long-term coexistence. By sequencing genomes and reconstructing the genetic history of one lineage, we showed that three mutations together were sufficient to produce the frequency-dependent fitness effects that allowed this lineage to invade and stably coexist with the other. These mutations all affected regulatory genes and collectively caused substantial metabolic changes. Moreover, the particular derived alleles were critical for the initial divergence and invasion, indicating that the establishment of this polymorphism depended on specific epistatic interactions.
Data from: Negative frequency-dependent selection of sexually antagonistic alleles in Myodes glareolus
Sexually antagonistic genetic variation, where optimal values of traits are sex-dependent, is known to slow the loss of genetic variance associated with directional selection on fitness-related traits. However, sexual antagonism alone is not sufficient to maintain variation indefinitely. Selection of rare forms within the sexes can help to conserve genotypic diversity. We combined theoretical models and a field experiment with Myodes glareolus to show that negative frequency-dependent selection on male dominance maintains variation in sexually antagonistic alleles. In our experiment, high-dominance male bank voles were found to have low-fecundity sisters, and vice versa. These results show that investigations of sexually antagonistic traits should take into account the effects of social interactions on the interplay between ecology and evolution, and that investigations of genetic variation should not be conducted solely under laboratory conditions.
Data from: Known mutator alleles do not markedly increase mutation rate in clinical Saccharomyces cerevisiae strains
Natural selection has the potential to act on all phenotypes, including genomic mutation rate. Classic evolutionary theory predicts that in asexual populations, mutator alleles, which cause high mutation rates, can fix due to linkage with beneficial mutations. This phenomenon has been demonstrated experimentally and may explain the frequency of mutators found in bacterial pathogens. By contrast, in sexual populations, recombination decouples mutator alleles from beneficial mutations, preventing mutator fixation. In the facultatively sexual yeast Saccharomyces cerevisiae, segregating alleles of MLH1 and PMS1 have been shown to be incompatible, causing a high mutation rate when combined. These alleles had never been found together naturally, but were recently discovered in a cluster of clinical isolates. Here we report that the incompatible mutator allele combination only marginally elevates mutation rate in these clinical strains. Genomic and phylogenetic analyses provide no evidence of a historically elevated mutation rate. We conclude that the effect of the mutator alleles is dampened by background genetic modifiers. Thus, the relationship between mutation rate and microbial pathogenicity may be more complex than once thought. Our findings provide rare observational evidence that supports evolutionary theory suggesting that sexual organisms are unlikely to harbour alleles that increase their genomic mutation rate.
Data from: An exceptionally high nucleotide and haplotype diversity and a signature of positive selection for the eIF4E resistance gene in barley are revealed by allele mining and phylogenetic analyses of natural populations.
In barley, the eukaryotic translation initiation factor 4E (eIF4E) gene situated on chromosome 3H is recognised as an important source of resistance to the bymoviruses Barley yellow mosaic virus and Barley mild mosaic virus. In modern barley cultivars two recessive eIF4E alleles, rym4 and rym5, confer different isolate-specific resistances. In this study the sequence of eIF4E was analysed in 1090 barley landraces and non-current cultivars originating from 84 countries. An exceptionally high nucleotide diversity was evident in the coding sequence of eIF4E but not in either the adjacent MCT-1 gene or the sequence related eIF(iso)4E gene situated on chromosome 1H. Surprisingly, all nucleotide polymorphisms detected in the coding sequence of eIF4E resulted in amino acid changes. A total of 47 eIF4E haplotypes were identified and phylogenetic analysis using maximum likelihood provided evidence of strong positive selection acting on this barley gene. The majority of eIF4E haplotypes were found to be specific to distinct geographic regions. Furthermore, the eIF4E haplotype diversity (uh) was found to be considerably higher in East Asia, whereas SNP genotyping identified a comparatively low degree of genome-wide genetic diversity in 16 out of 17 tested accessions (each carrying a different eIF4E haplotype) from this same region. In addition, selection statistic calculations using coalescent simulations showed evidence of non neutral variation for eIF4E in several geographic regions, including East Asia, the region with a long history of the bymovirus-induced yellow mosaic disease. Together these findings suggest eIF4E may play a role in barley adaptation to local habitats.
Data from: Parapatric speciation in three islands: dynamics of geographical configuration of allele sharing
We studied the time to speciation by geographical isolation for a species living on three islands connected by rare migration. We assumed that incompatibility was controlled by a number of quantitative loci and that individuals differing in loci by more than a threshold did not mix genetically with each other. For each locus, we defined the geographical configuration (GC), which specifies islands with common alleles, and traced the stochastic transitions between different GCs. From these results, we calculated the changes in genetic distances. As a single migration event provides an opportunity for transitions in multiple loci, the GCs of different loci are correlated, which can be evaluated by constructing the stochastic differential equations of the number of loci with different GCs. Our model showed that the low number of incompatibility loci facilitates parapatric speciation and that migrants arriving as a group shorten the waiting time to speciation compared with the same number of migrants arriving individually. We also discuss how speciation rate changes with geographical structure.
Data from: Joint allelic effects on fitness and metric traits
Theoretical explanations of empirically observed standing genetic variation, mutation, and selection suggest that many alleles must jointly affect fitness and metric traits. However, there are few direct demonstrations of the nature and extent of these pleiotropic associations. We implemented a mutation accumulation (MA) divergence experimental design in Drosophila serrata to segregate genetic variants for fitness and metric traits. By exploiting naturally occurring MA line extinctions as a measure of line-level total fitness, manipulating sexual selection, and measuring productivity we were able to demonstrate genetic covariance between fitness and standard metric traits, wing size and shape. Larger size was associated with lower total fitness and male sexual fitness, but higher productivity. Multivariate wing shape traits, capturing major axes of wing shape variation among MA lines, evolved only in the absence of sexual selection, and to the greatest extent in lines that went extinct, indicating that mutations contributing wing shape variation also typically had deleterious effects on both total fitness and male sexual fitness. This pleiotropic covariance of metric traits with fitness will drive their evolution, and generate the appearance of selection on the metric traits even in the absence of a direct contribution to fitness.
Data from: Sexual reproduction in Aspergillus flavus sclerotia: acquisition of novel alleles from soil populations and uniparental mitochondrial inheritance
Aspergillus flavus colonizes agricultural commodities worldwide and contaminates them with carcinogenic aflatoxins. The high genetic diversity of A. flavus populations is largely due to sexual reproduction characterized by the formation of ascospore-bearing ascocarps embedded within sclerotia. A. flavus is heterothallic and laboratory crosses between strains of the opposite mating type produce progeny showing genetic recombination. Sclerotia formed in crops are dispersed onto the soil surface at harvest and are predominantly produced by single strains of one mating type. Less commonly, sclerotia may be fertilized during co-infection of crops with sexually compatible strains. In this study, laboratory and field experiments were performed to examine sexual reproduction in single-strain and fertilized sclerotia following exposure of sclerotia to natural fungal populations in soil. Female and male roles and mitochondrial inheritance in A. flavus were also examined through reciprocal crosses between sclerotia and conidia. Single-strain sclerotia produced ascospores on soil and progeny showed biparental inheritance that included novel alleles originating from fertilization by native soil strains. Sclerotia fertilized in the laboratory and applied to soil before ascocarp formation also produced ascospores with evidence of recombination in progeny, but only known parental alleles were detected. In reciprocal crosses, sclerotia and conidia from both strains functioned as female and male, respectively, indicating A. flavus is hermaphroditic, although the degree of fertility depended upon the parental sources of sclerotia and conidia. All progeny showed maternal inheritance of mitochondria from the sclerotia. Compared to A. flavus populations in crops, soil populations would provide a higher likelihood of exposure of sclerotia to sexually compatible strains and a more diverse source of genetic material for outcrossing.
Data from: The impact of library preparation protocols on the accuracy of allele frequency estimates in Pool-Seq data
Sequencing pools of individuals (Pool-Seq) is a cost-effective method to determine genome-wide allele frequency estimates. Given the importance of meta-analyses combining data sets, we determined the influence of different genomic library preparation protocols on the consistency of allele frequency estimates. We found that typically no more than 1% of the variation in allele frequency estimates could be attributed to differences in library preparation. Also read length had only a minor effect on the consistency of allele frequency estimates. By far, the most pronounced influence could be attributed to sequence coverage. Increasing the coverage from 30- to 50-fold improved the consistency of allele frequency estimates by at least 27%. We conclude that Pool-Seq data can be easily combined across different library preparation methods, but sufficient sequence coverage is key to reliable results.
Data from: How species evolve collectively: implications of gene flow and selection for the spread of advantageous alleles
The traditional view that species are held together through gene flow has been challenged by observations that migration is too restricted among populations of many species to prevent local divergence. However, only very low levels of gene flow are necessary to permit the spread of highly advantageous alleles, providing an alternative means by which low-migration species might be held together. We re-evaluate these arguments given the recent and wide availability of indirect estimates of gene flow. Our literature review of Fst values for a broad range of taxa suggests that gene flow in many taxa is considerably greater than suspected from earlier studies and often is sufficiently high to homogenize even neutral alleles. However, there are numerous species from essentially all organismal groups that lack sufficient gene flow to prevent divergence. Crude estimates on the strength of selection on phenotypic traits and effect sizes of quantitative trait loci (QTL) suggest that selection coefficients for leading QTL underlying phenotypic traits may be high enough to permit their rapid spread across populations. Thus, species may evolve collectively at major loci through the spread of favourable alleles, while simultaneously differentiating at other loci due to drift and local selection.
Relate-estimated coalescence rates and allele ages for European beef cattle
<h1>Overview</h1> <p>Coalescence rates and allele ages calculated for five beef cattle breeds (Charolais, Simmental, Limousin, Hereford, Angus) using Relate.</p> <p>We estimated the joint genealogy of 684 individuals, including both <em>Bos taurus</em> and <em>Bos indicus</em>, then extracted the embedded genealogy for the five cattle breeds and re-estimated the population size history and branch lengths.</p> <p>We extracted the embedded genealogy for each of the five beef cattle breeds, jointly estimated the population size history, and re-estimated the branch lengths.</p> <p>Please email bft990914@163.com for any queries.</p> <h1>Coalescence Rates and Allele Ages</h1> <p>The *.coal files record coalescence rates for each of the five cattle breeds.</p> <p>The gzipped files allele_ages_*.gz record allele ages for each of the five cattle breeds.</p>
Sex-specific splicing of Z- and W-borne nr5a1 alleles suggests sex determination is controlled by chromosome conformation
<p><i>Pogona vitticeps</i> has female heterogamety (ZZ/ZW) but the master sex determining gene is unknown, as is the case for all reptiles. We show that <i>nr5a1</i>, a gene that is essential in mammalian sex determination, has alleles on the Z and W chromosomes (Z-<i>nr5a1</i> and W-<i>nr5a1</i>), which are both expressed and can recombine. Three transcript isoforms of Z-<i>nr5a1</i> were detected in gonads of adult ZZ males, two of which encode a functional protein. However, ZW females produced sixteen isoforms, most of which contained premature stop codons. The array of transcripts produced by the W-borne allele (W-<i>nr5a1</i>) is likely to produce truncated polypeptides that could act as a competitive inhibitor to the full-length intact protein. We hypothesize that an altered configuration of the W chromosomes affects the conformation of the primary transcript generating inhibitory W-borne isoforms that suppress testis determination. Under this hypothesis, the GSD system of <i>P. vitticeps</i> is a W-borne dominant female-determiner that may be controlled epigenetically.</p>
Fitness benefit plays a vital role in the retention of the Pi-ta susceptible alleles
<p>In plants, large numbers of <i>R</i> genes, which segregate as loci with alternative alleles conferring different resistance to pathogens, have been maintained over a long evolutionary time. In theory, there seem to be no reason for hosts to harbor susceptible alleles in view of their null contribution to resistance. As such, why should populations support disease-susceptible individuals along with disease-resistant individuals? In rice, a single copy gene <i>Pi-at</i> segregates for two expressed clades of alleles, one resistant and the other susceptible. We simulated loss-of-function of the <i>Pi-ta</i> susceptible allele using the CRISPR/Cas9 system to detect subsequent fitness changes and obtained insights into fitness effects on retention of the <i>Pi-ta</i> susceptible allele. Our creation of artificial knockout of the <i>Pi-ta</i> susceptible allele suffered a fitness decline of up to 49% in term of filled grains yield upon the loss of <i>Pi-ta</i>'s function. The <i>Pi-ta</i> susceptible alleles might serve as an off-switch to the downstream immune signaling, thus contributing to fine-tuning of plant defense response. These findings highlight the interplay between genetic architecture and fitness effects of segregating <i>R</i> gene alleles and also provide a plausible explanation how host genomes can tolerate the possible genetic load associated with a vast repertoire of <i>R</i> genes. This attempt to evaluate the fitness effect of the <i>R</i> gene in crop will bring some clues to researchers and breeders that not all disease resistant genes will bring fitness cost as universally acknowledged.</p>
Supporting data and code for: Distribution of invasive versus native whitefly species and their pyrethroid knock-down resistance allele in a context of interspecific hybridization
<p>This is the first release of the final data and code for the article accepted for publication in Scientific Reports journal. It contains all the necessary scripts to produce most of the analyses and figures of the manuscript. All the necessary data can be found in the 'data' folder.</p>
Raw allelic matrix and supplementary materials: Origin and dispersion pathways of guava in the Galapagos Islands inferred through genetics and historical records
<p>Guava (<i>Psidium guajava</i>) is an aggressive invasive plant in the Galapagos Islands. Determining its provenance and genetic diversity could explain its adaptability and spread, and how this relates to past human activities. With this purpose, we analyzed 11 SSR markers in guava individuals from Isabela, Santa Cruz, San Cristobal and Floreana islands in the Galapagos, as well as from mainland Ecuador. The mainland guava population appeared genetically differentiated from the Galapagos populations, with higher genetic diversity levels found in the former. We consistently found that the Central Highlands region of mainland Ecuador is one of the most likely origins of the Galapagos populations. Moreover, the guavas from Isabela and Floreana show a potential genetic input from southern mainland Ecuador, while the population from San Cristobal would be linked to the coastal mainland regions. Interestingly, the proposed origins for the Galapagos guava coincide with the first human settlings of the archipelago. Through Approximate Bayesian Computation, we propose a model where San Cristobal was the first island to be colonized by guava from the mainland, then it would have spread to Floreana and finally to Santa Cruz; Isabela would have been seeded from Floreana. An independent trajectory could also have contributed in the invasion of Floreana and Isabela. The pathway shown in our model agrees with the human colonization history of the different islands in the Galapagos. Our model, in conjunction with the clustering patterns of the individuals (based on genetic distances), suggests that guava introduction history in the Galapagos archipelago was driven by either a single event or a series of introduction events in rapid succession. We thus show that genetic analyses supported by historical sources can be used to track the arrival and spread of invasive species in novel habitats and the potential role of human activities in such processes.</p>
Recla 8-state founder allele dosages
<p>8-state founder allele dosages for the Recla data. Created by R/qtl2 (https://kbroman.org/qtl2/).</p>
Data from: Fine mapping of dominant X-linked incompatibility alleles in Drosophila hybrids
Sex chromosomes have a large effect on reproductive isolation and play an important role in hybrid inviability. In Drosophila hybrids, X-linked genes have pronounced deleterious effects on fitness in male hybrids, which have only one X chromosome. Several studies have succeeded at locating and identifying recessive X-linked alleles involved in hybrid inviability. Nonetheless, the density of dominant X-linked alleles involved in interspecific hybrid viability remains largely unknown. In this report, we study the effects of a panel of small fragments of the D. melanogaster X-chromosome carried on the D. melanogaster Y-chromosome in three kinds of hybrid males: D. melanogaster/D. santomea, D. melanogaster/D. simulans and D. melanogaster/D. mauritiana. D. santomea and D. melanogaster diverged over 10 million years ago, while D. simulans (and D. mauritiana) diverged from D. melanogaster over 3 million years ago. We find that the X-chromosome from D. melanogaster carries dominant alleles that are lethal in mel/san, mel/sim, and mel/mau hybrids, and more of these alleles are revealed in the most divergent cross. We then compare these effects on hybrid viability with two D. melanogaster intraspecific crosses. Unlike the interspecific crosses, we found no X-linked alleles that cause lethality in intraspecific crosses. Our results reveal the existence of dominant alleles on the X-chromosome of D. melanogaster which cause lethality in three different interspecific hybrids. These alleles only cause inviability in hybrid males, yet have little effect in hybrid females. This suggests that X-linked elements that cause hybrid inviability in males might not do so in hybrid females due to differing sex chromosome interactions.
A Streamlined and High-Throughput Error-Corrected Next-Generation Sequencing Method for Low Variant Allele Frequency Quantitation
<p></p><p>Quantifying mutant or variable allele frequencies (VAFs) of ≤10−3 using next-generation sequencing (NGS) has utility in both clinical and nonclinical settings. Two common approaches for quantifying VAFs using NGS are tagged single-strand sequencing and duplex sequencing. While duplex sequencing is reported to have sensitivity up to 10−8 VAF, it is not a quick, easy, or inexpensive method. We report a method for quantifying VAFs that are ≥10−4 that is as easy and quick for processing samples as standard sequencing kits, yet less expensive than the kits. The method was developed using PCR fragment-based VAFs of Kras codon 12 in log10 increments from 10−5 to 10−1, then applied and tested on native genomic DNA. For both sources of DNA, there is a proportional increase in the observed VAF to input VAF from 10−4 to 100% mutant samples. Variability of quantitation was evaluated within experimental replicates and shown to be consistent across sample preparations. The error at each successive base read was evaluated to determine if there is a limit of read length for quantitation of ≥10−4, and it was determined that read lengths up to 70 bases are reliable for quantitation. The method described here is adaptable to various oncogene or tumor suppressor gene targets, with the potential to implement multiplexing at the initial tagging step. While easy to perform manually, it is also suited for robotic handling and batch processing of samples, facilitating detection and quantitation of genetic carcinogenic biomarkers before tumor formation or in normal-appearing tissue.</p><p></p>
Data from: The genetics of adaptation to discrete heterogeneous environments: frequent mutation or large-effect alleles can allow range expansion
Range expansions are complex evolutionary and ecological processes. From an evolutionary standpoint, a populations' adaptive capacity can determine the success or failure of expansion. Using individual-based simulations, we model range expansion over a two-dimensional, approximately continuous landscape. We investigate the ability of populations to adapt across patchy environmental gradients and examine how the effect sizes of mutations influence the ability to adapt to novel environments during range expansion. We find that genetic architecture and landscape patchiness both have the ability to change the outcome of adaptation and expansion over the landscape. Adaptation to new environments succeeds via many mutations of small effect or few of large effect, but not via the intermediate between these cases. Higher genetic variance contributes to increased ability to adapt, but an alternative route of successful adaptation can proceed from low genetic variance scenarios with alleles of sufficiently large effect. Steeper environmental gradients can prevent adaptation and range expansion on both linear and patchy landscapes. When the landscape is partitioned into local patches with sharp changes in phenotypic optimum, the local magnitude of change between subsequent patches in the environment determines the success of adaptation to new patches during expansion.
Chinook salmon environmental data and allele frequency matrix
<p><span><span><span><span><span><span><span><span><span><span><span>Many species that undergo long breeding migrations, such as anadromous fishes, face highly heterogeneous environments along their migration corridors and at their spawning sites. These environmental challenges encountered at different life stages may act as strong selective pressures and drive local adaptation. However, the relative influence of environmental conditions along the migration corridor compared to the conditions at spawning sites on driving selection is still unknown. In this study, we performed genome-environment associations (GEA) to understand the relationship between landscape and environmental conditions driving selection in seven populations of the anadromous Chinook salmon (<i>Oncorhynchus tshawytscha)–</i>a species of important economic, social, cultural and ecological value–in the Columbia River basin. We extracted environmental variables for the shared migration corridors and at distinct spawning sites for each population, and used a Pool-seq approach to perform whole genome resequencing. Bayesian and univariate genome-environment association tests with migration-specific and spawning site-specific environmental variables indicated many more candidate SNPs associated with environmental conditions of the migration corridor compared to spawning sites. Specifically, variables associated with temperature, precipitation, terrain roughness, and elevation variables of the migration corridor were the most significant drivers of environmental selection. Additional analyses of neutral loci revealed two distinct clusters representing populations from different geographic regions of the drainage that also exhibit differences in adult migration timing (summer vs. fall). Tests for genomic regions under selection revealed a strong peak on chromosome 28, corresponding to the GREB1L/ROCK1 region that has been identified previously in salmonids as a region associated with adult migration timing. Our results show that environmental variation experienced throughout migration corridors imposed a greater selective pressure on Chinook salmon than environmental conditions at spawning sites.</span></span></span></span></span></span></span></span></span></span></span></p>
Dataset of laboratory research of paper "Mutation Analysis and Characteristics of the 5T Allele in the Cystic Fibrosis Trans Membrane Gene"
<p>Dataset contains</p> <p>1. Blood Lymphocite</p> <p>2. Utensils and raw Materials</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.