Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
899
datasets available to search
ShareScore release 0.9.0
Dataset results
899 results for “allele”
Data from: Temperature and sex related effects of serine protease alleles on larval development in the Glanville fritillary butterfly
Open the record for dataset details and reuse information.
Data from: MHC-mediated sexual selection on bird song: generic polymorphism, particular alleles and acoustic signals
Open the record for dataset details and reuse information.
Native American genetic ancestry and pigmentation allele contributions to skin color in a Caribbean population
Open the record for dataset details and reuse information.
Data from: Long-distance dispersal suppresses introgression of local alleles during range expansions
Open the record for dataset details and reuse information.
Data from: Natural allelic variations of xenobiotic metabolizing enzymes affect sexual dimorphism in Oryzias latipes
Open the record for dataset details and reuse information.
Data from: Controlling for p-value inflation in allele frequency change in experimental evolution and artificial selection experiments
Open the record for dataset details and reuse information.
Scoring of 13 microsatellite loci for Tetrastigma loheri in Cebu (Philippines) based on the fragment length size of their respective alleles
Open the record for dataset details and reuse information.
Natural hybridization reveals incompatible alleles that cause melanoma in swordtail fish
Open the record for dataset details and reuse information.
Is drought tolerance a domestication trait in tepary bean?: Allelic diversity at abiotic stress responsive genes in cultivated Phaseolus acutifolius A. Gray and its wild relatives
<p>Some of the major impacts of climate change are expected in the poorest regions of the world where drought stress and nutrient deficiency are already a main issue. Legumes are an essential food crop for the poorest because of their high dietary protein and micronutrient contents. However, they are generally drought susceptible. Therefore, our goal in this study was to explore allele diversity at abiotic stress responsive candidate genes in the only drought tolerant cultivated bean species of the genus <i>Phaseolus</i>, tepary bean (<i>P. acutifolius</i> A. Gray) and its related species <i>P. parvifolius </i>Freytag. Specifically, we estimated drought tolerance in 52 tepary bean <i>s.l.</i> geo-referenced germplasm accessions from the <i>P. acutifolius</i>–<i>parvifolius</i> clade using climate information, and used this estimated drought stress index to examine allele correlations with <i>Asr2</i>, <i>Dreb2B</i> and ERECTA-encoding candidate genes for drought tolerance. Genetic clustering showed that cultivated and wild <i>P. acutifolius</i> were intermingled with <i>P. acutifolius </i>var.<i> tenuifolius</i> and <i>P. parvifolius</i>, signifying that allele diversity at candidate genes for drought tolerance was not scarce in tepary bean <i>s.l</i>. <i>Dreb2B</i> and ERECTA-encoding genes harbored signatures of directional/purifying selection, likely in favor of adaptive alleles selectively advantageous because each had two SNPs significantly correlated (<i>p-value</i> < 0.05) with habitat drought stress at six and 12 months. These results suggest that tepary bean <i>s.l. </i>is a reservoir of novel alleles at candidate genes for drought tolerance, as expected for a drought-tolerant species that originated in warmer and arid environments. Abiotic stress responsive candidate genes also exhibit comparable patterns of selective signatures when comparing orthologous across species, which speaks for a predominant role of gene sub-functionalization likely due to ecological constrains. Our study therefore corroborates that the candidate gene approach is still an effective alternative for marker validation across a broader genetic basis of germplasm accessions. Further efforts to determine the genetic architecture of drought tolerance will unlock novel alleles hidden in a crop with limited modern relevance as tepary bean, but capable of acting as a donor in backcrossing and genome editing strategies with elite common bean lines aiming to meet the imminent demands of a drier world.</p>
Reference allele frequencies for populations pools of Atlantic Herring (Clupea harengus)
<p>Atlantic herring is widespread in North Atlantic and adjacent waters and is one of the most abundant vertebrates on earth. This species is well suited to explore genetic adaptation due to minute genetic differentiation at selectively neutral loci. Here we report hundreds of loci underlying ecological adaptation to different geographic areas and spawning conditions. Four of these represent megabase inversions confirmed by long read sequencing. The genetic architecture underlying ecological adaptation in herring deviates from expectation under a classical infinitesimal model for complex traits because of large shifts in allele frequencies at hundreds of loci under selection.</p>
Data from: Recombination and hitchhiking of deleterious alleles
When new advantageous alleles arise and spread within a population, deleterious alleles at neighbouring loci can hitchhike alongside them and spread to fixation in areas of low recombination, introducing a fixed mutation load. We use branching processes and diffusion equations to calculate the probability that a deleterious allele hitchhikes and fixes alongside an advantageous mutant. As expected, the probability of fixation of a deleterious hitchhiker rises with the selective advantage of the sweeping allele and declines with the selective disadvantage of the deleterious hitchhiker. We then use computer simulations of a genome with an infinite number of loci to investigate the increase in load after an advantageous mutant is introduced. We show that the appearance of advantageous alleles on genetic backgrounds loaded with deleterious alleles has two potential effects: it can fix deleterious alleles and also facilitate the persistence of recombinant lineages that happen to occur. The latter is expected to reduce the signals of selection in the surrounding region. We consider these results in light of human genetic data to infer how likely it is that such deleterious hitchhikers have occurred in our recent evolutionary past.
Data from: Contrasting dynamics of a mutator allele in asexual populations of differing size
Mutators have been shown to hitchhike in asexual populations when the anticipated beneficial mutation supply rate of the mutator subpopulation, NU_b (for subpopulation of size N and beneficial mutation rate U_b) exceeds that of the wild-type subpopulation. Here, we examine the effect of total population size on mutator dynamics in asexual experimental populations of Saccharomyces cerevisiae. Although mutators quickly hitchhike to fixation in smaller populations, mutator fixation requires more and more time as population size increases; this observed delay in mutator hitchhiking is consistent with the expected effect of clonal interference. Interestingly, despite their higher beneficial mutation supply rate, mutators are supplanted by the wild type in very large populations. We postulate that this striking reversal in mutator dynamics is caused by an interaction between clonal interference, the fitness cost of the mutator allele, and infrequent large-effect beneficial mutations in our experimental populations. Our work thus identifies a potential set of circumstances under which mutator hitchhiking can be inhibited in natural asexual populations, despite recent theoretical predictions that such populations should have a net tendency to evolve ever-higher genomic mutation rates.
Data from: Genotype-free estimation of allele frequencies reduces bias and improves demographic inference from RADSeq data
Restriction-site associated sequencing (RADSeq) facilitates rapid generation of thousands of genetic markers at relatively low cost; however, several sources of error specific to RADSeq methods often lead to biased estimates of allele frequencies and thereby to erroneous population genetic inference. Estimating the distribution of sample allele frequencies without calling genotypes was shown to improve population inference from whole genome sequencing data, but the ability of this approach to account for RADSeq-specific biases remains unexplored. Here we assess in how far genotype-free methods of allele frequency estimation affect demographic inference from empirical RADSeq data. Using the well-studied pied flycatcher (Ficedula hypoleuca) as a study system, we compare allele frequency estimation and demographic inference from whole genome sequencing data with that from RADSeq data matched for samples using both genotype-based and genotype free methods. The demographic history of pied flycatchers as inferred from RADSeq data was highly congruent with that inferred from WGS data when allele frequencies were estimated directly from the read data. In contrast, when allele frequencies were derived from called genotypes, RADSeq-based estimates of most model parameters fell outside the 95% confidence interval (CI) of estimates derived from WGS data. Notably, more stringent filtering of genotypes tended to increase the discrepancy between parameter estimates from WGS and RADSeq data, respectively. The results from this study demonstrate the ability of genotype-free methods to improve AFS-based demographic inference from RADSeq data and highlight the need to account for uncertainty in NGS data regardless of sequencing method.
Data from: Phylogeny of Leavenworthia S-alleles suggests unidirectional mating system evoution and enhanced positive selection following an ancient population bottleneck
The adoption of self-fertilization from an ancestral outcrossing state is one of the most common evolutionary transitions in the flowering plants. In the mustard family, outcrossing is typically enforced by sporophytic self-incompatibility (SI), but there are also many self-compatible species. The genus Leavenworthia contains taxa that either possess or lack SI. Here we present data showing that SI is associated with strict outcrossing and that there is widespread trans-specific sequence polymorphism at the locus involved in the recognition of self-pollen (the S-locus). This ancestral polymorphism is consistent with the presence of an outcrossing mating system in the common ancestor of Leavenworthia species, and suggests that there have been several independent losses of SI in the group. When compared with other mustard species, the bulk of Leavenworthia S-allele sequences are highly diverged from those found in other Brassicaceae and show relatively low levels of nucleotide diversity, a pattern that suggests the common ancestor of the genus likely underwent a strong population bottleneck. The hypothesis of post-bottleneck S-locus rediversification is supported by tests showing stronger positive selection acting on S-alleles from Leavenworthia than those found in other Brassicaceae.
Data from: Distribution of MICB diversity in the Zhejiang Han population: PCR sequence-based typing for exons 2–6 and identification of five novel MICB alleles
The polymorphism of major histocompatibility complex class I chain-related gene B (MICB) and variations in MICB alleles in a variety of populations have been characterized using several genotyping approaches. In the present study, a novel polymerase chain reaction sequence-based typing (PCR-SBT) method was established for the genotyping of MICB exons 2–6, and the allelic frequency of MICB in the Zhejiang Han population was investigated. Among 400 unrelated healthy Han individuals from Zhejiang Province, China, a total of 20 MICB alleles were identified, of which MICB*005:02:01, MICB*002:01:01, and MICB*004:01:01 were the most predominant alleles, with frequencies of 0.57375, 0.1225, and 0.08375, respectively. Nine MICB alleles were detected on only one occasion, giving a frequency of 0.00125. Of the 118 distinct MICB ∼ HLA-B haplotypes identified, 42 showed significant linkage disequilibrium (P < 0.05). Haplotypes MICB*005:02:01 ∼ B*46:01, MICB*005:02:01 ∼ B*40:01, and MICB*008 ∼ B*58:01 were the most common haplotypes, with frequencies of 0.0978, 0.0761, and 0.0616, respectively. Five novel alleles, MICB*005:07, MICB*005:08, MICB*027, MICB*028, and MICB*029 were identified. Compared with the MICB*005:02:01 sequence, a G > A substitution was observed at nucleotide position 210 in MICB*005:07, and a 1,134 T > C substitution in MICB*005:08 and an 862 G > A substitution in MICB*027 were detected. In addition, it appears that MICB*028 probably arose from MICB*004:01:01 with an A to G substitution at position 1,147 in exon 6. MICB*029 had a G > T transversion at nucleotide position 730 in exon 4, compared with that of MICB*002:01:01. On the basis of the new PCR-SBT assay, these observed results demonstrated MICB allelic variations in the Zhejiang Han population.
Data from: Minor allele frequency thresholds strongly affect population structure inference with genomic datasets
One common method of minimizing errors in large DNA sequence datasets is to drop variable sites with a minor allele frequency below some specified threshold. Though widespread, this procedure has the potential to alter downstream population genetic inferences and has received relatively little rigorous analysis. Here we use simulations and an empirical SNP dataset to demonstrate the impacts of minor allele frequency (MAF) thresholds on inference of population structure. We find that model-based inference of population structure is confounded when singletons are included in the alignment, and that both model-based and multivariate analyses infer less distinct clusters when more stringent MAF cutoffs are applied. We propose that this behavior is caused by the combination of a drop in the total size of the data matrix and by correlations between allele frequencies and mutational age. We recommend a set of best practices for applying MAF filters in studies seeking to describe population structure with genomic data.
Data from: Selection on structural allelic variation biases plasticity estimates
Wang and Althoff (2019) explored the capacity of Drosophila melanogaster to exhibit adaptive plasticity in a novel environment. In a full-sib, half-sib design, they scored the activity of the enzyme alcohol dehydrogenase (ADH) and plastic responses, measured as changes in ADH activity across ethanol concentrations in the range of 0-10% (natural variation) and 16% (the novel environment). ADH activity increased with alcohol concentration, and there was a positive association between larval viability and ADH activity in the novel environment. They also reported that families exhibiting greater plasticity had higher larval survival in the novel environment, concluding that ADH plasticity is adaptive. However, the four authors now concur that, since the study estimated plasticity from phenotypic differences across environments using full-sib families, it is not possible to disentangle the contributions of allele frequency changes at the Adh locus from regulatory control at loci known to influence ADH activity. Selective changes in allele frequencies may thus conflate estimates of plasticity; any type of "plasticity" (adaptive, neutral, or maladaptive) could be inferred depending on allele frequencies. The problem of scoring sib-groups after selection should be considered in any plasticity study that cannot use replicated genotypes. Researchers should monitor changes in allele frequencies as one mechanism to deal with this issue.
Data from: 'True' null allele detection in microsatellite loci: a comparison of methods, assessment of difficulties, and survey of possible improvements
Null alleles are alleles that for various reasons fail to amplify in a PCR assay. The presence of null alleles in microsatellite data is known to bias the genetic parameter estimates. Thus, efficient detection of null alleles is crucial, but the methods available for indirect null allele detection return inconsistent results. Here, our aim was to compare different methods for null allele detection, to explain their respective performance and to provide improvements. We applied several approaches to identify the 'true' null alleles based on the predictions made by five different methods, used either individually or in combination. First, we introduced simulated 'true' null alleles into 240 population data sets and applied the methods to measure their success in detecting the simulated null alleles. The single best-performing method was ML-NullFreq_frequency. Furthermore, we applied different noise reduction approaches to improve the results. For instance, by combining the results of several methods, we obtained more reliable results than using a single one. Rule-based classification was applied to identify population properties linked to the false discovery rate. Rules obtained from the classifier described which population genetic estimates and loci characteristics were linked to the success of each method. We have shown that by simulating 'true' null alleles into a population data set, we may define a null allele frequency threshold, related to a desired true or false discovery rate. Moreover, using such simulated data sets, the expected null allele homozygote frequency may be estimated independently of the equilibrium state of the population.
Data from: Environmental complexity and the purging of deleterious alleles
Sexual interactions among adults can generate selection on both males and females with genome-wide consequences. Sexual selection through males is one component of this selection that has been argued to play an important role in purging deleterious alleles. A common technique to assess the influence of sexual selection is by a comparison of experimental evolution under enforced monogamy vs. polygamy. Mixed results from past studies may be due to the use of highly simplified lab conditions that alter the nature of sexual interactions. Here we examine the rate of purging of 22 gene disruption mutations in experimental polygamous populations of Drosophila melanogaster in each of two mating environments: a simple, high density environment (i.e., typical fly vials) and a lower density, more spatially complex environment. Based on past work, we expect sexual interactions in the latter environment to result in stronger selection in both sexes. Consistent with this, we find that mutations tend to be purged more quickly in populations evolving in complex environments. We discuss possible mechanisms by which environmental complexity might modulate the rate at which deleterious alleles are purged and putatively ascribe a role for sexual interactions in explaining the treatment differences in our experiment.
Data from: Allelic diversity for neutral markers retains a higher adaptive potential for quantitative traits than expected heterozygosity
The adaptive potential of a population depends on the amount of additive genetic variance for quantitative traits of evolutionary importance. This variance is a direct function of the expected frequency of heterozygotes for the loci which affect the trait (QTL). It has been argued, but not demonstrated experimentally, that long-term response to selection is more dependent on QTL allelic diversity than on QTL heterozygosity. Conservation programmes, aimed at preserving this variation, usually rely on neutral markers rather than on quantitative traits for making decisions on management. Here, we address, both through simulation analyses and experimental studies with Drosophila melanogaster, the question of whether allelic diversity for neutral markers is a better indicator of a high adaptive potential than expected heterozygosity. In both experimental and simulation studies, we established synthetic populations for which either heterozygosity or allelic diversity was maximized using information from QTL (simulations) or unlinked neutral markers (simulations and experiment). The synthetic populations were selected for the quantitative trait to evaluate the evolutionary potential provided by the two optimization methods. Our results show that maximizing the number of alleles of a low number of markers implies higher responses to selection than maximizing their heterozygosity.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.