Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

63

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

63 results for “allele frequencies”

Learn how ShareScore rates datasets ↗
zenodo48/100

Gnomadv4.1 Enhanced Allele Frequencies (EAF) for use in PhyloFrame

<p>Source data to accompany manuscript: Equitable machine learning counteracts ancestral bias in precision medicine.</p> <pre><br><br><br></pre>

opencc-by-4.0Nov 2024View details →
zenodo44/100

Data for: Detecting Long-Term Balancing Selection Using Allele Frequency Correlation

<p>Genome-wide and top 1% scores for 1KG project data output from BetaScan reported in:</p> <p><a href="https://pubmed.ncbi.nlm.nih.gov/28981714/">Detecting Long-Term Balancing Selection Using Allele Frequency Correlation.</a></p> <p>Siewert KM, Voight BF. Mol Biol Evol. 2017 Nov 1;34(11):2996-3005. doi: 10.1093/molbev/msx209.</p> <p>PMID:&nbsp;28981714</p> <p>Code available at:&nbsp;https://github.com/ksiewert/BetaScan</p>

opencc-by-4.0Jul 2017View details →
zenodo44/100

A scalable, accurate, and universal analysis framework using individual-level allele frequency for large-scale genetic association studies in an admixed population

<p>Inclusion of individuals with diverse or admixed genetic ancestries is crucial to discover novel findings that may be missed by genomics analyses rooted solely in Caucasian population. Here, we present an analysis framework, SPAmix, which is scalable to a large-scale biobank data analysis including hundreds of thousands of admixed individuals and is universally applicable to various types of complex traits including binary trait, quantitative trait, time-to-event trait, longitudinal traits, etc. For each genetic variant, SPAmix uses genotype data and genetic principal components (PCs) to estimate individual-level allele frequency, which is subsequently used to calibrate p values via a retrospective analysis. A hybrid strategy including saddlepoint approximation (SPA) can greatly increase the accuracy to analyze rare genetic variants, especially if the phenotypic distribution is unbalanced or extremely unbalanced. Compared to Tractor, SPAmix does not require local ancestry information and can be straightforwardly applicable to a multi-way admixed population. Meanwhile, SPAmix can also be extended to SPAmix<sub>local</sub> in which the local ancestry can be incorporated if available. In addition, we propose SPAmix<sub>CCT</sub> to combine the p values of SPAmix and SPAmix<sub>local</sub> via Cauchy combination (CCT). SPAmix<sub>local</sub> performs close to Tractor when analyzing quantitative traits and is more accurate when analyzing binary traits with an unbalanced case-control ratio. And SPAmix<sub>CCT </sub>is an optimal unified approach for various cross-ancestry genetic architectures. Extensive simulation studies and real data analyses of 369,314 UK Biobank individuals from multiple ancestries demonstrated that SPAmix is scalable and can discover novel hits while controlling type I error rates well.</p>

opencc-by-4.0Sep 2023View details →
dryad40/100

Distinct signals of clinal and seasonal allele frequency change at eQTLs in Drosophila melanogaster

<p>Populations of short-lived organisms can respond to spatial and temporal environmental heterogeneity through local adaptation. However, the comparative signals of local adaptation across space and time remains poorly understood. Here, we examined patterns of allele frequency change across a latitudinal cline and between seasons at previously reported expression quantitative trait loci (eQTLs). We divided eQTLs into groups by utilizing differential expression profiles of fly populations collected across latitudinal clines or exposed to different environmental conditions. In general, we find that eQTLs are enriched for clinally varying polymorphisms, and that these eQTLs change in frequency in concordant ways across the cline and in response to starvation and chill-coma. The enrichment of eQTLs among seasonally varying polymorphisms is more subtle, and the direction of allele frequency change at eQTLs appears to be somewhat idiosyncratic. Taken together, we suggest that clinal adaptation at eQTLs is at least partially distinct from seasonal adaptation.</p>

opencc-zeroSep 2022View details →
zenodo40/100

Evolution of allele frequencies in two chicken lines divergently selected for meat ultimate pH

<p>Two lines of chicken were divergently selected during 5 generations for high or low meat ultimate pH. Genotypes at about 50K SNPs were obtained for a sample of individuals in each line and at each generation (including the founder population G0). The present dataset provides the allele frequencies for all SNP and generations in the two lines, at plink frq.strat format. The position of the SNP on the chicken genome are provided in another file at plink .map format.</p> <p>These data were first ued in the following publication:</p> <p>Le Bihan-Duval, E., Hennequet-Antier, C., Berri, C., Beauclercq, S. A., Bourin, M. C., Boulay, M., ... &amp; Boitard, S. (2018). Identification of genomic regions and candidate genes for chicken meat ultimate pH by combined detection of selection signatures and QTL. <em>BMC genomics</em>, <em>19</em>(1), 294.</p>

opencc-bySep 2019View details →
dryad40/100

Distinct signals of clinal and seasonal allele frequency change at eQTLs in Drosophila melanogaster

Open the record for dataset details and reuse information.

publicSep 2022View details →
zenodo36/100

Early Onset TAAD cohort logR Ratio and B allele frequency data

<p>Recurrent Rare Genomic Copy Number Variants and Bicuspid Aortic Valve Are Enriched in Early Onset Thoracic Aortic Aneurysms and Dissections</p> <p>Abstract:</p> <p>Thoracic Aortic Aneurysms and Dissections (TAAD) are a major cause of death in the United States. The spectrum of TAAD ranges from genetic disorders, such as Marfan syndrome, to sporadic isolated disease of unknown cause. We hypothesized that genomic copy number variants (CNVs) contribute causally to early onset TAAD (ETAAD). We conducted a genome-wide SNP array analysis of ETAAD patients of European descent who were enrolled in the National Registry of Genetically Triggered Thoracic Aortic Aneurysms and Cardiovascular Conditions (GenTAC). Genotyping was performed on the Illumina Omni-Express platform, using PennCNV, Nexus and CNVPartition for CNV detection. ETAAD patients (n = 108, 100% European American, 28% female, average age 20 years, 55% with bicuspid aortic valves) were compared to 7013 dbGAP controls without a history of vascular disease using downsampled Omni 2.5 data. For comparison, 805 sporadic TAAD patients with late onset aortic disease (STAAD cohort) and 192 affected probands from families with at least two affected relatives (FTAAD cohort) from our institution were screened for additional CNVs at these loci with SNP arrays. We identified 47 recurrent CNV regions in the ETAAD, FTAAD and STAAD groups that were absent or extremely rare in controls. Nine rare CNVs that were either very large (&gt;1 Mb) or shared by ETAAD and STAAD or FTAAD patients were also identified. Four rare CNVs involved genes that cause arterial aneurysms when mutated. The largest and most prevalent of the recurrent CNVs were at Xq28 (two duplications and two deletions) and 17q25.1 (three duplications). The percentage of individuals harboring rare CNVs was significantly greater in the ETAAD cohort (32%) than in the FTAAD (23%) or STAAD (17%) cohorts. We identified multiple loci affected by rare CNVs in one-third of ETAAD patients, confirming the genetic heterogeneity of TAAD. Alterations of candidate genes at these loci may contribute to the pathogenesis of TAAD.</p> <p>&nbsp;</p>

opencc-zeroJan 2016View details →
zenodo36/100

Evolution of allele frequencies in the cattle breed Asturiana de los Valles

<p>Genotype data for 153 animals from the Asturiana de los Valles bovine breed, with birth dates from 1980 to 2013. These genotypes were obtained from 3 different sources:&nbsp;</p> <p>- 88 sires were genotyped using the Illumina&rsquo;s BovineSNP50 v.2 chip; the resulting data are provided at plink ped/map format (asturiana_50K.tar.xz).</p> <p>- 50 animals (25 sires and 25 dams) were genotyped using the Illumina&rsquo;s Bovine High Density BeadChip 770K SNP; the resulting data are provided at plink ped/map format (asturiana_800K.tar.xz).</p> <p>- 15 sires were sequenced on a HiSeq 3,000; the resulting genotype calls are provided at vcf format (asturiana_WGS_SNP.vcf.gz).</p> <p>Boitard et al (2021) combined these 3 datasets in order to detect recent and historical selection signatures in this breed. The scripts used for this analysis can be found at https://github.com/sboitard/Asturiana_analysis.</p>

opencc-by-4.0May 2022View details →
zenodo36/100

Data and code from "Temporal allele frequency changes in large-effect loci reveal potential fishing impacts on salmon life-history diversity" (Miettinen et al. 2024)

<p>This archive contains code and data files to perform analyses detailed in Miettinen et al. (2024, Evolutionary Applications, https://doi.org/10.1111/eva.13690).</p>

opencc-by-4.0Apr 2024View details →
zenodo36/100

1000 Genomes Autosomal Allele Frequency Differences Across Ancestries (build 38)

<p>This dataset contains chromosome-specific allele frequency differences across different genetic ancestry groups for 73,159,508 autosomal SNPs captured in 2,548 individuals from the 1000 Genomes Project (build 38), available <a href="http://ftp.1000genomes.ebi.ac.uk/vol1/ftp/data_collections/1000_genomes_project/release/20190312_biallelic_SNV_and_INDEL/">here</a>.&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Jul 2014View details →
dryad36/100

Data for: DRD4 allele frequencies in greylag geese vary between urban and rural sites

<p>With the increasing urbanisation of the last decades, more and more bird species occur in urban habitats. Birds which thrive in urban habitats often have a higher tolerance towards human disturbance and show behaviours which differ from their rural counterparts. There is increasing evidence that many behaviours have a genetic basis. One candidate gene is the dopamine receptor D4 (DRD4), which has been associated with fear and thus, flight initiation distance (FID). In this study, we analysed a segment of DRD4 in greylag geese <em>Anser anser</em>, describing the variability of this gene across several geographically distant populations, and comparing its variability between an urban and a rural site in south-west Germany. We additionally measured FIDs of urban and rural geese to test for a possible correlation with DRD4 genotypes. We found a high variation of DRD4, with 10 variable sites leading to 11 alleles and 35 genotypes. Two genotypes occurred in 60% of all geese and were thus defined as common genotypes versus 33 rare genotypes. Population differentiation was very low between the urban and rural sites in Germany but common genotypes occurred more often in the urban area and rare genotypes more often in the rural area. FID was significantly higher at the rural site, but no significant correlation between FID and DRD4 genotypes could be detected. Nevertheless, our results suggest that local site selection may be related to DRD4 genotypes.</p>

opencc-zeroDec 2022View details →
dryad36/100

Data from: Combining allele frequency and tree-based approaches improves phylogeographic inference from natural history collections

Open the record for dataset details and reuse information.

publicDec 2017View details →
dryad36/100

Data for: DRD4 allele frequencies in greylag geese vary between urban and rural sites

Open the record for dataset details and reuse information.

publicDec 2022View details →
dryad32/100

Genomic regions influencing aggressive behavior in honey bees are defined by colony allele frequencies

For social animals, the genotypes of group members affect the social environment, and thus individual behavior, often indirectly. We used genome-wide association studies (GWAS) to determine the influence of individual vs. group genotypes on aggression in honey bees. Aggression in honey bees arises from the coordinated actions of colony members, primarily nonreproductive "soldier" bees, and thus, experiences evolutionary selection at the colony level. Here, we show that individual behavior is influenced by colony environment, which in turn, is shaped by allele frequency within colonies. Using a population with a range of aggression, we sequenced individual whole genomes and looked for genotype–behavior associations within colonies in a common environment. There were no significant correlations between individual aggression and specific alleles. By contrast, we found strong correlations between colony aggression and the frequencies of specific alleles within colonies, despite a small number of colonies. Associations at the colony level were highly significant and were very similar among both soldiers and foragers, but they covaried with one another. One strongly significant association peak, containing an ortholog of the Drosophila sensory gene dpr4 on linkage group (chromosome) 7, showed strong signals of both selection and admixture during the evolution of gentleness in a honey bee population. We thus found links between colony genetics and group behavior and also, molecular evidence for group-level selection, acting at the colony level. We conclude that group genetics dominates individual genetics in determining the fatal decision of honey bees to sting.

opencc-zeroAug 2020View details →
zenodo32/100

Data from: Signatures of introgression across the allele frequency spectrum

<p>This repository is associated with the article &quot;Signatures of introgression across the allele frequency spectrum&quot; by Simon H. Martin and William Amos, in Molecular Biology and Evolution (<a href="https://doi.org/10.1093/molbev/msaa239">https://doi.org/10.1093/molbev/msaa239</a>)</p> <p>Empirical genotype data from six different taxa are included. All are based on previously published data, but we provide the processed genotype files and frequency spectra used for our analyses for convenience. The repository also contains the plotted values underlying all figures (both empirical and simulated results).</p>

opencc-by-4.0Sep 2020View details →
dryad32/100

Geographic allele frequency variation in the 1000 Genomes hg38 NYGC dataset

<p>A key challenge in human genetics is to describe and understand the distribution of human genetic variation. Often genetic variation is described by showing rela tionships among populations or individuals, in each case drawing inferences over a large number of variants. Here, we present an alternative representation of human genetic variation that reveals the relative abundance of different allele frequency patterns across populations. This approach allows viewers to easily see several features of human genetic structure: (1) most variants are rare and geographically localized, (2) variants that are common in a single geographic region are more likely to be shared across the globe than to be private to that region, and (3) where two individuals differ, it is most often due to variants that are common globally, regardless of whether the individuals are from the same region or different regions. To guide interpretation of the results, we also apply the visualization to contrasting theoretical scenarios with varying levels of divergence and gene flow. Our variant-centric visualization clarifies the major geographic patterns of human variation and can be used to help correct potential misconceptions about the extent and nature of genetic differentiation among populations.</p>

opencc-zeroDec 2020View details →
dryad32/100

Data from: Accuracy of allele frequency estimation using pooled RNA-Seq

For non-model organisms, genome-wide information that describes functionally relevant variation may be obtained by RNA-Seq following de novo transcriptome assembly. While sequencing has become relatively inexpensive, the preparation of a large number of sequencing libraries remains prohibitively expensive for population genetic analyses of non-model species. Pooling samples may be then an attractive alternative. To test whether pooled RNA-Seq accurately predicts true allele frequencies, we analyzed the liver transcriptomes of 10 bank voles. Each sample was sequenced both as an individually barcoded library and as a part of a pool. Equal amounts of total RNA from each vole were pooled prior to mRNA selection and library construction. Reads were mapped onto the de novo assembled reference transcriptome. High-quality genotypes for individual voles, determined for 23,682 SNPs, provided information on "true" allele frequencies; allele frequencies estimated from the pool were then compared to these values. "True" frequencies and those estimated from the pool were highly correlated. Mean relative estimation error was 21% and did not depend on expression level. However, we also observed a minor effects of inter-individual variation in gene expression and allele specific gene expression influencing allele frequency estimation accuracy. Moreover we observed strong negative relationship between minor allele frequency and relative estimation error. Our results indicate that pooled RNA-Seq exhibits accuracy comparable to pooled genome resequencing, but variation in expression level between individuals should be assessed and accounted for. This should help in taking account the difference in accuracy between conservatively expressed transcripts and these which are variable in expression level.

opencc-zeroDec 2012View details →
dryad32/100

Data from: Contemporary evolution of sea urchin gamete-recognition proteins: experimental evidence of density-dependent gamete performance predicts shifts in allele frequencies over time

Species whose reproductive strategies evolved at one density regime might be poorly adapted to other regimes. Field and laboratory experiments on the sea urchin Strongylocentrotus franciscanus examined the influences of the two most common sperm bindin alleles, which differ at two amino acid sites, on fertilization success. In the field experiment, the Arginine/Glycine (RG) genotype performed best at low densities and the Glycine/Arginine (GR) genotype at high densities. In the lab experiment, the RG genotype had a higher affinity with available eggs, whereas the GR genotype was less likely to induce polyspermy. These sea urchins can reach 200 years of age. The RG allele dominates in old sea urchins, whereas younger sea urchins have near equal RG and GR allele frequencies. A latitudinal cline in RG and GR genotypes is consistent with longer survival of sea urchins in the north and with predominance of RG genotypes in older individuals. The oldest sea urchins were likely conceived at low densities, before sea-urchin predators, like sea otters, were overharvested and sea urchin densities exploded off the west coast. Contemporary evolution of gamete-recognition proteins might allow species to adapt to shifts in abundances and reduces the risk of reproductive failure in altered populations.

opencc-zeroDec 2011View details →
dryad32/100

Data from: Microevolution of S-allele frequencies in wild cherry populations: respective impacts of negative frequency dependent selection and genetic drift

Negative frequency dependent selection (NFDS) is supposed to be the main force controlling allele evolution at the gametophytic self-incompatibility locus (S-locus) in strictly outcrossing species. Genetic drift also influences S-allele evolution. In perennial sessile organisms, evolution of allelic frequencies over two generations is mainly shaped by individual fecundities and spatial processes. Using wild cherry populations between two successive generations, we tested whether S-alleles evolved following NFDS qualitative and quantitative predictions. We showed that allelic variation was negatively correlated with parental allelic frequency as expected under NFDS. However, NFDS predictions in finite population failed to predict more than half all S-allele quantitative evolution. We developed a spatially-explicit mating model which included the S-locus. We studied the effects of self-incompatibility and local drift within populations due to pollen dispersal in spatially distributed individuals, and variation in male fecundity on male mating success and allelic frequency evolution. Male mating success was negatively related to male allelic frequency as expected under NFDS. Spatial genetic structure combined with self-incompatibility resulted in higher effective pollen dispersal. Limited pollen dispersal in structured distributions of individuals and genotypes, non-random distribution of individuals and unequal pollen production significantly contributed to S-allele frequency evolution by creating local drift effects strong enough to counteract the NFDS effect on some alleles.

opencc-zeroDec 2010View details →
dryad32/100

Data from: Positive selection of deleterious alleles through interaction with a sex-ratio suppressor gene in African buffalo: a plausible new mechanism for a high frequency anomaly

Although generally rare, deleterious alleles can become common through genetic drift, hitchhiking or reductions in selective constraints. Here we present a possible new mechanism that explains the attainment of high frequencies of deleterious alleles in the African buffalo (Syncerus caffer) population of Kruger National Park, through positive selection of these alleles that is ultimately driven by a sex-ratio suppressor. We have previously shown that one in four Kruger buffalo has a Y-chromosome profile that, despite being associated with low body condition, appears to impart a relative reproductive advantage, and which is stably maintained through a sex-ratio suppressor. Apparently, this sex-ratio suppressor prevents fertility reduction that generally accompanies sex-ratio distortion. We hypothesize that this body-condition-associated reproductive advantage increases the fitness of alleles that negatively affect male body condition, causing genome-wide positive selection of these alleles. To investigate this we genotyped 459 buffalo using 17 autosomal microsatellites. By correlating heterozygosity with body condition (heterozygosity-fitness correlations), we found that most microsatellites were associated with one of two gene types: one with elevated frequencies of deleterious alleles that have a negative effect on body condition, irrespective of sex; the other with elevated frequencies of sexually antagonistic alleles that are negative for male body condition but positive for female body condition. Positive selection and a direct association with a Y-chromosomal sex-ratio suppressor are indicated, respectively, by allele clines and by relatively high numbers of homozygous deleterious alleles among sex-ratio suppressor carriers. This study, which employs novel statistical techniques to analyse heterozygosity-fitness correlations, is the first to demonstrate the abundance of sexually-antagonistic genes in a natural mammal population. It also has important implications for our understanding not only of the evolutionary and ecological dynamics of sex-ratio distorters and suppressors, but also of the functioning of deleterious and sexually-antagonistic alleles, and their impact on population viability.

opencc-zeroDec 2013View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record