Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

63

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

63 results for “allele frequency”

Learn how ShareScore rates datasets ↗
dryad28/100

Data from: Genotype-free estimation of allele frequencies reduces bias and improves demographic inference from RADSeq data

Restriction-site associated sequencing (RADSeq) facilitates rapid generation of thousands of genetic markers at relatively low cost; however, several sources of error specific to RADSeq methods often lead to biased estimates of allele frequencies and thereby to erroneous population genetic inference. Estimating the distribution of sample allele frequencies without calling genotypes was shown to improve population inference from whole genome sequencing data, but the ability of this approach to account for RADSeq-specific biases remains unexplored. Here we assess in how far genotype-free methods of allele frequency estimation affect demographic inference from empirical RADSeq data. Using the well-studied pied flycatcher (Ficedula hypoleuca) as a study system, we compare allele frequency estimation and demographic inference from whole genome sequencing data with that from RADSeq data matched for samples using both genotype-based and genotype free methods. The demographic history of pied flycatchers as inferred from RADSeq data was highly congruent with that inferred from WGS data when allele frequencies were estimated directly from the read data. In contrast, when allele frequencies were derived from called genotypes, RADSeq-based estimates of most model parameters fell outside the 95% confidence interval (CI) of estimates derived from WGS data. Notably, more stringent filtering of genotypes tended to increase the discrepancy between parameter estimates from WGS and RADSeq data, respectively. The results from this study demonstrate the ability of genotype-free methods to improve AFS-based demographic inference from RADSeq data and highlight the need to account for uncertainty in NGS data regardless of sequencing method.

opencc-zeroDec 2018View details →
dryad28/100

Data from: Minor allele frequency thresholds strongly affect population structure inference with genomic datasets

One common method of minimizing errors in large DNA sequence datasets is to drop variable sites with a minor allele frequency below some specified threshold. Though widespread, this procedure has the potential to alter downstream population genetic inferences and has received relatively little rigorous analysis. Here we use simulations and an empirical SNP dataset to demonstrate the impacts of minor allele frequency (MAF) thresholds on inference of population structure. We find that model-based inference of population structure is confounded when singletons are included in the alignment, and that both model-based and multivariate analyses infer less distinct clusters when more stringent MAF cutoffs are applied. We propose that this behavior is caused by the combination of a drop in the total size of the data matrix and by correlations between allele frequencies and mutational age. We recommend a set of best practices for applying MAF filters in studies seeking to describe population structure with genomic data.

opencc-zeroDec 2018View details →
dryad28/100

Data from: Negative frequency-dependent selection of sexually antagonistic alleles in Myodes glareolus

Sexually antagonistic genetic variation, where optimal values of traits are sex-dependent, is known to slow the loss of genetic variance associated with directional selection on fitness-related traits. However, sexual antagonism alone is not sufficient to maintain variation indefinitely. Selection of rare forms within the sexes can help to conserve genotypic diversity. We combined theoretical models and a field experiment with Myodes glareolus to show that negative frequency-dependent selection on male dominance maintains variation in sexually antagonistic alleles. In our experiment, high-dominance male bank voles were found to have low-fecundity sisters, and vice versa. These results show that investigations of sexually antagonistic traits should take into account the effects of social interactions on the interplay between ecology and evolution, and that investigations of genetic variation should not be conducted solely under laboratory conditions.

opencc-zeroDec 2010View details →
dryad28/100

Data from: The impact of library preparation protocols on the accuracy of allele frequency estimates in Pool-Seq data

Sequencing pools of individuals (Pool-Seq) is a cost-effective method to determine genome-wide allele frequency estimates. Given the importance of meta-analyses combining data sets, we determined the influence of different genomic library preparation protocols on the consistency of allele frequency estimates. We found that typically no more than 1% of the variation in allele frequency estimates could be attributed to differences in library preparation. Also read length had only a minor effect on the consistency of allele frequency estimates. By far, the most pronounced influence could be attributed to sequence coverage. Increasing the coverage from 30- to 50-fold improved the consistency of allele frequency estimates by at least 27%. We conclude that Pool-Seq data can be easily combined across different library preparation methods, but sufficient sequence coverage is key to reliable results.

opencc-zeroDec 2014View details →
dryad28/100

A Streamlined and High-Throughput Error-Corrected Next-Generation Sequencing Method for Low Variant Allele Frequency Quantitation

<p></p><p>Quantifying mutant or variable allele frequencies (VAFs) of ≤10−3 using next-generation sequencing (NGS) has utility in both clinical and nonclinical settings. Two common approaches for quantifying VAFs using NGS are tagged single-strand sequencing and duplex sequencing. While duplex sequencing is reported to have sensitivity up to 10−8 VAF, it is not a quick, easy, or inexpensive method. We report a method for quantifying VAFs that are ≥10−4 that is as easy and quick for processing samples as standard sequencing kits, yet less expensive than the kits. The method was developed using PCR fragment-based VAFs of Kras codon 12 in log10 increments from 10−5 to 10−1, then applied and tested on native genomic DNA. For both sources of DNA, there is a proportional increase in the observed VAF to input VAF from 10−4 to 100% mutant samples. Variability of quantitation was evaluated within experimental replicates and shown to be consistent across sample preparations. The error at each successive base read was evaluated to determine if there is a limit of read length for quantitation of ≥10−4, and it was determined that read lengths up to 70 bases are reliable for quantitation. The method described here is adaptable to various oncogene or tumor suppressor gene targets, with the potential to implement multiplexing at the initial tagging step. While easy to perform manually, it is also suited for robotic handling and batch processing of samples, facilitating detection and quantitation of genetic carcinogenic biomarkers before tumor formation or in normal-appearing tissue.</p><p></p>

opencc-zeroAug 2019View details →
dryad28/100

Chinook salmon environmental data and allele frequency matrix

<p><span><span><span><span><span><span><span><span><span><span><span>Many species that undergo long breeding migrations, such as anadromous fishes, face highly heterogeneous environments along their migration corridors and at their spawning sites. These environmental challenges encountered at different life stages may act as strong selective pressures and drive local adaptation. However, the relative influence of environmental conditions along the migration corridor compared to the conditions at spawning sites on driving selection is still unknown. In this study, we performed genome-environment associations (GEA) to understand the relationship between landscape and environmental conditions driving selection in seven populations of the anadromous Chinook salmon (<i>Oncorhynchus tshawytscha)–</i>a species of important economic, social, cultural and ecological value–in the Columbia River basin. We extracted environmental variables for the shared migration corridors and at distinct spawning sites for each population, and used a Pool-seq approach to perform whole genome resequencing. Bayesian and univariate genome-environment association tests with migration-specific and spawning site-specific environmental variables indicated many more candidate SNPs associated with environmental conditions of the migration corridor compared to spawning sites. Specifically, variables associated with temperature, precipitation, terrain roughness, and elevation variables of the migration corridor were the most significant drivers of environmental selection. Additional analyses of neutral loci revealed two distinct clusters representing populations from different geographic regions of the drainage that also exhibit differences in adult migration timing (summer vs. fall). Tests for genomic regions under selection revealed a strong peak on chromosome 28, corresponding to the GREB1L/ROCK1 region that has been identified previously in salmonids as a region associated with adult migration timing. Our results show that environmental variation experienced throughout migration corridors imposed a greater selective pressure on Chinook salmon than environmental conditions at spawning sites.</span></span></span></span></span></span></span></span></span></span></span></p>

opencc-zeroSep 2021View details →
dryad28/100

Data from: Genotype-free estimation of allele frequencies reduces bias and improves demographic inference from RADSeq data

Open the record for dataset details and reuse information.

publicJan 2019View details →
dryad28/100

Data from: Negative frequency-dependent selection of sexually antagonistic alleles in Myodes glareolus

Open the record for dataset details and reuse information.

publicNov 2011View details →
dryad28/100

Data from: Comparing van Oosterhout and Chybicki-Burczyk methods of estimating null allele frequencies for inbred populations

Open the record for dataset details and reuse information.

publicAug 2012View details →
dryad28/100

Data from: PoMo: an allele frequency-based approach for species tree estimation

Open the record for dataset details and reuse information.

publicJun 2015View details →
dryad28/100

Data from: Minor allele frequency thresholds strongly affect population structure inference with genomic datasets

Open the record for dataset details and reuse information.

publicFeb 2019View details →
dryad28/100

A Streamlined and High-Throughput Error-Corrected Next-Generation Sequencing Method for Low Variant Allele Frequency Quantitation

Open the record for dataset details and reuse information.

publicDec 2019View details →
dryad28/100

Data from: The impact of library preparation protocols on the accuracy of allele frequency estimates in Pool-Seq data

Open the record for dataset details and reuse information.

publicMay 2015View details →
dryad28/100

Reference allele frequencies for populations pools of Atlantic Herring (Clupea harengus)

Open the record for dataset details and reuse information.

publicDec 2020View details →
dryad28/100

Chinook salmon environmental data and allele frequency matrix

Open the record for dataset details and reuse information.

publicSep 2021View details →
dryad28/100

Data from: Accounting for genotype uncertainty in the estimation of allele frequencies in autopolyploids

Open the record for dataset details and reuse information.

publicNov 2015View details →
dryad28/100

The theory and applications of measuring broad-range and chromosome-wide recombination rate from allele frequency decay around a selected locus

Open the record for dataset details and reuse information.

publicSep 2020View details →
geo24/100

Definitive classification of ovarian mature cystic teratomas origin using B allele frequency plots with SNP array analysis

GEO Series GSE85424. Homo sapiens. 38 samples. Type: Genome variation profiling by SNP array.

openGEO-OpenJul 2018View details →
dryad24/100

Data from: Identifying consistent allele frequency differences in studies of stratified populations

1. With increasing application of pooled-sequencing approaches to population genomics robust methods are needed to accurately quantify allele frequency differences between populations. Identifying consistent differences across stratified populations can allow us to detect genomic regions under selection and that differ between populations with different histories or attributes. Current popular statistical tests are easily implemented in widely available software tools which make them simple for researchers to apply. However, there are potential problems with the way such tests are used ,which means that underlying assumptions about the data are frequently violated. 2. These problems are highlighted by simulation of simple but realistic population genetic models of neutral evolution and the performance of different tests are assessed. We present alternative tests (including GLMs with quasibinomial error structure) with attractive properties for the analysis of allele frequency differences and re-analyse a published dataset. 3. The simulations show that common statistical tests for consistent allele frequency differences perform poorly, with high false positive rates. Applying tests that do not confound heterogeneity and main effects significantly improves inference. Variation in sequencing coverage likely produces many false positives and re-scaling allele frequencies to counts out of a common value or an effective sample size reduces this effect. 4. Many researchers are interested in identifying allele frequencies that vary consistently across replicates to identify loci underlying phenotypic responses to selection or natural variation in phenotypes. Popular methods that have been suggested for this task perform poorly in simulations. Overall, quasibinomial GLMs perform better and also have the attractive feature of allowing correction for multiple testing by standard procedures and are easily extended to other designs.

opencc-zeroDec 2016View details →
ClinicalTrials.gov24/100

Human Leukocyte Antigen Class II (DRB1 and DQB1) Alleles and Haplotypes Frequencies in Patients With Pemphigus Vulgaris Among the Russian Population

ClinicalTrials.gov study NCT05284929. IPD Sharing: Not stated. Countries: 1. Publications: 0.

restrictedIPD-UNDECIDEDFeb 2026View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record