Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

18

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

18 results for “genotyping error”

Learn how ShareScore rates datasets ↗
zenodo36/100

Documentation and code for reproducing analyses presented in: vcferr: Development, Validation, and Application of a SNP Genotyping Error Simulation Framework

<p>Documentation and code for reproducing analyses presented in: vcferr: Development, Validation, and Application of a SNP Genotyping Error Simulation Framework. Please see the <strong>README.pdf</strong> for step-by-step instructions for reproducing the entire analysis described in the paper.</p>

opencc-by-4.0Jun 2022View details →
zenodo36/100

Distinguishing mutations and null alleles from genotyping errors using mother progeny comparisons in Brazilian pine (Araucaria angustifolia)

The use of microsatellite markers provides a window into the evolutionary processes of a given species. As such, these markers are widely used in scientific and applied research and are praised for their practicality and ease of use, however, the unavoidable incidence of genotyping deviations has been broadly neglected in the literature. Therefore, the present study aimed to estimate the rate of null alleles, mutations and genotyping errors in microsatellite loci, using Araucaria angustifolia, a threatened species, as a case study. We estimated the rates of the different types of genotyping deviations using mother-progeny genotype comparison from 50 seed-trees and their respective progeny (seeds). A total of 2336 A. angustifolia samples were genotyped, and we found that the rate of null alleles was 0.045. From the 1972 mother-progeny comparisons, the overall genotype deviation rate was 1.58%, consisting of 145 inconsistences (mutations), 339 null alleles and 210 genotyping errors. In terms of seed numbers, 128 (6.5%) showed inconsistencies in at least one locus, 118 (6.0%) null alleles, and 321 (16.3%) genotyping errors. This is the first study to describe the inconsistences (mutations) between mother-progeny genotypes for A. angustifolia, and the outcome makes it clear that an understanding of these genotyping deviations must be considered in assessing the accuracy of inferences made based on population genetics analyses.

opencc-zeroSep 2019View details →
zenodo36/100

Distinguishing mutations and null alleles from genotyping errors using mother progeny comparisons in Brazilian pine (Araucaria angustifolia)

the rate of null alleles, mutations and genotyping errors in microsatellite loci, using Araucaria angustifolia, a threatened species, as a case study. We estimated the rates of the different types of genotyping deviations using mother-progeny genotype comparison from 50 seed-trees and their respective progeny (seeds). A total of 2336 A. angustifolia samples were genotyped, and we found that the rate of null alleles was 0.045. From the 1972 mother-progeny comparisons, the overall genotype deviation rate was 1.58%, consisting of 145 inconsistences (mutations), 339 null alleles and 210 genotyping errors. In terms of seed numbers, 128 (6.5%) showed inconsistencies in at least one locus, 118 (6.0%) null alleles, and 321 (16.3%) genotyping errors. This is the first study to describe the inconsistences (mutations) between mother-progeny genotypes for A. angustifolia, and the outcome makes it clear that an understanding of these genotyping deviations must be considered in assessing the accuracy of inferences made based on population genetics analyses.

opencc-zeroOct 2019View details →
dryad36/100

Data from: RAD sequencing, genotyping error estimation and de novo assembly optimization for population genetic inference

Restriction site-associated DNA sequencing (RADseq) provides researchers with the ability to record genetic polymorphism across thousands of loci for non-model organisms, potentially revolutionising the field of molecular ecology. However, as with other genotyping methods, RADseq is prone to a number of sources of error that may have consequential effects for population genetic inferences, and these have received only limited attention in terms of the estimation and reporting of genotyping error rates. Here we use individual sample replicates, under the expectation of identical genotypes, to quantify genotyping error in the absence of a reference genome. We then use sample replicates to (1) optimize de novo assembly parameters within the program Stacks, by minimizing error and maximizing the retrieval of informative loci, and; (2) quantify error rates for loci, alleles and SNPs. As an empirical example we use a double digest RAD dataset of a non-model plant species, Berberis alpina, collected from high altitude mountains in Mexico.

opencc-zeroDec 2013View details →
dryad36/100

Data from: RAD sequencing, genotyping error estimation and de novo assembly optimization for population genetic inference

Open the record for dataset details and reuse information.

publicJun 2014View details →
dryad32/100

Data from: Finding the right coverage: The impact of coverage and sequence quality on SNP genotyping error rates

Restriction-enzyme-based sequencing methods enable the genotyping of thousands of single nucleotide polymorphism (SNP) loci in non-model organisms. However, in contrast to traditional genetic markers, genotyping error rates in SNPs derived from restriction-enzyme-based methods remain largely unknown. Here, we estimated genotyping error rates in SNPs genotyped with double digest RAD sequencing from Mendelian incompatibilities in known mother-offspring dyads of Hoffman's two-toed sloth (Choloepus hoffmanni) across a range of coverage and sequence quality criteria, for both reference-aligned and de novo-assembled datasets. Genotyping error rates were more sensitive to coverage than sequence quality and low coverage yielded high error rates, particularly in de novo-assembled datasets. For example, coverage ≥5 yielded median genotyping error rates of ≥0.03 and ≥0.11 in reference-aligned- and de novo-assembled datasets, respectively. Genotyping error rates declined to ≤0.01 in reference-aligned datasets with a coverage &gt;30, but remained &gt;0.04 in the de novo-assembled datasets. We observed approximately 10- and 13-fold declines in the number of loci sampled in the reference-aligned and de novo-assembled datasets when coverage was increased from &gt;5 to &gt;30 at quality score ≥30, respectively. Finally, we assessed the effects of genotyping coverage on a common population genetic application, parentage assignments, and showed that the proportion of incorrectly assigned maternities was relatively high at low coverage. Overall, our results suggest that the tradeoff between sample size and genotyping error rates be considered prior to building sequencing libraries, reporting genotyping error rates become standard practice, and that effects of genotyping errors on inference be evaluated in restriction-enzyme-based SNP studies.

opencc-zeroDec 2015View details →
dryad32/100

Data from: Estimation of genotyping error rate from repeat genotyping, unintentional recaptures and known parent-offspring comparisons in 16 microsatellite loci for brown rockfish (Sebastes auriculatus)

Genotyping errors are present in almost all genetic data and can affect biological conclusions of a study, particularly for studies based on individual identification and parentage. Many statistical approaches can incorporate genotyping errors, but usually need accurate estimates of error rates. Here, we used a new microsatellite data set developed for brown rockfish (Sebastes auriculatus) to estimate genotyping error using three approaches: (i) repeat genotyping 5% of samples, (ii) comparing unintentionally recaptured individuals and (iii) Mendelian inheritance error checking for known parent–offspring pairs. In each data set, we quantified genotyping error rate per allele due to allele drop-out and false alleles. Genotyping error rate per locus revealed an average overall genotyping error rate by direct count of 0.3%, 1.5% and 1.7% (0.002, 0.007 and 0.008 per allele error rate) from replicate genotypes, known parent–offspring pairs and unintentionally recaptured individuals, respectively. By direct-count error estimates, the recapture and known parent–offspring data sets revealed an error rate four times greater than estimated using repeat genotypes. There was no evidence of correlation between error rates and locus variability for all three data sets, and errors appeared to occur randomly over loci in the repeat genotypes, but not in recaptures and parent–offspring comparisons. Furthermore, there was no correlation in locus-specific error rates between any two of the three data sets. Our data suggest that repeat genotyping may underestimate true error rates and may not estimate locus-specific error rates accurately. We therefore suggest using methods for error estimation that correspond to the overall aim of the study (e.g. known parent–offspring comparisons in parentage studies).

opencc-zeroDec 2011View details →
dryad32/100

Data from: Detecting macroevolutionary genotype-phenotype associations using error-corrected rates of protein convergence

<p><span>On macroevolutionary timescales, extensive mutations and phylogenetic uncertainty mask the signals of genotype-phenotype associations underlying convergent evolution. To overcome this problem, we extended the widely used framework of nonsynonymous-to-synonymous substitution rate ratios and developed the novel metric <em>ω<sub>C</sub></em>, which measures the error-corrected convergence rate of protein evolution. While </span><span><em>ω<sub>C</sub></em></span><span> distinguishes natural selection from genetic noise and phylogenetic errors in simulation and real examples, its accuracy allows an exploratory genome-wide search of adaptive molecular convergence without phenotypic hypothesis or candidate genes. Using gene expression data, we explored over 20 million branch combinations in vertebrate genes and identified the joint convergence of expression patterns and protein sequences with amino acid substitutions in functionally important sites, providing hypotheses on undiscovered phenotypes. We further extended our method with a heuristic algorithm to detect highly repetitive convergence among computationally nontrivial higher-order phylogenetic combinations. Our approach allows bidirectional searches for genotype-phenotype associations, even in lineages that diverged for hundreds of millions of years.</span></p>

opencc-zeroOct 2022View details →
dryad32/100

Data from: Estimation of genotyping error rate from repeat genotyping, unintentional recaptures and known parent-offspring comparisons in 16 microsatellite loci for brown rockfish (Sebastes auriculatus)

Open the record for dataset details and reuse information.

publicAug 2012View details →
dryad32/100

Data from: Finding the right coverage: The impact of coverage and sequence quality on SNP genotyping error rates

Open the record for dataset details and reuse information.

publicFeb 2016View details →
dryad32/100

Data from: Dealing with AFLP genotyping errors to reveal genetic structure in Plukenetia volubilis (Euphorbiaceae) in the Peruvian Amazon

Open the record for dataset details and reuse information.

publicJan 2018View details →
dryad32/100

Data from: Detecting macroevolutionary genotype-phenotype associations using error-corrected rates of protein convergence

Open the record for dataset details and reuse information.

publicOct 2022View details →
dryad28/100

Data from: Identifying and reducing AFLP genotyping error: an example of tradeoffs when comparing population structure in broadcast spawning versus brooding oysters

Phylogeographic inferences about gene flow are strengthened through comparison of co-distributed taxa, but also depend on adequate genomic sampling. Amplified Fragment Length Polymorphisms (AFLP) provide a rapid and inexpensive source of multilocus allele frequency data for making genomically robust inferences. Every AFLP study initially generates markers with a range of locus-specific genotyping error rates and applies criteria to select a subset for analysis. However, there has been very little empirical evaluation of the best tradeoff between culling all but the lowest-error loci to minimize overall genotyping error versus the potential for increasing population genetic signal by retaining more loci. Here, we used AFLPs to compare population structure in co-distributed broadcast spawning (Crassostrea virginica) and brooding (Ostrea equestris) oyster species. Using existing methods for almost entirely automated marker selection and scoring, genotyping error tradeoffs were evaluated by comparing results across a nested series of datasets with mean mismatch errors of 0, 1, 2, 3, 4 and &gt;4%. Artifactual population structure was diagnosed in high-error datasets and we assessed the low-error point at which expected population substructure signal was lost. In both species we identified substructure patterns deemed to be inaccurate at error rates {less than or equal to}2% and &gt;4%. In the species comparison, the optimum datasets showed higher gene flow for the brooding oyster with more oceanic salinity tolerances. AFLP tradeoffs may differ among studies, but our results suggest that important signal may be lost in the pursuit of 'acceptable' error levels and our procedures provide a general method for empirically exploring these tradeoffs.

opencc-zeroDec 2010View details →
dryad28/100

Data from: Identifying and reducing AFLP genotyping error: an example of tradeoffs when comparing population structure in broadcast spawning versus brooding oysters

Open the record for dataset details and reuse information.

publicDec 2011View details →
geo24/100

Low-pass sequencing increases the power of GWAS and decreases measurement error of polygenic risk scores compared to genotyping arrays

GEO Series GSE165845. Homo sapiens. 360 samples. Type: Genome variation profiling by array.

openGEO-OpenJan 2021View details →
dryad24/100

Data from: Estimating genotyping errors from genotype and reconstructed pedigree data

1. Genotyping errors are rules rather than exceptions in reality, and are found in virtually all but very small datasets. These errors, even when occurring at an extremely low rate, can derail many genetic analyses such as parentage/sibship assignments and linkage/association studies. 2. Nonetheless, few robust and accurate methods are available for estimating the rate of occurrence of genotyping errors and for identifying individual erroneous genotypes at a locus. Methods based on duplicate genotyping are expensive, and estimate genotype inconsistency rather than error rate at a locus. Methods based on Hardy-Weinberg equilibrium tests have low robustness and low power, and apply only to those particular errors that cause excessive homozygosity. Methods based on pedigrees are powerful, robust and accurate. However, they rely on known and complete pedigrees that are unfortunately rarely available from natural populations in the wild. 3. I proposed a maximum likelihood method to reconstruct pedigrees from genotype data with errors occurring at a roughly estimated (presumed) rate. In this paper, I describe how to use the method and inferred pedigree in estimating allelic dropout (or null allele) rate and false allele rate jointly at each marker locus, in identifying the erroneous genotypes, and in inferring the most likely genotypes at each locus of each individual. I examine the power, accuracy and robustness of the method by extensive simulations, and demonstrate the usefulness of the method by analysing three empirical datasets. 4. It is concluded that, both pedigrees and the rates of genotyping errors at each locus can be reliably estimated from the same genotype data by the same likelihood method, when marker information is sufficient and some sampled individuals are first-degree relatives. The erroneous genotypes are however inferred conservatively, and are reliably detected only when they occur in large families and/or at highly polymorphic loci. Estimation of genotyping error rates per locus and identification of erroneous genotypes of each individual at each locus should be routinely conducted to assess and improve data quality, to highlight markers for optimization of genotyping protocols or for replacement, and to enable the integration of genotyping errors in a robust statistical analysis.

opencc-zeroDec 2016View details →
dryad24/100

Data from: Estimating genotyping errors from genotype and reconstructed pedigree data

Open the record for dataset details and reuse information.

publicJul 2018View details →
geo16/100

A robust machine learning approach for missing persons cases with high genotyping errors

GEO Series GSE209804. Homo sapiens. 24 samples. Type: Genome variation profiling by SNP array; SNP genotyping by SNP array.

openGEO-OpenJul 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record