Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
64
datasets available to search
ShareScore release 0.9.0
Dataset results
64 results for “population size estimation”
Data for: Improving population size estimation at Western Capercaillie leks: lek counts vs. genetic methods
Open the record for dataset details and reuse information.
Data from: A model-derived short-term estimation method of effective size for small populations with overlapping generations
Open the record for dataset details and reuse information.
Data from: Accounting for heterogeneity when estimating stopover duration, timing and population size of red knots along the Luannan Coast of Bohai Bay, China
Open the record for dataset details and reuse information.
Data from: Hierarchical distance sampling to estimate population sizes of common lizards across a desert ecoregion
Open the record for dataset details and reuse information.
Data from: Heritability estimates from genome wide relatedness matrices in wild populations: application to a passerine, using a small sample size
Open the record for dataset details and reuse information.
Data from: Effects of population characteristics and structure on estimates of effective population size in a house sparrow metapopulation
Open the record for dataset details and reuse information.
Data from: Estimating effective population size of guanacos in Patagonia: an integrative approach for wildlife conservation
Open the record for dataset details and reuse information.
Genomic prediction with non-additive effects in beef cattle: Stability of variance component and genetic effect estimates against population size
Open the record for dataset details and reuse information.
Data from: An evaluation of the methods to estimate effective population size from measures of linkage disequilibrium
In 1971, John Sved derived an approximate relationship between linkage disequilibrium and effective population size for an ideal finite population. This seminal work was extended by Sved and Feldman (1973) and Weir and Hill (1980) who derived additional equations with the same purpose. These equations yield useful estimates of effective population size, as they require a single sample in time. As these estimates of effective population size are now commonly used on a variety of genomic data, from arrays of single nucleotide polymorphisms to whole genome data, some authors have investigated their bias through simulation studies and proposed corrections for different mating systems. However, the cause of the bias remains elusive. Here we show the problems of using linkage disequilibrium as a statistical measure and, analogously, the problems in estimating effective population size from such measure. For that purpose, we compare three commonly used approaches with a transition probability based method that we develop here. It provides an exact computation of linkage disequilibrium. We show here that the bias in the estimates of linkage disequilibrium and effective population size are partly due to low frequency markers, tightly linked markers or to a small total number of crossovers per generation. These biases, however, do not decrease when increasing sample size or using unlinked markers. Our results show the issues of such measures of effective population based on linkage disequilibrium, and suggest which of the method here studied should be used in empirical studies as well as the optimal distance between markers for such estimates.
Data from: Estimations of linkage disequilibrium, effective population size and ROH-based inbreeding coefficients in Spanish Churra sheep using imputed high-density SNP genotypes
In this study, the availability of the Ovine HD SNP BeadChip (HD-chip) and the development of an imputation strategy provided an opportunity to further investigate the extent of linkage disequilibrium (LD) at short distances in the genome of the Spanish Churra dairy sheep breed. A population of 1686 animals, including 16 rams and their half-sib daughters, previously genotyped for the 50K-chip, was imputed to the HD-chip density based on a reference population of 335 individuals. After assessing the imputation accuracy for beagle v4.0 (0.922) and fimpute v2.2 (0.921) using a cross-validation approach, the imputed HD-chip genotypes obtained with beagle were used to update the estimates of LD and effective population size for the studied population. The imputed genotypes were also used to assess the degree of homozygosity by calculating runs of homozygosity and to obtain genomic-based inbreeding coefficients. The updated LD estimations provided evidence that the extent of LD in Churra sheep is even shorter than that reported based on the 50K-chip and is one of the shortest extents compared with other sheep breeds. Through different comparisons we have also assessed the impact of imputation on LD and effective population size estimates. The inbreeding coefficient, considering the total length of the run of homozygosity, showed an average estimate (0.0404) lower than the critical level. Overall, the improved accuracy of the updated LD estimates suggests that the HD-chip, combined with an imputation strategy, offers a powerful tool that will increase the opportunities to identify genuine marker-phenotype associations and to successfully implement genomic selection in Churra sheep.
Data from: Estimating national population sizes: methodological challenges and applications illustrated in the common nightingale, a declining songbird in the UK
1. Estimation of national population size can be important for setting conservation priorities but its methodology has received little critical attention. Sites for highly aggregated species are often prioritised if they contain 1% of national or biogeographical populations but the utility of this approach for other species is unclear. 2. To make recommendations for study design, we present methods used to estimate the UK population size of the common nightingale Luscinia megarhynchos. We assess the sensitivity of the population estimate to the analytical method used and identify sites of national importance for this territorial songbird. 3. Survey effort was directed by prior knowledge of the species' distribution and the survey design maximised detectability by focussing on the period of greatest song output. We used three different statistical methods to account for detectability, estimating that 55–65% of the national population was detected during surveys. 4. Birds in areas not known to contain the species accounted for 13–23% of the population estimate. Methods to account for these individuals contributed the greatest uncertainty to the results, due to the difficulty of surveying a very large sample of random sites and consequent need to stratify the sample. 5. The 12 derived estimates ranged between 5094 and 5938 territorial males, with the confidence limits ranging from 4764 to 6534. Site delimitation, using clustering based on nearest-neighbour distances, identified one site clearly of national importance and several others potentially nationally important, depending on the population threshold and clustering distance used. 6. Synthesis and applications. National population estimation is difficult and requires that species-specific variability in detectability and individuals present outside surveyed areas are accurately accounted for through survey design and statistical analysis. Accounting for these sources of error will not always be possible and will hamper efforts to assess true population size and consequently to determine whether sites, however defined, exceed critical thresholds of importance. Resources may be better invested in other activities, for example in generating population trends based on relative indices. The latter are generally easier to produce, potentially more robust and arguably more suitable for many conservation applications.
Data from: A comparison of single-sample estimators of effective population sizes from genetic marker data
In molecular ecology and conservation genetics studies, the important parameter of effective population size (Ne) is increasingly estimated from a single sample of individuals taken at random from a population and genotyped at a number of marker loci. Several estimators are developed, based on the information of linkage disequilibrium (LD), heterozygote excess (HE), molecular coancestry (MC) and sibship frequency (SF) in marker data. The most popular is the LD estimator, because it is more accurate than HE and MC estimators and is simpler to calculate than SF estimator. However, little is known about the accuracy of LD estimator relative to that of SF and about the robustness of all single-sample estimators when some simplifying assumptions (e.g. random mating, no linkage, no genotyping errors) are violated. This study fills the gaps and uses extensive simulations to compare the biases and accuracies of the four estimators for different population properties (e.g. bottlenecks, nonrandom mating, haplodiploid), marker properties (e.g. linkage, polymorphisms) and sample properties (e.g. numbers of individuals and markers) and to compare the robustness of the four estimators when marker data are imperfect (with allelic dropouts). Extensive simulations show that SF estimator is more accurate, has a much wider application scope (e.g. suitable to nonrandom mating such as selfing, haplodiploid species, dominant markers) and is more robust (e.g. to the presence of linkage and genotyping errors of markers) than the other estimators. An empirical data set from a Yellowstone grizzly bear population was analysed to demonstrate the use of the SF estimator in practice.
Mercury and Arsenic muscle concentration data as used in "Mixed model approaches can leverage database information to improve the estimation of size-adjusted contaminant concentrations in fish populations"
<p>These mercury and arsenic concentration data, as recieved from Gretchen Lescord, and downloaded from the MOE fish contaminant database, were used to create the publication Mixed model approaches can leverage database information to improve the estimation of size-adjusted contaminant concentrations in fish populations. The markdown and code used for the analysis of this data can be found on Github at https://github.com/GLFC-WET/HGAS_master.</p>
Data from: Estimation of contemporary effective population size and population declines using RAD sequence data
Large genomic datasets generated with restriction-site associated DNA sequencing (RADseq), in combination with demographic inference methods, are improving our ability to gain insights into the population history of species. We used a simulation approach to examine the potential for RADseq datasets to accurately estimate effective population size (Ne) over the course of stable and declining population trends, and we compare the ability of two methods of analysis to accurately distinguish stable from steadily declining populations over a contemporary time scale (20 generations). Using a linkage disequilibrium-based analysis, individual sampling (i.e., n ≥ 30) had the greatest effect on Ne estimation and the detection of population-size declines, with declines reliably detected across scenarios approximately 10 generations after they began. Coalescent-based inference required fewer sampled individuals (i.e., n = 15), and instead was most influenced by the size of the SNP dataset, with 25,000 to 50,000 SNPs required for accurate detection of population trends and at least 20 generations after decline began. The number of samples available and targeted number of RADseq loci are important criteria when choosing between these methods. Neither method suffered any apparent bias due to the effects of allele dropout typical of RAD data. With an understanding of the limitations and biases of these approaches, researchers can make more informed decisions when designing their sampling and analyses. Overall, our results reveal that demographic inference using RADseq data can be successfully applied to infer recent population size change and may be important tools for population monitoring and conservation biology.
Data from: A critical assessment of estimating census population size from genetic population size (or vice versa) in three fishes
Open the record for dataset details and reuse information.
Data from: Estimation of contemporary effective population size and population declines using RAD sequence data
Open the record for dataset details and reuse information.
Data from: An evaluation of the methods to estimate effective population size from measures of linkage disequilibrium
Open the record for dataset details and reuse information.
Data from: Estimating national population sizes: methodological challenges and applications illustrated in the common nightingale, a declining songbird in the UK
Open the record for dataset details and reuse information.
Data from: A comparison of single-sample estimators of effective population sizes from genetic marker data
Open the record for dataset details and reuse information.
Data from: Estimations of linkage disequilibrium, effective population size and ROH-based inbreeding coefficients in Spanish Churra sheep using imputed high-density SNP genotypes
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.