Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
158
datasets available to search
ShareScore release 0.9.0
Dataset results
158 results for “effective population size”
Data from: Effects of population size and isolation on heterosis, mean fitness, and inbreeding depression in a perennial plant
Open the record for dataset details and reuse information.
Data from: Gene flow and effective population sizes of the butterfly Maculinea alcon in a highly fragmented, anthropogenic landscape
Open the record for dataset details and reuse information.
Genomic prediction with non-additive effects in beef cattle: Stability of variance component and genetic effect estimates against population size
Open the record for dataset details and reuse information.
Data from: Stock enhancement or sea ranching? Insights from monitoring the genetic diversity, relatedness and effective size in a seeded great scallop population (Pecten maximus)
Open the record for dataset details and reuse information.
Data and scripts from: Experimental evidence of size-selective harvest and environmental stochasticity effects on population demography, fluctuations, and nonlinearity
Open the record for dataset details and reuse information.
Data from: An evaluation of the methods to estimate effective population size from measures of linkage disequilibrium
In 1971, John Sved derived an approximate relationship between linkage disequilibrium and effective population size for an ideal finite population. This seminal work was extended by Sved and Feldman (1973) and Weir and Hill (1980) who derived additional equations with the same purpose. These equations yield useful estimates of effective population size, as they require a single sample in time. As these estimates of effective population size are now commonly used on a variety of genomic data, from arrays of single nucleotide polymorphisms to whole genome data, some authors have investigated their bias through simulation studies and proposed corrections for different mating systems. However, the cause of the bias remains elusive. Here we show the problems of using linkage disequilibrium as a statistical measure and, analogously, the problems in estimating effective population size from such measure. For that purpose, we compare three commonly used approaches with a transition probability based method that we develop here. It provides an exact computation of linkage disequilibrium. We show here that the bias in the estimates of linkage disequilibrium and effective population size are partly due to low frequency markers, tightly linked markers or to a small total number of crossovers per generation. These biases, however, do not decrease when increasing sample size or using unlinked markers. Our results show the issues of such measures of effective population based on linkage disequilibrium, and suggest which of the method here studied should be used in empirical studies as well as the optimal distance between markers for such estimates.
Data from: Estimations of linkage disequilibrium, effective population size and ROH-based inbreeding coefficients in Spanish Churra sheep using imputed high-density SNP genotypes
In this study, the availability of the Ovine HD SNP BeadChip (HD-chip) and the development of an imputation strategy provided an opportunity to further investigate the extent of linkage disequilibrium (LD) at short distances in the genome of the Spanish Churra dairy sheep breed. A population of 1686 animals, including 16 rams and their half-sib daughters, previously genotyped for the 50K-chip, was imputed to the HD-chip density based on a reference population of 335 individuals. After assessing the imputation accuracy for beagle v4.0 (0.922) and fimpute v2.2 (0.921) using a cross-validation approach, the imputed HD-chip genotypes obtained with beagle were used to update the estimates of LD and effective population size for the studied population. The imputed genotypes were also used to assess the degree of homozygosity by calculating runs of homozygosity and to obtain genomic-based inbreeding coefficients. The updated LD estimations provided evidence that the extent of LD in Churra sheep is even shorter than that reported based on the 50K-chip and is one of the shortest extents compared with other sheep breeds. Through different comparisons we have also assessed the impact of imputation on LD and effective population size estimates. The inbreeding coefficient, considering the total length of the run of homozygosity, showed an average estimate (0.0404) lower than the critical level. Overall, the improved accuracy of the updated LD estimates suggests that the HD-chip, combined with an imputation strategy, offers a powerful tool that will increase the opportunities to identify genuine marker-phenotype associations and to successfully implement genomic selection in Churra sheep.
Data from: Sexual selection has minimal impact on effective population sizes in species with high rates of random offspring mortality: an empirical demonstration using fitness distributions
The effective population size (Ne) is a fundamental parameter in population genetics that influences the rate of loss of genetic diversity. Sexual selection has the potential to reduce Ne by causing the sex-specific distributions of individuals that successfully reproduce to diverge. To empirically estimate the effect of sexual selection on Ne, we obtained fitness distributions for males and females from an outbred, laboratory-adapted population of Drosophila melanogaster. We observed strong sexual selection in this population (the variance in male reproductive success was ∼14 times higher than that for females), but found that sexual selection had only a modest effect on Ne, which was 75% of the census size. This occurs because the substantial random offspring mortality in this population diminishes the effects of sexual selection on Ne, a result that necessarily applies to other high fecundity species. The inclusion of this random offspring mortality creates a scaling effect that reduces the variance/mean ratios for male and female reproductive success and causes them to converge. Our results demonstrate that measuring reproductive success without considering offspring mortality can underestimate Ne and overestimate the genetic consequences of sexual selection. Similarly, comparing genetic diversity among different genomic components may fail to detect strong sexual selection.
Data from: Effective size in density-dependent two-sex populations: the effect of mating systems
Density dependence in vital rates is a key feature affecting temporal fluctuations of natural populations. This has important implications for the rate of random genetic drift. Mating systems also greatly affect effective population sizes, but knowledge of how mating system and density regulation interact to affect random genetic drift is poor. Using theoretical models and simulations, we compare Ne in short-lived, density dependent animal populations with different mating systems. We study the impact of a fluctuating, density dependent sex ratio and consider both a stable and a fluctuating environment. We find a negative relationship between annual Ne/N and adult population size N due to density dependence, suggesting that loss of genetic variation is reduced at small densities. The magnitude of this decrease was affected by mating system and life history. A male-biased, density dependent sex ratio reduces the rate of genetic drift compared to an equal, density independent sex ratio, but a stochastic change towards male-bias reduces the Ne/N ratio. Environmental stochasticity amplifies temporal fluctuations in population size, and is thus vital to consider in estimation of effective population sizes over longer time periods. Our results on the reduced loss of genetic variation at small densities, particularly in polygamous populations, indicate that density regulation may facilitate adaptive evolution at small population sizes.
Data from: Temporal genetic stability and high effective population size despite fisheries-induced life-history trait evolution in the North Sea sole.
Heavy fishing and other anthropogenic influences can have profound impact on a species' resilience to harvesting. Besides the decrease of the census and effective population size, strong declines in mature adults and recruiting individuals may lead to almost irreversible genetic changes in life-history traits. Here, we investigated the evolution of genetic diversity and effective population size in the heavily exploited sole (Solea solea), through the analysis of historical DNA from a collection of 1379 sole (Solea solea) otoliths dating back from 1957. Despite documented shifts in life-history traits, neutral genetic diversity inferred from 11 microsatellite markers showed a remarkable stability over a period of 50 years of heavy fishing. Using simulations and corrections for fisheries induced demographic variation, both point and temporal estimates of effective population size (Ne) were always higher than 1000, suggesting that despite the severe census size decrease over a 50 year period of harvesting, genetic drift is probably not strong enough to significantly decrease the neutral diversity of this species in the North Sea. However the ratio of effective population size to the census size (Ne/Nc) was very small (10-5), suggesting that overall only few adults contribute to the next generation. The high Ne level together with the low Ne/Nc ratio is most likely caused by a combination of an equalized reproductive output of younger cohorts, a decrease in generation time and a large variance in reproductive success typical for marine species. Because strong evolutionary changes in age and size at first maturation have been observed for sole, changes in adaptive genetic variation should be further monitored to detect the evolutionary consequences of human-induced selection.
Data from: Evolution of mutation rates in hypermutable populations of Escherichia coli propagated at very small effective population size
Mutation is the ultimate source of the genetic variation—including variation for mutation rate itself—that fuels evolution. Natural selection can raise or lower the genomic mutation rate of a population by changing the frequencies of mutation rate modifier alleles associated with beneficial and deleterious mutations. Existing theory and observations suggest that where selection is minimized, rapid systematic evolution of mutation rate either up or down is unlikely. Here, we report systematic evolution of higher and lower mutation rates in replicate hypermutable Escherichia coli populations experimentally propagated at very small effective size—a circumstance under which selection is greatly reduced. Several populations went extinct during this experiment, and these populations tended to evolve elevated mutation rates. In contrast, populations that survived to the end of the experiment tended to evolve decreased mutation rates. We discuss the relevance of our results to current ideas about the evolution, maintenance and consequences of high mutation rates.
Data from: Linkage disequilibrium and effective population size when generations overlap
Estimates of effective population size are critical for species of conservation concern. Genetic datasets can be used to provide robust estimates of this important parameter. However, the methods used to obtain these estimates assume that generations are discrete. We used simulated data to assess the influences of overlapping generations on estimates of effective size provided by the linkage disequilibrium method. Our simulations focus on two factors: the degree of reproductive skew exhibited by the focal species and the generation time, without considering sample size or the level of polymorphism at marker loci. In situations where a majority of reproduction is achieved by a small fraction of the population, the effective number of breeders can be much smaller than the per generation effective population size. The linkage disequilibrium in samples of newborns can provide estimates of the former size, while our results indicate that the latter size is best estimated using random samples of reproductively mature adults. Using samples of adults, the downwards bias was less than ~15% across our simulated life histories. As noted in previous assessments, precision of the estimate depends on the magnitude of effective size itself, with greater precision achieved for small populations.
Data from: Modeling the growth and decline of pathogen effective population size provides insight into epidemic dynamics and drivers of antimicrobial resistance
Non-parametric population genetic modeling provides a simple and flexible approach for studying demographic history and epidemic dynamics using pathogen sequence data. Existing Bayesian approaches are premised on stochastic processes with stationary increments which may provide an unrealistic prior for epidemic histories which feature extended period of exponential growth or decline. We show that non-parametric models defined in terms of the growth rate of the effective population size can provide a more realistic prior for epidemic history. We propose a non-parametric autoregressive model on the growth rate as a prior for effective population size, which corresponds to the dynamics expected under many epidemic situations. We demonstrate the use of this model within a Bayesian phylodynamic inference framework. Our method correctly reconstructs trends of epidemic growth and decline from pathogen genealogies even when genealogical data is sparse and conventional skyline estimators erroneously predict stable population size. We also propose a regression approach for relating growth rates of pathogen effective population size and time-varying variables that may impact the replicative fitness of a pathogen. The model is applied to real data from rabies virus and Staphylococcus aureus epidemics. We find a close correspondence between the estimated growth rates of a lineage of methicillin-resistant S. aureus and population-level prescription rates of beta-lactam antibiotics. The new models are implemented in an open source R package called skygrowth which is available at https://github.com/mrc-ide/skygrowth.
Data from: A comparison of single-sample estimators of effective population sizes from genetic marker data
In molecular ecology and conservation genetics studies, the important parameter of effective population size (Ne) is increasingly estimated from a single sample of individuals taken at random from a population and genotyped at a number of marker loci. Several estimators are developed, based on the information of linkage disequilibrium (LD), heterozygote excess (HE), molecular coancestry (MC) and sibship frequency (SF) in marker data. The most popular is the LD estimator, because it is more accurate than HE and MC estimators and is simpler to calculate than SF estimator. However, little is known about the accuracy of LD estimator relative to that of SF and about the robustness of all single-sample estimators when some simplifying assumptions (e.g. random mating, no linkage, no genotyping errors) are violated. This study fills the gaps and uses extensive simulations to compare the biases and accuracies of the four estimators for different population properties (e.g. bottlenecks, nonrandom mating, haplodiploid), marker properties (e.g. linkage, polymorphisms) and sample properties (e.g. numbers of individuals and markers) and to compare the robustness of the four estimators when marker data are imperfect (with allelic dropouts). Extensive simulations show that SF estimator is more accurate, has a much wider application scope (e.g. suitable to nonrandom mating such as selfing, haplodiploid species, dominant markers) and is more robust (e.g. to the presence of linkage and genotyping errors of markers) than the other estimators. An empirical data set from a Yellowstone grizzly bear population was analysed to demonstrate the use of the SF estimator in practice.
Data from: Alternative reproductive tactics increase effective population size and decrease inbreeding in wild Atlantic salmon
While nonanadromous males (stream-resident and/or mature male parr) contribute to reproduction in anadromous salmonids, little is known about their impacts on key population genetic parameters. Here, we evaluated the contribution of Atlantic salmon mature male parr to the effective number of breeders (Nb) using both demographic (variance in reproductive success) and genetic (linkage disequilibrium) methods, the number of alleles, and the relatedness among breeders. We used a recently published pedigree reconstruction of a wild anadromous Atlantic salmon population in which 2548 fry born in 2010 were assigned parentage to 144 anadromous female and 101 anadromous females that returned to the river to spawn in 2009 and to 462 mature male parr. Demographic and genetic methods revealed that mature male parr increased population Nb by 1.79 and 1.85 times, respectively. Moreover, mature male parr boosted the number of alleles found among progenies. Finally, mature male parr were in average less related to anadromous females than were anadromous males, likely because of asynchronous sexual maturation between mature male parr and anadromous fish of a given cohort. By increasing Nb and allelic richness, and by decreasing inbreeding, the reproductive contribution of mature male parr has important evolutionary and conservation implications for declining Atlantic salmon populations.
Data from: The effect of neighborhood size on effective population size in theory and in practice
The distinction between the effective size of a population (Ne) and the effective size of its neighborhoods (Nn) has sometimes become blurred. Ne reflects the effect of random sampling on the genetic composition of a population of size N, whereas Nn is a measure of within-population spatial genetic structure and depends strongly on the dispersal characteristics of a species. Although Nn is independent of Ne, the reverse is not true. Using simulations of a population of annual plants, it was found that the effect of Nn on Ne was well approximated by Ne=N/(1−FIS), where FIS (determined by Nn) was evaluated population wide. Nn only had a notable influence of increasing Ne as it became smaller (less than or equal to16). In contrast, the effect of Nn on genetic estimates of Ne was substantial. Using the temporal method (a standard two-sample approach) based on 1000 single-nucleotide polymorphisms (SNPs), and varying sampling method, sample size (2–25% of N) and interval between samples (T=1–32 generations), estimates of Ne ranged from infinity to <0.1% of the true value (defined as Ne based on 100% sampling). Estimates were never accurate unless Nn and T were large. Three sampling techniques were tested: same-site resampling, different-site resampling and random sampling. Random sampling was the least biased method. Extremely low estimates often resulted when different-site resampling was used, especially when the population was large and the sample fraction was small, raising the possibility that this estimation bias could be a factor determining some very low Ne/N that have been published.
Data from: Reconstructing the phylogenetic history of long-term effective population size and life-history traits using patterns of amino acid replacement in mitochondrial genomes of mammals and birds
The nearly neutral theory, which proposes that most mutations are deleterious or close to neutral, predicts that the ratio of nonsynonymous over synonymous substitution rates (dN/dS), and potentially also the ratio of radical over conservative amino acid replacement rates (Kr/Kc), are negatively correlated with effective population size. Previous empirical tests, using life-history traits (LHT) such as body-size or generation-time as proxies for population size, have been consistent with these predictions. This suggests that large-scale phylogenetic reconstructions of dN/dS or Kr/Kc might reveal interesting macroevolutionary patterns in the variation in effective population size among lineages. In this work, we further develop an integrative probabilistic framework for phylogenetic covariance analysis introduced previously, so as to estimate the correlation patterns between dN/dS, Kr/Kc, and three LHT, in mitochondrial genomes of birds and mammals. Kr/Kc displays stronger and more stable correlations with LHT than does dN/dS, which we interpret as a greater robustness of Kr/Kc, compared with dN/dS, the latter being confounded by the high saturation of the synonymous substitution rate in mitochondrial genomes. The correlation of Kr/Kc with LHT was robust when controlling for the potentially confounding effects of nucleotide compositional variation between taxa. The positive correlation of the mitochondrial Kr/Kc with LHT is compatible with previous reports, and with a nearly neutral interpretation, although alternative explanations are also possible. The Kr/Kc model was finally used for reconstructing life-history evolution in birds and mammals. This analysis suggests a fairly large-bodied ancestor in both groups. In birds, life-history evolution seems to have occurred mainly through size reduction in Neoavian birds, whereas in placental mammals, body mass evolution shows disparate trends across subclades. Altogether, our work represents a further step toward a more comprehensive phylogenetic reconstruction of the evolution of life-history and of the population-genetics environment.
Data from: Fitness decline in spontaneous mutation accumulation lines of Caenorhabditis elegans with varying effective population sizes
The rate and fitness effects of new mutations have been investigated by mutation accumulation (MA) experiments in which organisms are maintained at a constant minimal population size to facilitate the accumulation of mutations with minimal efficacy of selection. We evolved 35 MA lines of Caenorhabditis elegans in parallel for 409 generations at three population sizes (N = 1, 10, and 100), representing the first spontaneous long-term MA experiment at varying population sizes with corresponding differences in the efficacy of selection. Productivity and survivorship in the N = 1 lines declined by 44% and 12%, respectively. The average effects of deleterious mutations in N = 1 lines are estimated to be 16.4% for productivity and 11.8% for survivorship. Larger populations (N = 10 and 100) did not suffer a significant decline in fitness traits despite a lengthy and sustained regime of consecutive bottlenecks exceeding 400 generations. Together, these results suggest that fitness decline in very small populations is dominated by mutations with large deleterious effects. It is possible that the MA lines at larger population sizes contain a load of cryptic deleterious mutations of small to moderate effects that would be revealed in more challenging environments.
Data from: Large fluctuations in the effective population size of the malaria mosquito Anopheles gambiae s.s. during vector control cycle
On Bioko Island, Equatorial Guinea, indoor residual spraying (IRS) has been part of the Bioko Island Malaria Control Project since early 2004. Despite success in reducing childhood infections, areas of high transmission remained on the island. We therefore examined fluctuations in the effective population size (N_e) of the malaria vector Anopheles gambiae in an area of persistent high transmission over two spray rounds. We analyzed data for 13 microsatellite loci from 791 An. gambiae specimens collected at 6 time points in 2009 and 2010 and reconstructed the demographic history of the population during this period using Approximate Bayesian Computation (ABC). Our analysis shows that IRS rounds have a big impact on N_e, reducing it by 65% to 92% from pre-spray round N_e. More importantly our analysis shows that after 3-5 months, the An. gambiae population rebounded by 2,818% compared to shortly following the spray round. Our study underscores the importance of adequate spray round frequency to provide continuous suppression of mosquito populations, and that increased spray round frequency should substantially improve the efficacy of IRS campaigns. It also demonstrates the ability of ABC to reconstruct a detailed demographic history across only a few tens of generations in a large population.
Data from: Effective population size in eusocial Hymenoptera with worker-produced males
In many eusocial Hymenoptera, a proportion of males are produced by workers. To assess the effect of male production by workers on the effective population size Ne, a general expression of Ne in Hymenoptera with worker-produced males is derived on the basis of the genetic drift in the frequency of a neutral allele. Stochastic simulation verifies that the obtained expression gives a good prediction of Ne under a wide range of conditions. Numerical computation with the expression indicates that worker reproduction generally reduces Ne. The reduction can be serious in populations with a unity or female biased breeding sex ratio. Worker reproduction may increase Ne in populations with a male biased breeding sex ratio, only if each laying worker produce a small number of males and the difference of male progeny number among workers is not large. Worker reproduction could be an important cause of the generally lower genetic variation found in Hymenoptera, through its effect on Ne.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.