Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
319
datasets available to search
ShareScore release 0.9.0
Dataset results
319 results for “Population estimation”
Data from: using camera traps and N-mixture models to estimate population abundance: model selection really matters
Open the record for dataset details and reuse information.
Absolute fish population censuses in ponds demonstrate eDNA metabarcoding provides biodiversity estimates comparable to conventional sampling methods
Open the record for dataset details and reuse information.
Estimation of breeding population size using DNA-based pedigree reconstruction in brown bears
Open the record for dataset details and reuse information.
Estimating spatio-temporal reproductive dynamics of fish populations with passive acoustic monitoring: A state-space model approach
Open the record for dataset details and reuse information.
Supporting R-code and data for "Why are population growth rate estimates of past and present hunter-gatherers so different?" (Tallavaara and Jørgensen, 2020)
<p>This submission contains data and R-code that enable to reproduce the data manipulations and analyses in the paper “Why are population growth rate estimates of past and present hunter-gatherers so different?” by Miikka Tallavaara and Erlend Kirkeng Jørgensen (Philosophical transactions of the Royal Society B). Please, cite the paper and this Zenodo repository if you use the files included in this Zenodo record in your work.</p> <p>The submission includes a html-file titled “Why are population growth rate estimates of past and present hunter-gatherers so different? - Data analyses” (TJ2020.html) that contains R-code and instructions and comments for running the code (open this file in your browser). In addition, the submission includes Rdata-file (dataTJ2020.Rdata) containing all the data that are not created within the code and pure R-code (TJ2020.R).</p>
Major inconsistencies of inferred population genetic structure estimated in a large set of domestic horse breeds using microsatellites
<p>STRUCTURE remains the most applied tool aimed at recovering the true, but unknown, population structure from observed microsatellite data or other genetic markers. About 30% of <span class="Program"><span>STRUCTURE</span></span>-based studies could not be reproduced (Gilbert et al., 2012). Here we use a large set of data from 2323 horses from 93 domestic breeds plus the Przewalski horse, typed at 15 microsatellite markers, to evaluate how program settings, in particular the so far insufficiently evaluated number of replicates, impact the estimation of the optimal number of population clusters <i>K</i><sub>opt</sub> that best describe the observed data. Domestic horses are suited as a test case as there is extensive knowledge of the history of many breeds, extensive phylogenetic analyses. Different methods based on different genetic assumptions and statistical procedures (<span class="Program"><span>DAPC</span></span>, <span class="Program"><span>FLOCK</span></span>, PCoA and <span class="Program"><span>STRUCTURE</span></span> with different run scenarios) all revealed the general, broad-scale relationships among the breeds that largely reflect known breed histories but diverged largely how they characterized small-scale patterns. <span class="Program"><span>STRUCTURE</span></span> failed to consistently identify <i>K</i><sub>opt</sub> using the most widespread approach, the ΔK method, despite very large numbers of MCMCs (3,000,000) and replicates (100). The interpretation of breed structure over increasing numbers of<i> K</i>, without assuming a <i>K</i><sub>opt</sub>, was consistent with known breed histories. The over-reliance on <i>K</i><sub>opt</sub> should be replaced by a qualitative description of clustering over increasing <i>K</i>, which is scientifically more honest and has the advantage of being much faster and less computer intensive as lower numbers of MCMC iterations and repetitions suffice for stable results. Very large data sets are highly challenging for cluster analyses, especially when populations with complex genetic histories are investigated.</p>
Data from: An integrated population model for estimating the relative effects of natural and anthropogenic factors on a threatened population of steelhead trout
<p>This collection includes data on the abundance, age composition, and harvest of adult steelhead trout, as well as the numbers of juveniles released from a hatchery, for steelhead trout (<em>Oncorhyncus mykiss</em>) from the Skagit River in Washington, USA. It also includes estimates of the marine survival of hatchery-origin steelhead.</p>
Data from: Estimating the molecular evolutionary rates of mitochondrial genes referring to Quaternary Ice Age events with inferred population expansions and dispersals in Japanese Apodemus
Background: Determining reliable evolutionary rates of molecular markers is essential in illustrating historical episodes with phylogenetic inferences. Although emerging evidence has suggested a high evolutionary rate for intraspecific genetic variation, it is unclear how long such high evolutionary rates persist because a recent calibration point is rarely available. Other than using fossil evidence, it is possible to estimate evolutionary rates by relying on the well-established temporal framework of the Quaternary glacial cycles that would likely have promoted both rapid expansion events and interisland dispersal events. Results: We examined mitochondrial cytochrome b (Cytb) and control region (CR) gene sequences in two Japanese wood mouse species, Apodemus argenteus and A. speciosus, of temperate origin and found signs of rapid expansion in the population from Hokkaido, the northern island of Japan. Assuming that global warming after the last glacial period 7–10 thousand years before present (kyr BP) was associated with the expansion, the evolutionary rates (sites per million years, myr) of Cytb and CR were estimated as 11–16% and 22–32%, respectively, for A. argenteus, and 12–17% and 17–24%, respectively, for A. speciosus. Additionally, the significant signature of rapid expansion detected in the mtDNA sequences of A. speciosus from the remaining southern main islands, Honshu, Shikoku, and Kyushu, provided an estimated Cytb evolutionary rate of 3.1%/site/myr under the assumption of a postglacial population expansion event long ago, most probably at 130 kyr BP. Bayesian analyses using the higher evolutionary rate of 11–17%/site/myr for Cytb supported the recent demographic or divergence events associated with the Last Glacial Maximum. However, the slower evolutionary rate of 3.1%/site/myr would be reasonable for several divergence events that were associated with glacial periods older than 130 kyr BP. Conclusions: The faster and slower evolutionary rates of Cytb can account for divergences associated with the last and earlier glacial maxima, respectively, in the phylogenetic inference of murine rodents. The elevated evolutionary rate seemed to decline within 100,000 years.
Data from: Interannual variation in effective number of breeders and estimation of effective population size in long-lived iteroparous lake sturgeon (Acipenser fulvescens)
Quantifying interannual variation in effective adult breeding number (Nb) and relationships between Nb, effective population size (Ne), adult census size (N) and population demographic characteristics are important to predict genetic changes in populations of conservation concern. Such relationships are rarely available for long-lived iteroparous species like lake sturgeon (Acipenser fulvescens). We estimated annual Nb and generational Ne using genotypes from 12 microsatellite loci for lake sturgeon adults (n = 796) captured during ten spawning seasons and offspring (n = 3925) collected during larval dispersal in a closed population over 8 years. Inbreeding and variance Nb estimated using mean and variance in individual reproductive success derived from genetically identified parentage and using linkage disequilibrium (LD) were similar within and among years (interannual range of Nb across estimators: 41–205). Variance in reproductive success and unequal sex ratios reduced Nb relative to N on average 36.8% and 16.3%, respectively. Interannual variation in Nb/N ratios (0.27–0.86) resulted from stable N and low standardized variance in reproductive success due to high proportions of adults breeding and the species' polygamous mating system, despite a 40-fold difference in annual larval production across years (437–16 417). Results indicated environmental conditions and features of the species' reproductive ecology interact to affect demographic parameters and Nb/N. Estimates of Ne based on three single-sample estimators, including LD, approximate Bayesian computation and sibship assignment, were similar to annual estimates of Nb. Findings have important implications concerning applications of genetic monitoring in conservation planning for lake sturgeon and other species with similar life histories and mating systems.
Data from: Estimated six percent loss of genetic variation in wild populations since the industrial revolution
Genetic variation is fundamental to population fitness and adaptation to environmental change. Human activities are driving declines in many wild populations and could have similar effects on genetic variation. Despite the importance of estimating such declines, no global estimate of the magnitude of ongoing genetic variation loss has been conducted across species. By combining studies that quantified recent changes in genetic variation across a mean of 27 generations for 91 species, we conservatively estimate a 5.4-6.5% decline in within-population genetic diversity of wild organisms since the industrial revolution. This loss has been most severe for island species, which show a 30% average decline. We identified taxonomic and geographic gaps in temporal studies that must be urgently addressed. Our results are consistent with single time-point meta-analyses, which indicated that genetic variation is likely declining. However, our results represent the first confirmation of a global decline, and provide an estimate of the magnitude of the genetic variation lost from wild populations.
Data from: Estimating population size in the presence of temporary migration using a joint analysis of telemetry and capture recapture data
1.Temporary migration – where individuals can leave and re-enter a sampled population – is a feature of many capture–mark–recapture (CMR) studies of mobile populations which, if unaccounted for, can lead to biased estimates of population capture probabilities and consequently biased estimates of population abundance. 2. We present a method for incorporating radiotelemetry data within a CMR study to eliminate bias due to temporary migration using a Bayesian state-space model. 3. Our results indicate that using a relatively small number of telemetry tags, it is possible to greatly reduce bias in estimates of capture probabilities using telemetry data to model transition probabilities in and out of the sampling area. In a capture–recapture data set for trout Cod in the Murray river, Australia, accounting for temporary migration led to overall higher estimates of capture probabilities than models assuming permanent or zero migration. Also, individual heterogeneity in detectability can be managed through explicit modelling. We show how accounting for temporary migration when estimating capture probabilities can be used to estimate the abundance and size distribution of a population as though it were closed. 4. Our model provides a basis for more complex models that might integrate telemetry data into other CMR scenarios, thus allowing for greater precision in estimates of vital rates that might otherwise be biased by temporary migration. Our results highlight the importance of accounting for migration in survey design and parameter estimation, and the potential scope for supplementing large-scale CMR data sets with a subset of auxiliary data that provide information on processes that are hidden to primary sampling processes.
Data from: Among- and within-population variation in flowering time of Iberian Arabidopsis thaliana estimated in field and glasshouse conditions
The study of the evolutionary and population genetics of quantitative traits requires the assessment of within- and among-population patterns of variation. We carried out experiments including eight Iberian Arabidopsis thaliana populations (10 individuals per population) in glasshouse and field conditions. We quantified among- and within-population variation for flowering time and for several field life-history traits. Individuals were genotyped with microsatellites, single nucleotide polymorphisms and four well-known flowering genes (FRI, FLC, CRY2 and PHYC). Phenotypic and genotypic data were used to conduct QST–FST comparisons. Life-history traits varied significantly among- and within-populations. Flowering time also showed substantial within- and among-population variation as well as significant genotype × environment interactions among the various conditions. Individuals bearing FRI truncations exhibited reduced recruitment in field conditions and differential flowering time behavior across experimental conditions, suggesting that FRI contributes to the observed significant genotype × environment interactions. Flowering time estimated in field conditions was the only trait showing significantly higher quantitative genetic differentiation than neutral genetic differentiation values. Overall, our results show that these A. thaliana populations are genetically more differentiated for flowering time than for neutral markers, suggesting that flowering time is likely to be under divergent selection.
Data from: A model-derived short-term estimation method of effective size for small populations with overlapping generations
If not actively managed, small and isolated populations lose their genetic variability and the inbreeding rate increases. Combined, these factors limit the ability of populations to adapt to environmental changes, increasing their risk of extinction. The effective population size (Ne) is proportional to the loss of genetic diversity and therefore of considerable conservation relevance. However, estimators of Ne that account for demographic parameters in species with overlapping generations require sampling of populations across generations, which is often not feasible in long-lived species. We created an individual-based model that allows calculation of Ne based on demographic parameters that can be obtained in a time period much shorter than a generation. It can be adapted to every life-history parameter combination. The model is freely available as an r-package NEff. The model was first used in a simulation experiment observing changes in Ne in response to different degrees of generational overlap. Results showed that increased generational overlap slowed annual rates of heterozygosity loss, resulting in higher annual effective sizes (Ny) but decreased Ne per generation. Adding the effect of different recruitment rates only affected Ne for populations with low generational overlap. The model was further tested using real population data of the Australian arboreal gecko Gehyra variegata. Simulation results were compared to genetic analyses and matched estimates of the real population very well. Unlike other estimation methods of Ne, NEff neither requires long time series of population monitoring nor genetic analyses of changes in gene frequencies. Thus, it seems to be the first method for calculating Ne within short time periods and comparably low costs facilitating the use of Ne in applied conservation and management.
Data from: Population genetic and field ecological analyses return similar estimates of dispersal over space and time in an endangered amphibian
The explosive growth of empirical population genetics has seen a proliferation of analytical methods leading to a steady increase in our ability to accurately measure key population parameters, including genetic isolation, effective population size, and gene flow in natural systems. Assuming they yield similar results, population genetic methods offer an attractive complement to, or replacement of, traditional field ecological studies. However, empirical assessments of the concordance between direct field ecological and indirect population genetic studies of the same populations are uncommon in the literature. In this study, we investigate genetic isolation, rates of dispersal, and population sizes for the endangered California tiger salamander, Ambystoma californiense, across multiple breeding seasons in an intact vernal pool network. We then compare our molecular results to a previously published study based on multi-year, mark-recapture data from the same breeding sites. We found that field and genetic estimates of population size were only weakly correlated, but dispersal rates were remarkably congruent across studies and methods. In fact, dispersal probability functions derived from genetic data and traditional field ecological data were a significant match, suggesting that either method can be used effectively to assess population connectivity. These results provide one of the first explicit tests of the correspondence between landscape genetic and field ecological approaches to measuring functional population connectivity and suggest that even single-year genetic samples can return biologically meaningful estimates of natural dispersal and gene flow.
Data from: Estimating the relative fitness of escaped farmed salmon offspring in the wild and modeling the consequences of invasion for wild populations
Throughout their native range, wild Atlantic salmon populations are threatened by hybridization and introgression with escapees from net-pen salmon aquaculture. Although domestic-wild hybrid offspring have shown reduced fitness in lab and field experiments, consequential impacts on population abundance and genetic integrity remain difficult to predict in the field, in part because the strength of selection against domestic offspring is often unknown and context-dependent. Here we follow a single large escape event of farmed Atlantic salmon in southern Newfoundland and monitor changes in the in-river proportions of hybrids and feral individuals over time using genetically-based hybrid identification. Over a three-year period following the escape, the overall proportion of wild parr increased consistently (total wild proportion of 71.6%, 75.1%, 87.5% each year, respectively), with subsequent declines in feral (genetically pure farmed individuals originating from escaped, farmed adults) and hybrid parr. We quantify the strength of selection against parr of aquaculture ancestry and explore the genetic and demographic consequences for populations in the region. Within-cohort changes in the relative proportions of feral and F1 parr suggest reduced relative survival compared to wild individuals over the first (0.15 and 0.81 for feral and F1, respectively), and second years of life (0.26, 0.83). These relative survivorship estimates were used to inform an individual-based salmon eco-genetic model to project changes in adult abundance and overall allele frequency across three invasion scenarios ranging from short-term to long-term invasion and three relative survival scenarios. Modeling results indicate that total population abundance and time to recovery were greatly affected by relative survivorship and predict significant declines in wild population abundance under continued large escape events and calculated survivorship. Overall this work demonstrates the importance of estimating the strength of selection against domestic offspring in the wild to predict the long-term impact of farmed salmon escape events on wild populations.
Data from: Estimating effects of species interactions on populations of endangered species
Global change causes community composition to change considerably through time, with ever-new combinations of interacting species. To study the consequences of newly established species interactions, one available source of data could be observational surveys from biodiversity monitoring. However, approaches using observational data would need to account for niche differences between species and for imperfect detection of individuals. To estimate population sizes of interacting species, we extended N-mixture models that were developed to estimate true population sizes in single species. Simulations revealed that our model is able to disentangle direct effects of dominant on subordinate species from indirect effects of dominant species on detection probability of subordinate species. For illustration, we applied our model to data from a Swiss amphibian monitoring program and showed that sizes of expanding water frog populations were negatively related to population sizes of endangered yellow-bellied toads and common midwife toads and partly of natterjack toads. Unlike other studies that analyzed presence and absence of species, our model suggests that the spread of water frogs in Central Europe is one of the reasons for the decline of endangered toad species. Thus, studying population impacts of dominant species on population sizes of endangered species using data from biodiversity monitoring programs should help to inform conservation policy and to decide whether competing species should be subject to population management.
Data from: Use of hidden Markov capture-recapture models to estimate abundance in presence of uncertainty: application to estimating the prevalence of hybrids in animal populations
Estimating the relative abundance (prevalence) of different population segments is a key step in addressing fundamental research questions in ecology, evolution, and conservation. The raw percentage of individuals in the sample (naive prevalence) is generally used for this purpose, but it is likely to be subject to two main sources of bias. First, the detectability of individuals is ignored; second, classification errors may occur due to some inherent limits of the diagnostic methods. We developed a hidden Markov (also known as multievent) capture–recapture model to estimate prevalence in free‐ranging populations accounting for imperfect detectability and uncertainty in individual's classification. We carried out a simulation study to compare naive and model‐based estimates of prevalence and assess the performance of our model under different sampling scenarios. We then illustrate our method with a real‐world case study of estimating the prevalence of wolf (Canis lupus) and dog (Canis lupus familiaris) hybrids in a wolf population in northern Italy. We showed that the prevalence of hybrids could be estimated while accounting for both detectability and classification uncertainty. Model‐based prevalence consistently had better performance than naive prevalence in the presence of differential detectability and assignment probability and was unbiased for sampling scenarios with high detectability. We also showed that ignoring detectability and uncertainty in the wolf case study would lead to underestimating the prevalence of hybrids. Our results underline the importance of a model‐based approach to obtain unbiased estimates of prevalence of different population segments. Our model can be adapted to any taxa, and it can be used to estimate absolute abundance and prevalence in a variety of cases involving imperfect detection and uncertainty in classification of individuals (e.g., sex ratio, proportion of breeders, and prevalence of infected individuals).
Data from: Heritability estimates from genome wide relatedness matrices in wild populations: application to a passerine, using a small sample size
Genomic developments have empowered the investigation of heritability in wild populations directly from genome wide relatedness matrices (GRM). Such GRM based approaches can in particular be used to improve or substitute approaches based on social pedigree (PED-social). However, measuring heritability from GRM in the wild has not been widely applied yet, especially using small samples and in non-model species. Here, we estimated heritability for four quantitative traits (tarsus length, wing length, bill length and body mass), using PED-social and a pedigree corrected by genetic data (PED-corrected) and GRM from a small sample (n = 494) of blue tits from natural populations in Corsica genotyped at nearly 50,000 filtered SNPs derived from RAD-seq. We also measured genetic correlations among traits and we performed chromosome partitioning. Heritability estimates were slightly higher when using GRM compared to PED-social, and PED-corrected yielded intermediate values, suggesting a minor underestimation of heritability in PED-social due to incorrect pedigree links, including extra-pair paternity, and to lower information content than the GRM. Genetic correlations among traits were similar between PED-social and GRM but credible intervals were very large in both cases, suggesting a lack of power for this small dataset. Although a positive linear relationship was found between the number of genes per chromosomes and the chromosome heritability for tarsus length, chromosome partitioning similarly showed a lack of power for the three other traits. We discuss the usefulness and limitations of the quantitative genetic inferences based on genomic data in small samples from wild populations.
Data from: Accounting for heterogeneity when estimating stopover duration, timing and population size of red knots along the Luannan Coast of Bohai Bay, China
1. To successfully perform their long-distance migrations, migratory birds require sites along their migratory routes to rest and refuel. Monitoring the use of so-called stopover and staging sites provides insights into (1) the timing of migration and (2) the importance of a site for migratory bird populations. A recently developed Bayesian superpopulation model that integrates mark-recapture data and ring density data enabled the estimation of stopover timing, duration and population size. Yet, this model did not account for heterogeneity in encounter (p) and staying (ϕ) probabilities. 2. Here we extended the integrated superpopulation model by implementing finite mixtures to account for heterogeneity in p and ϕ. We used simulations and real data on red knots Calidris canutus staging in Bohai Bay, China, during spring migration to (1) show the importance of accounting for heterogeneity in encounter and staying probabilities to get unbiased estimates of stopover timing, duration and numbers of migratory birds at staging sites and (2) get accurate stopover parameter estimates for a migratory bird species at a key staging site that is threatened by habitat destruction. 3. Our simulations confirmed that heterogeneity in p affected stopover parameter estimates more than heterogeneity in ϕ. Bias was particularly severe when most birds had both low ϕ and p. Bias was largest for population size, intermediate for stopover duration and negligible for stopover timing. 4. 50,000-100,000 red knots were estimated to annually stop for 5-9 days in Bohai Bay between 10 and 30 May. This shows the key importance of this staging site for this declining species. There were no clear changes in stopover parameters over time. 5. Our study shows the importance of accounting for heterogeneity in both encounter and staying probabilities for accurately estimating stopover duration and population size and provides an appropriate modelling framework.
Data from: Relationship type affects the reliability of dispersal distance estimated using pedigree inferences in partially sampled populations: a case study involving invasive American mink in Scotland
Estimating dispersal—a key parameter for population ecology and management—is notoriously difficult. The use of pedigree assignments, aided by likelihood-based software, has become popular to estimate dispersal rate and distance. However, the partial sampling of populations may produce false assignments. Further, it is unknown how the accuracy of assignment is affected by the genealogical relationships of individuals and is reflected by software-derived assignment probabilities. Inspired by a project managing invasive American mink (Neovison vison), we estimated individual dispersal distances using inferred pairwise relationships of culled individuals. Additionally, we simulated scenarios to investigate the accuracy of pairwise inferences. Estimates of dispersal distance varied greatly when derived from different inferred pairwise relationships, with mother–offspring relationship being the shortest (average = 21 km) and the most accurate. Pairs assigned as maternal half-siblings were inaccurate, with 64%–97% falsely assigned, implying that estimates for these relationships in the wild population were unreliable. The false assignment rate was unrelated to the software-derived assignment probabilities at high dispersal rates. Assignments were more accurate when the inferred parents were older and immigrants and when dispersal rates between subpopulations were low (1% and 2%). Using 30 instead of 15 loci increased pairwise reliability, but half-sibling assignments were still inaccurate (>59% falsely assigned). The most reliable approach when using inferred pairwise relationships in polygamous species would be not to use half-sibling relationship types. Our simulation approach provides guidance for the application of pedigree inferences under partial sampling and is applicable to other systems where pedigree assignments are used for ecological inference.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.