Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
115
datasets available to search
ShareScore release 0.9.0
Dataset results
115 results for “abundance estimation”
Estimating spawning Green Sturgeon (Acipenser medirostris Ayres, 1854) abundance using side scan sonar and N-mixture models
Open the record for dataset details and reuse information.
Data from: Estimating fish population abundance by integrating quantitative data on environmental DNA and hydrodynamic modeling
<p>Molecular analysis of DNA left in the environment, known as environmental DNA (eDNA), has proven to be a powerful and cost-effective approach to infer occurrence of species. Nonetheless, relating measurements of eDNA concentration to population abundance remains difficult because detailed knowledge on the processes that govern spatial and temporal distribution of eDNA should be integrated to reconstruct the underlying distribution and abundance of a target species. In this study, we propose a general framework of abundance estimation for aquatic systems on the basis of spatially replicated measurements of eDNA. The proposed method explicitly accounts for production, transport, and degradation of eDNA by utilizing numerical hydrodynamic models that can simulate the distribution of eDNA concentrations within an aquatic area. It turns out that, under certain assumptions, population abundance can be estimated via a Bayesian inference of a generalized linear model. Application to a Japanese jack mackerel (<em>Trachurus japonicus</em>) population in Maizuru Bay revealed that the proposed method gives an estimate of population abundance comparable to that of a quantitative echo sounder method. Furthermore, the method successfully identified a source of exogenous input of eDNA (a fish market), which may render a quantitative application of eDNA difficult to interpret unless its effect is taken into account. These findings indicate the ability of eDNA to reliably reflect population abundance of aquatic macroorganisms; when the "ecology of eDNA" is adequately accounted for, population abundance can be quantified on the basis of measurements of eDNA concentration.</p>
Data from: Estimating abundance of an open population with an N-mixture model using auxiliary data on animal movements
Accurate assessment of abundance forms a central challenge in population ecology and wildlife management. Many statistical techniques have been developed to estimate population sizes because populations change over time and space, and to correct for the bias resulting from animals that are present in a study area but not observed. The mobility of individuals makes it difficult to design sampling procedures that account for movement into and out of areas with fixed jurisdictional boundaries. Aerial surveys are the gold standard used to obtain data of large mobile species in geographic regions with harsh terrain, but these surveys can be prohibitively expensive and dangerous. Estimating abundance with ground based census methods have practical advantages, but it can be difficult to simultaneously account for temporary emigration and observer error to avoid biased results. Contemporary research in population ecology increasingly relies on telemetry observations of the states and locations of individuals to gain insight on vital rates, animal movements, and population abundance. Analytical models that use observations of movements to improve estimates of abundance have not been developed. Here we build upon existing multi-state mark recapture methods using a hierarchical N-mixture model with multiple sources of data, including telemetry data on locations of individuals, to improve estimates of population sizes. We used a state-space approach to model animal movements to approximate the number of marked animals present within the study area at any observation period, thereby accounting for a frequently changing number of marked individuals. We illustrate the approach using data on a population of elk (Cervus elaphus nelsoni) in Northern Colorado, USA. We demonstrate substantial improvement compared to existing abundance estimation methods and corroborate our results from the ground based surveys with estimates from aerial surveys during the same seasons. We develop a hierarchical Bayesian N-mixture model using multiple sources of data on abundance, movement and survival to estimate the population size of a mobile species that uses remote conservation areas. The model improves accuracy of inference relative to previous methods for estimating abundance of open populations.
Data from: using camera traps and N-mixture models to estimate population abundance: model selection really matters
<p>Estimating the abundance or density of wildlife populations is a critical part of species conservation and management, but estimates can vary greatly in precision and accuracy according to the data collection and statistical methods, sampling and ecological variation, and sample size. N-mixture models are a common method which has been applied to a wide range of taxa for estimating population abundance from non-invasive data representing the distribution of the species. We used population estimates from an aerial survey of moose and videos from camera traps to assess the sensitivity of N-mixture models to ecological conditions, the spatial scale at which they were measured, the criteria used to define independent detections, and model choice based on the common statistical criterion of parsimony. The most parsimonious N-mixture models were considerably biased, producing implausibly large and considerably imprecise estimates of the abundance of moose. Most of the other models produced estimates of abundance that were ecologically realistic and relatively accurate. The accuracy of population estimates produced by N-mixture models were not overly sensitive to the formulation of models, the scale at which ecological conditions were measured, or the criteria used to define independent detection and by extension sample size. Our results suggest that parsimony was a poor measure of the predictive accuracy of the population estimates produced with the N-mixture model. Collecting and processing data from the aerial survey was less expensive and took less time, but data from camera traps can provide valuable information on behavior of the target species as well as insights into multiple species in the community.</p>
Data from: N-mixture models estimate abundance reliably: a field test on Marsh Tit using time-for-space substitution
<p>Imperfect detection in field studies on animal abundance, including birds, is common and can be corrected for in various ways. The binomial N-mixture (hereafter binmix) model developed for this task is widely used in ecological studies owing to its simplicity: it requires replicated count results as the input. However, it may overestimate abundance and be sensitive to even small violations of its assumptions. We used a 33-year dataset on the Marsh Tit, Poecile palustris, a sedentary forest passerine, from Białowieża Forest, Poland to validate inference from binmix models by comparing model-estimated abundances to the true number of breeding pairs within the plots, determined by exhaustive population study. The abundance estimates, derived from six springtime (April-May) counts of males on each plot in each year, were highly reliable: 116 out of 132 year-plot estimates (88%) included the true number of pairs within the 95% confidence intervals. Over- and underestimations were thus rare and similarly frequent (9 and 12 cases, respectively), with a tendency to overestimate at low densities and underestimate at high densities. Marsh Tits sing rarely but the frequency of countersinging increases with abundance, leading to non-independence in detections. When accounted for in a submodel for detection, the per-survey number of countersinging events positively affected detection probability but only weakly affected abundance estimates. Simulations further demonstrate that this property, overestimation at low densities and underestimation at high densities, may be a systematic bias of binmix model even if density-dependent detection is absent. While the behaviour of binmix models in specific situations requires more study, we conclude that these models are a valid tool to estimate abundance reliably when intensive population monitoring is not feasible.</p>
Estimating wolf abundance from cameras
<p>Detection histories for wolves derived from cameras in 3 study areas in Idaho, USA, 2016-2018. </p>
Estimating abundance in unmarked populations of Golden Eagle
<p> 1. Estimates of species abundance are of key importance in population and ecosystem level research but can be hard to obtain. Study designs using camera-traps are increasingly being used for large-scale monitoring of species that are elusive and/or occur naturally at low densities.</p> <p>2. Golden eagle (Aquila chrysaetos) is one such species, and we investigate whether existing large-scale monitoring programs using baited camera-traps can be used to estimate the abundance of golden eagles, as an alternative to traditional labour-intensive searches for active territories and nest sites during the breeding period.</p> <p>3. The camera-trap data allowed two measures of abundance to be estimated within each of four main study areas in mid and northern Norway; occupancy was measured as the probability of camera site use, and population size was measured as the number of eagle individuals using the camera sites within a study area. Spatial and temporal patterns in occupancy and population size were explored and evaluated against independent estimates of the breeding pair density in the study areas.</p> <p>4. Annual estimates of golden eagle occupancy showed low precision, while estimates of population size were more precise in relation to both estimated and anticipated abundance fluctuations. Estimates of population size may therefore be suitable for monitoring within study area temporal abundance trends, while estimates of occupancy seem unsuitable for such in golden eagles. Across study areas, patterns in both average occupancy and average population density estimated from population size, were consistent with the spatial pattern in average breeding pair densities (r = 0.99, and r = 0.89 respectively). This suggests that camera-trap based estimates of occupancy and population density reflect territory density at large spatial scales. In conclusion, our results suggest that baited camera-traps can be a cost-effective strategy for monitoring the abundance of golden eagles.</p>
Estimation of species abundance based on the number of segregating sites using environmental DNA (eDNA)
<p>The advancement of environmental DNA (eDNA) has enabled rapid and non-invasive species detection in aquatic environments. While most studies focus on detecting species presence or absence, recent research has explored using eDNA data to quantify species abundance. This estimation usually is based on the concentration of targeted eDNA. However, eDNA concentration can be influenced by various factors, both biotic and abiotic, which can obscure the relationship between concentration and species abundance. In this study, we suggest using the number of segregating sites as a proxy for estimating species abundance. We investigated this relationship in silico, in vitro, and in situ (mesocosm experiments) using two brackish goby species, <em>Acanthogobius hasta</em> and <em>Tridentiger bifasciatus</em>. Analysis of simulated and in vitro data, where DNA was mixed from a known number of individuals, revealed a strong correlation between the number of segregating sites and species abundance (R<sup>2</sup> > 0.9; P < 0.01). Results from the mesocosm experiment confirmed this correlation (R<sup>2</sup> = 0.70, P < 0.01). This correlation remained consistent despite biotic factors such as body size and feeding behavior of the fish (P > 0.05). Cross-validation tests demonstrated that the number of segregating sites predicts species abundance more accurately and reliably than eDNA concentration. In conclusion, the number of segregating sites is a precise and robust indicator of species abundance compared to eDNA concentration, offering a significant enhancement to the quantitative capabilities of eDNA technology.</p>
Microdata on vector abundance and IRS quality assurance (Estimating the impact of indoor residual spraying on sandfly abundance and incidence of visceral leishmaniasis in India from 2016 to 2022: an interrupted time-series analysis and modelling study)
<p>This repository contains the microdata on vector abundance and quality assurance of indoor residual spraying (IRS) that was used to estimate the impact of IRS on sandfly abundance and incidence of visceral leishmaniasis (VL) in India, as described in the paper "Estimating the impact of indoor residual spraying on sandfly abundance and incidence of visceral leishmaniasis in India from 2016 to 2022: an interrupted time-series analysis and modelling study" by Coffeng et al (<a href="https://doi.org/10.1016/S1473-3099(24)00420-1">https://doi.org/10.1016/S1473-3099(24)00420-1</a>). These data were collected as part of a BMGF-funded project led by dr. Michael Coleman at the Liverpool School for Tropical Medicine, as described in an earlier paper by Deb et al (<a href="https://doi.org/10.1371/journal.pntd.0009101">https://doi.org/10.1371/journal.pntd.0009101</a>).</p> <p>This repository does not include microdata on VL cases as these are owned by India's National Center for Vector Borne Disease Control (NCVBDC, <a href="https://ncvbdc.mohfw.gov.in/" target="_blank" rel="nofollow noreferrer noopener">https://ncvbdc.mohfw.gov.in/</a>).</p>
Data for Estimate of water and hydroxyl abundance on asteroid (16) Psyche from JWST data
<p>The data as ASCII files used to produce the figures in the manuscript, " Estimate of water and hydroxyl abundance on asteroid (16) Psyche from JWST data". Almost all the files are ECSV files created from astropy Table. These files have complete metadata headers. Otherwise, the data, where appropriate, include s3d fits IFU data cubes used to produce 1D spectra, the solar spectra output by the Planetary Spectrum Generator, a thermophysical model output, asteroid reflectance measurements, and chondrite reflectance measurements all of which are described in further detail below. </p> <p>The Figures folder contains data behind the figures where data were used.</p> <p>For Figure 2 this corresponds to the NIRSpec and MIRI flux observations with a column for wavelength (in microns), flux (in Jansky), and error (in Jansky). </p> <p>For Figure 3 this corresponds to the normalized reflectance for the NIRSpec observations. The first column in each dataset corresponds to the wavelength in microns, the second column to the normalized reflectance (unitless), the third column to the error (unitless), and in the case of the .txt files the fourth column corresponds to the gaussian fit to the normalized reflectance. </p> <p>For Figure 4 these are the normalized reflectance divided by the continuum of the groups of data along wavelength ranges described in the manuscript for Figure 4. The first column corresponds to the wavelength in microns, the second column to the reflectance, the third column to the continuum fit to the data, the fourth column is the normalized reflectance divided by the continuum, and the error. </p> <p>For Figure 5 these are the data used to produce the 3-micron feature plot with the gaussian fit for each NIRSpec observation. The first column corresponds to the wavelength in microns, second column to the normalized reflectance divided by the continuum (unitless) and subtracted from the mean of the continuum between 3.6 and 3.7 microns such that the average of the continuum would be 0, the fourth column the error (unitless), and the fifth column the gaussian fit to the data. </p> <p>For Figure 6 the only additional data needed to produce this plot beyond what is provided for Figure 5 is the IRTF data which is provided as a text file. The first column corresponds to the wavelength in microns, the second column to the normalized reflectance divided by the continuum (unitless), and the third column to the error (unitless). </p> <p>For Figure 7 the additional data used to produce this plot beyond what is provided for Figure 5 includes laboratory reflectance measurements of various chondrites. The .scl files are the original laboratory measurements where the first column corresponds to the wavelength in microns, the second column to the normalized reflectance (unitless) and the third column to the error to 1 significant digit (unitless). The scaled_lab_reflectance.txt file has all the laboratory reflectance data used to produce the Figure 7 plot that includes the wavelength in microns, the normalized scaled reflectance for the CM chondrite, the normalized scaled reflectance for the CH/CBb chondrite, and the normalized scaled reflectance for the CY chondrite. </p> <p>For Figure 8 the additional data include original normalized reflectance data for asteroids interamnia, themis, bamberga, and europa as .trim files. The columns for these files correspond to wavelength in microns, normalized reflectance (unitless), and the error (unitless). The 'continuumremoved' .txt files correspond to the wavelength in microns and the normalized reflectance divided by the continuum (unitless). </p> <p>For Figure 9 these are the MIRI emissivity data for the two sets of observations. The first column corresponds to the wavelength in microns, the second column to the normalized emission by taking the flux subtracting the solar flux and dividing by a thermophysical model (unitless), the error (unitless), and the gaussian fit to the normalized emission. </p> <p>For the supplemental figures these are the additional fit absorption features that are located in the appendix of the manuscript. For the 1.25 - 4.8 micron text files the columns correspond to the wavelength in microns, the continuum divided normalized reflectance (unitless), the error (unitless), and the gaussian fit. For the 5.74 - 5.98 micron text files the columns correspond to wavelength in microns, normalized emission divided by a linear fit to a region outside any potential features (unitless), the error (unitless), and the gaussian fit to the feature. </p> <p>The MIRI 1D spectra folder contains the ASCII files corresponding to the 1 arcsec aperture summed spectra from the MIRI observations. The columns correspond to wavelength in microns, flux in Jansky, and error in Jansky. </p> <p>The NIRSpec 1D spectra folder contains the ASCII files corresponding to the 1 arcsec aperture summed spectra from the NIRSpec observations. The columns correspond to wavelength in microns, flux in Jansky, and error in Jansky. </p> <p>The Solar Spectra folder contains the solar spectra flux used to reduce the data to produce the reflectance and emission spectra. The lbl text files are equivalent to the files with the same name without the lbl but also include the column information. The first column corresponds to the wavelength in microns, the second column to the total spectral irradiance in Jansky, the third column to the flux from the object in Jansky, the fourth column to the reflected sunlight in Jansky, and the fifth column to the thermal contribution in Jansky. The only column used for the solar reflectance used in the analysis corresponds to the fourth column. </p> <p>The ThermalModel folder contains the ASCII files with the data corresponding to the thermophysical model used to produce the reflectance (by subtracting the thermal contribution) and the emission (by dividing by the thermal contribution). The first column corresponds to the wavelength in microns, the second to the thermal flux in Jansky, the third to the smoothed thermal flux in Jansky, and the fourth column to the normalized thermal flux. Only column two corresponding to the thermal flux in Jansky was used in the analysis associated with the manuscript. </p>
Simulated wastewater sequencing data for benchmarking SARS-CoV-2 variant abundance estimation
<p>To evaluate the accuracy of variant abundance predictions from wastewater sequencing, we built a collection of benchmarking datasets that resemble real wastewater samples. For each variant (B.1.1.7, B.1.351, B.1.427, B.1.429, P.1) we created a series of 33 benchmarks by simulating sequencing reads from a variant genome, as well as a collection of background (non-variant of concern/interest) sequences, such that the variant abundance ranges from 0.05% to 100%. Analogously, we created a second series of benchmarks, simulating reads only from the Spike gene of each SARS-CoV-2 genome. We refer to the first set of benchmarks as "whole genome" (WG) and to the second set of benchmarks as "S-only". We repeated these simulations at different sequencing depths: 100x and 1000x coverage for the whole genome benchmarks, and 100x, 1000x, and 10,000x coverage for the S-only benchmarks.</p>
Considering sampling bias in close-kin mark-recapture (CKMR) abundance estimates of Atlantic salmon
<p>Genetic methods for the estimation of population size can be powerful alternatives to conventional methods. Close-kin mark-recapture (CKMR) is based on the principles of conventional mark-recapture, but instead of being physically marked, individuals are marked through their close kin. The aim of this study was to evaluate the potential of CKMR for the estimation of spawner abundance in Atlantic salmon and how age, sex, spatial, and temporal sampling bias may affect CKMR estimates. Spawner abundance in a wild population was estimated from genetic samples of adults returning in 2018 and of their potential offspring collected in 2019. Adult samples were obtained in two ways. First, adults were sampled and released alive in the breeding habitat during spawning surveys. Second, genetic samples were collected from out-migrating smolts PIT tagged in 2017 and registered when returning as adults in 2018. CKMR estimates based on adult samples collected during spawning surveys were somewhat higher than conventional counts. Uncertainty was small (CV<0.15), due to the detection of a high number of parent-offspring-pairs. Sampling of adults was age- and size-biased and correction for those biases resulted in moderate changes in the CKMR estimate. Juvenile dispersal was limited, but spatially balanced sampling of adults rendered CKMR estimates robust to spatially biased sampling of juveniles. CKMR estimates based on returning PIT tagged adults were approximately twice as high as estimates based on samples collected during spawning surveys. We suggest that estimates based on PIT tagged fish reflect the total abundance of adults entering the river, while estimates based on samples collected during spawning surveys reflect the abundance of adults present in the breeding habitat at the time of spawning. Our study showed that CKMR can be used to estimate spawner abundance in Atlantic salmon, with a moderate sampling effort, but a carefully designed sampling regime is required.</p>
Including a spatial predictive process in band recovery models improves inference for Lincoln estimates of animal abundance
<p>Abundance estimation is a critical component of conservation planning, particularly for exploited species where managers set regulations to restrict harvest based on current population size. An increasingly common approach for abundance estimation is through integrated population modeling (IPM), which uses multiple data sources in a joint likelihood to estimate abundance and additional demographic parameters. Lincoln estimators are one commonly used IPM component for harvested species, which combine information on the rate and the total number of individuals harvested within an integrated band-recovery framework to estimate abundance at large scales.</p> <p>A major assumption of the Lincoln estimator is that banding and recoveries are representative of the whole population, which may be violated if major sources of spatial heterogeneity in survival or harvest rates are not incorporated into the model. We developed an approach to account for spatial variation in harvest rates using a spatial predictive process, which we incorporated into a Lincoln estimator IPM.</p> <p>We simulated data under different configurations of sample sizes, harvest rates, and sources of spatial heterogeneity in harvest rate to assess potential model bias in parameter estimates. We then applied the model to data collected from a field study of wild turkeys (<em>Meleagris gallapavo</em>) to estimate local and statewide abundance in Maine, USA.</p> <p>We found that the band recovery model that incorporated a spatial predictive process consistently provided estimates of adult and juvenile abundance with low bias across a variety of spatial configurations of harvest rate and sampling intensities. When applied to data collected on wild turkeys, a model that did not incorporate spatial heterogeneity underestimated the harvest rate in some sub-regions. Consistent with simulation results, this led to over-estimation of both local and statewide abundance.</p> <p>Our work demonstrates that a spatial predictive process is a viable mechanism to account for spatial variation in harvest rates and limit bias in abundance estimates. This approach could be extended to large-scale band recovery datasets and has applicability for the estimation of population parameters in other ecological models as well.</p>
Abundance and population growth estimates for bare-nosed wombats
<p><span>Wildlife managers often rely on population estimates, but estimates can be challenging to obtain for geographically widespread species. Spotlight surveys provide abundance data for many species and, when conducted over wide spatial scales, have the potential to provide population estimates of geographically widespread species. The bare-nosed wombat (<em>Vombatus</em> <em>ursinus</em>) has a broad geographic range and is subject to spotlight surveys. We used 19 years (2002–2020) of annual spotlight surveys to provide the first estimates of population abundance for two of the three extant bare-nosed wombat subspecies: <em>V. u. ursinus</em> on Flinders Island; and <em>V. u. tasmaniensis</em> on the Tasmanian mainland. Using distance sampling methods, we estimated annual rates of change and 2020 population sizes for both sub-species. Tasmanian mainland surveys included habitat data, which allowed us to also look for evidence of habitat associations for <em>V. u. tasmaniensis</em>. The average wombat density estimate was higher on Flinders Island (0.42 ha<sup>-1</sup>, 95% CrI = 0.25 – 0.79) than on the Tasmanian mainland (0.11 ha<sup>-1</sup>, CrI = 0.07 – 0.19) and both wombat subspecies increased over the 19-year survey period with an estimated annual growth rate of 2.90% (CrI = -1.7 – 7.3) on Flinders Island and 1.20% (CrI = -1.1 – 2.9) on mainland Tasmania. Habitat associations for <em>V. u. tasmaniensis</em> were weak, possibly owing to survey design; however, we detected regional variation in density for this subspecies. We estimated the population size of <em>V. u. ursinus </em>to be 71,826 (CrI = 43,913 – 136,761) on Flinders Island, which when combined with a previously published estimate of 2,599 (CI = 2,254 – 2,858) from Maria Island, where the subspecies was introduced, provides a total population estimate. We also estimated 840,665 (CrI = 531,104 – 1,201,547) <em>V. u. tasmaniensis </em>on mainland Tasmania. These estimates may be conservative, owing to individual heterogeneity in when wombats emerge from burrows. Although these two sub-species are not currently threatened, our population estimates provide an important reference when assessing their population status in the future, and demonstrate how spotlight surveys can be valuable to inform management of geographically widespread species.</span></p>
Survey data and code to estimate abundance of Brachyramphus murrelets, Icy Bay, Alaska, USA
Open the record for dataset details and reuse information.
Data from: A hierarchical population model for the estimation of latent prey abundance and demographic rates of a nomadic predator
Open the record for dataset details and reuse information.
Considering sampling bias in close-kin mark-recapture (CKMR) abundance estimates of Atlantic salmon
Open the record for dataset details and reuse information.
A modeling framework for quantifying spatial recruitment dynamics using abundance estimation and sibship analysis: code and simulation study output
Open the record for dataset details and reuse information.
Data from: N-mixture models estimate abundance reliably: a field test on Marsh Tit using time-for-space substitution
Open the record for dataset details and reuse information.
Temporal variability in effective size (Ne) identifies potential sources of discrepancies between mark recapture and close kin mark recapture estimates of population abundance
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.