Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
126
datasets available to search
ShareScore release 0.7.1
Dataset results
126 results for “occupancy model”
Data for: Occupancy–detection models with museum specimen data: Promise and pitfalls
Open the record for dataset details and reuse information.
Multi-species occupancy modeling provides novel insights into amphibian metacommunity structure and wetland restoration
<p>A fundamental goal of community ecology is to understand species-habitat relationships and how they shape metacommunity structure. Recent advances in occupancy modeling enable habitat relationships to be assessed for both common and rare species within metacommunities using multi-species occupancy models (MSOM). These models account for imperfect species detection and offer considerable advantages over other analytical tools commonly used for community analyses under the elements of metacommunity structure (EMS) framework. Here, we demonstrate that MSOM can be used to infer habitat relationships and test metacommunity theory, using amphibians. Repeated frog surveys were undertaken at 55 wetland sites in eastern Australia. We detected 11 frog species from three families (Limnodynastidae, Myobatrachidae and Pelodryadidae). The rarest species was detected at only one site whereas the most common species was detected at 42 sites (naïve occupancy rate: 0.02 – 0.76). Two models were assessed representing two competing hypotheses; the best-supported model included the covariates distance to the nearest site (connectivity), wetland area, presence of the non-native eastern mosquitofish (<i>Gambusia holbrooki</i>), proportion cover of emergent vegetation, an interaction term between Gambusia and emergent vegetation cover, and the proportion canopy cover over a site. Hydroperiod played no detectable role in metacommunity structure. We found species-habitat relationships that fit with current metacommunity theory – occupancy increased with wetland area and connectivity. There was a strong negative relationship between occupancy and the presence of predatory Gambusia, and a positive interaction between Gambusia and emergent vegetation. The presence of canopy cover strongly increased occupancy for several tree frog species, highlighting the importance of terrestrial habitat for amphibian community structure. We demonstrated how responses by amphibians to environmental covariates at the species level can be linked to occupancy patterns at the metacommunity scale. Our results have clear management implications – wetland restoration projects for amphibians and likely other taxa should maximize wetland area and connectivity, establish partial canopy cover, and eradicate Gambusia or provide aquatic vegetation to mitigate the impact of this non-native fish. We strongly advocate the use of MSOM to elucidate the habitat drivers behind animal occupancy patterns and to derive unbiased occupancy estimates for monitoring programmes.</p>
Data from: Validating dispersal distances inferred from autoregressive occupancy models with genetic parentage assignments
1.Dispersal distances are commonly inferred from occupancy data but have rarely been validated. Estimating dispersal from occupancy data is further complicated by imperfect detection and the presence of unsurveyed patches. 2.We compared dispersal distances inferred from seven years of occupancy data for 212 wetlands in a metapopulation of the secretive and threatened California black rail (Laterallus jamaicensis coturniculus) to distances between parent-offspring dyads identified with 16 microsatellites. 3.We used a novel autoregressive multi-season occupancy model that accounted for both unsurveyed patches and imperfect detection to quantify patch isolation using buffer radius (BRM) and incidence function (IFM) connectivity measures at 15 scales (1–10, 15, 20, 25, and 30 km). Connectivity measures were then fit as colonization covariates in occupancy models to estimate a model-averaged dispersal distance. 4.As predicted, colonization was more strongly related to connectivity at small spatial scales (< 10 km). AIC weights were greatest at 7 km for BRM and at 4 km for IFM. 5.Model-averaged dispersal distances (BRM = 7.46 km; IFM = 5.48 km) showed good agreement with the mean (± SE) dispersal distance from 23 parent-offspring dyads (5.58 ± 1.92 km), indicating reasonably accurate mean dispersal distances can be inferred from occupancy data when isolation strongly affects colonization.
Data from: Disentangling elevational richness: a multi-scale hierarchical Bayesian occupancy model of Colorado ant communities
Understanding the forces that shape the distribution of biodiversity across spatial scales is central in ecology and critical to effective conservation. To assess effects of possible richness drivers, we sampled ant communities on four elevational transects across two mountain ranges in Colorado, USA, with seven or eight sites on each transect and twenty repeatedly sampled pitfall trap pairs at each site each for a total of 90 days. With a multi-scale hierarchical Bayesian community occupancy model, we simultaneously evaluated the effects of temperature, productivity, area, habitat diversity, vegetation structure, and temperature variability on ant richness at two spatial scales, quantifying detection error and genus-level phylogenetic effects. We fit the model with data from one mountain range and tested predictive ability with data from the other mountain range. In total, we detected 105 ant species, and richness peaked at intermediate elevations on each transect. Species-specific thermal preferences drove richness at each elevation with marginal effects of site-scale productivity. Trap-scale richness was primarily influenced by elevation-scale variables along with a negative impact of canopy cover. Soil diversity had a marginal negative effect while daily temperature variation had a marginal positive effect. We detected no impact of area, land cover diversity, trap-scale productivity, or tree density. While phylogenetic relationships among genera had little influence, congeners tended to respond similarly. The hierarchical model, trained on data from the first mountain range, predicted the trends on the second mountain range better than multiple regression, reducing root mean squared error up to 65%. Compared to a more standard approach, this modeling framework better predicts patterns on a novel mountain range and provides a nuanced, detailed evaluation of ant communities at two spatial scales.
Data from: Distinguishing distribution dynamics from temporary emigration using dynamic occupancy models
1. Dynamic occupancy models are popular for estimating dynamic distribution rates (colonization and extinction) from repeated presence/absence surveys of unmarked animals. This approach assumes closure among repeated samples within primary periods, allowing estimation of dynamic rates between these periods. However, the impact of temporary emigration (reversible changes in sampling availability) on dynamic rate estimates, has not been tested. 2. Using simulated data, we investigated the degree to which temporary emigration could mislead researchers interested in quantifying dynamics. We then compared results from three avian point count datasets to evaluate the likelihood that temporary emigration confounds estimates of dynamics for 19 species under a popular sampling protocol. 3. Simulated experiments indicated that when secondary periods were open to temporary emigration, presence of dynamics was correctly identified ≥ 95.1% of the time, and dynamic rate estimates were accurate. However, dynamic rate estimates were biased when secondary periods were closed to temporary emigration. In empirical datasets, dynamic occupancy models had greater support than closed models for all species when secondary sampling periods occurred in immediate succession (i.e., 3 samples within 10 minutes); however, our results suggest that this is because these estimates were heavily influenced by temporary emigration. When counts within a primary period were separated by 24-48 hours, we found evidence of dynamics for less than half of these species. We recommend an alternative sampling approach that allows accurate estimation of dynamic rates when temporary emigration is of no interest, and introduce a novel model for estimating both processes simultaneously in rare cases where they are both of biological interest. 4. Concern for violating the occupancy modeling closure assumption has led to widespread recommendations that samples within primary periods be conducted extremely close in time. However, this may not be the best approach when interest is in quantifying dynamic rates. While dynamic occupancy models provide estimates of 'colonization' and 'extinction,' these values do not inherently represent dynamics unless temporary emigration has been explicitly modeled, or accounted for with sampling design. Naiveté to this fact can result in incorrect conclusions about biological processes.
Data from: Sequential use of niche and occupancy models identifies conservation and research priority areas for two data-poor endemic birds from the Colombian Andes
<p>The lack of high-quality information on data-poor species can hinder efforts to inform conservation actions via spatial distribution modeling. This is particularly true for tropical birds of conservation concern, for which ecological studies and assessments of their conservation status have received limited funding. Here we use a cost- and time-efficient protocol for assessing the distribution of range-restricted taxa and to identify priority areas for their conservation based on a sequential application of Environmental Niche Models (ENMs) and Occupancy-Detection Models. This approach first uses available geographical information and niche-theory to prioritize potential study sites, which can later be surveyed to obtain high-quality presence-absence data to accurately model distributional ranges with limited resources. We apply this protocol to identify priority areas for two Neotropical birds of conservation concern endemic to the Colombian Andes: Yellow-headed Brush-finch (<i>Atlapetes flaviceps</i>) and Tolima Dove (<i>Leptotila conoveri</i>). We first fitted ENMs using spatially-filtered datasets containing all available records up to 2018. We then conducted field surveys across climatically suitable areas identified for both species, carrying out a total of 1750 counts to generate input data for the occupancy models. Overall, our results suggested more extended and more continuous distribution ranges for both species than previously reported, but also identified population strongholds that are not currently represented within the national protected areas system. Both species occupied a narrow elevational belt (~1300–2600) of the Central Andes of Colombia primarily on the slopes of the Magdalena River valley, with isolated populations in the Western and Eastern Andes; these areas have undergone some of the most marked landscape transformations in Colombia. This straightforward protocol maximizes available information and minimizes costs, while allowing for estimation of occurrence probabilities for range-restricted, data-poor taxa.</p>
Ignoring species availability biases occupancy estimates in single-scale occupancy models
<p>1. Most applications of single-scale occupancy models do not differentiate between availability and detectability, even though species availability is rarely equal to one. Species availability can be estimated using multi-scale occupancy models, and the availability process includes elements of species movement, behavior, and phenology. However, for the practical application of multi-scale occupancy models, it can be unclear what a robust sampling design looks like and what the statistical properties of the multi-scale and single-scale occupancy models are when availability is less than one.</p> <p>2. Using simulations, we explore the following common questions asked by ecologists during the design phase of a field study: (Q1) what is a robust sampling design for the multi-scale occupancy model when there are <i>a priori</i> expectations of parameter estimates?, (Q2) what is a robust sampling design when we have no expectations of parameter estimates?, and (Q3) can a single-scale occupancy model with a random effects term adequately absorb the extra heterogeneity produced when availability is less than one and provide reliable estimates of occupancy probability?.</p> <p>3. Our results show that there is a tradeoff between the number of sites and surveys needed to achieve a specified level of acceptable error for occupancy estimates using the multi-scale occupancy model. We also document that when species availability is low (< 0.40 on the probability scale), then single-scale occupancy models underestimate occupancy by as much as 0.40 on the probability scale, produce overly precise estimates, and provide poor parameter coverage. This pattern was observed when a random effects term was and was not included in the single-scale occupancy model, suggesting that adding a random-effects term does not adequately absorb the extra heterogeneity produced by the availability process. In contrast, when species availability was high (> 0.60), single-scale occupancy models performed similarly to the multi-scale occupancy model.</p> <p>4. As a companion, we provide an RShiny app that allows users to further explore our results and sampling designs across a number of different scenarios <a href="https://gdirenzo.shinyapps.io/multi-scale-occ/"><span>https://gdirenzo.shinyapps.io/multi-scale-occ/</span></a>. Our results suggest that unaccounted for availability can lead to underestimating species distributions when using single-scale occupancy models, which can have large implications on ecological inference and predictions for practitioners, such as those working at the front lines of invasion ecology, disease emergence, and species conservation. </p>
Data from: Sharing detection heterogeneity information among species in community models of occupancy and abundance can strengthen inference
<p>1. The estimation of abundance and distribution and factors governing patterns in these parameters is central to the field of ecology. The continued development of hierarchical models that best utilize available information to inform these processes is a key goal of quantitative ecologists. However, much remains to be learned about simultaneously modeling true abundance, presence, and trajectories of ecological communities.</p> <p>2. Simultaneous modeling of the population dynamics of multiple species provides an interesting mechanism to examine patterns in community processes and, as we emphasize herein, to improve species-specific estimates by leveraging detection information among species. Here we demonstrate a simple but effective approach to share information about observation parameters among species in hierarchical community abundance and occupancy models, where we use shared random effects among species to account for spatiotemporal heterogeneity in detection probability.</p> <p>3. We demonstrate the efficacy of our modeling approach using simulated abundance data, where we recover well our simulated parameters using N-mixture models. Our approach substantially increases precision in estimates of abundance compared to models that do not share detection information among species. We then expand this model, and apply it to repeated detection/non-detection data collected on six species of tits (Paridae) breeding at 119 1 km<sup>2</sup> sampling sites across a <em>P. montanus</em> hybrid zone in northern Switzerland (2004-2020). We find strong impacts of forest cover and elevation on population persistence and colonisation in all species. We also demonstrate evidence for interspecific competition on population persistence and colonization probabilities, where the presence of marsh tits reduces population persistence and colonisation probability of sympatric willow tits, potentially decreasing gene flow among willow tit subspecies.</p> <p>4. While conceptually simple, our results have important implications for the future modeling of population abundance, colonization, persistence, and trajectories in community frameworks. We suggest potential extensions of our modeling in this paper, and discuss how leveraging data from multiple species can improve model performance and sharpen ecological inference.</p>
Supplementary information for: A continuous-score occupancy modeling framework for incorporating uncertain machine learning output in autonomous biodiversity surveys
<p><span>Ecologists often study biodiversity by evaluating species occupancy and the relationship between occupancy and other covariates. Occupancy models are now widely used to account for false absences in field surveys and to reduce bias in estimates of covariate relationships. Existing occupancy models take as inputs binary detection/non-detection observations of species at each visit to each site. However, autonomous sensing devices and machine learning models are increasingly used to survey biodiversity, generating a new type of observation record (i.e., continuous-score data) that reflects the model's confidence a species is present in each autonomously sensed file, instead of binary detection/non-detection data. These data are not directly compatible with traditional binary occupancy modeling methods.</span></p> <p><span>Here, we develop a new occupancy model that models continuous scores on a visit level as a Gaussian mixture, combining a distribution of scores for files that do contain the species of interest and a distribution of scores for files that do not. The model takes as input continuous scores for each autonomously sensed and classified file, along with an optional small number of binary, manually verified detection and non-detection annotations.</span></p> <p><span>We present a simulation study that shows that over a range of empirically realistic parameters, our model outperforms traditional occupancy models that are based on binary annotation alone. We also apply this new model to an empirical case study using data generated from five machine learning classifiers applied to autonomous acoustic recordings gathered in the eastern United States.</span></p> <p><span>Because our occupancy model generalizes allowable input data beyond binary observations, it is particularly well-suited to the increasing volume of machine learning classified data in ecology and conservation.</span></p>
Data from: Integrating niche and occupancy models to infer the distribution of an endemic fossorial snake (Atractus lasallei)
<p>Understanding species distribution and habitat preferences is crucial for effective conservation strategies. However, the lack of information about population responses to environmental change at different scales hinders effective conservation measures. In this study, we estimate the potential and realized distribution of <em>Atractus lasallei</em>, a semi-fossorial snake endemic to the northwestern region of Colombia. We modelled the potential distribution of <em>A. lasallei</em> based on ecological niche theory (using maxent), and habitat use was characterized while accounting for imperfect detection using a single-season occupancy model. Our results suggest that <em>A. lasallei</em> selects areas characterized by slopes below 10°, with high average annual precipitation (>2500mm/year) and herbaceous and shrubby vegetation. Its potential distribution encompasses the northern Central Cordillera and two smaller centers along the Western Cordillera, but its habitat is heavily fragmented within this potential distribution. When the two models are combined, the species' realized distribution sums up to 935 km<sup>2</sup>, highlighting its vulnerability. We recommend approaches that focus on variability at different spatio-temporal scales to better comprehend the variables that affect species' ranges and identify threats to vulnerable species. Prompt actions are needed to protect herbaceous and shrub vegetation in this region, highly demanded for agriculture and cattle grazing.</p>
Accompanying data for the paper "Making Sense of Wildlife Habitat Use on Active Oil Sands Mines: Quasi-experiments, Occupancy Models, Trends Assessments, and Upland Habitat Reclamation"
<p>This data set contains both the raw species detection records and the derived occupancy model data used to assess usage patterns for the nine species of wildlife. Data have been anonymized by using non-identifying company and lease names. These attributes are not required to reproduce the results in this paper and was done per contractual requirements between LGL Limited and its clients.</p> <p>Data is currently being reviewed by the client and will be shared publicly once final approval has been received.</p>
Occupant Simulation Data based on Honda Accord 2024 Simplified Passenger Model and Full-factorial Sampling with 3,125 samples and HIII05F, HIII50M, HIII95M
<p>Database and FE-models with 9,375 Honda Accord 2014 passenger occupant simulations. </p> <p> </p>
Occupant Simulation Database and FE-Model based on Honda Accord 2024 Simplified Passenger Model and SOBOL Sampling with 8,192 samples and HIII05F, HIII50M, HIII95M
<p>Database and FE-models with 24,576 Honda Accord 2014 passenger occupant simulations. </p> <p> </p>
Occupant Simulation Database and FE-Model based on Honda Accord 2024 Simplified Passenger Model and SOBOL Sampling with 256 samples and HIII05F, HIII50M, HIII95M
<p>Database and FE-models with 768 Honda Accord 2014 passenger occupant FE-simulations. </p>
Using machine learning to model nontraditional spatial dependence in occupancy data
<p>Spatial models for occupancy data are used to estimate and map the true presence of a species, which may depend on biotic and abiotic factors as well as spatial autocorrelation. Traditionally researchers have accounted for spatial autocorrelation in occupancy data by using a correlated normally distributed site-level random effect, which might be incapable of modeling nontraditional spatial dependence such as discontinuities and abrupt transitions. Machine learning approaches have the potential to model nontraditional spatial dependence, but these approaches do not account for observer errors such as false absences. By combining the flexibility of Bayesian hierarchal modeling and machine learning approaches, we present a general framework to model occupancy data that accounts for both traditional and nontraditional spatial dependence as well as false absences. We demonstrate our framework using six synthetic occupancy data sets and two real data sets. Our results demonstrate how to model both traditional and nontraditional spatial dependence in occupancy data which enables a broader class of spatial occupancy models that can be used to improve predictive accuracy and model adequacy.</p>
Data from: Multispecies site occupancy modeling and study design for spatially replicated environmental DNA metabarcoding
<p>Although environmental DNA (eDNA) metabarcoding has become widely applied to gauge ecosystems in a noninvasive and cost-efficient manner, false negatives can occur due to various factors in its inherent multistage workflow. It is therefore essential to deal with this kind of species detection errors in eDNA metabarcoding to achieve accurate assessment of species distribution and diversity. To address this issue, we proposed a variant of the multispecies site occupancy model for eDNA metabarcoding studies and applied it to an eDNA metabarcoding dataset of freshwater fish communities collected in the Kasumigaura watershed in Japan.</p> <ul> </ul>
Multi‐species occupancy modeling reveals methodological and environmental effects on eDNA detection of amphibians in temporary ponds
<p>Aquatic environmental DNA is increasingly used for biodiversity monitoring, such as surveying threatened and invasive species. Mainstreaming these methods in practical applications, however, still requires significant standardisation and optimisation, namely regarding DNA capture methods. Here we evaluated how filter type (standard disc filters vs high-capacity capsules), number of sampling sites, volume of water filtered and environmental factors affected amphibian detection in Mediterranean temporary ponds. The study involved water filtering until clogging at one (capsules) and five (discs) sites from 16 small and shallow ponds, where three urodele and seven anuran species were recorded through sweep-netting and adult observations. Detection probabilities were estimated from site occupancy models based on replicate sampling and from an adaptation of time-to-detection models relating detection probability to volume of water filtered. Discs filtered relatively small volumes (15–1250 mL), with detection probabilities of the two abundant species (<em>Pelobates</em> <em>cultripes</em>, <em>Hyla</em> <em>meridionalis</em>) increasing rapidly with sample size and water volume, reaching almost perfect detection (0.95) at four and seven discs, and 420 mL and 1860 mL, respectively. However, reaching high detection probabilities for rare species (<em>Pelodytes</em> <em>atlanticus</em>, <em>Pleurodeles</em> <em>waltl</em>, <em>Triturus</em> <em>pygmaeus</em>) would require larger sampling effort than that used in our study. Despite filtering much larger volumes (600–5300 mL), filtering with capsules at a single site per pond provided lower detection probabilities for abundant species than filtering with discs at five sites. Rarer species showed no difference between methods, which may be due to small sample sizes and reduced statistical power for species with few detection. The effect of conductivity on species detectability was largely negative, while the influence of water clarity varied across species, and pH had no effects. Overall, our results suggest that eDNA amphibian surveys in Mediterranean temporary ponds need to consider filter clogging, heterogeneous DNA distribution, and highly conductive waters.</p>
Development and Validation of the Client Centered Occupational Therapy Service Model
ClinicalTrials.gov study NCT04465422. IPD Sharing: NO. Countries: 1. Publications: 4.
Data from: Validating dispersal distances inferred from autoregressive occupancy models with genetic parentage assignments
Open the record for dataset details and reuse information.
Data from: Multispecies site occupancy modeling and study design for spatially replicated environmental DNA metabarcoding
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.