Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

242

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

242 results for “occupancy data”

Learn how ShareScore rates datasets ↗
edi60/100

Pika habitat occupancy survey data for Niwot Ridge and Green Lakes Valley, 2016 - ongoing

Long-term monitoring of habitat occupancy can reveal patterns of habitat use, population dynamics, and factors controlling species distribution. The American pika (Ochotona princeps), a small mammal found in rocky habitats throughout western North America, has been targeted for occupancy studies due to its relatively conspicuous behavior and its unusual adaptations for surviving long, cold winters without hibernation. These adaptations include an unusually high resting metabolic rate and maintenance of body temperatures near the lethal maximum for this species, which would appear to compromise the pika's ability to survive warmer summers. Recent monitoring as well as projections based on future climate scenarios have suggested this species is experiencing a period of range retraction due to warming summers and/or loss of insulating winter snow cover. Niwot Ridge is situated ideally to test competing hypotheses about the trajectory and drivers of pika range shift. The pika is still common throughout the Colorado Rockies, but published models differ markedly regarding projections of the pika’s future distribution in this region. Niwot Ridge has experienced warmer summers as well as shorter periods of insulating snow cover in recent years, and there is evidence that pikas are now less common than they once were in at least one area on the ridge. This study is designed to provide robust data on pika population trends through long-term monitoring of occupancy in a spatially balanced random sample of pika habitat patches centered on Niwot Ridge. Survey plots (n = 72) were selected according to a Generalized Random-Tessellation Stratified (GRTS) algorithm, stratified dichotomously by elevation, average annual snow accumulation (SWE), and probabilities of pika occurrence based on previous data. Each plot extends 12 m in radius from a GRTS point. To ensure that each plot contains at least 10% cover of talus, plot coordinates were adjusted (usually less than 50 m) or replaced

openCC (other)May 2025View details →
edi56/100

Long-term demographic dataset for Cladonia perforata, including fine-scale cover, occupancy, and subpopulation area data, 2011-2024

This dataset includes all data pertaining to a long-term demographic study of Cladonia perforata (perforate reindeer lichen), a federally endangered lichen endemic to Florida, including fine-scale cover, occupancy, and population area data, conducted by the Archbold Biological Station Plant Ecology Program. This includes 13 years of data (2011-2024) from nine subpopulation (including seven at Archbold Biological Station, and two at the Lake Wales Ridge Wildlife and Environmental Area, Royce Unit), all located in rosemary scrub habitat within the Lake Wales Ridge metapopulation. This study sought to characterize the fire ecology and long-term population trends for the species, and thus also includes data on prescribed burn severity and time since fire. Data were collected using a stratified random plot design, with occupancy plots (presence/absence within 1.5 meter radius) throughout the subpopulation and a subset of these designated as cover plots only, with this cover data collected as point intercept hits within a 48x48cm area. Cover data also includes microhabitat data – canopy cover in densiometer reading and dominant ground cover. Cover and occupancy data were taken every 3 years for each subpopulation (subpopulations were on different yearly schedules). Subpopulation area was mapped using a submeter GPS unit every 6 years. Subpopulations were resampled for all metrics as soon as possible following a fire, and the sampling schedule was then reset.

openCC (other)Aug 2025View details →
edi48/100

Common Raven (Corvus corax) Occupancy Survey and Habitat Selection Data in Cliff Habitat of the Central Appalachian Region, USA, 2009-2010

We identified 24 cliff sites across four states of the Central Appalachian Region of the eastern USA (Kentucky, North Carolina, Virginia, and West Virginia) with known raven occupancy at which to perform occupancy surveys for estimating detection probability and the effects of covariates. We surveyed each cliff site 2-4 times in either 2009 or 2010 and recorded time-to-first detection and time to confirmed cliff occupancy during a two-hour survey. Daily surveys were completed between 06:00 and local solar noon. During each survey, we recorded covariates, including air temperature at survey start time, cloud cover, wind speed, and day of year. We also calculated the distance of the observation point from the cliff being surveyed and the forest cover around the cliff. We also collected data thought to be pertinent for habitat selection by ravens on 26 cliffs occupied by ravens and 26 cliffs deemed unoccupied by ravens in 2010. For each cliff, we measured cliff physiographic characteristics, such as cliff length, cliff height, and occlusion by vegetation, and landscape characteristics, including percent forest and urban cover around the cliff and distances from the cliff to the nearest road and human habitation.

openCC (other)Dec 2021View details →
zenodo44/100

Data for "How do ecologists estimate occupancy in practice?" by Goldstein et al.

<p>&nbsp;Data for review of occupancy estimtaion methods by Goldstein et al.</p> <p>&nbsp;</p> <p>Please see the accompanying manuscript for full methodology. We will link to it when the manuscript is published.</p> <p>&nbsp;</p> <p>This upload contains three datasets: "binary_scores_phase1.csv", "binary_scores_phase2.csv" and "modsel_results.csv". Each is a .csv file containing data from a survey of occupancy estimation practices. Each row represents one peer-reviewed paper, while each column represents a characteristic of the paper. Most columns are TRUE/FALSE, indicating whether or not the paper satisfied the relevant criterion.</p> <p>&nbsp;</p> <p>One additional raw resource is provided. The .zip file "all_papers_2022-04-07.zip" contains 6 .xls files giving the full set of papers returned by the original Web of Science search. These are unmodified from the initial search.</p> <p>&nbsp;</p> <p>All datasets contain the column:</p> <p>ID - A unique ID for each paper; it most cases, a DOI. When Web of Science returned an invalid DOI, the ID is set to the paper title instead.</p> <p>&nbsp;</p> <p>Note that all papers in Phase 2 are also in Phase 1, and all model selection papers are in both Phase 1 and Phase 2. The ID column can be used to join datasets.</p> <p>&nbsp;</p> <p>Across both datasets, if all options in a category are FALSE or if a field is NA, that may mean that the review team was unable to determine what choices the authors made.</p> <p>&nbsp;</p> <p>binary_scores_phase1.csv gives the results of Phase 2 of the review. It contains the following columns:</p> <p>&nbsp;</p> <p>coll_newdata - Did the authors analyze newly collected data?</p> <p>coll_existing - Did the authors analyze existing, previously published data?</p> <p>coll_longterm - Did the authors analyze data produced by a long-term monitoring program?</p> <p>coll_particip - Did the authors analyze participatory science data?</p> <p>eco_frshwtr - Was the study system a freshwater ecosystem?</p> <p>eco_marine - Was the study system a marine ecosystem?</p> <p>eco_terra - &nbsp;Was the study system a terrestrial ecosystem?</p> <p>region_USA - &nbsp;Were the data collected in the USA?</p> <p>region_NoAm - &nbsp;Were the data collected in North America?</p> <p>region_CenAm - &nbsp;Were the data collected in Central America?</p> <p>region_SoAm - &nbsp;Were the data collected in South America?</p> <p>region_Africa - &nbsp;Were the data collected in Africa?</p> <p>region_Eur - &nbsp;Were the data collected in Europe?</p> <p>region_Asia - &nbsp;Were the data collected in Asia?</p> <p>region_Oceania - &nbsp;Were the data collected in Oceania?</p> <p>framework_MLE - &nbsp;Did the authors estimate models in a maximum likelihood framework?</p> <p>framework_ML - &nbsp;Did the authors estimate models in a machine learning framework?</p> <p>framework_Bayes - &nbsp;Did the authors estimate models in a Bayesian framework?</p> <p>gof_AUC - &nbsp;Did the authors use area-under-the-curve to evaluate their models?</p> <p>gof_CV - &nbsp;Did the authors use cross validation to evaluate their models?</p> <p>gof_PPC - &nbsp;Did the authors use poserior predictive checks to evaluate their models?</p> <p>gof_parboot - &nbsp;Did the authors use parametric bootstrapping to evaluate their models?</p> <p>gof_bayespv - &nbsp;Did the authors use Bayesian p-values to evaluate their models?</p> <p>gof_any - &nbsp;Did the authors conduct any model checking?</p> <p>taxon_mammal - Were some or all of the study species mammals?</p> <p>taxon_bird - &nbsp;Were some or all of the study species birds?</p> <p>taxon_herp - Were some or all of the study species herptiles (reptiles and amphibians)?</p> <p>taxon_fish - &nbsp;Were some or all of the study species fish?</p> <p>taxon_arthro - &nbsp;Were some or all of the study species arthropods?</p> <p>taxon_othinv - &nbsp;Were some or all of the study species non-arthropod invertebrates?</p> <p>soft_unmarked - &nbsp;Did the authors estimate models using the software unmarked?</p> <p>soft_PRESENCE - &nbsp;Did the authors estimate models using PRESENCE-family software?</p> <p>soft_MARK - &nbsp;Did the authors estimate models using MARK-family software?</p> <p>soft_JAGS - &nbsp;Did the authors estimate models using JAGS-family software?</p> <p>soft_lme4 - &nbsp;Did the authors estimate models using the software lme4?</p> <p>soft_MaxEnt - &nbsp;Did the authors estimate models using the software MaxEnt?</p> <p>soft_baseR - &nbsp;Did the authors estimate models using custom models written in base-R?</p> <p>soft_NR - &nbsp;Did the authors fail to clearly report what software they used to estimate models?</p> <p>soft_other - &nbsp;Did the authors estimate models using some other software?</p> <p>nspecies - &nbsp;How many species did the authors study?</p> <p>nspec_1 - &nbsp;Did the authors analyze data on exactly 1 species?</p> <p>nspec_2 - &nbsp;Did the authors analyze data on exactly 2 species?</p> <p>nspec_3_5 - &nbsp;Did the authors analyze data on 3-5 species?</p> <p>nspec_6_10 - &nbsp;Did the authors analyze data on 6-10 species?</p> <p>nspec_11_20 - &nbsp; Did the authors analyze data on 11-20 species?</p> <p>nspec_20plus - &nbsp; Did the authors analyze data on more than 20 species?</p> <p>compare_avg - &nbsp; Did the authors conduct model averaging?</p> <p>compare_sel - &nbsp; Did the authors conduct model selection?</p> <p>compare_other - Did the authors compare multiple models without averaging or selecting between them?</p> <p>modsel_AICc - Did the authors use AICc in model selection or averaging?</p> <p>modsel_AIC - &nbsp;Did the authors use AIC in model selection or averaging?</p> <p>modsel_other - &nbsp;Did the authors use another information criterion in model selection or averaging?</p> <p>dattype_DND - Were some or all of the data collected as detection-nondetection data?</p> <p>dattype_count - &nbsp;Were some or all of the data collected as count data?</p> <p>dattype_PO - &nbsp;Were some or all of the data collected as presence-only data?</p> <p>dattype_other - &nbsp;Were some or all of the data collected in some other form?</p> <p>survey_visual - &nbsp;Were some or all of the data collected using in-person visual surveys?</p> <p>survey_audio - &nbsp;Were some or all of the data collected using in-person audio surveys?</p> <p>survey_camtrap - &nbsp;Were some or all of the data collected using camera trap surveys?</p> <p>survey_capture - &nbsp;Were some or all of the data collected using animal capture surveys?</p> <p>survey_sign - &nbsp;Were some or all of the data collected using sign surveys?</p> <p>survey_passaudio - &nbsp;Were some or all of the data collected using passive acoustic surveys?</p> <p>survey_DNA - &nbsp;Were some or all of the data collected using eDNA surveys?</p> <p>survey_other - &nbsp;Were some or all of the data collected using some other survey protocol?</p> <p>modtype_SSOM - &nbsp;Did the authors analyze data with an SSOM?</p> <p>modtype_DynOcc - &nbsp;Did the authors analyze data with a dynamic occupancy model?</p> <p>modtype_GLM - &nbsp;Did the authors analyze data with a GLM?</p> <p>modtype_cooccur - &nbsp;Did the authors analyze data with a multispecies co-occurrence model?</p> <p>modtype_community - &nbsp;Did the authors analyze data with a multispecies community model?</p> <p>modtype_MaxEnt - &nbsp; Did the authors analyze data with a MaxEnt model?</p> <p>modtype_other - &nbsp;Did the authors analyze data with another model?</p> <p>affil_acad - &nbsp;Did any of the authors have an academic affiliation?</p> <p>affil_govt - &nbsp;Did any of the authors have a government affiliation?</p> <p>affil_other - Did any of the authors have a private or NGO affiliation?</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>binary_scores_phase2.csv gives the results of Phase 2 of the review. It contains the following columns:</p> <p>&nbsp;</p> <p>code_avail - Did we determine that the authors published their model fitting code?</p> <p>data_avail - Did we determine that the authors published their data?</p> <p>nsite - At how many sites were data collected?</p> <p>nsite_lt20 - Were data collected at 20 or fewer sites?</p> <p>nsite_21_50 - Were data collected at 20-50 sites?</p> <p>nsite_51_100 - Were data collected at 51-100 sites?</p> <p>nsite_101p - Were data collected at more than 100 sites?</p> <p>time_pds - Over how many primary time periods (e.g. sampling seasons) were data collected?</p> <p>time_pds_1 - Were data collected during a single time period?</p> <p>time_pds_2_3 - Were data collected during 2-3 time periods?</p> <p>time_pds_4p - Were data collected during 4 or more time periods?</p> <p>nrepl - Roughly how many replicate surveys were collected per site?</p> <p>nrepl_1 - Was only one replicate survey conducted per site?</p> <p>nrepl_2_3 - Were 2-3 replicate surveys conducted per site?</p> <p>nrepl_4_5 - Were 4-5 replicate surveys conducted per site?</p> <p>nrepl_6p - Were 6 or more replicate surveys conducted per site?</p> <p>window - What sampling window was used to discretize continuous-time sampling? (continuous-time studies only; otherwise NA)</p> <p>det_window_lt_day - Was a sampling unit of less than one day used to discretize sampling?</p> <p>det_window_1day - Was a sampling unit of one day used to discretize sampling?</p> <p>det_window_2_6day - Was a sampling unit of 2-6 days used to discretize sampling?</p> <p>det_window_7_14day - Was a sampling unit of l7-14 days used to discretize sampling?</p> <p>det_window_15p_day - Was a sampling unit of 15 days or more used to discretize sampling?</p> <p>homerange_is_bigger - Did the authors describe their survey area per site as bigger than the target species' home range?</p> <p>homerange_is_smaller -Did the authors describe their survey area per site as smaller than the target species' home range?</p> <p>informative_priors - Did the authors use informative priors in a Bayesian analysis?</p> <p>time_in_mod_as_covar - Did the authors include primary time periods in the model as a covariate?</p> <p>time_in_mod_separate_model - Did the authors use separate models to analyze data collected during different primary time periods?</p> <p>time_in_mod_other - Did the authors include primary time periods in the model in some other way?</p> <p>ncovar_det - How many covariates were included in the best model's detection submodel?</p> <p>ncovar_det_zero - Did the detection submodel include 0 covariates?</p> <p>ncovar_det_1_3 - &nbsp;Did the detection submodel include 1-3 covariates?</p> <p>ncovar_det_4p - &nbsp;Did the detection submodel include 4 or more covariates?</p> <p>ncovar_occ - How many covariates were included in the best model's occupancy submodel?</p> <p>ncovar_occ_zero - &nbsp;Did the occupancy submodel include 0 covariates?</p> <p>ncovar_occ_1_3 - &nbsp;Did the occupancy submodel include 1-3 covariates?</p> <p>ncovar_occ_4p - &nbsp;Did the occupancy submodel include 4 or more covariates?</p> <p>covars_both_mods - Were any variables considered in both submodels?</p> <p>model_ranefs - Did the authors include any random effects?</p> <p>model_expl_spatial - Did the model have an explicit spatial component?</p> <p>motivating_q_range - Was range estimation a main goal motivating the study?</p> <p>motivating_q_drivers - Was identifying drivers of occupancy a main goal motivating the study?</p> <p>motivating_q_trends - Was identifying trends in occupancy a main goal motivating the study?</p> <p>motivating_q_predict - Was predicting occupancy under new conditions a main goal motivating the study?</p> <p>motivating_q_theoretical - Was advancing ecological theory a main goal motivating the study?</p> <p>motivating_q_field_method - Was evaluation of a field method a main goal motivating the study?</p> <p>motivating_q_model_method - Was evaluation of a modeling method a main goal motivating the study?</p> <p>context_conservation - Did the authors contextualize their study as relevant to conservation?</p> <p>context_management - Did the authors contextualize their study as relevant to wildlife management?</p> <p>context_natural_hist - Did the authors contextualize their study as relevant to studying natural history of target species?</p> <p>context_methodology - Did the authors contextualize their study as advancing methodology?</p> <p>detdensity - Did the authors mention the possibility that detection and animal density were confounded?</p> <p>interp_top_mod_as_biol - Did the authors interpret model selection results as evidence for a biological process?</p> <p>nmod_reported_best - Did the authors report only the best model from a model selection workflow?</p> <p>nmod_reported_some - Did the authors report multiple model results from a model selection workflow?</p> <p>nmod_reported_all - Did the authors report all model results from a model selection workflow?</p> <p>priors_reported - Did the authors report their priors in a Bayesian workflow?</p> <p>violation_nonindependence - Do the authors acknowledge violating the assumption of independent data?</p> <p>violation_movement - Do the authors acknowledge violating the assumption of no animal movement?</p> <p>violation_demography - Do the authors acknowledge violating the assumption of no demographic change?</p> <p>violation_det_heterogeneity - Do the authors acknowledge violating the assumption of no unmodeled heterogeneity in detection?</p> <p>violation_yes_other - Do the authors acknowledge violating another assumption?</p> <p>violation_explicit_no - Do the authors state that all assumptions were met?</p> <p>hypotheses_all - Do the authors provide hypotheses for the effect of all covariates?</p> <p>hypotheses_some - Do the authors provide hypotheses for the effect of some covariates?</p> <p>interpret_det_literal - Do the authors interpret detection literally?</p> <p>interpret_det_biol - Do the authors interpret detection as confounded with biology?</p> <p>interpret_vars_significant - Do the authors interpret some variables as significant based on p-values?</p> <p>interpret_vars_credible - Do the authors interpret some variables as meaningful based on Bayesian credible intervals?</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>modsel_results.csv describes the model selection choices made by 64 Phase 2 papers that conducted model selection. In addition to ID, it contains two columns:</p> <p>&nbsp;</p> <p>Submodel approach - Did the authors use a separate-by-submodel approach to model selection, did they only conduct model selection on one submodel, or did they perform variable selection on both submodels simultaneously?</p> <p>Candidate set approach - Did the authors conduct model selection among a set of a priori candidate models, or did they select between arbitrary models based on combinations of all covariates?</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2023View details →
zenodo44/100

Example code and data for ubms: An R package for fitting hierarchical occupancy and N-mixture abundance models in a Bayesian framework

<p>This repository contains an R script (grouse_example.R) and data (grouse_data.csv) used to reproduce the grouse abundance analysis described in Kellner, K. F., et al. (2021) ubms: An R package for fitting hierarchical occupancy and N-mixture abundance models in a Bayesian framework. Methods in Ecology and Evolution. The R script requires installation of the ubms R package, which can be obtained from CRAN (https://cran.r-project.org/package=ubms).</p> <p>The repository also contains an additional example occupancy analysis (occupancy_example.R) using the crossbill dataset included with the unmarked R package.</p>

opencc-by-4.0Oct 2021View details →
zenodo44/100

Occupations on the map: Using a super learner algorithm to downscale labor statistics, data

<p>This repository contains all the input and output data (including maps) related to <a href="https://doi.org/10.1371/journal.pone.0278120">Van Dijk et al. (2022), Occupations on the map: Using a super learner algorithm to downscale labor statistics</a>. It does not contain several large (&gt; 4GB) intermediate files, which summarize the results of the large number of machine learning models that were trained and tuned as part of the super learner algorithm. These&nbsp; files can be created by running the scripts in the supplementary GitHub repository:&nbsp;https://github.com/michielvandijk/occupations_on_the_map.&nbsp;All input and output maps produced as part of &nbsp;this study can also be accessed by means of an interactive web application: https://shiny.wur.nl/occupation-map-vnm.</p> <p>In this paper, we&nbsp;demonstrated an approach to create fine-scale gridded occupation maps by means of downscaling district-level labor statistics informed by remote sensing and other spatial information. We applied a super-learner algorithm that combined the results of different machine learning models to predict the shares of six major occupation categories and the labor force participation rate at a resolution of 30 arc seconds (~1x1 km) in Vietnam. The results were subsequently combined with gridded information on the working-age population to produce maps of the number of workers per occupation. The proposed approach can also be applied to produce maps of other (labor) statistics, which are only available at aggregated levels.</p>

opencc-by-4.0Dec 2021View details →
zenodo44/100

Spatial data sets of the paper Heinrich Stadial 1 continental sand dunes and Middle to Late Holocene paleosol sequences in SE Iberia: implications for human occupation and site formation processes

<p>Spatial data sets of the paper Heinrich Stadial 1 continental sand dunes and Middle to Late Holocene paleosol sequences in SE Iberia: implications for human occupation and site formation processes. This data set is composed by 3 shapefiles:</p> <ol> <li>Dune_field:&nbsp;Feature class polygon shapefile geometry representing the individual dunes identified in the Villena dune field.</li> <li>Sampled dunes:&nbsp;Shapefile of point geometry representing the location of the stratigraphic sequences of CC1, CC2 and CC3 sampled for texture, soil chemistry, OSL and radiocarbon dating.&nbsp;</li> <li>Sediment sourcing samples: Shapefile of point geometry representing the location of the reference samples of El Moron, El Arenal de la Virgen and Sierra del Castellar.&nbsp;</li> </ol> <p>The spatial reference system is EPSG 25830.</p>

opencc-by-4.0Jul 2023View details →
dryad40/100

Data from: Occupancy patterns and upper range limits of lowland Bornean birds along an elevational gradient

<p>Aim: The traditional view of species' distributions is that they are less abundant near the edges of their ranges and more abundant toward the center. Testing this pattern is difficult because of the complexity of distributions across wide geographical areas. An alternative strategy, however, is to measure species' distributional patterns along elevational gradients. We applied this strategy to examine whether lowland forest birds are indeed less common near their upper range limits on a Bornean mountain, and tested co-occurrence patterns among species for potential causes of attenuation, including signatures of habitat selection and competition at the periphery of their ranges.</p> <p>Location: Mt. Mulu, Borneo</p> <p>Taxon: Rain forest birds Methods: We surveyed lowland forest birds on Mt. Mulu (2,376 m), classified their elevation-occupancy distributions using Huisman – Olff – Fresco (HOF) models, and examined co-occurrence patterns of species pairs for signatures of shared habitat patches and interspecific competition.</p> <p>Results: For 39 of 50 common species, occupancy was highest at sea level then gradually declined near their upper range edges, in keeping with a 'rare periphery' hypothesis. With respect to habitat selection, lowland species do not appear to cluster together at sites of patchy similar habitat near their upper range limits; neither are most lowland species segregated from potential montane competitors where ranges overlap.</p> <p>Main conclusions: High relative abundance at sea level implies that species inhabit 'truncated niches' and are not currently near the limits of their fundamental niche, unless unknown critical response thresholds exist. However, indirect effects of increasing temperature predicted under climate change scenarios could still influence lower range limits of lowland species indirectly by altering habitat, precipitation regimes, and competitive interactions. The lack of non-random co-occurrence patterns implies that patchy habitat and simple pairwise species interactions are unlikely to be responsible for upper range limits in most species; diffuse competition across diverse rain forest bird communities could still play a role.</p>

opencc-zeroJul 2020View details →
dryad40/100

Data from: Improving inferences and predictions of species environmental responses with occupancy data

<p>Occupancy models represent a useful tool to estimate species distribution throughout the landscape. Among them, MacKenzie et al.'s model (2002, MC), is frequently used to infer species environmental responses. However, the assumption that detection probability is homogeneous or fully explained by covariates may limit its performance. Species should be more easily observed at sites with a higher number of individuals. We simulated data following Royle and Nichols (2003) occupancy model (RN) that accounts for abundance-driven heterogeneous detection and two variants with overdispersion in the detection probability and local abundances. Then, we compared the performance of the MC model against that of RN.</p> <p> In addition to model misspecifications, insufficient information in data (i.e. infrequent detections) can limit our ability to detect existing effects with affordable sampling designs. To deal with this source of error, we extended RN approach to a community-level joint species model (RN-JSM), where species responses and detectability depended on their traits and phylogeny. Then, we tested RN-JSM performance in simulated and out-of-sample field data.</p> <p>High abundance-driven heterogeneity in detection (i.e. common and secretive species) limited the ability of the MC model to quantify covariate effects; especially, when the number of visits was low. Both models (MC and RN), often failed to detect existing effects when data were overdispersed. Moreover, the RN model consistently lacked sufficient power when analyzing data from uncommon species (even when simulations and model specifications perfectly matched). This problem was solved by our RN-JSM, which yielded more precise and accurate estimates of species environmental responses. Increased accuracy in rare species held when the RN-JSM was tested with real and out-of-sample datasets.</p> <p>In the light of our results, we propose: (i) for common and secretive species analyze occupancy data with the RN model and prioritize revisiting sites; (ii) for species that may have overdispersed detectability or local abundances (e.g. with correlated behaviors or occurring in clusters), apply RN extensions that account for this extra variation (e.g. Poisson-beta or zero-inflated models). Finally, (iii) for uncommon species (mean abundances &lt; 1), whenever possible, gather data at the community level and apply joint-species modeling techniques.</p>

opencc-zeroApr 2022View details →
dryad40/100

Data from: Occupancy winners in tropical protected forests: a pantropical analysis

<p class="MsoNormal"><span>The structure of forest mammal communities appears surprisingly consistent across the continental tropics<span>, presumably due to convergent evolution in similar environments. W</span>hether such consistency extends to mammal occupancy, despite variation in species characteristics and context, remains unclear. Here we ask whether we can predict occupancy patterns and, if so, whether these relationships are consistent across biogeographic regions. Specifically, we assessed how mammal feeding guild, body mass and ecological specialization relate to occupancy in protected forests across the tropics. We used standardized camera-trap data (</span><span>1,002 camera-trap locations and 2-10 years of data)</span><span> and a hierarchical Bayesian occupancy model. We found that occupancy varied by regions, and </span><span>certain species characteristics</span><span> explained much of this variation. Herbivores consistently had the highest occupancy. However, only in the Neotropics did we detect a significant effect of body mass on occupancy: large mammals had lowest occupancy. Importantly, habitat specialists generally had higher occupancy than generalists, though this was reversed in the Indo-Malayan sites. We conclude that </span><span>habitat specialization is key for understanding variation in mammal occupancy across regions, and that habitat specialists often benefit more from protected areas, than do generalists. </span><span>The contrasting examples seen in the Indo-Malayan region likely reflect distinct anthropogenic pressures.</span></p>

opencc-zeroDec 2021View details →
zenodo40/100

Synthetic Indoor Climate and Occupancy Data from Office and Meeting Room Simulations

<p>This is the dataset used for the publication "Coddora: CO2-based Occupancy Detection model<br>trained via DOmain RAndomization". The goal is to provide training data for occupancy detection.<br><br>The dataset contains one million days of data including 10 occupied days for each of 100,000 randomized room models (50,000 rooms considering office activity and 50,000 meeting room activity). Data were generated in EnergyPlus simulations according to the methodology described in the paper.<br><br>When using the dataset, please cite:</p> <blockquote> <p><em>Manuel Weber, Farzan Banihashemi, Davor Stjelja, Peter Mandl, Ruben Mayer, and Hans-Arno Jacobsen. 2024. Coddora: CO2-Based Occupancy Detection Model Trained via Domain Randomization. In International Joint Conference on Neural Networks (IJCNN). June 30 - July 5, 2024, Yokohama, Japan.</em></p> </blockquote> <h2>Dataset Structure</h2> <p>The following files are provided:<br><br>&nbsp; &nbsp; 1. dataset_office_rooms.h5&nbsp; &nbsp;(provided as zip file)<br>&nbsp; &nbsp; 2. dataset_meeting_rooms.h5&nbsp; &nbsp;(provided as zip file)<br>&nbsp; &nbsp; 3. simulated_occupancy_office_rooms.csv<br>&nbsp; &nbsp; 4. simulated_occupancy_meeting_rooms.csv</p> <p>Please use an archiving tool such as 7zip to unzip the hdf5 files.<br>Both hdf5 files contain two datasets with the following keys:<br><br>&nbsp; &nbsp; 1. "<em>data</em>": contains the simulated indoor climate and occupancy data<br>&nbsp; &nbsp; 2. "metadata": contains the metadata that were used for each simulation</p> <p>The csv files contain the time series of occupancy that were used for the simulations.<br><br></p> <h2>Data</h2> <p><em>Data</em> includes the following fields:</p> <p><em>Datetime:</em> day of the year (may be relevant due to seasonal differences) and time of the day<br><em>Zone Air CO2 Concentration:</em> CO2 level in ppm<br><em>Zone Mean Air Temperature:</em> temperature in &deg;C<br><em>Zone Air Relative Humidity: </em>relative humidity in %<br><em>Occupancy: </em>level of occupancy relative to the maximum capacity of the room (in the range [0-1])<br><em>Ventilation:</em> fraction of window opening in the range [0.01, 1]<br><em>SimID:</em> foreign key to reference the room properties the simulation was based on<br><em>BinaryOccupancy:</em> 0 or 1 denoting absence or presence (for binary classification)</p> <p>&nbsp;</p> <p>Example row:</p> <table> <tbody> <tr> <th><em>Datetime</em></th> <th><em>Zone Air CO2 Concentration</em></th> <th><em>Zone Mean Air Temperature</em></th> <th><em>Zone Air Relative Humidity</em></th> <th><em>Occupancy</em></th> <th><em>Ventilation</em></th> <th><em>simID</em></th> <th><em>BinaryOccupancy</em></th> </tr> <tr> <td> <p>10/09 11:21:00</p> </td> <td> <p>1084.5624647371608</p> </td> <td> <p>24.545635909907148</p> </td> <td> <p>41.18393114737054</p> </td> <td> <p>0.7</p> </td> <td> <p>0.0</p> </td> <td>99</td> <td>1</td> </tr> </tbody> </table> <pre>&nbsp;</pre> <h2>Metadata</h2> <p><em>Metadata</em> includes the following fields. <br>Underscores denote that the field was not selected during randomization but calculated from the other values.</p> <p>width: room width in m<br>length: room length in m<br>height: hoom height in m<br>infiltration: &nbsp;infiltration per exterior area in m&sup3;/m&sup2;s<br>outdoor_co2: co2 concentration in the outdoor air in ppm (set to a random value between [300, 500])<br>orientation: angle between the room's facade orientation and the north direction in degrees<br>maxOccupants: room occupation limit, i.e. the maximum number of occupants<br>_floorArea: floor area in m&sup2; (calculated from room dimensions)<br>_volume: room volume in m&sup3; (calculated from room dimensions)<br>_exteriorSurfaceArea: surface area of the facade wall (calculated from room dimensions)<br>_winToFloorRatio: ratio between total window area and floor area (calculated from room model)<br>firstDayUsedOfOccupancySequence: selected starting day in the sequence of occupancy data for rooms with the respective maxOccupants value<br>simID: unique identifier of the simulation to relate between simulation metadata and resulting simulated data</p> <p>&nbsp;</p> <p>Example row:</p> <table> <tbody> <tr> <th>width</th> <th>length</th> <th>height</th> <th>infiltration</th> <th>outdoor_co2</th> <th>orientation</th> <th>maxOccupants</th> <th>_floorArea</th> <th>_volume</th> <th>_exteriorSurfaceArea</th> <th>_winToFloorRatio</th> <th>firstDayOfUsedOccupancySequence</th> <th>simID</th> </tr> <tr> <td>5.481</td> <td>5.190</td> <td>3.264</td> <td>0.000214</td> <td>438.0</td> <td>316.0</td> <td>4.0</td> <td>28.446</td> <td>92.849</td> <td>16.940</td> <td>0.216</td> <td>192</td> <td>0</td> </tr> </tbody> </table> <p>&nbsp;</p> <h2>Occupancy Data</h2> <p>The occupancy data provided through the separate csv files contain the data from the upfront occupancy simulations that the climate simulation was based on. For each level of considered room occupancy limit (maxOccupants), the datasets provide minute values of occupancy throughout 1000 days.</p> <p><em>Datetime, </em><em>Date, </em><em>Timestamp: fictive time of simulated occupancy record (sequences are in 1-minute resolution)</em><br><em>Occupants: number of present occupants</em><br><em>Occupancy: binary occupancy state (0=unoccupied, 1=occupied)</em><br><em>WindowState: binary state of ventilation (0=windows closed, 1=room is ventilated)</em><br><em>maxOccupants: maximum number of occupants considered for the simulated sequence</em><br><em>WindowOpeningFraction: fractional extent to which windows are opened, within the interval [0.01, 1]<br><br></em></p> <p>Example row:</p> <table> <tbody> <tr> <th>Datetime</th> <th>Date</th> <th>Timestamp</th> <th>Occupants</th> <th>Occupancy</th> <th>WindowState</th> <th>maxOccupants</th> <th>WindowOpeningFraction</th> </tr> <tr> <td>2023-01-01 00:00:00</td> <td>2023-01-01</td> <td>1.672531e+09</td> <td>0</td> <td>0</td> <td>0</td> <td>1</td> <td>0.0</td> </tr> </tbody> </table> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Apr 2024View details →
zenodo40/100

France under German occupation. The German and French administration 1940-1945 – full data

<p>At <a href="http://www.adresses-france-occupee.fr">www.adresses-france-occupee.fr</a>, the GHIP provides an interactive map showing the German and French authorities in France during the German occupation of France between 1940 and 1945. In addition to the historical and current address, the website provides information about the tasks, responsibilities and the structure of each department of the authorities, as well as photos if available. Depending on the type and scope of the query, it gives an impression of everyday life, the presence of the German occupation administration and forces and, last but not least, their cooperation with the French authorities. The information is based on a systematic evaluation of the telephone books of the German authorities and French administration directories from the time of the war. They were completed by research in German and French archives, in particular the <em>Archives municipales</em>, as well as press media, historical map collections and research works.</p> <p>With this entry we give access to the data of this database in four different tables (csv UTF-8 and excel).</p> <p>1) &nbsp;&nbsp; The table "services" relates to the departments (German: <em>Dienststellen</em>). The data is maintained in the state closest to that of the phone directories and address books: a department is valid in a hierarchy at a given address and at a given time. The table contains:</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; the ID number (id), its linking to its superior department (parent_service_id)</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; the name of the department (name),</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; the ID number of the source from which the information originates (source_id),</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; the ID number of the place where the departments office is located (place_id),</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; the status on the website (status: visible, pending, deleted),</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; the display of the hierarchy (d_breadcrumb) and</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; the last edit of the entry by the project team (last_edit_date).</p> <p>There are 68,044 entries in total.</p> <p>&nbsp;</p> <p>2) &nbsp;&nbsp; The table "service_bridge" is a cross table, relating the departments to each other by showing their hierarchy. IT contains the parent_service_id, the child_service_id, and the breadcrumb, that indicates the department's position in the hierarchy. There are 28,162 entries in total</p> <p>&nbsp;</p> <p>3) &nbsp;&nbsp; The table "places" is used to localise the departments. A place is a geographical point identified by its latitude and longitude. The address (in text format) of this place may have changed over time. The database contains both old and new names. The table contains the ID number of the place (id), the name of the building/accommodation (name), the status of the entry on the website (status: visible, pending, deleted), the current address (current_country, current_zip, current_city, current_street, current_house_number), former address (fromer_street, former_house_number) as well as the longitude and latitude of the place. There are 15,454 entries in total.</p> <p>&nbsp;</p> <p>4) &nbsp;&nbsp;&nbsp; The table "sources_details" relates to the sources used to build the database. By source we understand &lsquo;data source&rsquo;, i.e. any source used to obtain information on the German and French authorities of the time. We mostly used German and French telephone directories of the years 1940&ndash;1944. The table contains:</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; the ID number of the source (id),</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; the name of the source in short (name),</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; the country of publication or origin of the source (country),</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; the full edition date (edition_date) if available, otherwise approximate,</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; the year of the publication or year of the source (d_edition_year),</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; the month of the publication or month of the source (d_edition_month) and</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; the full bibliographical citation and/or explanations (source_name_complete).</p> <p>There are 57 entries in total.</p> <p>You can get in touch with us via the mail dh [at] dhi-paris.fr</p>

opencc-by-4.0May 2024View details →
zenodo40/100

Occupant Simulation Data based on Honda Accord 2024 Simplified Passenger Model and Full-factorial Sampling with 243 samples and VIRTHUMAN 5, 50, 95 Percentiles

<p>Database with 729 Honda Accord 2014 passenger occupant simulations featuring VIRTHUMAN.&nbsp;</p>

opencc-by-4.0Oct 2024View details →
dryad40/100

Data for: Occupancy–detection models with museum specimen data: Promise and pitfalls

<p>Historical museum records provide potentially useful data for identifying drivers of change in species occupancy. However, because museum records are typically obtained via many collection methods, methodological developments are needed in order to enable robust inferences. Occupancy-detection models, a relatively new and powerful suite of statistical methods, are a potentially promising avenue because they can account for changes in collection effort through space and time.</p> <p>We use simulated datasets to identify how and when patterns in data and/or modelling decisions can bias inference. We focus primarily on the consequences of contrasting methodological approaches for dealing with species' ranges and inferring species' non-detections in both space and time. </p> <p>We find that not all datasets are suitable for occupancy-detection analysis but, under the right conditions (namely, datasets that are broken into more time periods for occupancy inference and that contain a high fraction of community-wide collections, or collection events that focus on communities of organisms), models can accurately estimate trends. Finally, we present a case-study on eastern North American odonates where we calculate long-term trends of occupancy by using our most robust workflow. </p> <p>These results indicate that occupancy-detection models are a suitable framework for some research cases and expand the suite of available tools for macroecological analysis available to researchers, especially where structured datasets are unavailable.</p>

opencc-zeroDec 2021View details →
dryad40/100

Data from: N-dimensional hypervolumes in trait-based ecology: does occupancy rate matter?

<p>Many methods for estimating functional diversity of biological communities rely on measuring geometrical properties of n-dimensional hypervolumes in a trait space. To date, these properties are calculated from individual hypervolumes or from their pairwise combinations. Our capacity to detect functional diversity patterns due to the overlap of multiple hypervolumes is thus limited.</p> <p>Here, we propose a new approach for estimating functional diversity from a set of hypervolumes. We rely on the concept of occupancy rate, defined as the mean or absolute number of hypervolumes enclosing a given point in the trait space. Furthermore, we describe a permutation test to identify regions of the trait space in which the occupancy rate of two sets of hypervolumes differs.</p> <p>We illustrate the utility of our approach over existing methods with two examples on aquatic macroinvertebrates. The first example shows how occupancy rate relates to the stability of trait space utilisation due to increased flow intermittency and allows the identification of taxa in regions of the trait space with low occupancy rates. The second example shows how the permutation test based on occupancy rates can detect differences in trait space utilisation due to river morphology variation even with a high degree of overlap among input hypervolumes.</p> <p>Our newly developed approach is particularly suitable in functional diversity analysis when investigating patterns of overlap among multiple hypervolumes. We thus emphasise the need to consider analyses based on occupancy rate into functional diversity estimation.</p>

opencc-zeroApr 2023View details →
dryad40/100

Data for: Considerations for fitting occupancy models to data from eBird and similar volunteer-collected data

<p>An occupancy model makes use of data that are structured as sets of repeated visits to each of many sites, in order estimate the actual probability of occupancy (i.e., proportion of occupied sites) after correcting for imperfect detection using the information contained in the sets of repeated observations. We explore the conditions under which preexisting, volunteer-collected data from the citizen science project eBird can be used for fitting occupancy models. The data archived here are used to explore two ways in which the single-visit records could be used in occupancy models. First, we use empirical data contained within this archive to assess the potential for space-for-time substitution: aggregating single-visit records from different locations within a region into pseudo-repeat visits. The archived data are used to illustrate that the locations chosen for data collection by observers were not always representative of the habitat in the surrounding area, which would lead to biased estimates of occupancy probabilities when using space-for-time substitution. Second, create a large set of simulated data (output from the simulations contained in this archive) that we used to explore the utility of including data from single-visit records to supplement sets of repeated-visit data.</p>

opencc-zeroJul 2023View details →
zenodo40/100

Loss of socioemotional and occupational roles in people with Long COVID according to sociodemographic and clinical factors: Secondary data from Randomized Clinical Trial.

<p>This is a&nbsp;cross-sectional study was carried out with the participation of 100 patients diagnosed with Long-COVID, over 18 years of age and attended by Primary Health Care in the Autonomous Community of Aragon.&nbsp;The purpose of this study is to analyse the loss of socioemotional and occupational roles that people with Long COVID have suffered in their lives as a consequence of the disease. As a secondary objective, it was proposed to analyze the sociodemographic and clinical factors associated with this loss of roles. The main study variable was the loss of significant socioemotional and occupational roles of the participants. Sociodemographic and clinical data were also collected through a structured interview.</p>

opencc-by-4.0Aug 2023View details →
dryad40/100

Data from: Occupancy patterns and upper range limits of lowland Bornean birds along an elevational gradient

Open the record for dataset details and reuse information.

publicApr 2023View details →
dryad40/100

Data from: Large-scale eDNA sampling and hierarchical modeling elucidates the importance of stream habitat for eastern hellbender (<em>Cryptobranchus a. alleganiensis</em>) occupancy and eDNA detection

Open the record for dataset details and reuse information.

publicSep 2025View details →
dryad40/100

Data from: N-dimensional hypervolumes in trait-based ecology: does occupancy rate matter?

Open the record for dataset details and reuse information.

publicApr 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record