Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,042

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,042 results for “model species”

Learn how ShareScore rates datasets ↗
dryad36/100

Multi-species occupancy modeling provides novel insights into amphibian metacommunity structure and wetland restoration

<p>A fundamental goal of community ecology is to understand species-habitat relationships and how they shape metacommunity structure. Recent advances in occupancy modeling enable habitat relationships to be assessed for both common and rare species within metacommunities using multi-species occupancy models (MSOM). These models account for imperfect species detection and offer considerable advantages over other analytical tools commonly used for community analyses under the elements of metacommunity structure (EMS) framework. Here, we demonstrate that MSOM can be used to infer habitat relationships and test metacommunity theory, using amphibians. Repeated frog surveys were undertaken at 55 wetland sites in eastern Australia. We detected 11 frog species from three families (Limnodynastidae, Myobatrachidae and Pelodryadidae). The rarest species was detected at only one site whereas the most common species was detected at 42 sites (naïve occupancy rate: 0.02 – 0.76). Two models were assessed representing two competing hypotheses; the best-supported model included the covariates distance to the nearest site (connectivity), wetland area, presence of the non-native eastern mosquitofish (<i>Gambusia holbrooki</i>), proportion cover of emergent vegetation, an interaction term between Gambusia and emergent vegetation cover, and the proportion canopy cover over a site. Hydroperiod played no detectable role in metacommunity structure. We found species-habitat relationships that fit with current metacommunity theory – occupancy increased with wetland area and connectivity. There was a strong negative relationship between occupancy and the presence of predatory Gambusia, and a positive interaction between Gambusia and emergent vegetation. The presence of canopy cover strongly increased occupancy for several tree frog species, highlighting the importance of terrestrial habitat for amphibian community structure. We demonstrated how responses by amphibians to environmental covariates at the species level can be linked to occupancy patterns at the metacommunity scale. Our results have clear management implications – wetland restoration projects for amphibians and likely other taxa should maximize wetland area and connectivity, establish partial canopy cover, and eradicate Gambusia or provide aquatic vegetation to mitigate the impact of this non-native fish. We strongly advocate the use of MSOM to elucidate the habitat drivers behind animal occupancy patterns and to derive unbiased occupancy estimates for monitoring programmes.</p>

opencc-zeroOct 2020View details →
zenodo36/100

Data and Code for: Comparative species abundance modeling of Capitellidae (Annelida) in Tampa Bay, Florida, USA

<p>This publication contains the R-Scripts, RData, data frames, and other supplementary files necessary to reproduce the analyses and figures of&nbsp;Hilliard J, Karlen D, Dix T, Markham S, Schulze A (2020) Comparative species abundance modeling of Capitellidae (Annelida) in Tampa Bay, Florida, USA. Mar Ecol Prog Ser 653:105-119.&nbsp;<a href="https://doi.org/10.3354/meps13484">https://doi.org/10.3354/meps13484</a></p> <p>All files are compressed into a single ZIP folder. The main directory contains a README file that explains the contents of the subdirectories.</p>

opencc-by-4.0Nov 2020View details →
dryad36/100

Data from: Climate limitation at the cold edge – contrasting perspectives from species distribution modelling and a transplant experiment

<p>The role of climate in determining range margins is often studied using species distribution models (SDMs), which are easily applied but have well-known limitations, e.g. due to their correlative nature and colonization and extinction time lags. Transplant experiments can give more direct information on environmental effects, but often cover small spatial and temporal scales. We simultaneously applied an SDM using high-resolution spatial predictors and an integral projection (demographic) model based on a transplant experiment at 58 sites to examine the effects of microclimate, light and soil conditions on the distribution and performance of a forest herb, <i>Lathyrus vernus</i>, at its cold range margin in central Sweden. In the SDM, occurrences were strongly associated with warmer climates. In contrast, only weak effects of climate were detected in the transplant experiment, whereas effects of soil conditions and light dominated. The higher contribution of climate in the SDM is likely a result from its correlation with soil quality, forest type, and potentially historic land use, which were unaccounted for in the model. Predicted habitat suitability and population growth rate, yielded by the two approaches, were not correlated across the transplant sites. We argue that the ranking of site habitat suitability is probably more reliable in the transplant experiment than in the SDM because predictors in the former better describe understory conditions, but that ranking might vary among years, e.g. due to differences in climate. Our results suggest that <i>L. vernus</i> is limited by soil and light rather than directly by climate at its northern range edge, where conifers dominate forests and create suboptimal conditions of soil and canopy-penetrating light. A general implication of our study is that to better understand how climate change influences range dynamics, we should not only strive to improve existing approaches but also to use multiple approaches in concert.</p>

opencc-zeroDec 2020View details →
dryad36/100

Simulation models from: Can CRISPR-mediated gene drive work in pest and beneficial haplodiploid species?

<p>Gene drives based on CRISPR/Cas9 have the potential to reduce the enormous harm inflicted by crop pests and insect vectors of human disease, as well as to bolster valued species. In contrast with extensive empirical and theoretical studies in diploid organisms, little is known about CRISPR gene drive in haplodiploids, despite their immense global impacts as pollinators, pests, natural enemies of pests, and invasive species in native habitats. Here we analyze mathematical models demonstrating that, in principle, CRISPR homing gene drive can work in haplodiploids, as well as at sex-linked loci in diploids. However, relative to diploids, conditions favoring the spread of alleles deleterious to haplodiploid pests by CRISPR gene drive are narrower, the spread is slower, and resistance to the drive evolves faster. By contrast, the spread of alleles that impose little fitness cost or boost fitness was not greatly hindered in haplodiploids relative to diploids. Therefore, altering traits to minimize damage caused by harmful haplodiploids, such as interfering with transmission of plant pathogens, may be more likely to succeed than control efforts based on introducing traits that reduce pest fitness. Enhancing fitness of beneficial haplodiploids with CRISPR gene drive is also promising.</p>

opencc-zeroMay 2020View details →
dryad36/100

Hierarchical multi-grain models improve descriptions of species' environmental associations, distribution, and abundance

<p>The characterization of species' environmental niches and spatial distribution predictions based on them are now central to much of ecology and conservation, but implicitly requires decisions about the appropriate spatial scale (i.e. <i>grain</i>) of analysis. Ecological theory and empirical evidence suggest that range-resident species respond to their environment at two characteristic, hierarchical spatial grains: (i) <i>response grain</i>, the (relatively fine) grain at which an individual uses environmental resources, and (ii) <i>occupancy grain</i>,<i> </i>the (relatively coarse) grain equivalent to a typical home range. We use a multi-grain (MG) occupancy model, aided by fine-grain remotely sensed imagery, to simultaneously estimate species-environment associations at both grains, conduct grain optimization to measure response grain, and apply this analysis framework to an example species: a medium-sized bird (<i>Tockus deckeni</i>) in a heterogeneous East African landscape. Based on home range analysis of movement data, we calculate an occupancy grain of 1km for <i>T. deckeni</i>. Using a grain optimization procedure across 32 grains from 10m to 500m, we identify 60m as the most strongly supported response grain for a suite of environmental variables, slightly coarser than opportunistic behavioral observations would have suggested. Validation confirms that the accuracy of the optimized MG occupancy model substantially exceeds that of equivalent single-grain (SG) occupancy models. We further use a simulation approach to assess the potential impacts of accounting for the multi-scale structure of species' environmental requirements on estimates of population size. We find that the more strongly supported MG approach consistently predicts a minimum population sizes in the study landscape that is much lower than that provided by the SG model. This suggests that SG approaches commonly used in conservation applications could lead to overly optimistic abundance and population estimates and that the MG approach may be more appropriate for supporting species conservation goals. More generally, we conclude that multi-grain approaches of the sort presented, and increasingly enabled by growing high-resolution remotely sensed data, hold great promise for offering a more mechanistic framework for assessing the appropriate grain(s) for population monitoring and management and enable more reliable estimates of abundances and species' distributions.</p>

opencc-zeroJan 2020View details →
dryad36/100

Data from: A new null model approach to quantify performance and significance for ecological niche models of species distributions

Aim: Ecological niche modelling requires robust estimation of model performance and significance, but common evaluation approaches often yield biased estimates. Null models provide a solution but are rarely used in this field. We implemented an important modification to existing null-model tests, evaluating null models with the same withheld records that were used to evaluate the real model. We built and evaluated models across a range of modelling scenarios and for various performance measures using the algorithm Maxent and the monk parakeet (Myiopsitta monachus). Location: Native range in Southern America and global invasions predominantly in North/Central America and Europe Methods: We tested the ability of models built under 15 scenarios (five sets of calibration records and three settings that varied the level of model complexity) to predict spatially independent evaluation data in the invaded range (in effect, testing the models under spatial transfer). We quantified performance with measures of discriminatory ability and overfitting based on AUC and the omission error rate. We estimated null distributions of these measures and calculated effect size and significance. We determined how these estimates varied across modelling scenarios, comparing with two tests existing in the literature. Results: Performance varied starkly across modelling scenarios. As expected, the measures of overfitting agreed with each other and provided different information than that of discriminatory ability. However, high performance per se did not show strong association with high effect size and significance. Main Conclusions: Ecological niche models should be assessed with measures of effect size and significance based on appropriate null distributions, in contrast to several approaches existing in the literature. The proposed approach using independent evaluation data, implemented with our accompanying code, allows such estimates for either the same or a different region/time period, and it merits use and continued development.

opencc-zeroDec 2018View details →
dryad36/100

Data from: Effectiveness of joint species distribution models in the presence of imperfect detection

<p>Joint species distribution models (JSDMs) are a recent development in biogeography and enable the spatial modelling of multiple species and their interactions and dependencies. However, most models do not consider imperfect detection, which can significantly bias estimates. This is one of the first papers to account for imperfect detection when fitting data with JSDMs and to explore the complications that may arise.</p> <p>A multivariate probit JSDM that explicitly accounts for imperfect detection is proposed, and implemented using a Bayesian hierarchical approach. We investigate the performance of the JSDM in the presence of imperfect detection for a range of factors, including varied levels of detection and species occupancy, and varied numbers of survey sites and replications. To understand how effective this JSDM is in practice, we also compare results to those from a JSDM that does not explicitly model detection but instead makes use of  "collapsed data". A case study of owls and gliders in Victoria Australia is also illustrated.</p> <p>Using simulations, we found that the JSDMs explicitly accounting for detection can accurately estimate intrinsic correlation between species with enough survey sites and replications. Reducing the number of survey sites decreases the precision of estimates, while reducing the number of survey replications can lead to biased estimates. For low probabilities of detection, the model may require a large number of survey replications to remove bias from estimates. However, JSDMs not explicitly accounting for detection may have a limited ability to disentangle detection from occupancy, which substantially reduces their ability to accurately infer the species distribution spatially. Our case study showed positive correlation between Sooty Owls and Greater Gliders, despite a low number of survey replications.</p> <p>To avoid biased estimates of inter-species correlations and species distributions, imperfect detection needs to be considered. However, for low probability of detection, the JSDMs explicitly accounting for detection is data hungry. Estimates from such models may still be subject to bias. To overcome the bias, researchers need to carefully design surveys and choose appropriate modelling approaches. The survey design should ensure sufficient survey replications for unbiased inferences on species inter-dependencies and occupancy.</p>

opencc-zeroJun 2021View details →
dryad36/100

Data from: Effects of grain size and niche breadth on species distribution modeling

Scale is a vital component to consider in ecological research, and spatial resolution or grain size is one of its key facets. Species distribution models (SDMs) are prime examples of ecological research in which grain size is an important component. Despite this, SDMs rarely explicitly examine the effects of varying the grain size of the predictors for species with different niche breadths. To investigate the effect of grain size and niche breadth on SDMs, we simulated four virtual species with different grain sizes/niche breadths using three environmental predictors (elevation, aspect, and percent forest) across two real landscapes of differing heterogeneity in predictor values. We aggregated these predictors to seven different grain sizes and modeled the distribution of each of our simulated species using MaxEnt and GLM techniques at each grain size. We examined model accuracy using the AUC statistic, Pearson's correlations of predicted suitability with the true suitability, and the binary area of presence determined from suitability above the maximum True Skill Statistic (TSS) threshold. Habitat specialists were more accurately modeled than generalist species, and the models constructed at the grain size from which a species was derived generally performed the best. The accuracy of models in the homogenous landscape deteriorated with increasing grain size to a greater degree than models in the heterogenous landscape. Variable effects on the model varied with grain size, with elevation increasing in importance as grain size increased while aspect lost importance. The area of predicted presence was drastically affected by grain size, with larger grain sizes over predicting this value by up to a factor of 14. Our results have implications for species distribution modeling and conservation planning, and we suggest more studies include analysis of grain size as part of their protocol.

opencc-zeroDec 2016View details →
dryad36/100

Accounting for imperfect detection in data from museums and herbaria when modeling species distributions: Combining and contrasting data-level versus model-level bias correction

The digitization of museum collections as well as an explosion in citizen science initiatives has resulted in a wealth of data that can be useful for understanding the global distribution of biodiversity, provided that the well-documented biases inherent in unstructured opportunistic data are accounted for. While traditionally used to model imperfect detection using structured data from systematic surveys of wildlife, occupancy models provide a framework for modelling the imperfect collection process that results in digital specimen data. In this study, we explore methods for adapting occupancy models for use with biased opportunistic occurrence data from museum specimens and citizen science platforms using 7 species of Anacardiaceae in Florida as a case study. We explored two methods of incorporating information about collection effort to inform our uncertainty around species presence: (1) filtering the data to exclude collectors unlikely to collect the focal species and (2) incorporating collection covariates (collection type, time of collection, and history of previous detections) into a model of collection probability. We found that the best models incorporated both the background data filtration step as well as collector covariates. Month, method of collection and whether a collector had previously collected the focal species were important predictors of collection probability. Efforts to standardize meta-data associated with data collection will improve efforts for modeling the spatial distribution of a variety of species.

opencc-zeroJun 2021View details →
zenodo36/100

Data and scripts for "Joint species-trait distribution modelling: The role of intraspecific trait variation in community assembly"

<p>The README explains how to reproduce the analyses presented in the paper <strong>"Joint species-trait distribution modelling: The role of intraspecific trait variation in community assembly"</strong> by Abrego et al.</p> <p>The input data for the script pipeline is the file &ldquo;Kilpisjarvi_plant_data.csv&rdquo;. This file includes the data on the plants and their traits in the long format. Hence, each row of the data matrix corresponds to measurements on one plant species in one study plot. The joint species-trait distribution modelling (JSTDM) pipeline that analyses these data consists of the following R-scripts.</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; <strong>S1_define_JSTDM_models.R</strong>. This script defines the JSDTM models (null model and environmental model) that include five response types for each species: the presence-absence, abundance conditional on presence, and the plot-level trait values of specific leaf area (SLA), leaf area (LA) and mean height (MH). The model is defined in the Hierarchical Modelling of Species Communities (HMSC) framework utilizing the R-package Hmsc. The models are saved in the file &ldquo;unfitted_models.RData&rdquo;.</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; <strong>S2_fit_models.R. </strong>This script loads the unfitted models and fits them using the posterior sampling methods implemented in the R-package Hmsc. The models are fitted with increasing thinning until thin=100, which value was used to generate the results of the paper. The fitted models are saved in the file "models_thin_100_samples_250_chains_4.Rdata".</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; <strong>S3_plot_Omega_matrices.R. </strong>This script loads the fitted models and plots the association matrices (Fig. 2 of the paper). The csv file containing the values used to construct Fig. 2 is also given (figure2Cdata.csv and figure2Ddata.csv).</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; <strong>S4_show_VP_Beta_Gamma.R. </strong>This script loads the fitted models and extracts information on the variance partitionings (VP; Figs. S3 and S4 of the paper), the relationships between response types and environmental predictors (beta; Fig. S2 of the paper), and the relationships between response types and species-level traits (gamma; Fig. S5 of the paper). The csv file containing the values used to construct Fig. S2 (figureS2Adata.csv and figureS2Bdata.csv), Fig. S3 (figureS3data.csv), Fig. S4 (figureS4data.csv) and Fig S5 (figureS5data.csv) are also given.</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; <strong>S5_conditional_cross_validation.R</strong>. This script performs 10-fold cross validation to the data to test the predictive power related to the modelled plant traits. The script performs both regular (unconditional) cross-validation where all data are masked for the test fold, and conditional cross-validation where only the trait data (but not the abundance data) are masked for the test fold.</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; <strong>S6_show_conditional_cross_validation_results.R. </strong>This script plots the results of cross-validation (Fig. 3 of the paper). The csv file containing the values used to construct Fig. 3 is also given (figure3Adata.csv and figure3Bdata.csv).</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; <strong>S7_scenario_predictions.R</strong>. This script performs the scenario simulations described and shown in Fig. 4 of the paper. The csv file containing the values used to construct Fig. 4 is also given (figure4Bdata.csv and figure4Cdata.csv).</p>

opencc-by-4.0May 2024View details →
dryad36/100

Data from: How far can I extrapolate my species distribution model? Exploring Shape, a novel method

<p>Species distribution and ecological niche models (hereafter SDMs) are popular tools with broad applications in ecology, biodiversity conservation, and environmental science. Many SDM applications require projecting models in environmental conditions non-analog to those used for model training (extrapolation), giving predictions that may be statistically unsupported and biologically meaningless. We introduce a novel method, Shape, a model-agnostic approach that calculates the extrapolation degree for a given projection data point by its multivariate distance to the nearest training data point. Such distances are relativized by a factor that reflects the dispersion of the training data in environmental space. Distinct from other approaches, Shape incorporates an adjustable threshold to control the binary discrimination between acceptable and unacceptable extrapolation degrees. We compared Shape's performance to five extrapolation metrics based on their ability to detect analog environmental conditions in environmental space and improve SDMs suitability predictions. To do so, we used 760 virtual species to define different modeling conditions determined by species niche tolerance, distribution equilibrium condition, sample size, and algorithm. All algorithms had trouble predicting species niches. However, we found a substantial improvement in model predictions when model projections were truncated independently of extrapolation metrics. Shape's performance was dependent on extrapolation threshold used to truncate models. Because of this versatility, our approach showed similar or better performance than the previous approaches and could better deal with all modeling conditions and algorithms. Our extrapolation metric is simple to interpret, captures the complex shapes of the data in environmental space, and can use any extrapolation threshold to define whether model predictions are retained based on the extrapolation degrees. These properties make this approach more broadly applicable than existing methods for creating and applying SDMs. We hope this method and accompanying tools support modelers to explore, detect, and reduce extrapolation errors to achieve more reliable models.</p>

opencc-zeroOct 2023View details →
zenodo36/100

Figure 4 in Conservation gaps identification through patterns of species richness established from species niche models of mammals in a sector of Chaco Seco ecoregion

Figure 4. Response graphs of habitat suitability (ordinate axis) according to the explanatory variables that intervened in the adjustment of the model for brown brocket deer (A, B, C). The temperature is expressed in degrees Celsius.Source of bioclimatic variables (bio), site https://www.worldclim.org/data/bioclim.html.

opencc-by-nc-4.0Oct 2023View details →
zenodo36/100

Figure 8 in Conservation gaps identification through patterns of species richness established from species niche models of mammals in a sector of Chaco Seco ecoregion

Figure 8. Species richness maps obtained using three algorithms, (A) "fuzzy union″, (B) "species richness″ and (C) "total beta″.

opencc-by-nc-4.0Oct 2023View details →
zenodo36/100

Figure 7 in Conservation gaps identification through patterns of species richness established from species niche models of mammals in a sector of Chaco Seco ecoregion

Figure 7. Response graphs of habitat suitability (ordinate axis) according to the explanatory variables that intervened in theadjustment of themodel for collared peccary (A,B, C).Temperature is expressed in degrees Celsius and altitude in meters.Source of bioclimatic variables (bio), site https://www.worldclim.org/data/bioclim.html.

opencc-by-nc-4.0Oct 2023View details →
dryad36/100

Data for: The meta-analysis of the effects of spatial sampling bias correction on presence only species distribution models

<p>This dataset contains information extracted from 70 studies identified through a systematic review of the peer-reviewed literature (Web of Science and SCOPUS databases both searched on the 13/02/2023) to evaluate the effect of spatial sampling bias correction methods in presence-only species distribution models.</p>

opencc-zeroDec 2023View details →
zenodo36/100

The global distribution of plants used by humans datasets: list of utilised species, occurrence data and model outputs at 10 arc-minutes spatial resolution

<p>Datasets and model outputs used to map the global distribution of utilised plants by humans. The folder is composed of two subfolders <em>raw_data</em> and <em>processed_data</em> containing respectively the list of utilised plant species modelled -<em>utilised_plants_species_list.csv</em>-, and their occurrence data -<em>occurrence_data.zip-</em> and predicted distribution -<em>species_proba_per_cell.rds-.</em></p> <p>&nbsp;</p> <ul> <li>The file <em>utilised_plants_species_list.csv</em> in the <em>raw_data</em> folder contains a<strong> </strong>list of 35687 plant species (and hybrids) used by humans and 10 plant use categories with the following 14 fields:</li> </ul> <p><strong>plant_ID:<em> </em></strong>plant identifier number ranging from between 1-35687</p> <p><strong>binomial_acc_name:</strong> binomial accepted name of the plant species</p> <p><strong>author_acc_name</strong>: &nbsp;name of the author(s)</p> <p><strong>is_hybrid:</strong> logical TRUE or FALSE indicating whether the species is an hybrid or not.</p> <p><strong>AnimalFood:</strong> forage and fodder for vertebrate animals only.</p> <p><strong>EnvironmentalUses:</strong> examples include intercrops and nurse crops, ornamentals, barrier hedges, shade plants, windbreaks, soil improvers, plants for revegetation and erosion control, wastewater purifiers, indicators of the presence of metals, pollution, or underground water.</p> <p><strong>Fuels:</strong> charcoal, petroleum substitutes, fuel alcohols, etc. Given the importance of energy plants for people, those were distinguished from Materials.</p> <p><strong>GeneSources:</strong> wild relatives of major crops which may possess traits associated with biotic or abiotic resistance and may be valuable for breeding programs.</p> <p><strong>HumanFood:</strong> food for humans only, including beverages and food additives.</p> <p><strong>InvertebrateFood:</strong> plants consumed by invertebrates used by humans, such as bees, silkworms, lac insects and edible grubs.</p> <p><strong>Materials:</strong> woods, fibers, cork, cane, tannins, latex, resins, gums, waxes, oils, lipids, etc. and their derived products.</p> <p><strong>Medicines:</strong> both human and veterinary.</p> <p><strong>Poisons:</strong> plants which are poisonous to both vertebrates and invertebrates, both accidentally and intentionally, e.g., for hunting and fishing, molluscicides, herbicides, insecticides.</p> <p><strong>SocialsUses:</strong> plants used for social purposes, which cannot be defined as food or medicine, for instance, masticatories, smoking materials, narcotics, hallucinogens and psychoactive drugs, and plants with ritual or religious significance.</p> <p><strong>Totals:</strong> total number of uses recorded for a species</p> <p>&nbsp;</p> <ul> <li>The zipfile <em>occurrence_data.zip</em> in the <em>processed_data</em> folder contains 35687 Comma Separated Values (CSV) files, one for each species, containing curated geographic occurrence records used to &nbsp;build species distribution models with the following 14 fields:</li> </ul> <p><strong>Species:</strong> the binomial accepted name of the species</p> <p><strong>Fullname:</strong> &nbsp;same as species</p> <p><strong>decimalLongitude:</strong> the geographic longitude of the occurrence records of the species in decimal degrees</p> <p><strong>decimalLatitude:</strong> the geographic latitude of the occurrence records of the species in decimal degrees</p> <p><strong>countryCode:</strong> a three-letter standard abbreviation for the country of the occurrence locality</p> <p><strong>coordinateUncertaintyinMeters</strong>: indicator for the accuracy of the coordinate location, described as the radius of a circle around the stated point location</p> <p><strong>year:</strong> year of the observation of the occurrence record of the species</p> <p><strong>individualCount:</strong> the number of individuals present at the time of the observation</p> <p><strong>gbifID:</strong> unique identifier number for the occurrence from the original database</p> <p><strong>basisOfRecords:</strong> the type of the individual record, e.g. observation, physical specimen, fossil, living ex-situ, culture collection specimen</p> <p><strong>institutionCode</strong>: the name of the institution or organization listed as the data publisher on GBIF</p> <p><strong>establishmentMeans:</strong> statement about whether an organism has been introduced to a given place and time through the direct or indirect activity of modern humans</p> <p><strong>is_cultivated_observation:</strong> whether or not an organism is cultivated</p> <p><strong>sourceID:</strong> name of the source database</p> <p>&nbsp;</p> <ul> <li>The file <em>species_proba_per_cell.rds</em> in the <em>processed_data</em> folder is<em> a R Data Serialization </em>(RDS) file containing a data.table object with the following 3 fields:</li> </ul> <p><strong>plant_ID:</strong><em> </em>plant identifier number ranging from between 1-35687</p> <p><strong>proba:</strong> species occurrence probability</p> <p><strong>cell:</strong><em> </em>raster grid cell number between 1-2251762</p> <p>This object can be used in combination with a raster layer to reconstruct the modelled distribution of each species or retrieve species richness and endemism.</p>

opencc-by-4.0Dec 2022View details →
zenodo36/100

Evaluation data for: Adaptive sampling by citizen scientists improves species distribution model performance: a simulation study

<p>All of the evaluation data for the simulations in the paper: Adaptive sampling by citizen scientists improves species distribution model performance: a simulation study. We considered the impact of five adaptive sampling methods on the performance of species distribution models (SDMs), please see the paper for more information. Contained in this repository are the evaluation metrics (AUC, mean square error (MSE) and correlation) for SDMs before and after adaptive sampling has taken place. The MSE and correlation evaluation metrics were calculated against the true distributions of the species. These files are those with "combined_outputs" in the titles. The repository also contains the observations of all the species in the simulations both before and after adaptive sampling (the files with "all_observations" in the title.</p> <p>These datasets are to be used with the plotting and evaluation scripts in the GitHub repository associated with the paper.</p>

opencc-by-4.0Feb 2024View details →
dryad36/100

Data for: A new threshold selection method for species distribution models with presence-only data: extracting the mutation point of the P/E curve by threshold regression

<p>Selecting thresholds to convert continuous predictions of species distribution models proves critical for many real-world applications and model assessments. Prevalent threshold selection methods for presence-only data require unproven pseudo-absence data or subjective researchers' decisions. This study proposes a new method, Boyce-Threshold Quantile Regression (BTQR), to determine thresholds objectively without pseudo-absence data. We summarize that the mutation point is a typical shape feature of the predicted-to-expected (P/E) curve after reviewing relevant articles. Analysis based on source-sink theory suggests that this mutation point may represent a transition in habitat types and serve as an appropriate threshold. Threshold regression is introduced to accurately locate the mutation point.</p> <p>To validate the effectiveness of BTQR, we used four virtual species of varying prevalence and a real species with reliable distribution data. Six different species distribution models were employed to generate continuous suitability predictions. BTQR and nine other traditional methods transformed these continuous outputs into binary results. Comparative experiments show that BTQR has advantages in terms of accuracy, applicability, and consistency over the existing methods.</p>

opencc-zeroMar 2024View details →
dryad36/100

Incorporating plant phenological responses into species distribution models (SDMs) reduces estimates of future species loss and turnover

<p>Anthropogenetic climate change has caused distribution shifts of many species, and species distribution models (SDMs) are central for documenting this relationship. However, most SDMs rarely consider the evolution of climate-sensitive functional traits, such as phenology, which strongly affect species fitness. Using &gt;120,000 herbarium specimens representing 360 plant species across the eastern United States, we developed a novel "phenology-informed" SDM that integrates dynamic phenological responses to changing climates. Compared to standard SDMs, our phenology-informed SDMs forecast lower species habitat loss and less species turnover under climate change. These results suggest that phenotypic plasticity or local adaptation in phenology may help species adjust their ecological niches and persist in their habitats under rapid environmental change. Our findings reveal how phenology variation mediates species distributions and affects regional biodiversity patterns. Our newly developed model also circumvents the need for mechanistic models, facilitating the deployment of trait-based SDMs across unprecedented spatial and taxonomic scales.</p>

opencc-zeroMar 2024View details →
dryad36/100

Evaluating the predictors of habitat use and successful reproduction in a model bird species using a large scale automated acoustic array

<p>The emergence of continental to global scale biodiversity data has led to growing understanding of patterns in species distributions, and the determinants of these distributions, at large spatial scales. However, identifying the specific mechanisms, including demographic processes, and determining species distributions remains difficult, as large-scale data are typically restricted to observations of only species presence. New remote automated approaches for collecting data, such as automated recording units (ARUs), provide a promising avenue towards direct measurement of demographic processes, such as reproduction, that cannot feasibly be measured at scale by traditional survey methods. In this study, we analyze data collected by ARUs from 452 survey points across an approximately 1500 km study region to compare patterns in adult and juvenile distributions in the Great Horned Owl (<em>Bubo virginianus</em>). We specifically examine whether habitat associated with successful reproduction is the same as that associated with adult presence. We postulated that congruence between these two distributions would suggest that all areas of the species' range contribute equally to maintenance of the population, whereas significant differences would suggest more specificity in the species' requirements for successful reproduction. We filtered adult and juvenile calls of the species for manual review using automated classification and constructed single season occupancy models to compare land cover and vegetation covariates which significantly predicted presence of each life stage. We found that habitat use by adults was significantly predicted by increasing amounts of forest cover, reduced forest basal area, and lower elevations whereas juvenile presence was significantly predicted only by decreasing amounts of forest cover, a pattern opposite that of adults. These results show that presence of adult Great Horned Owls is not a sufficient proxy for locations at which reproduction occurs, and also demonstrate a highly scalable workflow that could be used for similar analyses in other sound-producing species.</p>

opencc-zeroApr 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record