Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

487

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

487 results for “species distribution modeling”

Learn how ShareScore rates datasets ↗
dryad40/100

Assessing patterns and risk to Chilean freshwater fish distributions using multi-species occupancy models

Open the record for dataset details and reuse information.

publicAug 2025View details →
dryad40/100

Data from: Complementary strengths of spatially-explicit and multi-species distribution models

Open the record for dataset details and reuse information.

publicDec 2019View details →
dryad40/100

The past, present, and future of predator-prey interactions in a warming world: using species distribution modeling to forecast ectotherm-endotherm niche overlap

Open the record for dataset details and reuse information.

publicAug 2024View details →
dryad40/100

Data and code for: Building use-inspired species distribution models: using multiple data types to examine and improve model performance

Open the record for dataset details and reuse information.

publicMay 2023View details →
dryad40/100

Spatial confounding in Bayesian species distribution modeling

Open the record for dataset details and reuse information.

publicAug 2022View details →
dryad40/100

Improving distribution models of sparsely-documented disease vectors by incorporating information on related species via joint modeling

Open the record for dataset details and reuse information.

publicApr 2024View details →
dryad40/100

Data from: Integrated species distribution models to account for sampling biases and improve range wide occurrence predictions

Open the record for dataset details and reuse information.

publicNov 2023View details →
dryad40/100

Resources for: Spatio-temporal integrated Bayesian species distribution models reveal lack of broad relationships between traits and range shifts

Open the record for dataset details and reuse information.

publicMar 2024View details →
dryad40/100

Data from: Species distribution models of the Spotted Wing Drosophila (Drosophila suzukii, Diptera: Drosophilidae) in its native and invasive range reveal an ecological niche shift

Open the record for dataset details and reuse information.

publicOct 2018View details →
dryad40/100

Data from: Integrating genomic data and simulations to evaluate alternative species distribution models and improve predictions of glacial refugia and future responses to climate change

Open the record for dataset details and reuse information.

publicJun 2024View details →
dryad40/100

Data from: Integrated SDM database: Enhancing the relevance and utility of species distribution models in conservation management

Open the record for dataset details and reuse information.

publicOct 2021View details →
dryad40/100

Code and data for Bayesian joint species distribution model selection for community-level prediction

Open the record for dataset details and reuse information.

publicNov 2023View details →
dryad36/100

Disentangling drivers of spatial autocorrelation in species distribution models

<p>Species distribution models (SDMs) are frequently used to understand the influence of site properties on species occurrence. For robust model inference, SDMs need to account for the spatial autocorrelation of virtually all species occurrence data. Current methods do not routinely distinguish between extrinsic and intrinsic drivers of spatial autocorrelation, although these may have different implications for conservation. Here, we present and test a method that disentangles extrinsic and intrinsic drivers of spatial autocorrelation using repeated observations of a species. We focus on unknown habitat characteristics and conspecific interactions as extrinsic and intrinsic drivers, respectively. We model the former with spatially correlated random effects and the latter with an autocovariate, such that the spatially correlated random effects are constant across the repeated observations whereas the autocovariate may change. We tested the performance of our model on virtual species data and applied it to observations of the corncrake Crex crex in the Netherlands. Applying our model to virtual species data revealed that it was well able to distinguish between the two different drivers of spatial autocorrelation, outperforming models with no or a single component for spatial autocorrelation. This finding was independent of the direction of the conspecific interactions (i.e., conspecific attraction versus competitive exclusion). The simulations confirmed that the ability of our model to disentangle both drivers of autocorrelation depends on repeated observations. In the case study, we discovered that the corncrake has a stronger response to habitat characteristics compared to a model that did not include spatially correlated random effects, whereas conspecific interactions appeared to be less important. This implies that future conservation efforts should primarily focus on maximizing habitat availability. Our study shows how to systematically disentangle extrinsic and intrinsic drivers of spatial autocorrelation. The method we propose can help to correctly identify the main drivers of species distributions.</p>

opencc-zeroAug 2020View details →
dryad36/100

Data from: Climate limitation at the cold edge – contrasting perspectives from species distribution modelling and a transplant experiment

<p>The role of climate in determining range margins is often studied using species distribution models (SDMs), which are easily applied but have well-known limitations, e.g. due to their correlative nature and colonization and extinction time lags. Transplant experiments can give more direct information on environmental effects, but often cover small spatial and temporal scales. We simultaneously applied an SDM using high-resolution spatial predictors and an integral projection (demographic) model based on a transplant experiment at 58 sites to examine the effects of microclimate, light and soil conditions on the distribution and performance of a forest herb, <i>Lathyrus vernus</i>, at its cold range margin in central Sweden. In the SDM, occurrences were strongly associated with warmer climates. In contrast, only weak effects of climate were detected in the transplant experiment, whereas effects of soil conditions and light dominated. The higher contribution of climate in the SDM is likely a result from its correlation with soil quality, forest type, and potentially historic land use, which were unaccounted for in the model. Predicted habitat suitability and population growth rate, yielded by the two approaches, were not correlated across the transplant sites. We argue that the ranking of site habitat suitability is probably more reliable in the transplant experiment than in the SDM because predictors in the former better describe understory conditions, but that ranking might vary among years, e.g. due to differences in climate. Our results suggest that <i>L. vernus</i> is limited by soil and light rather than directly by climate at its northern range edge, where conifers dominate forests and create suboptimal conditions of soil and canopy-penetrating light. A general implication of our study is that to better understand how climate change influences range dynamics, we should not only strive to improve existing approaches but also to use multiple approaches in concert.</p>

opencc-zeroDec 2020View details →
dryad36/100

Hierarchical multi-grain models improve descriptions of species' environmental associations, distribution, and abundance

<p>The characterization of species' environmental niches and spatial distribution predictions based on them are now central to much of ecology and conservation, but implicitly requires decisions about the appropriate spatial scale (i.e. <i>grain</i>) of analysis. Ecological theory and empirical evidence suggest that range-resident species respond to their environment at two characteristic, hierarchical spatial grains: (i) <i>response grain</i>, the (relatively fine) grain at which an individual uses environmental resources, and (ii) <i>occupancy grain</i>,<i> </i>the (relatively coarse) grain equivalent to a typical home range. We use a multi-grain (MG) occupancy model, aided by fine-grain remotely sensed imagery, to simultaneously estimate species-environment associations at both grains, conduct grain optimization to measure response grain, and apply this analysis framework to an example species: a medium-sized bird (<i>Tockus deckeni</i>) in a heterogeneous East African landscape. Based on home range analysis of movement data, we calculate an occupancy grain of 1km for <i>T. deckeni</i>. Using a grain optimization procedure across 32 grains from 10m to 500m, we identify 60m as the most strongly supported response grain for a suite of environmental variables, slightly coarser than opportunistic behavioral observations would have suggested. Validation confirms that the accuracy of the optimized MG occupancy model substantially exceeds that of equivalent single-grain (SG) occupancy models. We further use a simulation approach to assess the potential impacts of accounting for the multi-scale structure of species' environmental requirements on estimates of population size. We find that the more strongly supported MG approach consistently predicts a minimum population sizes in the study landscape that is much lower than that provided by the SG model. This suggests that SG approaches commonly used in conservation applications could lead to overly optimistic abundance and population estimates and that the MG approach may be more appropriate for supporting species conservation goals. More generally, we conclude that multi-grain approaches of the sort presented, and increasingly enabled by growing high-resolution remotely sensed data, hold great promise for offering a more mechanistic framework for assessing the appropriate grain(s) for population monitoring and management and enable more reliable estimates of abundances and species' distributions.</p>

opencc-zeroJan 2020View details →
dryad36/100

Data from: A new null model approach to quantify performance and significance for ecological niche models of species distributions

Aim: Ecological niche modelling requires robust estimation of model performance and significance, but common evaluation approaches often yield biased estimates. Null models provide a solution but are rarely used in this field. We implemented an important modification to existing null-model tests, evaluating null models with the same withheld records that were used to evaluate the real model. We built and evaluated models across a range of modelling scenarios and for various performance measures using the algorithm Maxent and the monk parakeet (Myiopsitta monachus). Location: Native range in Southern America and global invasions predominantly in North/Central America and Europe Methods: We tested the ability of models built under 15 scenarios (five sets of calibration records and three settings that varied the level of model complexity) to predict spatially independent evaluation data in the invaded range (in effect, testing the models under spatial transfer). We quantified performance with measures of discriminatory ability and overfitting based on AUC and the omission error rate. We estimated null distributions of these measures and calculated effect size and significance. We determined how these estimates varied across modelling scenarios, comparing with two tests existing in the literature. Results: Performance varied starkly across modelling scenarios. As expected, the measures of overfitting agreed with each other and provided different information than that of discriminatory ability. However, high performance per se did not show strong association with high effect size and significance. Main Conclusions: Ecological niche models should be assessed with measures of effect size and significance based on appropriate null distributions, in contrast to several approaches existing in the literature. The proposed approach using independent evaluation data, implemented with our accompanying code, allows such estimates for either the same or a different region/time period, and it merits use and continued development.

opencc-zeroDec 2018View details →
dryad36/100

Data from: Effectiveness of joint species distribution models in the presence of imperfect detection

<p>Joint species distribution models (JSDMs) are a recent development in biogeography and enable the spatial modelling of multiple species and their interactions and dependencies. However, most models do not consider imperfect detection, which can significantly bias estimates. This is one of the first papers to account for imperfect detection when fitting data with JSDMs and to explore the complications that may arise.</p> <p>A multivariate probit JSDM that explicitly accounts for imperfect detection is proposed, and implemented using a Bayesian hierarchical approach. We investigate the performance of the JSDM in the presence of imperfect detection for a range of factors, including varied levels of detection and species occupancy, and varied numbers of survey sites and replications. To understand how effective this JSDM is in practice, we also compare results to those from a JSDM that does not explicitly model detection but instead makes use of  "collapsed data". A case study of owls and gliders in Victoria Australia is also illustrated.</p> <p>Using simulations, we found that the JSDMs explicitly accounting for detection can accurately estimate intrinsic correlation between species with enough survey sites and replications. Reducing the number of survey sites decreases the precision of estimates, while reducing the number of survey replications can lead to biased estimates. For low probabilities of detection, the model may require a large number of survey replications to remove bias from estimates. However, JSDMs not explicitly accounting for detection may have a limited ability to disentangle detection from occupancy, which substantially reduces their ability to accurately infer the species distribution spatially. Our case study showed positive correlation between Sooty Owls and Greater Gliders, despite a low number of survey replications.</p> <p>To avoid biased estimates of inter-species correlations and species distributions, imperfect detection needs to be considered. However, for low probability of detection, the JSDMs explicitly accounting for detection is data hungry. Estimates from such models may still be subject to bias. To overcome the bias, researchers need to carefully design surveys and choose appropriate modelling approaches. The survey design should ensure sufficient survey replications for unbiased inferences on species inter-dependencies and occupancy.</p>

opencc-zeroJun 2021View details →
dryad36/100

Data from: Effects of grain size and niche breadth on species distribution modeling

Scale is a vital component to consider in ecological research, and spatial resolution or grain size is one of its key facets. Species distribution models (SDMs) are prime examples of ecological research in which grain size is an important component. Despite this, SDMs rarely explicitly examine the effects of varying the grain size of the predictors for species with different niche breadths. To investigate the effect of grain size and niche breadth on SDMs, we simulated four virtual species with different grain sizes/niche breadths using three environmental predictors (elevation, aspect, and percent forest) across two real landscapes of differing heterogeneity in predictor values. We aggregated these predictors to seven different grain sizes and modeled the distribution of each of our simulated species using MaxEnt and GLM techniques at each grain size. We examined model accuracy using the AUC statistic, Pearson's correlations of predicted suitability with the true suitability, and the binary area of presence determined from suitability above the maximum True Skill Statistic (TSS) threshold. Habitat specialists were more accurately modeled than generalist species, and the models constructed at the grain size from which a species was derived generally performed the best. The accuracy of models in the homogenous landscape deteriorated with increasing grain size to a greater degree than models in the heterogenous landscape. Variable effects on the model varied with grain size, with elevation increasing in importance as grain size increased while aspect lost importance. The area of predicted presence was drastically affected by grain size, with larger grain sizes over predicting this value by up to a factor of 14. Our results have implications for species distribution modeling and conservation planning, and we suggest more studies include analysis of grain size as part of their protocol.

opencc-zeroDec 2016View details →
dryad36/100

Accounting for imperfect detection in data from museums and herbaria when modeling species distributions: Combining and contrasting data-level versus model-level bias correction

The digitization of museum collections as well as an explosion in citizen science initiatives has resulted in a wealth of data that can be useful for understanding the global distribution of biodiversity, provided that the well-documented biases inherent in unstructured opportunistic data are accounted for. While traditionally used to model imperfect detection using structured data from systematic surveys of wildlife, occupancy models provide a framework for modelling the imperfect collection process that results in digital specimen data. In this study, we explore methods for adapting occupancy models for use with biased opportunistic occurrence data from museum specimens and citizen science platforms using 7 species of Anacardiaceae in Florida as a case study. We explored two methods of incorporating information about collection effort to inform our uncertainty around species presence: (1) filtering the data to exclude collectors unlikely to collect the focal species and (2) incorporating collection covariates (collection type, time of collection, and history of previous detections) into a model of collection probability. We found that the best models incorporated both the background data filtration step as well as collector covariates. Month, method of collection and whether a collector had previously collected the focal species were important predictors of collection probability. Efforts to standardize meta-data associated with data collection will improve efforts for modeling the spatial distribution of a variety of species.

opencc-zeroJun 2021View details →
zenodo36/100

Data and scripts for "Joint species-trait distribution modelling: The role of intraspecific trait variation in community assembly"

<p>The README explains how to reproduce the analyses presented in the paper <strong>"Joint species-trait distribution modelling: The role of intraspecific trait variation in community assembly"</strong> by Abrego et al.</p> <p>The input data for the script pipeline is the file &ldquo;Kilpisjarvi_plant_data.csv&rdquo;. This file includes the data on the plants and their traits in the long format. Hence, each row of the data matrix corresponds to measurements on one plant species in one study plot. The joint species-trait distribution modelling (JSTDM) pipeline that analyses these data consists of the following R-scripts.</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; <strong>S1_define_JSTDM_models.R</strong>. This script defines the JSDTM models (null model and environmental model) that include five response types for each species: the presence-absence, abundance conditional on presence, and the plot-level trait values of specific leaf area (SLA), leaf area (LA) and mean height (MH). The model is defined in the Hierarchical Modelling of Species Communities (HMSC) framework utilizing the R-package Hmsc. The models are saved in the file &ldquo;unfitted_models.RData&rdquo;.</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; <strong>S2_fit_models.R. </strong>This script loads the unfitted models and fits them using the posterior sampling methods implemented in the R-package Hmsc. The models are fitted with increasing thinning until thin=100, which value was used to generate the results of the paper. The fitted models are saved in the file "models_thin_100_samples_250_chains_4.Rdata".</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; <strong>S3_plot_Omega_matrices.R. </strong>This script loads the fitted models and plots the association matrices (Fig. 2 of the paper). The csv file containing the values used to construct Fig. 2 is also given (figure2Cdata.csv and figure2Ddata.csv).</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; <strong>S4_show_VP_Beta_Gamma.R. </strong>This script loads the fitted models and extracts information on the variance partitionings (VP; Figs. S3 and S4 of the paper), the relationships between response types and environmental predictors (beta; Fig. S2 of the paper), and the relationships between response types and species-level traits (gamma; Fig. S5 of the paper). The csv file containing the values used to construct Fig. S2 (figureS2Adata.csv and figureS2Bdata.csv), Fig. S3 (figureS3data.csv), Fig. S4 (figureS4data.csv) and Fig S5 (figureS5data.csv) are also given.</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; <strong>S5_conditional_cross_validation.R</strong>. This script performs 10-fold cross validation to the data to test the predictive power related to the modelled plant traits. The script performs both regular (unconditional) cross-validation where all data are masked for the test fold, and conditional cross-validation where only the trait data (but not the abundance data) are masked for the test fold.</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; <strong>S6_show_conditional_cross_validation_results.R. </strong>This script plots the results of cross-validation (Fig. 3 of the paper). The csv file containing the values used to construct Fig. 3 is also given (figure3Adata.csv and figure3Bdata.csv).</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; <strong>S7_scenario_predictions.R</strong>. This script performs the scenario simulations described and shown in Fig. 4 of the paper. The csv file containing the values used to construct Fig. 4 is also given (figure4Bdata.csv and figure4Cdata.csv).</p>

opencc-by-4.0May 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record