Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
24
datasets available to search
ShareScore release 0.9.0
Dataset results
24 results for “joint species distribution modelling”
Code and data for Bayesian joint species distribution model selection for community-level prediction
<p>Code and data for reproducing the analysis in the manuscript "Bayesian joint species distribution model selection for community-level prediction." Provided data include percent cover observations for 39 modeled vascular plant species within boreal forest understory communities and environmental model covariates. R code is provided to generate model inputs, apply alternative models, generate out-of-sample predictions, and calculate associated community and species log scores and alternative model evaluation metrics. Further, R source code is provided to implement the multinomial joint species distribution model defined in the manuscript. Details on the data, its processing, and the alternative model definitions and structure can be found in the main text of the manuscript. Provided data are currently being used in ongoing analyses and coordination with authors may be warranted to avoid duplicate publication. Potential users are encouraged to consider collaboration with authors when useful and appropriate. Misinterpretation of data may occur if used outside the context of the original analysis. All data are made available in their current state. While significant efforts have been made to ensure data accuracy, complete accuracy cannot be guaranteed. Data may be updated periodically. It is the responsibility of the data user to check for updated versions of the data.</p>
Improving distribution models of sparsely-documented disease vectors by incorporating information on related species via joint modeling
<p>A necessary component of understanding vector-borne disease risk is the accurate characterization of the distributions of their vectors. Species distribution models have been successfully applied to data-rich species but may produce inaccurate results for sparsely-documented vectors. In light of global change, vectors that are currently not well-documented could become increasingly important, requiring tools to predict their distributions. One way to achieve this could be to leverage data on related species to inform the distribution of a<strong> </strong>sparsely-documented vector based on the assumption that the environmental niches of related species are not independent. Relatedly, there is a natural dependence of the spatial distribution of a disease on the spatial dependence of its vector. Here, we propose to exploit these correlations by fitting a hierarchical model jointly to data on multiple vector species and their associated human diseases to improve distribution models of sparsely-documented species. To demonstrate this approach, we evaluated the ability of twelve models—which differed in their pooling of data from multiple vector species and inclusion of disease data—to improve distribution estimates of sparsely-documented vectors. We assessed our models on two simulated data sets, which allowed us to generalize our results and examine their mechanisms. We found that when the focal species is sparsely documented, incorporating data on related vector species reduces uncertainty and improves accuracy by reducing overfitting. When data on vector species are already incorporated, disease data only marginally improve model performance. However, when data on other vectors are not available, disease data can improve model accuracy and reduce overfitting and uncertainty. We then assessed the approach on empirical data on ticks and tick-borne diseases in Florida and found that incorporating data on other vector species improved model performance. This study illustrates the value of exploiting correlated data via joint modeling to improve distribution models of data-limited species.</p>
Fig. 2 in Parasite species co-occurrence patterns on Peromyscus: Joint species distribution modelling
Fig. 2. Results of variance partitioning for variation in ectoparasite prevalence explained by fixed and random effects for each ectoparasite species (columns). Explained variance presented for the constrained model for deer mice (n = 229 individuals). DM, deer mice; RBV, southern red-backed vole; WJM, woodland jumping mouse; PA, population abundance. Population abundance of small mammal species measured as captures per 100 trap nights. (For interpretation of the references to colour in this figure legend, the reader is referred to the Web version of this article.)
Improving distribution models of sparsely-documented disease vectors by incorporating information on related species via joint modeling
Open the record for dataset details and reuse information.
Code and data for Bayesian joint species distribution model selection for community-level prediction
Open the record for dataset details and reuse information.
Data from: Effectiveness of joint species distribution models in the presence of imperfect detection
<p>Joint species distribution models (JSDMs) are a recent development in biogeography and enable the spatial modelling of multiple species and their interactions and dependencies. However, most models do not consider imperfect detection, which can significantly bias estimates. This is one of the first papers to account for imperfect detection when fitting data with JSDMs and to explore the complications that may arise.</p> <p>A multivariate probit JSDM that explicitly accounts for imperfect detection is proposed, and implemented using a Bayesian hierarchical approach. We investigate the performance of the JSDM in the presence of imperfect detection for a range of factors, including varied levels of detection and species occupancy, and varied numbers of survey sites and replications. To understand how effective this JSDM is in practice, we also compare results to those from a JSDM that does not explicitly model detection but instead makes use of "collapsed data". A case study of owls and gliders in Victoria Australia is also illustrated.</p> <p>Using simulations, we found that the JSDMs explicitly accounting for detection can accurately estimate intrinsic correlation between species with enough survey sites and replications. Reducing the number of survey sites decreases the precision of estimates, while reducing the number of survey replications can lead to biased estimates. For low probabilities of detection, the model may require a large number of survey replications to remove bias from estimates. However, JSDMs not explicitly accounting for detection may have a limited ability to disentangle detection from occupancy, which substantially reduces their ability to accurately infer the species distribution spatially. Our case study showed positive correlation between Sooty Owls and Greater Gliders, despite a low number of survey replications.</p> <p>To avoid biased estimates of inter-species correlations and species distributions, imperfect detection needs to be considered. However, for low probability of detection, the JSDMs explicitly accounting for detection is data hungry. Estimates from such models may still be subject to bias. To overcome the bias, researchers need to carefully design surveys and choose appropriate modelling approaches. The survey design should ensure sufficient survey replications for unbiased inferences on species inter-dependencies and occupancy.</p>
Data and scripts for "Joint species-trait distribution modelling: The role of intraspecific trait variation in community assembly"
<p>The README explains how to reproduce the analyses presented in the paper <strong>"Joint species-trait distribution modelling: The role of intraspecific trait variation in community assembly"</strong> by Abrego et al.</p> <p>The input data for the script pipeline is the file “Kilpisjarvi_plant_data.csv”. This file includes the data on the plants and their traits in the long format. Hence, each row of the data matrix corresponds to measurements on one plant species in one study plot. The joint species-trait distribution modelling (JSTDM) pipeline that analyses these data consists of the following R-scripts.</p> <p>· <strong>S1_define_JSTDM_models.R</strong>. This script defines the JSDTM models (null model and environmental model) that include five response types for each species: the presence-absence, abundance conditional on presence, and the plot-level trait values of specific leaf area (SLA), leaf area (LA) and mean height (MH). The model is defined in the Hierarchical Modelling of Species Communities (HMSC) framework utilizing the R-package Hmsc. The models are saved in the file “unfitted_models.RData”.</p> <p>· <strong>S2_fit_models.R. </strong>This script loads the unfitted models and fits them using the posterior sampling methods implemented in the R-package Hmsc. The models are fitted with increasing thinning until thin=100, which value was used to generate the results of the paper. The fitted models are saved in the file "models_thin_100_samples_250_chains_4.Rdata".</p> <p>· <strong>S3_plot_Omega_matrices.R. </strong>This script loads the fitted models and plots the association matrices (Fig. 2 of the paper). The csv file containing the values used to construct Fig. 2 is also given (figure2Cdata.csv and figure2Ddata.csv).</p> <p>· <strong>S4_show_VP_Beta_Gamma.R. </strong>This script loads the fitted models and extracts information on the variance partitionings (VP; Figs. S3 and S4 of the paper), the relationships between response types and environmental predictors (beta; Fig. S2 of the paper), and the relationships between response types and species-level traits (gamma; Fig. S5 of the paper). The csv file containing the values used to construct Fig. S2 (figureS2Adata.csv and figureS2Bdata.csv), Fig. S3 (figureS3data.csv), Fig. S4 (figureS4data.csv) and Fig S5 (figureS5data.csv) are also given.</p> <p>· <strong>S5_conditional_cross_validation.R</strong>. This script performs 10-fold cross validation to the data to test the predictive power related to the modelled plant traits. The script performs both regular (unconditional) cross-validation where all data are masked for the test fold, and conditional cross-validation where only the trait data (but not the abundance data) are masked for the test fold.</p> <p>· <strong>S6_show_conditional_cross_validation_results.R. </strong>This script plots the results of cross-validation (Fig. 3 of the paper). The csv file containing the values used to construct Fig. 3 is also given (figure3Adata.csv and figure3Bdata.csv).</p> <p>· <strong>S7_scenario_predictions.R</strong>. This script performs the scenario simulations described and shown in Fig. 4 of the paper. The csv file containing the values used to construct Fig. 4 is also given (figure4Bdata.csv and figure4Cdata.csv).</p>
Joint species distribution modeling reveals a changing prey landscape for North Pacific right whales on the Bering shelf
<p>The eastern North Pacific right whale (NPRW) is the most endangered population of whale and has been observed north of its core feeding ground in recent years with low sea ice extent. Sea ice and water temperature are important drivers for zooplankton dynamics within the whale's core feeding ground in the southeastern Bering Sea, seasonally forming stable fronts along the shelf that give rise to distinct zooplankton communities. A northward shift in NPRW distribution driven by changing distribution of prey resources could put this species at increased risk of entanglement and vessel strikes. We modeled the abundance of NPRW prey, <em>Calanus glacialis</em>, <em>Neocalanus</em>, and <em>Thysanoessa</em> species, using a dynamic biophysical food web model of nine zooplankton guilds in the Bering shelf zooplankton community during a period of warming (2006–2016). This model is unique from prior zooplankton studies from the region in that it includes density dependence, thereby allowing us to ask whether species interactions influence zooplankton dynamics. Modeling confirmed the importance of sea ice and ocean temperature to zooplankton dynamics in the region. Density-independent growth drove community dynamics while dependent factors were comparatively minimal. Overall, <em>Calanus</em> responded to environmental terms, with the strength and direction of response driven by copepodite stage. <em>Neocalanus</em> and <em>Thysanoessa</em> responses were weaker, likely due to their primary occurrence on the outer shelf. We also modeled the steady-state (equilibrium) abundance of <em>Calanus</em> in conditions with and without wind gusts to test whether advection of outer shelf species might disrupt steady-state dynamics of <em>Calanus</em> abundance; results did not support disruption. Given the annual fall sampling design, we interpret our results as follows: low ice-extent winters induced stronger spring winds and weakened fronts on the shelf, thereby advecting some outer shelf species into the study region; increased development rates in these warm conditions influenced the proportion of <em>C. glacialis</em> copepodite stages over the season. Residual correlation suggests missing drivers, possibly predators and phytoplankton bloom composition. Given the continued loss of sea ice in the region and projected continued warming, our findings suggest that <em>C. glacialis</em> will move northward, and thus, whales may move northward to continue targeting them.</p>
Scale-dependence of ecological assembly rules: insights from empirical datasets and joint species distribution modelling
Open the record for dataset details and reuse information.
The impact of varying spatiotemporal scales on different joint species distribution models: A case study of pelagic fish species in the northwest Pacific Ocean
Open the record for dataset details and reuse information.
Data from: Using joint species distribution modelling to identify climatic and non-climatic drivers of Afrotropical ungulate distributions
Open the record for dataset details and reuse information.
Data from: Effectiveness of joint species distribution models in the presence of imperfect detection
Open the record for dataset details and reuse information.
Joint species distribution modeling reveals a changing prey landscape for North Pacific right whales on the Bering shelf
Open the record for dataset details and reuse information.
Scale dependency of joint species distribution models challenges interpretation of biotic interactions
Open the record for dataset details and reuse information.
Data from: Understanding co-occurrence by modelling species simultaneously with a Joint Species Distribution Model (JSDM)
A primary goal of ecology is to understand the fundamental processes underlying the geographic distributions of species. Two major strands of ecology – habitat modelling and community ecology – approach this problem differently. Habitat modellers often use species distribution models (SDMs) to quantify the relationship between species' and their environments without considering potential biotic interactions. Community ecologists, on the other hand, tend to focus on biotic interactions and, in observational studies, use co‐occurrence patterns to identify ecological processes. Here, we describe a joint species distribution model (JSDM) that integrates these distinct observational approaches by incorporating species co‐occurrence data into a SDM. JSDMs estimate distributions of multiple species simultaneously and allow decomposition of species co‐occurrence patterns into components describing shared environmental responses and residual patterns of co‐occurrence. We provide a general description of the model, a tutorial and code for fitting the model in R. We demonstrate this modelling approach using two case studies: frogs and eucalypt trees in Victoria, Australia. Overall, shared environmental correlations were stronger than residual correlations for both frogs and eucalypts, but there were cases of strong residual correlation. Frog species generally had positive residual correlations, possibly due to the fact these species occurred in similar habitats that were not fully described by the environmental variables included in the JSDM. Eucalypt species that interbreed had similar environmental responses but had negative residual co‐occurrence. One explanation is that interbreeding species may not form stable assemblages despite having similar environmental affinities. Environmental and residual correlations estimated from JSDMs can help indicate whether co‐occurrence is driven by shared environmental responses or other ecological or evolutionary process (e.g. biotic interactions), or if important predictor variables are missing. JSDMs take into account the fact that distributions of species might be related to each other and thus overcome a major limitation of modelling species distributions independently.
Data from: Testing species assemblage predictions from stacked and joint species distribution models
Aim: Predicting the spatial distribution of species assemblages remains an important challenge in biogeography. Recently, it has been proposed to extend correlative species distribution models (SDMs) by taking into account (a) covariance between species occurrences in so-called joint species distribution models (JSDMs) and (b) ecological assembly rules within the SESAM (spatially explicit species assemblage modelling) framework. Yet, little guidance exists on how these approaches could be combined. We, thus, aim to compare the accuracy of assemblage predictions derived from stacked and from joint SDMs. Location: Switzerland Taxon Birds, tree species Methods: Based on two monitoring schemes (national forest inventory and Swiss breeding bird atlas), we built SDMs and JSDMs for tree species (at 100m resolution) and forest birds (at 1km resolution). We tested accuracy of species assemblage and richness predictions on holdout data using different stacking procedures and ecological assembly rules. Results Despite minor differences, results were consistent between birds and tree species. Cross-validated species-level model performance was generally higher in SDMs than JSDMs. Differences in species richness and assemblage predictions were larger between stacking procedures and ecological assembly rules than between stacked SDMs and JSDMs. On average, predictions were slightly better for stacked SDMs compared to JSDMs, probabilistic stacks outperformed binary stacks, and ecological assembly rules yielded best predictions. Main conclusions: When predicting the composition of species assemblages, the choice of stacking procedure and ecological assembly rule seems more decisive than differences in underlying model type (SDM vs. JSDM). JSDMs do not seem to improve community predictions compared to SDMs or improve predictions for rare species. Still, JSDMs may provide additional insights into community assembly and may help deriving hypotheses about prevailing biotic interactions in the system. We provide simple rules of thumb for choosing appropriate modelling pathways. Future studies should test these preliminary guidelines for other taxa and biogeographic realms as well as for other JSDM algorithms.
Data from: Comparing the prediction of joint species distribution models with respect to characteristics of sampling data
Biotic interactions have been rarely included in traditional species distribution models, wherein Joint Species Distribution Models (JSDMs) emerge as a feasible approach to incorporate environmental factors and interspecific interactions simultaneously, making it a powerful tool for analyzing the structure and assembly processes of biotic communities. However, the predictability and statistical robustness of JSDMs are largely unknown because of the lack of research efforts for those newly developed models. This study systematically evaluated the performances of five JSDMs in predicting the occurrence and biomass of multiple species, with a particular focus on diverse characteristics of sampling data, including type of response variables, number of sampling sites, and the number of species included in models. In general, most models yielded satisfactory performances on fitting to observed data and on the estimation of environmental effects; however, they showed less well performances in evaluating species associations, and their predictability had large variations. The JSDMs showed inconsistent performances between the goodness-of-fit and predictability in cross-validation, and the Boral model was relatively robust than others. The predictability of JSDMs was less influenced by sample sizes and substantially improved by incorporating rare species. This study contributes to an appropriate model selection and application of JSDMs.
Data from: Understanding co-occurrence by modelling species simultaneously with a Joint Species Distribution Model (JSDM)
Open the record for dataset details and reuse information.
Data from: Testing species assemblage predictions from stacked and joint species distribution models
Open the record for dataset details and reuse information.
Data from: Does trait-based joint species distribution modelling reveal the signature of competition in stream macroinvertebrate communities?
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.