Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
487
datasets available to search
ShareScore release 0.7.1
Dataset results
487 results for “species distribution modeling”
Data from: How far can I extrapolate my species distribution model? Exploring Shape, a novel method
<p>Species distribution and ecological niche models (hereafter SDMs) are popular tools with broad applications in ecology, biodiversity conservation, and environmental science. Many SDM applications require projecting models in environmental conditions non-analog to those used for model training (extrapolation), giving predictions that may be statistically unsupported and biologically meaningless. We introduce a novel method, Shape, a model-agnostic approach that calculates the extrapolation degree for a given projection data point by its multivariate distance to the nearest training data point. Such distances are relativized by a factor that reflects the dispersion of the training data in environmental space. Distinct from other approaches, Shape incorporates an adjustable threshold to control the binary discrimination between acceptable and unacceptable extrapolation degrees. We compared Shape's performance to five extrapolation metrics based on their ability to detect analog environmental conditions in environmental space and improve SDMs suitability predictions. To do so, we used 760 virtual species to define different modeling conditions determined by species niche tolerance, distribution equilibrium condition, sample size, and algorithm. All algorithms had trouble predicting species niches. However, we found a substantial improvement in model predictions when model projections were truncated independently of extrapolation metrics. Shape's performance was dependent on extrapolation threshold used to truncate models. Because of this versatility, our approach showed similar or better performance than the previous approaches and could better deal with all modeling conditions and algorithms. Our extrapolation metric is simple to interpret, captures the complex shapes of the data in environmental space, and can use any extrapolation threshold to define whether model predictions are retained based on the extrapolation degrees. These properties make this approach more broadly applicable than existing methods for creating and applying SDMs. We hope this method and accompanying tools support modelers to explore, detect, and reduce extrapolation errors to achieve more reliable models.</p>
Data for: The meta-analysis of the effects of spatial sampling bias correction on presence only species distribution models
<p>This dataset contains information extracted from 70 studies identified through a systematic review of the peer-reviewed literature (Web of Science and SCOPUS databases both searched on the 13/02/2023) to evaluate the effect of spatial sampling bias correction methods in presence-only species distribution models.</p>
The global distribution of plants used by humans datasets: list of utilised species, occurrence data and model outputs at 10 arc-minutes spatial resolution
<p>Datasets and model outputs used to map the global distribution of utilised plants by humans. The folder is composed of two subfolders <em>raw_data</em> and <em>processed_data</em> containing respectively the list of utilised plant species modelled -<em>utilised_plants_species_list.csv</em>-, and their occurrence data -<em>occurrence_data.zip-</em> and predicted distribution -<em>species_proba_per_cell.rds-.</em></p> <p> </p> <ul> <li>The file <em>utilised_plants_species_list.csv</em> in the <em>raw_data</em> folder contains a<strong> </strong>list of 35687 plant species (and hybrids) used by humans and 10 plant use categories with the following 14 fields:</li> </ul> <p><strong>plant_ID:<em> </em></strong>plant identifier number ranging from between 1-35687</p> <p><strong>binomial_acc_name:</strong> binomial accepted name of the plant species</p> <p><strong>author_acc_name</strong>: name of the author(s)</p> <p><strong>is_hybrid:</strong> logical TRUE or FALSE indicating whether the species is an hybrid or not.</p> <p><strong>AnimalFood:</strong> forage and fodder for vertebrate animals only.</p> <p><strong>EnvironmentalUses:</strong> examples include intercrops and nurse crops, ornamentals, barrier hedges, shade plants, windbreaks, soil improvers, plants for revegetation and erosion control, wastewater purifiers, indicators of the presence of metals, pollution, or underground water.</p> <p><strong>Fuels:</strong> charcoal, petroleum substitutes, fuel alcohols, etc. Given the importance of energy plants for people, those were distinguished from Materials.</p> <p><strong>GeneSources:</strong> wild relatives of major crops which may possess traits associated with biotic or abiotic resistance and may be valuable for breeding programs.</p> <p><strong>HumanFood:</strong> food for humans only, including beverages and food additives.</p> <p><strong>InvertebrateFood:</strong> plants consumed by invertebrates used by humans, such as bees, silkworms, lac insects and edible grubs.</p> <p><strong>Materials:</strong> woods, fibers, cork, cane, tannins, latex, resins, gums, waxes, oils, lipids, etc. and their derived products.</p> <p><strong>Medicines:</strong> both human and veterinary.</p> <p><strong>Poisons:</strong> plants which are poisonous to both vertebrates and invertebrates, both accidentally and intentionally, e.g., for hunting and fishing, molluscicides, herbicides, insecticides.</p> <p><strong>SocialsUses:</strong> plants used for social purposes, which cannot be defined as food or medicine, for instance, masticatories, smoking materials, narcotics, hallucinogens and psychoactive drugs, and plants with ritual or religious significance.</p> <p><strong>Totals:</strong> total number of uses recorded for a species</p> <p> </p> <ul> <li>The zipfile <em>occurrence_data.zip</em> in the <em>processed_data</em> folder contains 35687 Comma Separated Values (CSV) files, one for each species, containing curated geographic occurrence records used to build species distribution models with the following 14 fields:</li> </ul> <p><strong>Species:</strong> the binomial accepted name of the species</p> <p><strong>Fullname:</strong> same as species</p> <p><strong>decimalLongitude:</strong> the geographic longitude of the occurrence records of the species in decimal degrees</p> <p><strong>decimalLatitude:</strong> the geographic latitude of the occurrence records of the species in decimal degrees</p> <p><strong>countryCode:</strong> a three-letter standard abbreviation for the country of the occurrence locality</p> <p><strong>coordinateUncertaintyinMeters</strong>: indicator for the accuracy of the coordinate location, described as the radius of a circle around the stated point location</p> <p><strong>year:</strong> year of the observation of the occurrence record of the species</p> <p><strong>individualCount:</strong> the number of individuals present at the time of the observation</p> <p><strong>gbifID:</strong> unique identifier number for the occurrence from the original database</p> <p><strong>basisOfRecords:</strong> the type of the individual record, e.g. observation, physical specimen, fossil, living ex-situ, culture collection specimen</p> <p><strong>institutionCode</strong>: the name of the institution or organization listed as the data publisher on GBIF</p> <p><strong>establishmentMeans:</strong> statement about whether an organism has been introduced to a given place and time through the direct or indirect activity of modern humans</p> <p><strong>is_cultivated_observation:</strong> whether or not an organism is cultivated</p> <p><strong>sourceID:</strong> name of the source database</p> <p> </p> <ul> <li>The file <em>species_proba_per_cell.rds</em> in the <em>processed_data</em> folder is<em> a R Data Serialization </em>(RDS) file containing a data.table object with the following 3 fields:</li> </ul> <p><strong>plant_ID:</strong><em> </em>plant identifier number ranging from between 1-35687</p> <p><strong>proba:</strong> species occurrence probability</p> <p><strong>cell:</strong><em> </em>raster grid cell number between 1-2251762</p> <p>This object can be used in combination with a raster layer to reconstruct the modelled distribution of each species or retrieve species richness and endemism.</p>
Evaluation data for: Adaptive sampling by citizen scientists improves species distribution model performance: a simulation study
<p>All of the evaluation data for the simulations in the paper: Adaptive sampling by citizen scientists improves species distribution model performance: a simulation study. We considered the impact of five adaptive sampling methods on the performance of species distribution models (SDMs), please see the paper for more information. Contained in this repository are the evaluation metrics (AUC, mean square error (MSE) and correlation) for SDMs before and after adaptive sampling has taken place. The MSE and correlation evaluation metrics were calculated against the true distributions of the species. These files are those with "combined_outputs" in the titles. The repository also contains the observations of all the species in the simulations both before and after adaptive sampling (the files with "all_observations" in the title.</p> <p>These datasets are to be used with the plotting and evaluation scripts in the GitHub repository associated with the paper.</p>
Data for: A new threshold selection method for species distribution models with presence-only data: extracting the mutation point of the P/E curve by threshold regression
<p>Selecting thresholds to convert continuous predictions of species distribution models proves critical for many real-world applications and model assessments. Prevalent threshold selection methods for presence-only data require unproven pseudo-absence data or subjective researchers' decisions. This study proposes a new method, Boyce-Threshold Quantile Regression (BTQR), to determine thresholds objectively without pseudo-absence data. We summarize that the mutation point is a typical shape feature of the predicted-to-expected (P/E) curve after reviewing relevant articles. Analysis based on source-sink theory suggests that this mutation point may represent a transition in habitat types and serve as an appropriate threshold. Threshold regression is introduced to accurately locate the mutation point.</p> <p>To validate the effectiveness of BTQR, we used four virtual species of varying prevalence and a real species with reliable distribution data. Six different species distribution models were employed to generate continuous suitability predictions. BTQR and nine other traditional methods transformed these continuous outputs into binary results. Comparative experiments show that BTQR has advantages in terms of accuracy, applicability, and consistency over the existing methods.</p>
Incorporating plant phenological responses into species distribution models (SDMs) reduces estimates of future species loss and turnover
<p>Anthropogenetic climate change has caused distribution shifts of many species, and species distribution models (SDMs) are central for documenting this relationship. However, most SDMs rarely consider the evolution of climate-sensitive functional traits, such as phenology, which strongly affect species fitness. Using >120,000 herbarium specimens representing 360 plant species across the eastern United States, we developed a novel "phenology-informed" SDM that integrates dynamic phenological responses to changing climates. Compared to standard SDMs, our phenology-informed SDMs forecast lower species habitat loss and less species turnover under climate change. These results suggest that phenotypic plasticity or local adaptation in phenology may help species adjust their ecological niches and persist in their habitats under rapid environmental change. Our findings reveal how phenology variation mediates species distributions and affects regional biodiversity patterns. Our newly developed model also circumvents the need for mechanistic models, facilitating the deployment of trait-based SDMs across unprecedented spatial and taxonomic scales.</p>
Data from: "Cross-realm transferability of species distribution models – species characteristics and prevalence matter more than modelling methods applied"
<h2>Abstract</h2> <p>This data contains occurrence observations (presence-absence) of 11 aquatic macrophytes from Bothnian Sea and Lake Puruvesi, and environmental covariates used to build species distribution models (SDMs) in paper "Cross-realm transferability of species distribution models – species characteristics and prevalence matter more than modelling methods applied" in Ecological Modelling.</p> <p>In addition to data files, also R code for fitting the SDMs is supplied, as is the R code to replicate the analysis conducted in the paper. The data is stored in rdata format (point data), without coordinate information due to data policy restrictions. The species in the data are <em>Isoëtes lacustris, Isoëtes echinospora, Ranunculus reptans, Ranunculus schmalhausenii, Potamogeton berchtoldii, Potamogeton perfoliatus, Potamogeton gramineus, Myriophyllum alterniflorum, Equisetum fluviatile, Eleocharis acicularis and Elodea canadensis. </em>The environmental covariates are bottom water salinity, turbidity, sandy substrate occurrence, colored dissolved organic matter (CDOM), surface fetch, sampling depth, total nitrogen and total phosphorus, and distance to closest <em>Phragmites australis </em>reed bed. </p> <h2>Objective of the study </h2> <p>The modelling objective of the paper was species distribution model (SDM) transferability assesment. Transferability was assessed using models built in marine areas in projecting the distributions of the target species in Lake Puruvesi, Saimaa, Eastern Finland. Macrophyte mapping data from Lake Puruvesi was used as independent test data, against which transferability of the models was assessed.</p> <h2>Location</h2> <p>The species data was collected from two geographic areas: Bothnian Bay (Baltic Sea) and Lake Puruvesi (Eastern Finland). The marine observations from Bothnian Bay were split into three overlapping areas (areas 1-3), to test the effect of input data gradient length to SDM transferability. The largest marine sampling area (Area 3) ranged from 62.95, 65.91 latitude and 19.14, 27.99 longitude. Area 2 ranged from 63.95, 65.91 latitude and 21.53, 27.99 longitude. Area 1 ranged from 64.91, 65.91 latitude and 23.82, 27.99 longitude. The Hummonselkä subbasin of Lake Puruvesi, where the macrophyte test data was collected, is located at 61.89, 62.05 latitude and 29.58, 29.78 longitude.</p> <h2>Species data </h2> <p>The species observations were collected using diving transects placed in the floor of the sea or lake, and species observations were recorded in 2 m22 grid cells separated by 10 meters along the transect or 1 meters depth, depending which criteria was met first. The species data was collected in 2010 - 2020 from marine area, and 2017 from Lake Puruvesi. All macrophytes in 2 x 1 m frames were identified to species level by the diver, and the data contained information on species presence or absence in each grid cell. The diving transects were conducted using systematic survey protocol used in the underwater inventories of the Finnish Underwater Biodiversity Survey Program (VELMU) (Frosblom & Virtanen et al. 2024). The locations of the diving transects were not randomly distributed, but were placed using expert judgement. As all our study species are macroscopic and relatively easily identifiable in the field (with the exception of possibility of mixing <em>I. echinospora</em> and <em>I. lacustris</em>), we consider the absences in our observation data to indicate true absences. That said, as the observation area is rather small (2 m2), it is possible that a species may be found in the site of investigation (e.g. a small lagoon) but be located outside the vegetation sampling grid.</p> <h2>Environmental data</h2> <h3>Bottom water salinity</h3> <p>The seasonal mean bottom water salinity was modeled using a generalized additive model with mean salinity as response, with log-link and gamma distribution for the errors. This was necessary to keep the resulting predictions positive. Bottom depth, CDOM, river influence and spatial location were used as predictors. Data from 448 locations were used and each location had a minimum of three observations. The model was validated using 30 % of the data left outside of the model fitting. The explained deviance of the model was 0.94 and the correlation between raw data and predicted values was 0.95 with few outliers.</p> <h3>Turbidity</h3> <p>Maps of turbidity (in FNU, Formazin Nephelometric Unit) were generated from Sentinel-2 Multi-Spectral Imager (MSI) observations using the Case-2 Regional Coast Colour (C2RCC) bio-optical inversion model, containing separate atmospheric correction and water quality parts. Before computing C2RCC, the original 10-meter input data was downsampled to 60 meters. The output variable of the C2RCC processor correlative to turbidity is the backscattering of total suspended sediments at 443 nm, which was further calibrated into turbidity (FNU) values using SYKE's empirical equations for coastal waters and clear lakes (for a similar approach, see Attila et al. (2013) and Sagerman, Hansen, and Wikström (2020)). Monthly observations of turbidity were aggregated into median composites to reduce the effects of cloud cover and other disturbances. Due to low solar elevation and ice cover in winter, the turbidity distribution maps are generated only for the summer months (May to September). The current processing covers years 2017 to 2021. An average raster layer was created from monthly observations as input for SDM building.</p> <h3>Probability of sandy substrate </h3> <p>Random forest model was used to classify sandy bottoms from Sentinel 2 MSI satellite images in shallow water areas. Identifying sandy substrate is based on the higher reflectance compared to other substrates. The model was trained and validated using diver recorded field observations in the Baltic area, and diver recorded and echo sounding observations in the freshwater area. For full coverage including areas beyond the shallow water, the satellite image classification was combined with boosted regression tree modelling result in the Baltic, and echo sounding based product in the freshwater region. The resulting layers were probabilities of sandy substrate with 10-meter cell resolution.</p> <h3>CDOM</h3> <p>We used different methods to estimate the CDOM levels in Bothnian Bay and Lake Puruvesi, based on biogeochemical model data and satellite images. For Lake Puruvesi, we applied the Finnish Environment Institute's (Syke) in-house CDOM algorithm to the Sentinel-2 MSI images processed by the C2RCC bio-optical processor (Brockmann et al. 2016). The observations in 10 m resolution were aggregated as monthly averages for each month of the summer season (May to October) from 2017 to 2021. For Bothnian Bay, we used Syke's in-house Sentinel-2 MSI CDOM layers (resolution: 60 m) aggregated as seasonal averages (1 Jul to 7 Sep). CDOM values are given as absorption coefficient of CDOM at 400 nm [m⁻¹].</p> <h3>Surface fetch</h3> <p>A surface fetch raster was produced to the Puruvesi and Bothnian Bay. The analysis required a feature layer of shorelines from Puruvesi and Baltic Sea. First, we created polyline from north to south spanning over the whole area of interest with a gap of 20 meters which is also the resolution of the output raster. These lines were then cut each time they hit the shoreline and the part of the line that was overlapping land was removed. The distance of the remaining lines was then calculated and a point with the distance value was created every 20 meters. Each time the line was cut when hitting an island for example and starting again from the other side of the island, the distance calculation started from 0. This created a point dataset with a distance value in each point. We repeated the procedure for 15 times for different compass directions with 22.5 degree intervals and calculated average fetch for each point location on 20 meters grid from these 15 point layers.</p> <h3>Depth</h3> <p>Depth was measured by a diver using a dive computer while surveying each vegetation grid cell, and measured depth was used when projecting model results to Puruvesi (transferability performance). In addition, a depth model for the freshwater region was created from Sentinel 2 MSI satellite image using the logarithmic band ratio model of blue and red band. The model was calibrated using diver recorded field observations and validated against echo sounding measurements. For more complete coverage and to include deep areas, echo soundings from multiple sources were combined with the satellite derived bathymetry. The cell resolution of the resulting depth layer was 10 meters.</p> <h3>Total nitrogen and phosphorus</h3> <p>Mean total nitrogen and phosphorus layers for marine area were produced using ArcGIS "splines with barriers" tool for the EEZ of Finland with 20 meters spatial resolution (Virtanen et al. 2018). Summer (July - September) nutrient measurements from 0 to 10 meters depth between 2010 and 2020, obtained from the VESLA database, were used as input data for the interpolation.</p> <p>Nitrogen and phosphorus measurements in Puruvesi between 2010 and 2020 was gathered from the VESLA database. Data from July to September was selected to represent the growing season. A mean value of NTOT and PTOT was then calculated for each location. Spline with Barriers (SwB) tool was used to interpolate the values (Arcmap 10.7.1). The tool uses a feature layer as barrier to create the raster representing only the area of interest. For the barrier and the extent of the interpolated raster we used a shapefile representing Lake Puruvesi shoreline. The resolution was set to 5x5 meters. SwB tool created an "extent box" around the area of interest which was removed with Extract by Mask tool using the shoreline feature layer. After the interpolation we noticed that either one of the locations was situated on land or the polygon used as barrier was "leaking". SwB doesn´t interpolate areas that doesn't have locations with values or aren´t connected to the main body of water. To fix this, the raster was extended outwards based on the values of nearby cells and after that the raster was masked again to remove any cells on land. The phosphorus interpolation provided negative values in southern parts on Enanlahti in Kontiolahti and Muholanlahti. These negative values were caused by considerably larger phosphorus values in Enanlahti Lamminniemi (9m) Enanlahti Lamminniemi (4m) locations when compared with the nearby Puruvesi Enanlahti location. The interpolation apparently continued to decrease the values according to the trend set by the difference between these locations and caused it to reach negative values. The southern parts of the bay, about 750 meters, was removed and new values were calculated based on the surrounding cells with Focal Statistics tool. The interpolations were validated by removing 20 % of the locations and reproducing the interpolation. The removed locations and their values were then compared to the interpolated raster. R∗2∗2 value from phosphorus interpolation model was 0.91 after removing two outliers and R22 value from nitrogen interpolation model was 0.715 after removing one outlier.</p> <h3>Distance to closest reed </h3> <p>The aquatic vegetation (<em>Phragmites australis</em> reeds) presence/absence maps were also generated from Sentinel-2 MSI data. The processing included extracting one month of data (July 2019) from green and near-infra-red bands from Sentinel-2 Global Mosaic (S2GM) service and transforming those to normalized-difference vegetation indices (NDVIs). After that, Bayesian statistics were used to predict the posterior probability of vegetation occurrence when distance from shore and NDVI were used as predictor variables. The posterior variable was thresholded and the resulting vegetation presence areas were sieved so that both too small vegetation areas (fewer than 5 pixels) or areas that were not directly attached to shoreline were removed. The resulting map has 10 m pixel size and tentatively represents the locations of reed belts or other shoreline-attached vegetation. This EO-based layer could also be referred to as helophytes or helophytic macrophytes, as it denotes a specific zone of vegetation with emergent aquatic plants containing leaf-green, particularly those that grow densely and have horizontally oriented leaves. In some lakes, this layer can represent, for example, thick stands of <em>Equisetum fluviatile</em>, although in most cases, it is associated with common reed belts. The approach is described in more detail in Koponen et al. (2022).</p> <h2>Data partitioning </h2> <p>Data was partitioned with 70/30 splitting into training and test (interpolation accuracy) data. The splitting was repeated 100 times for each species by randomly selecting 70 % of observations which were used to build each of the SDMs (GLM, GAM, BRT and BART). The partitioning was repeated for each of the three input data areas and 11 species. The input data indexes for replicating the split are supplied in the data files. </p> <h2>R code </h2> <p>Code files contain scripts for fitting the SDM models described in the paper using the data. Also code for calculating calidation statistics (modelling results) and code for statistical analyses for making inferences in the paper, are supplied. </p> <h2>Additional info</h2> <p>More details on modelling protocol may be found in the original paper in Ecological Modelling, and in the Supplementary Information file of the original paper, which follows the model reporting template "ODMAP" by Zurell et al. (2020).</p> <h2>References </h2> <p><span>Attila, J., et al. 2013. MERIS Case II water processor comparison on coastal sites of the northern Baltic Sea. - Remote Sensing of Environment 128: 138-149.</span></p> <p><span>Brockmann, C., et al. 2016. Evolution of the C2RCC neural network for Sentinel 2 and 3 for the retrieval of ocean colour products in normal and extreme optically complex waters. - In: Living Planet Symposium. p. 54.</span></p> <p><span>Forsblom, L., et al. 2024. Finnish inventory data of underwater marine biodiversity<span> </span>-Scientic data</span></p> <p><span>Koponen, S., et al. 2022. Blue Carbon Habitats: – a comprehensive mapping of Nordic salt marshes for estimating Blue Carbon storage potential. - Nordisk Ministerråd.</span></p> <p><span>Sagerman, J., et al. 2020. Effects of boat traffic and mooring infrastructure on aquatic vegetation: A systematic review and meta-analysis. - Ambio 49: 517-530.</span></p> <p><span>Zurell, D., et al. 2020. A standard protocol for reporting species distribution models. - Ecography 43: 1261-1277.</span></p> <p></p> <p> </p> <p> </p> <p> </p> <p> </p> <p> </p>
Data from: Effects of input data sources on species distribution model predictions across species with different distributional ranges
<p>Species distribution models (SDMs) are a popular tool in theoretical and quantitative ecology, and constitute the most widely used modelling framework in global change science and biodiversity conservation. As main data sources, SDMs require georeferenced biodiversity observations as a response or dependent variable (e.g. species occurrence, species richness, etc) and geographic layers of environmental information as predictors or independent variables (e.g. climate, land cover, vegetation indices derived from remote sensing, etc). However, although SDMs have become one of the most important quantitative tools for addressing regular and timely biodiversity assessments worldwide, these techniques are still subject to different sources of uncertainty that have been unequally assessed. Thus, despite uncertainty related to niche-based or distribution-based models has been addressed at different stages in the modelling process, an analysis of the effect of uncertainty coming from alternative data sources on the predictive ability of SDMs is still limited.</p> <p>Citizen-collected species occurrence data (e.g. eBird) are often used for fitting SDMs when data from standardized and expert-supported surveys (e.g. Atlases) are unavailable. On the other hand, macroclimate variables are much more commonly used as predictors in SDMs than other sources of information coming from remote sensing data. We assessed the effects of using different data sources (in both response and predictor variables) on SDM performance across a wide range of bird species with contrasting distributional ranges in the Iberian Peninsula (Portugal and Spain). To do that, a SDM ensemble-forecasting approach was implemented by using bird data from two different data sources: the semi-structured eBird project and standardized Atlases. We fitted SDMs with three predictor types: macroclimate, remotely sensed ecosystem functional attributes (EFAs) from vegetation indices, and their combination. Species were grouped in four range size classes. We also used different evaluation metrics to better assess the uncertainty of model predictions. We then applied generalized linear mixed-effects models to test the effect on model performance of input data source across distributional range sizes while accounting for different accuracy metrics. Pairwise comparisons between range projections were used to assess their spatial similarity.</p> <p>Our models demonstrated the usefulness and complementarity of different input data sources when modelling species distribution across different distributional ranges. Citizen science and remote sensing data contribute to update the knowledge of the distribution of the most threatened bird species by increasing the model accuracy. These findings highlight the need to integrate different data sources to improve the model predictions at regional scale. Our framework also underlines that model uncertainty should be examined more exhaustively at early stages of the modelling process.</p> <p>To perfom and replicate this study, this dataset provides all needed files (as tables) to fit SDMs: i) the Iberian bird species occurrences at 10km UTM square as a response or dependent variable; ii) the geographic layers of environmental information at 10km UTM square for the Iberian Peninsula as predictors or independent variables, such as climate data, ecosystem functioning attributes (EFAs) and the combined climate and EFA data. The dataset is provided by four <em>*.csv</em> files named as:</p> <p><em>1) The_Iberian_bird_species_occurrences_dataset_10km.csv</em></p> <p><em>2) CHELSA_bioclimate_variables_IP10km.csv</em></p> <p><em>3) MODIS_EVI-based_EFAs_IP10km.csv</em></p> <p><em>4) Combined_bioclimate_EFA_dataset_IP10km.csv</em></p> <p>For a more detailed description of the main dataset and each of these subdatasets, please refer to the attached README file.</p> <p><strong>Keywords:</strong> bird atlas, eBird data, ecosystem functional attributes (EFAs), Iberian Peninsula, IUCN categories, Model accuracy, MODIS EVI, narrow-ranged species, remote sensing, species distribution models (SDMs), widespread species</p>
Fig.1 in Associations Between Habitat Quality And Body Size In The Carpathian-Podolian Land Snail Vestia Turgida: Species Distribution Model Selection And Assessment Of Performance
Fig.1. Sampling sites for Vestia turgida in Ukraine (photo by O. Baidashnikov).
Fig. 5 in Associations Between Habitat Quality And Body Size In The Carpathian-Podolian Land Snail Vestia Turgida: Species Distribution Model Selection And Assessment Of Performance
Fig. 5. Partial dependence plot for terrain roughness index (tri).
Fig. 7 in Associations Between Habitat Quality And Body Size In The Carpathian-Podolian Land Snail Vestia Turgida: Species Distribution Model Selection And Assessment Of Performance
Fig. 7. Partial dependence plot for silt content (SLT).
Fig. 6 in Associations Between Habitat Quality And Body Size In The Carpathian-Podolian Land Snail Vestia Turgida: Species Distribution Model Selection And Assessment Of Performance
Fig. 6. Partial dependence plot for pH water (phh2o).
Fig. 2 in Interspecific Interactions as a Factor of Limitation of Geographical Distribution: Evidence Obtained by Modeling Home Ranges of Vole Twin Species Microtus Arvalis – M. Levis (Rodentia, Microtidae)
Fig. 2. Potential distribution of the East European vole (Microtus levis). Captions as in fig.1.
Publication release: How well do species distribution models predict occurrences in exotic ranges?
<div class="record-description"> <p>Species distribution models (SDMs) are widely used predictive tools to forecast potential biological invasions. However, the reliability of SDMs extrapolated to exotic ranges remains understudied, with most analyses restricted to few species and equivocal results. We examined the spatial transferability of SDMs for 647 non-indigenous species extrapolated across 1,867 invaded ranges, and identify what factors may help differentiate predictive success from failure. We performed a large-scale assessment of the transferability of SDMs using two modelling approaches: generalized additive models (GAMs) and MaxEnt. We fitted SDMs on the native ranges of species and extrapolated them to exotic ranges. We examined the influence of general factors and factors related to biological invasions on spatial transferability.</p> <p>Here, we provide the code and data for publication in Global Ecology and Biogeography as part of Nguyen and Leung 2022 "How well do species distribution models predict occurrences in exotic ranges?". Provided are the files and scripts necessary to fit and validate the SDMs using distirbutional data from their native and exotic ranges, respectively, formulated as generalized additive models (GAMs) or MaxEnt models. Additionally, provided is a script to validate the SDMs on their native fitting range using 10-fold cross-validation, and to fit the transferability model, as a linear mixed model (LMM), with a provided cleaned data.frame. The dataset provided includes a full species list with GBIF occurrence records, target-group background (TGB) records to use with model fitting and validation, as well as environmental data associated with the sightings.</p> </div>
Soil chemical variables improve models of understory plant species distributions
<div class="page"> <div class="section"> <div class="layoutArea"> <div class="column"><strong>Aim</strong></div> <div class="column">To determine the importance of soil variables relative to more commonly used topo-climatic or remotely sensed variables in species distribution models (SDMs) for understory plants.</div> <div class="column"> </div> <div class="column"><strong>Location</strong></div> <div class="column">White Mountain National Forest, New Hampshire, U.S.A.</div> <div class="column"> </div> <div class="column"><strong>Methods</strong></div> <div class="column">We fit models for presence of 41 forest understory plant species across 158 plots using soil, topographic, and spectral predictors to determine the relative contribution of different predictor types. We determined (a) if the potential importance of soil variables is greater than generally described in SDM literature, (b) which predictors are most important, and (c) if a standard subset of predictors can be used to effectively model all species.</div> <div class="column"> </div> <div class="column"><strong>Results</strong></div> <div class="column">Models containing all three predictor types performed best. Soil and topographic variables had comparable importance; spectral variables were of lesser importance. The best predictor variable was B horizon carbon to nitrogen ratio (B C:N), followed by topographic position index, elevation, and B horizon exchangeable calcium (B Ca). No standard subset effectively modeled all species.</div> <div class="column"> </div> <div class="column"><strong>Main conclusions</strong></div> <div class="column"> Our results and those of other SDMs that include in-situ soil geochemical data suggest that soil variables are increasingly important with more detailed descriptions of soils. Soil fertility data, such as B C:N and B Ca, are particularly important in acidic, forest soils where pH is a poor indicator of fertility. Commonly used topo-climatic variables provide meaningful predictions but are limited by their use of indirect predictor variables, inhibiting transferability and interpretability. The poor performance of models created using standard subsets of variables highlights the uniqueness of each species' niche and the need to combine flexible model building techniques with a variety of predictor variables.</div> </div> </div> </div>
Positional errors in species distribution modelling are not overcome by the coarser grains of analysis
<p>The performance of species distribution models is known to be affected by the analysis grain and the positional error of species occurrences. Coarsening of the spatial analysis grain has been suggested to compensate for positional errors. Nevertheless, this way of dealing with positional errors has never been thoroughly tested. With increasing use of fine-scale environmental data in predictive models developed for conservation and climate change studies it is increasingly important to test this assumption. Species distribution models using fine-scale environmental data are more likely to be negatively affected by positional error as the inaccurate species occurrences might easier end up in unsuitable environment, which can result in inappropriate conservation actions.</p> <p>Here, we examine the trade-offs between positional error and analysis grain and provide recommendations for best practice. We generated virtual species using tree canopy height, topography wetness index, and altitude derived from LiDAR point clouds at 5 x 5 m fine-resolution. We simulated the positional error in the range of 5 m to 99 m and evaluated the effects of several spatial grains in the range of 5 m to 500 m. In total, we assessed 49 combinations of positional accuracy and analysis grain. We used three common modelling techniques (MaxEnt, BRT and GLM) and four discrimination metrics to evaluate model performance (Sørensen index, overprediction and underprediction rate, AUC and TSS).</p> <p>We found that model performance decreased with increasing positional error in species occurrences and coarsening of the analysis grain. Most importantly, we showed that coarsening the analysis grain to compensate for positional error did not improve model performance. Our results reject coarsening of the analysis grain as a solution to address the negative effects of positional error on model performance.</p> <p>We recommend fitting models with the finest possible analysis grain (i.e., depending on data availablity) even when available species occurrences suffer from positional errors. If there are significant positional errors in species occurrence data, users are unlikely to benefit from making additional efforts to obtain higher resolution environmental data unless they also minimize the positional errors of species occurrences.</p>
Data from: Hindcast-validated species distribution models reveal future vulnerabilities of mangroves and salt marsh species
<p>Rapid climate change threatens biodiversity via habitat loss, range shifts, increases in invasive species, novel species interactions, and other unforeseen changes. Coastal and estuarine species are especially vulnerable to the impacts of climate change due to sea level rise and may be severely impacted in the next several decades. Species distribution modeling can project the potential future distributions of species under scenarios of climate change using bioclimatic data and georeferenced occurrence data. However, models projecting suitable habitat into the future are impossible to ground truth. One solution is to develop species distribution models for the present and project them to periods in the recent past where distributions are known to test model performance before making projections into the future. Here, we develop models using abiotic environmental variables to quantify the current suitable habitat available to eight Neotropical coastal species: four mangrove species and four salt marsh species. Using a novel model validation approach that leverages newly available monthly climatic data from 1960-2018, we project these niche models into two time periods in the recent past (i.e., within the past half-century) when either mangrove or salt marsh dominance was documented via other data sources. Models were hindcast-validated and then used to project the suitable habitat of all species at four time periods in the future under a model of climate change. For all future time periods, the projected suitable habitat of mangrove species decreased, and suitable habitat declined more severely in salt marsh species.</p>
Supporting data for article comparison and Uncertainty Analysis of Species Distribution Models
<p>Downloaded from Web of Science for the supporting data of article comparison and Uncertainty Analysis of Species Distribution Models.</p>
Dataset: Selecting tree species to restore forest under climate change conditions: complementing species distribution models with field experimentation
<p>This repository contains the files associated with the following article:</p> <p>Jesús Sandoval-Martínez, Ernesto I. Badano, Francisco A. Guerra-Coss, Jorge A. Flores Cano, Joel Flores, Sandra Milena Gelviz-Gelvez, Felipe Barragán-Torres, “Selecting tree species to restore forest under climate change conditions: complementing species distribution models with field experimentation”, submitted to <em>Journal of Environmental Management</em>.</p> <p><strong>Supplementary material 01 </strong>is a compressed file that contains two Microsoft Excel files with data that support the results of the study. A file correspond to <em>Vachellia pennatula</em> and the another file correspond to <em>Prosopis laevigata</em>. In both files, the first spreadsheet shows the occurrence data (latitude and longitude) used to calibrate the distribution model (SDM) of the corresponding species, the current values of the 19 bioclimatic variables associated with these coordinates and the Spearman correlation coefficients used to select the variables included in the SDM (selected variables are indicated in green). The second spreadsheet shows the current habitat occupancy probabilities of the target species estimated with the SDM at the geographic coordinates of occurrence points, while the table on the side shows the fraction of true presences dropping at the following probability categories: (1) habitat occupancy probabilities below 0.1 = unsuitable spatial units for the species, (2) habitat occupancy probabilities between 0.1 and 0.4 = barely suitable spatial units for the species, (3) habitat occupancy probabilities between 0.4 and 0.7 = moderately suitable spatial units for the species, and (4) habitat occupancy probabilities above 0.7 = highly suitable spatial units for the species. The third spreadsheet shows the one-thousand random geographic coordinates and the corresponding current and future habitat occupancy probabilities of each species. Future habitat occupancy probabilities are provided for three time periods (2041-2060, 2061-2080 and 2081-2100) at four radiative forcing levels each (2.6, 4.5, 7.0 and 8.5 W/m<sup>2</sup>).</p> <p><strong>Supplementary material 02 </strong>is a compressed file that contains a folder for <em>Vachellia pennatula</em> and another folder for <em>Prosopis laevigata</em>. Each of these folders contains the summaries of the MaxEnt outputs that support the results of the corresponding SDM.</p> <p><strong>Supplementary material 03 </strong>is a compressed Keyhole Markup Language file (KMZ) that contains interactive maps that are optimized for the desktop version of Google Earth. To accelerate visualization of maps, we recommend installing this software in a computer meeting the following requirements: CPU Intel Core i5 9<sup>th</sup> generation or higher, CPU clock speed 1.8 GHz or higher, random-access memory (RAM) 8 GB or higher, and video random access memory (VRAM) 1 GB or higher. Otherwise, opening this file may take several minutes. These maps are organized in a folder for <em>Vachellia pennatula</em> and another folder for <em>Prosopis laevigata</em>, which must be expanded for accessing the following information (click on the arrow on the left of folders to expand them):</p> <ul> <li><strong>Current climate </strong>– Activating this folder (click the fox on the left of the folder) display the map of habitat occupancy probabilities of species across Mexico under the current climate.</li> <li><strong>Period 2041-2060, 2061-2080 and 2081-2100 </strong>– Expanding each of these folders (click on the arrow on the left of folders) shows four subfolders that correspond to different radiative forcing levels (2.6, 4.5, 7.0 and 8.5 W/m<sup>2</sup>). Activating each of these sub folders (click the fox on the left of subfolders) display the map of habitat occupancy probabilities of species across Mexico expected on the corresponding time period and radiative forcing level. These maps also show the areas classified as climatically unsuitable in the multivariate environmental similarity surface (MESS) analysis. Clicking on the names of subfolders displays a figure showing the relationship between current and future habitat occupancy probabilities of the species on the corresponding time period and radiative forcing level. In these figures, the red line is the empirical relationship between these variables and the solid blue line is the theoretical relationship with intercept = 0 and slope = 1. The statistical results that support these relationships are also shown in these figures.</li> </ul> <p><strong>Supplementary material 04 </strong>is a compressed file that contains two Microsoft Excel files with data that support the results of the study. the file labeled as “Microclimate data” contains two spreadsheets, which correspond to the temperature and rainfall values measured in controls under the current climate and climate change simulation plots located of the field experiments. The file levelled as “Seedling emergence and survival” contains a spreadsheet for <em>Vachellia pennatula</em> and another one for <em>Prosopis laevigata</em>, which contains the data used to estimate the seedling emergence and survival rates in controls and climate change simulation plots.</p>
Data from: Climatically robust multi-scale species distribution models to support pronghorn recovery in California
<p>We combined two climate-based distribution models with three finer-scale suitability models to identify habitat for pronghorn recovery in California now and into the future.</p> <p>Location: California, United States </p> <p>Methods: We used a consensus approach to identify areas of suitable climate now (1980-2010) and future (2031-2060) for pronghorn in California. We compared the results of models from two separate hypotheses about their historical ecology in the state, specifically the migration hypothesis and the niche reduction hypothesis. We combined occurrences from GPS collars distributed across three populations of pronghorn in the state to create three distinct habitat models: (1) an ensemble model using Random Forests, Maxent, Classification and Regression Trees, and a Generalized Linear Model; (2) a step selection function; and (3) an expert-driven model. We evaluated consensus among both the climate models and the suitability models to prioritize areas for, and evaluate the prospects of, pronghorn recovery. </p> <p>Results: Climate suitability for pronghorn in the future depends heavily on model assumptions. Under the migration hypothesis, our model predicted that there will be on suitable climate in California in the future. Under the niche reduction hypothesis, by contrast, suitable climate will expand. Habitat also depended on the methods used, but areas of consensus among all three exist in large patches throughout the state.</p> <p>Main Conclusions: Identifying habitat for a species which has undergone extreme range collapse, and which has very fine scale habitat needs, presents novel challenges for spatial ecologists. Our multi-method, multi-hypothesis approach can allow habitat modelers to identify areas of consensus and, perhaps more importantly, critical knowledge gaps that could resolve disagreements among the models. For pronghorn, a better understanding of their upper thermal tolerances and whether historical populations migrated will be crucial to their potential recovery in California and throughout the arid Southwest.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.