Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

487

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

487 results for “species distribution models”

Learn how ShareScore rates datasets ↗
edi52/100

Species Distribution Modeling of Carnivorous Plants Worldwide

Forecasting how carnivorous plant species will respond to climatic change is a key issue in their conservation and management but presents a number of challenges. These challenges derive from interactions between the relatively simplistic statistical methods typically used to forecast species responses to climatic change, which to date have been limited mainly to species distribution models (“SDMs) and particular aspects of the ecology of carnivorous plants, including their rarity, habitat specialization, and limited dispersal ability. The small ranges and oftentimes low local abundance of carnivorous plants provide few occurrence records, which increase the potential for poorly or over-fitted SDMs and misspecification of relationships with their “optimal” environments. The unique habitats in which carnivorous plants often grow also are difficult to characterize using the basic temperature and precipitation data that often undergird SDMs. Rather, habitats in which carnivorous plants are common often are decoupled from broader climatic patterns (e.g., many retain high soil moisture even during seasonal drought) and may be associated with frequent disturbance. Last, dispersal limitation also may constrain range shifts of carnivorous plants as the climate changes. These three issues raise two related questions that are critical for understanding and forecasting the future of carnivorous plants. First, to what extent are current carnivorous plants distributions constrained by climate; and second, how readily, if at all, might carnivorous plants disperse to colonize new habitat as it becomes climatically suitable? We estimated the vulnerability of carnivorous plants to climatic change in light of challenges identified with SDMs in general and their particular application to these unique species. We combined two approaches: “ensembles of small models”, which attempt to deal with the challenges of fitting SDMs for data-limited species; and “bioclimatic velocity”, which is

openCC0Dec 2023View details →
zenodo48/100

Global taxonomic occurrence grids using GBIF data for species distribution models.

<p>To achieve large geographic coverage, species occurrence databases that are composed of ad hoc species data collections such as that provided by the Global Biodiversity Information Facility (GBIF) are often used. A drawback to using these data is their geographic sampling bias, in which some regions are more intensively sampled than others, while other areas have very little to none reported sampling effort. Uneven sampling effort can mislead conclusions about biodiversity patterns and species distributions (Gotelli &amp; Colwell, 2001; Lobo, 2008).</p> <p>Here we provide taxonomic occurrence grids to help mitigate the effects of sampling bias in species distribution modeling. These grids can be used to exclude areas of (a custom-defined) low sampling effort from the background when sampling for pseudo-absences&rsquo; (Phillips et al., 2009; Barbet-Massin et al.,2012). The occurrence grids have a 1 degree spatial resolution using WGS 84 as the geographic coordinate system. Each 1 degree grid cell contains the number of records present in GBIF corresponding to a specific taxonomic group: plants, mammals, reptiles, amphibians, birds and molluscs.</p> <p>To construct the occurrence grids, we used the 1- by 1-degree world latitude and longitude vector grid provided by ESRI (Redlands, California). It has a custom license which permits it reuse as long as ESRI is cited. It was downloaded from : <a href="https://www.arcgis.com/home/item.html?id=f11bcdc5d484400fa926dcce68de3df7">https://www.arcgis.com/home/item.html?id=f11bcdc5d484400fa926dcce68de3df7</a></p> <p>To map spatial sampling effort, the number of georeferenced occurrences corresponding to each taxonomic group contained by each 1- by 1-degree grid cell were counted. The grids were then converted to GeoTIFFs. The raster values correspond to the number of occurrences reported for the grid cells. For the purposes of the <a href="https://osf.io/7dpgr/">TrIAS project</a>, grid cells with fewer than 5 occurrences were removed. The TrIAS taxonomic occurrence grids are used as inputs to the TrIAS risk modelling and mapping workflow: https://github.com/trias-project/risk-modelling-and-mapping. Full (with all grid cells containing at least one occurrence) taxonomic occurrence grids are also provided.</p> <p>GBIF data for each taxonomic group were downloaded using the following criteria: &ldquo;Basis of Record&rdquo;: Observation, Machine Observation, Human Observation, Specimen, Material sample, Literature Occurrence, Unknown evidence., &quot;HasCoordinate is true&quot;, &quot;HasGeospatialIssue is false&quot;, &quot;TaxonKey is Amphibia&quot;, &quot;Year 1975-2005&quot;.</p> <p><strong>Raster Attributes</strong></p> <table> <tbody> <tr> <td> <p>Attribute</p> </td> <td> <p>Description</p> </td> </tr> <tr> <td> <p>OID</p> </td> <td> <p>numeric row ID</p> </td> </tr> <tr> <td> <p>Value</p> </td> <td> <p>the number of records contained in the grid cell</p> </td> </tr> <tr> <td> <p>Count</p> </td> <td> <p>the number of times the value appears in the raster</p> </td> </tr> </tbody> </table> <p>&nbsp;</p> <p>&nbsp;</p> <p>The extent of each taxonomic occurrence grid:</p> <ul> <li> <p>longitude -180.0; latitude -90.0 (southwest corner)</p> </li> <li> <p>longitude 180.0; latitude 90.0 (northeast corner)</p> </li> </ul> <p>&nbsp;</p> <p><strong>Files:</strong></p> <p>TrIAS taxonomic occurrence grids</p> <p>amphib_1deg_min5.tif</p> <p>birds_1deg_min5.tif</p> <p>mammals_1deg_min5.tif</p> <p>molluscs_1deg_min5.tif</p> <p>reptiles_1deg_min5.tif</p> <p>&nbsp;</p> <p>Raw taxonomic occurrence grids</p> <p>amphib_1deg_grid.tif</p> <p>birds_1deg_grid.tif</p> <p>mammals_1deg_grid.tif</p> <p>molluscs_1deg_grid.tif</p> <p>reptiles_1deg_grid.tif</p> <p><br> &nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Jan 2023View details →
zenodo44/100

Remote sensing based species distribution modelling based on GLCM and vegetation fractions for the city of Leipzig

<p>Modelling dataset and fractional vegetation cover dataset used in the study &quot;Earth observation based indication for avian species distribution models using the spectral trait concept and machine learning in an urban setting&quot; Wellmann et al. 2020.</p> <p>&nbsp;</p> <p>Reference:</p> <p></p> <p>Wellmann, T., Lausch, A., Scheuer, S., &amp; Haase, D. (2020). Earth observation based indication for avian species distribution models using the spectral trait concept and machine learning in an urban setting. <em>Ecological Indicators</em>, <em>111</em>(April 2020), 106029. https://doi.org/10.1016/j.ecolind.2019.106029</p> <p></p>

opencc-by-4.0Jan 2020View details →
zenodo44/100

Occurrence data used to create species distribution models and apply an evaluation method

<p>These two files containing&nbsp;a table with three columns: species names, longitude, latitude. Each row of the tables represents a georeferenced presence record for the corresponding species. The original presence data were downloaded from the GBIF database and after going through a cleaning process, we ended with these records that passed all the tests.</p> <p>These datasets were used to create species distribution models (SDMs) that were then used to apply a new method to evaluate the performance of different SDMs. Jim&eacute;nez &amp; Sober&oacute;n (2020)</p>

opencc-by-4.0Aug 2020View details →
zenodo44/100

Background data 'Effect of biotic dependencies in species distribution models: The future distribution of Thymallus thymallus under consideration of Allogamus auricollis'

<p>Background data of the paper 'Effect of biotic dependencies in species distribution models: The future distribution of Thymallus thymallus under consideration of Allogamus auricollis'</p>

opencc-by-nd-4.0May 2017View details →
zenodo44/100

Presence-Absence Points for Tree Species Distribution Modelling for Europe

<p>The dataset is a collection of presence and absence points for forest tree species for Europe. Each unique combination of longitude, latitude and year was considered as an independent sample. Presence data was obtained from the harmonized tree species occurrence dataset by <a href="https://zenodo.org/record/5524611">Heisig and Hengl (2020)</a> and absence data from the <a href="https://ec.europa.eu/eurostat/web/lucas">LUCAS</a> (in-situ source) dataset.</p> <p>A set of <strong>50</strong> different forest tree species was selected from the harmonized tree species dataset and data lacking a temporal observation was overlaid with yearly forest masks derived from land cover maps produced by <a href="https://zenodo.org/record/4725429">Parente et al. (2021)</a>. We overlaid the points with the probability maps for the classes:</p> <ul> <li>311: Broad-leaved forest,</li> <li>312: Coniferous forest,</li> <li>313: Mixed forest,</li> <li>323: Sclerophyllous forest,</li> <li>324: Transitional woodland-shrub,</li> <li>333: Sparsely vegetated area.</li> </ul> <p>Points were included in the dataset only if the probability value extracted for at least one of the above classes was <strong>&ge; 50%</strong> for all the years considered. An additional quality flag was added to distinguish points coming from this operation and the points with original year of observation coming from source datasets.</p> <p>The final dataset contains <strong>4,359,999</strong> observations for and a total of <strong>630 </strong>columns.&nbsp;<br> <br> The first <strong>8 </strong>columns of the dataset contain metadata information used to uniquely identify the points:</p> <ul> <li><strong>id</strong>: unique point identifier,</li> <li><strong>year</strong>: year of observation,</li> <li><strong>postprocess</strong>: quality flag to identify if the temporal reference of an observation comes from the original dataset or is the result of spatiotemporal overlay with forest masks,</li> <li><strong>Tile_ID</strong>: contains the tile id from the eu_tiling_system (30 km grid),</li> <li><strong>easting</strong>: longitude coordinates in Coordinate Reference System ETRS89 / LAEA Europe (= EPSG code 3035),</li> <li><strong>northing</strong>: latitude coordinates in Coordinate Reference System ETRS89 / LAEA Europe (= EPSG code 3035),</li> <li><strong>Atlas_class</strong>: name of the tree species according to the European Atlas of Forest Tree Species or NULL in case of absence point,</li> <li><strong>lc1</strong>: contains original LUCAS land cover class or NULL if it&#39;s a presence point.</li> </ul> <p>The remaining columns contain the extracted values of a series of predictor variables (temperature, precipitation, elevation, topographical information, spectral reflectance) useful for species distribution modeling applications. These points were used to model the potential and realized distribution of a series of <strong>16 target species </strong>for the period 2000 - 2020. The approach involved training three ML models to predict probability of presence (<em>i.e.</em> <a href="http://link.springer.com/article/10.1023/A:1010933404324">Random Forest</a>,&nbsp;<a href="http://dl.acm.org/doi/abs/10.1145/2939672.2939785">XGBoost</a>, <a href="https://rss.onlinelibrary.wiley.com/doi/abs/10.2307/2344614">GLM</a>), which served as input to train a linear meta-model (<em>i.e.</em> <a href="http://papers.nips.cc/paper/2014/file/ede7e2b6d13a41ddf9f4bdef84fdc737-Paper.pdf">Logistic regression classifier</a>), responsible for predicting the final probability of presence for each species.</p> <p>The <em>RDS </em>file is created from a data.table object and suitable for fast reading in the R-programming environment. The <em>CSV.GZ</em> file contains records as a table with easting and northing in Coordinate Reference System ETRS89 / LAEA Europe (= EPSG code 3035) and can be fed in a GIS after being unzipped.</p> <p>We provide <em>RDS </em>files for a 30km tile as an example containing raster stacks at 30m resolution of all the covariates included in the regression matrix. You can find the specific geographical location of the tile in Europe using the attached <em>GeoPackage&nbsp;</em>(&quot;eu_tiling_system_30km&quot;): open it in QGIS and filter by &quot;ID&quot;.</p> <p>In our approach we considered both static and dynamic covariates: dynamic covariates are calculated as averages of a 4 years time window (example: 2004 contains averages from 2002 to 2006). To get the predictions for a specific year, covariates contained in the <em>static</em> RDS file need to be bound with the respective year.</p> <p>To access our predictions (probabilities and uncertainties) produced for the target species access:</p> <ul> <li><strong>Open Data Science Europe viewer: <a href="https://maps.opendatascience.eu">https://maps.opendatascience.eu</a></strong></li> <li>Check the <strong>Related identifiers </strong>section of this repository to access each species individually</li> </ul> <p>If you instead would like to know more about the creation of this dataset and the modeling:</p> <ul> <li><strong>watch</strong> the talk at Open Data Science Workshop 2021 (<a href="https://doi.org/10.5446/55256">TIB AV-PORTAL</a>)</li> <li><strong>access </strong>the repository with our R/Python scripts and follow the instructions (<a href="https://gitlab.com/geoharmonizer_inea/spatial-layers/-/tree/master/veg_tree.species_anv.pnv.eml">GitLab</a>)</li> </ul> <p>A publication describing, in detail, all processing steps, accuracy assessment and general analysis of species distribution maps is available on <a href="https://doi.org/10.7717/peerj.13728">PeerJ</a>. To suggest any improvement/fix&nbsp;use&nbsp;<a href="https://gitlab.com/geoharmonizer_inea/spatial-layers/-/issues">https://gitlab.com/geoharmonizer_inea/spatial-layers/-/issues</a>.</p>

opencc-by-4.0Dec 2021View details →
zenodo44/100

Landuse/Landcover predictors for invasive species distribution modelling in Europe.

<p><strong>Description</strong></p> <p>This data set contains a set of predictors characterizing land use/land cover derived from the CORINE dataset, anthropogenic pressure from the global terrestrial human footprint dataset, and&nbsp;the distance to&nbsp; the nearest waterbody, for continental Europe. All have been aligned with the 1 km<sup>2</sup>&nbsp;EEA Reference Grid. The climate variables based on historical (1976-2005) and future (2040-2070) scenarios are available from De Troch et al., 2020 also via Zenodo. These rasters represent the habitat and anthropogenic predictors needed in the Tracking Invasive Alien Species (TrIAS) workflow for invasive species distribution modelling (wiSDM).</p> <p><strong>Geographic coverage</strong></p> <p>Europe</p> <p><strong>Methods</strong></p> <p>Land use classes were extracted from&nbsp;the CORINE06 100 m GeoTiff downloaded from Copernicus. The percentage of each 1 km<sup>2</sup> EEA Reference Grid cell occupied by coniferous forest, deciduous forest, wetlands, grasslands and agriculture was calculated. Multiple land use sub-classes were aggregated for the following categories: agriculture,&nbsp;wetlands, grasslands (Table 1). &nbsp;These data layers have been processed in R to replace all NAs that are within the European landmass, with zeros to distinguish them from the ocean, which remain NA, as in the CORINE dataset. In this context, a zero reflects the absence of a given land cover attribute. &nbsp;</p> <p>The mean anthropogenic pressure per 1km<sup>2&nbsp;&nbsp;</sup>EEA Reference Grid cell was extracted from the global terrestrial human footprint dataset (Venter et al, 2016). Distance to the nearest waterbody within each 1km<sup>2</sup>&nbsp; EEA Reference Grid cell was calculated using the 2016 Surface Water Bodies shapefile available from the EEA (https://www.eea.europa.eu/data-and-maps/data/wise-wfd-spatial/surface-water-body).&nbsp;</p> <p>&nbsp;</p> <table> <tbody> <tr> <td>Land Use Class</td> <td>CORINE LABEL</td> </tr> <tr> <td>Agriculture</td> <td>Non-irrigated arable land (211),&nbsp; Rice fields (213),Vineyards (221),Fruit trees and berry plantations (222),Olive groves (223),Pastures (231),Annual crops associated with permanent crops (241),Complex cultivation patterns (242),Land principally occupied by agriculture, with significant areas of natural vegetation (243)</td> </tr> <tr> <td>&nbsp;</td> </tr> <tr> <td>&nbsp;</td> </tr> <tr> <td>Coniferous forest</td> <td>Coniferous forest (312)</td> </tr> <tr> <td>Deciduous forest</td> <td>Broad-leaved forest (311)</td> </tr> <tr> <td>Grassland</td> <td>Natural grasslands (321), Moors and heathland, (322) Sclerophyllous vegetation (323)</td> </tr> <tr> <td>Wetland</td> <td>Inland marshes (411), Peat bogs (412)</td> </tr> </tbody> </table> <p>Table 1. How the&nbsp;the original land use/land cover types as labelled in CORINE were combined (or not).</p> <p><strong>Files</strong></p> <p>distance2water_EEA_1km.tif &nbsp;(distance to nearest waterbody)</p> <p>ESM1000m.tif&nbsp; (mean anthropogenic pressure)</p> <p>corine_perAgriculture.tif</p> <p>corine_perWetland.tif</p> <p>corine_pergrass.tif</p> <p>corine_perdeciduous.tif</p> <p>corine_perConiferous.tif</p> <p>&nbsp;</p>

opencc-by-4.0Apr 2023View details →
zenodo44/100

Data from: Identifying priority areas for spatial management of mixed fisheries using ensemble of multi-species distribution models. Panzeri D. et al., 2023, Fish and Fisheries

<p>Panzeri D.<sup>1</sup>, Russo T., Arneri E., Carlucci R., Cossarini G., Isajlović I., Krstulović &Scaron;ifner S., Manfredi C., Masnadi F., Reale M., Scarcella G., Solidoro C., Spedicato M.T., Vrgoč N., W. Zupa, Libralato S<sup>2</sup>.</p> <p><sup>1&nbsp;</sup>dpanzeri@ogs.it<br> <sup>2&nbsp;</sup>slibralato@ogs.it</p> <p>Spatial fisheries management is widely used to reduce overfishing, rebuild stocks, and protect biodiversity. However, the&nbsp;effectiveness and optimization of spatial measures depend on accurately identifying ecologically meaningful areas, which can be difficult in mixed fisheries. To apply a method generally to a range of target species, we developed an ensemble of species distribution models (e-SDM) that combines general additive models, generalized linear mixed models, random forest, and gradient-boosting machine methods in a training and testing protocol. The e-SDM was used to integrate density indices from two scientific bottom trawl surveys with the geopositional data, relevant oceanographic variables from the three-dimensional physical-biogeochemical operational model, and fishing effort from the vessel monitoring system. The determined best distributions for juveniles and adults are used to determine hot spots of aggregation based on single or multiple target species. We applied e-SDM to juvenile and adult stages of 10 marine demersal species representing 60% of the total demersal landings in the central areas of the Mediterranean Sea. Using the e-SDM results, hot spots of aggregation and grounds potentially more selective were identified for each species and for the target species group of otter trawl and beam trawl fisheries. The results confirm the ecological appropriateness of existing fishery restriction areas and support the identification of locations for new spatial management measures.</p> <p>Data (csv)&nbsp;for Panzeri et al. 2023</p> <p>1.&nbsp;<a href="https://zenodo.org/api/files/0b1b7af4-6a3b-481d-8d5f-57cf02d20eaa/Ensemble_density_F%26F_D.Panzeri_et_al_2023.csv">Ensemble_density_F&amp;F_D.Panzeri_et_al_2023.csv: CSV file with density values&nbsp; (column pred) in terms of number of individuals (log N/km2) for each species (column sp) and life stage (column age) for each grid cell (X = longitude and Y = latitude).</a>&nbsp;</p> <p>2.&nbsp;<a href="https://zenodo.org/api/files/0b1b7af4-6a3b-481d-8d5f-57cf02d20eaa/Ensemble_density_F%26F_D.Panzeri_et_al_2023.csv">Getis_hotspot_F&amp;F_D.Panzeri_et_al_2023.csv: CSV file with Getis ord Gi* values (column Gi) derived from the previous file 1, developed for each species and life stage for each grid cell (X = longitude and Y = latitude).</a></p> <p>3.&nbsp;<a href="https://zenodo.org/api/files/0b1b7af4-6a3b-481d-8d5f-57cf02d20eaa/Ensemble_density_F%26F_D.Panzeri_et_al_2023.csv">Multispecies_HotSpot_F&amp;F_D.Panzeri_et_al_2023.csv: Frequency map expressed as the number of species for each grid cell (column freq) that has the hotspot (previous file 2) above the third quartile.</a></p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Sep 2023View details →
zenodo40/100

Bioclimatic data for species distribution modelling in the Amazon Basin

<p>In this dataset, bioclimatic data regarding the Amazon Basin, in the near of the cities of Manaus and Manacapuru are available. There are 11 environmental data variables, referring to temperature, atmospheric pressure, concentration of pollutants and aerosols, such as carbon monoxide, ozone, carbon dioxide, among others. These were collected by the G-159 Gulfstream aircraft during its two periods of operation (IOP1 and IOP2), available in the GOAmazon (Green Ocean Amazon) project&#39;s data repository. A spatial interpolation methodology (linear barycentric interpolation) was applied to each variable, in order to obtain a larger area of data. The species occurrence data were collected from the repositories of the ICMBio (Instituto Chico Mendes de Conserva&ccedil;&atilde;o da Biodiversidade) Portal da Biodiversidade and GBIF (Global Biodiversity Information Facility), referring to the same date and location of the environmental data.&nbsp;<br> &nbsp;</p>

opencc-by-4.0Nov 2020View details →
dryad40/100

Data from: Species distribution models of the Spotted Wing Drosophila (Drosophila suzukii, Diptera: Drosophilidae) in its native and invasive range reveal an ecological niche shift

<p>The Spotted Wing Drosophila (<em>Drosophila</em> <em>suzukii</em>) is native to Southeast Asia. Since its first detection in 2008 in Europe and North America, it has been a pest to the fruit production industry as it feeds and oviposits on ripening fruit. Here we aim to model the potential geographical distribution of <em>D. suzukii</em>. We performed an extensive literature review to map the current records. In total, 517 documented occurrences (96 native and 421 invasive) were identified spanning 52 countries. Next, we constructed three species distribution models (SDMs) based on occurrence records in: 1) the native range (SDMnative), 2) the invasive range in Europe (SDMEurope) and 3) a global model of all records (SDMglobal). The models aimed to investigate, whether this species will be able to occupy additional ecological niches beyond its native range and expand its current geographic distribution both globally and in Europe. The SDMs were generated using Maximum Entropy algorithms (Maxent) based on present occurrence records and bioclimatic variables (WorldClim). Predictions of habitat suitability vary greatly depending on the origins of occurrence records. According to all models, precipitation and low temperatures were key limiting factors for the distribution of <em>D. suzukii</em>, which suggests that this species requires a humid environment with mild winters in order to establish a permanent population in its invasive range. Several regions in the invasive range, not presently occupied by this species, were predicted highly suitable, especially in northern Europe, suggesting that <em>D. suzukii</em> is not occupying its full fundamental niche yet. Synthesis and applications. Based on these models of potential geographic distribution of the Spotted Wing Drosophila (<em>Drosophila</em> <em>suzukii</em>), we show a shift in the ecological niche in <em>D. suzukii</em> populations, emphasizing the importance of using presence and local environmental data. Further investigation regarding new occurrences is recommended to secure optimal pest management. Despite a continuing expansion, many countries still lack proper surveillance schemes, and we urge policymakers to initiate appropriate management programs.</p>

opencc-zeroDec 2017View details →
dryad40/100

Data from: Complementary strengths of spatially-explicit and multi-species distribution models

<p><span><span><span><span><span><span><span><span><span><span><span>         Species distribution models (SDMs) project the outcome of community assembly processes - dispersal, the abiotic environment, and biotic interactions - onto geographic space. Recent advances in SDMs account for these processes by simultaneously modeling the species that comprise a community in a multivariate statistical framework or by incorporating residual spatial autocorrelation in SDMs. However, the effects of combining both multivariate and spatially-explicit model structures on the ecological inferences and the predictive abilities of a model are largely unknown. We used data on eastern hemlock  (<i>Tsuga canadensis</i>L.) and five additional co-occurring overstory tree species in 35,569 forest stands across Michigan, USA to evaluate how the choice of model structure, including spatial and non-spatial forms of univariate and multivariate models, affects ecological inference about the processes that shape community composition as well as model predictive ability.</span></span></span></span></span></span></span></span></span></span></span></p> <p><span><span><span><span><span><span><span><span><span><span>            Incorporating residual spatial autocorrelation via spatial random effects did not improve out-of-sample prediction for the six tree species, although in-sample model fit was higher in the spatial models. Spatial models attributed less variation in occurrence probability to environmental covariates than the non-spatial models for all six tree species, and estimated higher (more positive) residual co-occurrence values for most species pairs. The non-spatial multivariate model was better suited for evaluating habitat suitability and hypotheses about the processes that shape community composition.  Environmental correlations and residual correlations among species pairs were positively related, perhaps indicating that residual correlations were due to shared responses to unmeasured environmental covariates. This work highlights the importance of choosing a non-spatial model formulation to address research questions about the species-environment relationship or residual co-occurrence patterns, and a spatial model formulation when within-sample prediction accuracy is the main goal.</span></span></span></span></span></span></span></span></span></span></p>

opencc-zeroDec 2019View details →
dryad40/100

Code and data for Bayesian joint species distribution model selection for community-level prediction

<p>Code and data for reproducing the analysis in the manuscript "Bayesian joint species distribution model selection for community-level prediction."  Provided data include percent cover observations for 39 modeled vascular plant species within boreal forest understory communities and environmental model covariates. R code is provided to generate model inputs, apply alternative models, generate out-of-sample predictions, and calculate associated community and species log scores and alternative model evaluation metrics. Further, R source code is provided to implement the multinomial joint species distribution model defined in the manuscript. Details on the data, its processing, and the alternative model definitions and structure can be found in the main text of the manuscript.  Provided data are currently being used in ongoing analyses and coordination with authors may be warranted to avoid duplicate publication. Potential users are encouraged to consider collaboration with authors when useful and appropriate. Misinterpretation of data may occur if used outside the context of the original analysis. All data are made available in their current state. While significant efforts have been made to ensure data accuracy, complete accuracy cannot be guaranteed. Data may be updated periodically. It is the responsibility of the data user to check for updated versions of the data.</p>

opencc-zeroNov 2023View details →
dryad40/100

Data from: Integrated species distribution models to account for sampling biases and improve range wide occurrence predictions

<p><strong><span>Aim</span></strong></p> <p><span>Species distribution models (SDMs) that integrate presence-only and presence-absence data offer a promising avenue to improve information on species' geographic distributions. The use of such 'integrated SDMs' on a species range-wide extent has been constrained by the often-limited presence-absence data and by the heterogeneous sampling of the presence-only data. Here, we evaluate integrated SDMs for studying species ranges with a novel expert range map-based evaluation. We build a new understanding about how integrated SDMs address issues of estimation accuracy and data deficiency and thereby offer advantages over traditional SDMs.</span></p> <p><strong><span>Location</span></strong></p> <p><span>South and Central America.</span></p> <p><strong><span>Time period</span></strong></p> <p><span>1979-2017.</span></p> <p><strong><span>Major taxa studied</span></strong></p> <p><span>Hummingbirds.</span></p> <p><strong><span>Methods</span></strong></p> <p><span>We build integrated SDMs by linking two observation models – one for each data type – to the same underlying spatial process.</span> <span>We validate SDMs with two schemes: i) cross-validation with presence-absence data and ii) comparison with respect to the species' whole range as defined with IUCN range maps. We also compare models relative to the estimated response curves and compute the association between the benefit of the data integration and the number of presence records in each data set.</span></p> <p><strong><span>Results</span></strong></p> <p><span>The integrated SDM accounting for the spatially varying sampling intensity of the presence-only data was one of the top-performing models in both model validation schemes. Presence-only data alleviated overly large niche estimates, and data integration was beneficial compared to modelling solely presence-only data for species that had few presence points when predicting the species' whole range. On the community level, integrated models improved the species richness prediction.</span></p> <p><strong><span>Main conclusions</span></strong></p> <p><span>Integrated SDMs combining presence-only and presence-absence data are successfully able to borrow strengths from both data types and offer improved predictions of species' ranges. Integrated SDMs can potentially alleviate the impacts of taxonomically and geographically uneven sampling and to leverage the detailed sampling information in presence-absence data.</span></p>

opencc-zeroNov 2023View details →
dryad40/100

Resources for: Spatio-temporal integrated Bayesian species distribution models reveal lack of broad relationships between traits and range shifts

<p><strong>Aim</strong>: Climate change and habitat loss or degradation are some of the greatest threats that species face today, often resulting in range shifts. Species traits have been discussed as important predictors of range shifts, with the identification of general trends being of great interest for conservation efforts. However, studies reviewing relationships between traits and range shifts have questioned the existence of such generalized trends, due to mixed results and weak correlations, as well as analytical shortcomings. The aim of this study was to test this relationship empirically, using analytical approaches that account for common sources of bias when assessing range trends.<br><strong>Location</strong>: Tanzania, East Africa.<br><strong>Time period</strong>: 1980-1999 and 2000-2020.<br><strong>Major taxa studied</strong>: 57 savannah specialist birds found in Tanzania, belonging to 26 families and 11 orders.<br><strong>Methods</strong>: We applied recently developed integrated spatio-temporal species distribution models in R-INLA, combining citizen science and bird atlas data to estimate ranges of species, quantify range shifts, and test the predictive power of traditional trait groups, as well as exposure-related and sensitivity traits. We based our study on 40 years of bird observations in East African savannahs, a biome that has experienced increasing climatic and non-climatic pressures over recent decades. We correlated patterns of change with species traits.<br><strong>Results</strong>: We find indications of relationships identified by previous research, but low average explanatory power of traits from an ecological perspective, confirming the lack of meaningful general associations. However, our analysis finds compelling species-specific results.<br><strong>Main conclusions</strong>: We highlight the importance of individual assessments, while demonstrating the usefulness of our analytical approach for analyses of range shifts.</p>

opencc-zeroMar 2024View details →
zenodo40/100

Fig. 8. Heat map resulting from the Species Distribution Model using MaxEnt, where 1 in A revision of the genus Armillipora Quate (Diptera: Psychodidae) with the descriptions of two new species

Fig. 8. Heat map resulting from the Species Distribution Model using MaxEnt, where 1 is equal to the highest probability of distribution, while 0 is the lowest probability.

opencc-by-4.0Mar 2024View details →
dryad40/100

Improving distribution models of sparsely-documented disease vectors by incorporating information on related species via joint modeling

<p>A necessary component of understanding vector-borne disease risk is the accurate characterization of the distributions of their vectors. Species distribution models have been successfully applied to data-rich species but may produce inaccurate results for sparsely-documented vectors. In light of global change, vectors that are currently not well-documented could become increasingly important, requiring tools to predict their distributions. One way to achieve this could be to leverage data on related species to inform the distribution of a<strong> </strong>sparsely-documented vector based on the assumption that the environmental niches of related species are not independent. Relatedly, there is a natural dependence of the spatial distribution of a disease on the spatial dependence of its vector. Here, we propose to exploit these correlations by fitting a hierarchical model jointly to data on multiple vector species and their associated human diseases to improve distribution models of sparsely-documented species. To demonstrate this approach, we evaluated the ability of twelve models—which differed in their pooling of data from multiple vector species and inclusion of disease data—to improve distribution estimates of sparsely-documented vectors. We assessed our models on two simulated data sets, which allowed us to generalize our results and examine their mechanisms. We found that when the focal species is sparsely documented, incorporating data on related vector species reduces uncertainty and improves accuracy by reducing overfitting. When data on vector species are already incorporated, disease data only marginally improve model performance.  However, when data on other vectors are not available, disease data can improve model accuracy and reduce overfitting and uncertainty. We then assessed the approach on empirical data on ticks and tick-borne diseases in Florida and found that incorporating data on other vector species improved model performance. This study illustrates the value of exploiting correlated data via joint modeling to improve distribution models of data-limited species.</p>

opencc-zeroApr 2024View details →
zenodo40/100

Dataset for Integrated Species Distribution Model for pikeperch larvae in the Porvoo-Sipoo archipelago

<p>This record contains the data required to run the code for fitting the Integrated Species Distribution Model described in <a href="https://arxiv.org/abs/2206.08817">arXiv:2206.08817 [stat.ME].</a></p> <h1>Files in this record</h1> <ul> <li><strong>transect_data.csv</strong> Line transect observations from Porvoo-Sipoo archipelago, Finland on June 2017.</li> <li><strong>expert_assessments.tif</strong> Rasterized, anonymous expert assessments. Categorical values denoting how likely a given location is to be a spawning location for pikeperch. 4 categories, with smaller values corresponding to higher probabilities.</li> <li><strong>covariate_raster_example.tif</strong> Rasterized example environmental covariate values. These are similarly structured as the covariate data used in the study and compatible with the analysis code. However, since we do not have the permission to release the original data set, these values are instead generated based on the projected planar coordinates such that they have roughly similar spatial gradients as the original covariates.</li> </ul> <h1>Detailed descriptions</h1> <h2>Transect data</h2> <h3>Location and replicate identifiers</h3> <ul> <li> <p><strong>id</strong> : transect identifier. Replicates of the same transect have the same identifier.</p> </li> <li> <p><strong>id2</strong> : alternate transect identifier, unique for each transect.</p> </li> <li> <p><strong>repeated</strong> : whether transect was replicated or not.</p> </li> <li> <p><strong>X_euref</strong> : easting coordinate, EUREF_FIN_TM35FIN, for the transect starting location in [meters]</p> </li> <li> <p><strong>Y_euref</strong> : northing coordinate, EUREF_FIN_TM35FIN, for the transect starting location in [meters]</p> </li> <li><strong>date</strong> : date of the measurement, DD/MM/YYYY</li> <li><strong>week</strong> : week number of the measurement date</li> </ul> <h3>In situ measurements</h3> <ul> <li> <p><strong>volume</strong> : Transect water volume [m^3]. Transect length (500m) multiplied by sampler surface area. Used as survey effort.</p> </li> <li> <p><strong>heading</strong> : compass heading (direction) for the transect, in [degrees].</p> </li> <li> <p><strong>SumKUHA</strong> : total pikeperch (<em>Sander lucioperca</em>, kuha in Finnish) larvae count in each transect [scalar]</p> </li> </ul> <h2>Expert assessments</h2> <p>The raster contains assessments from 10 local experts encoded as separate raster layers (Expert_1, Expert_2, ..., Expert_10). Raster resolution is 50m x 50m and the planar coordinates are based on the same coordinate reference system as the transect observations (UTM zone 35).</p> <p>The assessments are coded as integers with values between 1 and 4, with smaller values corresponding to higher probabilities.</p> <h2>Covariate raster example</h2> <p>This raster has the same spatial dimensions and uses the same coordinate reference system as the expert assessment raster and has three layers, one for each covariate. The covariate values are generated based on the spatial coordinates such that each covariate has similar spatial gradient as the original covariate. The covarites have the same names as in the original covariate data (<strong>dptLUKE</strong>, <strong>dist10m</strong> and <strong>lined3km</strong>).</p> <h1>Creators</h1> <p>Transect data collected and curated by Sanna Kuningas.</p> <p>Original covariate rasters curated by Sanna Kuningas from data sets collected by the Finnish Environment Institute and the Natural Resources Institute Finland.</p> <p>Expert assessments originally digitized and rasterized by Jussi M&auml;kinen.&nbsp; Additional refinement to assessment rasters by Karel Kaurila.</p> <p>Preparation for publishing on Zenodo for all of the data sets&nbsp; by Karel Kaurila.</p> <h2>Change log</h2> <ul> <li>&nbsp;2025 Jan 31: Included columns <strong>date</strong> and&nbsp;<strong>week</strong> for <strong>transect_data.csv</strong>.</li> </ul>

opencc-by-4.0Nov 2024View details →
dryad40/100

Data from: Integrated SDM database: Enhancing the relevance and utility of species distribution models in conservation management

<p><span>1. Species' ranges are changing at accelerating rates. Species distribution models (SDMs) are powerful tools that help rangers and decision-makers prepare for reintroductions, range shifts, reductions, and/or expansions by predicting habitat suitability across landscapes. Yet, range-expanding or -shifting species in particular face other challenges that traditional SDM procedures cannot quantify, due to large differences between a species' currently-occupied range and potential future range. The realism of SDMs is thus lost and not as useful for conservation management in practice. Here, we address these challenges with an extended assessment of habitat suitability through an <i>integrated SDM database (iSDMdb)</i>.</span></p> <p><span>2. The<i> iSDMdb</i> is a spatial database of predicted sites in a species' prediction range, derived from SDM results, and is a single spatial feature that contains additional, user-friendly data fields that synthesise and summarise SDM predictions and uncertainty, human impacts, restoration features, novel preferences in novel spaces, and management priorities. To illustrate its utility<i>,</i> we used the endangered New Zealand sea lion (<i>Phocarctos hookeri</i>). We consulted with wildlife rangers, decision-makers, and sea lion experts to supplement SDM predictions with additional, more realistic, and applicable information for management. </span></p> <p><span>3. Almost half the data fields included in this database resulted from engaging with these end-users during our study. The SDM found 395 predicted sites. However, the <i>iSDMdb</i>'s additional assessments showed that the actual suitability of most sites (90%) was questionable due to human impacts. &gt;50% of sites contained unnatural barriers (fences, grazing grasslands), and 75% of sites had roads located within the species' range of inland movement. Just 5% of the predicted sites were mostly (&gt;80%) protected.</span></p> <p><span>4. Integrating SDM results with supplemental assessments provides a way to address SDM limitations, especially for range-expanding or -shifting species. SDM products for conservation applications have been critiqued for lacking transparency and interpretation support, and ineffectively communicating uncertainty. The <i>iSDMdb</i> addresses these issues and enhances the practical relevance and utility of SDMs for stakeholders, rangers, and decision-makers. We exemplify how to build an <i>iSDMdb</i> using open-source tools, and how to make diverse, complex assessments more accessible for end-users.</span></p>

opencc-zeroOct 2021View details →
dryad40/100

Using species distribution models and decision tools to direct surveys and identify potential translocation sites for a critically endangered species

<p>Aim: Occurrence records for cryptic species are typically limited or highly uncertain, leaving their distributions poorly resolved and hampering conservation. This can apply to well‐studied species, and increased survey effort and/or novel methods are required to improve distribution data. Here, we paired species distribution modelling (SDM) with decision tools to direct surveys for the critically endangered Leadbeater's possum (Gymnobelideus leadbeateri) outside its current restricted range. We also assessed survey areas for their suitability to host translocations.</p> <p>Location: Victoria, Australia.</p> <p>Method: We used both recent and historic records (now out of range and spatially uncertain) of Leadbeater's possum to build SDMs using MaxEnt. The SDMs informed an initial multi‐criteria decision analysis (MCDA) that enabled prioritization of 80 survey sites across seven forest patches (13–145 km outside the known range), which we surveyed using camera traps. Site and vegetation data were used in a post‐survey MCDA to rank their potential translocation suitability.</p> <p>Results: The SDM predictions were consistent with the species' ecology, identifying cold areas with high rainfall that had not recently burnt as suitable. The spatial uncertainty of records did not exert a strong influence on either model predictions or the ranking of patches for surveys. Camera trap surveys yielded records of 19 native species, with Leadbeater's possum detected in only one survey patch, 13 km outside of its previously known range. The post‐survey MCDA identified three forest patches as potentially suitable for conservation translocations, and these priorities were not sensitive to the decision criteria used.</p> <p>Main conclusions: The approach outlined here prioritized survey effort over a large area, resulting in detection of Leadbeater's possum in one new patch. The potential translocation sites identified could present an important risk‐spreading measure for the species given the threat posed by bushfire. Combining SDMs and decision tools can help target surveys and guide subsequent conservation strategies.</p>

opencc-zeroJan 2022View details →
zenodo40/100

Fig. 2 in Associations Between Habitat Quality And Body Size In The Carpathian-Podolian Land Snail Vestia Turgida: Species Distribution Model Selection And Assessment Of Performance

Fig. 2. Linear relationship (solid line) and 95 % confidence interval (gray area) between habitat quality predicted by the BART model (x-axis) and shell height (H in millimeters, y-axis), derived from the linear mixed model.

opencc-by-4.0Jan 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record