Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

35

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

35 results for “spatial statistics”

Learn how ShareScore rates datasets ↗
zenodo52/100

ICARIA: spatially distributed climate projections from statistical downscaling

<p><strong>ICARIA </strong>project had as one of its main purposes to develop coherent, reliable and usable downscaled climate projections from the last CMIP6 in order to construct the basis for efficient support to climate adaptation and decision-making of the related stakeholders, supporting the adaptation of critical assets within the project. These projections were obtained with also the purpose of being freely available for further use in subsequent studies and, hence, foster adaptation to climate change in more areas. Therefore, ICARIA&rsquo;s climate information is already based on CMIP6 models and incorporating in its workflow the current SSPs. The presented high-resolution future climate projections display a unique dataset, being obtained from a high-quality and high-density set of weather observations that are then interpolated to the case studies of interest in a <strong>100x100m resolution grid,&nbsp;</strong>which is the main outcome offered in this publication. These models will provide the scenarios to be considered within the Risk Assessment and the design and development of all adaptation measures coming as ICARIA outcomes.</p> <p>For further details, find here a brief of the <strong>methodology </strong>followed:<strong> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</strong></p> <p><strong>----- &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp;</strong></p> <p><em>The statistical downscaling methodology applied in ICARIA by FIC, named FICLIMA (Ribalaygua et al. 2013), consists of a two-step analogue/regression statistical method which has been used in national and international projects with good verification results (i.e.: Monjo et al. 2016). The first step is common for all simulated climate variables and it is based on an analogue stratification (Zorita et al. 1993). An analogue method was applied based on the hypothesis that &lsquo;analogue&rsquo; atmospheric patterns (predictors) should cause analogue local effects (predictands), which means that the number of days that were most similar to the day to be downscaled was selected. The similarity between any two days was measured according to three nested synoptic windows (with different weights) and four large-scale fields using a pseudo-Euclidean distance between the large-scale fields used as predictors. For each predictor, the weighted Euclidean distance was calculated and standardised by substituting it with the closest percentile of a reference population of weighted Euclidean distances for that predictor. This method is a good method for reproducing nonlinear relationships between predictors and the predictands, but it could not be used to simulate values outside of the range of observed values. In order to overcome this problem and obtain a better simulation, a second step was required.</em></p> <p><em>For this second step, the procedures applied depend on the variable of interest. To determine the temperature, multiple linear regression analysis for the selected number of most analogous days was performed for each station and for each problem day. From a group of potential predictors, the linear regression selected those with the highest correlation, using a forward and backward stepwise approach.</em></p> <p><em>For precipitation, a group of m problem days (we use the whole days of a month) is downscaled. For each problem day we obtain a &ldquo;preliminary precipitation amount&rdquo; averaging the rain amount of its n most analogous days, so we can sort the m problem days from the highest to the lowest &ldquo;preliminary precipitation amount&rdquo;. For assigning the final precipitation amount, all amounts of the m&times;n analogous days are sorted and clustered in m groups. Every quantity is finally assigned, orderly, to the m days previously sorted by the &ldquo;preliminary precipitation amount&rdquo;.</em></p> <p><em>For wind or relative humidity, the second step is a transfer function between the observed probability distribution and the simulated one using the averaged values from the n = 30 analogous days. Particularly, a parametric bias correction was performed to the time series obtained from the analogue stratification (first step). In order to estimate the improvement of this procedure, the bias correction was also applied to the direct model outputs.</em></p> <p><em>This second step done at a daily scale with an inner thorough verification procedure is essential and the main differentiating process of FICLIMA method. It extends beyond mean values to include extremes and covers all time scales, including daily intervals. With the verification it can be proven If the method correctly simulates changes from one day to the next, indicating an effective capture of the underlying physical connections between predictors and predictands. These physical links remain relatively consistent, even in the face of climate change (as opposed to purely empirical relationships that might shift). In essence, this approach theoretically addresses the primary challenge in statistical downscaling known as the non-stationarity problem. This problem questions the stability of predictor/predictand relationships established in the past, probing whether these relationships will persist in the future.</em></p> <p>-----</p> <p>The dataset shared here includes information for the three case studies tackled in ICARIA: <strong>Barcelona Metropolitan Area (AMB), Salzburg Region (SLZ), and South Aegean Region (SAR)</strong>. The information provided covers data and outcomes by 10 models belonging to CMIP6. Each model has a historical archive, from 01/01/1950 to 31/12/2014 and 4 future scenarios (ssp126, ssp245, ssp370 and ssp585) ranging from 01/01/2015 to 31/12/2100. The relation of the selected models is detailed in the next Table:</p> <p><strong>Table 1</strong>.<em> Information about the 10 climate models belonging to the 6 Coupled Model Intercomparison Project (CMIP6) corresponding to the IPCC AR6. Models were retrieved from the Earth System Grid Federation (ESGF) portal in support of the Program for Climate Model Diagnosis and Intercomparison (PCMDI).</em></p> <div> <div> <table> <tbody> <tr> <td> <p><strong>CMIP6 MODELS</strong></p> </td> <td> <p><strong>Resolution</strong></p> </td> <td> <p><strong>Responsible Centre</strong></p> </td> <td> <p><strong>References</strong></p> </td> </tr> <tr> <td> <p>ACCESS-CM2</p> </td> <td> <p>1,875&ordm; x 1,250&ordm;</p> </td> <td> <p>Australian Community Climate and Earth System Simulator (ACCESS), Australia</p> </td> <td> <p>Bi, D. et al (2020)</p> </td> </tr> <tr> <td> <p>BCC-CSM2-MR</p> </td> <td> <p>1,125&ordm; x 1,121&ordm;</p> </td> <td> <p>Beijing Climate Center (BCC), China Meteorological Administration, China.</p> </td> <td> <p>Wu T. et al. (2019)</p> </td> </tr> <tr> <td> <p>CanESM5</p> </td> <td> <p>2,812&ordm; x 2,790&ordm;</p> </td> <td> <p>Canadian Centre for Climate Modeling and Analysis (CC-CMA), Canad&aacute;.</p> </td> <td> <p>Swart, N.C. et al. (2019)</p> </td> </tr> <tr> <td> <p>CMCC-ESM2</p> </td> <td> <p>1,000&ordm; x 1,000&ordm;</p> </td> <td> <p>Centro Mediterraneo sui Cambiamenti Climatici (CMCC).</p> </td> <td> <p>Cherchi et al, 2018</p> </td> </tr> <tr> <td> <p>CNRM-ESM2-1</p> </td> <td> <p>1,406&ordm; x 1,401&ordm;</p> </td> <td> <p>CNRM (Centre National de Recherches Meteorologiques), Meteo-France, Francia.</p> </td> <td> <p>Seferian, R. (2019)</p> </td> </tr> <tr> <td> <p>EC-EARTH3</p> </td> <td> <p>0,703&ordm; x 0,702&ordm;</p> </td> <td> <p>EC-EARTH Consortium</p> </td> <td> <p>EC-Earth Consortium. (2019)</p> </td> </tr> <tr> <td> <p>MPI-ESM1-2-HR</p> </td> <td> <p>0,938&ordm; x 0,935&ordm;</p> </td> <td> <p>Max-Planck Institute for Meteorology (MPI-M), Germany.</p> </td> <td> <p>M&uuml;ller et al., (2018)</p> </td> </tr> <tr> <td> <p>MRI-ESM2-0</p> </td> <td> <p>1,125&ordm; x 1,121&ordm;</p> </td> <td> <p>Meteorological Research Institute (MRI), Japan.</p> </td> <td> <p>Yukimoto, S. et al. (2019)</p> </td> </tr> <tr> <td> <p>NorESM2-MM</p> </td> <td> <p>1,250&ordm; x 0,942&ordm;</p> </td> <td> <p>Norwegian Climate Centre (NCC), Norway.</p> </td> <td> <p>Bentsen, M. et al. (2019)</p> </td> </tr> <tr> <td> <p>UKESM1-0-LL</p> </td> <td> <p>1,875&ordm; x 1,250&ordm;</p> </td> <td> <p>UK Met Office, Hadley Centre, United Kingdom</p> </td> <td> <p>Good, P. et al. (2019)</p> </td> </tr> </tbody> </table> <p>The climate projections have been developed over each of the observational locations that were retrieved to run the statistical downscaling. The results from these projections have been&nbsp;<strong>spatially interpolated into a 100x100m grid with a Multi-lineal Regression Model</strong> considering diverse adjustments and topographic corrections. The results presented here are the<strong> median of the 10 models used, obtained for each of the 4 SSP</strong>s and each of the time periods considered in ICARIA until the year 2100. The variables treated belong to the main climate variables and their related extreme indicators as they were defined during the ICARIA project. You can find here a summary table of all the variables and indicators that were used to develop the projections.</p> <strong>Table 2.</strong> <em>Summary of selected thermal and precipitation indicators, grouped aligned with the main hazards they feed. &ldquo;nd&rdquo; = number of days; &ldquo;ne&rdquo; = number of events.</em> <div> <table> <tbody> <tr> <td> <p><strong>Index/name</strong></p> </td> <td> <p><strong>Short description</strong></p> </td> <td> <p><strong>Source</strong></p> </td> <td> <p><strong>Variable</strong></p> </td> <td> <p><strong>Units</strong></p> </td> <td> <p><strong>Threshold</strong></p> </td> </tr> <tr> <td> <p><strong>Thermal indicators</strong></p> </td> </tr> <tr> <td> <p>TX90 / TX10</p> </td> <td> <p>Warm/cold days</p> </td> <td> <p>Zhang et al. (2011)</p> </td> <td> <p>TX</p> </td> <td> <p>nd</p> </td> <td> <p>90 / 10%</p> </td> </tr> <tr> <td> <p>HD</p> </td> <td> <p>Heat day</p> </td> <td> <p>ICARIA</p> </td> <td> <p>TX</p> </td> <td> <p>nd</p> </td> <td> <p>&gt; 30 &deg;C</p> </td> </tr> <tr> <td> <p>EHD</p> </td> <td> <p>Extreme heat day</p> </td> <td> <p>ICARIA</p> </td> <td> <p>TX</p> </td> <td> <p>nd</p> </td> <td> <p>&gt; 35 &deg;C</p> </td> </tr> <tr> <td> <p>TR</p> </td> <td> <p>Tropical nights</p> </td> <td> <p>Zhang et al. (2011)</p> </td> <td> <p>TN</p> </td> <td> <p>nd</p> </td> <td> <p>&gt; 20 &deg;C</p> </td> </tr> <tr> <td> <p>EQ</p> </td> <td> <p>Equatorial nights</p> </td> <td> <p>AEMet 2020, ICARIA</p> </td> <td> <p>TN</p> </td> <td> <p>nd</p> </td> <td> <p>&gt; 25 &deg;C</p> </td> </tr> <tr> <td> <p>IN</p> </td> <td> <p>Infernal nights</p> </td> <td> <p>ICARIA</p> </td> <td> <p>TN</p> </td> <td> <p>nd</p> </td> <td> <p>&gt; 30 &deg;C</p> </td> </tr> <tr> <td> <p>FD</p> </td> <td> <p>Frost days</p> </td> <td> <p>Zhang et al. (2011)</p> </td> <td> <p>TN</p> </td> <td> <p>nd</p> </td> <td> <p>&lt; 0 &deg;C</p> </td> </tr> <tr> <td> <p>Max consec</p> </td> <td> <p>Max spell length for above thermal indicators</p> </td> <td> <p>ICARIA</p> </td> <td> <p>-</p> </td> <td> <p>nd</p> </td> <td> <p>-</p> </td> </tr> <tr> <td> <p>N&ordm; events</p> </td> <td> <p>Number of above thermal indicators events</p> </td> <td> <p>ICARIA</p> </td> <td> <p>-</p> </td> <td> <p>ne</p> </td> <td> <p>&gt; 3 days</p> </td> </tr> <tr> <td> <p>TXm</p> </td> <td> <p>Mean maximum temperatures</p> </td> <td> <p>ICARIA</p> </td> <td> <p>TX</p> </td> <td> <p>&deg;C</p> </td> <td> <p>-</p> </td> </tr> <tr> <td> <p>TNm</p> </td> <td> <p>Mean minimum temperatures</p> </td> <td> <p>ICARIA</p> </td> <td> <p>TN</p> </td> <td> <p>&deg;C</p> </td> <td> <p>-</p> </td> </tr> <tr> <td> <p>TM</p> </td> <td> <p>Mean temperatures</p> </td> <td> <p>ICARIA</p> </td> <td> <p>TA</p> </td> <td> <p>&deg;C</p> </td> <td> <p>-</p> </td> </tr> <tr> <td> <p>HWle</p> </td> <td> <p>Heatwave length</p> </td> <td> <p>ICARIA</p> </td> <td> <p>TX</p> </td> <td> <p>nd</p> </td> <td> <p>3d &gt; 95% TX</p> </td> </tr> <tr> <td> <p>HWim/HWix</p> </td> <td> <p>Mean and maximum heatwave intensity</p> </td> <td> <p>ICARIA</p> </td> <td> <p>TX</p> </td> <td> <p>&deg;C</p> </td> <td> <p>3d &gt; 95% TX</p> </td> </tr> <tr> <td> <p>HWf</p> </td> <td> <p>Heatwave frequency</p> </td> <td> <p>ICARIA</p> </td> <td> <p>TX</p> </td> <td> <p>ne</p> </td> <td> <p>3d &gt; 95% TX</p> </td> </tr> <tr> <td> <p>HWd</p> </td> <td> <p>Heatwave days</p> </td> <td> <p>ICARIA</p> </td> <td> <p>TX</p> </td> <td> <p>nd</p> </td> <td> <p>3d &gt; 95% TX</p> </td> </tr> <tr> <td> <p>HI - P90</p> </td> <td> <p>Heat Index (percentile 90)</p> </td> <td> <p>NWS (1994)</p> </td> <td> <p>TX, RH</p> </td> <td> <p>&deg;C</p> </td> <td> <p>TX&gt;27 &deg;C, HR&gt; 40%</p> </td> </tr> <tr> <td> <p>UTCI</p> </td> <td> <p>Universal Thermal Climate Index</p> </td> <td> <p>Br&ouml;de et al. (2012)</p> </td> <td> <p>TA<br>RH, W</p> </td> <td> <p>-</p> </td> <td> <p>-</p> </td> </tr> <tr> <td> <p>UHI</p> </td> <td> <p>Isla de calor (BCN) anual y estacional</p> </td> <td> <p>AMB, Metrobs 2015</p> </td> <td> <p>T</p> </td> <td> <p>&deg;C</p> </td> <td> <p>TM1-TM2 &gt; 0 &deg;C</p> </td> </tr> <tr> <td> <p><strong>Precipitation indicators</strong></p> </td> </tr> <tr> <td> <p>R20</p> </td> <td> <p>Number of heavy precipitation days</p> </td> <td> <p>Zhang et al. (2011)</p> </td> <td> <p>P</p> </td> <td> <p>nd</p> </td> <td> <p>&gt;20 mm</p> </td> </tr> <tr> <td> <p>R50, R100</p> </td> <td> <p>Days with extreme heavy rain</p> </td> <td> <p>AMB et al. (2017)</p> </td> <td> <p>P</p> </td> <td> <p>nd</p> </td> <td> <p>&gt;50mm</p> <p>&gt;100mm</p> </td> </tr> <tr> <td> <p>Ra</p> </td> <td> <p>Yearly and seasonal rainfall relative change</p> </td> <td> <p>ICARIA</p> </td> <td> <p>P</p> </td> <td> <p>mm</p> </td> <td> <p>&ge; 0.1mm</p> </td> </tr> <tr> <td> <p>IDF - CCF</p> </td> <td> <p>IDF Curves - Climate Change Factor</p> </td> <td> <p>Arnbjerg-Nielsen (2012)</p> </td> <td> <p>P</p> </td> <td> <p>-</p> </td> <td> <p>&ge; 0.1mm</p> </td> </tr> <tr> <td> <p><strong>Forest fire indicators</strong></p> </td> </tr> <tr> <td> <p>Mean FWI</p> </td> <td> <p>Mean Canadian FWI in fire season</p> </td> <td> <p>Stock, B.J. et al. (1989)</p> </td> <td> <p>RHn, TX, P, W</p> </td> <td> <p>.</p> </td> <td> <p>June-<br>September</p> </td> </tr> <tr> <td> <p>Very High FWI</p> </td> <td> <p>Very High Canadian FWI</p> </td> <td> <p>Stock, B.J. et al. (1989)</p> </td> <td> <p>RHn, TX, P, W</p> </td> <td> <p>nd</p> </td> <td> <p>FWI &gt; 38</p> </td> </tr> </tbody> </table> </div> <p>&nbsp;</p> <p><strong>Table 3</strong>. <em>Summary of selected drought, oceanic and wind indicators, grouped aligned with the main hazards they feed. &ldquo;nd&rdquo; = number of days; &ldquo;ne&rdquo; = number of events.</em></p> <div> <table> <tbody> <tr> <td> <p><strong>Index/name</strong></p> </td> <td> <p><strong>Short description</strong></p> </td> <td> <p><strong>Source</strong></p> </td> <td> <p><strong>Variable</strong></p> </td> <td> <p><strong>Units</strong></p> </td> <td> <p><strong>Threshold</strong></p> </td> </tr> <tr> <td> <p><strong>Drought indicators</strong></p> </td> </tr> <tr> <td> <p>CDDx</p> </td> <td> <p>Maximum dry spell duration</p> </td> <td> <p>Zhang et al. (2011)</p> </td> <td> <p>P</p> </td> <td> <p>nd</p> </td> <td> <p>&lt; 1 mm</p> </td> </tr> <tr> <td> <p>CDDm</p> </td> <td> <p>Mean dry spell duration</p> </td> <td> <p>Zhang et al. (2011)</p> </td> <td> <p>P</p> </td> <td> <p>nd</p> </td> <td> <p>&lt; 1 mm</p> </td> </tr> <tr> <td> <p>SPI</p> </td> <td> <p>SPI&nbsp;</p> <p>of 1, 3, 6, 12, 24 &amp; 36 months</p> </td> <td> <p>McKee et al. (1993)&nbsp;</p> </td> <td> <p>P, TA</p> </td> <td> <p>mm</p> </td> <td> <p>&ge; 0.1mm</p> </td> </tr> <tr> <td> <p>SPEI</p> </td> <td> <p>SPEI&nbsp;</p> <p>of 1, 3, 6, 12, 24 &amp; 36 months</p> </td> <td> <p>Vicente-Serrano et al.&nbsp; (2010)</p> </td> <td> <p>P, TA</p> </td> <td> <p>mm</p> </td> <td> <p>&ge; 0.1mm</p> </td> </tr> <tr> <td> <p><strong>Oceanic indicators</strong></p> </td> </tr> <tr> <td> <p>SS</p> </td> <td> <p>Storm surge</p> </td> <td> <p>Bryant et al. (2016)</p> </td> <td> <p>MT</p> </td> <td> <p>cm</p> </td> <td> <p>-</p> </td> </tr> <tr> <td> <p>OW</p> </td> <td> <p>Significant/maximum wave height</p> </td> <td> <p>ICARIA</p> </td> <td> <p>WH</p> </td> <td> <p>m</p> </td> <td> <p>-</p> </td> </tr> <tr> <td> <p>Wind indicators</p> </td> </tr> <tr> <td> <p>EWG</p> </td> <td> <p>Extreme wind gusts</p> </td> <td> <p>ICARIA</p> </td> <td> <p>W</p> </td> <td> <p>km/h</p> </td> <td> <p>-</p> </td> </tr> </tbody> </table> </div> <p>&nbsp;</p> </div> </div>

opencc-by-4.0Jul 2024View details →
zenodo44/100

Data to reproduce the results: Statistical power of spatial earthquake forecast tests

<p>We provide data needed to reproduce the figures from the publication titled &quot;Statistical power of spatial earthquake forecast tests&quot;.</p>

opencc-by-4.0Oct 2022View details →
zenodo40/100

Dataset for "Topography-based statistical modelling reveals high spatial variability and seasonal emission patches in forest floor methane flux"

<p>This dataset provides measured and upscaled forest floor methane (CH4) fluxes and soil moisture.</p> <p>This dataset is related to the following manuscript:</p> <p>Vainio et al., Topography-based statistical modelling reveals high spatial variability and seasonal emission patches in forest floor methane flux, Biogeosciences, in review. (The discussion preprint is available at https://doi.org/10.5194/bg-2020-263.)</p>

opencc-by-4.0Dec 2020View details →
zenodo40/100

Hyperparameter tuning and performance assessment of statistical and machine-learning models using spatial data.

<p>This is a research compendium (RC) for the publication &quot;Hyperparameter tuning and performance assessment of statistical and machine-learning algorithms using spatial data&quot;.</p> <p>The code (including figures, appendices and the manuscript) is packed in <strong>pathogen-modeling-3.zip&nbsp;</strong>or can be found directly in the <a href="https://github.com/pat-s/pathogen-modeling">Github repository</a>.</p> <ul> <li><strong>Publication figures</strong>:&nbsp;analysis/paper/submission/3/latex-source-files/</li> <li><strong>Appendices</strong>: analysis/paper/submission/3/</li> </ul> <p>This RC represents a static snapshot at the time of submission. The Github repository will receive changes after the publication was published.</p> <p><strong>Data sources</strong></p> <ul> <li>Atlas Climatico:&nbsp;<a href="http://opengis.uab.es/wms/iberia/index.htm">http://opengis.uab.es/wms/iberia/index.htm</a></li> <li>DEM:&nbsp;ftp://ftp.geo.euskadi.eus/lidar/MDE_LIDAR_2016_ETRS89/</li> <li>Lithology:&nbsp;<a href="http://www.geo.euskadi.eus/geonetwork/srv/spa/main.home">http://www.geo.euskadi.eus/geonetwork/srv/spa/main.home</a></li> <li>pH:&nbsp;<a href="https://esdac.jrc.ec.europa.eu/content/soil-ph-europe#tabs-0-description=0">https://esdac.jrc.ec.europa.eu/content/soil-ph-europe#tabs-0-description=0</a></li> <li>soil:&nbsp;<a href="https://www.isric.org/explore/soilgrids">https://www.isric.org/explore/soilgrids</a></li> </ul> <p><strong>Licenses</strong></p> <p>All files are shared via the given license with the exception of &quot;soil.tif&quot; which is shared via the&nbsp;<strong>ODbL </strong>license<strong>.</strong></p>

opencc-by-4.0Apr 2019View details →
zenodo40/100

Codes in R for spatial statistics analysis, ecological response models and spatial distribution models

<p>In the last decade, a plethora of algorithms have been developed for spatial ecology studies. In our case, we use some of these codes for underwater research work in applied ecology analysis of threatened endemic fishes and their natural habitat. For this, we developed codes in Rstudio&reg; script environment to run spatial and statistical analyses for ecological response and spatial distribution models (e.g., Hijmans &amp; Elith, 2017; Den Burg <em>et al.</em>, 2020). The employed R packages are as follows: caret (Kuhn et al., 2020), corrplot (Wei &amp; Simko, 2017), devtools (Wickham, 2015), dismo (Hijmans &amp; Elith, 2017), gbm (Freund &amp; Schapire, 1997; Friedman, 2002), ggplot2 (Wickham et al., 2019), lattice (Sarkar, 2008), lattice (Musa &amp; Mansor, 2021), maptools (Hijmans &amp; Elith, 2017), modelmetrics (Hvitfeldt &amp; Silge, 2021), pander (Wickham, 2015), plyr (Wickham &amp; Wickham, 2015), pROC (Robin et al., 2011), raster (Hijmans &amp; Elith, 2017), RColorBrewer (Neuwirth, 2014), Rcpp (Eddelbeuttel &amp; Balamura, 2018), rgdal (Verzani, 2011), sdm (Naimi &amp; Araujo, 2016), sf (e.g., Zainuddin, 2023), sp (Pebesma, 2020) and usethis (Gladstone, 2022).</p> <p>It is important to follow all the codes in order to obtain results from the ecological response and spatial distribution models. In particular, for the ecological scenario, we selected the Generalized Linear Model (GLM) and for the geographic scenario we selected DOMAIN, also known as Gower&#39;s metric (Carpenter <em>et al.</em>, 1993). We selected this regression method and this distance similarity metric because of its adequacy and robustness for studies with endemic or threatened species (<em>e.g.</em>, Naoki <em>et al.</em>, 2006). Next, we explain the statistical parameterization for the codes immersed in the GLM and DOMAIN running:</p> <p>In the first instance, we generated the background points and extracted the values of the variables (<a href="https://zenodo.org/api/files/fdd5446b-dee9-4b52-ad4f-cf556443d3dd/Code2_Extract_values_DWp_SC.R?versionId=c1ea0c61-53fe-4f95-ab88-0c1cb28399cb">Code2_Extract_values_DWp_SC.R</a>). Barbet-Massin <em>et al. </em>(2012) recommend the use of 10,000 background points when using regression methods (<em>e.g.</em>, Generalized Linear Model) or distance-based models (<em>e.g.</em>, DOMAIN). However, we considered important some factors such as the extent of the area and the type of study species for the correct selection of the number of points (Pers. Obs.).&nbsp; Then, we extracted the values of predictor variables (<em>e.g.</em>, bioclimatic, topographic, demographic, habitat) in function of presence and background points (<em>e.g.</em>, Hijmans and Elith, 2017).</p> <p>Subsequently, we subdivide both the presence and background point groups into 75% training data and 25% test data, each group, following the method of Sober&oacute;n &amp; Nakamura (2009) and Hijmans &amp; Elith (2017). For a training control, the 10-fold (cross-validation) method is selected, where the response variable presence is assigned as a factor. In case that some other variable would be important for the study species, it should also be assigned as a factor (Kim, 2009).</p> <p>After that, we ran the code for the GBM method (Gradient Boost Machine; <a href="https://zenodo.org/api/files/fdd5446b-dee9-4b52-ad4f-cf556443d3dd/Code3_GBM_Relative_contribution.R?versionId=1656bbae-66aa-409e-bb91-d8007dee8f95">Code3_GBM_Relative_contribution.R</a> and <a href="https://zenodo.org/api/files/fdd5446b-dee9-4b52-ad4f-cf556443d3dd/Code4_Relative_contribution.R?versionId=0e1d9352-e6b2-43da-984b-d6853a914258">Code4_Relative_contribution.R</a>), where we obtained the relative contribution of the variables used in the model. We parameterized the code with a Gaussian distribution and cross iteration of 5,000 repetitions (<em>e.g.</em>, Friedman, 2002; kim, 2009; Hijmans and Elith, 2017). In addition, we considered selecting a validation interval of 4 random training points (Personal test). The obtained plots were the partial dependence blocks, in function of each predictor variable.</p> <p>Subsequently, the correlation of the variables is run by Pearson&#39;s method (<a href="https://zenodo.org/api/files/fdd5446b-dee9-4b52-ad4f-cf556443d3dd/Code5_Pearson_Correlation.R?versionId=275f8dd4-b056-44d2-bfe5-f6264bc3298b">Code5_Pearson_Correlation.R</a>) to evaluate multicollinearity between variables (Guisan &amp; Hofer, 2003). It is recommended to consider a bivariate correlation &plusmn; 0.70 to discard highly correlated variables (<em>e.g.</em>, Awan <em>et al.</em>, 2021).</p> <p>Once the above codes were run, we uploaded the same subgroups (<em>i.e.</em>, presence and background groups with 75% training and 25% testing) (<a href="https://zenodo.org/api/files/fdd5446b-dee9-4b52-ad4f-cf556443d3dd/Code6_Presence&amp;backgrounds.R?versionId=d797b528-782f-4a19-bd61-cfb197f38513">Code6_Presence&amp;backgrounds.R</a>) for the GLM method code (<a href="https://zenodo.org/api/files/fdd5446b-dee9-4b52-ad4f-cf556443d3dd/Code7_GLM_model.R?versionId=e4aca276-d601-49ec-a62c-a9223b05a7ed">Code7_GLM_model.R</a>). Here, we first ran the GLM models per variable to obtain the <em>p</em>-significance value of each variable (alpha &le; 0.05); we selected the value one (<em>i.e.</em>, presence) as the likelihood factor. The generated models are of polynomial degree to obtain linear and quadratic response (<em>e.g.</em>, Fielding and Bell, 1997; Allouche <em>et al.</em>, 2006). From these results, we ran ecological response curve models, where the resulting plots included the probability of occurrence and values for continuous variables or categories for discrete variables. The points of the presence and background training group are also included.</p> <p>On the other hand, a global GLM was also run, from which the generalized model is evaluated by means of a 2 x 2 contingency matrix, including both observed and predicted records. A representation of this is shown in Table 1 (adapted from Allouche et al., 2006). In this process we select an arbitrary boundary of 0.5 to obtain better modeling performance and avoid high percentage of bias in type I (omission) or II (commission) errors (e.g., Carpenter et al., 1993; Fielding and Bell, 1997; Allouche et al., 2006; Kim, 2009; Hijmans and Elith, 2017).</p> <p>Table 1. Example of 2 x 2 contingency matrix for calculating performance metrics for GLM models. A represents true presence records (true positives), B represents false presence records (false positives - error of commission), C represents true background points (true negatives) and D represents false backgrounds (false negatives - errors of omission).</p> <table align="center"> <tbody> <tr> <td> <p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</p> </td> <td> <p>Validation set</p> </td> </tr> <tr> <td> <p>Model</p> </td> <td> <p>True</p> </td> <td> <p>False</p> </td> </tr> <tr> <td> <p>Presence</p> </td> <td> <p>A</p> </td> <td> <p>B</p> </td> </tr> <tr> <td> <p>Background</p> </td> <td> <p>C</p> </td> <td> <p>D</p> </td> </tr> </tbody> </table> <p>We then calculated the Overall and True Skill Statistics (TSS) metrics. The first is used to assess the proportion of correctly predicted cases, while the second metric assesses the prevalence of correctly predicted cases (Olden and Jackson, 2002). This metric also gives equal importance to the prevalence of presence prediction as to the random performance correction (Fielding and Bell, 1997; Allouche <em>et al.</em>, 2006).</p> <p>The last code (<em>i.e.</em>, <a href="https://zenodo.org/api/files/fdd5446b-dee9-4b52-ad4f-cf556443d3dd/Code8_DOMAIN_SuitHab_model.R?versionId=d951a8f2-d3a4-4804-b862-1b2762061876">Code8_DOMAIN_SuitHab_model.R</a>) is for species distribution modelling using the DOMAIN algorithm (Carpenter <em>et al.</em>, 1993). Here, we loaded the variable stack and the presence and background group subdivided into 75% training and 25% test, each. We only included the presence training subset and the predictor variables stack in the calculation of the DOMAIN metric, as well as in the evaluation and validation of the model.</p> <p>Regarding the model evaluation and estimation, we selected the following estimators:</p> <p>1) partial ROC, which evaluates the approach between the curves of positive (<em>i.e.</em>, correctly predicted presence) and negative (i.e., correctly predicted absence) cases. As farther apart these curves are, the model has a better prediction performance for the correct spatial distribution of the species (Manzanilla-Qui&ntilde;ones, 2020).</p> <p>2) ROC/AUC curve for model validation, where an optimal performance threshold is estimated to have an expected confidence of 75% to 99% probability (De Long <em>et al.</em>, 1988).</p>

opencc-by-4.0Feb 2023View details →
zenodo36/100

Supporting data for boundary layer water vapour statistics from high-spatial-resolution spaceborne imaging spectroscopy

<p>This dataset includes the properties necessary to reproduce the analysis of water vapour statistics derived from imaging spectroscopy as in:</p> <p>Richardson et al. (2021a) DOI: 10.5194/amt-14-5555-2021<br> Richardson et al. (2021b) DOI:&nbsp;10.5194/amt-2021-163 (pre-acceptance DOI, follow links to published version)</p> <p>Files include the retrieval emulator parameters, atmospheric profiles used in the emulator development, column-mean water vapour and cloud water both for the total column water vapour (TCWV) and &quot;effective&quot; TCWV, which accounts for the water vapour integrated along the direct solar path at a range of solar zenith angles, see Richardson 2021b, Eq. (7).</p>

opencc-by-4.0Nov 2021View details →
zenodo36/100

Replication data for: Spatial and temporal origins of the La Perouse low oxygen pool: A combined Lagrangian statistical approach

<p>This dataset contains:</p> <p>a) <strong>NEP36 </strong>daily Model (NEMO) Output from 20130228 till 20131005 in NetCDF format,</p> <p>b) Moving Vessel Profiler (<strong>MVP</strong>) Survey data gathered onboard the R/V Falkor during August 2013 in ASCII .mat files,</p> <p>c) Files required to initialize and run Lagrangian Particle tracking model <strong>ARIANE </strong>i.e. one mesh_mask file in netCDF format and one text file containing initial positions based on the Eulerian grid of the sliced NEP36 model</p> <p>d) the output files from running the particle tracking model ARIANE in NetCDF format</p>

opencc-by-4.0Oct 2021View details →
zenodo32/100

Video S1: Evolution of CO2 spatial statistical distribution.

<p>Spatio-temporal evolution of CO2 levels in an exam at room 4.1 at the ETSE, in University of Valencia</p>

opencc-by-4.0Jan 2021View details →
zenodo32/100

Understanding Citywalk Through Social Media: A Spatial-Statistical Analysis in Shanghai

<p>The uploaded file contains the original Citywalk social media posts, coordinates and classifications of different Citywalk POIs, sentiment scores, <span>Shannon Diversity Index</span> for different grids, public transport statistics, public transport accessibility statistics, green area statistics, and the Grid Citywalk Activity Index.</p>

opencc-by-4.0Nov 2024View details →
zenodo32/100

Supporting information for Spatial scales of the velocity shear layer and Kelvin-Helmholtz waves on the magnetopause: First statistical results

<p>Kelvin-Helmholtz (KH) waves observed by ARTEMIS P1 from 2015 to 2020. From the first column to the last column is: probe, date, time interval of KH waves, the location of KH waves in GSM coordinate, and the dominate period, phase speed, and wavelength of KH waves.</p>

opencc-by-4.0Jan 2022View details →
zenodo32/100

Application of GIS and Spatial Statistics in identifying the rental value of shop lot in a shopping mall based on location factor

<p>GIS Property Data</p>

opencc-by-4.0Mar 2022View details →
zenodo32/100

DeepCellMap: a Deep Learning Approach Coupled to Spatial Statistics to unravel Microglial Spatial Organization in the Developing Human Brain

<p>DeepCellMap is a deep-learning-assisted tool that integrates multi-scale image processing with advanced spatial and clustering statistics. This pipeline is designed to map microglial organization during normal and pathological brain development but can be adapted to any cell type.} Using DeepCellMap, we can capture the morphological diversity of microglia,&nbsp; identify strong coupling between proliferative and phagocytic phenotypes, and show that distinct spatial clusters rarely overlap as human brain development progresses. Additionally, we uncover a novel association between microglia and blood vessels in fetal brains exposed to maternal SARS-CoV-2. These findings offer insights into whether various microglial phenotypes form networks in the developing brain to occupy space, and in conditions involving haemorrhages, whether microglia respond to, or influence changes in blood vessel integrity. DeepCellMap is available as open-source software and is a powerful tool for extracting spatial statistics and analyzing cellular organization in large tissue sections, accommodating various imaging modalities. </p>

opencc-by-4.0Oct 2024View details →
dryad28/100

Data from: Spatially structured statistical network models for landscape genetics

A basic understanding of how the landscape impedes, or creates resistance to, the dispersal of organisms and hence gene flow is paramount for successful conservation science and management. Spatially structured ecological networks are often used to represent spatial landscape-genetic relationships, where nodes represent individuals or populations and resistance to movement is represented using non-binary edge weights. Weights are typically assigned or estimated by the user, rather than observed, and validating such weights is challenging. We provide a synthesis of current methods used to estimate edge weights and an overview of common model types, stressing the advantages and disadvantages of each approach and their ability to model landscape-genetic data. We further explore a set of spatial-statistical methods that provide ecologists with alternative approaches for modeling spatially explicit processes that may affect genetic structure. This includes an overview of spatial autoregressive models, with a particular focus on how correlation and partial correlation are used to represent neighborhood structure with the inverse of the covariance matrix (i.e., precision matrix). We then demonstrate how to model resistance by specifying an appropriate statistical model on the nodes, conditioned on the edge weights, through the precision matrix. This integration of network ecology and spatial statistics provides a practical analytical framework for landscape-genetic studies. The results can be used to make statistical inferences about the relative importance of individual landscape characteristics, such as the vegetative cover, hillslope, or the presence of roads or rivers, on gene flow. In addition, the R code we include allows readers to explore landscape-genetic structure in their own datasets, which will potentially provide new insights into the evolutionary processes that generated ecological networks, as well as valuable information about the optimal characteristics of conservation corridors.

opencc-zeroDec 2017View details →
zenodo28/100

Figure 8b from: Briz-Redón Á (2019) SpNetPrep: An R package using Shiny to facilitate spatial statistics on road networks. Research Ideas and Outcomes 5: e33521. https://doi.org/10.3897/rio.5.e33521

Figure 8b "Point Pattern Edition" features. - Information that is displayed (marks of the point pattern, if available, as defined by the user) when an event is clicked

opencc-by-4.0Feb 2019View details →
zenodo28/100

Figure 5a from: Briz-Redón Á (2019) SpNetPrep: An R package using Shiny to facilitate spatial statistics on road networks. Research Ideas and Outcomes 5: e33521. https://doi.org/10.3897/rio.5.e33521

Figure 5a Example of use of the SimplifyLinearNetwork function. - A road network introduced as input in which there is an excess of road segments and vertex

opencc-by-4.0Feb 2019View details →
zenodo28/100

Figure 1 from: Briz-Redón Á (2019) SpNetPrep: An R package using Shiny to facilitate spatial statistics on road networks. Research Ideas and Outcomes 5: e33521. https://doi.org/10.3897/rio.5.e33521

Figure 1 Workflow that describes all the steps that could be carried out in order to perform a spatial analysis on a point pattern that lies on a linear network. Some of these steps which lead to the final statistical analysis may be skipped but, at least, all of them should be considered. The blocks pointing the steps of the process include some of the R packages that would allow to successfully achieve each of them.

opencc-by-4.0Feb 2019View details →
zenodo28/100

Figure 3b from: Briz-Redón Á (2019) SpNetPrep: An R package using Shiny to facilitate spatial statistics on road networks. Research Ideas and Outcomes 5: e33521. https://doi.org/10.3897/rio.5.e33521

Figure 3b "Network Edition" example of use (I). - Network resulting from clicking on "Rebuild linear network" in the situation of a

opencc-by-4.0Feb 2019View details →
zenodo28/100

Figure 2a from: Briz-Redón Á (2019) SpNetPrep: An R package using Shiny to facilitate spatial statistics on road networks. Research Ideas and Outcomes 5: e33521. https://doi.org/10.3897/rio.5.e33521

Figure 2a "Network Edition" features. - Overview of the "Network Edition" section of the SpNetPrep application

opencc-by-4.0Feb 2019View details →
zenodo28/100

Figure 6b from: Briz-Redón Á (2019) SpNetPrep: An R package using Shiny to facilitate spatial statistics on road networks. Research Ideas and Outcomes 5: e33521. https://doi.org/10.3897/rio.5.e33521

Figure 6b "Network Direction" features. - Manual addition of traffic flow to the network by using the options "Add flow" and "Add long flow"

opencc-by-4.0Feb 2019View details →
zenodo28/100

Figure 3a from: Briz-Redón Á (2019) SpNetPrep: An R package using Shiny to facilitate spatial statistics on road networks. Research Ideas and Outcomes 5: e33521. https://doi.org/10.3897/rio.5.e33521

Figure 3a "Network Edition" example of use (I). - Use of the "Join vertex" (in green), "Remove edge" (in red) and "Add point" options (in green) in the SpNetPrep application

opencc-by-4.0Feb 2019View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record