Presence-Absence Points for Tree Species Distribution Modelling for Europe
<p>The dataset is a collection of presence and absence points for forest tree species for Europe. Each unique combination of longitude, latitude and year was considered as an independent sample. Presence data was obtained from the harmonized tree species occurrence dataset by <a href="https://zenodo.org/record/5524611">Heisig and Hengl (2020)</a> and absence data from the <a href="https://ec.europa.eu/eurostat/web/lucas">LUCAS</a> (in-situ source) dataset.</p> <p>A set of <strong>50</strong> different forest tree species was selected from the harmonized tree species dataset and data lacking a temporal observation was overlaid with yearly forest masks derived from land cover maps produced by <a href="https://zenodo.org/record/4725429">Parente et al. (2021)</a>. We overlaid the points with the probability maps for the classes:</p> <ul> <li>311: Broad-leaved forest,</li> <li>312: Coniferous forest,</li> <li>313: Mixed forest,</li> <li>323: Sclerophyllous forest,</li> <li>324: Transitional woodland-shrub,</li> <li>333: Sparsely vegetated area.</li> </ul> <p>Points were included in the dataset only if the probability value extracted for at least one of the above classes was <strong>≥ 50%</strong> for all the years considered. An additional quality flag was added to distinguish points coming from this operation and the points with original year of observation coming from source datasets.</p> <p>The final dataset contains <strong>4,359,999</strong> observations for and a total of <strong>630 </strong>columns. <br> <br> The first <strong>8 </strong>columns of the dataset contain metadata information used to uniquely identify the points:</p> <ul> <li><strong>id</strong>: unique point identifier,</li> <li><strong>year</strong>: year of observation,</li> <li><strong>postprocess</strong>: quality flag to identify if the temporal reference of an observation comes from the original dataset or is the result of spatiotemporal overlay with forest masks,</li> <li><strong>Tile_ID</strong>: contains the tile id from the eu_tiling_system (30 km grid),</li> <li><strong>easting</strong>: longitude coordinates in Coordinate Reference System ETRS89 / LAEA Europe (= EPSG code 3035),</li> <li><strong>northing</strong>: latitude coordinates in Coordinate Reference System ETRS89 / LAEA Europe (= EPSG code 3035),</li> <li><strong>Atlas_class</strong>: name of the tree species according to the European Atlas of Forest Tree Species or NULL in case of absence point,</li> <li><strong>lc1</strong>: contains original LUCAS land cover class or NULL if it's a presence point.</li> </ul> <p>The remaining columns contain the extracted values of a series of predictor variables (temperature, precipitation, elevation, topographical information, spectral reflectance) useful for species distribution modeling applications. These points were used to model the potential and realized distribution of a series of <strong>16 target species </strong>for the period 2000 - 2020. The approach involved training three ML models to predict probability of presence (<em>i.e.</em> <a href="http://link.springer.com/article/10.1023/A:1010933404324">Random Forest</a>, <a href="http://dl.acm.org/doi/abs/10.1145/2939672.2939785">XGBoost</a>, <a href="https://rss.onlinelibrary.wiley.com/doi/abs/10.2307/2344614">GLM</a>), which served as input to train a linear meta-model (<em>i.e.</em> <a href="http://papers.nips.cc/paper/2014/file/ede7e2b6d13a41ddf9f4bdef84fdc737-Paper.pdf">Logistic regression classifier</a>), responsible for predicting the final probability of presence for each species.</p> <p>The <em>RDS </em>file is created from a data.table object and suitable for fast reading in the R-programming environment. The <em>CSV.GZ</em> file contains records as a table with easting and northing in Coordinate Reference System ETRS89 / LAEA Europe (= EPSG code 3035) and can be fed in a GIS after being unzipped.</p> <p>We provide <em>RDS </em>files for a 30km tile as an example containing raster stacks at 30m resolution of all the covariates included in the regression matrix. You can find the specific geographical location of the tile in Europe using the attached <em>GeoPackage </em>("eu_tiling_system_30km"): open it in QGIS and filter by "ID".</p> <p>In our approach we considered both static and dynamic covariates: dynamic covariates are calculated as averages of a 4 years time window (example: 2004 contains averages from 2002 to 2006). To get the predictions for a specific year, covariates contained in the <em>static</em> RDS file need to be bound with the respective year.</p> <p>To access our predictions (probabilities and uncertainties) produced for the target species access:</p> <ul> <li><strong>Open Data Science Europe viewer: <a href="https://maps.opendatascience.eu">https://maps.opendatascience.eu</a></strong></li> <li>Check the <strong>Related identifiers </strong>section of this repository to access each species individually</li> </ul> <p>If you instead would like to know more about the creation of this dataset and the modeling:</p> <ul> <li><strong>watch</strong> the talk at Open Data Science Workshop 2021 (<a href="https://doi.org/10.5446/55256">TIB AV-PORTAL</a>)</li> <li><strong>access </strong>the repository with our R/Python scripts and follow the instructions (<a href="https://gitlab.com/geoharmonizer_inea/spatial-layers/-/tree/master/veg_tree.species_anv.pnv.eml">GitLab</a>)</li> </ul> <p>A publication describing, in detail, all processing steps, accuracy assessment and general analysis of species distribution maps is available on <a href="https://doi.org/10.7717/peerj.13728">PeerJ</a>. To suggest any improvement/fix use <a href="https://gitlab.com/geoharmonizer_inea/spatial-layers/-/issues">https://gitlab.com/geoharmonizer_inea/spatial-layers/-/issues</a>.</p>
ShareScore
44/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 8
- Harmonization
- 4
- Access
- 20
- Reuse readiness
- 8
- Engagement
- 4