HDBSCAN Clusters Maize Crop Stress - Adige River-Fed Downstream Irrigated Plain, 2022-2023
<div dir="ltr"> <table style="width: 99.5207%;"><colgroup><col style="width: 27.9117%;"><col style="width: 72.0629%;"></colgroup> <tbody> <tr> <td> <p dir="ltr"><strong>Demonstration Case Name</strong></p> </td> <td> <p dir="ltr">Multi-Hazards in the Downstream Area of the Adige River Basin.</p> </td> </tr> <tr> <td> <p dir="ltr"><strong>Dataset Name/Title</strong></p> </td> <td> <p dir="ltr">HDBSCAN Clusters Maize Crop Stress - Adige River-Fed Downstream Irrigated Plain, 2022-2023</p> </td> </tr> <tr> <td> <p dir="ltr"><strong>Dataset Description</strong></p> </td> <td> <p dir="ltr">The dataset contains HDBSCAN (Hierarchical Density-Based Spatial Clustering of Applications with Noise) clusters based on a synthetic stress indicator obtained through a PCA (Principal Component Analysis) on NDVI (Normalized Difference Vegetation Index) and NDMI (Normalized Difference Moisture Index) Sentinel-2 based indices. The dataset contains identified vegetation stress clusters on the following dates: 02/07/2022 17/07/2022, 22/07/2022, 01/08/2022, 06/08/2022, 11/08/2022, 16/08/2022, 07/07/2023, 17/07/2023, 27/07/2023, 06/08/2023, 11/08/2023, 16/08/2023, covering portion of 2022 and 2023 maize cropping seasons (July, August) in the selected area. All observations are located in a plain area that relies on the Adige River for cropland irrigation. For each row, the dataset contains the following columns: </p> <ul> <li> <p dir="ltr"><em>date</em>: date of the clustering, in dd/mm/YYYY format</p> </li> <li> <p dir="ltr"><em>cluster</em>: cluster number, -1 identifies the noise (unclassified) cluster</p> </li> <li> <p dir="ltr"><em>mean_ndvi</em>: mean NDVI across all maize fields in the cluster</p> </li> <li> <p dir="ltr"><em>mean_ndmi</em>: mean NDMI across all maize fields in the cluster</p> </li> <li> <p dir="ltr"><em>prevalent_hsg</em>: most frequent Hydrologic Soil Group </p> </li> <li> <p dir="ltr"><em>SPEI90</em>: Standardized Precipitation Evapotranspiration Index (SPEI) calculated on a 90 days time frame (see <em>Key Methodologies </em>for details)</p> </li> <li> <p dir="ltr"><em>SPEI180</em>: Standardized Precipitation Evapotranspiration Index (SPEI) calculated on a 180 days time frame (see <em>Key Methodologies </em>for details)</p> </li> <li> <p dir="ltr"><em>SPEI365</em>: Standardized Precipitation Evapotranspiration Index (SPEI) calculated on a 365 days time frame (see <em>Key Methodologies </em>for details)</p> </li> <li> <p dir="ltr"><em>temp_anom_X</em>: temperature anomaly (° C) X days before the considered date (see <em>Key Methodologies </em>for details)</p> </li> <li> <p dir="ltr"><em>SWI005_X</em>: Soil Water Index (%) with T-value = 5 (indicating the model water infiltration time) X days before the considered date</p> </li> <li> <p dir="ltr"><em>irr_channel_distance_m</em>: average distance (m) of maize fields in the cluster from the closest irrigation channel</p> </li> <li> <p dir="ltr"><em>geometry</em>: polygon geometry of the cluster in EPSG:32632</p> </li> </ul> </td> </tr> <tr> <td> <p dir="ltr"><strong>Key Methodologies</strong></p> </td> <td> <p dir="ltr">The 1981-2023 period was used as reference for the computation of the SPEI index, using the Hargreaves equation (Hargreaves, 1994) to estimate the daily potential evapotranspiration, with the extra-terrestrial radiation evaluated from the latitude and the day of the year. The gamma distribution was used for standardizing the water balance time series.</p> <p dir="ltr">Temperature anomalies were defined for each calendar day as the difference between the mean daily temperature and the 50th percentile of the 1991–2020 calendar day. Percentiles were computed using a centred 15‑day moving window to smooth short-term variability. </p> <p dir="ltr">Maize crop field-level NDVI and NDMI values were calculated by averaging pixel-level NDVI/NDMI derived from Sentinel-2 L2A observations within the irrigated districts fed by the Adige River waters (data courtesy of ANBI Veneto). The observations result from a multi-step data cleaning process. Images were cloud masked using the Sentinel-2 Scene Classification Layer (SCL). Crop field observations were excluded if more than 50% of their pixels were unavailable e.g., due to cloud cover; remaining observations were filtered using a Bare Soil Index (BSI) threshold of 0.08 (Mzid et al., 2021), to exclude non vegetated (bare soil) pixels. Finally, fields associated with same year alternating crops were disentangled by analyzing temporal BSI profiles to detect two green-up periods separated by at least one observation identified as bare soil (BSI > 0.08).</p> <p dir="ltr">A synthetic stress indicator was defined as the first principal component (PC1) derived from a one-component PCA applied to NDVI and NDMI values. HDBSCAN algorithm was used on the synthetic stress indicator to identify clusters following hyperparameters optimization via grid search.</p> <p dir="ltr"> </p> </td> </tr> <tr> <td> <p dir="ltr"><strong>Temporal Domain</strong></p> </td> <td> <p dir="ltr">2022-2023</p> </td> </tr> <tr> <td> <p dir="ltr"><strong>Spatial Domain</strong></p> </td> <td> <p dir="ltr">The dataset is provided over the [10.7, 45.0, 12.3, 45.6] spatial domain (min longitude, min latitude, max longitude, max latitude in WGS84, EPSG:4326).</p> </td> </tr> <tr> <td> <p dir="ltr"><strong>Key Variables/Indicators</strong></p> </td> <td> <p dir="ltr">Date, cluster, mean_ndvi, mean_ndmi, n_fields, prevalent_hsg, SPEI90, SPEI180, SPEI365, temp_anom_1, temp_anom_2, temp_anom_3, temp_anom_4, temp_anom_5, temp_anom_6, SWI005_1, SWI005_2, SWI005_3, irr_channel_distance_m, geometry</p> <p dir="ltr">See <em>Dataset Description</em> section for details on the variables.</p> </td> </tr> <tr> <td> <p dir="ltr"><strong>Data Format</strong></p> </td> <td> <p dir="ltr">csv</p> </td> </tr> <tr> <td> <p dir="ltr"><strong>Source Data</strong></p> </td> <td> <ul> <li> <p dir="ltr">NDVI, NDMI and BSI were obtained from ESA Copernicus Sentinel-2 L2A</p> </li> <li> <p dir="ltr">Crop field level information on farmer declared crop were obtained from the Veneto Agency for Payments in Agriculture (AVEPA, Agenzia Veneta per i Pagamenti, https://www.avepa.it/web/avepa)</p> </li> <li> <p dir="ltr">Hydrologic Soil Group at 1:250000 scale data were obtained from the Agenzia Regionale per la prevenzione e protezione ambientale Veneto (ARPAV, Agenzia Regionale per la Prevenzione e Protezione Ambientale del Veneto - Regional Agency for Environmental Prevention and Protection of Veneto, https://www.arpa.veneto.it/), and follow the classification scheme outlined in USDA National Engineering Handbook (USDA-NRCS, 2009)</p> </li> <li> <p dir="ltr">Irrigated districts were provided by courtesy of ANBI Veneto (Associazione Nazionale Bonifiche Irrigazioni - National Association of Land Reclamation and Irrigation)</p> </li> <li> <p dir="ltr">Temperature and precipitation data for SPEI and temperature anomaly were obtained from the SCIA-ISPRA dataset (ISPRA, <a href="https://scia.isprambiente.it/dati-e-indicatori/">https://scia.isprambiente.it/dati-e-indicatori/</a>)</p> </li> </ul> </td> </tr> <tr> <td> <p dir="ltr"><strong>Accessibility</strong></p> </td> <td> <p dir="ltr">https://doi.org/10.5281/zenodo.15301314</p> </td> </tr> <tr> <td> <p dir="ltr"><strong>Stakeholder Relevance</strong></p> </td> <td> <p dir="ltr">The dataset provides valuable post-disaster information on crop vegetation dynamics during hot and dry events. Impacted maize cultivated areas (clusters) across multiple dates during 2022 and 2023 cropping seasons are reported. The inclusion of two years, one characterized by severe drought (2022) and one by non-drought conditions (2023) provides insights into how the considered cropland area was affected under different abiotic stressor conditions. Additional information is provided by the inclusion of meteorological (SPEI, temperature anomaly), soil and territorial characteristic layers, thus enabling a more detailed analysis of the underlying impact drivers to support the understanding of the root causes of impacts. Moreover, the dataset provides a spatially and temporally explicit representation of maize stress clusters, highlighting areas that were more impacted during the course of the 2022 maize cropping season thus providing a valuable overview of the most vulnerable areas. This approach represents a promising tool to aid adaptation and management strategies, particularly related to water management (e.g., irrigation) and crop suitability under different meteorological and physical conditions.</p> </td> </tr> <tr> <td> <p dir="ltr"><strong>Limitations/Assumptions</strong></p> </td> <td> <p dir="ltr">In cases where a field was associated with more than one crop, a disaggregation technique was applied based on assumptions about crop growth phases. Additionally, NDVI and NDMI are not exclusively influenced by plant responses to drought, extreme heat and their combinations, as other factors (e.g., pests, soil/crop management) can affect plant vigour. However, the spatial extension and intensity of the 2022 drought event, the comparison with 2023 and the large number of fields considered potentially limit these uncertainties.</p> </td> </tr> <tr> <td> <p dir="ltr"><strong>Additional Outputs/Information</strong></p> </td> <td> <p dir="ltr">The dataset access is currently restricted due to pending related publication.</p> </td> </tr> <tr> <td> <p dir="ltr"><strong>Contact Information</strong></p> </td> <td> <p dir="ltr">Furlanetto, Jacopo (CMCC Foundation - Euro-Mediterranean Center on Climate Change, National Biodiversity Future Center) - Data curator</p> <p dir="ltr">Albergo, Edoardo (CMCC Foundation - Euro-Mediterranean Center on Climate Change, National Biodiversity Future Center) - Data curator</p> <p dir="ltr">Masina, Marinella (CMCC Foundation - Euro-Mediterranean Center on Climate Change)- Data curator</p> <p dir="ltr">Maraschini, Margherita (CMCC Foundation - Euro-Mediterranean Center on Climate Change) - Data curator</p> <p dir="ltr">Ferrario, Davide Mauro (CMCC Foundation - Euro-Mediterranean Center on Climate Change) - Data curator</p> <p dir="ltr">Torresan, Silvia (CMCC Foundation - Euro-Mediterranean Center on Climate Change, National Biodiversity Future Center) - Data manager</p> </td> </tr> </tbody> </table> </div>
ShareScore
20/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 4
- Harmonization
- 8
- Access
- 8
- Reuse readiness
- 0
- Engagement
- 0