Skip to main content
zenodoopen

A dataset of Earth Observation Data for Lithological Mapping using Machine Learning

<p><strong>Dataset Information</strong></p> <p>Machine Learning (ML) algorithms had successfully contributed in the creation of automated methods of recognizing patterns in high-dimensional data. Remote sensing data&nbsp; covers&nbsp; wide&nbsp; geographical areas and could be used to solve the problem of the demand of various&nbsp; in-situ data.&nbsp; Lithologicall mapping using remotely sensed data&nbsp; is one of the most challenging&nbsp; applications of ML algorithms. In the framework of the &ldquo;AI for Geoapplications&rdquo; project , ML and especially Deep Learning (DL) methodologies are investigated&nbsp; for&nbsp; the identification and characterization of the lithology based on remote sensing data in various&nbsp; pilot areas&nbsp; in Greece.&nbsp; In order to train and test the various ML algorithms, a dataset consisting of&nbsp; 30 ROIs selected&nbsp; mainly&nbsp; from low -vegetated areas,&nbsp; that cover 2% of the total&nbsp; area of Greece was created</p> <p><strong>Dataset Preprocessing</strong></p> <p>Dataset preprocessing was executed using a combination of SNAP, QGIS and ENVI tools.</p> <p>Preprocessing steps:</p> <p>Defining areas with the following properties:</p> <ul> <li> <p>Zero cloud and snow coverage</p> </li> <li> <p>No water bodies</p> </li> <li> <p>Minimum vegetation</p> </li> </ul> <p>For the Aster Images:</p> <ul> <li> <p>Subset on defined areas</p> </li> <li> <p>Mosaic images when needed</p> </li> <li> <p>Digitising clouds</p> </li> </ul> <p>For the Labels:</p> <ul> <li> <p>We got the Soil map from YPEN (<a href="https://ypen.gov.gr/">https://ypen.gov.gr/</a>)</p> </li> <li> <p>Subset on defined areas</p> </li> <li> <p>All categories are represented with good analogies</p> </li> <li> <p>Clip label files with digitised clouds</p> </li> <li> <p>Rasterize</p> </li> </ul> <p>&nbsp;</p> <p>For the Labels we have eighteen categories for the twenty-eight areas that we collected data.&nbsp;We use the following coding&nbsp;for the&nbsp;Labels of our <strong>Dataset</strong>:</p> <table> <tbody> <tr> <td> <p><strong>Alluvial deposits</strong></p> </td> <td> <p><strong>0</strong></p> </td> </tr> <tr> <td> <p><strong>Limestone colluvial deposits</strong></p> </td> <td> <p><strong>1</strong></p> </td> </tr> <tr> <td> <p><strong>Limestones</strong></p> </td> <td> <p><strong>2</strong></p> </td> </tr> <tr> <td> <p><strong>Schists</strong></p> </td> <td> <p><strong>3</strong></p> </td> </tr> <tr> <td> <p><strong>Quaternary sediments</strong></p> </td> <td> <p><strong>4</strong></p> </td> </tr> <tr> <td> <p><strong>Gneiss</strong></p> </td> <td> <p><strong>5</strong></p> </td> </tr> <tr> <td> <p><strong>Slope fan debris</strong></p> </td> <td> <p><strong>6</strong></p> </td> </tr> <tr> <td> <p><strong>Mixed flysch</strong></p> </td> <td> <p><strong>7</strong></p> </td> </tr> <tr> <td> <p><strong>Flysch shale and cherts</strong></p> </td> <td> <p><strong>8</strong></p> </td> </tr> <tr> <td> <p><strong>Dolomites</strong></p> </td> <td> <p><strong>9</strong></p> </td> </tr> <tr> <td> <p><strong>Granite</strong></p> </td> <td> <p><strong>10</strong></p> </td> </tr> <tr> <td> <p><strong>Sandstone flysch</strong></p> </td> <td> <p><strong>11</strong></p> </td> </tr> <tr> <td> <p><strong>Flysch colluvial deposits</strong></p> </td> <td> <p><strong>12</strong></p> </td> </tr> <tr> <td> <p><strong>Peridotite and Gabbro</strong></p> </td> <td> <p><strong>13</strong></p> </td> </tr> <tr> <td> <p><strong>River bed deposits</strong></p> </td> <td> <p><strong>14</strong></p> </td> </tr> <tr> <td> <p><strong>Gneiss colluvial deposits</strong></p> </td> <td> <p><strong>15</strong></p> </td> </tr> <tr> <td> <p><strong>Not available</strong></p> </td> <td> <p><strong>-100</strong></p> </td> </tr> <tr> <td> <p><strong>cloud coverage</strong></p> </td> <td> <p><strong>-999</strong></p> </td> </tr> </tbody> </table> <p>The following table lists the available <strong>areas </strong>and the <strong>categories </strong>that each contains<strong>:&nbsp;<a href="https://docs.google.com/spreadsheets/d/17q0L5Ltz7V4uBY9i6DhULJsJCtf7BOY1nbblB-hf3Pw/edit?usp=share_link">Lithology_Dataset</a> </strong></p> <p>&nbsp;</p> <p>For the <strong>Sentinel-2 images</strong>, we made the following process:</p> <ul> <li> <p><strong>Resampling 10m</strong></p> </li> <li> <p><strong>Subset on defined areas</strong></p> </li> </ul> <p>The Sentinel-2 map contains: Sentinel 2 false colour composite 11/8/4 with OSM background</p> <p>The Final step is the collocation of the previous into a datacube i.e a multidimensional array with 25 bands (datacube dimensions differentiate for every area) using the Aster image as base (15m spatial resolution).&nbsp;</p> <ul> <li> <p>Bands 1-14: Aster</p> </li> <li> <p>Bands 15-24: S2</p> </li> <li> <p>Band 25: Label</p> </li> </ul> <p>The code for preprocessing the dataset in order to be used for machine learning algorithms can be found in the following link:&nbsp;&nbsp;</p> <p><a href="https://github.com/georgegiannop/Lithology">https://github.com/georgegiannop/Lithology</a></p> <p><strong>Citation</strong></p> <p>If you use this dataset in your work, please cite our paper:</p> <p>Vernikos, I., Giannopoulos, G., Christopoulou, A., Begaj, A., Stefouli, M., Bratsolis, E., and Charou, E.: A dataset of Earth Observation Data for Lithological Mapping using Machine Learning, EGU General Assembly 2023, Vienna, Austria, 24&ndash;28 Apr 2023, EGU23-17570,&nbsp;<a href="https://doi.org/10.5194/egusphere-egu23-17570">https://doi.org/10.5194/egusphere-egu23-17570</a>, 2023.</p> <p>&nbsp;</p> <p>&nbsp;</p>

ShareScore

44/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
8
Harmonization
4
Access
20
Reuse readiness
8
Engagement
4

Topics