Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

181

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

181 results for “SENTINEL-2”

Learn how ShareScore rates datasets ↗
edi60/100

Spectral Vegetation Indices from Harmonized Landsat and Sentinel-2 Data for Harvard Forest 2015-2020

The goal of this work is to exploit time series of remotely sensed data sets with ground observations to improve our understanding of how seasonal variation in canopy and environmental conditions affect the relationship between vegetation indices and leaf area index (LAI) and fraction of absorbed photosynthetically active radiation (fAPAR). Using three different common vegetation indices (EVI2, NDVI, NIRV), we can estimate LAI, fAPAR, and daily absorbed photosynthetically active radiation (APAR) using a semi-empirical model.

openCC0Dec 2023View details →
zenodo48/100

Ground Truth and Automated Classification from Copernicus Sentinel-2 Imagery

<p>Ground-Truth and Sentinel2 imagery classification of <em>Trees Outside Forest</em> in an agroforestry landscape in Umbria,&nbsp;Italy.</p> <p>Location:&nbsp;Alfina plains, Castelgiorgio area, Umbria, Italy.&nbsp;Reference system:&nbsp;EPSG:32632&nbsp;(WGS84, UTM zone 32 North)&nbsp;Extent: West 740609 &mdash; East 750828,&nbsp;South 4726490 &mdash; North 4737250</p> <p>Dataset&nbsp;format: geopackage, a single file&nbsp;<strong>data.gpkg</strong>&nbsp;containing 9 vector layers (in alphabetical order):</p> <ol> <li>Areas&nbsp;&mdash; Areas of interest, 2 polygons</li> <li>Classification&nbsp;&mdash; Automated classification from Sentinel2 imagery, 11781 polygons</li> <li>Hedgerows1&nbsp;&mdash; Ground truth, hedgerows of Area1, 148 lines</li> <li>Hedgerows2&nbsp;&mdash; Ground truth, hedgerows of Area2, 135 lines</li> <li>Sentinel2&nbsp;&mdash; Sentinel2 scenes footprint, one&nbsp;polygon</li> <li>Trees1&nbsp;&mdash; Ground truth, isolated trees of Area1, 55 points</li> <li>Trees2&nbsp;&mdash; Ground truth, isolated trees of Area2, 64 points</li> <li>Woods1&nbsp;&mdash; Ground truth, small forest patches of Area1, 33 polygons</li> <li>Woods2&nbsp;&mdash; Ground truth, small forest patches of Area2, 37 polygons</li> </ol> <p>Accompanying map:&nbsp;<strong>map.qgz</strong>, Qgis 3.6 format. The geopackage&nbsp;dataset is supposed to be stored in the same directory of the map (relative path = ./)</p> <p>Dataset description and metadata: <strong>meta.pdf</strong>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2019View details →
zenodo48/100

Subsample of the maximum Water Area Extent of Telangana Rainwater Harvesting System from Sentinel-2

<p>Small Reservoirs Maximum Water Area Extent polygones (MWAE) composing the Rainwater Harvesting System (RHS) derieved from Sentinel-2 Multispectral data in the Telangana state, South-India. MWAE is extracted from Sentinel-2 cloud free images time serie collected from 2016 to 2021 (last access in 2021) over the area covered by stereoscopic images acquired from Pl&eacute;iades satellites (DEM available 10.5281/zenodo.10403040). A random forest classification is used with a set of training and validation samples. These samples are Sentinel pixel locations (10 x 10 meters) corresponding to permanent water pixels extracted from Global Surface Water datasets (doi:10.1038/nature20584) and never flooded pixels derived from Height Above Nearest Drainage data-set (10.1016/j.jhydrol.2011.03.051).</p>

opencc-by-4.0Dec 2023View details →
zenodo48/100

DeepOrchidSeries: A Sentinel-2 Dataset to inform convolutional SDMs with twelve-month Sentinel-2 image time-series, Orchid family

<p><strong>Deep Species Distribution Modelling from Sentinel-2 Image Time-series: a Global Scale Analysis on the Orchid Family</strong>&nbsp;</p> <ul> <li><strong><em>DeepOrchidSeries</em></strong> dataset gathers Sentinel-2 image time-series around geolocated orchid occurrences. Seasonal evolutions of the habitats are captured in the twelve-month RGB/IR time-series with 640x640m spatial resolution. It allows novel Species Distribution Models (SDMs) coupled with convolutional networks to take advantage of both spatial and temporal information.</li> <li>Our <strong>associated article</strong> is describing the modeling choices made to shape this ambitious dataset. It is submitted to <a href="https://www.frontiersin.org/research-topics/18336/plant-biodiversity-science-in-the-era-of-artificial-intelligence">https://www.frontiersin.org/research-topics/18336/plant-biodiversity-science-in-the-era-of-artificial-intelligence</a>. We believe such global data, methods and scripts are valuable to the conservation ecology community and especially deep-SDMs users. To our knowledge, no similar ready-to-use dataset is available. In the article, the dataset&#39;s temporal dimension is proven to significantly improve SDMs performances.</li> <li><strong><em>sen2patch</em></strong> is the gitlab project gathering the code to create such dataset. It is available at <a href="https://gitlab.inria.fr/jestopin/sen2patch">https://gitlab.inria.fr/jestopin/sen2patch</a>.</li> <li><strong><em>DeepOrchidSeries.csv</em></strong> contains all occurrences-level information. <ul> <li>We advice to load it with: <pre><code class="language-python">import pandas as pd df = pd.read_csv("path/to/DeepOrchidSeries.csv", sep=';') df.columns ['gbifid', 'canonical_name', 'decimallatitude', 'decimallongitude', 'speciesKey', 'cell_index', 'bot_country', 'bot_code', 'lvl2_code', 'continent_code']</code></pre> <ul> <li>&#39;gbifid&#39; is the occurrences GBIF ID</li> <li>&#39;canonical_name&#39;, is the species canonical name</li> <li>&#39;decimallatitude&#39;, &#39;decimallongitude&#39; are the species coordinates in decimal degrees</li> <li>&#39;speciesKey&#39; is the species GBIF unique identifier</li> <li>&#39;cell_index&#39; is&nbsp;a unique cell ID in a 0.0025&deg; lon/lat grid partitioning the Earth (used to stratify train/val/test set by geographic blocks)</li> <li>&#39;bot_country&#39;, &#39;bot_code&#39;, &#39;lvl2_code&#39;, &#39;continent_code&#39; are geographic subdivisions defined in <a href="https://github.com/tdwg/wgsrpd">https://github.com/tdwg/wgsrpd</a> (code and string for WGSRPD level 1, the botanical countries)</li> </ul> </li> </ul> </li> <li> <p>Initial <a href="https://www.gbif.org/">GBIF</a> query DOI is <a href="http://https://doi.org/10.15468/dl.4bijtu">https://doi.org/10.15468/dl.4bijtu</a> (26 August 2019).</p> </li> <li><strong><em>DeepOrchidSeries.tar</em></strong> file contains the satellite image time-series and is available at <a href="https://lab.plantnet.org/deeporchidseries/">https://lab.plantnet.org/deeporchidseries/</a> <ul> <li><em>.tar</em> archive measure 286 GB and extends to 432 GB once decompressed.</li> <li>Image time-series relative tree paths are constructed from the occurrences unique GBIF IDs.</li> <li>For a given occurence <em>gbifid</em>, matching patches are located in: <em>final_dataset_by_gbifid/gbifid[-2:]/gbifid[-4:-2]</em>, <em>i.e.</em> in a first folder named with the <em>gbifid</em> last two numbers and a subfolder with the previous two ones. Example: the time-series files matching occurrence 2236837714 are located at <em>final_dataset_by_gbifid/14/77/</em>.&nbsp;</li> <li>Image time-series are composed of twelve 16 bits RGB <em>.png</em>&nbsp; and twelve 16 bits IR <em>.png</em> files containing data identical to the original L1C products, no lossy compression was made. There are one RGB and one IR .png file per month.</li> <li>Patches from month MM/YYYY of occurrence <em>gbifid</em> are named<em> </em><em>RGB_YYYY_MM_gbifid_.png</em> and <em>IR0_YYYY_MM_gbifid_.png</em>.</li> </ul> </li> <li><em><strong>models.zip</strong></em> is the archive containing the four PyTorch models weights described in our article and<strong><em> </em></strong><em><strong>inception_env.py</strong></em> the used Inception V3 architecture. <em><strong>index.json</strong></em> contains the dictionnary linking the models class indexes from 0 to 14128 with our labels <em>speciesKey</em>: {&quot;class_index&quot;:speciesKey}.</li> </ul> <p>&nbsp;</p> <ul> <li><strong>ACKNOWLEDGMENTS</strong>: We warmly thank Alexander Zizka et al. for providing us the geographically and taxonomically curated set of Orchids occurrences. This dataset contains modified Copernicus Sentinel data and Copernicus Service information (2018). Sentinel-2 MSI data used were available at no cost from ESA Sentinels Scientific Data Hub.</li> </ul>

opencc-by-4.0Dec 2021View details →
zenodo48/100

Data supporting 'Empirical correction of systematic orthorectification error in Sentinel-2 velocity fields for Greenlandic outlet glaciers'

<p><strong>Note:&nbsp;An updated dataset covering the majority of Greenland&#39;s marine-terminating glaciers is available as part of the NASA Making Earth System Data Records for Use in Research Environments (MEaSUREs) project through the National Snow and Ice Data Center (NSIDC) at&nbsp;<a href="https://doi.org/10.5067/B28FM2QVVYWY">https://doi.org/10.5067/B28FM2QVVYWY</a>.&nbsp;</strong></p> <p>Data supporting the paper:</p> <blockquote> <p>Chudley, T. R., Howat, I. M., Yadav, B. N., &amp; Noh, M. J. (2022). Empirical correction of systematic orthorectification error in Sentinel-2 velocity fields for Greenlandic outlet glaciers.&nbsp;<em>The Cryosphere.&nbsp;</em>16, 2629&ndash;2642, https://doi.org/10.5194/tc-16-2629-2022</p> </blockquote> <p>Dataset consists of four netCDF files containing stacked Sentinel-2 velocity data of four Greenlandic outlet glaciers (Helheim Glacier, Jakobshavn Isbr&aelig;, Store Glacier, and Kangerlussuaq) between 2017 and 2021. Velocity data are derived and corrected following the methods outlined in Chudley&nbsp;<em>et al.</em>&nbsp;(2022).&nbsp;&nbsp;</p> <p>NetCDF files are created by, and tested to be&nbsp;readable by, Python&#39;s xarray package.</p> <p>The dimensions of the netCDF file are as follows:</p> <ul> <li><strong>X</strong> - <em>x&nbsp;</em>coordinates in NSDIC Sea Ice Polar Stereographic North (EPSG:3413).</li> <li><strong>Y</strong> -&nbsp;<em>y</em>&nbsp;coordinates in NSDIC Sea Ice Polar Stereographic North (EPSG:3413).</li> <li><strong>time</strong> - temporal midpoint of velocity field.</li> </ul> <p>The variables of the netCDF file are as follows:</p> <ul> <li><strong>dmag</strong> - the absolute magnitude of the velocity, in metres per day.</li> <li><strong>dx</strong> - the velocity in the&nbsp;<em>x</em>&nbsp;direction, in metres per day.</li> <li><strong>dy</strong> - the velocity in the&nbsp;<em>y</em>&nbsp;direction, in metres per day.</li> <li><strong>date1</strong> - the date and time of the first scene acquisition.</li> <li><strong>date2</strong> - the date and time of the second scene acquisition.</li> <li><strong>baseline</strong> - the temporal baseline, in days, between scene acquisitions.</li> <li><strong>orbit_pair</strong> - the combination of orbital pathways in the string format &#39;RXXX_RYYY&#39;, where XXX is relative orbit number of the first scene and YYY the relative orbit number of the second scene.</li> <li><strong>mag_rmse</strong> - the root mean square error of the absolute velocity of the off-ice area.&nbsp;</li> <li><strong>dx_mean</strong> - the mean velocity of the off-ice area in the&nbsp;<em>x</em>&nbsp;direction.</li> <li><strong>dx_sd</strong> - the standard deviation of the velocity of the off-ice area in the&nbsp;<em>x</em>&nbsp;direction.</li> <li><strong>dy_mean</strong> -&nbsp;the mean velocity of the off-ice area in the&nbsp;<em>x</em>&nbsp;direction.</li> <li><strong>dy_sd</strong> -&nbsp;the standard deviation of the velocity of the off-ice area in the <em>y</em>&nbsp;direction.</li> </ul>

opencc-by-4.0Jun 2022View details →
zenodo48/100

Sentinel2GlobalLULC: A dataset of Sentinel-2 georeferenced RGB imagery annotated for global land use/land cover mapping with deep learning (License CC BY 4.0)

<p>Sentinel2GlobalLULC is a deep learning-ready dataset of RGB images from the Sentinel-2 satellites designed for global land use and land cover (LULC) mapping. Sentinel2GlobalLULC v2.1&nbsp;contains 194,877 images in GeoTiff and JPEG format corresponding to 29 broad LULC classes. Each image has 224 x 224 pixels at 10 m spatial resolution and was produced by assigning the 25th percentile of all available observations in the Sentinel-2 collection between June 2015 and October 2020 in order to remove atmospheric effects (i.e., clouds, aerosols, shadows, snow, etc.). A spatial purity value was assigned to each image based on the consensus across 15 different global LULC products available in Google Earth Engine (GEE).&nbsp;</p> <p>&nbsp;</p> <p>Our dataset is structured into 3 main zip-compressed folders, an Excel file with a dictionary for class names and descriptive statistics per LULC class, and a python script to convert RGB GeoTiff images into JPEG format. The first folder called &quot;Sentinel2LULC_GeoTiff.zip&quot;&nbsp;contains 29 zip-compressed subfolders where each one corresponds to a specific LULC class with hundreds to thousands of GeoTiff Sentinel-2 RGB images. The second folder called &quot;Sentinel2LULC_JPEG.zip&quot; contains 29 zip-compressed subfolders with a JPEG formatted version of the same images provided in the first main folder. The third folder called &quot;Sentinel2LULC_CSV.zip&quot; includes 29 zip-compressed CSV files with as many rows as provided images and with 12&nbsp;columns containing the following metadata (this same metadata is provided in the image filenames):&nbsp;</p> <ul> <li>Land Cover Class ID: is the identification number of each LULC class</li> <li>Land Cover Class Short Name: is the short name of each LULC class</li> <li>Image ID: is the identification number of each image within its corresponding LULC class&nbsp;</li> <li>Pixel purity Value: is the spatial purity of each pixel for its corresponding LULC class calculated as the spatial consensus across up to 15 land-cover products&nbsp;</li> <li>GHM Value: is the spatial average of the Global Human Modification index (gHM) for each image</li> <li>Latitude: is the latitude of the center point of each image</li> <li>Longitude: is the longitude of the center point of each image</li> <li>Country Code: is the Alpha-2 country code of each image as described in the ISO 3166 international standard. To understand the country codes, we recommend the user to visit the following website where they present the Alpha-2 code for each country as described in the ISO 3166 international standard:https: //www.iban.com/country-codes</li> <li>Administrative Department Level1: is the administrative level 1 name to which each image belongs</li> <li>Administrative Department Level2: is the administrative level 2 name to which each image belongs</li> <li>Locality: is the name of the locality to which each image belongs</li> <li>Number of S2 images : is&nbsp;the number of found instances in the corresponding Sentinel-2 image collection between June 2015 and October 2020, when compositing&nbsp;and exporting&nbsp;its corresponding&nbsp;image tile</li> </ul> <p>For seven LULC classes, we could not export from GEE all images that fulfilled a spatial purity of 100% since there were millions of them. In this case, we exported a stratified random sample of 14,000 images and provided an additional CSV file with the images actually contained in our dataset. That is, for these seven LULC classes, we provide these 2 CSV files:</p> <ul> <li>A CSV file that contains all exported images for this class&nbsp;</li> <li>A CSV file that contains all images available for this class at spatial purity of 100%, both the ones exported and the ones not exported, in case the user wants to export them. These CSV filenames end with &quot;including_non_downloaded_images&quot;.</li> </ul> <p>To clearly state the geographical coverage of images available in this dataset,&nbsp; we&nbsp;included in the version v2.1, &nbsp;a compressed folder called &quot;Geographic_Representativeness.zip&quot;. This zip-compressed folder&nbsp;contains a csv file&nbsp;for each LULC class that provides the complete list of countries represented in that class. Each csv file has two columns, the first one gives the country code and the second one gives the number of images provided in that country for that LULC class. In addition to these 29 csv files, we provided another csv file that maps each ISO Alpha-2 country code to its original full country name.</p> <p>&copy;&nbsp;<a href="https://doi.org/10.5281/zenodo.5055632">Sentinel2GlobalLULC Dataset&nbsp;</a>by&nbsp;&nbsp;Yassir Benhammou, Domingo Alcaraz-Segura, Emilio Guirado, Rohaifa Khaldi, Boujem&acirc;a Achchab, Francisco Herrera &amp; Siham Tabik&nbsp;is marked with Attribution 4.0 International&nbsp;(CC-BY 4.0)</p>

opencc-by-4.0Jul 2022View details →
zenodo48/100

A Map of Land Use and Land Cover in Southern Malawi Derived from Sentinel-2 Data (2023)

<h3><strong>Overview</strong></h3> <p>The land use and land cover map comprises the Mulanje and Phalombe districts, in Southern Malawi. It includes five classes: forest, natural vegetation, cropland, wetland, and other lands. The map is derived from Sentinel-2 mosaics, resulting in a spatial resolution of 10 meters, for 2023.&nbsp;</p> <p>&nbsp;</p> <h3><strong>Map Accuracy</strong></h3> <p>The land use and land cover map achieves an overall accuracy of 89%. Details of user and producer accuracies are provided in Table 1.</p> <p>Table 1:&nbsp; Land use and land cover classification validation,including overall, producer (PA) and user (UA) accuracies values for each class.</p> <div> <table> <tbody> <tr> <td> <p><strong>Class&nbsp;</strong></p> </td> <td> <p><strong>Producer Accuracy</strong></p> </td> <td> <p><strong>User Accuracy</strong></p> </td> </tr> <tr> <td> <p>Cropland</p> </td> <td> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 93%</p> </td> <td> <p>&nbsp; &nbsp; &nbsp; &nbsp;85%</p> </td> </tr> <tr> <td> <p>Wetland</p> </td> <td> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp;100%</p> </td> <td> <p>&nbsp; &nbsp; &nbsp; 100%</p> </td> </tr> <tr> <td> <p>Other Lands</p> </td> <td> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;90%</p> </td> <td> <p>&nbsp; &nbsp; &nbsp; &nbsp;95%</p> </td> </tr> <tr> <td> <p>Forest</p> </td> <td> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;79%</p> </td> <td> <p>&nbsp; &nbsp; &nbsp; 90%</p> </td> </tr> <tr> <td> <p>Natural Vegetation</p> </td> <td> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 90%</p> </td> <td> <p>&nbsp; &nbsp; &nbsp; 90%</p> </td> </tr> <tr> <td> <p><strong>Overall Accuracy</strong></p> </td> <td><br> <p><strong>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;89%</strong></p> </td> </tr> </tbody> </table> </div> <h3>&nbsp;</h3> <h3><strong>Files descripion</strong></h3> <ul> <li>MLW_Sentinel_LULC_2023.tif / .qml: land use and land cover map and QGIS style file</li> <li>training_samples.gpkg: training samples with class labels</li> <li>validation_samples.gpkg: validation samples with class labels</li> </ul>

opencc-by-4.0May 2024View details →
zenodo48/100

FNEWs Kalman-gefilterte Sentinel-2 Bilder 2017 bis 2022

<p>Im Projekt Fernerkundungsbasiertes Nationales Erfassungssystem f&uuml;r Waldsch&auml;den (FNEWS) werden Waldsch&auml;den auf Grundlage eines strukturelles Zeitreihenmodells in Kombination mit dem Kalman-Filter erfasst. Das auf Sentinel-2-Satellitendaten basierende Waldschadenerfassungssystem erm&ouml;glicht differenzierte Ver&auml;nderungs- und Schadanalysen f&uuml;r die vier Untersuchungsgebiete des Projektes, die in Sachsen, Niedersachsen, Bayern und Baden-W&uuml;rttemberg liegen.</p> <p>F&uuml;r die j&auml;hrliche Kartierung von Waldsch&auml;den wird jeweils zum Stichtag 31.08. ein wolkenfreies Kalman-gefiltertes Sentinel-2 Bild (KFB) in einer CIR-Falschfarbendarstellung erzeugt. Die Kalman-Filterung bewirkt, dass atmosph&auml;rische St&ouml;reinfl&uuml;sse unterdr&uuml;ckt werden und gleichzeitig der Zustand am Boden m&ouml;glichst wirklichkeitsgetreu im KFB abgebildet wird. Die KFB sind die Datengrundlage f&uuml;r die Ableitung der FNEWs-Jahresprodukte (https://doi.org/10.3220/DATA20230907171359-0). Hier werden die Sentinel-2 KFB aus den Jahren 2017 bis einschlie&szlig;lich 2022 zur Verf&uuml;gung gestellt. Sie k&ouml;nnen allgemein als Hintergrund-Layer verwendet werden oder speziell bei der Interpretation der FNEWS Jahresprodukte helfen. Die KFB werden als COG - cloud optimized GeoTIFF zum Download bereitgestellt.</p> <p>Bei Verwendung der KFB sind folgende Einschr&auml;nkungen zu ber&uuml;cksichtigen: Wald&auml;nderungen sind m&ouml;glicherweise noch nicht oder noch nicht in ihrer vollen Auspr&auml;gung im KFB abgebildet, wenn sie zeitlich erst kurz vor dem Stichdatum eingetreten sind. Das strukturelle Zeitreihenmodell ist auf Wald optimiert, d.h. au&szlig;erhalb des Waldes ist mit gr&ouml;&szlig;eren Abweichungen zu rechnen.</p>

opencc-by-4.0May 2024View details →
zenodo48/100

MUDDAT: A SENTINEL-2 IMAGE-BASED MUDDY WATER BENCHMARK DATASET FOR ENVIRONMENTAL MONITORING.

<p>This is a dataset for mapping muddy waters based on Sentinel-2 (L2A products) satellite imagery. The image data are saved as GeoTIFF files and metadata files are provided in json format. There are 19 images in total, based on 16 distinct European Areas of Interest (AOIs), covering a total of 9 countries such as:</p> <ul> <li>Greece</li> <li>Italy</li> <li>France</li> <li>Spain</li> <li>Belgium</li> <li>UK</li> <li>Sweden</li> <li>Finland and</li> <li>Serbia</li> </ul> <p>From the Sentinel-2 L2A products were extracted 10 spectral bands and then resampled to a 10m spatial resolution. All spectral bands used can be found in the Metadata/Source files. The annotated images comprise 3 classes, "Non-muddy", "Muddy" and "Ambiguous". More details about the annotation methodology can be found on the accepted abstract (file:&nbsp;<a href="../api/records/11220437/draft/files/Accepted_Abstract_03_15_2024.pdf/content" target="_blank" rel="noopener noreferrer">Accepted_Abstract_03_15_2024.pdf</a>) or the published paper, that you can find here: <a href="https://doi.org/10.1109/IGARSS53475.2024.10642051" target="_blank" rel="noopener">10.1109/IGARSS53475.2024.10642051</a>.</p>

opencc-by-4.0May 2024View details →
zenodo48/100

S1S2-Water: A global dataset for semantic segmentation of water bodies from Sentinel-1 and Sentinel-2 satellite images

<p>The S1S2-Water dataset is a global reference dataset for training, validation and testing of convolutional neural networks for semantic segmentation of surface water bodies in publicly available Sentinel-1 and Sentinel-2 satellite images. The dataset consists of 65 triplets of Sentinel-1 and Sentinel-2 images with quality checked binary water mask. Samples are drawn globally on the basis of the Sentinel-2 tile-grid (100 x 100 km) under consideration of pre-dominant landcover and availability of water bodies. Each sample is complemented with metadata and Digital Elevation Model (DEM) raster from the Copernicus DEM.</p><p>This work was supported by the German Federal Ministry of Education and Research (BMBF) through the project "Künstliche Intelligenz zur Analyse von Erdbeobachtungs- und Internetdaten zur Entscheidungsunterstützung im Katastrophenfall" (AIFER) under Grant 13N15525, and by the Helmholtz Artificial Intelligence Cooperation Unit through the project "AI for Near Real Time Satellite-based Flood Response" (AI4FLOOD) under Grant ZT-IPF-5-39.&nbsp;</p>

opencc-by-4.0Sep 2023View details →
zenodo48/100

Sentinel-2 Satellite Imagery Based Forest Fire Monitoring

<p><strong>Forest Fire in Villages near Berlin - Normalized Burn Ratio (NBR)</strong></p> <p>Villages in Treuenbrietzen (Frohnsdorf, Klausdorf and Tiefenbrunnen) around 50 km southwest of Berlin have been severely affected by recent unpredicted wildfire and the size of the burned area is about of 400 hectares, which started to spread on 23rd of August, 2018. More than 500 people had to leave their homes as a result of the fire in Treuenbrietzen and the burning fire with dense smoke continued for days. This year Europe has faced a long hot dry summer with almost no rain and as a consequence some European countries like Germany are on high alert regarding possible forest fires.</p>

opencc-by-4.0Feb 2019View details →
zenodo48/100

A harmonized Landsat Sentinel-2 (HLS) dataset for benchmarking time series reconstruction methods of vegetation indices

<p>Satellite images can be used to derive time series of vegetation indices, such as normalized difference vegetation index (NDVI) or enhanced vegetation index (EVI), at global scale. Unfortunately, recording artifacts, clouds, and other atmospheric contaminants impacts a significant portion of the produced images, requiring the usage of ad-hoc techniques to reconstruct the time series in the affected regions. In literature, several methods have been proposed to fill the gaps present in the images, and some works also presented performance comparisons between them (Roerink et al., 2000; Moreno-Mart&iacute;nez et al., 2020; Siabi et al., 2022). Because of the lack of a ground truth for the reconstructed images, the performance evaluation requires the creation of datasets where artificial gaps are introduced in a reference image, such that metrics like the root mean square error (RMSE) can be computed comparing the reconstructed images with the reference one. Different approaches have been used to create the reference images and the artificial gaps, but in most cases, the artificial gaps are introduced using arbitrary patterns and/or the reference image is produced artificially and not using real satellite images (e.g. Kandasamy et al., 2013; Liu et al., 2017; Julien &amp; Sobrino, 2018). In addition, to the best of our knowledge, few of them are openly available and directly accessible allowing for fully reproducible research.</p> <p>We provide here a benchmark dataset for time series reconstruction method based on the<strong>&nbsp;<a href="https://hls.gsfc.nasa.gov/">harmonized Landsat Sentinel-2 (HLS)</a> </strong>collection where the artificial gaps are introduced with a realistic spatio-temporal distribution. In particular, we selected six tiles that we considered representative for most of the main climate classes (e.g. equatorial, arid, warm temperature, boreal and polar), as depicted in the preview.</p> <p>Specifically, following the&nbsp;<strong><a href="https://hls.gsfc.nasa.gov/products-description/tiling-system/">relative tiling system</a></strong> shown above, we downloaded the Red, NIR and F-mask bands from both the HLSL30 and HLSS30 collections for the tiles 19FCV, 22LEH, 32QPK, 31UFS, 45WFV and 49MWM. From the Red and NIR band we derived the NDVI as:</p> <p><span class="math-tex">\(NDVI = {NIR - Red \over NIR + Red}\)</span></p> <p>only for clear-sky on lend pixels (F-mask bits 1, 3, 4 and 5 equal zero), setting as not a number the remaining pixels. The images are then aggregated on a 16 days base, averaging the available values for each pixel in each temporal range. The so obtained data, are considered from us as the reference data for the benchmarking, and stored following the file naming convention</p> <p><em>HLS.T&lt;TILE_NAME&gt;.&lt;YYYYDDD&gt;.v2.0.NDVI.tif</em></p> <p>where <em>TILE_NAME</em> is one between the above specified ones, <em>YYYY</em> is the corresponding year (spanning from 2015 to 2022) and <em>DDD</em> is the day of the year from which the corresponding 16 days range starts. Finally, for each tile, we have a time series composed of <strong>184</strong> images (23 images for 8 years) that can be easily manipulated, for example using the <strong><a href="https://github.com/scikit-map/scikit-map/tree/master">Scikit-Map library</a></strong> in Python.</p> <p>Starting from those data, for each image we considered the mask of currently present gaps, we randomly rotated it by 90, 180 or 270 degrees and we added artificial gaps in the pixels of the rotated mask. Doing so, we believe that the spatio-temporal distribution will be still realistic, providing a solid benchmark for gap-filling methods that work on time series, on spatial pattern or combination of the both.</p> <p>The data including the artificial gaps are stored with the naming structure</p> <p><em>HLS.T&lt;TILE_NAME&gt;.&lt;YYYYDDD&gt;.v2.0.NDVI_art_gaps.tif</em></p> <p>following the previously mentioned convention. The performance metrics, such as RMSE or normalized RMSE (NRMSE), can be computed by applying a reconstruction method on the images with artificial gaps, and then comparing the reconstructed time series with the reference one only on the artificially created gaps locations.&nbsp;</p> <p>This dataset was used to compare the performance of some gap-filling methods and we provide a&nbsp;<strong><a href="https://github.com/OpenGeoHub/EO-benchmark/blob/main/gap_filling_methods/gap_filling_comparison.ipynb">Jupyter notebook</a></strong> that shows how to access and use the data. The files are provided in GeoTIFF format and projected in the coordinate reference system WGS 84 / UTM zone 19N (EPSG:32619).&nbsp;</p> <p>If you succeed to produce higher accuracy or develop a new algorithm for gap filling, please contact authors or post on our GitHub repository. May the force be with you!</p> <p>References:</p> <ol> <li> <p>Julien, Y., &amp; Sobrino, J. A. (2018). TISSBERT: A benchmark for the validation and comparison of NDVI time series reconstruction methods. Revista de Teledetecci&oacute;n, (51), 19-31.&nbsp;<a href="https://doi.org/10.4995/raet.2018.9749">https://doi.org/10.4995/raet.2018.9749</a>&nbsp;</p> </li> <li> <p>Kandasamy, S., Baret, F., Verger, A., Neveux, P., &amp; Weiss, M. (2013). A comparison of methods for smoothing and gap filling time series of remote sensing observations&ndash;application to MODIS LAI products. Biogeosciences, 10(6), 4055-4071.&nbsp;<a href="https://doi.org/10.5194/bg-10-4055-2013">https://doi.org/10.5194/bg-10-4055-2013</a>&nbsp;</p> </li> <li> <p>Liu, R., Shang, R., Liu, Y., &amp; Lu, X. (2017). Global evaluation of gap-filling approaches for seasonal NDVI with considering vegetation growth trajectory, protection of key point, noise resistance and curve stability. Remote Sensing of Environment, 189, 164-179.&nbsp;<a href="https://doi.org/10.1016/j.rse.2016.11.023">https://doi.org/10.1016/j.rse.2016.11.023</a>&nbsp;</p> </li> <li> <p>Moreno-Mart&iacute;nez, &Aacute;., Izquierdo-Verdiguier, E., Maneta, M. P., Camps-Valls, G., Robinson, N., Mu&ntilde;oz-Mar&iacute;, J., ... &amp; Running, S. W. (2020). Multispectral high resolution sensor fusion for smoothing and gap-filling in the cloud. Remote Sensing of Environment, 247, 111901.<a href="https://doi.org/10.1016/j.rse.2020.111901"> https://doi.org/10.1016/j.rse.2020.111901</a>&nbsp;</p> </li> <li> <p>Roerink, G. J., Menenti, M., &amp; Verhoef, W. (2000). Reconstructing cloudfree NDVI composites using Fourier analysis of time series. International Journal of Remote Sensing, 21(9), 1911-1917.&nbsp;<a href="https://doi.org/10.1080/014311600209814">https://doi.org/10.1080/014311600209814</a></p> </li> <li> <p>Siabi, N., Sanaeinejad, S. H., &amp; Ghahraman, B. (2022). Effective method for filling gaps in time series of environmental remote sensing data: An example on evapotranspiration and land surface temperature images. Computers and Electronics in Agriculture, 193, 106619.<a href="https://doi.org/10.1016/j.compag.2021.106619"> https://doi.org/10.1016/j.compag.2021.106619</a></p> </li> </ol>

opencc-by-4.0Dec 2022View details →
zenodo44/100

Sentinel-2 Cloud Mask Catalogue

<p><strong>Overview</strong></p> <p>This dataset comprises cloud masks for 513 1022-by-1022 pixel subscenes, at 20m resolution, sampled random from the 2018 Level-1C Sentinel-2 archive. The design of this dataset follows from some observations about cloud masking: (i) performance over an entire product is highly correlated, thus subscenes provide more value per-pixel than full scenes, (ii) current cloud masking datasets often focus on specific regions, or hand-select the products used, which introduces a bias into the dataset that is not representative of the real-world data, (iii) cloud mask performance appears to be highly correlated to surface type and cloud structure, so testing should include analysis of failure modes in relation to these variables.</p> <p>The data was annotated semi-automatically, using the <a href="https://github.com/ESA-PhiLab/iris">IRIS toolkit</a>, which allows users to dynamically train a Random Forest (implemented using <a href="https://github.com/microsoft/LightGBM">LightGBM</a>), speeding up annotations by iteratively improving it&#39;s predictions, but preserving the annotator&#39;s ability to make final manual changes when needed. This hybrid approach allowed us to process many more masks than would have been possible manually, which we felt was vital in creating a large enough dataset to approximate the statistics of the whole Sentinel-2 archive.</p> <p>In addition to the pixel-wise, 3 class (CLEAR, CLOUD, CLOUD_SHADOW) segmentation masks, we also provide users with binary<br> classification &quot;tags&quot; for each subscene that can be used in testing to determine performance in specific circumstances. These include:</p> <ul> <li><strong>SURFACE TYPE</strong>: <em>11 categories</em></li> <li><strong>CLOUD TYPE</strong>: <em>7 categories</em></li> <li><strong>CLOUD HEIGHT</strong>: <em>low, high</em></li> <li><strong>CLOUD THICKNESS</strong>: <em>thin, thick</em></li> <li><strong>CLOUD EXTENT</strong>: <em>isolated, extended</em></li> </ul> <p>&nbsp;</p> <p>Wherever practical, cloud shadows were also annotated, however this was sometimes not possible due to high-relief terrain, or large ambiguities. In total, 424 were marked with shadows (if present), and 89 have shadows that were not annotatable due to very ambiguous shadow boundaries, or terrain that cast significant shadows. If users wish to train an algorithm specifically for cloud shadow masks, we advise them to remove those 89 images for which shadow was not possible, however, bear in mind that this will systematically reduce the difficulty of the shadow class compared to real-world use, as these contain the most difficult shadow examples.</p> <p>In addition to the 20m sampled subscenes and masks, we also provide users with shapefiles that define the boundary of the mask on the original Sentinel-2 scene. If users wish to retrieve the L1C bands at their original resolutions, they can use these to do so.</p> <p>Please see the README for further details on the dataset structure&nbsp;and more.</p> <p>&nbsp;</p> <p><strong>Contributions &amp; Acknowledgements</strong></p> <p>The data were collected, annotated, checked, formatted and published by Alistair Francis and John Mrziglod.</p> <p>Support and advice was provided by Prof. Jan-Peter Muller and Dr. Panagiotis Sidiropoulos, for which we are grateful.</p> <p>We would like to extend our thanks to Dr. Pierre-Philippe Mathieu and the rest of the team at <em>ESA PhiLab</em>, who provided the environment in which this project was conceived, and continued to give technical support throughout.</p> <p>Finally, we thank the <em>ESA Network of Resources</em> for sponsoring this project by providing ICT resources.</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Oct 2020View details →
zenodo44/100

Robust Damage Estimation of Typhoon Goni on Coconut Crops with Sentinel-2 Imagery

<p>Damage estimation status of coconut trees plantation in the Phillippines derived from Sentinel-2 Imagery. Overall we estimated that 14.1 M coconut trees were affected by the typhoon inside our area of study. Please refer to our <a href="https://www.mdpi.com/2072-4292/13/21/4302">original</a> publication for more details.</p> <p>0: Uncertain</p> <p>1: No Data</p> <p>2: Background class</p> <p>3: Unchanged coconut plantation</p> <p>4: Damaged coconut plantation</p> <p>5: New coconut plantation</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Oct 2021View details →
zenodo44/100

Reference dataset for comparison of cloud detection algorithms for Sentinel-2 imagery

<p>Sentinel-2 cloud mask reference dataset generated and analyzed as part of Tarrio, K., Tang, X., Masek, J.G., Claverie, M., Ju, J., Qiu, S., Zhu, Z. and Woodcock, C.E., 2020. Comparison of cloud detection algorithms for Sentinel-2 imagery. Science of Remote Sensing, 2, p.100010. [https://www.sciencedirect.com/science/article/pii/S2666017220300092](https://www.sciencedirect.com/science/article/pii/S2666017220300092)</p> <p><strong>1. Reference masks</strong></p> <p><strong>Algorithms:</strong></p> <ul> <li>Fmask 1.x</li> <li>Fmask 2.x</li> <li>Fmask 4.x</li> <li>Tmask</li> <li>Sen2Cor</li> <li>MAJA</li> <li>LaSRC</li> </ul> <p><strong>Locations:</strong></p> <ul> <li>South Africa (35JPM)</li> <li>Senegal (28PDC)</li> <li>Switzerland (32TLT)</li> <li>France (31TCJ, 31TFJ)</li> <li>Morocco (29RNQ)</li> </ul> <p><strong>Standardized legend:</strong></p> <p>Original algorithm outputs were standardized to the same categorical legend.</p> <ul> <li>0 = clear land</li> <li>1 = clear water</li> <li>2 = cloud shadow</li> <li>3 = snow/ice</li> <li>4 = cloud</li> </ul> <p>All reference masks processed to both 10m and 30m resolution, with the exception of Tmask, which is available only at a 30m resolution.</p> <p><strong>Mask naming convention:</strong></p> <p>All processed masks are named according to the following convention:<br> M&lt;*resolution*&gt;&lt;*S2 MGRS tile ID*&gt;&lt;*YYYY*&gt;&lt;*DOY*&gt;&lt;*algorithm*&gt;<br> e.g. **M30T28PDC2016351TMASK**</p> <p><br> <strong>2. Interpreted sample points</strong></p> <p>Sample points were selected based on agreement among different map products. This record includes a shapefile with the final interpretations for each of the sampled sites. (See publication for additional information.)</p>

opencc-by-4.0Nov 2021View details →
zenodo44/100

Outputs of the Jupyter Notebook - Detecting floating objects using Deep Learning and Sentinel-2 imagery

<p>The dataset contains the outputs of the notebook &quot;Detecting floating objects using Deep Learning and Sentinel-2 imagery&quot;&nbsp;published in the ocean modelling section of The Environmental Data Science Book.</p> <p><strong>Contributions</strong></p> <p><em>Notebook</em></p> <ul> <li> <p>Jamila Mifdal (author), European Space Agency &Phi;-lab,&nbsp;<a href="https://github.com/jmifdal">@jmifdal</a></p> </li> <li> <p>Raquel Carmo (author), European Space Agency &Phi;-lab,&nbsp;<a href="https://github.com/raquelcarmo">@raquelcarmo</a></p> </li> <li> <p>Alejandro Coca-Castro (reviewer), The Alan Turing Institute,&nbsp;<a href="https://github.com/acocac">@acocac</a></p> </li> </ul> <p><em>Modelling codebase</em></p> <ul> <li> <p>Jamila Mifdal (author), European Space Agency &Phi;-lab,&nbsp;<a href="https://github.com/jmifdal">@jmifdal</a></p> </li> <li> <p>Raquel Carmo (author), European Space Agency &Phi;-lab,&nbsp;<a href="https://github.com/raquelcarmo">@raquelcarmo</a></p> </li> <li> <p>Marc Ru&szlig;wurm (author), EPFL-ECEO,&nbsp;<a href="https://github.com/MarcCoru">@marccoru</a></p> </li> </ul>

opencc-by-4.0Jan 2022View details →
zenodo44/100

Paired Sentinel-1 and Sentinel-2 Images for 2 Locations in Scotland and India for 2019 and 2020

<p>The dataset contains two years of coverage (2019 and 2020) for two distant geographical areas in India and in Scotland.</p> <p>If using this dataset, please cite the paper where it has been introduced:</p> <pre><code>@article{rs14061342, author = {Czerkawski, Mikolaj and Upadhyay, Priti and Davison, Christopher and Werkmeister, Astrid and Cardona, Javier and Atkinson, Robert and Michie, Craig and Andonovic, Ivan and Macdonald, Malcolm and Tachtatzis, Christos}, title = {Deep Internal Learning for Inpainting of Cloud-Affected Regions in Satellite Imagery}, journal = {Remote Sensing}, volume = {14}, year = {2022}, number = {6}, article-number = {1342}, url = {https://www.mdpi.com/2072-4292/14/6/1342}, ISSN = {2072-4292}, DOI = {10.3390/rs14061342} }</code></pre> <p>&nbsp;</p>

opencc-by-4.0Jan 2022View details →
zenodo44/100

Satellite-derived chlorophyll-a concentrations for Lake Harsha (USA) using Mixture Density Networks and Sentinel-2 and Landsat 8 imagery

<p>This dataset contains satellite-derived chlorophyll-a data of Lake Harsha (USA) for the period 21 Mar. 2013 - 01 Feb. 2021. Chlorophyll-a concentrations&nbsp;have been calculated using Mixture Density Networks and Sentinel-2 and Landsat 8 imagery.</p> <p>Mixture Density Networks are a class of neural networks that tackle the inverse problem by modelling the multimodal distribution of target variables using a mixture of Gaussians. For more information, please refer to the following:</p> <ul> <li>Pahlevan, N., Smith, B., Alikas, K., Anstee, J., et al. (2022). Simultaneous retrieval of selected optical water quality indicators from Landsat-8, Sentinel-2, and Sentinel-3. <em>Remote Sensing of Environment, 270</em>, 112860</li> <li>Smith, B., Pahlevan, N., Schalles, J., et al. (2021). A Chlorophyll-a Algorithm for Landsat-8 Based on Mixture Density Networks. <em>Frontiers in Remote Sensing, 1</em></li> <li>Pahlevan, N., Smith, B., Schalles, J., et al. (2020). Seamless retrievals of chlorophyll-a from Sentinel-2 (MSI) and Sentinel-3 (OLCI) in inland and coastal waters: A machine-learning approach. <em>Remote Sensing of Environment, 240</em>, 111604</li> </ul>

opencc-by-4.0Jun 2022View details →
zenodo44/100

Codes and dataset of the publication "Effectiveness of Sentinel-1 and Sentinel-2 for Flood Detection Assessment in Europe"

<p>The folder contains the codes, input and output of the analysis carried out for supporting the publication of the paper:</p> <p>Tarpanelli A., Mondini A., Camici S.:Effectiveness of Sentinel-1 and Sentinel-2 for Flood Detection Assessment in Europe, Natural Hazards and Earth System Sciences, https://doi.org/10.5194/nhess-2022-63, 2022.</p> <p>&nbsp;</p> <p>The codes should be run in order A1-A7 to generate all the figures of the paper.</p> <p>For details please send an email to:</p> <p>angelica.tarpanelli@irpi.cnr.it</p> <p>&nbsp;</p>

opencc-by-4.0Jul 2022View details →
zenodo44/100

Olive orchard stress assessment with supplementary Sentinel-2 data

<p>This file contains&nbsp;ground truthing data from olive orchards in Halkidiki, N.Greece and Sentinel-2 data for the noted samples. The samples collected for ground truthing are polygons inside the borders of olive orchards in Halkidiki. Polygons or Sampling units contain information associated with biotic and abiotic stress-related assessments carried out by visual inspection and laboratory analysis of samples with ongoing symptoms. Assessments were recorded as percentages of symptoms observed in the total vegetation surface present in each sampling unit. Symptom percentages were attributed to three possible classes: Verticillium dahliae,&nbsp;Spilocaea oleaginea,&nbsp;Unidentified Stress Factors. These percentage values were used to characterize healthy trees and the incidence of V. dahliae, S. oleaginea and unidentified stress factors (USF) in the sample. USF was used for all other non-classified surveyed symptoms attributed to diseases, pests, or abiotic-related damage. Percentages were summed to compute the total stress present in each sampling unit. &ldquo;Total stress&rdquo; refers to the stress incidence value used together with different thresholds to create binary labels for each sample of &ldquo;stressed&rdquo; or &ldquo;not stressed&rdquo;.</p> <p>The polygon geographical information for each sample were recorded and stored in shapefile format&nbsp;using SW maps, a mobile mapping and GIS app.</p> <p>Sentinel-2 data was paired with&nbsp;each sample using the Feature Info Service (FIS) available from sentinel hub, now upgraded into the Statistical API tool (&nbsp;&amp; ).&nbsp;This API enables&nbsp;acquisition of statistics calculated based on satellite imagery without having to download images. In the Statistical API request&nbsp;the area of interest, time period, evalscript and statistical measures of interest can be calculated. The requested statistics are returned in the API response.</p> <p>Statistical API deployments:</p> <p><a href="https://creodias.sentinel-hub.com/api/v1/statistics">https://creodias.sentinel-hub.com/api/v1/statistics</a></p> <p><a href="https://services.sentinel-hub.com/api/v1/statistics">https://services.sentinel-hub.com/api/v1/statistics</a></p> <p><a href="https://services-uswest2.sentinel-hub.com/api/v1/statistics">https://services-uswest2.sentinel-hub.com/api/v1/statistics</a></p>

opencc-by-4.0Oct 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record