Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
181
datasets available to search
ShareScore release 0.7.1
Dataset results
181 results for “SENTINEL-2”
Doodleverse/Segmentation Zoo Res-UNet models for 2-class (water, other) segmentation of Sentinel-2 and Landsat-7/8 3-band (RGB) images of coasts.
<p><em><strong>Doodleverse/Segmentation Zoo Res-UNet models for 2-class (water, other) segmentation of Sentinel-2 and Landsat-7/8 3-band (RGB) images of coasts.</strong></em></p> <p><em><strong>Version 3: Updated 2023-04-25</strong></em></p> <p>These Residual-UNet model data are based on RGB (red, green, and blue) images of coasts and associated labels.</p> <p>Models have been created using Segmentation Gym* using the following dataset**: <a href="https://doi.org/10.5281/zenodo.7384242">https://doi.org/10.5281/zenodo.7384242</a></p> <p>Classes: {0=other, 1=water}</p> <p><strong>File descriptions</strong></p> <p>For each model, there are 5 files with the same root name:</p> <p>1. <strong>'.json' </strong>config file: this is the file that was used by Segmentation Gym* to create the weights file. It contains instructions for how to make the model and the data it used, as well as instructions for how to use the model for prediction. It is a handy wee thing and mastering it means mastering the entire Doodleverse.</p> <p>2.<strong> '.h5'</strong> weights file: this is the file that was created by the Segmentation Gym* function `train_model.py`. It contains the trained model's parameter weights. It can called by the Segmentation Gym* function `seg_images_in_folder.py`. Models may be ensembled.</p> <p>3.<strong> '_modelcard.json'</strong> model card file: this is a json file containing fields that collectively describe the model origins, training choices, and dataset that the model is based upon. There is some redundancy between this file and the `config` file (described above) that contains the instructions for the model training and implementation. The model card file is not used by the program but is important metadata so it is important to keep with the other files that collectively make the model and is such is considered part of the model</p> <p>4. <strong> '_model_history.npz'</strong> model training history file: this numpy archive file contains numpy arrays describing the training and validation losses and metrics. It is created by the Segmentation Gym function `train_model.py`</p> <p>5. <strong> '.png'</strong> model training loss and mean IoU plot: this png file contains plots of training and validation losses and mean IoU scores during model training. A subset of data inside the .npz file. It is created by the Segmentation Gym function `train_model.py`</p> <p>Additionally, BEST_MODEL.txt contains the name of the model with the best validation loss and mean IoU</p> <p> </p> <p><strong>References</strong></p> <p>*Segmentation Gym: Buscombe, D., & Goldstein, E. B. (2022). A reproducible and reusable pipeline for segmentation of geoscientific imagery. Earth and Space Science, 9, e2022EA002332. <a href="https://doi.org/10.1029/2022EA002332">https://doi.org/10.1029/2022EA002332</a> See: <a href="https://github.com/Doodleverse/segmentation_gym">https://github.com/Doodleverse/segmentation_gym</a></p> <p>** Buscombe, D. (2022). Images and 2-class labels for semantic segmentation of Sentinel-2 and Landsat RGB satellite images of coasts (water, other) (v1.0) [Data set]. Zenodo. <a href="https://doi.org/10.5281/zenodo.7384242">https://doi.org/10.5281/zenodo.7384242</a></p>
FORMS: Forest Multiple Source height, wood volume, and biomass maps in France at 10 to 30 m resolution based on Sentinel-1, Sentinel-2, and GEDI data with a deep learning approach.
<p>The products can be vizualized at <a href="https://martinschwartz0.users.earthengine.app/view/forms-height-biomass-volume-viewer">https://martinschwartz0.users.earthengine.app/view/forms-height-biomass-volume-viewer</a></p> <p>- FORMS-H: Canopy height map of France at 10 m resolution. The units are in centimeter (10^-2 m).</p> <p>- FORMS-B: Above-ground biomass density map of France at 30 m resolution. The units are in Mg ha-1</p> <p>- FORMS-V: Wood volume density map of France at 30 m resolution. The units are in m3 ha-1</p> <p>Please refer to the paper <a href="https://doi.org/10.5194/essd-15-4927-2023">https://doi.org/10.5194/essd-15-4927-2023</a> for further details.</p>
June 2023 Supplement Images and 4-class labels for semantic segmentation of Sentinel-2 and Landsat RGB, NIR, and SWIR satellite images of coasts (water, whitewater, sediment, other)
<p><strong>June 2023 Supplement of Images and 4-class labels for semantic segmentation of Sentinel-2 and Landsat RGB, NIR, and SWIR satellite images of coasts (water, whitewater, sediment, other)</strong></p> <p><strong>Description</strong></p> <p>Supplementary dataset to:</p> <p>Buscombe, Daniel, Goldstein, Evan, Bernier, Julie, Bosse, Stephen, Colacicco, Rosa, Corak, Nick, Fitzpatrick, Sharon, del Jesús González Guillén, Anais, Ku, Venus, Paprocki, Julie, Platt, Lindsay, Steele, Bethel, Wright, Kyle, & Yasin, Brandon. (2022). Images and 4-class labels for semantic segmentation of Sentinel-2 and Landsat RGB satellite images of coasts (water, whitewater, sediment, other) (v1.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.7335647</p> <p>This supplemental dataset consists of 283 RGB images and 283 associated labels for semantic segmentation of Sentinel-2 and Landsat RGB satellite images of coasts. Of these, 77 images-label pairs also have a corresponding NIR and SWIR satellite image. The 4 classes are 0=water, 1=whitewater, 2=sediment, 3=other</p> <p>These images and labels have been made using the Doodleverse software package, Doodler*. These images and labels could be used within numerous Machine Learning frameworks for image segmentation, but have specifically been made for use with the Doodleverse software package, Segmentation Gym**.</p> <p>Imagery are a mixture of 10-m Sentinel-2 and 15-m pansharpened Landsat 7, 8, and 9 visible-band imagery of various sizes. NIR, SWIR, Red, Green, and Blue bands only</p> <p><strong>File descriptions</strong></p> <ol> <li>classes.txt, a file containing the class names</li> <li>images.zip, a zipped folder containing the 3-band images of varying sizes and extents</li> <li>labels.zip, a zipped folder containing the 1-band label images</li> <li>overlays.zip, a zipped folder containing a semi-transparent overlay of the color-coded label on the image (blue=0=water, red=1=whitewater, yellow=2=sediment, green=3=other)</li> <li>nir.zip</li> <li>swir.zip</li> </ol> <p><strong>References</strong></p> <p>Buscombe, Daniel, Goldstein, Evan, Bernier, Julie, Bosse, Stephen, Colacicco, Rosa, Corak, Nick, Fitzpatrick, Sharon, del Jesús González Guillén, Anais, Ku, Venus, Paprocki, Julie, Platt, Lindsay, Steele, Bethel, Wright, Kyle, & Yasin, Brandon. (2022). Images and 4-class labels for semantic segmentation of Sentinel-2 and Landsat RGB satellite images of coasts (water, whitewater, sediment, other) (v1.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.7335647</p> <p>*Doodler: Buscombe, D., Goldstein, E.B., Sherwood, C.R., Bodine, C., Brown, J.A., Favela, J., Fitzpatrick, S., Kranenburg, C.J., Over, J.R., Ritchie, A.C. and Warrick, J.A., 2021. Human‐in‐the‐Loop Segmentation of Earth Surface Imagery. Earth and Space Science, p.e2021EA002085<a href="https://doi.org/10.1029/2021EA002085">https://doi.org/10.1029/2021EA002085</a>. See <a href="https://github.com/Doodleverse/dash_doodler">https://github.com/Doodleverse/dash_doodler.</a></p> <p>**Segmentation Gym: Buscombe, D., & Goldstein, E. B. (2022). A reproducible and reusable pipeline for segmentation of geoscientific imagery. Earth and Space Science, 9, e2022EA002332. <a href="https://doi.org/10.1029/2022EA002332">https://doi.org/10.1029/2022EA002332</a> See: <a href="https://github.com/Doodleverse/segmentation_gym">https://github.com/Doodleverse/segmentation_gym</a></p> <p> </p>
Doodleverse/CoastSeg Segformer models for 4-class (water, whitewater, sediment and other) segmentation of Sentinel-2 and Landsat-7/8 3-band (RGB) images of coasts.
<p><em><strong>Doodleverse/CoastSeg Segformer models for 4-class (water, whitewater, sediment and other) segmentation of Sentinel-2 and Landsat-7/8 3-band (RGB) images of coasts.</strong></em></p> <p>These Segformer model data are based on RGB (red, green, and blue) images of coasts and associated labels.</p> <p>Models have been created using Segmentation Gym* using the following datasets**: <a href="https://doi.org/10.5281/zenodo.7335647">https://doi.org/10.5281/zenodo.7335647</a> and ***https://doi.org/10.5281/zenodo.8011926. Those datasets have been combined and the training and validation images and labels are provided here.</p> <p>Classes: {0=water, 1=whitewater, 2=sediment, 3=other}</p> <p><strong>Model validation accuracy statistics</strong></p> <p>model name: overall accuracy, mean frequency weighted IoU, mean IoU, Matthews correlation. Bold indicates best overall</p> <ul> <li><strong>v5: .94, .90, .64, .87</strong></li> <li>v6: .93, .89, .63, .87</li> <li>v7: .92, .88, .61, .84</li> <li>v8: .93, .89, .63, .87</li> <li>v9: .92, .88, .62, .85</li> <li>v10: .93, .89, .63, .86</li> </ul> <p> </p> <p><strong>File descriptions</strong></p> <p>For each model, there are 5 files with the same root name:</p> <p>1. <strong>'.json' </strong>config file: this is the file that was used by Segmentation Gym* to create the weights file. It contains instructions for how to make the model and the data it used, as well as instructions for how to use the model for prediction. It is a handy wee thing and mastering it means mastering the entire Doodleverse.</p> <p>2.<strong> '.h5'</strong> weights file: this is the file that was created by the Segmentation Gym* function `train_model.py`. It contains the trained model's parameter weights. It can called by the Segmentation Gym* function `seg_images_in_folder.py`. Models may be ensembled.</p> <p>3.<strong> '_modelcard.json'</strong> model card file: this is a json file containing fields that collectively describe the model origins, training choices, and dataset that the model is based upon. There is some redundancy between this file and the `config` file (described above) that contains the instructions for the model training and implementation. The model card file is not used by the program but is important metadata so it is important to keep with the other files that collectively make the model and is such is considered part of the model</p> <p>4. <strong> '_model_history.npz'</strong> model training history file: this numpy archive file contains numpy arrays describing the training and validation losses and metrics. It is created by the Segmentation Gym function `train_model.py`</p> <p>5. <strong> '.png'</strong> model training loss and mean IoU plot: this png file contains plots of training and validation losses and mean IoU scores during model training. A subset of data inside the .npz file. It is created by the Segmentation Gym function `train_model.py`</p> <p>Additionally, BEST_MODEL.txt contains the name of the model with the best validation loss and mean IoU</p> <p>This is a sister model to this set of Residual UNets: Buscombe, Daniel. (2022). Doodleverse/Segmentation Zoo Res-UNet models for 4-class (water, whitewater, sediment and other) segmentation of Sentinel-2 and Landsat-7/8 3-band (RGB) images of coasts. (v1.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.6950472</p> <p><strong>References</strong></p> <p>*Segmentation Gym: Buscombe, D., & Goldstein, E. B. (2022). A reproducible and reusable pipeline for segmentation of geoscientific imagery. Earth and Space Science, 9, e2022EA002332. <a href="https://doi.org/10.1029/2022EA002332">https://doi.org/10.1029/2022EA002332</a> See: <a href="https://github.com/Doodleverse/segmentation_gym">https://github.com/Doodleverse/segmentation_gym</a></p> <p>** Buscombe, Daniel, Goldstein, Evan, Bernier, Julie, Bosse, Stephen, Colacicco, Rosa, Corak, Nick, Fitzpatrick, Sharon, del Jesús González Guillén, Anais, Ku, Venus, Paprocki, Julie, Platt, Lindsay, Steele, Bethel, Wright, Kyle, & Yasin, Brandon. (2022). Images and 4-class labels for semantic segmentation of Sentinel-2 and Landsat RGB satellite images of coasts (water, whitewater, sediment, other) (v1.0) [Data set]. Zenodo. <a href="https://doi.org/10.5281/zenodo.7335647">https://doi.org/10.5281/zenodo.7335647</a></p> <p>***Buscombe, Daniel. (2023). June 2023 Supplement Images and 4-class labels for semantic segmentation of Sentinel-2 and Landsat RGB, NIR, and SWIR satellite images of coasts (water, whitewater, sediment, other) (v1.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.8011926 </p> <p> </p> <p> </p>
Doodleverse/CoastSeg Segformer models for 4-class (water, whitewater, sediment and other) segmentation of Sentinel-2 and Landsat-7/8 MNDWI images of coasts.
<p><em><strong>Doodleverse/CoastSeg Segformer models for 4-class (water, whitewater, sediment and other) segmentation of Sentinel-2 and Landsat-7/8 MNDWI images of coasts.</strong></em></p> <p>Models have been created using Segmentation Gym* using the following datasets ** https://zenodo.org/record/7384263 and ***: https://doi.org/10.5281/zenodo.7335647. Those datasets have been combined and the training and validation images and labels are provided here.</p> <p>Classes: {0=water, 1=whitewater, 2=sediment, 3=other}</p> <p><strong>Model validation accuracy statistics</strong></p> <p>model name: overall accuracy, mean frequency weighted IoU, mean IoU, Matthews correlation. Bold indicates best overall</p> <p> v2: 0.808, 0.7309, 0.47864, 0.656<br> v3: 0.809, 0.7302, 0.4982, 0.664</p> <p><strong>File descriptions</strong></p> <p>For each model, there are 5 files with the same root name:</p> <p>1. '.json' config file: this is the file that was used by Segmentation Gym* to create the weights file. It contains instructions for how to make the model and the data it used, as well as instructions for how to use the model for prediction. It is a handy wee thing and mastering it means mastering the entire Doodleverse.</p> <p>2. '.h5' weights file: this is the file that was created by the Segmentation Gym* function `train_model.py`. It contains the trained model's parameter weights. It can called by the Segmentation Gym* function `seg_images_in_folder.py`. Models may be ensembled.</p> <p>3. '_modelcard.json' model card file: this is a json file containing fields that collectively describe the model origins, training choices, and dataset that the model is based upon. There is some redundancy between this file and the `config` file (described above) that contains the instructions for the model training and implementation. The model card file is not used by the program but is important metadata so it is important to keep with the other files that collectively make the model and is such is considered part of the model</p> <p>4. '_model_history.npz' model training history file: this numpy archive file contains numpy arrays describing the training and validation losses and metrics. It is created by the Segmentation Gym function `train_model.py`</p> <p>5. '.png' model training loss and mean IoU plot: this png file contains plots of training and validation losses and mean IoU scores during model training. A subset of data inside the .npz file. It is created by the Segmentation Gym function `train_model.py`</p> <p>Additionally, BEST_MODEL.txt contains the name of the model with the best validation loss and mean IoU</p> <p>This is a sister model to these sets of Residual UNets:</p> <p> https://zenodo.org/record/7352850<br> https://zenodo.org/record/7557080</p> <p><strong>References</strong></p> <p>*Segmentation Gym: Buscombe, D., & Goldstein, E. B. (2022). A reproducible and reusable pipeline for segmentation of geoscientific imagery. Earth and Space Science, 9, e2022EA002332. https://doi.org/10.1029/2022EA002332 See: https://github.com/Doodleverse/segmentation_gym</p> <p>** Buscombe, Daniel. (2022). Images and 2-class labels for semantic segmentation of Sentinel-2 and Landsat RGB, NIR, and SWIR satellite images of coasts (water, other) (v1.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.7384263</p> <p>***Coast Train data release: Wernette, P.A., Buscombe, D.D., Favela, J., Fitzpatrick, S., and Goldstein E., 2022, Coast Train--Labeled imagery for training and evaluation of data-driven models for image segmentation: U.S. Geological Survey data release, https://doi.org/10.5066/P91NP87I. See https://coasttrain.github.io/CoastTrain/ for more information</p> <p>***Buscombe, Daniel. (2023). June 2023 Supplement Images and 4-class labels for semantic segmentation of Sentinel-2 and Landsat RGB, NIR, and SWIR satellite images of coasts (water, whitewater, sediment, other) (v1.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.8011926 </p>
Doodleverse/CoastSeg Segformer models for 4-class (water, whitewater, sediment and other) segmentation of Sentinel-2 and Landsat-7/8 NDWI images of coasts.
<p><strong>Doodleverse/CoastSeg Segformer models for 4-class (water, whitewater, sediment and other) segmentation of Sentinel-2 and Landsat-7/8 NDWI images of coasts.</strong></p> <p>Models have been created using Segmentation Gym* using the following datasets ** https://zenodo.org/record/7384263 and ***: https://doi.org/10.5281/zenodo.7335647. Those datasets have been combined and the training and validation images and labels are provided here.</p> <p>Classes: {0=water, 1=whitewater, 2=sediment, 3=other}</p> <p><strong>Model validation accuracy statistics</strong></p> <p>model name: overall accuracy, mean frequency weighted IoU, mean IoU, Matthews correlation. Bold indicates best overall</p> <table> <tbody> <tr> <td>0.896016693115234</td> <td>0.832759195999637</td> <td>0.565748652153519</td> <td>0.806139409136944</td> </tr> </tbody> </table> <table> <tbody> <tr> <td>0.906008201175266</td> <td>0.847625161837392</td> <td>0.593821675882991</td> <td>0.819790462192222</td> </tr> </tbody> </table> <table> <tbody> <tr> <td>0.903999212053087</td> <td>0.844255821722932</td> <td>0.577444030164045</td> <td>0.813646408575</td> </tr> </tbody> </table> <p> </p> <p><strong>File descriptions</strong></p> <p>For each model, there are 5 files with the same root name:</p> <p>1. '.json' config file: this is the file that was used by Segmentation Gym* to create the weights file. It contains instructions for how to make the model and the data it used, as well as instructions for how to use the model for prediction. It is a handy wee thing and mastering it means mastering the entire Doodleverse.</p> <p>2. '.h5' weights file: this is the file that was created by the Segmentation Gym* function `train_model.py`. It contains the trained model's parameter weights. It can called by the Segmentation Gym* function `seg_images_in_folder.py`. Models may be ensembled.</p> <p>3. '_modelcard.json' model card file: this is a json file containing fields that collectively describe the model origins, training choices, and dataset that the model is based upon. There is some redundancy between this file and the `config` file (described above) that contains the instructions for the model training and implementation. The model card file is not used by the program but is important metadata so it is important to keep with the other files that collectively make the model and is such is considered part of the model</p> <p>4. '_model_history.npz' model training history file: this numpy archive file contains numpy arrays describing the training and validation losses and metrics. It is created by the Segmentation Gym function `train_model.py`</p> <p>5. '.png' model training loss and mean IoU plot: this png file contains plots of training and validation losses and mean IoU scores during model training. A subset of data inside the .npz file. It is created by the Segmentation Gym function `train_model.py`</p> <p>Additionally, BEST_MODEL.txt contains the name of the model with the best validation loss and mean IoU</p> <p>This is a sister model to these sets of Residual UNets:</p> <p>https://zenodo.org/record/7557072<br> https://zenodo.org/record/7352859</p> <p><strong>References</strong></p> <p>*Segmentation Gym: Buscombe, D., & Goldstein, E. B. (2022). A reproducible and reusable pipeline for segmentation of geoscientific imagery. Earth and Space Science, 9, e2022EA002332. https://doi.org/10.1029/2022EA002332 See: https://github.com/Doodleverse/segmentation_gym</p> <p>** https://zenodo.org/record/7384263</p> <p>***Buscombe, Daniel. (2023). June 2023 Supplement Images and 4-class labels for semantic segmentation of Sentinel-2 and Landsat RGB, NIR, and SWIR satellite images of coasts (water, whitewater, sediment, other) (v1.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.8011926 </p>
Texas 2022 water clarity and color (FLAMe and Sentinel-2)
Water clarity and color were determined for six reservoirs using rapid spatial surveys from a sensor equipped boat and concurrent Sentinel-2 satellite imagery across Texas during drought conditions between the months of July and August 2022. From west to east, these systems include Red Bluff Reservoir, O.H. Ivie Lake, Lake Arrowhead, Lake Brownwood, Lake Waco, and Lake Bonham. For the water year leading up to the sampling dates, the precipitation ranged from 182 mm in Red Bluff Reservoir to 1036 mm in Lake Bonham. A total of 254 km of boat path were covered across the six reservoirs with a mean boat speed of 19.17 km/h. The data for this study covers three spatial approaches 1) along the boat path 2) longitudinal transects from dam to river arm and 3) whole system. For the boat path, data variables include turbidity measured continuously with a YSI EXO2 sonde, Secchi disk depth predicted from the turbidity values, normalized difference turbidity index (NDTI), and dominant wavelength. For both the longitudinal transects and whole system data, variables include the two remotely derived measures of clarity and color, NDTI and dominant wavelength. Data is also categorized by zone as either "arm" (reservoir arm) or "body" (main body) determined by a 4m depth threshold to compare between zones.
Dataset of processed Sentinel-2 images for chlorophyll-a estimation in high-altitude lakes in the Sierra Nevada, Spain
<p>This dataset contains Sentinel 2 satellite images clipped to 5 high-altitude lakes in the Sierra Nevada Mountain Range, Spain. The images were processed with the following atmospheric correction algorithms:</p><ul><li><a href="https://c2rcc.org/">C2RCC</a> (<a href="https://ui.adsabs.harvard.edu/abs/2016ESASP.740E..54B/abstract">Brockmann et al. 2016</a>)</li><li><a href="https://github.com/MarcYin/SIAC">SIAC</a> (<a href=" https://doi.org/10.5194/gmd-15-7933-2022">Yin et al. 2022)</a></li><li><a href="https://github.com/acolite/acolite/releases/tag/20221114.0">ACOLITE</a> (<a href="https://doi.org/10.1016/j.rse.2018.07.015">Vanhellemont & Ruddick, 2018</a>)</li><li><a href="https://grass.osgeo.org/grass83/manuals/i.atcorr.html">6SV</a> (<a href="https://doi.org/10.1109/36.581987">Vermote et al. 2006</a>)</li></ul><p><strong>Included Lakes and and their IDs:</strong></p><ul><li>Laguna de la Caldera (ID = P-2)</li><li>Laguna-embalse de las Yeguas (ID = D-6)</li><li>Laguna de Río Seco (ID = P-8)</li><li>Laguna Larga (ID = G-7)</li><li>Laguna de la Mosca (ID = G-11)</li></ul>
Cartographie fine de la couverture de Macrocystis pyrifera sur l'archipel de Kerguelen à partir d'images Sentinel-2 acquises en 2020
<p>Cette donnée représente une cartographie fine de la couverture de Macrocystis pyrifera (MP) sur l’archipel de Kerguelen à partir d'images Sentinel-2 acquises en 2020. Etant donné la proximité de MP avec Durvilleae Antarctica (DA) (du point de vue de leur signature spectrale), les résultats intègrent aussi cette dernière.</p> <p>Les données d’Observation de la Terre utilisées sont les images satellitaires libres et gratuites Sentinel-2 (A et B). La chaîne de traitement mise en place est entièrement basée sur le langage R. Les téléchargements et prétraitements ont été effectués grâce à une librairie complète, sen2r, créée par Luigi Ranghetti (Ranghetti et al., 2020).<br>Les résultats de détection des algues brunes à partir de la chaîne sont proches de l’exhaustivité avec un Kappa global de 0.993. C’est l’indice du grNDVI de Wang et al. (2007) qui a été utilisé pour réaliser ces produits.</p> <p>Un produit final utilisable par un utilisateur n’ayant pas forcément été initié à la télédétection a été favorisé. Dans ce but une couche shapefile représentant la présence moyenne d’algue brune sur la zone (c.à.d. sur toute l’étendue temporelle disponible) est ici mise à disposition.</p> <p>Cette donnée a été réalisée dans le cadre du stage de master2 d'Alexis Pré co-encadré par l'UMR Espace-Dev à la station SEAS-OI et le service SIG des Terres Australes et Antarctiques Françaises.</p> <p>Wang, F. M., Huang, J. F., Tang, Y. L., & Wang, X. Z. (2007). New vegetation index and its application in estimating leaf area index of rice. <em>Rice Science</em>, <em>14</em>(3), 195-203.</p> <p>Ranghetti, L., Boschetti, M., Nutini, F., & Busetto, L. (2020). “sen2r”: An R toolbox for automatically downloading and preprocessing Sentinel-2 satellite data. <em>Computers & Geosciences</em>, <em>139</em>, 104473.</p>
Mapping sugarcane globally at 10 m resolution using GEDI and Sentinel-2
<p><strong>Dataset Abstract:</strong><br>Sugarcane is an important source of food, biofuel, and farmer income in many countries. At the same time, sugarcane is implicated in many social and environmental challenges, including water scarcity and nutrient pollution. Currently, few of the top sugar-producing countries generate reliable maps of where sugarcane is cultivated. To fill this gap, we introduce a dataset of detailed sugarcane maps for the top 13 producing countries in the world, comprising nearly 90% of global production. Maps were generated for the 2019-2022 period by combining data from the Global Ecosystem Dynamics Investigation (GEDI) and Sentinel-2 (S2). GEDI data were used to provide training data on where tall and short crops were growing each month, while S2 features were used to map tall crops for all cropland pixels each month. Sugarcane was then identified by leveraging the fact that sugar is typically the only tall crop growing for a substantial fraction of time during the study period. Comparisons with field data, pre-existing maps, and official government statistics all indicated high precision and recall of our maps. Agreement with field data at the pixel level exceeded 80% in most countries, and sub-national sugarcane areas from our maps were consistent with government statistics. Exceptions appeared mainly due to problems in underlying cropland masks, or to under-reporting of sugarcane area by governments. <br>The final maps should be useful in studying the various impacts of sugarcane cultivation and producing maps of related outcomes such as sugarcane yields.</p> <p><strong>USAGE: Users must mask the provided sugarcane map with the most appropriate crop mask from the ones provided. If none of the provided crop masks are suitable, users can use an external crop mask instead.</strong></p> <p>Validation results for the sugarcane maps are detailed in Section 4.3 of the paper. For Indonesia and Guatemala, no field-level data or raster datasets were available for validation of our sugarcane maps.</p> <p><br><strong>Dataset:</strong> <br>5 bands<br>b1: Number of tall months<br>b2: Sugarcane Map: 0 = non-sugarcane, 1 = sugarcane<br>b3: ESA crop mask: 0 = non-cropland, 1 = cropland<br>b4: ESRI crop mask: 0 = non-cropland, 1 = cropland<br>b5: GLAD crop mask: 0 = non-cropland, 1 = cropland</p> <p> </p> <p>The dataset can be accessed on Google Earth Engine (GEE) at <br><strong><a href="https://code.earthengine.google.com/?asset=projects/lobell-lab/gedi_sugarcane/maps/imgColl_10m_ESAESRIGLAD">https://code.earthengine.google.com/?asset=projects/lobell-lab/gedi_sugarcane/maps/imgColl_10m_ESAESRIGLAD</a><br></strong><br>Example GEE script for visualizing and masking the sugarcane maps by country available at:<br><strong><a href="https://code.earthengine.google.com/545a87ce9bc29f2b5ad180955d974f8c?asset=projects%2Flobell-lab%2Fgedi_sugarcane%2Fmaps%2FimgColl_10m_ESAESRIGLAD">https://code.earthengine.google.com/545a87ce9bc29f2b5ad180955d974f8c?asset=projects%2fl Bell-lab%2Fgedi_sugarcane%2 Maps%2FimgColl_10m_ESAESRIGLAD</a></strong></p>
Seasonal RGB composites from Sentinel-2 (2017-2024) for Catalonia, Spain; Sétif, Algeria; Behia and Kafr Elsheihk Governates, Egypt; Marseille, France; Sicily, Italy.
<p>A dataset containing seasonal Sentinel-2 RGB images (2017 to 2024) for five case study areas of the TRANSITION project (https://www.transition-med.org/), funded by PRIMA (https://prima-med.org/). The areas are Catalonia, Spain; Sétif, Algeria; Behia and Kafr Elsheihk Governates, Egypt; Marseille, France; Sicily, Italy. The dataset can be useful for anyone looking to conduct agriculture-related research using Earth Observation data in these five areas.</p>
Evaluation of Spatiotemporal Fusion Methods Using Sentinel-2 And Sentinel-3: A New Benchmark Dataset And Comparison
<p>In Earth observation, data fusion is important to generate high temporal and spatial resolution images. Nevertheless, existing research on data fusion primarily concentrates on merging two sources of data (mostly MODIS and Landsat). Therefore, we offer the community a new benchmark dataset for evaluating data fusion using new European sensors (Sentinel-2 and Sentinel-3).</p> <p>The dataset is composed of three different sites located in different parts of the world to ensure the diversity of the ecosystem. The two components of the dataset are collected from operating missions ( Sentinel-2 and Sentinel-3). We also provide 10 bands for Sentinel-2 ranging from blue to SWIR, 4 bands at 10m resolution and 6 at 20m resolution. For Sentinel-3 16 bands are provided with a spatial resolution of 300m. The multiple bands allow for different applications for this dataset such as testing data fusion methods, etc.</p>
Mapping Tree Species Fractions in Temperate Mixed Forests Using Sentinel-2 Time Series and Synthetically Mixed Training Data
<p>This dataset contains the latest version of a selection of result data of the paper "Mapping Tree Species Fractions in Temperate Mixed Forests Using Sentinel-2 Time Series and Synthetically Mixed Training Data" (DOI: https://doi.org/10.1016/j.rse.2025.114740 )</p> <p>The dataset contains:</p> <ol> <li>A geopackage of training points of pure tree species</li> <li>The resulting 12-band tree species fraction map of Rhineland-Palatinate</li> <li>HSV-colored map of dominant tree species. For information which tree species are represented by the different colors, refer to the Supplemental in the original paper.</li> <li>CSV-table of predicted and reference propotion of the tree species in the validation polygon (the original polygon data can not be published due to data privacy regulations) </li> </ol> <p> </p>
Mapping canopy cover in African dry forests from combined use of Sentinel-1 and Sentinel-2 data: 2018 maps for Tanzania
<p>The monitoring of tropical forests has benefited from the increased availability of high-resolution earth observation data. However, the seasonality and openness of the canopy of dry tropical forests remains a challenge for optical sensors. The availability of time series of remote sensing images at 10-meters is changing this paradigm.</p> <p>In the context of REDD+ national reporting requirements, we investigated a methodology that is reproducible and adaptable in order to ensure user appropriation. The overall methodology consists of three main steps: (i) the generation of Sentinel-1 (S1) and Sentinel-2 (S2) layers, (ii) the collection of an ad-hoc training/validation dataset and (iii) the classification of the satellite data. Three different classification workflows are compared in terms of their capability to capture the canopy cover of forests in East Africa. Two types of maps are derived from these mapping approaches: i) binary tree cover/no tree cover (TC/NTC) maps, and ii) maps of canopy cover classes. The method is applied at scale, over Tanzania and one final map for each workflow is shared. Two big data computing platforms are combined to exploit the important volume of satellite data available over a yearly period.</p> <p>The reference dataset (training and validation), the three best maps and the codes to produce the S1 and S2 composites on Google Earth Engine are shared here.</p> <p>The folder “reference_dataset.zip” contains the expert based training and validation dataset. The point shapefile corresponding to the center of the plot as well as the 3x3 and 5x5 polygon shapefile are shared together with qml layer file for each type of shapefile.</p> <p>Three maps (binary TC-NTC “pixel” RF, forest type “pixel” RF and “window” ETC) are shared. A 40 km buffer from national boundaries is kept in order to let users refine their area of interest. The qml style file are also shared.</p> <p>In the “script.zip” folder, the javascript codes to generate the S1 and S2 mosaics are shared.</p>
Sentinel-2 Madagascar mangrove cover density
<p>These datasets are Madagascar mangrove cover density products, called Sentinel-2 Madagascar mangrove cover density (2016 and 2018), based on 64 tiles analysis of high resolution satellite image Sentinel-2 (L2A Level). Data images are acquired at 1C level and pre-processed at the SEAS-OI station, Réunion island. We used an object-based image analysis method based on NDVI value to classify three level of mangrove density (very dense, dense and less dense). The attribute table is composed of the object identifier, the vegetation density and the administrative region to which the object belongs. The analysis method at national scale was based on field data collection in three reference sites: in north-western of Madagascar (Mahajamba Bay representing 10% of the malagasy mangrove and Mahavavy Delta ) and in south-western of Madagascar (Tsingilofilo Bay). The overall accuracy of the analysis exceeds 90% for all three sites making it an operational product. This data has been realized in the framework of a thesis at the University of Reunion within the UMR Espace-Dev. This thesis was financed by the Région Réunion</p>
Building fraction map of Germany (Sentinel-1/-2 based, 10m and 100m resolution)
<p>This dataset features a map of building fractions (as opposed to built-up fractions including other impervious surfaces such as roads) for Germany on a 10m grid based on Sentinel-1A/B and Sentinel-2A/B time series. The data were created by using machine learning regression and spectral unmixing, using synthetically mixed training data. The dataset is completely based on freely accessible satellite imagery, and was validated with freely available building footprint reference data for three federal states.</p> <p>We recommend to use data at an aggregated resolution of 20m, 50m, or 100m, and to clip data at about 20% building fraction when using 10m resolution maps (or roughly the corresponding RMSE at any other resolution).</p> <p><strong>Temporal extent</strong><br> Used Sentinel-2 data were acquired in 2018, and Sentinel-1 data were acquired in 2017 (see publication). The map is, thus, representative for 2017/2018. Validation results can be affected by building footprint reference data from different years.</p> <p><strong>Data format</strong><br> The data come in tiles of 30x30km (see shapefile). The projection is EPSG:3035. The images are compressed GeoTiff files (*.tif). There is a mosaic in GDAL Virtual format (*.vrt), which can readily be opened in most Geographic Information Systems. Building fraction values are in percent, from 0 to 100. In the original dataset with 10m spatial resolution, fraction values are equivalent to area in m². In the aggregated dataset with 100m spatial resolution, the values must be multiplied with 100 in order to see area in m².</p> <p><strong>Further information</strong><br> For further information, please see the publication or contact Franz Schug (franz.schug@geo.hu-berlin.de). A web-visualization of this dataset is available <a href="https://ows.geo.hu-berlin.de/webviewer/building-area/">here</a>.</p> <p><strong>Publication</strong><br> Schug, F.; Frantz, D.; Okujeni, A.; Hostert, P. (2022). Sub-pixel building area mapping based on synthetic training data and regression-based unmixing using Sentinel-1 and -2 data. Remote Sensing Letters. DOI: 10.1080/2150704X.2022.2088253</p> <p><strong>Acknowledgements</strong><br> The dataset was generated by FORCE v. 3.6.1 (<a href="https://doi.org/10.3390/rs11091124">paper</a>, <a href="https://github.com/davidfrantz/force">code</a>), which is freely available software under the terms of the GNU General Public License v. >= 3. Sentinel imagery were obtained from the <a href="https://scihub.copernicus.eu/">European Space Agency and the European Commission</a>. Sentinel-1 data were provided by <a href="https://eodc.eu/">EODC</a>. We thank the providers of the building footprint reference data (see publication).</p> <p><strong>Funding</strong><br> This dataset was produced with funding from the European Research Council (ERC) under the European Union's Horizon 2020 research and innovation programme (<a href="https://boku.ac.at/understanding-the-role-of-material-stock-patterns-for-the-transformation-to-a-sustainable-society-mat-stocks">MAT_STOCKS</a>, grant agreement No 741950).</p> <p> </p>
Satellite-derived chlorophyll-a concentrations for Lake Hume (Australia) using Mixture Density Networks and Sentinel-2 and Landsat 8 imagery
<p>This dataset contains satellite-derived chlorophyll-a data of Lake Hume (Australia) for the period 21 Mar. 2013 - 01 Feb. 2021. Chlorophyll-a concentrations have been calculated using Mixture Density Networks and Sentinel-2 and Landsat 8 imagery.</p> <p>Mixture Density Networks are a class of neural networks that tackle the inverse problem by modelling the multimodal distribution of target variables using a mixture of Gaussians. For more information, please refer to the following:</p> <ul> <li>Pahlevan, N., Smith, B., Alikas, K., Anstee, J., et al. (2022). Simultaneous retrieval of selected optical water quality indicators from Landsat-8, Sentinel-2, and Sentinel-3. <em>Remote Sensing of Environment, 270</em>, 112860</li> <li>Smith, B., Pahlevan, N., Schalles, J., et al. (2021). A Chlorophyll-a Algorithm for Landsat-8 Based on Mixture Density Networks. <em>Frontiers in Remote Sensing, 1</em></li> <li>Pahlevan, N., Smith, B., Schalles, J., et al. (2020). Seamless retrievals of chlorophyll-a from Sentinel-2 (MSI) and Sentinel-3 (OLCI) in inland and coastal waters: A machine-learning approach. <em>Remote Sensing of Environment, 240</em>, 111604</li> </ul>
Satellite-derived chlorophyll-a concentrations for Western Water Treatment Plant (Melbourne, Australia) using Mixture Density Networks and Sentinel-2 and Landsat 8 imagery
<p>This dataset contains satellite-derived chlorophyll-a data of the Western Water Treatment Plant (Melbourne, Australia) for the period 21 Mar. 2013 - 01 Feb. 2021. Chlorophyll-a concentrations have been calculated using Mixture Density Networks and Sentinel-2 and Landsat 8 imagery.</p> <p>Mixture Density Networks are a class of neural networks that tackle the inverse problem by modelling the multimodal distribution of target variables using a mixture of Gaussians. For more information, please refer to the following:</p> <ul> <li>Pahlevan, N., Smith, B., Alikas, K., Anstee, J., et al. (2022). Simultaneous retrieval of selected optical water quality indicators from Landsat-8, Sentinel-2, and Sentinel-3. <em>Remote Sensing of Environment, 270</em>, 112860</li> <li>Smith, B., Pahlevan, N., Schalles, J., et al. (2021). A Chlorophyll-a Algorithm for Landsat-8 Based on Mixture Density Networks. <em>Frontiers in Remote Sensing, 1</em></li> <li>Pahlevan, N., Smith, B., Schalles, J., et al. (2020). Seamless retrievals of chlorophyll-a from Sentinel-2 (MSI) and Sentinel-3 (OLCI) in inland and coastal waters: A machine-learning approach. <em>Remote Sensing of Environment, 240</em>, 111604</li> </ul>
Generating Imperviousness Maps from Multispectral Sentinel-2 Satellite Imagery
<p>This dataset contains a list of Sentinel-2 tiles covering Italy for the year 2017. For each tile, a corresponding ground truth GeoTIFF is present which contains a clip of the soil consumption provided by ISPRA (<a href="https://www.isprambiente.gov.it">https://www.isprambiente.gov.it</a>).</p> <p>Dataset can be used to train a Machine Learning model to extract imperviousness maps using Sentinel-2 satellite images.</p> <p>More details can be found reading the paper: </p> <p>Giacco, G., Marrone, S., Langella, G., & Sansone, C. (2022). ReFuse: Generating Imperviousness Maps from Multi-Spectral Sentinel-2 Satellite Imagery. <em>Future Internet</em>, <em>14</em>(10), 278.</p>
KappaSet: Sentinel-2 KappaZeta Cloud and Cloud Shadow Masks
<p><strong>General information</strong></p> <p>The dataset consists of 9251 labelled sub-tiles from 1038 Sentinel-2 (S2) Level-1C (L1C) products distributed over the globe. In terms of seasonal distribution, S2 products can be divided into the following groups:</p> <ul> <li> <p>Winter products: 29 austral and 142 boreal S2 products</p> </li> <li> <p>Sprint products: 45 austral and 257 boreal S2 products</p> </li> <li> <p>Summer products: 30 austral and 293 boreal S2 products</p> </li> <li> <p>Autumn products: 29 austral and 213 boreal S2 products</p> </li> </ul> <p>Each S2 product was oversampled at 10 m resolution for 512 x 512 pixels sub-tiles. From each S2 product, the most challenging ~5 sub-tiles per product were selected for labelling. Each selected L1C S2 product represents different clouds, such as cumulus, stratus, or cirrus, which are spread over various geographical locations around the world. The classification pixel-wise map consists of the following categories:</p> <ul> <li> <p>0 – UNDEFINED: pixels that the labeler is not sure which class they belong to;</p> </li> <li> <p>1 – CLEAR: pixels without clouds or cloud shadows;</p> </li> <li> <p>2 – CLOUD SHADOW: pixels with cloud shadows;</p> </li> <li> <p>3 – SEMI TRANSPARENT CLOUD: pixels with thin clouds through which the land is visible; include cirrus clouds that are on the high cloud level (5-15km).</p> </li> <li> <p>4 – CLOUD: pixels with cloud; include stratus and cumulus clouds that are on the low cloud level (from 0-0.2km to 2km).</p> </li> <li> <p>5 – MISSING: missing or invalid pixels.</p> </li> </ul> <p>The dataset was labelled using Computer Vision Annotation Tool (CVAT) and Segments.ai. With the possibility of integrating an active learning process in Segments.ai, the labelling was performed semi-automatically. The distribution of the dataset is presented in the Figure below. Color represents the season from which the product was chosen.</p> <p>The dataset limitations must be considered: the data mostly covers terrestrial regions (around 91%) and includes some water areas (around 9%); only around 7% of the dataset contains snow. Current sub-tiles do not have georeferencing. </p> <p><strong>Contributions and Acknowledgements</strong></p> <p>The data were annotated by Olga Wold, Mariana Rohtsalu, Nikita Murin, Joosep Truupõld and Fariha Harun. The data verification and Software Development were performed by Indrek Sünter, Heido Trofimov, Anton Kostiukhin, Marharyta Domnich, Mihkel Järveoja, Olga Wold and Tetiana Shtym. The methodology was developed by Kaupo Voormansik, Indrek Sünter, Marharyta Domnich and Tetiana Shtym.</p> <p>The data were collected, processed, and checked as a part of “KappaMask: AI-based Cloudmask Processor for Sentinel-2” project. We thank Segments.ai team for providing a wonderful annotation tool that was actively used to prepare the dataset. In the end, we thank European Space Agency (ESA) for supporting, advising, and funding the project.</p> <p>The project was funded by <em><strong>European Space Agency,</strong></em> Contract No. 4000132124/20/I-DT.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.