Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
86
datasets available to search
ShareScore release 0.7.1
Dataset results
86 results for “RGB Images”
SCLabels: Labelled rectified RGB images from the Spanish CoastSnap network
<h1>Training dataset</h1> <p><span>The SCLabels dataset is intended to be used in the exploring and development of Artificial Intelligence (AI) applications aimed at the automation of the shoreline extraction process from rectified images. SCLabels includes rectified RGB images from the Spanish CoastSnap network and their corresponding masks, together with a metadata file and a README file. RGB images encompass variable geographic locations, fields of view, beach types and degrees of occupation, tidal regimes, meteoceanic and lightning conditions, and a variety of environmental characteristics. Masks account for dense pixel labels including 5 categories: i) No data; ii) Not classified; iii) Landwards; iv) Seawards; and v) Shoreline. In the metadata file, images are linked to their corresponding masks, and information about the geographic location of each image, capture characteristics and image source, shoreline position and other auxiliary data are provided. The README file enhances the explainability and comprehension of the dataset, elaborating on the context and contents, and providing detailed explanations of the metadata, potential limitations, technical aspects of the image processing and annotation stages, usage recommendations, and related works. </span></p> <h1>Technical details</h1> <p>The SCLabels dataset version 1.0.0 is packaged in a compressed file (SCLabels_v1.0.0.zip). A total of 1717 RGB images are shared in JPG format, corresponding masks in PNG format, a metadata file in JSON format, and the README file in PDF format.</p> <h2>Data preprocessing</h2> <p><span>To generate the SCLabels masks, rectified RGB images and their corresponding shorelines were used. RGB images were cropped to the minimum and maximum alongshore pixel coordinates of the shoreline (vertical axis) plus 10 additional pixels above and below to preserve contextual information. A grayscale image was then derived from each cropped RGB image for subsequent pixel labelling. First, a binary mask was derived, marking "NoData'' for black and white padded pixels resulting from the registration and rectification steps. Subsequently, the shoreline was densified, ensuring at least one pixel per row was assigned the "Shoreline" label. Next, "Landwards" and "Seawards" labels were assigned to the right and left of the shoreline. Pixels left unlabelled were categorised as "NotClassified". Finally, masks’ values were reclassified to align with the predefined labels, and the grayscale masks were exported. For additional information, please consult the README file. </span></p> <h2>Data splitting</h2> <p><span>Data splitting requirements may vary depending on the chosen AI approach (e.g., splitting by entire images, image patches, or image rows). Researchers should use a consistent data splitting method and document the approach and splits used in publications. This transparency enables reproducible results and facilitates comparisons between studies.</span></p> <h2>Classes, labels and annotations</h2> <p><span>The SCLabels dataset includes one mask per rectified RGB image, sharing the same width and height. These masks are in greyscale and PNG format, and consist of five different labels:</span></p> <table> <tbody> <tr> <td><strong> Mask value</strong></td> <td><strong> Label</strong></td> <td><strong> Description</strong></td> </tr> <tr> <td>0</td> <td>NoData</td> <td>High probability of being black or white padded pixels, used to pad non-rectangular images within the image registration and rectification processes</td> </tr> <tr> <td>25</td> <td>NotClassified</td> <td>Not labeled pixels</td> </tr> <tr> <td>75</td> <td>Landwards</td> <td>All pixels that are towards the landside with respect to the shoreline (row-wise), excluding “NoData” ones</td> </tr> <tr> <td>150</td> <td>Seawards</td> <td>All pixels that are towards the seaside with respect to the shoreline (row-wise), excluding “NoData” ones</td> </tr> <tr> <td>255</td> <td>Shoreline</td> <td>Pixels intersected by the mapped shoreline densified to cover one pixel per row, at least</td> </tr> </tbody> </table> <h2>Parameters</h2> <p><span>RGB values or any transformation in the colour space can be used as parameters.</span><span> </span></p> <h2>Data sources</h2> <p><span>In the CoastSnap initiative, citizens capture images (oblique smartphone photos) from fixed CoastSnap stations and share them with the scientific managers. Images are subjected to a quality control process, spatially registered to a designated target image, and rectified (georeferencing). The shoreline is subsequently digitised from each rectified image.</span><span> </span></p> <h2>Data quality</h2> <p><span>All images included have been supervised by CSs’ scientific managers. However, citizen scientists take images by smartphones (different camera quality) at irregular intervals across various sites with varying weather and illumination conditions. Users of SCLabels dataset must be aware of this variance. </span></p> <h2>Image resolution</h2> <p><span>The resolution of the images depends on the CoastSnap station and the length of the shoreline, ranging from 241x188 pixels to 801x796 pixels.</span></p> <h2>Spatial coverage</h2> <p><span>The SCLabels dataset version 1.0.0 contains data from five Spanish CoastSnap stations, including sandy beaches in the northwest (</span><span>agrelo</span><span>), the Cíes Islands (</span><span>cies</span><span>), the south (</span><span>cadiz</span><span>), and the Balearic Islands (</span><span>samarador </span><span>and </span><span>arenaldentem</span><span>).</span></p> <table> <tbody> <tr> <td><strong> CoastSnap station</strong></td> <td><strong> Longitude</strong></td> <td><strong> Latitude</strong></td> </tr> <tr> <td><em>agrelo</em></td> <td>-8.772</td> <td>42.331</td> </tr> <tr> <td><em>cies</em></td> <td>-8.900</td> <td>42.226</td> </tr> <tr> <td><em>cadiz</em></td> <td>-6.288</td> <td>36.522</td> </tr> <tr> <td><em>samarador</em></td> <td>3.185</td> <td>39.350</td> </tr> <tr> <td><em>arenaldentem</em></td> <td>2.974</td> <td>39.353</td> </tr> </tbody> </table> <h2>Contact information</h2> <p><span>For further technical inquiries or additional information about the annotated dataset, please contact jsoriano@socib.es.</span></p>
Thermal Bridges on Building Rooftops - Hyperspectral (RGB + Thermal + Height) drone images of Karlsruhe, Germany, with thermal bridge annotations
<p><strong>Overview:</strong></p> <p>The dataset of <strong>Thermal Bridges on Building Rooftops (TBBR dataset)</strong> consists of annotated combined RGB and thermal drone images with a height map. All images were converted to a uniform format of 3000x4000 pixels, aligned, and cropped to <strong>2680x3370</strong> to remove empty borders. See the "Usage" section below for details about the stored formats made available here.</p> <p>The raw images for our dataset were recorded with a normal (RGB) and a FLIR-XT2 (thermal) camera on a DJI M600 drone. They show six large building blocks of around 20 buildings per block recorded in the city centre of the German city Karlsruhe east of the market square. Because of a high overlap rate of the images, the same buildings are on average recorded from different angles in different images about 20 times.</p> <p>All images were recorded during a drone flight on March 19, 2019 from 7 a.m. to 8 a.m. At this time, temperatures were between 3.78 ° C and 4.97 ° C, humidity between 80% and 98%. There was no rain on the day of the flight, but there was 2.3mm/m² 48 hours beforehand. For recording the thermographic images an emissivity of 1.0 was set. The global radiation during this period was between 38.59 W / m² and 120.86 W / m². No direct sunlight can be seen visually on any of the recordings.</p> <p>The dataset contains <strong>926 images</strong> with a total of <strong>6,927 annotations</strong> of thermal bridges on rooftops, split into train and test subsets with 723 (5,614) and 203 (1,313) images (annotations), respectively. The annotations only include thermal bridges that are visually identifiable with the human eye. Because of the aforementioned image overlap, each thermal bridge is annotated multiple times from different angles.</p> <p>For the annotation of the thermal images the image processing program <em>VGG Image Annotator </em>from the Visual Geometry Group, version 2.0.10, was used. The thermal bridge annotations are outlined with polygon shapes. These polygon lines were placed as close as possible but outside the area of significant temperature increase. If a detected thermal bridge was partially covered by another building component located in the foreground, the thermal bridge was also marked across the covering in case of minor coverings. Adjacent thermal bridges, which affect different rooftop components, were annotated separately. For example, a window with poor insulation of the window reveal located in the area of a poorly insulated roof is annotated individually. There is no overlap between annotated areas. While each image contains annotations, they also include thermal bridges present that are not annotated.</p> <p><strong>Usage:</strong></p> <p>Each compressed archive file represents one of the six flight paths. For the related publication the final path (Flug1_105Media) was used as a hold-out test sample. The archives contain Numpy files (one per image) of shape (2680, 3370, 5), where the final dimension is the colour channel in the format [B, G, R, Thermal, Height].</p> <p>Archives were compressed using <a href="https://facebook.github.io/zstd/">ZStandard</a> compression. They can be decompressed in a terminal by running e.g.</p> <pre><code class="language-bash">tar -I zstd -xvf Flug1_105Media.tar.zst</code></pre> <p>these will be decompressed into the file structure:</p> <pre><code>images/ └── Flug1_105Media/ └── DJI_0004_R.npy └── DJI_0006_R.npy └── ...</code></pre> <p>Corresponding annotations are provided in the COCO JSON format. There is one file for training (Flug1_100Media - Flug1_104Media blocks) and one for test (Flug1_105Media block). They contain a single class (thermal bridge) and expect the folder structure shown below.</p> <p>Note: The annotation files contain <em>relative</em> paths to numpy files, in case of problems please convert to <em>absolute</em> paths (i.e. insert the containing directory before each file path in the JSON annotation files).</p> <p>We provide the <a href="https://github.com/Helmholtz-AI-Energy/TBBRDet"><strong>TBBRDet software</strong></a> which includes a dataloader and dataset inspection tools which make use of the <a href="https://github.com/facebookresearch/detectron2">Detectron2</a> and <a href="https://github.com/open-mmlab/mmdetection">MMDetection</a> libraries.</p> <p>We recommend the following folder structure for use:</p> <pre><code>├── train/ │ ├── Flug1_100-104Media_coco.json │ └── images/ │ ├── Flug1_100Media/ │ │ ├── DJI_XXXX_R.npy │ │ └── ... │ ├── ... │ └── Flug1_104Media/ │ ├── DJI_XXXX_R.npy │ └── ... └── test/ ├── Flug1_105Media_coco.json └── images/ └── Flug1_105Media/ ├── DJI_XXXX_R.npy └── ...</code></pre> <p><strong>Metadata:</strong></p> <p>The experimental metadata was structured with the <strong>Spatio Temporal Asset Catalog (STAC)</strong> specification family. This specification provides a standardized way for describing geospatial assets. It defines related JSON object types of Item, Catalog, and Catalog, extending on Collection as the basis.</p> <p>One STAC Collection JSON object provides information about the recorded images and the environmental conditions during recordings. It also contains information about the overall bounding box of the entire area in which images were recorded.</p> <p>This object links to related STAC Item JSON objects containing information about the recorded city blocks and the cameras. The objects for the city blocks contain the GeoJSON geometry of the respective block and the<br> corresponding bounding box. The objects containing the camera information are based on an existing STAC extension for camera related metadata.</p> <p>Metadata of the archived NumPy files for each image was structured using the <strong>Data Package</strong> schema from the <strong>Frictionless Standards</strong>. This standard describes a collection of data files. Therefore, metadata about all containerized NumPy files of the six flight paths (Flug1_100Media - Flug1_104Media blocks and Flug1_105Media block) is provided within a JSON-based file.</p> <p>Note that <strong>camera1</strong> corresponds to the <strong>RGB camera</strong> and <strong>camera2</strong> the <strong>thermal</strong>.</p> <p><strong>FAIR Digital Objects:</strong></p> <p>All files are represented in a standardized way as <strong>FAIR Digital Objects<br> (FAIR DOs)</strong> to enable machine actionable decisions on the data in spirit of<br> the FAIR principles.</p> <p><strong>Persistent Identifier (PID):</strong></p> <p>Persistent Identifiers (PIDs) are resolvable with the <a href="https://hdl.handle.net/">Handle.Net Registry (HNR)</a>.</p> <table> <thead> <tr> <th scope="col">File</th> <th scope="col">Persistent Identifier (PID)</th> </tr> </thead> <tbody> <tr> <td>Flug1_100-104Media_coco.json</td> <td>21.11152/6ea60288-d895-414e-80c0-26c9fdd662b2</td> </tr> <tr> <td>Flug1_105Media_coco.json</td> <td>21.11152/58d43ddc-5e29-4980-8675-ae579b50a1e2</td> </tr> <tr> <td>Flug1_100.tar.zst</td> <td>21.11152/6858a0b5-cc60-40e9-afef-8c2dd8b35e8e</td> </tr> <tr> <td>Flug1_101.tar.zst</td> <td>21.11152/e670f510-7e00-4d3a-9b90-3bac7a7c069e</td> </tr> <tr> <td>Flug1_102.tar.zst</td> <td>21.11152/3ab9f444-05f6-445e-a691-62fae4021bea</td> </tr> <tr> <td>Flug1_103.tar.zst</td> <td>21.11152/365fd8cf-8e86-41b8-9d0e-b816fdd01d29</td> </tr> <tr> <td>Flug1_104.tar.zst</td> <td>21.11152/041a6111-644a-4617-afb3-3c421a88e8e3</td> </tr> <tr> <td>Flug1_105.tar.zst</td> <td>21.11152/f48bf4e7-3879-4216-8f64-45a060b8f658</td> </tr> <tr> <td>Flug1_100-105_frictionless_standards.json</td> <td>21.11152/7b58b3b5-75eb-4417-ac4d-abe025e159f6</td> </tr> <tr> <td>Flug1_collection_stac_spec.json</td> <td>21.11152/ba370aa3-6422-428c-9ff7-c2ef429df603</td> </tr> <tr> <td>Flug1_100_stac_spec.json</td> <td>21.11152/09cb76fc-b8cb-4116-a22a-68c5bdfa77b0</td> </tr> <tr> <td>Flug1_101_stac_spec.json</td> <td>21.11152/24a55398-b96b-43dd-b0fb-cd8ce302c7ce</td> </tr> <tr> <td>Flug1_102_stac_spec.json</td> <td>21.11152/721234ac-4b5a-4d02-9944-82a08ef2db35</td> </tr> <tr> <td>Flug1_103_stac_spec.json</td> <td>21.11152/ebaeb5bc-0514-47c9-bcd2-98f0253843d8</td> </tr> <tr> <td>Flug1_104_stac_spec.json</td> <td>21.11152/9854677c-77c5-4a0b-916b-57dd9ec20198</td> </tr> <tr> <td>Flug1_105_stac_spec.json</td> <td>21.11152/cfd0fc0e-f5ea-464e-a57f-28e882924860</td> </tr> <tr> <td>Flug1_camera1_stac-spec.json</td> <td>21.11152/976fcf28-f924-4a21-b53d-5d054ad8198d</td> </tr> <tr> <td>Flug1_camera2_stac-spec.json</td> <td>21.11152/37833c54-1d36-42e4-858d-831447122863</td> </tr> </tbody> </table>
Hyperspectral (RGB + Thermal) drone images of Karlsruhe, Germany - Raw images for the Thermal Bridges on Building Rooftops (TBBR) dataset
<p><strong>Overview:</strong></p> <p>This repository contains the <strong>raw images</strong> for the dataset of <a href="https://doi.org/10.5281/zenodo.4767771"><strong>Thermal Bridges on Building Rooftops (TBBR) dataset</strong></a>.</p> <p>This dataset contains <strong>5696 drone images</strong> (2848 RGB and 2848 thermal) of building rooftops, recorded with a normal (RGB) and a FLIR-XT2 (thermal) camera on a DJI M600 drone. They show six large building blocks of around 20 buildings per block recorded in the city centre of the German city Karlsruhe east of the market square. Because of a high overlap rate of the images, the same buildings are on average recorded from different angles in different images about 20 times.</p> <p>All images were recorded during a drone flight on March 19, 2019 from 7 a.m. to 8 a.m. At this time, temperatures were between 3.78 ° C and 4.97 ° C, humidity between 80% and 98%. There was no rain on the day of the flight, but there was 2.3mm/m² 48 hours beforehand. For recording the thermographic images an emissivity of 1.0 was set. The global radiation during this period was between 38.59 W / m² and 120.86 W / m². No direct sunlight can be seen visually on any of the recordings.</p> <p><strong>Usage:</strong></p> <p>Each zip archive file represents one of the six drone flight paths. The archives contain JPG files of size 4000x3000 pixels (RGB) and 640x512 (Thermal), separated into individual directories for RGB and Thermal:</p> <pre><code>├── Flug_100/ │ ├── RGB/ │ │ ├── DJI_0004.jpg │ │ ├── DJI_0006.jpg │ │ └── ... │ └── Thermal/ │ ├── DJI_0003_R.JPG │ ├── DJI_0005_R.JPG │ └── ... ├── Flug_101/ │ ├── RGB/ │ │ ├── DJI_0001.jpg │ │ ├── DJI_0003.jpg │ │ └── ... │ └── Thermal/ │ ├── DJI_0000_R.JPG │ ├── DJI_0002_R.JPG │ └── ... └── ...</code></pre> <p><strong>File Numbering/Naming Scheme:</strong></p> <p>The pairs of RGB + Thermal images follow the simple numbering scheme of: <strong>RGB = Thermal + 1</strong>.<br> For example, DJI_0003_R.jpg and DJI_0004.JPG are the matching Thermal and RGB images, respectively, that can be merged to form a single hyperspectral drone image.</p> <p>To perform the merging, we recommend using the <strong>merge_image_layers.py</strong> script provided by the associated <strong><a href="https://github.com/Helmholtz-AI-Energy/TBBRDet">TBBRDet software</a></strong> (see the scripts/alignment/ directory).</p> <p>For convenience, we have provided a CSV listing all annotated images in the <a href="https://doi.org/10.5281/zenodo.4767771"><strong>Thermal Bridges on Building Rooftops (TBBR) dataset</strong></a>. The CSV format is as follows:</p> <pre><code>Flight,RGB,Thermal Flug_100,DJI_0048.jpg,DJI_0047_R.JPG Flug_100,DJI_0050.jpg,DJI_0049_R.JPG ...</code></pre> <p> </p>
Doodleverse/Segmentation Zoo Res-UNet model for NOAA ERI/4-class segmentation of RGB 512x512 images
<p>This Residual-UNet model is trained on 1,179 pairs of human-generated segmentation labels and images from Emergency Response Imagery (ERI) collected by US National Oceanic and Atmospheric Administration (NOAA) after Hurricane Barry, Delta, Dorian, Florence, Ida, Laura, Michael, Sally, Zeta, and Tropical Storm Gordon. The dataset is available here: https://doi.org/10.5281/zenodo.7268082</p> <p>Models have been created using Segmentation Gym:</p> <p>Code - https://github.com/Doodleverse/segmentation_gym</p> <p>Paper - https://doi.org/10.1029/2022EA002332</p> <p> </p> <p>The model takes input images that are 512 x 512 x 3 pixels, and the output is 512 x 512 x 4, corresponding to 4 classes:</p> <ol> <li>water</li> <li>bare sediment</li> <li>vegetation</li> <li>development (roads, buildings, power lines, parking lots, etc.)</li> </ol> <p> </p> <p>Included here are 6 files with the same root name:</p> <ol> <li> '.json' config file: this is the file that was used by Segmentation Gym to create the weights file. It contains instructions for how to make the model and the data it used, as well as instructions for how to use the model for prediction.</li> <li>'.h5' weights file: this is the file that was created by the Segmentation Gym function `train_model.py`. It contains the trained model's parameter weights. It can called by the Segmentation Gym function `seg_images_in_folder.py`.</li> <li> '_model_history.npz' model training history file: this numpy archive file contains numpy arrays describing the training and validation losses and metrics. It is created by the Segmentation Gym function `train_model.py`</li> <li> '.png' model training loss and mean IoU plot: this png file contains plots of training and validation losses and mean IoU scores during model training. A subset of data inside the .npz file. It is created by the Segmentation Gym function `train_model.py`</li> <li>'.zip' of the model in the Tensorflow ‘saved model’ format. It is created by the Segmentation Gym function `utils/gen_saved_model.py`</li> <li>'_modelcard.json' model card file: this is a json file containing fields that collectively describe the model origins, training choices, and dataset that the model is based upon. There is some redundancy between this file and the `config` file (described above) that contains the instructions for the model training and implementation. The model card file is not used by the program but is important metadata so it is important to keep with the other files that collectively make the model and is such is considered part of the model</li> </ol> <p>Additionally, BEST_MODEL.txt contains the name of the model with the best validation loss and mean IoU</p>
Sonar-to-RGB Image Translation for Diver Monitoring in Poor Visibility Environments
<p><strong>Context</strong></p><p>This dataset is part of the paper "Sonar-to-RGB Image Translation for Diver Monitoring in Poor Visibility Environments" presented at Oceans 2022, Hampton Roads,<strong> </strong>DOI: <a href="https://doi.org/10.1109/OCEANS47191.2022.9977024">10.1109/OCEANS47191.2022.9977024</a></p><p>This dataset consists of paired camera and multi-beam sonar images of technical divers performing different underwater tasks in two locations: an indoor test basin and a lake. The general goal is to assist emergency operators that monitor the safety of divers operating in bad visibility conditions.</p><p>This data was used to train image-to-image translation models in order to generate realistic optical-like images given only sonar images as input or a combination of a sonar image and a dark or turbid optical image.</p><p> </p><p><strong>Content</strong></p><p>This repository contains three .zip folders each containing data collected in a different lab or field trial.</p><ul><li>'basin-dataset-1.zip' and 'basin-dataset-2.zip' contain data that were collected in an indoor testing facility at DFKI - Robotics Innovation Center, Bremen, Germany.</li><li>'lake-dataset-1.zip' and 'lake-dataset-2.zip' contains data collected at lake Kreidesee, Hemmoor, Germany.</li></ul><p>Each .zip file contains two subfolders labelled as 'camera' and 'sonar', each containing the images in png format. Data files under these subfolders with matching names composes a pair of time-synchronized images. For example, 'camera/0001.png' corresponds to 'sonar/0001.png'. The acquisition timestamp represented in seconds since epoch for every data file is recorded in 'sample.csv' include in each .zip file.</p><p>For more details and meta-information on the collected data please refer to "data_description.json" included in this repository.</p><p>Additional tools for handling and preparing the data can be found under <a href="https://github.com/DeeperSense/oceans_2022">https://github.com/DeeperSense/oceans_2022</a></p><p> </p><p><strong>Acknowledgements</strong></p><p>The data in this repository were collected as a joint effort between the German Center for Artificial Intelligence (DFKI), the German Federal Agency for technical Relief (THW), and Kraken Robotics GmbH. This work is part of the project DeeperSense that received funding from the European Commission. Program H2020-ICT-2020-2 ICT-47-2020 Project Number: 101016958.</p><p>The authors would like to thank the Federal Government and the Heads of Government of the Länder, as well as the Joint Science Conference (GWK), for their initiative within the framework of the NFDI4Ing consortium (German Research Foundation (DFG) - project number 442146713).</p>
Metadata of a Large Sonar and Stereo Camera Dataset Suitable for Sonar-to-RGB Image Translation
<h1>Metadata of a Large Sonar and Stereo Camera Dataset Suitable for Sonar-to-RGB Image Translation</h1> <h2>Introduction</h2> <p>This is a set of metadata describing a large dataset of synchronized sonar and stereo camera recordings, that were captured between August 2021 and September 2023 during the project <a href="https://robotik.dfki-bremen.de/en/research/projects/deepersense/">DeeperSense</a> (https://robotik.dfki-bremen.de/en/research/projects/deepersense/), as training data for Sonar-to-RGB image translation. <a href="../records/7728089">Parts</a> <a href="../records/10220989">of</a> the sensor data have been published (https://zenodo.org/records/7728089, https://zenodo.org/records/10220989). Due to the size of the sensor data corpus, it is currently impractical to make the entire corpus accessible online. Instead, this metadatabase serves as a relatively compact representation, allowing interested researchers to inspect the data, and select relevant portions for their particular use case, which will be made available on demand. This is an effort to comply with the <a href="https://www.go-fair.org/fair-principles/">FAIR</a> principle A2 (https://www.go-fair.org/fair-principles/) that metadata shall be accessible, even when the base data is not immediately.</p> <h3>Locations and sensors</h3> <p>The sensor data was captured at four different locations, including one laboratory (Maritime Exploration Hall at DFKI RIC Bremen) and three field locations (Chalk Lake Hemmoor, Tank Wash Basin Neu-Ulm, Lake Starnberg). At all locations, a ZED camera and a Blueprint Oculus M1200d sonar were used. Additionally, a SeaVision camera was used at the Maritime Exploration Hall at DFKI RIC Bremen and at the Chalk Lake Hemmoor. The <code>examples/</code> directory holds a typical output image for each sensor at each available location.</p> <h3>Data volume per session</h3> <p>Six data collection sessions were conducted. The table below presents an overview of the amount of data captured in each session:</p> <table> <tbody> <tr> <th>Session dates</th> <th>Location</th> <th>Number of datasets</th> <th>Total duration of datasets [h]</th> <th>Total logfile size [GB]</th> <th>Number of images</th> <th>Total image size [GB]</th> </tr> <tr> <td>2021-08-09 - 2021-08-12</td> <td>Maritime Exploration Hall at DFKI RIC Bremen</td> <td>52</td> <td>10.8</td> <td>28.8</td> <td>389’047</td> <td>88.1</td> </tr> <tr> <td>2022-02-07 - 2022-02-08</td> <td>Maritime Exploration Hall at DFKI RIC Bremen</td> <td>35</td> <td>4.4</td> <td>54.1</td> <td>629’626</td> <td>62.3</td> </tr> <tr> <td>2022-04-26 - 2022-04-28</td> <td>Chalk Lake Hemmoor</td> <td>52</td> <td>8.1</td> <td>133.6</td> <td>1’114’281</td> <td>97.8</td> </tr> <tr> <td>2022-06-28 - 2022-06-29</td> <td>Tank Wash Basin Neu-Ulm</td> <td>42</td> <td>6.7</td> <td>144.2</td> <td>824’969</td> <td>26.9</td> </tr> <tr> <td>2023-04-26 - 2023-04-27</td> <td>Maritime Exploration Hall at DFKI RIC Bremen</td> <td>55</td> <td>7.4</td> <td>141.9</td> <td>739’613</td> <td>9.6</td> </tr> <tr> <td>2023-09-01 - 2023-09-02</td> <td>Lake Starnberg</td> <td>19</td> <td>2.9</td> <td>40.1</td> <td>217’385</td> <td>2.3</td> </tr> <tr> <th> </th> <th> </th> <th>255</th> <th>40.3</th> <th>542.7</th> <th>3’914’921</th> <th>287.0</th> </tr> </tbody> </table> <h2>Data and metadata structure</h2> <h3>Sensor data corpus</h3> <p>The sensor data corpus comprises two processing stages:</p> <ul> <li>raw data streams stored in ROS bagfiles (aka <strong>logfiles</strong>),</li> <li>camera and sonar images (aka <strong>datafiles</strong>) extracted from the logfiles.</li> </ul> <p>The files are stored in a file tree hierarchy which groups them by session, dataset, and modality:</p> <pre><code>${session_key}/ ${dataset_key}/ ${logfile_name} ${modality_key}/ ${datafile_name}</code></pre> <p>A typical logfile path has this form:</p> <pre><code>2023-09_starnberg_lake/ 2023-09-02-15-06_hydraulic_drill/ stereo_camera-zed-2023-09-02-15-06-07.bag</code></pre> <p>A typical datafile path has this form:</p> <pre><code>2023-09_starnberg_lake/ 2023-09-02-15-06_hydraulic_drill/ zed_right/ 1693660038_368077993.jpg</code></pre> <p>All directory and file names, and their particles, are designed to serve as identifiers in the metadatabase. Their formatting, as well as the definitions of all terms, are documented in the file <code>entities.json</code>.</p> <h3>Metadatabase</h3> <p>The metadatabase is provided in two equivalent forms:</p> <ul> <li>as a standalone <a href="https://www.sqlite.org/index.html">SQLite</a> (https://www.sqlite.org/index.html) database file <code>metadata.sqlite</code> for users familiar with SQLite,</li> <li>as a collection of CSV files in the <code>csv/</code> directory for users who prefer other tools.</li> </ul> <p>The database file has been generated from the CSV files, so each database table holds the same information as the corresponding CSV file. In addition, the metadatabase contains a series of convenience views that facilitate access to certain aggregate information.</p> <p>An entity relationship diagram of the metadatabase tables is stored in the file <code>entity_relationship_diagram.png</code>. Each entity, its attributes, and relations are documented in detail in the file <code>entities.json</code></p> <p>Some general design remarks:</p> <ul> <li>For convenience, timestamps are always given in both a human-readable form (ISO 8601 formatted datetime strings with explicit local time zone), and as seconds since the UNIX epoch.</li> <li>In practice, each logfile always contains a single stream, and each stream is stored always in a single logfile. Per database schema however, the entities <code>stream</code> and <code>logfile</code> are modeled separately, with a “many-streams-to-one-logfile” relationship. This design was chosen to be compatible with, and open for, data collections where a single logfile contains multiple streams.</li> <li>A <code>modality</code> is not an attribute of a <code>sensor</code> alone, but of a <code>datafile</code>: Because a <code>sensor</code> is an attribute of a <code>stream</code>, and a single stream may be the source of multiple modalities (e.g. RGB vs. grayscale images from the same camera, or cartesian vs. polar projection of the same sonar output). Conversely, the same modality may originate from different sensors.</li> </ul> <p>As a usage example, the data volume per session which is tabulated at the top of this document, can be extracted from the metadatabase with the following SQL query:</p> <div> <pre><code><span><span>SELECT</span></span> <span> PRINTF(</span> <span> <span>'%s - %s'</span>,</span> <span> <span>SUBSTR</span>(session_start, <span>1</span>, <span>10</span>),</span> <span> <span>SUBSTR</span>(session_end, <span>1</span>, <span>10</span>)) <span>AS</span> <span>'Session dates'</span>,</span> <span> location_name_english <span>AS</span> Location,</span> <span> number_of_datasets <span>AS</span> <span>'Number of datasets'</span>,</span> <span> total_duration_of_datasets_h <span>AS</span> <span>'Total duration of datasets [h]'</span>,</span> <span> total_logfile_size_gb <span>AS</span> <span>'Total logfile size [GB]'</span>,</span> <span> number_of_images <span>AS</span> <span>'Number of images'</span>,</span> <span> total_image_size_gb <span>AS</span> <span>'Total image size [GB]'</span></span> <span><span>FROM</span></span> <span> location</span> <span> <span>JOIN</span> <span>session</span> <span>USING</span> (location_id)</span> <span> <span>JOIN</span> (</span> <span> <span>SELECT</span></span> <span> session_id,</span> <span> <span>COUNT</span>(dataset_id) <span>AS</span> number_of_datasets,</span> <span> <span>ROUND</span>(</span> <span> <span>SUM</span>(dataset_duration) <span>/</span> <span>3600</span>,</span> <span> <span>1</span>) <span>AS</span> total_duration_of_datasets_h,</span> <span> <span>ROUND</span>(</span> <span> <span>SUM</span>(total_logfile_size) <span>/</span> <span>10e9</span>,</span> <span> <span>1</span>) <span>AS</span> total_logfile_size_gb</span> <span> <span>FROM</span></span> <span> location</span> <span> <span>JOIN</span> <span>session</span> <span>USING</span> (location_id)</span> <span> <span>JOIN</span> dataset <span>USING</span> (session_id)</span> <span> <span>JOIN</span> view__dataset_total_logfile_size <span>USING</span> (dataset_id)</span> <span> <span>GROUP</span> <span>BY</span></span> <span> session_id</span> <span> ) <span>USING</span> (session_id)</span> <span> <span>JOIN</span> (</span> <span> <span>SELECT</span></span> <span> session_id,</span> <span> <span>COUNT</span>(datafile_id) <span>AS</span> number_of_images,</span> <span> <span>ROUND</span>(<span>SUM</span>(datafile_size) <span>/</span> <span>10e9</span>, <span>1</span>) <span>AS</span> total_image_size_gb</span> <span> <span>FROM</span></span> <span> <span>session</span></span> <span> <span>JOIN</span> dataset <span>USING</span> (session_id)</span> <span> <span>JOIN</span> stream <span>USING</span> (dataset_id)</span> <span> <span>JOIN</span> <span>datafile</span> <span>USING</span> (stream_id)</span> <span> <span>GROUP</span> <span>BY</span></span> <span> session_id</span> <span> ) <span>USING</span> (session_id)</span> <span><span>ORDER</span> <span>BY</span> session_id;</span></code></pre> </div>
Ultra-high-resolution modified RGB UAV-imaging of Alternaria solani
<p>This dataset is collected from both symptomatic and non-symptomatic plants during the growing seasons of 2019 and 2022, on 40x20 m experimental fields in Lemberge (Merelbeke), Belgium (50.986544°N, 3.774066°E) using a DJI M600 PRO unmanned aerial vehicle equiped with a modified Sony Alpha 7III camera with 135 mm lens. The field trial is conducted in analogy to the method described by Van De Vijver et al. (2020, 2022), using two different cultivers, Spunta (2019) and Fontane (2022) respectively. The dataset of 2019 comprises data from three different flights (3, 6 and 9 days after inoculation) and the dataset of 2022 from four different flights (5,7, 9 and 13 days after inoculation). </p> <p>This dataset consists out of 7660 patches of 256x256 pixels, cropped out of the original images, labeled and sorted in two categories (1: Alternaria, 0: no Alternaria), accompagned by a csv file containing the following information:</p> <ul> <li>Original patch name</li> <li>Random patch name (used during the labeling process)</li> <li>Row patch number</li> <li>Column patch number</li> <li>Block number, column block number and row block number</li> <li>Original mage name</li> <li>Coordinates of original image: latitude, longitude, altitude</li> <li>Date of flight</li> <li>Label (0: no Alternaria, 1: Alternaria)</li> </ul> <p>More detailed information about this dataset (both the collection and the preprocessing) can be found in the corresponding article 'Ultra-high-resolution UAV-Imaging and Supervised Deep Learning for Accurate Detection of Alternaria Solani in Potato Fields.' </p> <p> </p> <p>If you use this dataset, please refer to the related journal paper as follows: "Wieme J, Leroux S, Cool SR, Van Beek J, Pieters JG and Maes WH (2024) Ultra-highresolution UAV-imaging and supervised deep learning for accurate detection of Alternaria solani in potato fields. Front. Plant Sci. 15:1206998. doi: 10.3389/fpls.2024.1206998"</p> <p> </p> <p>This dataset was gathered within the Proeftuin Smart Farming 4.0 project (180503) within the Industry 4.0 Living Labs with funding from Flanders innovation & entrepreneurship (VLAIO, Belgium) and in the Horizon 2020 project SmartAgriHubs - Connecting the dots to unleash the innovation potential for digital transformation of the European agrifood sector with funding from the European Union under grant agreement No. 818182. Jana Wieme is funded by grant 1SE3921N of Research Foundation Flanders (FWO).</p> <p> </p>
Dataset of polygons with the contour of 900 juniper shrubs used to track shrub growth from 1977 to 2020 in Sierra Nevada (Spain) using very high resolution aerial and satellite RGB images.
<p><strong>This database provides as polygons the contours of 900 juniper shrubs (<em>Juniperus communis L.</em> and <em>Juniperus sabina L.</em>) along 5 decades (years 1977, 1984, 2001, 2010 and 2020). The contour of each of 900 shrubs manually mapped using the Google Satellite composite for the year 2020) was tracked back in time using orthophotos provided by REDIAM. Contours were obtained by manual annotation as polygon shapefiles in QGIS 3.10.3. Additionally, for the year 2020, the polygons were characterized with five attributes that gather ecological information: Morphotype (Hemispherical, Striped, Senescent, With rock), Presence of surrounding vegetation (Bare Soil, Surrounding Vegetation), Presence of nearby human land-uses (Surrounded by human facilities within 250 meters, Non-anthropized environment) Health status (as percentage of canopy cover with brown foliage: values between 0-5, where 0 corresponds to 100% photosynthetically active cover, decreasing the photosynthetically active cover until category 5 which corresponds to 100% damaged cover), and the subjective annotation certainty of the GIS technician (values between 0-5, where the value 0 corresponds to a very uncertain annotation up to the value 5 which corresponds to a fairly certain annotation). </strong></p>
RT-Trees: Evaluation and RGB training images with masks
<p>This is the RT-Trees dataset proposed and used in the paper titled, "Shadowsense: Unsupervised Domain Adaptation and Feature Fusion for Shadow-Agnostic Tree Crown Detection From RGB-Thermal Drone Imagery", published at the <a href="https://openaccess.thecvf.com/content/WACV2024/html/Kapil_ShadowSense_Unsupervised_Domain_Adaptation_and_Feature_Fusion_for_Shadow-Agnostic_Tree_WACV_2024_paper.html">IEEE/CVF WACV 2024</a> conference. Due to the size of the dataset and Zenodo's 50GB limit, the dataset is partitioned into two separate uploads. This upload contains the evaluation splits (test & val), along with the labelled subset of RGB training images used for a supervised training experiment, and the much larger set of unlabelled RGB images used for fully-unsupervised training. </p> <p>The second upload includes the corresponding unlabelled thermal images used for unsupervised training. </p>
Doodleverse/Segmentation Zoo Res-UNet models for 4-class (water, whitewater, sediment and other) segmentation of Sentinel-2 and Landsat-7/8 5-band (RGB+NIR+SWIR) images of coasts.
<p>These Residual-UNet model data are based on 5-band RGB+NIR+SWIR (red, green, blue, near-infrared, and short-wave infrared) images of coasts and associated labels.</p> <p> </p> <p>Models have been created using Segmentation Gym* using the following dataset**: <a href="https://doi.org/10.5281/zenodo.7344571">https://doi.org/10.5281/zenodo.7344571 </a></p> <p>Classes: {0=water, 1=whitewater, 2=sediment, 3=other}</p> <p><strong>File descriptions</strong></p> <p>For each model, there are 5 files with the same root name:</p> <p>1. <strong>'.json' </strong>config file: this is the file that was used by Segmentation Gym* to create the weights file. It contains instructions for how to make the model and the data it used, as well as instructions for how to use the model for prediction. It is a handy wee thing and mastering it means mastering the entire Doodleverse.</p> <p>2.<strong> '.h5'</strong> weights file: this is the file that was created by the Segmentation Gym* function `train_model.py`. It contains the trained model's parameter weights. It can called by the Segmentation Gym* function `seg_images_in_folder.py`. Models may be ensembled.</p> <p>3.<strong> '_modelcard.json'</strong> model card file: this is a json file containing fields that collectively describe the model origins, training choices, and dataset that the model is based upon. There is some redundancy between this file and the `config` file (described above) that contains the instructions for the model training and implementation. The model card file is not used by the program but is important metadata so it is important to keep with the other files that collectively make the model and is such is considered part of the model</p> <p>4. <strong> '_model_history.npz'</strong> model training history file: this numpy archive file contains numpy arrays describing the training and validation losses and metrics. It is created by the Segmentation Gym function `train_model.py`</p> <p>5. <strong> '.png'</strong> model training loss and mean IoU plot: this png file contains plots of training and validation losses and mean IoU scores during model training. A subset of data inside the .npz file. It is created by the Segmentation Gym function `train_model.py`</p> <p>Additionally, BEST_MODEL.txt contains the name of the model with the best validation loss and mean IoU</p> <p> </p> <p><strong>References</strong></p> <p>*Segmentation Gym: Buscombe, D., & Goldstein, E. B. (2022). A reproducible and reusable pipeline for segmentation of geoscientific imagery. Earth and Space Science, 9, e2022EA002332. <a href="https://doi.org/10.1029/2022EA002332">https://doi.org/10.1029/2022EA002332</a> See: <a href="https://github.com/Doodleverse/segmentation_gym">https://github.com/Doodleverse/segmentation_gym</a></p> <p>** Buscombe, Daniel. (2022). Images and 4-class labels for semantic segmentation of Sentinel-2 and Landsat RGB, NIR, and SWIR satellite images of coasts (water, whitewater, sediment, other) (v1.0) [Data set]. Zenodo. <a href="https://doi.org/10.5281/zenodo.7344571">https://doi.org/10.5281/zenodo.7344571</a></p>
Images and 4-class labels for semantic segmentation of Sentinel-2 and Landsat RGB, NIR, and SWIR satellite images of coasts (water, whitewater, sediment, other)
<p><em><strong>Images and 4-class labels for semantic segmentation of Sentinel-2 and Landsat RGB, NIR, and SWIR satellite images of coasts (water, whitewater, sediment, other)</strong></em></p> <p><strong>Description</strong></p> <p>579 images and 579 associated labels for semantic segmentation of Sentinel-2 and Landsat RGB satellite images of coasts. The 4 classes are 0=water, 1=whitewater, 2=sediment, 3=other</p> <p>These images and labels have been made using the Doodleverse software package, Doodler*. These images and labels could be used within numerous Machine Learning frameworks for image segmentation, but have specifically been made for use with the Doodleverse software package, Segmentation Gym**.</p> <p>Some (422) of these images and labels were originally included in the Coast Train*** data release, and have been modified from their original by reclassifying from the original classes to the present 4 classes.</p> <p>The label images are a subset of the following data release**** <a href="https://doi.org/10.5281/zenodo.7335647">https://doi.org/10.5281/zenodo.7335647</a></p> <p>Imagery comes from the following 10 sand beach sites:</p> <ol> <li>Duck, NC, Hatteras NC, USA</li> <li>Santa Cruz CA, USA</li> <li>Galveston TX, USA</li> <li>Truc Vert,France</li> <li>Sunset State Beach CA, USA</li> <li>Torrey Pines CA, USA</li> <li>Narrabeen, NSW, Australia</li> <li>Elwha WA, USA</li> <li>Ventura region, CA, USA</li> <li>Klamath region, CA USA</li> </ol> <p>Imagery are a mixture of 10-m Sentinel-2 and 15-m pansharpened Landsat 7, 8, and 9 visible-band imagery of various sizes. Red, Green, Blue, NIR, and SWIR bands only</p> <p><strong>File descriptions</strong></p> <ol> <li>classes.txt, a file containing the class names</li> <li>images.zip, a zipped folder containing the 3-band RGB images of varying sizes and extents</li> <li>nir.zip, a zipped folder containing the corresponding near-infrared (NIR) imagery</li> <li>swir.zip, a zipped folder containing the corresponding shortwave-infrared (SWIR) imagery</li> <li>labels.zip, a zipped folder containing the 1-band label images</li> <li>overlays.zip, a zipped folder containing a semi-transparent overlay of the color-coded label on the image (blue=0=water, red=1=whitewater, yellow=2=sediment, green=3=other)</li> <li>resized_images.zip, RGB images resized to 512x512x3 pixels</li> <li>resized_nir.zip, NIR images resized to 512x512x3 pixels</li> <li>resized_swir.zip, SWIR images resized to 512x512x3 pixels</li> <li>resized_labels.zip, label images resized to 512x512 pixels</li> </ol> <p><strong>References</strong></p> <p>*Doodler: Buscombe, D., Goldstein, E.B., Sherwood, C.R., Bodine, C., Brown, J.A., Favela, J., Fitzpatrick, S., Kranenburg, C.J., Over, J.R., Ritchie, A.C. and Warrick, J.A., 2021. Human‐in‐the‐Loop Segmentation of Earth Surface Imagery. Earth and Space Science, p.e2021EA002085<a href="https://doi.org/10.1029/2021EA002085">https://doi.org/10.1029/2021EA002085</a>. See <a href="https://github.com/Doodleverse/dash_doodler">https://github.com/Doodleverse/dash_doodler.</a></p> <p>**Segmentation Gym: Buscombe, D., & Goldstein, E. B. (2022). A reproducible and reusable pipeline for segmentation of geoscientific imagery. Earth and Space Science, 9, e2022EA002332. <a href="https://doi.org/10.1029/2022EA002332">https://doi.org/10.1029/2022EA002332</a> See: <a href="https://github.com/Doodleverse/segmentation_gym">https://github.com/Doodleverse/segmentation_gym</a></p> <p>***Coast Train data release: Wernette, P.A., Buscombe, D.D., Favela, J., Fitzpatrick, S., and Goldstein E., 2022, Coast Train--Labeled imagery for training and evaluation of data-driven models for image segmentation: U.S. Geological Survey data release, <a href="https://doi.org/10.5066/P91NP87I">https://doi.org/10.5066/P91NP87I</a>. See <a href="https://coasttrain.github.io/CoastTrain/">https://coasttrain.github.io/CoastTrain/ </a>for more information</p> <p>**** Buscombe, Daniel, Goldstein, Evan, Bernier, Julie, Bosse, Stephen, Colacicco, Rosa, Corak, Nick, Fitzpatrick, Sharon, del Jesús González Guillén, Anais, Ku, Venus, Paprocki, Julie, Platt, Lindsay, Steele, Bethel, Wright, Kyle, & Yasin, Brandon. (2022). Images and 4-class labels for semantic segmentation of Sentinel-2 and Landsat RGB satellite images of coasts (water, whitewater, sediment, other) (v1.0) [Data set]. Zenodo. <a href="https://doi.org/10.5281/zenodo.7335647">https://doi.org/10.5281/zenodo.7335647</a></p>
Images and 4-class labels for semantic segmentation of Sentinel-2 and Landsat RGB satellite images of coasts (water, whitewater, sediment, other)
<p><strong>Description</strong></p> <p>1018 images and 1018 associated labels for semantic segmentation of Sentinel-2 and Landsat RGB satellite images of coasts. The 4 classes are 0=water, 1=whitewater, 2=sediment, 3=other</p> <p>These images and labels have been made using the Doodleverse software package, Doodler*. These images and labels could be used within numerous Machine Learning frameworks for image segmentation, but have specifically been made for use with the Doodleverse software package, Segmentation Gym**.</p> <p>Some (473) of these images and labels were originally included in the Coast Train*** data release, and have been modified from their original by reclassifying from the original classes to the present 4 classes.</p> <p>Imagery comes from the following 10 sand beach sites:</p> <ol> <li>Duck, NC, Hatteras NC, USA</li> <li>Santa Cruz CA, USA</li> <li>Galveston TX, USA</li> <li>Truc Vert,France</li> <li>Sunset State Beach CA, USA</li> <li>Torrey Pines CA, USA</li> <li>Narrabeen, NSW, Australia</li> <li>Elwha WA, USA</li> <li>Ventura region, CA, USA</li> <li>Klamath region, CA USA</li> </ol> <p>Imagery are a mixture of 10-m Sentinel-2 and 15-m pansharpened Landsat 7, 8, and 9 visible-band imagery of various sizes. Red, Green, and Blue bands only</p> <p><strong>File descriptions</strong></p> <ol> <li>classes.txt, a file containing the class names</li> <li>images.zip, a zipped folder containing the 3-band images of varying sizes and extents</li> <li>labels.zip, a zipped folder containing the 1-band label images</li> <li>overlays.zip, a zipped folder containing a semi-transparent overlay of the color-coded label on the image (blue=0=water, red=1=whitewater, yellow=2=sediment, green=3=other)</li> <li>resized_images.zip, RGB images resized to 512x512x3 pixels</li> <li>resized_labels.zip, label images resized to 512x512 pixels</li> </ol> <p><strong>References</strong></p> <p>*Doodler: Buscombe, D., Goldstein, E.B., Sherwood, C.R., Bodine, C., Brown, J.A., Favela, J., Fitzpatrick, S., Kranenburg, C.J., Over, J.R., Ritchie, A.C. and Warrick, J.A., 2021. Human‐in‐the‐Loop Segmentation of Earth Surface Imagery. Earth and Space Science, p.e2021EA002085<a href="https://doi.org/10.1029/2021EA002085">https://doi.org/10.1029/2021EA002085</a>. See <a href="https://github.com/Doodleverse/dash_doodler">https://github.com/Doodleverse/dash_doodler.</a></p> <p>**Segmentation Gym: Buscombe, D., & Goldstein, E. B. (2022). A reproducible and reusable pipeline for segmentation of geoscientific imagery. Earth and Space Science, 9, e2022EA002332. <a href="https://doi.org/10.1029/2022EA002332">https://doi.org/10.1029/2022EA002332</a> See: <a href="https://github.com/Doodleverse/segmentation_gym">https://github.com/Doodleverse/segmentation_gym</a></p> <p>***Coast Train data release: Wernette, P.A., Buscombe, D.D., Favela, J., Fitzpatrick, S., and Goldstein E., 2022, Coast Train--Labeled imagery for training and evaluation of data-driven models for image segmentation: U.S. Geological Survey data release, <a href="https://doi.org/10.5066/P91NP87I">https://doi.org/10.5066/P91NP87I</a>. See <a href="https://coasttrain.github.io/CoastTrain/">https://coasttrain.github.io/CoastTrain/ </a>for more information</p> <p> </p>
Doodleverse/Segmentation Zoo Res-UNet models for 4-class (water, whitewater, sediment and other) segmentation of Sentinel-2 and Landsat-7/8 3-band (RGB) images of coasts.
<p><em><strong>Doodleverse/Segmentation Zoo Res-UNet models for 4-class (water, whitewater, sediment and other) segmentation of Sentinel-2 and Landsat-7/8 3-band (RGB) images of coasts.</strong></em></p> <p> </p> <p>These Residual-UNet model data are based on RGB (red, green, and blue) images of coasts and associated labels.</p> <p> </p> <p>Models have been created using Segmentation Gym* using the following dataset**: <a href="https://doi.org/10.5281/zenodo.7335647">https://doi.org/10.5281/zenodo.7335647</a></p> <p>Classes: {0=water, 1=whitewater, 2=sediment, 3=other}</p> <p><strong>File descriptions</strong></p> <p>For each model, there are 5 files with the same root name:</p> <p>1. <strong>'.json' </strong>config file: this is the file that was used by Segmentation Gym* to create the weights file. It contains instructions for how to make the model and the data it used, as well as instructions for how to use the model for prediction. It is a handy wee thing and mastering it means mastering the entire Doodleverse.</p> <p>2.<strong> '.h5'</strong> weights file: this is the file that was created by the Segmentation Gym* function `train_model.py`. It contains the trained model's parameter weights. It can called by the Segmentation Gym* function `seg_images_in_folder.py`. Models may be ensembled.</p> <p>3.<strong> '_modelcard.json'</strong> model card file: this is a json file containing fields that collectively describe the model origins, training choices, and dataset that the model is based upon. There is some redundancy between this file and the `config` file (described above) that contains the instructions for the model training and implementation. The model card file is not used by the program but is important metadata so it is important to keep with the other files that collectively make the model and is such is considered part of the model</p> <p>4. <strong> '_model_history.npz'</strong> model training history file: this numpy archive file contains numpy arrays describing the training and validation losses and metrics. It is created by the Segmentation Gym function `train_model.py`</p> <p>5. <strong> '.png'</strong> model training loss and mean IoU plot: this png file contains plots of training and validation losses and mean IoU scores during model training. A subset of data inside the .npz file. It is created by the Segmentation Gym function `train_model.py`</p> <p>Additionally, BEST_MODEL.txt contains the name of the model with the best validation loss and mean IoU</p> <p> </p> <p><strong>References</strong></p> <p>*Segmentation Gym: Buscombe, D., & Goldstein, E. B. (2022). A reproducible and reusable pipeline for segmentation of geoscientific imagery. Earth and Space Science, 9, e2022EA002332. <a href="https://doi.org/10.1029/2022EA002332">https://doi.org/10.1029/2022EA002332</a> See: <a href="https://github.com/Doodleverse/segmentation_gym">https://github.com/Doodleverse/segmentation_gym</a></p> <p>** Buscombe, Daniel, Goldstein, Evan, Bernier, Julie, Bosse, Stephen, Colacicco, Rosa, Corak, Nick, Fitzpatrick, Sharon, del Jesús González Guillén, Anais, Ku, Venus, Paprocki, Julie, Platt, Lindsay, Steele, Bethel, Wright, Kyle, & Yasin, Brandon. (2022). Images and 4-class labels for semantic segmentation of Sentinel-2 and Landsat RGB satellite images of coasts (water, whitewater, sediment, other) (v1.0) [Data set]. Zenodo. <a href="https://doi.org/10.5281/zenodo.7335647">https://doi.org/10.5281/zenodo.7335647</a></p> <p> </p>
Images and 2-class labels for semantic segmentation of Sentinel-2 and Landsat RGB satellite images of coasts (water, other)
<p><em><strong>Images and 2-class labels for semantic segmentation of Sentinel-2 and Landsat RGB satellite images of coasts (water, other)</strong></em></p> <p>Images and 2-class labels for semantic segmentation of Sentinel-2 and Landsat RGB satellite images of coasts (water, other)</p> <p><strong>Description</strong></p> <p>4088 images and 4088 associated labels for semantic segmentation of Sentinel-2 and Landsat RGB satellite images of coasts. The 2 classes are 1=water, 0=other. Imagery are a mixture of 10-m Sentinel-2 and 15-m pansharpened Landsat 7, 8, and 9 visible-band imagery of various sizes. Red, Green, Blue bands only</p> <p>These images and labels could be used within numerous Machine Learning frameworks for image segmentation, but have specifically been made for use with the Doodleverse software package, Segmentation Gym**.</p> <p>Two data sources have been combined</p> <p><strong>Dataset 1</strong></p> <ul> <li>1018 image-label pairs from the following data release**** https://doi.org/10.5281/zenodo.7335647</li> <li>Labels have been reclassified from 4 classes to 2 classes.</li> <li>Some (422) of these images and labels were originally included in the Coast Train*** data release, and have been modified from their original by reclassifying from the original classes to the present 2 classes.</li> <li>These images and labels have been made using the Doodleverse software package, Doodler*.</li> </ul> <p><strong>Dataset 2</strong></p> <ul> <li>3070 image-label pairs from the Sentinel-2 Water Edges Dataset (SWED)***** dataset, https://openmldata.ukho.gov.uk/, described by Seale et al. (2022)******</li> <li>A subset of the original SWED imagery (256 x 256 x 12) and labels (256 x 256 x 1) have been chosen, based on the criteria of more than 2.5% of the pixels represent water</li> </ul> <p><strong>File descriptions</strong></p> <ul> <li> classes.txt, a file containing the class names</li> <li> images.zip, a zipped folder containing the 3-band RGB images of varying sizes and extents</li> <li> labels.zip, a zipped folder containing the 1-band label images</li> <li> overlays.zip, a zipped folder containing a semi-transparent overlay of the color-coded label on the image (red=1=water, bllue=0=other)</li> <li> resized_images.zip, RGB images resized to 512x512x3 pixels</li> <li> resized_labels.zip, label images resized to 512x512x1 pixels</li> </ul> <p><strong>References</strong></p> <p>*Doodler: Buscombe, D., Goldstein, E.B., Sherwood, C.R., Bodine, C., Brown, J.A., Favela, J., Fitzpatrick, S., Kranenburg, C.J., Over, J.R., Ritchie, A.C. and Warrick, J.A., 2021. Human‐in‐the‐Loop Segmentation of Earth Surface Imagery. Earth and Space Science, p.e2021EA002085https://doi.org/10.1029/2021EA002085. See https://github.com/Doodleverse/dash_doodler.</p> <p>**Segmentation Gym: Buscombe, D., & Goldstein, E. B. (2022). A reproducible and reusable pipeline for segmentation of geoscientific imagery. Earth and Space Science, 9, e2022EA002332. https://doi.org/10.1029/2022EA002332 See: https://github.com/Doodleverse/segmentation_gym</p> <p>***Coast Train data release: Wernette, P.A., Buscombe, D.D., Favela, J., Fitzpatrick, S., and Goldstein E., 2022, Coast Train--Labeled imagery for training and evaluation of data-driven models for image segmentation: U.S. Geological Survey data release, https://doi.org/10.5066/P91NP87I. See https://coasttrain.github.io/CoastTrain/ for more information</p> <p>****Buscombe, Daniel, Goldstein, Evan, Bernier, Julie, Bosse, Stephen, Colacicco, Rosa, Corak, Nick, Fitzpatrick, Sharon, del Jesús González Guillén, Anais, Ku, Venus, Paprocki, Julie, Platt, Lindsay, Steele, Bethel, Wright, Kyle, & Yasin, Brandon. (2022). Images and 4-class labels for semantic segmentation of Sentinel-2 and Landsat RGB satellite images of coasts (water, whitewater, sediment, other) (v1.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.7335647</p> <p>*****Seale, C., Redfern, T., Chatfield, P. 2022. Sentinel-2 Water Edges Dataset (SWED) https://openmldata.ukho.gov.uk/</p> <p>******Seale, C., Redfern, T., Chatfield, P., Luo, C. and Dempsey, K., 2022. Coastline detection in satellite imagery: A deep learning approach on new benchmark data. Remote Sensing of Environment, 278, p.113044.</p>
Doodleverse/Segmentation Zoo Res-UNet models for 2-class (water, other) segmentation of Sentinel-2 and Landsat-7/8 3-band (RGB) images of coasts.
<p><em><strong>Doodleverse/Segmentation Zoo Res-UNet models for 2-class (water, other) segmentation of Sentinel-2 and Landsat-7/8 3-band (RGB) images of coasts.</strong></em></p> <p>These Residual-UNet model data are based on RGB (red, green, and blue) images of coasts and associated labels.</p> <p>Models have been created using Segmentation Gym* using the following dataset**: <a href="https://doi.org/10.5281/zenodo.7384242">https://doi.org/10.5281/zenodo.7384242</a></p> <p>Classes: {0=other, 1=water}</p> <p><strong>File descriptions</strong></p> <p>For each model, there are 5 files with the same root name:</p> <p>1. <strong>'.json' </strong>config file: this is the file that was used by Segmentation Gym* to create the weights file. It contains instructions for how to make the model and the data it used, as well as instructions for how to use the model for prediction. It is a handy wee thing and mastering it means mastering the entire Doodleverse.</p> <p>2.<strong> '.h5'</strong> weights file: this is the file that was created by the Segmentation Gym* function `train_model.py`. It contains the trained model's parameter weights. It can called by the Segmentation Gym* function `seg_images_in_folder.py`. Models may be ensembled.</p> <p>3.<strong> '_modelcard.json'</strong> model card file: this is a json file containing fields that collectively describe the model origins, training choices, and dataset that the model is based upon. There is some redundancy between this file and the `config` file (described above) that contains the instructions for the model training and implementation. The model card file is not used by the program but is important metadata so it is important to keep with the other files that collectively make the model and is such is considered part of the model</p> <p>4. <strong> '_model_history.npz'</strong> model training history file: this numpy archive file contains numpy arrays describing the training and validation losses and metrics. It is created by the Segmentation Gym function `train_model.py`</p> <p>5. <strong> '.png'</strong> model training loss and mean IoU plot: this png file contains plots of training and validation losses and mean IoU scores during model training. A subset of data inside the .npz file. It is created by the Segmentation Gym function `train_model.py`</p> <p>Additionally, BEST_MODEL.txt contains the name of the model with the best validation loss and mean IoU</p> <p> </p> <p><strong>References</strong></p> <p>*Segmentation Gym: Buscombe, D., & Goldstein, E. B. (2022). A reproducible and reusable pipeline for segmentation of geoscientific imagery. Earth and Space Science, 9, e2022EA002332. <a href="https://doi.org/10.1029/2022EA002332">https://doi.org/10.1029/2022EA002332</a> See: <a href="https://github.com/Doodleverse/segmentation_gym">https://github.com/Doodleverse/segmentation_gym</a></p> <p>** Buscombe, D. (2022). Images and 2-class labels for semantic segmentation of Sentinel-2 and Landsat RGB satellite images of coasts (water, other) (v1.0) [Data set]. Zenodo. <a href="https://doi.org/10.5281/zenodo.7384242">https://doi.org/10.5281/zenodo.7384242</a></p>
Doodleverse/Segmentation Zoo/Seg2Map Res-UNet models for FloodNet/10-class segmentation of RGB 768x512 UAV images
<p><em><strong>Doodleverse/Segmentation Zoo/Seg2Map Res-UNet models for FloodNet/10-class segmentation of RGB 768x512 UAV images</strong></em></p> <p>These Residual-UNet model data are based on [FloodNet](https://github.com/BinaLab/FloodNet-Challenge-EARTHVISION2021) images and associated labels.</p> <p>Models have been created using Segmentation Gym* using the following dataset**: https://github.com/BinaLab/FloodNet-Challenge-EARTHVISION2021</p> <p>Image size used by model: 768 x 512 x 3 pixels</p> <p><em>classes:</em><br> 1. Background<br> 2. Building-flooded<br> 3. Building-non-flooded<br> 4. Road-flooded<br> 5. Road-non-flooded<br> 6. Water<br> 7. Tree<br> 8. Vehicle<br> 9. Pool<br> 10. Grass</p> <p><em>File descriptions</em></p> <p>For each model, there are 5 files with the same root name:</p> <p>1. '.json' config file: this is the file that was used by Segmentation Gym* to create the weights file. It contains instructions for how to make the model and the data it used, as well as instructions for how to use the model for prediction. It is a handy wee thing and mastering it means mastering the entire Doodleverse.</p> <p>2. '.h5' weights file: this is the file that was created by the Segmentation Gym* function `train_model.py`. It contains the trained model's parameter weights. It can called by the Segmentation Gym* function `seg_images_in_folder.py`. Models may be ensembled.</p> <p>3. '_modelcard.json' model card file: this is a json file containing fields that collectively describe the model origins, training choices, and dataset that the model is based upon. There is some redundancy between this file and the `config` file (described above) that contains the instructions for the model training and implementation. The model card file is not used by the program but is important metadata so it is important to keep with the other files that collectively make the model and is such is considered part of the model</p> <p>4. '_model_history.npz' model training history file: this numpy archive file contains numpy arrays describing the training and validation losses and metrics. It is created by the Segmentation Gym function `train_model.py`</p> <p>5. '.png' model training loss and mean IoU plot: this png file contains plots of training and validation losses and mean IoU scores during model training. A subset of data inside the .npz file. It is created by the Segmentation Gym function `train_model.py`</p> <p>Additionally, BEST_MODEL.txt contains the name of the model with the best validation loss and mean IoU</p> <p>images.zip and labels.zip contain the images and labels, respectively, used to train the model</p> <p><em>References</em><br> *Segmentation Gym: Buscombe, D., & Goldstein, E. B. (2022). A reproducible and reusable pipeline for segmentation of geoscientific imagery. Earth and Space Science, 9, e2022EA002332. https://doi.org/10.1029/2022EA002332 See: https://github.com/Doodleverse/segmentation_gym</p> <p>** Rahnemoonfar, M., Chowdhury, T., Sarkar, A., Varshney, D., Yari, M. and Murphy, R.R., 2021. Floodnet: A high resolution aerial imagery dataset for post flood scene understanding. IEEE Access, 9, pp.89644-89654.</p>
Doodleverse/Segmentation Zoo/Seg2Map Res-UNet models for CoastTrain water/other segmentation of RGB 768x768 orthomosaic images
<p><em><strong>Doodleverse/Segmentation Zoo/Seg2Map Res-UNet models for CoastTrain water/other segmentation of RGB 768x768 orthomosaic images</strong></em></p> <p>These Residual-UNet model data are based on Coast Train images and associated labels. https://coasttrain.github.io/CoastTrain/docs/Version%201:%20March%202022/data</p> <p>Models have been created using Segmentation Gym* using the following dataset**: https://doi.org/10.1038/s41597-023-01929-2</p> <p>Image size used by model: 768 x 768 x 3 pixels</p> <p><em>classes:</em><br> 1. Water<br> 2. Other</p> <p><em>File descriptions</em></p> <p>For each model, there are 5 files with the same root name:</p> <p>1. '.json' config file: this is the file that was used by Segmentation Gym* to create the weights file. It contains instructions for how to make the model and the data it used, as well as instructions for how to use the model for prediction. It is a handy wee thing and mastering it means mastering the entire Doodleverse.</p> <p>2. '.h5' weights file: this is the file that was created by the Segmentation Gym* function `train_model.py`. It contains the trained model's parameter weights. It can called by the Segmentation Gym* function `seg_images_in_folder.py`. Models may be ensembled.</p> <p>3. '_modelcard.json' model card file: this is a json file containing fields that collectively describe the model origins, training choices, and dataset that the model is based upon. There is some redundancy between this file and the `config` file (described above) that contains the instructions for the model training and implementation. The model card file is not used by the program but is important metadata so it is important to keep with the other files that collectively make the model and is such is considered part of the model</p> <p>4. '_model_history.npz' model training history file: this numpy archive file contains numpy arrays describing the training and validation losses and metrics. It is created by the Segmentation Gym function `train_model.py`</p> <p>5. '.png' model training loss and mean IoU plot: this png file contains plots of training and validation losses and mean IoU scores during model training. A subset of data inside the .npz file. It is created by the Segmentation Gym function `train_model.py`</p> <p>Additionally, BEST_MODEL.txt contains the name of the model with the best validation loss and mean IoU</p> <p><em>References</em><br> *Segmentation Gym: Buscombe, D., & Goldstein, E. B. (2022). A reproducible and reusable pipeline for segmentation of geoscientific imagery. Earth and Space Science, 9, e2022EA002332. https://doi.org/10.1029/2022EA002332 See: https://github.com/Doodleverse/segmentation_gym</p> <p>**Buscombe, D., Wernette, P., Fitzpatrick, S. <em>et al.</em> A 1.2 Billion Pixel Human-Labeled Dataset for Data-Driven Classification of Coastal Environments. <em>Sci Data</em> <strong>10</strong>, 46 (2023). https://doi.org/10.1038/s41597-023-01929-2</p>
Doodleverse/Segmentation Zoo/Seg2Map Res-UNet models for FloodNet/10-class segmentation of RGB 1024x768 UAV images
<p><em><strong>Doodleverse/Segmentation Zoo/Seg2Map Res-UNet models for FloodNet/10-class segmentation of RGB 1024x768<strong> </strong>UAV images</strong></em></p> <p>These Residual-UNet model data are based on [FloodNet](https://github.com/BinaLab/FloodNet-Challenge-EARTHVISION2021) images and associated labels.</p> <p>Models have been created using Segmentation Gym* using the following dataset**: https://github.com/BinaLab/FloodNet-Challenge-EARTHVISION2021</p> <p>Image size used by model: 1024 x 768 x 3 pixels</p> <p><em>classes:</em><br> 1. Background<br> 2. Building-flooded<br> 3. Building-non-flooded<br> 4. Road-flooded<br> 5. Road-non-flooded<br> 6. Water<br> 7. Tree<br> 8. Vehicle<br> 9. Pool<br> 10. Grass</p> <p><em>File descriptions</em></p> <p>For each model, there are 5 files with the same root name:</p> <p>1. '.json' config file: this is the file that was used by Segmentation Gym* to create the weights file. It contains instructions for how to make the model and the data it used, as well as instructions for how to use the model for prediction. It is a handy wee thing and mastering it means mastering the entire Doodleverse.</p> <p>2. '.h5' weights file: this is the file that was created by the Segmentation Gym* function `train_model.py`. It contains the trained model's parameter weights. It can called by the Segmentation Gym* function `seg_images_in_folder.py`. Models may be ensembled.</p> <p>3. '_modelcard.json' model card file: this is a json file containing fields that collectively describe the model origins, training choices, and dataset that the model is based upon. There is some redundancy between this file and the `config` file (described above) that contains the instructions for the model training and implementation. The model card file is not used by the program but is important metadata so it is important to keep with the other files that collectively make the model and is such is considered part of the model</p> <p>4. '_model_history.npz' model training history file: this numpy archive file contains numpy arrays describing the training and validation losses and metrics. It is created by the Segmentation Gym function `train_model.py`</p> <p>5. '.png' model training loss and mean IoU plot: this png file contains plots of training and validation losses and mean IoU scores during model training. A subset of data inside the .npz file. It is created by the Segmentation Gym function `train_model.py`</p> <p>Additionally, BEST_MODEL.txt contains the name of the model with the best validation loss and mean IoU</p> <p> </p> <p><em>References</em><br> *Segmentation Gym: Buscombe, D., & Goldstein, E. B. (2022). A reproducible and reusable pipeline for segmentation of geoscientific imagery. Earth and Space Science, 9, e2022EA002332. https://doi.org/10.1029/2022EA002332 See: https://github.com/Doodleverse/segmentation_gym</p> <p>** Rahnemoonfar, M., Chowdhury, T., Sarkar, A., Varshney, D., Yari, M. and Murphy, R.R., 2021. Floodnet: A high resolution aerial imagery dataset for post flood scene understanding. IEEE Access, 9, pp.89644-89654.</p>
Doodleverse/Segmentation Zoo/Seg2Map Res-UNet models for CoastTrain/8-class segmentation of RGB 768x768 NAIP images
<p><em><strong>Doodleverse/Segmentation Zoo/Seg2Map Res-UNet models for CoastTrain 8-class segmentation of RGB 768x768 NAIP images</strong></em></p> <p>These Residual-UNet model data are based on Coast Train images and associated labels. https://coasttrain.github.io/CoastTrain/docs/Version%201:%20March%202022/data</p> <p>Models have been created using Segmentation Gym* using the following dataset**: https://doi.org/10.1038/s41597-023-01929-2</p> <p>Image size used by model: 768 x 768 x 3 pixels</p> <p>classes:</p> <p>water<br> whitewater<br> sediment<br> other_bare_natural_terrain<br> marsh_vegetation<br> terrestrial_vegetation<br> agricultural<br> development</p> <p>File descriptions</p> <p>For each model, there are 5 files with the same root name:</p> <p>1. '.json' config file: this is the file that was used by Segmentation Gym* to create the weights file. It contains instructions for how to make the model and the data it used, as well as instructions for how to use the model for prediction. It is a handy wee thing and mastering it means mastering the entire Doodleverse.</p> <p>2. '.h5' weights file: this is the file that was created by the Segmentation Gym* function `train_model.py`. It contains the trained model's parameter weights. It can called by the Segmentation Gym* function `seg_images_in_folder.py`. Models may be ensembled.</p> <p>3. '_modelcard.json' model card file: this is a json file containing fields that collectively describe the model origins, training choices, and dataset that the model is based upon. There is some redundancy between this file and the `config` file (described above) that contains the instructions for the model training and implementation. The model card file is not used by the program but is important metadata so it is important to keep with the other files that collectively make the model and is such is considered part of the model</p> <p>4. '_model_history.npz' model training history file: this numpy archive file contains numpy arrays describing the training and validation losses and metrics. It is created by the Segmentation Gym function `train_model.py`</p> <p>5. '.png' model training loss and mean IoU plot: this png file contains plots of training and validation losses and mean IoU scores during model training. A subset of data inside the .npz file. It is created by the Segmentation Gym function `train_model.py`</p> <p>Additionally, BEST_MODEL.txt contains the name of the model with the best validation loss and mean IoU</p> <p>References<br> *Segmentation Gym: Buscombe, D., & Goldstein, E. B. (2022). A reproducible and reusable pipeline for segmentation of geoscientific imagery. Earth and Space Science, 9, e2022EA002332. https://doi.org/10.1029/2022EA002332 See: https://github.com/Doodleverse/segmentation_gym</p> <p>**Buscombe, D., Wernette, P., Fitzpatrick, S. et al. A 1.2 Billion Pixel Human-Labeled Dataset for Data-Driven Classification of Coastal Environments. Sci Data 10, 46 (2023). https://doi.org/10.1038/s41597-023-01929-2</p> <p> </p>
Doodleverse/Segmentation Zoo/Seg2Map Res-UNet models for OpenEarthMap/9-class segmentation of RGB 512x512 high-res. images
<p><em><strong>Doodleverse/Segmentation Zoo/Seg2Map Res-UNet models for OpenEarthMap/9-class segmentation of RGB 512x512 high-res. images</strong></em></p> <p>These Residual-UNet model data are based on the [OpenEarthMap dataset](https://open-earth-map.org/)</p> <p>Models have been created using Segmentation Gym* using the following dataset**: https://zenodo.org/record/7223446#.Y9gtWHbMIuV </p> <p>Image size used by model: 512 x 512 x 3 pixels</p> <p>classes:<br> 1. bareland<br> 2. rangeland<br> 3. development<br> 4. road<br> 5. tree<br> 6. water<br> 7. agricultural<br> 8. building<br> 9. nodata</p> <p>File descriptions</p> <p>For each model, there are 5 files with the same root name:</p> <p>1. '.json' config file: this is the file that was used by Segmentation Gym* to create the weights file. It contains instructions for how to make the model and the data it used, as well as instructions for how to use the model for prediction. It is a handy wee thing and mastering it means mastering the entire Doodleverse.</p> <p>2. '.h5' weights file: this is the file that was created by the Segmentation Gym* function `train_model.py`. It contains the trained model's parameter weights. It can called by the Segmentation Gym* function `seg_images_in_folder.py`. Models may be ensembled.</p> <p>3. '_modelcard.json' model card file: this is a json file containing fields that collectively describe the model origins, training choices, and dataset that the model is based upon. There is some redundancy between this file and the `config` file (described above) that contains the instructions for the model training and implementation. The model card file is not used by the program but is important metadata so it is important to keep with the other files that collectively make the model and is such is considered part of the model</p> <p>4. '_model_history.npz' model training history file: this numpy archive file contains numpy arrays describing the training and validation losses and metrics. It is created by the Segmentation Gym function `train_model.py`</p> <p>5. '.png' model training loss and mean IoU plot: this png file contains plots of training and validation losses and mean IoU scores during model training. A subset of data inside the .npz file. It is created by the Segmentation Gym function `train_model.py`</p> <p>Additionally, BEST_MODEL.txt contains the name of the model with the best validation loss and mean IoU</p> <p>References<br> *Segmentation Gym: Buscombe, D., & Goldstein, E. B. (2022). A reproducible and reusable pipeline for segmentation of geoscientific imagery. Earth and Space Science, 9, e2022EA002332. https://doi.org/10.1029/2022EA002332 See: https://github.com/Doodleverse/segmentation_gym</p> <p>**Xia, Yokoya, Adriano, & Broni-Bediako. (2022). OpenEarthMap: A Benchmark Dataset for Global High-Resolution Land Cover Mapping [Data set]. Zenodo. https://doi.org/10.5281/zenodo.7223446</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.