Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

70

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

70 results for “Semantic Segmentation”

Learn how ShareScore rates datasets ↗
zenodo40/100

Dataset for generating LOD3 building models from structure-from-motion and semantic segmentation

<p>This repository contains the codes for computing geometrical digital twins as LOD3 models for buildings, using a structure from motion and semantic segmentation. The methodology hereby implements was presented in the paper [Generating LOD3 building models from structure-from-motion and semantic segmentation&quot; by Pantoja-Rosero et., al. (2022)] (<a href="https://doi.org/10.1016/j.autcon.2022.104430">https://doi.org/10.1016/j.autcon.2022.104430</a>)</p>

opencc-by-4.0Jun 2022View details →
zenodo40/100

A Dataset of Synthetic Images of Outdoor Scenes Taken from Sidewalks, for Temporal Semantic Segmentation Applications

<p>This dataset has been generated using the CARLA simulator (release 0.9.11), an open-source 3D simulator for experiments in autonomous vehicle, based on the Unreal Engine game engine. It comes with pre-made city environment maps. CARLA is distributed with several integrated maps as well as parameters to increase the variety in the dataset. In the release that we have used, there are 13 semantic segmentation classes: None, Building, Fence, Other, Pedestrian, Pole, Lane-marking, Road, Sidewalk, Vegetation, Vehicle, Wall, and Traffic sign. The &quot;None&quot; category corresponds to textures that are not part of an object, such as lawns which are not part of &quot;Vegetation&quot;, or sky. In the &ldquo;Other&rdquo; category are found objects that are not included in the other classes like plant and flower pots. For smart mobility applications, the &ldquo;Sidewalks&rdquo; and &ldquo;Road&rdquo; classes are of particular importance to find the way forward, as well as &ldquo;Buildings&rdquo; and &ldquo;Poles&rdquo; for obstacle avoidance. Sequences are made of 4 images. The dataset is composed of 46436 frames (11609 sequences) partitioned in 41024 frames (10256 sequences) for train, 2696 frames (674 sequences) for validation, and 2716 for test (679 sequences). The size of the images is 800 x 600 (resp. width x height).</p> <p>Additionaly, we have generated another smaller dataset with images taken from 2 different viewpoints: one located on the road and the other located on the sidewalk. The number of frames for train/validation/test is respectively 7288 (1822 sequences) partitioned in 6344 (1687 sequences) for train, 416 frames (104 sequences) for validation, and 424 for test (106 sequences). This smaller dataset is aimed at showing the importance of the viewpoint in the result of semantic segmentation. This can be done by cross-validation: learning on images taken from a viewpoint located on the road and test on images with a viewpoint located on the sidewalk, and vice versa.</p>

opencc-by-4.0Jul 2022View details →
zenodo40/100

Train and Evaluation Code, Road Classification Models and Test set of the paper "Insights into the Effects of Image Overlap and Image Size on Semantic Segmentation Models Trained for Road Surface Area Extraction from Aerial Orthophotography"

<p>This repository contains the Python scripts built for training and evaluation of the implementation, together with the test data and the resulting road segmentation models corresponding to the paper "Insights into the Effects of Image Overlap and Image Size on Semantic Segmentation Models Trained for Road Surface Area Extraction from Aerial Orthophotography". The scripts make use of the Tensorflow with Keras framework and their additional required dependencies.</p> <p>The training and validation set is based on the binary SROADEX dataset (<a href="../records/6482346">https://zenodo.org/records/6482346</a>) that was re-split into tiles that feature the image resolutions (256 x 256, 512 x 512, and 1024 x 1024 pixels) and image overlaps (0% and 12.5%) considered in this study. The data have been generated using scripts developed in Python using Open Source libraries (GDAL/OGR and MapScript) for rasterization of vector cartography that represents the axes of the different types of roads (urban, interurban and rural). This binary road data contains information from 16 full orthoimages (28.5 km * 18.5 km) with spatial resolution of 0.5 m/pixel from the insular and peninsular Spanish territory. Due to the size on disk of approximately 492 gigabytes, this training and validation data is only available upon request from the corresponding author. The test set has been generated from a novel area from Palencia (Spain) and features 18 million pixels labelled with the positive "Road" class. The test sets are provided in the repository for each resolution (with no overlap), so that additional DL models can be evaluated on the same data and compared with the results achieved in this study.</p> <p>The structure of the information shared in this repository is as follows:<br>The scripts have been grouped by tile resolution (256, 512 and 1024). First, the test set and the evaluation script can be found. For each tile resolution, there are two subfolders (corresponding to the "no overlap" and "12.5% overlap"). In each case, the Python scripts for training the models in the three repetitions are shared, and the trained models (H5 format) are shared in compressed form. Finally, for each resolution we also share the testing dataset which consists of two folders.</p> <p>The material is distributed under a CC-BY 4.0 license.</p>

opencc-by-4.0Jun 2024View details →
zenodo40/100

A semantic segmentation dataset of Arctic sea ice from Operation IceBridge data

<p>This dataset is a semantic segmentation dataset of Arctic sea ice based on deep learning method from Operation IceBridge images. It contains 29,372 labeled images, each of which corresponds to an image of Operation IceBridge and are stored in TIFF format. The dataset can be accessed using ArcGIS, ENVI, and the GDAL library in Python easily. Where label 1 represents melt ponds, label 2 represents sea ice/snow, label 3 represents submerged ice, and label 4 represents open ocean water, respectively.</p>

opencc-by-4.0Sep 2024View details →
zenodo40/100

Dataset for semantic segmentation in NDT with step-heating thermography for CFRP laminates

<p>Dataset composed of 36 images (640x480 pixels) of 30 bands or channels from step-heating for the same Carbon Fiber Reinforced Polymer (CFRP) laminate.</p> <p>In each image the specimen is rotated 10&ordm; to generate new data, with different illumination, background, and heating/cooling sequences.</p> <p>Images are generated from a video composed of the heating process, which takes 10 seconds, and the cooling process, which takes another 10 seconds, to a total of 1000 frames per video. This video is then reduced to 30 images using different processing methods explained below.</p> <p>The 30 bands or channels consist of the following bands repeated for the heating and cooling processes:</p> <ol> <li>First PCT component</li> <li>Second PCT component</li> <li>Third PCT component</li> <li>Fourth PCT component</li> <li>PPT</li> <li>Kurtosis</li> <li>Skewness</li> <li>TSR Coefficient 7</li> <li>TSR Coefficient 6</li> <li>TSR Coefficient 5</li> <li>TSR Coefficient 4</li> <li>TSR Coefficient 3</li> <li>TSR Coefficient 2</li> <li>TSR Coefficient 1</li> <li>TSR Coefficient 0</li> </ol> <p>Please cite the original paper:</p> <p>LINK: TODO</p> <p>BibTex:</p> <p>TODO</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Sep 2021View details →
zenodo40/100

WE3DS: An RGB-D image dataset for semantic segmentation in agriculture

<p>Here, we introduce a novel RGB-D image database (WE3DS) for semantic segmentation in crop farming. It contains 2,568 RGB-D images (color image and distance map) and hand-annotated ground-truth masks for semantic segmentation and is the first RGB-D image dataset for multi-class plant species semantic segmentation task. Images were taken under natural light conditions using an RGB-D sensor consisting of two RGB cameras in a stereo setup.</p> <p>&nbsp;</p> <p><strong>Please cite the original source when using this dataset.</strong></p> <p>Kitzler, F.; Barta, N.; Neugschwandtner, R.W.; Gronauer, A.; Motsch, V. WE3DS: An RGB-D Image Dataset for Semantic Segmentation in Agriculture. <em>Sensors</em> <strong>2023</strong>, <em>23</em>, 2713. <a href="https://doi.org/10.3390/s23052713">https://doi.org/10.3390/s23052713 </a></p>

opencc-by-4.0Feb 2023View details →
zenodo40/100

SPVPANELEX: Dataset containing aerial orthoimages (covering 257.93 km2 of the Spanish territory, with a spatial resolution of 0.5 m) labelled with photovoltaic panel information for binary recognition and semantic segmentation

<p>The data have been generated using scripts developed in Python with Open-Source libraries (GDAL/OGR and MapScript) to rasterize of vector cartography representing the photovoltaic (PV) panels instalations in urban, industrial, and rural areas. This PV panels cartography has been generated by manual digitalizing the PV panels found latest aerial orthofotographs available on June 1, 2021 from Plano Nacional de Ortofotograf&iacute;a A&eacute;rea (PNOA), produced by the National Geographic Institute of Spain, using the Web Map Service PNOA-MA.<br> <br> The dataset consists of 239,680 images of 256 &times; 256 pixels in size, in png format, labelled with Class_1: &ldquo;Contains PV panel&rdquo; and Class_2: &ldquo;Does not contain PV panel&rdquo;, that were pre-divided with a split criterion of 70:10:20%. in train, validation and test folders, respectively.<br> <br> The structure of the data is as follows:<br> 1-Panels-Ortho and 1-Panels-Mask contain the images featuring PV panels and their corresponding ground truth mask for training the semantic segmentation networks.<br> 1-Panels-Ortho and 2-NoPanels-Ortho contain images containing and not containing PV panels, for the training of binary recognition models of PV panels.<br> <br> Moreover, in each folder the structure is the same: train, test, validation containing 70%, 10% and 20% of the total images and masks of each type.<br> <br> 1-Panels-Ortho<br> &nbsp; &nbsp; |----Train<br> &nbsp; &nbsp; |----Test<br> &nbsp; &nbsp; -----Validation<br> <br> 1-Panels-Mask<br> &nbsp; &nbsp; |----Train<br> &nbsp; &nbsp; |----Test<br> &nbsp; &nbsp; -----Validation<br> <br> 2-NoPanels-Ortho<br> &nbsp; &nbsp; |----Train<br> &nbsp; &nbsp; |----Test<br> &nbsp; &nbsp; -----Validation</p>

opencc-by-4.0Apr 2023View details →
zenodo40/100

NCCD-PF - A pre-failure narrow concrete cracks dataset for engineering structures damage classification and semantic segmentation

<p>The&nbsp;NCCD-PF dataset was developed for the classification and semantic segmentation of narrow concrete cracks in engineering structures elements at the pre-failure state. It only includes cracks whose width is narrower than 0.3 mm, i.e. the limit value specified in EC 1992-1-1 for typical elements of engineering structures and environmental conditions.</p> <p>This dataset is dedicated to the early crack detection at a stage when the serviceability limit state has not yet been exceeded and the failure of a structural element has not occurred. By implementing the early crack detection approach, it is possible to protect cracks in order to stop or slow down their propagation and thus to extend the structure's lifespan.</p> <p>This dataset contains images of cracks appearing on various elements of engineering structures (bridges, viaducts, tunnels) made of reinforced concrete (including abutments, tunnel walls, concrete barriers, pillars). The images were captured on construction sites and during inspections of engineering structures, at different stages of the reinforced concrete structure's working conditions - from the construction stage (when the elements are loaded only by their own weight) to the structure's use stage (when the elements are loaded by most of the design loads). The images are also differentiated by the cause of the cracking (ex., thermal and shrinkage stresses in young concrete, excessive stresses). The images were acquired using fixed-focus cameras without prior conditioning in order to represent the real working conditions of a bridge engineer during structural inspections. The images are characterised by a high degree of complexity due to the quality of the concrete surface finish (e.g. presence of formwork marks, concrete trowel marks), which could potentially be recognised&nbsp;as cracks.</p> <p>This dataset is dedicated to researchers working in the fields of computer vision, machine learning and deep learning. In particular, it contains domain knowledge in structural health monitoring, so that it can support the work of engineers in detecting cracks of concrete elements in a pre-failure state.</p> <p>A detailed description of the dataset is presented in <a href="https://www.nature.com/articles/s41597-023-02839-z" target="_blank" rel="noopener">A pre-failure narrow concrete cracks dataset for engineering structures damage classification and segmentation</a> (DOI: 10.1038/s41597-023-02839-z).</p>

opencc-by-4.0Aug 2023View details →
zenodo40/100

FOR-instance: a UAV laser scanning benchmark dataset for semantic and instance segmentation of individual trees

<p>The challenge of accurately segmenting individual trees from laser scanning data hinders the assessment of crucial tree parameters necessary for effective forest management, impacting many downstream applications. While dense laser scanning offers detailed 3D representations, automating the segmentation of trees and their structures from point clouds remains difficult. The lack of suitable benchmark datasets and reliance on small datasets have limited method development. The emergence of deep learning models exacerbates the need for standardized benchmarks.&nbsp;Addressing these gaps, the FOR-instance data represent a novel benchmarking dataset to enhance forest measurement using dense airborne laser scanning data, aiding researchers in advancing segmentation methods for forested 3D scenes.</p> <p>In this repository, users will&nbsp;find forest laser scanning point clouds from unamnned aerial vehicle (using Riegl sensors) that are manually segmented according to the individual trees (1130 trees) and semantic classes. The point clouds are subdivided into five data collections representing different forests in Norway, the Czech Republic, Austria, New Zealand, and Australia.&nbsp;</p> <p>These data are meant to be used either for developement of new methods (using the dev data) or for testing of exisitng methods (test data). The data splits are provided in the&nbsp;data_split_metadata.csv file.</p> <p>A full description of the FOR-instance data can be found at&nbsp;<a href="http://arxiv.org/abs/2309.01279">http://arxiv.org/abs/2309.01279</a>&nbsp;</p>

opencc-by-4.0Aug 2023View details →
zenodo40/100

WhiteRoadLines: Dataset of 27,025 images (256x256 pixeles at 0.15 m/ pixel) containing representative road lines and markings labelled for multi-class semantic segmentation

<p>The dataset consists of 27,025 PNG images (256x256 pixels) of high resolution aerial orthoimages at 0,15 m/pixel of resolution. The images contain information related to representative road lines and markings found on highway pavement and is labelled for multi-class semantic segmentation with tree classes of white road<br>lines and markings: (1) continuous line (black color), (2) dashed line (dark gray color) and (3) separation of entry and exit lanes (light gray color), together with (4) the background (white color).&nbsp;<br>&nbsp;</p><p>The dataset has been created in the framework of the SROADEX project to train a multiclass semantic segmentation process based on Deep Learning.<br>The data have been generated using scripts developed in Python using Open Source libraries (GDAL/OGR and MapScript) for rasterization of vector cartography representing the three different types of white road lines. This cartography has been obtained from Spanish official sources (National Geographic Institute) that we have<br>revised and edited in a meticulous and systematic way to verify that the road lines are represented on the cartography according to the orthoimages, available on January 1, 2022 in the download center of the National Center of Geographic Information (CNIG).&nbsp;</p><p>In the digitisation process, 46 homogeneously distributed areas of Spain have been selected. The orthoimages used have been resampled from the original resolution of 0,25m/pixel to 0,15m/pixel, as this is closer to the width of two of the three classes of white lines in the dataset. It resulted in 80% of the images for training (21622), 10% for validation (2702) and 10% for testing (2701). The following table summarises the number of pixels of each category included in each of the three sub-datasets</p><p>&nbsp;</p><p>Set&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Nº images &nbsp; Class_1 (continuous line) &nbsp; Class_2 (discontinuous line) Class_3 (line defining highway entrance or exit) &nbsp;Class_4 (background)</p><p>Train &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 21,622 &nbsp; &nbsp; &nbsp;27,633,537 &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 4,543,552 &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 3,284,380 &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 1,381,557,923</p><p>Validation &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 2,702 &nbsp; &nbsp; &nbsp; &nbsp;3,433,103 &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;570,741 &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;395,646 &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 172,678,782</p><p>Test &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;2,701 &nbsp; &nbsp; &nbsp; &nbsp;3,435,072 &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;536,838 &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;429,527 &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 172,611,299</p><p>Total &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 27,025 &nbsp; &nbsp; &nbsp;34,501,712 &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 5,651,131 &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 4,109,553 &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;1,726,848,004</p>

opencc-by-4.0Oct 2023View details →
zenodo36/100

Clouds dataset for semantic segmentation

<p>This database contains images used for training a fully convolutional neural network for the semantic segmentation of clouds in images from the Sentinel-2 Level 2A Satellite.</p> <p>The images refer to the RGB composition (bands 4, 3, and 2).</p> <p>After that, each band was converted to a byte type (0-255).</p> <p>The images are still divided into three main sets: training, validation and testing:</p> <ol> <li><strong>Training dataset: </strong>it contains 1.160 GeoTIFF images with 512x512 pixels and associated PNG masks (clouds indicated in white and background in black color).</li> <li><strong>Validation dataset</strong>: it contains 100 GeoTIFF images with 512x512 pixels and associated PNG masks used for validation step.</li> <li><strong>Test dataset:&nbsp;</strong>it contains 100 GeoTIFF images 512x512 pixels for testing.</li> </ol>

opencc-by-4.0Aug 2020View details →
zenodo36/100

Brazil's South Region landslide dataset for semantic segmentation

<p>This database contains images used for the semantic segmentation of landslide scars from a fully convolutional neural network U-Net of the three States of&nbsp;Brazil&#39;s South Region: Rio Grande do Sul, Santa Catarina and Paran&aacute;. Each .rar file contains 3 folders:</p> <p><strong>Image: </strong>GeoTIFF 8 bits images of locations with landslide scars.</p> <p><strong>Masks: </strong>PNG masks (scars indicated in white and background in black color).</p> <p><strong>Slope: </strong>GeoTIFF float rasters of slope (in degrees).</p>

opencc-by-4.0Dec 2020View details →
zenodo36/100

Amazon and Atlantic Forest image datasets for semantic segmentation

<p>This database contains images from<strong> Amazon </strong>and <strong>Atlantic Forest </strong>brazilian biomes used for training a fully convolutional neural network for the semantic segmentation of forested areas in images from the Sentinel-2 Level 2A Satellite.</p> <p>The images refer to the composition of bands 4, 3, 2 and 8. Each band was converted to a byte type (0-255).</p> <p>The images are still divided into three main sets: training, validation and testing:</p> <ol> <li><strong>Training dataset: </strong>it contains 499 and 485 GeoTIFF images (Amazon and Atlantic Forest, respectively) with 512x512 pixels and associated PNG masks (forest indicated in white and background in black color).</li> <li><strong>Validation dataset</strong>: it contains 100 GeoTIFF images for each biome with 512x512 pixels and associated PNG masks used for validation step.</li> <li><strong>Test dataset:&nbsp;</strong>it contains 20 GeoTIFF images for each biome with 512x512 pixels for testing.</li> </ol>

opencc-by-4.0Feb 2021View details →
zenodo36/100

Test Dataset for 3D semantic image segmentation of the various organs from CT and MR scans

<p>These test cases are for the <a href="https://github.com/MIC-DKFZ/nnUNet/releases/tag/v1.7.1">nnUnet v1</a> models trained on the following datasets:<br><br></p> <table> <tbody> <tr> <td>Dataset&nbsp;</td> <td>Task</td> <td>Model Details on Zenodo</td> </tr> <tr> <td>&nbsp;<a href="../record/6802614">TotalSegmentator</a>&nbsp;and&nbsp;<a href="../record/5903672">FLARE21</a> datasets</td> <td>Segment Liver from CT scans</td> <td>https://zenodo.org/record/8274976</td> </tr> <tr> <td><a href="https://kits-challenge.org/kits23/">KiTS23</a> datasets and a subset of the<a href="https://wiki.cancerimagingarchive.net/pages/viewpage.action?pageId=5800386#5800386566e265abf95408aa64c4917f0cbe5d9">&nbsp;TCGA-KIRC&nbsp;</a>dataset</td> <td>Segment Kidney, Cyst, and Tumors from CT Scans</td> <td>https://zenodo.org/records/8277846</td> </tr> <tr> <td><a href="http://ji%20yuanfeng.%20(2022).%20amos%20a%20large-scale%20abdominal%20multi-organ%20benchmark%20for%20versatile%20medical%20image%20segmentation%20[data%20set].%20zenodo.%20https">AMOS</a>&nbsp;and&nbsp;<a href="http://macdonald,%20jacob%20a.,%20zhu,%20zhe,%20konkel,%20brandon,%20mazurowski,%20maciej,%20wiggins,%20walter,%20&amp;%20bashir,%20mustafa.%20(2020).%20duke%20liver%20dataset%20(mri)%20v2%20(2.0.0)%20[data%20set].%20zenodo.%20https//doi.org/10.5281/zenodo.7774566">DUKE Liver</a> datasets</td> <td>Segment Liver from the MR scans</td> <td>https://zenodo.org/record/8290124</td> </tr> <tr> <td>Data from m&nbsp;<a href="../record/6624726">pi-cai</a></td> <td>Segment Prostate region from MR scans</td> <td>https://zenodo.org/record/8290093</td> </tr> </tbody> </table>

opencc-by-4.0Dec 2023View details →
zenodo36/100

A dataset for semantic segmentation of typical oceanic and atmospheric phenomena from Sentinel-1 images

<p>We have constructed a SAR (Synthetic Aperture Radar) image semantic segmentation dataset that includes 12 oceanic and atmospheric phenomena: Atmospheric Front (AF), Oceanic Front (OF), Rainfall (RF), Iceberg (IC), Sea Ice (SI), Pure Ocean Wave (POW), Wind Streak (WS), Low Wind Area (LWA), Biological Slick (BS), Micro Convective Cells (MCC), Internal Wave (IW), and Eddy.</p> <p>This dataset is built using Sentinel-1 IW and WV mode images. For WV mode data, we referenced TenGeoP-SARwv and SAR_WV_SemanticSegmentation and selected 2,383 images for semantic segmentation and annotation. For IW mode images, we incorporated some images from Tao et al.'s internal wave detection dataset. We selected 484 Sentinel-1 IW mode images obtained from 2015 to 2022 and divided them into 2,628 sub-images.</p> <p>The dataset contains a total of 5,011 image slices, with approximately 400 images for each phenomenon. All images are 16-bit .tiff files with a resolution of 100m and a size of 256x256 pixels. The images were manually annotated using the Labelme software, generating corresponding JSON files, which were then used to create the related annotation .png files.</p> <p>The updated version(V2) provides geographic information for each image.</p> <p>Thank you for your interest in our dataset. Here are the meanings of each label:</p> <p>1. BG: The unlabelled parts in JSON files are "BG" (Background)<br>2. AF: Atmospheric Front<br>3. BS: Biological Slick<br>4. I: &ldquo;I&rdquo; is equivalent to &ldquo;IB&rdquo;, representing icebergs<br>5. LWA: Low Wind Area<br>6. MCC: Micro Convective Cells<br>7. OF: Oceanic Front<br>8. POW: Pure Ocean Wave<br>9. RC: &ldquo;RC&rdquo; (Rain Cells) is equivalent to &ldquo;RF&rdquo; (Rainfall), both representing the&nbsp; rainfall phenomenon in the SAR image.&nbsp;<br>10. SI: Sea Ice<br>11. WS: Wind Streak<br>12. Eddy<br>13. IW: Internal Wave<br><em>14. HM: Represents the artificial objects appearing in the image, such as ships, aquaculture floating rafts, wind power facilities, etc.</em><br><em>15. OS: Unlike &ldquo;BS&rdquo;,&ldquo;OS&rdquo; represents mineral oil spills appearing in the SAR image (currently, there is insufficient data available for training, which will be supplemented in the future).</em></p>

opencc-by-4.0Jun 2024View details →
zenodo36/100

Tango Spacecraft Dataset for Region of Interest Estimation and Semantic Segmentation

<p><strong>Reference Paper:&nbsp;</strong></p> <p><a href="https://doi.org/10.1016/j.actaastro.2023.01.012"><strong>M. Bechini, M. Lavagna, P. Lunghi, Dataset generation and validation for spacecraft pose estimation via monocular images processing, Acta Astronautica 204 (2023) 358&ndash;369</strong></a></p> <p><a href="https://www.researchgate.net/publication/361924362_Spacecraft_Pose_Estimation_via_Monocular_Image_Processing_Dataset_Generation_and_Validation">M. Bechini, P. Lunghi, M. Lavagna. &quot;Spacecraft Pose Estimation via Monocular Image Processing: Dataset Generation and Validation&quot;. In 9th European Conference for Aeronautics and Aerospace Sciences (EUCASS)</a></p> <p><strong>General Description:</strong></p> <p>The &quot;<em>Tango Spacecraft Dataset for Region of Interest Estimation and Semantic Segmentation</em>&quot; dataset here published should be used for Region of Interest (ROI) and/or semantic segmentation tasks. It is split into 30002 train images and 3002 test images representing the Tango spacecraft from Prisma mission, being the largest publicly available dataset of synthetic space-borne noise-free images tailored to ROI extraction and Semantic Segmentation tasks (up to our knowledge). The label of each image gives, for the Bounding Box annotations, the filename of the image, the ROI top-left corner (minimum x, minimum y) in pixels, the ROI bottom-right corner (maximum x, maximum y) in pixels,&nbsp;and the center point of the ROI in pixels. The annotation are taken in image reference frame with the origin located at the top-left corner of the image, positive x rightward and positive y downward. Concerning the Semantic Segmentation, RGB masks are provided. Each RGB mask correspond to a single image in both train and test dataset. The RGB images are such that the R channel corresponds to the spacecraft, the G channel corresponds to the Earth (if present), and the B channel corresponds to the background (deep space). Per each channel the pixels have non-zero value only in correspondence of the object that they represent (Tango, Earth, Deep Space).&nbsp;More information on the dataset split and on the label format are reported below.&nbsp;</p> <p><strong>Images Information:</strong></p> <p>The dataset comprises 30002 synthetic grayscale images of Tango spacecraft from Prisma mission that serves as train set, while the test set is formed by 3002 synthetic grayscale images of Tango spacecraft from Prisma mission in PNG format.&nbsp;About 1/6 of the images both in the train and in the test set have a non-black background, obtained by rendering an Earth-like model in the raytracing process used to define the images reported.&nbsp;The images are noise-free to increase the flexibility of the dataset. The illumination direction of the spacecraft in the scene is uniformly distributed in the 3D space in agreement with the Sun position constraints.</p> <p><br> <strong>Labels Information:</strong></p> <p>Labels for the bounding box extraction are here provided in separated JSON files. The files are formatted per each image as in the following example:</p> <ul> <li>&nbsp; &nbsp; filename &nbsp; &nbsp;: tango_img_1 &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;# name of the image to which the data are referred</li> <li>&nbsp;&nbsp; &nbsp;rol_tl&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; : [x, y] &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; # ROI top-left corner (minimum x, minimum y) in pixels</li> <li>&nbsp; &nbsp;&nbsp;roi_br &nbsp; &nbsp; &nbsp; &nbsp;&nbsp;: [x, y] &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;# ROI bottom-right corner (maximum x, maximum y) in pixels</li> <li>&nbsp; &nbsp; roi_cc &nbsp; &nbsp; &nbsp; &nbsp; : [x, y] &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; #&nbsp;center point of the ROI in pixels</li> </ul> <p>Notice that the annotation are taken in image reference frame with the origin located at the top-left corner of the image, positive x rightward and positive y downward.To make&nbsp;the usage of the dataset easier, both the training set and the test set are split in two folders containing the images with earth as background and without background.</p> <p>Concerning the Semantic Segmentation Labels, they are provided as RGB masks named as &quot;filename_mask.png&quot; where &quot;filename&quot; is the filename of the image of the training set or the test set to which a specific mask is referred.&nbsp;The RGB images are such that the R channel corresponds to the spacecraft, the G channel corresponds to the Earth (if present), and the B channel corresponds to the background (deep space). Per each channel the pixels have non-zero value only in correspondence of the object that they represent (Tango, Earth, Deep Space).&nbsp;</p> <p><strong>VERSION CONTROL</strong></p> <ul> <li>v1.0: This version contains&nbsp;the dataset (both train and test) of full scale images with ROI annotations and RGB masks for Semantic Segmentation tasks. These images have width=height=1024 pixels. The position of tango with respect to the camera is randomly selected from a uniform distribution, but it is ensured the full visibility in all the images.&nbsp;</li> </ul> <p>Note: this dataset contains the same images of the&nbsp;<em>&quot;Tango Spacecraft Wireframe Dataset Model for Line Segments Detection&quot;</em>&nbsp;v2.0 full-scale&nbsp;(DOI:&nbsp;<a href="https://doi.org/10.5281/zenodo.6372848">https://doi.org/10.5281/zenodo.6372848</a>) and also &quot;<em>Tango Spacecraft Dataset for Monocular Pose Estimation</em>&quot; v1.0 (DOI: <a href="https://doi.org/10.5281/zenodo.6499007">https://doi.org/10.5281/zenodo.6499007</a>)&nbsp;and they can be used&nbsp;together by combining the annotations of the relative pose and the ones of the reprojected wireframe model of Tango, with also the ones of the ROI. <strong>These three datasets give the most comprehensive dataset of space borne synthetic images ever published</strong> (up to our knowledge).</p>

opencc-by-nc-4.0Apr 2022View details →
zenodo36/100

Strawberry dataset for Semantic Segmentation

<p>The dataset was annotated using the labelme tool, and it was trained using the <a href="https://github.com/ayoolaolafenwa/PixelLib">pixellib</a>.</p> <p>&nbsp;</p>

opencc-by-4.0Jun 2022View details →
zenodo36/100

Data, code, models for "Weakly Supervised Semantic Segmentation for Joint Key Local Structure Localization and Classification of Aurora Image"

<p>Data, code and models for https://ieeexplore.ieee.org/document/8410588/</p>

opencc-by-4.0Jul 2018View details →
zenodo36/100

Amazon Rainforest dataset for semantic segmentation

<p>This database contains images used for the semantic segmentation of forest and non-forest areas from a fully convolutional neural network U-Net .</p> <p>1. <strong>Training dataset: </strong>it contains 30 GeoTIFF images with 512x512 pixels and associated PNG masks (forest indicated in white and non-forest in black color).</p> <p>2. <strong>Validation dataset</strong>: it contains 15&nbsp;GeoTIFF images with 512x512 pixels and associated PNG masks used for U-Net validation step.</p> <p>3. <strong>Test dataset:&nbsp;</strong>it contains 15&nbsp;GeoTIFF images 512x512 pixels for testing.</p>

opencc-by-4.0May 2019View details →
zenodo36/100

Semantic2D: A Semantic Dataset for 2D Lidar Semantic Segmentation: Dataset

<p>The 2D lidar semantic segmentation datasets for the paper titled "Semantic2D: A Semantic Dataset for 2D Lidar Semantic Segmentation" by Zhanteng Xie and Philip Dames.</p> <p>The relevant code is available at: https://github.com/TempleRAIL/semantic2d</p> <p>&nbsp;</p>

opencc-by-4.0Sep 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record