Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
38,240
datasets available to search
ShareScore release 0.7.1
Dataset results
38,240 results for “image”
SCLabels: Labelled rectified RGB images from the Spanish CoastSnap network
<h1>Training dataset</h1> <p><span>The SCLabels dataset is intended to be used in the exploring and development of Artificial Intelligence (AI) applications aimed at the automation of the shoreline extraction process from rectified images. SCLabels includes rectified RGB images from the Spanish CoastSnap network and their corresponding masks, together with a metadata file and a README file. RGB images encompass variable geographic locations, fields of view, beach types and degrees of occupation, tidal regimes, meteoceanic and lightning conditions, and a variety of environmental characteristics. Masks account for dense pixel labels including 5 categories: i) No data; ii) Not classified; iii) Landwards; iv) Seawards; and v) Shoreline. In the metadata file, images are linked to their corresponding masks, and information about the geographic location of each image, capture characteristics and image source, shoreline position and other auxiliary data are provided. The README file enhances the explainability and comprehension of the dataset, elaborating on the context and contents, and providing detailed explanations of the metadata, potential limitations, technical aspects of the image processing and annotation stages, usage recommendations, and related works. </span></p> <h1>Technical details</h1> <p>The SCLabels dataset version 1.0.0 is packaged in a compressed file (SCLabels_v1.0.0.zip). A total of 1717 RGB images are shared in JPG format, corresponding masks in PNG format, a metadata file in JSON format, and the README file in PDF format.</p> <h2>Data preprocessing</h2> <p><span>To generate the SCLabels masks, rectified RGB images and their corresponding shorelines were used. RGB images were cropped to the minimum and maximum alongshore pixel coordinates of the shoreline (vertical axis) plus 10 additional pixels above and below to preserve contextual information. A grayscale image was then derived from each cropped RGB image for subsequent pixel labelling. First, a binary mask was derived, marking "NoData'' for black and white padded pixels resulting from the registration and rectification steps. Subsequently, the shoreline was densified, ensuring at least one pixel per row was assigned the "Shoreline" label. Next, "Landwards" and "Seawards" labels were assigned to the right and left of the shoreline. Pixels left unlabelled were categorised as "NotClassified". Finally, masks’ values were reclassified to align with the predefined labels, and the grayscale masks were exported. For additional information, please consult the README file. </span></p> <h2>Data splitting</h2> <p><span>Data splitting requirements may vary depending on the chosen AI approach (e.g., splitting by entire images, image patches, or image rows). Researchers should use a consistent data splitting method and document the approach and splits used in publications. This transparency enables reproducible results and facilitates comparisons between studies.</span></p> <h2>Classes, labels and annotations</h2> <p><span>The SCLabels dataset includes one mask per rectified RGB image, sharing the same width and height. These masks are in greyscale and PNG format, and consist of five different labels:</span></p> <table> <tbody> <tr> <td><strong> Mask value</strong></td> <td><strong> Label</strong></td> <td><strong> Description</strong></td> </tr> <tr> <td>0</td> <td>NoData</td> <td>High probability of being black or white padded pixels, used to pad non-rectangular images within the image registration and rectification processes</td> </tr> <tr> <td>25</td> <td>NotClassified</td> <td>Not labeled pixels</td> </tr> <tr> <td>75</td> <td>Landwards</td> <td>All pixels that are towards the landside with respect to the shoreline (row-wise), excluding “NoData” ones</td> </tr> <tr> <td>150</td> <td>Seawards</td> <td>All pixels that are towards the seaside with respect to the shoreline (row-wise), excluding “NoData” ones</td> </tr> <tr> <td>255</td> <td>Shoreline</td> <td>Pixels intersected by the mapped shoreline densified to cover one pixel per row, at least</td> </tr> </tbody> </table> <h2>Parameters</h2> <p><span>RGB values or any transformation in the colour space can be used as parameters.</span><span> </span></p> <h2>Data sources</h2> <p><span>In the CoastSnap initiative, citizens capture images (oblique smartphone photos) from fixed CoastSnap stations and share them with the scientific managers. Images are subjected to a quality control process, spatially registered to a designated target image, and rectified (georeferencing). The shoreline is subsequently digitised from each rectified image.</span><span> </span></p> <h2>Data quality</h2> <p><span>All images included have been supervised by CSs’ scientific managers. However, citizen scientists take images by smartphones (different camera quality) at irregular intervals across various sites with varying weather and illumination conditions. Users of SCLabels dataset must be aware of this variance. </span></p> <h2>Image resolution</h2> <p><span>The resolution of the images depends on the CoastSnap station and the length of the shoreline, ranging from 241x188 pixels to 801x796 pixels.</span></p> <h2>Spatial coverage</h2> <p><span>The SCLabels dataset version 1.0.0 contains data from five Spanish CoastSnap stations, including sandy beaches in the northwest (</span><span>agrelo</span><span>), the Cíes Islands (</span><span>cies</span><span>), the south (</span><span>cadiz</span><span>), and the Balearic Islands (</span><span>samarador </span><span>and </span><span>arenaldentem</span><span>).</span></p> <table> <tbody> <tr> <td><strong> CoastSnap station</strong></td> <td><strong> Longitude</strong></td> <td><strong> Latitude</strong></td> </tr> <tr> <td><em>agrelo</em></td> <td>-8.772</td> <td>42.331</td> </tr> <tr> <td><em>cies</em></td> <td>-8.900</td> <td>42.226</td> </tr> <tr> <td><em>cadiz</em></td> <td>-6.288</td> <td>36.522</td> </tr> <tr> <td><em>samarador</em></td> <td>3.185</td> <td>39.350</td> </tr> <tr> <td><em>arenaldentem</em></td> <td>2.974</td> <td>39.353</td> </tr> </tbody> </table> <h2>Contact information</h2> <p><span>For further technical inquiries or additional information about the annotated dataset, please contact jsoriano@socib.es.</span></p>
Wollestraat 29, Bruges (BE): high-resolution images of dry wood cores taken form a medieval floor joists, for tree-ring analysis
<ul><li>Dry-wood cores taken from historical timbers of a floor joists in the medieval building 'De Oude Steen', Wollestraat 29, Bruges (Belgium).</li><li><a href="https://id.erfgoed.net/erfgoedobjecten/29956 ">https://id.erfgoed.net/erfgoedobjecten/29956 </a></li><li>The cores were sampled at 22/02/2023 with a dry-wood borer (internal diameter 12 mm, external diameter 19 mm).</li><li>The cores were surfaced with increasingly finer sanding papers, from P60 up to P4000.</li><li>The cores were photograpphed with a Sony alpha7R IV full frame camera and FE 90 mm F/2.8G macro lens.</li><li>The<a href="https://www.wsl.ch/en/services-produkte/skippy/"> Skippy</a> system served as the image capturing platform.</li><li>The individual digital macro-photos were stitched with PTGui into a mosaic image (.tiff).</li><li>The mosaic images have a resolution of ~4 µm.</li></ul>
VoroCrack3d: An annotated data set of 3d CT concrete images with synthetic crack structures
<p>VoroCrack3d is an annotated data set of 3d CT images of concrete with synthetic crack structures. Its main purpose is the training and testing of machine learning models for 3d crack segmentation. The data set comprises 1344 images together with their corresponding ground truths. The concrete backgrounds are cropped out sections of size 400x400x400 voxels of CT images of concrete. To this end, several different concrete samples were scanned (normal concrete (NC), high-performance concrete (HPC), ultra-high-performance concrete (UHPC), air pore concrete; without and with reinforcements (straight steel fibers, crimped steel fibers, hooked-end steel fibers, polypropylene fibers, fibers made of glass fiber-reinforced polymer). The original concrete images have a resolution between 2.8 and 106 micrometers.</p> <p>The crack structures are modeled via minimum-weight surfaces in Voronoi diagrams according to the paper</p> <p>[1] C. Jung, C. Redenbach, Crack Modeling via Minimum-Weight Surfaces in 3d Voronoi Diagrams, Journal of Mathematics in Industry, 13, 10 (2023). https://doi.org/10.1186/s13362-023-00138-1.</p> <p>The surfaces are discretized, dilated and superimposed on the concrete backgrounds.</p> <p>The data set offers a high variety regarding concrete types, noise levels and crack widths, shapes, regularity and branching. This makes it suitable for studying the generalizability and robustness of 3d crack segmentation methods.</p> <p>______________________________________________________________________________________________</p> <p>The folder 'data' contains seven subfolders, each containing the data generated from a specific concrete type (NC, HPC, air pore concrete, polypropylene fiber-reinforced concrete, steel fiber-reinforced concrete (straight, crimped and hooked-end steel fibers)).</p> <p>Each subfolder again contains four subfolders according to the point process model that was used for generating the 3d Voronoi diagrams. The point processes and Voronoi diagrams are restricted to windows of size 400x150x400. </p> <p>- 'hc': Hard core point process with 60% volume density and intensity 0.000025 obtained from force-biased sphere packing.<br>- 'matclust': Matérn cluster process with parent intensity 0.0002/50, offspring intensity 50 and cluster radius 20.<br>- 'ppp': Poisson point process with intensity 0.0002.<br>- 'ppp-scaled': Poisson point process with intensity 0.0002 (but inside 200x150x200 window). The resulting Voronoi diagram is stretched in x- and z- direction by a factor of 2.</p> <p>Each of these contains five subfolders: one for the 3d input images, two for the corresponding labels (ground truths; one with and one without pores/fibers), one for the input and label previews (slice z=200 for each of the images) and a misc folder containing the concrete background without crack and, if applicable, the pore/fiber segmentation image.</p> <p>The data itself then contains 48 images:<br>1a-1d: crack with up to seven branches; fixed crack width (~1 voxel).<br>2a-2d: crack with up to four branches; fixed crack width (~1 voxel).<br>3a-3d: crack with up to one branch; fixed crack width (~1 voxel).<br>4a-4d: crack with no branches; fixed crack width (~1 voxel).<br>5a-5d: crack with no branches; fixed crack width (~3 voxels).<br>6a-6d: crack with no branches; fixed crack width (~5 voxels).<br>7a-7d: crack with no branches; fixed crack width (~7 voxels).<br>8a-8d: crack with up to seven branches; multiscale crack (bernoulli parameter 0.01);<br>9a-9d: crack with up to seven branches; multiscale crack (bernoulli parameter 0.02);<br>10a-10d: crack with up to seven branches; multiscale crack (bernoulli parameter 0.05);<br>11a-11d: crack with up to seven branches; multiscale crack (bernoulli parameter 0.1);<br>12a-12d: crack with up to seven branches; multiscale crack (bernoulli parameter 0.2);</p> <p>The names 'a'-'d' indicate level of added noise added to the image:<br>a: None.<br>b: Uniformly on [-sigma,sigma] <br>c: Uniformly on [-2*sigma,2*sigma] <br>d: Uniformly on [-4*sigma,4*sigma] <br>Negative values are mapped to 0. <br>For inputs of type int, noise values are rounded to the nearest integer.<br>(sigma = standard deviation of voxel greyvalues in image)</p> <p>Note that the grey values in the ground truths correspond to the local crack width. They can be thresholded to obtain binary masks.</p> <p>For more details, we refer to [1].</p>
Hyperspectral Imaging Dataset for Laser Thermal Ablation Monitoring in Vital Organs
<p><strong>Objectives:</strong> The objective of the research was to use hyperspectral imaging (HSI) to detect thermal damage induced in vital organs (such as the liver, pancreas, and stomach) during laser thermal therapy. The experimental study was conducted during thermal ablation procedures on live pigs.</p> <p><strong>Ethical Approval:</strong> The experiments were performed at the Institute for Image Guided Surgery in Strasbourg, France. This experimental study was approved by the local Ethical Committee on Animal Experimentation (ICOMETH No. 38.2015.01.069) and by the French Ministry of Higher Education and Research (protocol №APAFiS-19543-2019030112087889, approved on March 14, 2019). All animals were treated in accordance with the ARRIVE guidelines, the French legislation on the use and care of animals, and the guidelines of the Council of the European Union (2010/63/EU).</p> <p><strong>Description:</strong> During our experimental study, we used a TIVITA hyperspectral camera to acquire hypercubes of size 640x480x100 voxels, indicating 640x480 pixels for 100 bands, and regular RGB images at each acquisition step. These bands were acquired directly from the hyperspectral camera without additional pre-processing. The hypercube was acquired in approximately 6 seconds and synchronized with the absence of breathing motion using a protocol implemented for animal anesthesia. Polyurethane markers were placed around the target area to serve as references for superimposing the hyperspectral images, which were acquired using target areas selected according to the hyperspectral camera manufacturer's guidelines.</p> <p>As part of our investigation, we included hyperspectral cubes from 20 experiments conducted under identical conditions in our study. The hyperspectral cubes were collected in three distinct stages. In the first stage, the cubes were gathered before laparotomy at a temperature of 37°C. In the second stage, we obtained the cubes as the temperature gradually increased from 60°C to 110°C at 10°C intervals. Finally, in the last stage, the cubes were collected after turning off the laser during the post-ablation phase. Thus, we obtained a total of 233 hyperspectral cubes, each consisting of 100 wavelengths, resulting in a dataset of 23,300 two-dimensional images. The temperature changes were recorded, and the “<em>Temperature profile during laser ablation</em>” image illustrates the corresponding profile, highlighting the specific time intervals during which the hyperspectral camera and laser were activated and deactivated. To provide a visual representation of the collected data, we have included several examples of images captured from different organs in the “<em>Examples of ablation areas</em>” figure.</p> <p>The raw dataset, comprising 233 hyperspectral cubes of 100 wavelengths each, was transformed into 699 single-channel images using PCA and t-SNE decompositions. These images were then divided into training and test subsets and prepared in the COCO object detection format. This COCO dataset can be used for training and testing different neural networks.</p> <p><strong>Access to the Study:</strong> Further information about this study, including curated source code, dataset details, and trained models, can be accessed through the following repositories:</p> <ul> <li><strong>Source code:</strong> <a href="https://github.com/ViacheslavDanilov/hsi_analysis" target="_blank" rel="noopener">https://github.com/ViacheslavDanilov/hsi_analysis</a></li> <li><strong>Dataset:</strong> <a href="https://doi.org/10.5281/zenodo.10444212" target="_blank" rel="noopener">https://doi.org/10.5281/zenodo.10444212</a></li> <li><strong>Models:</strong> <a href="https://doi.org/10.5281/zenodo.10444269" target="_blank" rel="noopener">https://doi.org/10.5281/zenodo.10444269</a></li> </ul>
High-Voltage Disconnector State Identification: synthetic and real images of substation disconnectors
<p>This dataset contains the training and test images used in the work detailed in the article: Barpp Gomes, V., Marchesi, B., Gruber, Y.A. <em>et al.</em> Exploring Synthetic Data for Training Deep Learning Models for High-Voltage Disconnector State Identification. <em>J Control Autom Electr Syst</em> (2025). <a href="https://doi.org/10.1007/s40313-025-01204-2">https://doi.org/10.1007/s40313-025-01204-2</a></p> <p>Contais about 940,000 synthetic (CGI-rendered) and 60,000 real (camera-captured) samples of four types of substation disconnectors, on both open and closed states:</p> <ul> <li>230 kV center break (type 1, as indicated in the article);</li> <li>230 kV center break (type 2);</li> <li>230 kV double side break</li> <li>525 kV horizontal semi-pantograph</li> </ul> <p>Each zip file contains images of one type of substation disconnector. Images are sized 320x128 and are organized in folders, as follows:</p> <ul> <li>00_train_synth: Synthetic training images.</li> <li>01_train_real: A small set of real training images, as indicated in the article.</li> <li>02_test_real_normal1: One set of real test images.</li> <li>03_test_real_normal2: Another set of real test images, from a different time period.</li> <li>04_test_real_maneuvers: A special set of real test images in which the switches have been operated (are in different states).</li> </ul>
Data from: A FAIR and modular image-based workflow for knowledge discovery in the emerging field of imageomics
<p>Data and results from the Imageomics Workflow. These include data files from the Fish-AIR repository (https://fishair.org/) for purposes of reproducibility and outputs from the application-specific imageomics workflow contained in the Minnow_Segmented_Traits repository (https://github.com/hdr-bgnn/Minnow_Segmented_Traits).</p> <p>Fish-AIR:<br> This is the dataset downloaded from Fish-AIR, filtering for Cyprinidae and the Great Lakes Invasive Network (GLIN) from the Illinois Natural History Survey (INHS) dataset. These files contain information about fish images, fish image quality, and path for downloading the images. The data download ARK ID is dtspz368c00q. (2023-04-05). The following files are unaltered from the Fish-AIR download. We use the following files:</p> <p>extendedImageMetadata.csv: A CSV file containing information about each image file. It has the following columns: ARKID, fileNameAsDelivered, format, createDate, metadataDate, size, width, height, license, publisher, ownerInstitutionCode. Column definitions are defined https://fishair.org/vocabulary.html and the persistent column identifiers are in the meta.xml file.</p> <p>imageQualityMetadata.csv: A CSV file containing information about the quality of each image. It has the following columns: ARKID, license, publisher, ownerInstitutionCode, createDate, metadataDate, specimenQuantity, containsScaleBar, containsLabel, accessionNumberValidity, containsBarcode, containsColorBar, nonSpecimenObjects, partsOverlapping, specimenAngle, specimenView, specimenCurved, partsMissing, allPartsVisible, partsFolded, brightness, <br> uniformBackground, onFocus, colorIssue, quality, resourceCreationTechnique. Column definitions are defined https://fishair.org/vocabulary.html and the persistent column identifiers are in the meta.xml file.</p> <p>multimedia.csv: A CSV file containing information about image downloads. It has the following columns: ARKID, parentARKID, accessURI, createDate, modifyDate, fileNameAsDelivered, format, scientificName, genus, family, batchARKID, batchName, license, source, ownerInstitutionCode. Column definitions are defined https://fishair.org/vocabulary.html and the persistent column identifiers are in the meta.xml file.</p> <p>meta.xml: A XML file with the metadata about the column indices and URIs for each file contained in the original downloaded zip file. This file is used in the fish-air.R script to extract the indices for column headers.</p> <p>The outputs from the Minnow_Segmented_Traits workflow are:</p> <p>sampling.df.seg.csv: Table with tallies of the sampling of image data per species during the data cleaning and data analysis. This is used in Table S1 in Balk et al. </p> <p>presence.absence.matrix.csv: The Presence-Absence matrix from segmentation, not cleaned. This is the result of the combined outputs from the presence.json files created by the rule “create_morphological_analysis”. The cleaned version of this matrix is shown as Table S3 in Balk et al.</p> <p>heatmap.avg.blob.png and heatmap.sd.blob.png: Heatmaps of average area of biggest blob per trait (heatmap.avg.blob.png) and standard deviation of area of biggest blob per trait (heatmap.sd.blob.png). These images are also in Figure S3 of Balk et al.</p> <p>minnow.filtered.from.iqm.csv: Filtered fish image data set after filtering (see methods in Balk et al. for filter categories).</p> <p>burress.minnow.sp.filtered.from.iqm.csv: Fish image data set after filtering and selecting species from Burress et al. 2017.</p>
Dataset of 400 pomegranate tree (Punica granatum L. 'Wonderful') images
<p>Dataset of 400 pomegranate tree (Punica granatum L. ‘Wonderful’) images, with the corresponding fruit masks.</p> <p>The dataset is designed for training artificial intelligence models for instance segmentation.</p> <p>The pictures were collected by means of mobile devices (smartphones), in random trees, from different distances, orientations and in varying lighting conditions. The resolution of the images and masks is 640x480 pixels. The dataset is divided into training (70%), validation (15%) and test set (15%). Stratification was performed in 3 periods of the season to ensure that all fruit ripening stages were present in each subset. Masks consist of a very detailed manual annotation of the visible part for each of the fruits in the images.</p>
R-CAUSTIC: Rippling CAUSTICs underwater Image dataset
<p><strong>Description</strong></p><p>Rippling caustics seem to be the main factor degrading the underwater RGB image quality and affecting the image- based 3D reconstruction process in very shallow waters. These effects are adversely affecting image matching algorithms by throwing off most of them, leading to less accurate matches and causing issues in the Simultaneous Localization and Mapping (SLAM) based navigation of the Remotely Operated Vehicles (ROV) and Autonomous Underwater Vehicles (AUV) on shallow waters. Also, they are the main cause for dissimilarities in the generated textures and orthoimages. In order to fill the gap in the literature regading underwater rippling caustics imagery with real ground truth and reference images, the first real-world underwater caustics benchmark dataset which contains 1465 underwater images is presented. Together with the RGB imagery, the corresponding generated ground truth images are delivered for facilitating the training and testing of machine learning and deep learning methods for image classification. R-CAUSTIC dataset also provides the necessary data to evaluate, at least to some extent, the performance of 3D reconstruction approaches. Data were acquired using a GoPro Hero 4 Black action camera with image dimensions of 4000 x 3000 pixels, focal length of 2.77mm and pixel size of 1.55μm and a tripod. Action cameras are widely used for underwater image acquisition. The dataset was captured in near-shore underwater sites at depths varying from 0.5 to 2m. No artificial light sources were used. Due to the wind, the turbulent surface of the water created dynamic rippling caustics on the seabed. In total 1465 RGB images were collected, separated in 7 different datasets; five of them containing stereo images, one of them tri-stereo images and one consists of multi-stereo imagery acquired in 7 different camera poses.</p><p> </p><p><strong>Publication</strong></p><p>The paper is availbale in Open Access here: https://ieeexplore.ieee.org/document/10172291</p><p><strong>If you use this dataset please cite it as R-CAUSTIC</strong> [Reference].<br>[Reference]: <strong>P. Agrafiotis, K. Karantzalos and A. Georgopoulos, "Seafloor-Invariant Caustics Removal From Underwater Imagery," in </strong><i><strong>IEEE Journal of Oceanic Engineering</strong></i><strong>, vol. 48, no. 4, pp. 1300-1321, Oct. 2023, doi: 10.1109/JOE.2023.3277168.</strong></p><p>BibTeX:</p><p>@ARTICLE{10172291, author={Agrafiotis, Panagiotis and Karantzalos, Konstantinos and Georgopoulos, Andreas}, journal={IEEE Journal of Oceanic Engineering}, title={Seafloor-Invariant Caustics Removal From Underwater Imagery}, year={2023}, volume={48}, number={4}, pages={1300-1321}, doi={10.1109/JOE.2023.3277168}}</p><p> </p>
High-resolution images from a low-cost imaging device for hyphae in soil
<p>This dataset contains high-resolution images produced by a low-cost imaging device for hyphae in soil called <em>Hyphascope</em>. Using a digital microscope camera (DMC; 600× magnification),<em> </em>the device takes detailed images (0.83 × 0.62 mm imaged area) of a soil profile from evenly spaced camera positions within a user-defined volume. Repeated imaging of a soil profile with <em>Hyphascope</em> enables researchers to observe and quantify changes in the amount, distribution, and morphology of hyphae.</p> <p>Individual images were combined using the <em>Grid/Collection stitching</em> plugin of the <em>Fiji</em> distribution of <em>imageJ</em> (Preibisch et al. 2009). All images are supplied in the JPG format to limit their file size. For more details on the assembly and application of <em>Hyphascope</em>, see <a href="https://doi.org/10.17504/protocols.io.bp2l6xo3zlqe/v1">this protocol on protocols.io</a>. For information on the development, limitations, and expected outcomes of the protocol, see <a href="https://doi.org/10.1371/journal.pone.0318083">this article</a> published in PLOS ONE. </p> <p> </p> <div> <h2>Image set 1: 10 × 10 mm soil profile area at 20 - 30 mm soil depth</h2> <p>Imaged at 0.65 μm px<sup>-1</sup> (39200 dpi)* in a <em>Quercus serrata</em> grove on 2023/05/25 during a period of high hyphal density in the soil.</p> <h3>Individual images (18 rows × 14 images each)</h3> <ul> <li> <p><em>set1_foc00000.zip</em> (focus depth 0 mm)</p> </li> </ul> <h3>Combined images</h3> <ul> <li> <p><em>set1_foc00000_combined.jpg</em> (focus depth 0 mm)</p> </li> </ul> <h2>Image set 2: 5 × 5 mm soil profile area at 100 - 105 mm soil depth</h2> <p>Imaged at 0.52 μm px<sup>-1</sup> (49000 dpi) in a <em>Quercus serrata</em> grove on 2023/09/25.</p> <h3>Individual images (9 rows × 7 images each)</h3> <ul> <li> <p><em>set1_foc00000.zip</em> (focus depth 0 mm)<em><br></em></p> </li> <li> <p><em>set1_foc00025zip</em> (focus depth 0.025 mm)</p> </li> <li><em>set1_foc00050.zip</em> (focus depth 0.05 mm)</li> </ul> <h3>Combined images</h3> <ul> <li> <p><em>set1_foc00000_combined.jpg</em> (focus depth 0 mm)</p> </li> <li> <p><em>set1_foc00025_combined.jpg</em> (focus depth 0.025 mm)</p> </li> <li><em>set1_foc00050_combined.jpg</em> (focus depth 0.05 mm)</li> </ul> <h2>Image set 3: 5 × 5 mm soil profile area at soil surface level</h2> <p>Imaged at 0.52 μm px<sup>-1</sup> (49000 dpi) in a <em>Quercus serrata</em> grove on 2023/10/14 during a rain event.</p> <h3>Individual images (9 rows × 7 images each)</h3> <ul> <li> <p><em>set2_foc00000.zip</em> (focus depth 0 mm)<em><br></em></p> </li> <li> <p><em>set2_foc00025zip</em> (focus depth 0.025 mm)</p> </li> <li><em>set2_foc00050.zip</em> (focus depth 0.05 mm)</li> </ul> <h3>Combined images</h3> <ul> <li> <p><em>set2_foc00000_combined.jpg</em> (focus depth 0 mm)</p> </li> <li> <p><em>set2_foc00025_combined.jpg</em> (focus depth 0.025 mm)</p> </li> <li><em>set2_foc00050_combined.jpg</em> (focus depth 0.05 mm)</li> </ul> <p> </p> <p> </p> <p><em>*Units of imaging resolution: </em></p> <ol> <li><em>pixel width (μm px-1), i.e. the horizontal or vertical distance on the imaged surface covered by a single pixel; </em></li> <li><em>dots per inch (dpi), i.e. the number of pixels along a horizontal or vertical distance of 25.4 mm on the imaged surface.</em></li> </ol> </div>
Dataset of pomegranate tree (Punica granatum L. 'Wonderful') image times series
<p>Dataset of pomegranate tree (Punica granatum L. ‘Wonderful’) image times series. The pictures were collected by means of Raspberry Pi cameras with OV5647 sensor (5 MP, f2.9). Sensors were installed on fixed platforms for continuous measurement with zenithal orientation at a distance of approximately 1 metre from the canopy. Images were captured daily at 9 a.m. (GMT+2) from July to mid-October in 2021 and 2022. The resolution of the images is 640x480 pixels.</p>
Micro-CT images of deep brain stimulation leads
<p>The dataset contain micro-CT images of leads used in deep brain stimulation. A lead comprises multiple electrodes and enables the delivery of electrical pulses to the brain to treat medical conditions such as Parkinson's disease, essential tremor or epilepsy. Images were acquired with a Skyscan 1276 micro-CT system from Bruker. Each image is provided in Nifti format (.nii) along with its corresponding log file (.log) generated by the scanner. The file names indicate the manufacturer and sample model. 'BS' denotes Boston Scientific.<br><br>Images can be visualized at:<br>https://activgroup.github.io/DBS-lead-microCT/<br><br>To contribute, please contact thomas.billoud@uniklinik-freiburg.de</p>
Alpha-Galactosaminidase family GH114 protein from Fusarium solani: X-ray diffraction images
<p>This submission includes h5-files with diffraction images recorded using the Dectris EIGER X 16M detector at the DIAMOND beamline I04. The model of the crystal structure and associated information can be found in the Protein Data Bank entry 9EP6. The model has P 31 2 1 symmetry and three molecules per asymmetric unit. This is a case of crystal pathology – partial disorder. There is electron density for the fourth molecule which could be modelled with occupancy 1/2 and would overlap with a symmetry-related molecule.</p>
Videos of the processed microscope images and time series of the petrophysical parameters from image processing and geochemical simulation and of the measured induced polarisation [Video][Dataset]
<p>Supporting Information for the manuscript <em>Microfluidics and spectral induced polarization for direct observation and petrophysical modeling of calcite dissolution</em> published in Geophysical Research Letters</p> <ul> <li><strong>Data Set S1.</strong> Porosity, water saturation, and calcite sample perimeter from image<br>processing.</li> <li><strong>Data Set S2.</strong> Porosity, water conductivity, and pH from geochemical simulation.</li> <li><strong>Data Set S3.</strong> Real and imaginary components of the complex electrical conductivity at<br>2.5 Hz and CEC from petrophysical modeling.</li> <li><strong>Movie S1.</strong> Dissolution of the calcite sample with the detected contour superimposed in<br>white on the grayscale images. Time, length scale, and flow direction are indicated. In<br>case of problems launching the file, we recommend using VLC Media Player software.</li> <li><strong>Movie S2.</strong> Segmented images of the CO2 bubbles produced by the calcite dissolution.<br>Time, length scale, and flow direction are indicated. In case of problems launching the<br>file, we recommend using VLC Media Player software.</li> </ul>
Training and test dataset of STED images of microtubules in fixed cells
<p>Training and test dataset of microtubule used in the manuscript "Denoising diffusion models for high-resolution microscopy image restoration".</p>
Live-cell STED dataset of mitochondria containing ground truth and corresponding low intensity noisy images
<p>The dataset was acquired as part of the manuscript "Denoising diffusion models for high-resolution microscopy image restoration". The dataset contains ground truth and low intensity STED images of mitochondria acquired in live U2-OS cells stably expressing TOM20 coupled to the dead mutant of HaloTag7 which was made fluorescent by using the exchangeable ligand Hy4 bound to the fluorophore SiR. </p>
A dataset of colorectal cancer histopathological images
<p>The dataset contains the histopathological images of the ColoPola dataset (https://doi.org/10.5281/zenodo.10068018).</p> <p>CLCXYYZZNN_Hx</p> <p>CLC: colorectal (cancer) tissue</p> <p>NLC: normal tissue</p> <p>X - Times<br>YY - Sample number<br>ZZ - Serial number<br>NN - Image number<br>H - Magnification</p>
Raw planetary images and boulder labels data (as shapefiles) collected during the BOULDERING Marie Skłodowska-Curie Global fellowship
<p>This database contains 64 large images of craters on the lunar and martian surfaces and 3 images of boulder fields on Earth (see manuscript <a href="https://agupubs.onlinelibrary.wiley.com/doi/full/10.1029/2023JE008013">https://agupubs.onlinelibrary.wiley.com/doi/full/10.1029/2023JE008013</a> for more information on those terrestrial locations). The data was collected during the BOULDERING Marie Skłodowska-Curie Global fellowship between October 2021 and 2024.</p> <p>For each image, the boulder outlines within specific tiles within the image were carefully mapped in QGIS. More information about the labelling procedure can be found in the following manuscript (<a href="https://agupubs.onlinelibrary.wiley.com/doi/full/10.1029/2023JE008013">https://agupubs.onlinelibrary.wiley.com/doi/full/10.1029/2023JE008013</a>). This dataset differs from the previous dataset included along with the manuscript <a href="https://zenodo.org/records/8171052">https://zenodo.org/records/8171052</a>, as it contains more mapped images, especially of boulder populations around young impact structures on the Moon (cold spots). </p> <p>For each location, you will find a raster with a .tif format, and three shapefiles:</p> <ul> <li> <p>a boulder-mapping file, which is the manually digitized outline of boulders.</p> </li> <li> <p>a tiles-completely-mapped file, which depicts the patches/tiles/windows on which the boulder mapping has been conducted.</p> </li> <li> <p>a global-tiles file, which shows all of the image patches/tiles/windows (pick the term you are the most familiar with) within a raster.</p> </li> </ul> <p>In addition you will find .pkl (which stands for pickle), which contains some information about the patches/tiles/windows if you would need to clip those windows out from the original raster. You can find more information in the way we process this raw data into a format which can be ingested in a deep learning model (see <a href="https://zenodo.org/records/14250874" target="_blank" rel="noopener">https://zenodo.org/records/14250874</a>) in the two following github repositories (<a href="https://github.com/astroNils/YOLOv8-BeyondEarth" target="_blank" rel="noopener">https://github.com/astroNils/YOLOv8-BeyondEarth</a> and <a href="https://github.com/astroNils/MLtools/tree/main" target="_blank" rel="noopener">https://github.com/astroNils/MLtools</a>). If you don't plan in adding more training data, you can directly used the pre-processed database (see <a href="https://zenodo.org/records/14250874" target="_blank" rel="noopener">https://zenodo.org/records/14250874</a>).</p> <p>There are multiple locations/images per planetary body. Cold spots are located on the Moon, but they are saved in a folder of their own. </p> <p>Note that the cold spots boulder mapping shapefiles are partially manually mapped, and partially originating from predictions made from a deep learning model (which explains the outline of boulders are predicted within one pixel).</p> <p><strong>How to cite:</strong></p> <p>Please refer to the "how to cite" section of the readme file of <a href="https://github.com/astroNils/YOLOv8-BeyondEarth" target="_blank" rel="noopener">https://github.com/astroNils/YOLOv8-BeyondEarth.</a></p> <p><strong>Structure:</strong></p> <pre><code>. └── raw_data/ ├── coldspots/ │ └── image_name/ │ ├── shp/ │ │ ├── <image_name>-tiles-completely-mapped.shp │ │ ├── <image_name>-boulder-mapping.shp │ │ └── <image_name>-global-tiles.shp │ └── raster/ │ └── <image_name>.tif ├── earth/ │ └── image_name/ │ ├── shp/ │ │ ├── <image_name>-tiles-completely-mapped.shp │ │ ├── <image_name>-boulder-mapping.shp │ │ └── <image_name>-global-tiles.shp │ └── raster/ │ └── <image_name>.tif ├── mars/ │ └── image_name/ │ ├── shp/ │ │ ├── <image_name>-tiles-completely-mapped.shp │ │ ├── <image_name>-boulder-mapping.shp │ │ └── <image_name>-global-tiles.shp │ └── raster/ │ └── <image_name>.tif └── moon/ └── image_name/ ├── shp/ │ │ ├── <image_name>-tiles-completely-mapped.shp │ │ ├── <image_name>-boulder-mapping.shp │ │ └── <image_name>-global-tiles.shp └── raster/ └── <image_name>.tif</code></pre>
Labeled Images at OBSEA for Object Detection Algorithms
<p>Images from OBSEA underwater cameras labeled with marine species to train AI-based Object Detection algorithms.</p>
Training Images for "ImmuNet" Convolutional Neural Network
<p>This dataset contains all annotations and images for training the machine learning architecture presented in this manscript:</p> <p>Shabaz Sultan, Mark A. J. Gorris, Lieke L. van der Woude, Franka Buytenhuijs, Evgenia Martynova, Sandra van Wilpe, Kiek Verrijp, Carl G. Figdor, I. Jolanda M. de Vries, Johannes Textor:<br>ImmuNet: a segmentation-free machine learning pipeline for immune landscape phenotyping in tumors by multiplex imaging.<br>Biology Methods and Protocols 10(1), bpae094, 2025. doi: 10.1093/biomethods/bpae094</p> <p>The .tar.gz file contains several multichannel images stored as TIFF files, and arranged in a folder structure that is convenient for matching the files to the annotations provided in the .json.gz file. We also provide an .h5 file that contains the final trained network that was used to generate the figures in this manuscript.</p> <p>Further information on the data can be found in the manuscript cited above. Instructions on how to use the annotations and the code can be found on our GitHub page at: https://github.com/jtextor/immunet</p>
LigPCDS: Labeled Dataset of X-ray Protein Ligand Images in 3D Point Cloud and Validated Deep Learning Models
<p>The difference electron density from X-ray protein crystallography was used to create the first dataset of labeled ligand images in 3D point clouds, named <strong>LigPCDS</strong>. The dataset contain 244,226 entries of free organic ligands containing 3D representations labeled with two major labeling approaches: SP-based and AtomSymbol-based.</p> <p> </p> <p>The data from free organic molecules (non-covalent ligands) was retrieved from the Research Collaboratory for Structural Bioinformatics Protein Data Bank (RCSB PDB) in december 2019 with resolutions ranging from 1.5 to 2.2 Å. The ligand images (blobs) were interpolated from their calculated difference electron density map in a 3D grid-like bounding box, around their atomic positions, and stored in point clouds. These ligand grid representations were further processed to retrive the final ligands representation in 3D point clouds using a mask of the shape of the ligand. A grid spacing of 0.5 Å gave the best results. The density value of the grid points was used as feature. The labeling approach used the structure of the ligands to propose vocabularies of chemical classes based on the chemical atoms themselves and their cyclic substructures. These structure annotations were applied pointwise to the ligand 3D representations using an atomic sphere model. Four proposed vocabularies were validated by successfully training good performance deep learning models for the semantic segmentation of a stratified dataset from LigPCDS, using 78902 entries.</p> <p>The four validated deep learning models are: (i) the LigandRegion, composed by generic atoms of any type; (ii) the AtomCycle, composed by generic atoms outside cycles and generic cycles; (iii) the AtomC347CA56, composed by generic atoms outside cycles, not aromatic cycles of size 3 to 7 and aromatic cycles of size 5 and 6; and (iv) the AtomSymbolGroups, composed by the atoms symbols with groupings. The mean accuracy of these models in their cross-validation was between 49.7% <span lang="EN-GB">[-19.4,20.</span><span lang="EN-GB">2]</span> and 77.4% <span lang="EN-GB">[-11.7,12.1]</span> in terms of Intersection over Union (mIoU) metric and between 62.4% <span lang="EN-GB">[-18.8,19.</span><span lang="EN-GB">7]</span> and 87.0% <span lang="EN-GB">[-8.4,8.8]</span> in F1-score (mF1), confidence interval between squared brackets. The models i, ii and iii and the used labeled representations in 3D point cloud are contained in the SP-based record; and model iv and its used labeled representations are contained in the AtomSymbol-based record.</p> <p>The dataset and validated models may be used to tackle problems regarding known and unknown ligand building to drug discovery and fragment screening pipelines. </p> <p>The code used to create and validated the LigPCDS is available at the following repository: https://github.com/danielatrivella/np3_ligand</p> <p>This repository also contains the NP³ Blob Label application for ligand building using the validated deep learning models from LigPCDS.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.