Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

979

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

979 results for “image dataset”

Learn how ShareScore rates datasets ↗
zenodo40/100

CytoNuke Dataset: Towards reliable whole-cell segmentation in bright-field histological images

<p>This is the dataset from the preprint "Cyto R-CNN and CytoNuke Dataset: Towards reliable whole-cell segmentation in bright-field histological images" by Raufeisen et al. (2024). It contains 6,683 annotations (3,991 nuclei and 2,607 whole cells) of head and neck squamous cell carcinoma cells in hematoxylin and eosin stained histological images. The annotations are in COCO format and distributed over 83 PNG images. Cyto R-CNN was trained on this dataset and compared with other state-of-the-art methods. The CytoNuke dataset is released under the CC BY 4.0 license.</p> <p>The histological images are from the CPTAC dataset:<br>National Cancer Institute Clinical Proteomic Tumor Analysis Consortium (CPTAC). (2018). The Clinical Proteomic Tumor Analysis Consortium Head and Neck Squamous Cell Carcinoma Collection (CPTAC-HNSCC) (Version 15) [Data set]. The Cancer Imaging Archive. https://doi.org/10.7937/K9/TCIA.2018.UW45NH81</p> <p>Funding: Behrus Puladi was funded by the Medical Faculty of RWTH Aachen University as part of the Clinician Scientist Program. We acknowledge FWF enFaced 2.0 [KLI 1044, https://enfaced2.ikim.nrw/] and KITE (Plattform f&uuml;r KI-Translation Essen) from the REACT-EU initiative [https://kite.ikim.nrw/, EFRE-0801977]. Fabian H&ouml;rst, Jianning Li, Jens Kleesiek and Jan Egger received funding from the Cancer Research Center Cologne Essen (CCCE).</p>

opencc-by-4.0Jan 2024View details →
zenodo40/100

BioTISR: a time-lapse biological image dataset for super-resolution microscopy

<p>BioTISR is a biological image dataset for super-resolution microscopy, currently including 2D and 3D time-lapse image pairs of low-and-high resolution images of a variety of biology structures, aiming to provide a high-quality dataset of time-lapse biological SR images for the community to spark more developments of computational SR methods.</p> <p>At present, 2D dataset includes five specimens (clathrin-coated pits, lysosomes, outer mitochondrial membrane, microtubules, and F-actin) acquired with the GI/TIRF-SIM mode and nonlinear SIM mode of our Multi-SIM system, and 3D dataset includes three specimens (outer mitochondrial membrane, microtubules, and F-actin) acquired with 3D-SIM mode of the Multi-SIM system. For each type of specimen and each imaging modality, we acquired the raw data from at least 50 distinct regions-of-interest (ROI). For each ROI, we acquired two (3D data) or three (2D data) groups of N-phase &times; M-orientation &times; T-timepoint raw images with a constant exposure time but increasing the excitation light intensity, where (N, M, T) are (3, 3, 20) for TIRF-SIM and GI-SIM, (5, 5, 10) for nonlinear SIM, and (3, 5, 10) for 3D-SIM. Specific imaging conditions and scripts for reading MRC file are provided in Supplement Files.</p> <p>The BioTISR dataset is related to the following paper:<a href="https://www.nature.com/articles/s41587-025-02553-8#citeas">Qiao, C., Liu, S., Wang, Y.&nbsp;<em>et al.</em>&nbsp;A neural network for long-term super-resolution imaging of live cells with reliable confidence quantification.&nbsp;<em>Nat Biotechnol</em> (2025). https://doi.org/10.1038/s41587-025-02553-8</a>, which is an extension of our previously published <a href="https://doi.org/10.6084/m9.figshare.13264793.v9">BioSR dataset</a> (https://www.nature.com/articles/s41592-020-01048-5).</p> <p>Limited by quota, the original images uploaded in the current 3D dataset are wide-field images obtained after averaging 15 images, where (N, M, T) are (1, 1, 10). We will update them to raw SIM images after the quota is expanded.</p> <p>2D dataset's url:</p> <p><a href="https://doi.org/10.5281/zenodo.13843670" target="_blank" rel="noopener">https://doi.org/10.5281/zenodo.13843670</a></p> <p>3D dataset's urls:</p> <p>F-actin:</p> <p>WF input: <a href="https://doi.org/10.5281/zenodo.13843673">https://doi.org/10.5281/zenodo.13843673</a></p> <p>Raw SIM input:<a href="https://doi.org/10.5281/zenodo.13994464" target="_blank" rel="noopener">https://doi.org/10.5281/zenodo.13994464</a></p> <p>Microtubules:</p> <p>WF input:&nbsp;<a href="https://doi.org/10.5281/zenodo.13932988" target="_blank" rel="noopener">https://doi.org/10.5281/zenodo.13932988</a></p> <p>Raw SIM input: <a href="https://doi.org/10.5281/zenodo.13989327" target="_blank" rel="noopener">https://doi.org/10.5281/zenodo.13989327</a></p> <p>Mitochondria:</p> <p>WF input: <a href="https://doi.org/10.5281/zenodo.13843183" target="_blank" rel="noopener">https://doi.org/10.5281/zenodo.13843183</a></p> <p>Raw SIM input: <a href="https://doi.org/10.5281/zenodo.14000502" target="_blank" rel="noopener">https://doi.org/10.5281/zenodo.14000502</a></p> <p>&nbsp;</p> <p>Update 2025.5.6</p> <p>Add optical transfer function(OTF) of the microscopy system and the pixel size of each data to the supplementary files.</p>

opencc-by-4.0Oct 2024View details →
zenodo40/100

CloudTracks: A Dataset for Localizing Ship Tracks in Satellite Images of Clouds

<p>The CloudTracks dataset consists of 1,780 MODIS satellite images hand-labeled for the presence of more than 12,000 ship tracks. More information about how the dataset was constructed may be found at&nbsp;<a href="http://github.com/stanfordmlgroup/CloudTracks">github.com/stanfordmlgroup/CloudTracks</a>. The file structure of the dataset is as follows:</p><p>CloudTracks/<br>&nbsp; &nbsp; full/<br>&nbsp; &nbsp; &nbsp; &nbsp;images/<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; (sample image name) mod2002121.1920D.png<br>&nbsp; &nbsp; &nbsp; &nbsp;jsons/<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; (sample json name) mod2002121.1920D.json</p><p>The naming convention is as follows:<br>mod2002121.1920D: the first 3 letters specify which of the sensors on the two MODIS satellites captured the image, mod for Terra and myd for Aqua. This is followed by a 4 digit year (2002) and a 3 digit day of the year (121). The following 4 digits specify the time of day (1920; 24 hour format in the UTC timezone), followed by D or N for Day or Night.</p><p>The 1,780 MODIS Terra and Aqua images were collected between 2002 and 2021 inclusive over various stratocumulus cloud regions (such as the East Pacific and East Atlantic) where ship tracks have commonly been observed. Each image has dimension 1354 x 2030 and a spatial resolution of 1km. Of the 36 bands collected by the instruments, we selected channels 1, 20, and 32 to capture useful physical properties of cloud formations.</p><p>The labels are found in the corresponding JSON files for each image. The following keys in the json are particularly important:</p><p>imagePath: the filename of the image.<br>shapes: the list of annotations corresponding to the image, where each element of the list is a dictionary corresponding to a single instance annotation. The dictionary has a key with value "shiptrack" or "uncertain" which is the label of the annotation and the corresponding value is a linestrip detailing the ship track path.</p><p>Further pre-processing details may be found at the GitHub link above. If you have any questions about the dataset, contact us at:<br><a href="mailto:mahmedch@stanford.edu">mahmedch@stanford.edu</a>,&nbsp;<a href="mailto:lynakim@stanford.edu">lynakim@stanford.edu</a>,&nbsp;<a href="mailto:jirvin16@cs.stanford.edu">jirvin16@cs.stanford.edu</a></p>

opencc-by-4.0Oct 2023View details →
zenodo40/100

Dataset with segmentations of 117 important anatomical structures in 1228 CT images

<p>Info: This is version 2 of the TotalSegmentator dataset.<br><br>In 1228&nbsp;CT images we segmented 117&nbsp;anatomical structures&nbsp;covering a majority of relevant classes for most use cases.&nbsp;The CT images were randomly sampled from clinical routine, thus representing a real world dataset which generalizes to clinical application. The dataset contains a wide range of different pathologies, scanners, sequences and institutions.</p><p>Link to a copy of this dataset on Dropbox for much quicker download: <a href="https://www.dropbox.com/scl/fi/oq0fsz8oauory204g8o6f/Totalsegmentator_dataset_v201.zip?rlkey=afnl2ixhqca2ukkf1v9p6jz7p&amp;dl=0">Dropbox Link</a></p><p>Overview of differences to v1 of this dataset: <a href="https://github.com/wasserth/TotalSegmentator/blob/master/resources/improvements_in_v2.md">here</a></p><p>A small subset of this dataset with only 102&nbsp;subjects for quick download+exploration can be found here: <a href="https://doi.org/10.5281/zenodo.8367169">here</a></p><p>You can find a segmentation model trained on this dataset <a href="https://github.com/wasserth/TotalSegmentator">here</a>.<br><br>More details about the dataset can be found in the corresponding <a href="https://doi.org/10.1148/ryai.230024">paper</a>&nbsp;(the paper describes v1 of the dataset). Please cite this paper if you use the dataset.</p><p>This dataset was created by the department of <a href="https://www.unispital-basel.ch/en/radiologie-nuklearmedizin/forschung-radiologie-nuklearmedizin">Research and Analysis at University Hospital Basel</a>.</p><p><strong>UPDATE</strong>: On 2023-10-27 we uploaded version 2.0.1 which fixes broken files.</p>

opencc-by-4.0Sep 2023View details →
zenodo40/100

Dataset of processed Sentinel-2 images for chlorophyll-a estimation in high-altitude lakes in the Sierra Nevada, Spain

<p>This dataset contains Sentinel 2 satellite images clipped to 5 high-altitude lakes in the Sierra Nevada Mountain Range, Spain. The images were processed with the following atmospheric correction algorithms:</p><ul><li><a href="https://c2rcc.org/">C2RCC</a> (<a href="https://ui.adsabs.harvard.edu/abs/2016ESASP.740E..54B/abstract">Brockmann et al. 2016</a>)</li><li><a href="https://github.com/MarcYin/SIAC">SIAC</a> (<a href=" https://doi.org/10.5194/gmd-15-7933-2022">Yin et al. 2022)</a></li><li><a href="https://github.com/acolite/acolite/releases/tag/20221114.0">ACOLITE</a> (<a href="https://doi.org/10.1016/j.rse.2018.07.015">Vanhellemont &amp; Ruddick, 2018</a>)</li><li><a href="https://grass.osgeo.org/grass83/manuals/i.atcorr.html">6SV</a> (<a href="https://doi.org/10.1109/36.581987">Vermote et al. 2006</a>)</li></ul><p><strong>Included Lakes and and their IDs:</strong></p><ul><li>Laguna de la Caldera (ID = P-2)</li><li>Laguna-embalse de las Yeguas (ID = D-6)</li><li>Laguna de Río Seco (ID = P-8)</li><li>Laguna Larga (ID = G-7)</li><li>Laguna de la Mosca (ID = G-11)</li></ul>

opencc-by-4.0Oct 2023View details →
zenodo40/100

An urban traffic dataset composed of visible images and their semantic segmentation generated by the CARLA simulator

<p><strong>If you use this dataset please cite this paper: Rosende, S.B.; Gavil&aacute;n, D.S.J.; Fern&aacute;ndez-Andr&eacute;s, J.; S&aacute;nchez-Soriano, J. An Urban Traffic Dataset Composed of Visible Images and Their Semantic Segmentation Generated by the CARLA Simulator.&nbsp;<em>Data</em>&nbsp;2024,&nbsp;<em>9</em>, 4. <a href="https://doi.org/10.3390/data9010004">https://doi.org/10.3390/data9010004</a></strong></p> <p>A dataset of aerial urban traffic images and their semantic segmentation is presented to be used to train computer vision algorithms, among which those based on convolutional neural networks stand out. The images have been generated using the CARLA simulator (but would be like those that could be obtained with fixed aerial cameras or by using AUVs) in the field of intelligent transportation management. The presented dataset is available and accessible to improve the performance of vision and road traffic management systems, especially for the detection of incorrect or dangerous maneuvers.</p>

opencc-by-4.0Oct 2023View details →
zenodo40/100

Dataset related to article "Phantom‑based analysis of variations in automatic exposure control across three mammography systems: implications for radiation dose and image quality in mammography, DBT, and CEM"

<p>The dataset comprises &nbsp;information from several DICOM tags extracted from digital mammography (DM), digital breast tomosynthesis (DBT), and contrast-enhanced mammography (CEM) images acquired in a phantom study aimed at characterizing the automatic exposure control (AEC) behavior of diverse mammography equipment. The final ten columns of the datasets encompass signal (mena pixel values, MPV) and noise (standard deviation, SD) measurements derived from phantom images. These measurements are used to compute several image quality metrics, including contrast, signal-to-noise ratio (SNR), contrast-to-noise ratio (CNR), CNR relative difference in comparison to the 45 mm reference thickness, and a figure of merit (FOM) obtained by diving the squared CNR by the mean glandular dose (MGD).</p>

opencc-by-4.0Apr 2024View details →
zenodo40/100

SPIDER - Lumbar spine segmentation in MR images: a dataset and a public benchmark

<p>This is a large publicly available multi-center lumbar spine magnetic resonance imaging (MRI) dataset with reference segmentations of vertebrae, intervertebral discs (IVDs), and spinal canal. The dataset&nbsp;includes 447&nbsp;sagittal T1 and T2 MRI series from 218&nbsp;studies of 218 patients with a history of low back pain. The data was collected from four different hospitals. There is an additional&nbsp;hidden test set, not available here, used in the accompanying SPIDER challenge on spider.grand-challenge.org. We share this data&nbsp;to encourage wider participation and collaboration in the field of spine segmentation, and ultimately improve the diagnostic value of lumbar spine MRI.</p> <p>Which MRI studies are assigned to the training and validation sets can be found in the overview file. This file also provides the biological sex for all patients and the age for the patients for which this was available. It also includes a number of scanner and acquisition parameters for each individual MRI study. The dataset also comes with radiological gradings found in a separate file for the following degenerative changes:</p> <p>1.&ensp;&ensp;&ensp;&ensp;Modic changes (type I, II or III)</p> <p>2.&ensp;&ensp;&ensp;&ensp;Upper and lower endplate changes / Schmorl nodes (binary)</p> <p>3.&ensp;&ensp;&ensp;&ensp;Spondylolisthesis (binary)</p> <p>4.&ensp;&ensp;&ensp;&ensp;Disc herniation (binary)</p> <p>5.&ensp;&ensp;&ensp;&ensp;Disc narrowing (binary)</p> <p>6.&ensp;&ensp;&ensp;&ensp;Disc bulging (binary)</p> <p>7.&ensp;&ensp;&ensp;&ensp;Pfirrman grade (grade 1 to 5).&nbsp;</p> <p>All radiological gradings are provided per IVD level.</p> <div>This dataset, and the associated public benchmark, are described in this paper: <a href="https://www.nature.com/articles/s41597-024-03090-w" target="_blank" rel="noopener">https://www.nature.com/articles/s41597-024-03090-w</a></div> <div>The public segmenation challenge can be found here: <a href="https://spider.grand-challenge.org/" target="_blank" rel="noopener">https://spider.grand-challenge.org/</a></div> <div>&nbsp;</div> <div>When using this dataset, please cite this dataset with the correct DOI, and also cite the afformentioned paper.</div>

opencc-by-4.0Nov 2021View details →
zenodo40/100

Organ-on-a-Chip (OOC) Image Dataset

<p><strong>Overview: </strong>This dataset contains 3000+ images generated from OOC (organ-on-a-chip) setup with different cell types. The images were generated by an automated brightfield microscopy setup; for each image, such parameters as cell type, time after seeding, and class label ('good' or 'bad' sample quality as assessed by a biology expert) are provided. Furthermore, for some images, seeding density and flow rate are given as well. The dataset can be used for training machine learning classifiers for the automated analysis of the data generated with OOC setup, allowing to create more reliable tissue models and automate decision making processes for growing OOC.</p><p>The dataset comprises images of OOC samples from the following cell lines:</p><ul><li>A549 (human lung adenocarcinoma alveolar basal epithelial cells, CCL-185, ATTC)</li><li>Caco-2 (colorectal adenocarcinoma epithelial cells, HTB-37, ATCC)</li><li>HPMEC (human pulmonary microvascular endothelial cells; 3000, ScienCell)</li><li>HUVEC (human umbilical vein endothelial cells, CRL-1730, ATCC)</li><li>NHBE (normal human bronchial epithelial cells, CC-2541, Lonza)</li><li>HSAEC (human small airway epithelial cells, PCS-301-010, ATCC)</li></ul><p><strong>Structure of the dataset:</strong> The dataset is split into three main folders that correspond to the data split for training machine learning models, i.e., 'train', 'val', and 'test'. The train/val/test split is done proportionally with respect to the class labels, cell lines, and time after seeding (see below), yet the data can be split or merged in other ways to suit the needs of prospective users of the dataset. Within each of the main folders, there are a 'bad' and a 'good' folder with the images corresponding to the respective class labels (see 'Overview' above). The images in 'bad' / 'ģood' folders are further subdivided into folders corresponding to respective cell lines, which are in their turn subdivided into folders corresponding to the different times after seeding. Therefore, it is easy to find images of interest, e.g., '4+ days' 'good' images of the cell line A549 from the 'train' dataset. Further information about the images is available in the file 'OOC_datasheet.xlsx'.&nbsp;</p><p><strong>Acknowledgement:</strong> The work presented in this paper was supported by the project 'AI-improved organ on chip cultivation for personalised medicine (AimOOC)' (contract with Central Finance and Contracting Agency of Republic of Latvia no. 1.1.1.1/21/A/079; the project is co-financed by REACT-EU funding for mitigating the consequences of the pandemic crisis).</p>

opencc-by-4.0Nov 2023View details →
zenodo40/100

CryoVirusDB: An Expert Labelled Cryo-EM Image Dataset for AI-Driven Virus Particle recognition and Extraction

<p><span>With the advancements in instrumentation, image processing algorithms, and computational capabilities, single-particle electron cryo-microscopy (cryo-EM) has achieved nearly atomic resolutions in the 3D reconstruction of viruses. These detailed structures play a crucial role in comprehending the biological functions and advancing the development of more precise vaccines and antiviral treatments. Despite the effectiveness of deep learning in analyzing microscopic images, its potential in identifying and extracting virus particles from cryo-EM micrographs has been hindered by the limited availability of diverse and high-quality datasets. In this study, we introduce 'CryoVirusDB,' a labeled dataset containing coordinates of accurately selected virus particles in cryo-EM micrographs. CryoVirusDB comprises 9,941 micrographs featuring 9 different viruses along with the coordinates of 0.2 million virus particles in total. We anticipate that CryoVirusDB will enhance the capabilities of deep learning in accurately identifying virus particles in cryo-EM micrographs, thereby facilitating the subsequent 2D-3D reconstruction process.</span></p> <p><span>Instructions to download and use dataset: https://github.com/BioinfoMachineLearning/CryoVirusDB</span></p>

opencc-by-4.0Dec 2023View details →
zenodo40/100

RafanoSet: Dataset of raw, manual and automatically annotated Raphanus Raphanistrum weed images for object detection and segmentation in Heterogenous Agriculture Environment

<p>This dataset is a collection of raw and annotated Multispectral (MS) images acquired in a heterogenous agricultural environment with MicaSense RedEdge-M camera. The spectra particularly&nbsp;Green,&nbsp;Blue,&nbsp;Red,&nbsp;Red Edge and Near Infrared (NIR) were acquired at sub-metre level..&nbsp;<br><br>The MS images were labelled manually using VIA and automatically using Grounding DINO in combination with Segment Anything Model. The segmentation masks obtained using these two annotation techniqes over as well as the source code to perform necessary image processing operations are provided in the repository. The images are focussed over Horseradish (Raphanus Raphanistrum) infestations in Triticum Aestivum (wheat) crops.</p> <p>The nomenclature of sequecncing and naming images and annotations has been in this format: IMG_&lt;scene number&gt;_&lt;spectral channel number&gt;<br><strong>_1</strong>: Blue<br><strong>_2</strong>: Green<br><strong>_3</strong>: Red<br><strong>_4</strong>: Near Infrared<br><strong>_5</strong>: RedEdge<br><br>Example: An image name&nbsp; <strong>IMG_0200_3 </strong>represents the scene number<strong> 200</strong> in <strong>Red channel</strong></p> <p>This dataset 'RafanoSet'is categorized in 6 directories namely 'Raw Images', 'Manual Annotations', 'Automated Annotations', 'Binary Masks - Manual', 'Binary Masks - Automated' and 'Codes'. The sub-directory 'Raw Images' consists of manually acquired 85 images in .PNG format. over 17 different scenes. The sub-directory 'Manual Annotations' consists of annotation file 'region_data' in COCO segmentation format. The sub-directory 'Automated Annotations' consists of 80 automatically annotated images in .JPG format and 80 .XML files in Pascal VOC annotation format.</p> <p>The scientific framework of image acquisition and annotations are explained in the Data in Brief paper which is the course of peer review. This is just a prerequisite to the data article.&nbsp;<br><br>Field experimentation roles:</p> <p>The image acquisition was performed by Mariano Crimaldi, a researcher, on behalf of Department of Agriculture and the hosting institution University of Naples Federico II, Italy.</p> <p>Shubham Rana has been the curator and analyst for the data under the supervision of his PhD supervisor Prof. Salvatore Gerbino. They are affiliated with Department of Engineering, University of Campania 'Luigi Vanvitelli'.&nbsp;</p> <p>Domenico Barretta, Department of Engineering has been associated in consulting and brainstorming role particularly with data validation, annotation management and litmus testing of the datasets.</p>

opencc-by-4.0Jan 2024View details →
zenodo40/100

Experimental datasets and CT-images for dilatant hardening study - Williams and French 2024

<p>Experimental datasets include collected and calculated parameters for the suite of experiments conducted. CT image datasets are the raw core scans.</p>

opencc-by-4.0Feb 2024View details →
zenodo40/100

Deep Learning Annotation Dataset and Images of Pea Aphids

<p><span>The small size and extensive polymorphisms of aphids make it difficult to identify larvae and adults solely based on their morphology. Here, we present an identification tool for the developmental stages of <em>Acyrthosiphon</em> <em>pisum</em> (Hemiptera: Aphididae) based on deep learning as a proof of concept. You Only Look Once (YOLO) algorithm is one of the most effective deep learning techniques for object detection. Although several studies have been conducted using deep learning technology for the detection and counting of tiny pests, the type of light source and size of the images were the limiting factors, as training was highly focused on uniform datasets and small insects. One way to overcome this problem is to introduce many types of datasets obtained from various light sources and microscopic magnifications. This strategy minimizes errors and omissions in aphid detection across all developmental stages in aphid individuals to the greatest extent possible. The experimental results showed that our modified YOLOv8 model could obtain over 95.9% and 99% accuracy for mean average precision (mAP) and Recall, respectively, under various light sources, such as yellow, white, and natural light, and stereomicroscope magnifications. This study showed an improved accuracy of aphid recognition at all developmental stages.</span><span> </span><span>The study presents a novel deep learning model utilizing the YOLO algorithm to identify developmental stages of </span><em><span>A</span></em><span>. </span><em><span>pisum</span></em><span>. This model achieves high accuracy across various light sources and magnifications, thereby enhancing aphid biology studies.</span></p>

opencc-by-4.0Mar 2024View details →
zenodo40/100

MatSeg DataSet and Benchmark For Zero-Shot Material States Segmentation From images

<h2>This is an old version for the new version see&nbsp;<a href="../records/11331618">https://zenodo.org/records/11331618</a></h2> <p>&nbsp;</p> <p>A Dataset and Benchmark for zero-shot segmentation of materials states described in: &ldquo;Learning Zero-Shot Material States Segmentation, by Implanting Natural Image Patterns in Synthetic Data&rdquo; Described in&nbsp;<strong><a href="https://arxiv.org/pdf/2403.03309.pdf">https://arxiv.org/pdf/2403.03309.pdf</a>&nbsp;</strong></p> <p>See ReadMe in the zip file for technical details.</p> <p>&nbsp;</p> <h2><strong>MatSeg Benchmark&nbsp;</strong></h2> <p>A benchmark for zero-shot material state segmentation. The benchmark contains 820 real-world images with a wide range of material states and settings. For example: food states (cooked/burned..), plants (infected/dry.), to rocks/soil (minerals/sediment),&nbsp; construction/metals (rusted, worn),&nbsp; liquids&nbsp; (foam/sediment), and many other states in a class-agnostic manner.&nbsp; The goal is to evaluate the segmentation of material materials without knowledge or pretraining on the material or setting. The focus is on materials with complex scattered boundaries, and gradual transition&nbsp; (like the level of wetness of the surface). The annotation of the benchmark is point-based and similarity-based. Hence, for each image, we select several points and regions (Figure 4). We group the points of the same materials into the same label, we also define a group of points that have partial similarity. For example points in group A are more similar to points in group B than to points in group C (In case materials A and B are similar to each other but not identical). This approach allows us to capture the complexity of gradual transition and partial similarities in the world. While also enabling dealing with complex scattered and blurry shapes without needing to annotate the full shape which in many cases is unclear or very hard.</p> <p>Files <a href="../api/records/10801191/draft/files/MatSegBenchmarkPart1of3.zip/content" target="_blank" rel="noopener noreferrer">MatSegBenchmark</a>*.zip</p> <h2><strong>MatSeg synthetic Dataset Samples&nbsp;</strong></h2> <p>Synthethic dataset of images of materials spread on object surfaces and their segmentation map.</p> <p>The synthetic dataset is a very big, sample of the dataset as been uploaded.</p> <p>Files:&nbsp; &nbsp; &nbsp; &nbsp;MatSegSynthehticDataSample*.zip</p> <p>The full dataset can be found in this URLS:</p> <p><a href="https://e.pcloud.link/publink/show?code=kZHCcnZOfzqInb3anSl7xzFBoqCDmkr2JKV">https://e.pcloud.link/publink/show?code=kZHCcnZOfzqInb3anSl7xzFBoqCDmkr2JKV</a></p> <p><a href="https://icedrive.net/s/SBb3g9WzQ5wZuxX9892Z3R4bW8jw">https://icedrive.net/s/SBb3g9WzQ5wZuxX9892Z3R4bW8jw</a></p> <p>&nbsp;</p> <p>Generation Script for the synthetic data:</p> <p><a href="https://github.com/sagieppel/MatSeg-Synthethic-Dataset-Generation-Script">https://github.com/sagieppel/MatSeg-Synthethic-Dataset-Generation-Script</a></p> <p><a href="../records/10822596/files/sagieppel/MatSeg-Synthethic-Dataset-Generation-Script-3.zip?download=1">https://zenodo.org/records/10822596</a></p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-zeroMar 2024View details →
zenodo40/100

Oracle bone fragment image dataset

<p>Oracle bone inscriptions record valuable information about ancient Chinese history in Shang Dynasty. However, many complete oracle bones have broken into pieces over the years, resulting in disjointing sentences. This poses a serious challenge in rejoining oracle bone fragments and interpreting the inscriptions. With oracle bone fragments scattered around the world, researchers have turned their attention to artificial intelligence algorithms to piece together oracle bone fragment images. This paper presents a benchmark datasets for rejoining oracle bone fragment. The benchmark consists of three parts. The first part consists of 5374 high-resolution oracle bone fragment images. The second part consists of 110 oracle bone fragment image pairs that can be rejoined together, including the rejoinable location. The third part consists of the images of oracle bone pairs with matching shapes from multiple sources, divided into two categories: 23,072 rejoinable images and 116003 unrejoinable images. The dataset can be used to evaluate the performance oracle bone rejoining algorithms, the dataset also contains undiscovered rejoinable fragments, which offer promising opportunities for further analysis.&nbsp;</p>

opencc-by-4.0Mar 2024View details →
zenodo40/100

Images of dust and datasets obtained from them, cloister of Santa María del Paular, in positions 1, 3 and 6 (2021-2022)

<p>This dataset is related to the research article entitled "Assessment of dust deposition through image analysis in complex and remote exhibition sites &ndash; study in the cloister of the Santa Mar&iacute;a de El Paular Monastery in the Sierra de Guadarrama in Spain "</p> <p>These are photographs of slides located in 3 different points of the cloister, on which the dust was allowed to settle for 4 different periods of time.</p> <p>Folders *_x10 contain original microscope photos</p> <p>Folders *_x10_Split contain the result of dividing the original images (from Folders *_x10) into nine rectangles of equal size</p> <p>Folders *_x10_BW contain the result of of treating images (from Folders *_x10_Split) with an automatic threshold (triangle - ImageJ)</p> <p>Folders *_x10_DT contains csv files with dust particles detected in each image (from folders *_x10_BW)</p> <p>Also included are macros (for splitting and thresholding)</p>

opencc-by-4.0Mar 2024View details →
zenodo40/100

Accompanying dataset for: "IBEX: A versatile multiplex optical imaging approach for deep phenotyping and spatial analysis of cells in complex tissues"

<p>Mouse datasets were acquired using the manual IBEX multiplex imaging protocol and accompany the manuscript &ldquo;IBEX: A versatile multiplex optical imaging approach for deep phenotyping and spatial analysis of cells in complex tissues&rdquo;, A. Radtke&nbsp;<em>et al.</em>, 2020, PNAS.</p> <p>All image data are stored using the&nbsp;<a href="https://imaris.oxinst.com/support/imaris-file-format">Imaris file format</a>. To view these multi-channel images, you can either use one of these&nbsp;free&nbsp;viewers,&nbsp;<a href="https://imaris.oxinst.com/imaris-viewer">Imaris viewer</a>,&nbsp;<a href="https://imagej.net/Fiji">Fiji</a>.</p> <p>Each experiment has an associated imaging meta-data file in xlsx format and the resulting image in Imaris format.</p> <p><strong>Mouse spleen (Manual)</strong></p> <p>Dataset is a 16 parameter&nbsp;IBEX experiment performed on a mouse spleen section labeled with the nuclear marker JOJO-1 and antibodies directed against the indicated markers. Images were acquired using an inverted Leica TCS SP8 X confocal microscope equipped with a 40X objective (NA 1.3), 4 HyD and 1 PMT detectors, a white light laser that produces a continuous spectral output between&nbsp;470 and 670 nm as well as 405, 685, and 730 nm lasers. All images were captured at an 8-bit depth, with a line average of 3, and 1024x1024 format with the following pixel dimensions: x (0.284 &micro;m), y (0.284 &micro;m), and z (1 &micro;m). Images were tiled and merged using the LAS X Navigator software (LAS X 3.5.5.19976).</p> <p><strong>Mouse thymus (Manual)</strong></p> <p>Dataset is a 26 parameter&nbsp;IBEX experiment performed on a mouse thymus section labeled with the nuclear marker JOJO-1 and antibodies directed against the indicated markers. Images were acquired using an inverted Leica TCS SP8 X confocal microscope equipped with a 40X objective (NA 1.3), 4 HyD and 1 PMT detectors, a white light laser that produces a continuous spectral output between&nbsp;470 and 670 nm as well as 405, 685, and 730 nm lasers. All images were captured at an 8-bit depth, with a line average of 3, and 1024x1024 format with the following pixel dimensions: x (0.284 &micro;m), y (0.284 &micro;m), and z (1 &micro;m). Images were tiled and merged using the LAS X Navigator software (LAS X 3.5.5.19976).</p> <p><strong>Mouse lung (Manual)</strong></p> <p>Dataset is a 23 parameter&nbsp;IBEX experiment performed on a mouse lung section labeled with the nuclear marker JOJO-1 and antibodies directed against the indicated markers. Images were acquired using an inverted Leica TCS SP8 X confocal microscope equipped with a 40X objective (NA 1.3), 4 HyD and 1 PMT detectors, a white light laser that produces a continuous spectral output between&nbsp;470 and 670 nm as well as 405, 685, and 730 nm lasers. All images were captured at an 8-bit depth, with a line average of 3, and 1024x1024 format with the following pixel dimensions: x (0.379 &micro;m), y (0.379 &micro;m), and z (1 &micro;m). Images were tiled and merged using the LAS X Navigator software (LAS X 3.5.5.19976).</p> <p><strong>Mouse small intestine (Manual)</strong></p> <p>Dataset is a 20 parameter&nbsp;IBEX experiment performed on a mouse small intestine section labeled with the nuclear marker JOJO-1 and antibodies directed against the indicated markers. Images were acquired using an inverted Leica TCS SP8 X confocal microscope equipped with a 40X objective (NA 1.3), 4 HyD and 1 PMT detectors, a white light laser that produces a continuous spectral output between&nbsp;470 and 670 nm as well as 405, 685, and 730 nm lasers. All images were captured at an 8-bit depth, with a line average of 3, and 1024x1024 format with the following pixel dimensions: x (0.284 &micro;m), y (0.284 &micro;m), and z (1 &micro;m). Images were tiled and merged using the LAS X Navigator software (LAS X 3.5.5.19976).</p> <p><strong>Mouse liver (Manual)</strong></p> <p>Dataset is an 18 parameter&nbsp;IBEX experiment performed on a liver section from a LysM-tdtomato reporter mouse labeled with antibodies directed against the indicated markers. Images were acquired using an inverted Leica TCS SP8 X confocal microscope equipped with a 40X objective (NA 1.3), 4 HyD and 1 PMT detectors, a white light laser that produces a continuous spectral output between&nbsp;470 and 670 nm as well as 405, 685, and 730 nm lasers. All images were captured at an 8-bit depth, with a line average of 3, and 1024x1024 format with the following pixel dimensions: x (0.284 &micro;m), y (0.284 &micro;m), and z (1 &micro;m). Images were tiled and merged using the LAS X Navigator software (LAS X 3.5.5.19976).</p> <p><strong>Mouse naive lymph node (Manual)</strong></p> <p>Dataset is a 41 parameter&nbsp;IBEX experiment performed on a mouse lymph node section labeled with the nuclear marker JOJO-1 and antibodies directed against the indicated markers. Images were acquired using an inverted Leica TCS SP8 X confocal microscope equipped with a 40X objective (NA 1.3), 4 HyD and 1 PMT detectors, a white light laser that produces a continuous spectral output between&nbsp;470 and 670 nm as well as 405, 685, and 730 nm lasers. All images were captured at an 8-bit depth, with a line average of 3, and 1024x1024 format with the following pixel dimensions: x (0.284 &micro;m), y (0.284 &micro;m), and z (1 &micro;m). Images were tiled and merged using the LAS X Navigator software (LAS X 3.5.5.19976).</p> <p><strong>Mouse immunized lymph node (Manual)</strong></p> <p>Dataset is a 41 parameter&nbsp;IBEX experiment performed on a mouse lymph node section labeled with the nuclear marker JOJO-1 and antibodies directed against the indicated markers. Images were acquired using an inverted Leica TCS SP8 X confocal microscope equipped with a 40X objective (NA 1.3), 4 HyD and 1 PMT detectors, a white light laser that produces a continuous spectral output between 470 and 670 nm as well as 405, 685, and 730 nm lasers. All images were captured at an 8-bit depth, with a line average of 3, and 1024x1024 format with the following pixel dimensions: x (0.284 &micro;m), y (0.284 &micro;m), and z (1 &micro;m). Images were tiled and merged using the LAS X Navigator software (LAS X 3.5.5.19976).</p>

opencc-by-4.0Mar 2024View details →
zenodo40/100

Companinon Dataset for PASTIS : VHR satellite images (SPOT 6-7)

<p>To enhance the spatial resolution and utility of <a href="https://github.com/VSainteuf/pastis-benchmark">PASTIS-R dataset</a>, we introduce PASTIS-HD, which integrates contemporaneous VHR satellite images (SPOT 6-7), resampled to a 1m resolution and converted to 8 bits. This enhancement significantly improves the dataset's spatial content, providing more granular information for agricultural parcel segmentation.</p> <p>This folder can be added to the PASTIS-R dataset to get the PASTIS-HD version.<br><br>The SPOT images are opendata thanks to the Dataterra Dinamis initiative in the case of the <a href="https://dinamis.data-terra.org/opendata/">"Couverture France DINAMIS" program</a>.<br><br></p> <p>If you use PASTIS please cite the&nbsp;<a href="https://arxiv.org/abs/2107.07933" rel="nofollow">related paper</a>:</p> <blockquote> <p>@article{garnot2021panoptic,<br>&nbsp; title={Panoptic Segmentation of Satellite Image Time Series<br>with Convolutional Temporal Attention Networks},<br>&nbsp; author={Sainte Fare Garnot, Vivien &nbsp;and Landrieu, Loic },<br>&nbsp; journal={ICCV},<br>&nbsp; year={2021}<br>}</p> </blockquote> <p><br><br>For the PASTIS-R optical-radar fusion dataset, please also cite&nbsp;<a href="https://arxiv.org/abs/2112.07558v1" rel="nofollow">this paper</a>:</p> <blockquote> <pre>@article{garnot2021mmfusion, title = {Multi-modal temporal attention models for crop mapping from satellite time series}, journal = {ISPRS Journal of Photogrammetry and Remote Sensing}, year = {2022}, doi = {https://doi.org/10.1016/j.isprsjprs.2022.03.012}, author = {Vivien {Sainte Fare Garnot} and Loic Landrieu and Nesrine Chehata}, }</pre> </blockquote> <p>For the PASTIS-HD with the 3 modality optical-radar time series plus VHR images dataset, please also cite <a href="https://arxiv.org/abs/2404.08351">this paper</a>:</p> <blockquote> <p>@article{astruc2024omnisat,<br>&nbsp; title={Omni{S}at: {S}elf-Supervised Modality Fusion for {E}arth Observation},<br>&nbsp; author={Astruc, Guillaume and Gonthier, Nicolas and Mallet, Clement and Landrieu, Loic},<br>&nbsp; journal={arXiv preprint arXiv:2404.08351},<br>&nbsp; year={2024}<br>}</p> </blockquote>

openetalab-2.0Apr 2024View details →
zenodo40/100

RoadTrafficMARKS: Dataset of 10057 images (256x256 pixels) of Curved arrow, Straight arrow, Pedestrian crossing, Straight-right merge arrow, Merge arrow, Left arrow, Yield, Stop, Right arrow, Straight-left merge arrow, Speed limit and Straight-right-left merge arrow markings and their BoundingBox labels

<p><strong>The dataset consists of 10057 PNG images (256x256&nbsp;pixels) of </strong><strong>high resolution aerial orthoimages </strong><strong>taggged with </strong><strong>twelve classes of traffic signals</strong><strong>: (1) Curved arrow, (2) Straight arrow, (3) Pedestrian crossing, (4) Straight-right merge arrow, (5) Merge arrow, (6) Left arrow, (7) Yield, (8) Stop, (9) Right arrow, (10) Straight-left merge arrow, (11) Speed limit and (12) Straight-right-left merge arrow markings, together with their corresponding Bounding Boxes. The dataset has been created in the framework of the SROADEX project to train an identification process based on artificial neural networks.</strong></p> <p><strong>The dataset was created by manually tagging the twelve class of marks on&nbsp;orthoimage tiles of 256x256 pixels </strong><strong><strong>with the LabelMe tool</strong>. After the semantic labeling (manual digitalization of the contour of the signals), a&nbsp;transformation and random splitting process has been carried to prepare the data for the neural networks</strong><strong><strong> training</strong>. It resulted in 80% of the images for training (8031), 10% for validation (1000) and 10% for testing (1021).</strong></p> <p><strong>The next table presents the number of images of each class on the </strong><strong><strong>&quot;train&quot;, &quot;valid&quot; and &quot;test&quot; </strong>sets.</strong></p> <table> <tbody> <tr> <td>&nbsp;</td> <td>Train</td> <td>Valid</td> <td>Test</td> <td>Total</td> </tr> <tr> <td>CURVED ARROW</td> <td>319</td> <td>43</td> <td>48</td> <td>410</td> </tr> <tr> <td>STRAIGHT ARROW</td> <td>6730</td> <td>806</td> <td>843</td> <td>8379</td> </tr> <tr> <td>PEDESTRIAN CROSSING</td> <td>3881</td> <td>465</td> <td>485</td> <td>4831</td> </tr> <tr> <td>STRIGNT-RIGHT MERGE ARROW</td> <td>1373</td> <td>187</td> <td>177</td> <td>1737</td> </tr> <tr> <td>MERGE ARROW</td> <td>550</td> <td>91</td> <td>59</td> <td>700</td> </tr> <tr> <td>LEFT ARROW&nbsp;</td> <td>238</td> <td>33</td> <td>31</td> <td>302</td> </tr> <tr> <td>YIELD</td> <td>356</td> <td>47</td> <td>40</td> <td>443</td> </tr> <tr> <td>STOP</td> <td>141</td> <td>19</td> <td>15</td> <td>175</td> </tr> <tr> <td>RIGHT ARROW</td> <td>836</td> <td>93</td> <td>111</td> <td>1040</td> </tr> <tr> <td>STRIGNT-LEFT MERGE ARROW</td> <td>152</td> <td>19</td> <td>21</td> <td>192</td> </tr> <tr> <td>SPEED LIMIT</td> <td>392</td> <td>50</td> <td>45</td> <td>487</td> </tr> <tr> <td>STRIGNT-RIGHT-LEFT MERGE ARROW</td> <td>82</td> <td>9</td> <td>9</td> <td>100</td> </tr> </tbody> </table>

opencc-by-4.0Jul 2023View details →
zenodo40/100

BanglaWriting Words Dataset: A Collection of Isolated Word Images from the BanglaWriting multi-purpose Bangla offline-handwriting dataset (WoBW)

<p>The WoBW (Words from BanglaWriting) dataset is a curated collection of isolated word images, adapted from the original BanglaWriting dataset (url:&nbsp;https://data.mendeley.com/datasets/r43wkvdk4w/1).</p> <p>WoBW focuses on individual words extracted from handwritten Bangla text samples in the BanglaWriting corpus, making it a valuable resource for research in word-level Bangla handwriting recognition and related natural language processing tasks.</p> <p>Mridha, Dr. M. F.; Quwsar Ohi, Abu; Ali, M. Ameer; Emon, Mazedul Islam; Kabir, Md Mohsin (2020), &ldquo;BanglaWriting: A multi-purpose offline Bangla handwriting dataset&rdquo;, Mendeley Data, V1, doi: 10.17632/r43wkvdk4w.1</p>

opencc-by-4.0Nov 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record