Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
979
datasets available to search
ShareScore release 0.9.0
Dataset results
979 results for “image dataset”
(08)-Strobl2021A-DS0002 – Tribolium castaneum ACOS{ATub'H2B-mRuby} #1 subline long-term live imaging dataset of embryonic development acquired with light sheet fluorescence microscopy
<p>(08)-Strobl2021A-DS0002 – <em>Tribolium castaneum</em> ACOS{ATub'H2B-mRuby} #1 subline long-term live imaging data of embryonic development acquired with light sheet fluorescence microscopy</p>
(08)-Strobl2021A-DS0001 – Tribolium castaneum AGOC{Zen1'#O(LA)-mEmerald} #1 subline long-term live imaging dataset of embryonic development acquired with light sheet fluorescence microscopy
<p>(08)-Strobl2021A-DS0001 – <em>Tribolium castaneum</em> AGOC{Zen1'#O(LA)-mEmerald} #1 subline long-term live imaging dataset of embryonic development acquired with light sheet fluorescence microscopy</p>
(07)-Ratke2020A-DS0005 – Tribolium castaneum AGOC{Zen1'#O(LA)-mEmerald} #2 subline long-term live imaging dataset of embryonic development acquired with light sheet fluorescence microscopy
<p>(07)-Ratke2020A-DS0005 – <em>Tribolium castaneum</em> AGOC{Zen1'#O(LA)-mEmerald} #2 subline long-term live imaging dataset of embryonic development acquired with light sheet fluorescence microscopy</p>
(07)-Ratke2020A-DS0004 – Tribolium castaneum AGOC{Zen1'#O(LA)-mEmerald} #1 subline long-term live imaging dataset of embryonic development acquired with light sheet fluorescence microscopy
<p>(07)-Ratke2020A-DS0004 – <em>Tribolium castaneum</em> AGOC{Zen1'#O(LA)-mEmerald} #1 subline long-term live imaging dataset of embryonic development acquired with light sheet fluorescence microscopy</p>
(07)-Ratke2020A-DS0003 – Drosophila melanogaster w[*]; P{w[+mC]=His2Av-EGFP.C}2/SM6a line long-term live imaging dataset of embryonic development acquired with light sheet fluorescence microscopy
<p>(07)-Ratke2020A-DS0003 – <em>Drosophila melanogaster</em> w[*]; P{w[+mC]=His2Av-EGFP.C}2/SM6a line long-term live imaging dataset of embryonic development acquired with light sheet fluorescence microscopy</p>
(07)-Ratke2020A-DS0001 – Drosophila melanogaster y[1] w[67c23]; P{w[+mC]=Ubi-GFP.nls}ID-2; P{Ubi-GFP.nls}ID-3 line long-term live imaging dataset of embryonic development acquired with light sheet fluorescence microscopy
<p>(07)-Ratke2020A-DS0001 – <em>Drosophila melanogaster</em> y[1] w[67c23]; P{w[+mC]=Ubi-GFP.nls}ID-2; P{Ubi-GFP.nls}ID-3 line long-term live imaging dataset of embryonic development acquired with light sheet fluorescence microscopy</p>
(07)-Ratke2020A-DS0002 – Drosophila melanogaster w[*]; P{w[+mC]=Tub84B-EGFP.NLS}3 long-term live imaging dataset acquired with light sheet fluorescence microscopy
<p>(07)-Ratke2020A-DS0002 <em>–</em> <em>Drosophila melanogaste</em>r y[1] w[67c23]; P{w[+mC]=Ubi-GFP.nls}ID-2; P{Ubi-GFP.nls}ID-3 (Bloomington <em>Drosophila</em> Stock Center #29724) long-term live imaging dataset acquired with light sheet fluorescence microscopy</p>
(08)-Strobl2021A-DS0003 – Tribolium castaneum Gruul #1 hybrid line long-term live imaging dataset of embryonic development acquired with light sheet fluorescence microscopy
<p>(08)-Strobl2021A-DS0003 – <em>Tribolium castaneum</em> Gruul #1 hybrid line long-term live imaging dataset of embryonic development acquired with light sheet fluorescence microscopy</p>
(07)-Ratke2020A-DS0006 – Tribolium castaneum AGOC{Zen1'#O(LA)-mEmerald} #3 subline long-term live imaging dataset of embryonic development acquired with light sheet fluorescence microscopy
<p>(07)-Ratke2020A-DS0006 – <em>Tribolium castaneum</em> AGOC{Zen1'#O(LA)-mEmerald} #3 subline long-term live imaging dataset of embryonic development acquired with light sheet fluorescence microscopy</p>
Detection of Areas with Human Vulnerability Using Public Satellite Images and Deep Learning (Dataset)
<div> <h2>Overview</h2> <a href="https://github.com/fbvidal/HumanVulnerabilityDetectionDL#overview"></a></div> <p>This repository contains the code and resources for the project titled <strong>"Detection of Areas with Human Vulnerability Using Public Satellite Images and Deep Learning"</strong>. The goal of this project is to identify regions where individuals are living under precarious conditions and facing neglected basic needs, a situation often seen in Brazil. This concept is referred to as "human vulnerability" and is exemplified by families living in inadequate shelters or on the streets in both urban and rural areas.</p> <p>Focusing on the Federal District of Brazil as the research area, this project aims to develop two novel public datasets consisting of satellite images. The datasets contain imagery captured at 50m and 100m scales, covering regions of human vulnerability, traditional areas, and improperly disposed waste sites.</p> <p>The project also leverages these datasets for training deep learning models, including <strong>YOLOv7</strong> and other state-of-the-art models, to perform image segmentation. A comparative analysis is conducted between the models using two training strategies: training from scratch with random weight initialization and fine-tuning using pre-trained weights through <strong>transfer learning</strong>.</p> <div> <h3>Key Achievements</h3> <a href="https://github.com/fbvidal/HumanVulnerabilityDetectionDL#key-achievements"></a></div> <ul> <li>Two new satellite image datasets focusing on human vulnerability and improperly disposed waste sites, available in public domains.</li> <li>Comparison of image segmentation models, including <strong>YOLOv7</strong> and <strong>Segmentation Models</strong>, with performance metrics.</li> <li>Best F1-scores: 0.55 for <strong>YOLOv7</strong> and 0.64 for <strong>Segmentation Models</strong>.</li> </ul> <p>This repository provides the code, models, and data pipelines used for training, evaluation, and performance comparison of these deep learning models.<br><br></p> <div> <h2>Citation (Bibtex)</h2> <a href="https://github.com/fbvidal/HumanVulnerabilityDetectionDL?tab=readme-ov-file#citation-bibtex"></a></div> <pre><code>@TECHREPORT {TechReport-Julia-Laura-HumanVulnerability-2024, author = "Julia Passos Pontes, Laura Maciel Neves Franco, Flavio De Barros Vidal", title = "Detecção de Áreas com Atividades de Vulnerabilidade Humana utilizando Imagens Públicas de Satélites e Aprendizagem Profunda", institution = "University of Brasilia", year = "2024", type = "Undergraduate Thesis", address = "Computer Science Department - University of Brasilia - Asa Norte - Brasilia - DF, Brazil", month = "aug", note = "People living in precarious conditions and with their basic needs neglected is an unfortunate reality in Brazil. This scenario will be approached in this work according to the concept of \"human vulnerability\" and can be exemplified through families who live in inadequate shelters, without basic structures and on the streets of urban or rural centers. Therefore, assuming the Federal District as the research scope, this project proposes to develop two new databases to be made available publicly, considering the map scales of 50m and 100m, and composed by satellite images of human vulnerability areas, regions treated as traditional and waste disposed inadequately. Furthermore, using these image bases, trainings were done with the YOLOv7 model and other deep learning models for image segmentation. By adopting an exploratory approach, this work compares the results of different image segmentation models and training strategies, using random weight initialization (from scratch) and pre-trained weights (transfer learning). Thus, the present work was able to reach maximum F1 score values of 0.55 for YOLOv7 and 0.64 for other segmentation models." } </code></pre> <div> </div> <div> <h2>License</h2> <a href="https://github.com/fbvidal/HumanVulnerabilityDetectionDL?tab=readme-ov-file#license"></a></div> <p>This project is licensed under the MIT License - see the LICENSE file for details.</p>
FireSafetyNet: An Image-Based Dataset with Pretrained Weights for Machine Learning-Driven Fire Safety Inspection
<p>This dataset offers a diverse collection of images curated to support the development of computer vision models for detecting and inspecting Fire Safety Equipment (FSE) and related components. Images were collected from a variety of public buildings in Germany, including university buildings, student dormitories, and shopping malls. The dataset consists of self-captured images using mobile cameras, providing a broad range of real-world scenarios for FSE detection.</p> <p>In the journal paper associated with these image datasets, the open-source dataset FireNet (Boehm et al. 2019) was additionally utilized for training. However, to comply with licensing and distribution regulations, images from <a href="https://www.firenet.xyz/">FireNet</a> have been excluded from this dataset. Interested users can visit the FireNet repository directly to access and download those images if additional data is required. The provided weights (.pt), however, are trained on the provided self-made images and FireNet using YOLOv8.</p> <p>The dataset is organized into six sub-datasets, each corresponding to a specific FSE-related machine learning service:</p> <ol> <li> <p><strong>Service 1: FSE Detection</strong> - This sub-dataset provides the foundation for FSE inspection, focusing on the detection of primary FSE components like fire blankets, fire extinguishers, manual call points, and smoke detectors.</p> </li> <li> <p><strong>Service 2: FSE Marking Detection</strong> - Building on the first service, this sub-dataset includes images and annotations for detecting FSE marking signs.</p> </li> <li> <p><strong>Service 3: Condition Check - Modal</strong> - This sub-dataset addresses the inspection of FSE condition in a modal manner, focusing on instances where fire extinguishers might be blocked or otherwise non-compliant. This dataset includes semantic segmentation annotations of fire extinguishers. For upload reasons, this set is split into <em>3_1_FSE Condition Check_modal_train_data (containing training images and annotations) </em>and <em>3_1_FSE Condition Check_modal_val_data_and_weights (containing validation images, annotations </em>and<em> the best weights).</em></p> </li> <li> <p><strong>Service 4: Condition Check - Amodal</strong> - Extending the modal condition check, this sub-dataset involves amodal detection to identify and infer the state of FSE components even when they are partially obscured. This dataset includes semantic segmentation annotations of fire extinguishers. This dataset includes semantic segmentation annotations of fire extinguishers. For upload reasons, this set is split into <em>4_1_FSE Condition Check_amodal_train_data (containing training images and annotations) </em>and <em>4_1_FSE Condition Check_amodal_val_data_and_weights (containing validation images, annotations </em>and<em> the best weights).</em></p> </li> <li> <p><strong>Service 5: Details Extraction - Inspection Tags</strong> - This sub-dataset provides a detailed examination of the inspection tags on fire extinguishers. It includes annotations for extracting semantic information such as the next maintenance date, contributing to a thorough evaluation of FSE maintenance practices.</p> </li> <li> <p><strong>Service 6: Details Extraction - Fire Classes Symbols</strong> - The final sub-dataset focuses on identifying fire class symbols on fire extinguishers.</p> </li> </ol> <p>This dataset is intended for researchers and practitioners in the field of computer vision, particularly those engaged in building safety and compliance initiatives.</p>
Datasets used in a Transformer network for image inversion of multi-dimensional nonuniform aperture synthesis radiometers
<p><span>该数据集于 2023 年 11 月在中南大学生成,并通过 matlab 仿真软件进行仿真和收集。主要用于图像重建网络的训练和测试。</span>该数据集由原始场景亮度数据、能见度数据和一维、二维和三维非均匀天线阵列图像重建的能见度函数对应的频域采样点位置数据,以及使用其他一些常规方法进行图像重建获得的亮度数据组成。此外,为了验证所提方法的有效性,在工作频率为 33.5 Ghz 的原型 8 元一维非均匀天线阵列上进行了室内实验,并生成了测量数据集。</p> <p>具体来说,名为 Tb_in、V2_noise、T3_AAF 和 T2_idft 的四个仿真数据集存储在名为 1d 的 zip 包中。</p> <p>Tb_IN_1d存储了一维非均匀天线对应的原始场景亮温数据,该数据选自西北工业大学制作的遥感影像场景分类公共数据集。</p> <p>V2_noise存储了包含各种误差的能见度函数值,主要是通过将原始场景亮温图像输入到运行在接收频率为 33.5 GHz 的非均匀积分孔径辐射计模拟程序中得到的。</p> <p>T2_idft 和 T3_AAF 分别是使用逆离散傅里叶变换和阵列因子形成方法进行图像重建获得的明亮温度数据。这两组数据都可用于后续的比较实验。</p> <p>名为 Tb_IN_2d、Tb_out_2d、VS_2d 和 VS_P_2d 的四个数据集存储在名为 2d 的 zip 包中。</p> <p>Tb_IN_2d内部存储的是二维非均匀天线对应的原始场景亮温数据,该数据选自西北工业大学制作的遥感影像场景分类公共数据集。</p> <p>VS_2d为二维非均匀天线阵列对应的能见度函数值,主要是将原始场景亮温图像输入到接收频率为 33.5 GHz 的非均匀集成孔径辐射计模拟程序中得到的。</p> <p>VS_P_2d存储了二维非均匀天线阵列的能见度函数对应的频域采样点位置,该值主要通过计算能见度函数值得到。</p> <p>Tb_out_2d文件存储了使用传统方法进行图像重建得到的亮温值,该数据也用于后续与所提方法获得的数据的比较实验。</p> <p>名为 Tb_IN_3d、Tb_out_3d、VS_3d 和 VS_P_3d 的四个数据集存储在名为 3d 的 zip 包中。</p> <p>Tb_IN_3d内部存储的是 3D 非均匀天线对应的原始场景亮温数据,该数据选自西北工业大学制作的遥感图像场景分类公共数据集。</p> <p>VS_3d是三维非均匀天线阵列对应的能见度函数值,是将原始场景亮温图像输入到接收频率为 33.5 GHz 的非均匀积分孔径辐射计仿真程序中得到的。</p> <p>VS_P_3d存储了三维非均匀天线阵列的能见度函数对应的频域采样点位置,该位置是通过计算能见度函数值得到的。</p> <p>Tb_out_3d文件存储了使用常规方法进行图像重建得到的亮温值,该数据也用于后续与所提方法获得的数据的比较实验。</p> <p>名为 Array2_R、Array3_R、V2_noise 和 Tb_out 的四个数据集存储在名为 8mm8 channel 的 zip 包中。</p> <p>存储在 Array2_R 和 Array3_R 中的是系统在不同位置测量的目标点源的相关矩阵,矩阵中元素的值反映了测试点对目标点源的检测能力。根据此相关矩阵,可以计算可见性值。</p> <p>存储在 V2_noise 内部的是包含各种误差的测量可见性函数的样本。对这些数据进行实验主要是为了验证所提方法的有效性。</p> <p>存储在 Tb_out 中的是使用测试数据集获得的亮温结果数据,用于测试训练的网络。</p>
Hexaploid Wheat Spike Image Dataset (7 species and one amphidiploid; 190 images)
<p>The sample included plants of seven species of hexaploid wheat and one amphidiploid from the collection of Dr. N.P. Goncharov (Institute of Cytology and genetics, SB RAS, Novosibirsk, Russia). Plants were grown in a hydroponic greenhouse under individual isolation and standard conditions of humidity, temperature and light for several seasons. In total, the sample included 19 accessions and 190 plants (1 spike per plant).</p> <p>Spike images were obtained under laboratory conditions using a “table” protocol as described in previous work (Genaev et al. (2019) Morphometry of the Wheat Spike by Analyzing 2D Images. Agronomy 2019, 9, 390). </p> <p>This is supplementary material for the paper: Komyshev, E.G.; Genaev, M.A.; Kruchinina, Y.V.; Koval, V.S.; Goncharov, N.P.; Afonnikov, D.A. Evaluation of the Spike Diversity of Seven Hexaploid Wheat Species and an Artificial Amphidiploid Using a Quadrangle Model Obtained from 2D Images. Plants 2024, 13, 2736. Please cite this work when you publishing material related to this dataset.</p>
GalaxiesML: an imaging and photometric dataset of galaxies for machine learning
<div># GalaxiesML README</div> <p> </p> <div>Version 6.1</div> <p> </p> <div>## Overview</div> <p> </p> <div>GalaxiesML is a machine learning-ready dataset of galaxy images, photometry, redshifts, and structural parameters. It is designed for machine learning applications in astrophysics, particularly for tasks such as redshift estimation and galaxy morphology classification. The dataset comprises **286,401 galaxy images** from the Hyper-Suprime-Cam (HSC) Survey PDR2 in five filters: g, r, i, z, y, with spectroscopically confirmed redshifts as ground truth.</div> <p> </p> <div>This dataset is particularly useful for developing machine learning models for upcoming large-scale surveys like **LSST** and **Euclid**.</div> <p> </p> <div>## Features</div> <p> </p> <div>- **286,401 galaxy images** in five photometric bands (g, r, i, z, y).</div> <div>- Spectroscopic redshifts for each galaxy, with redshift values ranging from **0.01 to 4**.</div> <div>- Morphological parameters derived from galaxy images, including **Sérsic index**, **half-light radius**, and **ellipticity**.</div> <div>- **Machine learning-friendly formats**: images are provided in **HDF5** format, along with CSV metadata.</div> <p><br><br></p> <div>## Examples of Using GalaxiesML</div> <p> </p> <div>Examples of uses of GalaxiesML are outlined in Do et al. (2024). The repository for example code are here:</div> <p> </p> <div><a href="https://github.com/astrodatalab/galaxiesml_examples">https://github.com/astrodatalab/galaxiesml_examples</a></div> <p> </p> <div>## Citation</div> <p> </p> <div>Please cite the following papers if you use this dataset in your work:</div> <p> </p> <div>1. **GalaxiesML Dataset**:</div> <div> <div> <div>Do, T. et al., *GalaxiesML: A Dataset of Galaxy Images, Photometry, Redshifts, and Structural Parameters for Machine Learning*. arXiv:2410.00271, <a href="https://arxiv.org/abs/2410.00271">https://arxiv.org/abs/2410.00271</a> (2024)</div> </div> </div> <p> </p> <div>2. **Hyper Suprime-Cam Subaru Strategic Program (HSC PDR2)**:</div> <div>- Aihara, H., et al., *Second Data Release of the Hyper Suprime-Cam Subaru Strategic Program*. Publications of the Astronomical Society of Japan, 71(6), 114 (2019). DOI: [10.1093/pasj/psz103](https://doi.org/10.1093/pasj/psz103)</div> <p> </p> <div>3. **Spectroscopic Surveys**:</div> <div>- Several publicly available spectroscopic redshift catalogs were used in creating this dataset. Notable sources include:</div> <div>- **zCOSMOS Survey**: Lilly, S. J., et al., *The zCOSMOS 10k-Bright Spectroscopic Sample*. The Astrophysical Journal Supplement Series, 184(2), 218-229 (2009). DOI: [10.1088/0067-0049/184/2/218](https://doi.org/10.1088/0067-0049/184/2/218)</div> <div>- **VIMOS Public Extragalactic Survey (VIPERS)**: Garilli, B., et al., *The VIMOS Public Extragalactic Survey (VIPERS): First Data Release of 57,204 Spectroscopic Measurements*. Astronomy & Astrophysics, 562, A23 (2014). DOI: [10.1051/0004-6361/201322790](https://doi.org/10.1051/0004-6361/201322790)</div> <div>- **DEEP2 Survey**: Newman, J. A., et al., *The DEEP2 Galaxy Redshift Survey: Design, Observations, Data Reduction, and Redshifts*. The Astrophysical Journal Supplement Series, 208(1), 5 (2013). DOI: [10.1088/0067-0049/208/1/5](https://doi.org/10.1088/0067-0049/208/1/5)</div> <p><br><br></p> <div>## How to Access</div> <p> </p> <div>The dataset is publicly available on **Zenodo** with the DOI: **[10.5281/zenodo.11117528](https://doi.org/10.5281/zenodo.11117528)**.</div> <p> </p> <div>## License</div> <p> </p> <div>This dataset is licensed under a **Creative Commons Attribution 4.0 International License (CC BY 4.0)**. You are free to share and adapt the dataset as long as appropriate credit is given. For more details, visit: **[CC BY 4.0 License](https://creativecommons.org/licenses/by/4.0/)**.</div> <p> </p> <div>Please cite the references mentioned above if you use this dataset in your work.</div>
Ocular Toxoplasmosis Fundus Images Dataset
<p>The <strong>Ocular Toxoplasmosis (OT) Fundus Images Dataset</strong> contains a set of eye images collected at the <em>Hospital de Clínicas</em> and <em>Hospital General Pedriático Acosta Ñu </em>medical centers from Asunción, Paraguay.</p> <p>The dataset is used in the generation of models for automatic detection of ocular toxoplasmosis. Used as a tool for OT diagnosis, a predictive model could save time, help diagnose atypical cases and also assist ophthalmologists, being particularly useful for those with less experience.</p> <p>Dataset structure:</p> <pre><code>+-- Ocular_Toxoplasmosis_Data | +-- masks | +-- images | +-- dataset_labels.csv</code></pre> <p>The dataset contains two major folders, one with all the collected images and another one with the masks for lesions of eyes images with ocular toxoplasmosis. The dataset also includes a <strong>csv</strong> file with the labels for each image: healthy, active and inactive lesions, this two being non healthy.</p> <p>Notes about the masks: the dataset includes masks for all non healthy images (active and inactive lesions). In order to differentiate the active lesions which are less common that the inactive ones, the mask includes the suffix <strong>-a </strong>for all the masks of active lesions.</p>
X-ray tomography 3D image dataset of natural fibre reinforced polypropylene
<p>Natural fibre composites have potential sustainability benefits over traditional composites, but their irregular shapes and mechanical properties require more thorough examination compared to glass or carbon fibre composites. 3D X-ray imaging allows for non-destructive examination of the structure and shape. The dataset includes 3D images obtained using micro X-ray computed tomography of natural fibre composites. The images provide valuable insights into the material's characteristics. Since there is limited open-access 3D image data on natural fibre composites, this dataset lays the groundwork for future image analysis and numerical modelling.</p>
Dataset for Medical Image Processing in Python Carpentries lesson
<p>This dataset contains a collection of medical imaging files for use in the <a href="https://github.com/esciencecenter-digital-skills/medical-image-processing">"Medical Image Processing with Python" lesson</a>, originally developed by the <a href="https://www.esciencecenter.nl/">Netherlands eScience Center</a>. </p> <p>The dataset includes:</p> <ol> <li>SimpleITK compatible files: MRI T1 and CT scans (<em>training_001_mr_T1.mha, training_001_ct.mha</em>), digital X-ray (<em>digital_xray.dcm</em> in DICOM format), neuroimaging data (<em>A1_grayT1.nrrd, A1_grayT2.nrrd</em>). Data have been downloaded from <a href="https://insightsoftwareconsortium.github.io/SimpleITK-Notebooks/Python_html/00_Setup.html">here</a>. </li> <li>MRI data: a T2-weighted image (<em>OBJECT_phantom_T2W_TSE_Cor_14_1.nii</em> in NIfTI-1 format). Data have been downloaded from <a href="https://zenodo.org/records/6467772">here</a>. </li> <li>Example images for the machine learning lesson: chest X-rays (<em>rotatechest.png, other_op.png</em>), cardiomegaly example (<em>cardiomegaly_cc0.png</em>).</li> <li>Array data: Array data for the Intro to Medical Imaging lesson. Numpy arrays were created by processing and manipulation of publicly available data i.e. from <a href="https://doi.org/10.1109/TNS.1974.6499235">the Schepp Logan phantom</a> and from the <a href="https://fastmri.med.nyu.edu/">NYU FastMRI dataset</a> <div> </div> </li> <li>Additional data: to be added</li> </ol> <p>These files represent various medical imaging modalities and formats commonly used in clinical research and practice. They are intended for educational purposes, allowing students to practice image processing techniques, machine learning applications, and statistical analysis of medical images using Python libraries such as scikit-image, pydicom, and SimpleITK.</p>
BlueberryDCM: A Canopy Image Dataset for Detection, Counting, and Maturity Assessment of Blueberries
<p>The <strong>BlueberryDCM</strong> dataset consists of <strong>140 RGB images</strong> of blueberry canopies captured at varied spatial scales. All the images were acquired using smartphones in natural field light conditions in different orchards in the season of 2022, with 134 images in Mississippi and 6 images in Michigan. A total of <strong>17,955 bounding box annotations</strong> were manually done in the <a href="https://www.robots.ox.ac.uk/~vgg/software/via/">VGG Image Annotator</a> (VIA) (v2.0.12) for the blueberry instances of two fruit maturity classes, "<strong>Blue</strong>" and "<strong>Unblue</strong>", representing ripe and unripe fruit, respectively. In addition, for each maturity class, there are two sub-categories in the annotation, "<strong>visible</strong>", and "<strong>occluded</strong>", to indicate whether the fruit is fully visible in the canopy or partially occluded. The original annotation format exported from the VGG is <a href="https://www.robots.ox.ac.uk/~vgg/software/via/">VIA .json.</a> The derived annotation files in two other formats, .xml (<a href="https://docs.cvat.ai/docs/manual/advanced/formats/format-voc/">Pascal VOC format</a>) and .txt (<a href="https://docs.ultralytics.com/datasets/detect/#ultralytics-yolo-format">YOLO format with noralized xywh</a>, with 0, 1, 2, and 3 denoting the four categories of "<strong>Unblue_visible</strong>", "<strong>Unblue_occluded</strong>", "<strong>Blue_visible</strong>", and "<strong>Blue_occluded</strong>" bluerries, respectively) are provided in the dataset for the compatibility of a wide range of object detectors. Hence, the dataset contains both the raw images (.jpg) and three corresponding annotations files (.json, .xml, and .txt) with the same file names, totaling about 107 MB in file size. </p> <p> </p> <p>The dataset was used for in a study (see below) on the <a href="https://www.sciencedirect.com/science/article/pii/S2772375524002259">evaluation of YOLOv8 and YOLOv9 models for blueberry detection, counting, and maturity assessment</a>. The detection accuracy of 93% mAP@50 was achieved by YOLOv8l, with an error of about 10 blueberries in fruit counting and an error of 3.6% in estimating the "Blue" fruit percentage. Software programs for the modeling work are made publicly available at: <a href="https://github.com/vicdxxx/BlueberryDetectionAndCounting">https://github.com/vicdxxx/BlueberryDetectionAndCounting</a>. In addition, the blueberry dataset was also used as a preliminary database for developing an iOS-based mobile application, which is described in <a href="https://doi.org/10.13031/aim.202401022">Deng, B., Lu, Y., WanderWeide, J., 2024. Development and preliminary evaluation of a deep learning-based fruit counting mobile application for highbush Blueberries. 2024 ASABE Annual International Meeting 2401022</a></p> <p> </p> <p>Details about the dataset curation and statistics as well as modeling experiments are described in the journal article: <a href="https://www.sciencedirect.com/science/article/pii/S2772375524002259">Deng, B., Lu, Y., 2024</a>. <a href="https://doi.org/10.1016/j.atech.2024.100620">Detection, Counting, and Maturity Assessment of Blueberries in Canopy Images using YOLOv8 and YOLOv9</a><a href="https://www.sciencedirect.com/science/article/pii/S2772375524002259">. Smart Agricultural Technology.</a> <a title="Persistent link using digital object identifier" href="https://doi.org/10.1016/j.atech.2024.100620" target="_blank" rel="noreferrer noopener">https://doi.org/10.1016/j.atech.2024.100620</a>. If you use the dataset in published research, please consider citing the dataset or the journal article. Hopefully, you find the dataset useful. </p>
Dataset for "LOROS: Laboratory Simulations of the Optical RadiOmeter composed of CHromatic Imagers (OROCHI) Experiment of the Martian Moons eXploration (MMX) Mission"
<p>This dataset hosts the image and numerical data analysed and derived in the accompanying Stabbins & Kameda article for the special issue of Progress in Earth and Planetary Science on instrumentation and preparations for the JAXA Martian Moons eXploration (MMX) mission. The paper describes and validates the performance of the Laboratory OROCHI Simulator (LOROS).</p> <p>OROCHI (Optical RadiOmeter composed of CHromatic Imagers) is a multispectral multi-view imaging system for the JAXA MMX spacecraft, that will image Phobos and Deimos across 8 visible and near-infrared spectral channels with unprecedented spatial resolution, recording data that in synergy with the other instruments of the MMX spacecraft and rover will constrain hypotheses on the origin of the Martian moons.</p> <p>LOROS is a laboratory simulator of OROCHI, constructed from commercial off-the-shelf parts.</p> <p>The dataset for the characterisation and validation of LOROS is composed of the following sub-sets:</p> <p>A. Modulation Transfer Function<br>B. Expected Reflectance of Carbonaceous Chondrite & Dark Spectralon<br>C. Radiometric Calibration<br>D. Dark Spectralon Validation</p> <div> <h2>Dataset A: Modulation Transfer Function</h2> This dataset includes the table of results of MTF measurements of the slant-edge target at 5 different random orientations in the range of ~7--10°: <div>- <code>mtf_results_07122023.csv</code></div> <br> <div>and the region-of-interest images, for each orientation and each LOROS channel, used to perform the analysis via the <a href="https://sourceforge.net/p/mtfmapper/home/Home/" target="_blank" rel="noopener">MTF Mapper software</a>:</div> <div>- <code>mtf_measurements_07122023</code></div> <br> <div>The directory tree of measurements, for the <em>n</em>th orientation, is illustrated below. Region-of-interest images are stored under <code>img</code>, and are averaged over 25 repeat images to minimise random noise, have had dark frames subtracted, and have been converted from 12-bit to 8-bit grayscale images for compatibility with the MTF Mapper software. Modulation Transfer Function (MTF) and Spatial Frequency Response (SFR) diagnostics generated by MTF Mapper are stored in the <code>results</code> directory.</div> <div> </div> <div><code>mtf_measurements_07122023</code></div> <div><code>├── mtf_knifeedge_low_07122023_*n*</code></div> <div><code>│ ├── img</code></div> <div><code>│ │ ├── 0_850_img_ave.tif</code></div> <div><code>│ │ ├── 1_475_img_ave.tif</code></div> <div><code>│ │ ├── ...</code></div> <div><code>│ ├── results</code></div> <div><code>│ │ ├── 0_850_img_ave_annotated.jpg</code></div> <div><code>│ │ ├── 0_850_img_ave_edge_mtf_values.txt</code></div> <div><code>│ │ ├── 0_850_img_ave_edge_sfr_values.txt</code></div> <div><code>│ │ ├── 1_475_img_ave_annotated.jpg</code></div> <div><code>│ │ ├── ...</code></div> <div><code>├── mtf_knifeedge_low_07122023_*n+1*</code></div> <div><code>│ ├── img</code></div> <div><code>│ │ ├── ...</code></div> <div> </div> <div>This data constitutes part of <strong>Table 1</strong> and <strong>Figure 2</strong> of the manuscript.</div> <div> <h2>Dataset B: Expected Reflectance of Carbonaceous Chondrite & Dark Spectralon</h2> This dataset includes the high-resolution ($\delta\lambda$=1 nm) reference reflectance spectra of the representative Carbonaceous Chondrite meteorite (<a href="https://westernreflectancelab.com/visor/graph/?results-selection=16136&results-item=16136&results-item=15972&results-item=231&results-item=230&graph=&form-TOTAL_FORMS=1&form-INITIAL_FORMS=0&form-MIN_NUM_FORMS=0&form-MAX_NUM_FORMS=1000&form-0-sample_name=nogoya&form-0-any_field=meteorite&form-0-id=&sort_params=-sample_name&page_selected=1&jump-to-page=" target="_blank" rel="noopener">Nogoya)</a> and the 5% reflectance Spectralon calibration target (<a href="https://www.labsphere.com/wp-content/uploads/2021/09/SpectralonStandards.pdf" target="_blank" rel="noopener">SCT5</a>):<br> <div>- <code>highres_input.csv</code></div> <br> <div>and the resampled spectra of these materials expected for OROCHI and LOROS filter wavelengths:</div> <br> <div>- <code>loros_observation.csv</code></div> <div>- <code>orochi_observation.csv</code></div> <br> <div><code>B_expected_reflectance</code></div> <div><code>├── README.md</code></div> <div><code>├── highres_input.csv</code></div> <div><code>├── loros_observation.csv</code></div> <div><code>└── orochi_observation.csv</code></div> <br> <div>This data constitutes <strong>Table 1</strong> and <strong>Figure 10</strong> of the manuscript.</div> <div> </div> <div> <div> <h2>Dataset C: Radiometric Calibration</h2> This dataset contains the image and derived data for 4 experiments with different illumination conditions for characterising the radiometric response of each of the 8 channels of LOROS.</div> <div><br> <div>This dataset contributes to <strong>Tables 2 - 4</strong> and <strong>Figures 3 - 9</strong> of the manuscript.</div> <br> <div>The final derived metrics are hosted in the spreadsheet:</div> <br> <div>- <code>measured_sensor_properties.csv</code></div> <br> <div>and image data and intermediary derived properties for each experiment are stored in the</div> <br> <div>- <code>experiments</code></div> <br> <div>directory.</div> <br> <div><code>C_radiometric_calibration</code></div> <div><code>├── README.md</code></div> <div><code>├── experiments</code></div> <div><code>│ ├── F*S5L10</code></div> <div><code>│ ├── F*S99L10</code></div> <div><code>│ ├── FGS99L2</code></div> <div><code>│ └── FGS99L10</code></div> <div><code>└── measured_sensor_properties.csv</code></div> <br> <h3><code>experiments</code> Directories</h3> In the directory of each experiment are sub-directories hosting Photon Transfer and Dark Transfer datasets, and a spreadsheet of derived metrics of these.<br> <div> </div> <div><code>C_radiometric_calibration</code></div> <div><code>├── README.md</code></div> <div><code>├── experiments</code></div> <div><code>│ ├── F*S5L10</code></div> <div><code>│ │ ├── dark_transfer_curve</code></div> <div><code>│ │ ├── photo_transfer_curve</code></div> <div><code>│ │ └── F*S5L10_derived_properties.csv</code></div> <div><code>│ └── ...</code></div> <div><code>└── measured_sensor_properties.csv</code></div> <div> </div> </div> <div> </div> <div><strong>Derived Properties</strong><br> <div> </div> <div>The spreadsheet (<code>[experiment]_derived_properties.csv</code>) collecting the properties derived from each experiment holds the following information, that has been extracted from the Photon Transfer and Dark Transfer curves as described in §4.2 of the manuscript:</div> <br> <div><code>camera # The camera number and wavelength</code></div> <div><code>k_adc # Sensitivity (e-/DN)</code></div> <div><code>full_well_e # Saturation Capacity (electrons)</code></div> <div><code>full_well_dn # Saturation Capacity (Digital Numbers)</code></div> <div><code>read_noise_e # Read Noise (electrons)</code></div> <div><code>read_noise_dn # Read Noise (Digital Numbers)</code></div> <div><code>bias_e # Offset (electrons)</code></div> <div><code>bias_dn # Offset (Digital Numbers)</code></div> <div><code>dark_current_e # Dark Current (electrons/second)</code></div> <div><code>dark_current_dn # Dark Current (Digital Numbers/second)</code></div> <div><code>DR # Dynamic Range</code></div> <div><code>lin_min # Minimum Linearity Error</code></div> <div><code>lin_max # Maximum Linearity Error</code></div> <div><code>linearity # Average Linearity Error</code></div> <div><code>snr_max # Maximum Signal-to-Noise Ratio</code></div> <div><code>t_exp_min # Minimum Exposure used in experiment (seconds)</code></div> <div><code>t_exp_max # Maximum Exposure used in experiment (seconds)</code></div> <div><code>expected_response # Expected Response (or 'Digital Flux') for OROCHI^12 at Phobos (Digital Numbers/second)</code></div> <div><code>response # Fitted Response (or 'Digital Flux') (Digital Numbers/second)</code></div> <br> <div>These values are given for each channel of LOROS, as well as the expected values for LOROS in off-the-shelf configuration (with no gain adjustment), LOROS with the gain adjustment, and OROCHI if downsampled to 12-bit resolution digital numbers.</div> <br> <div>This data constitutes <strong>Table 2</strong> of the manuscript.</div> <br> <div><strong>Dark Transfer Curve</strong></div> <br> <div>The <code>dark_transfer_curve</code> directory hosts the derived Dark Transfer Curve data (<code>derived_data</code>) and the source region-of-interest dark image pair data (<code>raw_data</code>) for each LOROS channel.</div> <br> <div><code>dark_transfer_curve</code></div> <div><code>├── derived_data</code></div> <div><code>│ ├── F*S5L10_0_850_dtc.csv</code></div> <div><code>│ ├── F*S5L10_1_475_dtc.csv</code></div> <div><code>│ ├── F*S5L10_2_400_dtc.csv</code></div> <div><code>│ ├── F*S5L10_3_550_dtc.csv</code></div> <div><code>│ ├── F*S5L10_4_725_dtc.csv</code></div> <div><code>│ ├── F*S5L10_5_950_dtc.csv</code></div> <div><code>│ ├── F*S5L10_6_650_dtc.csv</code></div> <div><code>│ └── F*S5L10_7_550_dtc.csv</code></div> <div><code>└── raw_data</code></div> <div><code>├── 0_850</code></div> <div><code>│ ├── 850_10095570us_1_calibration.tif</code></div> <div><code>│ ├── 850_10095570us_2_calibration.tif</code></div> <div><code>│ ├── 850_104us_1_calibration.tif</code></div> <div><code>│ ├── 850_104us_2_calibration.tif</code></div> <div><code>│ ├── ...</code></div> <div><code>├── 1_475</code></div> <div><code>├── 2_400</code></div> <div><code>├── 3_550</code></div> <div><code>├── 4_725</code></div> <div><code>├── 5_950</code></div> <div><code>├── 6_650</code></div> <div><code>├── 7_550</code></div> <div><code>└── camera_config.csv</code></div> <br> <div>The <code>raw_data</code> directory hosts a dark image pair for each exposure time used, and the <code>camera_config.csv</code> spreadsheet gives metadata for the system configuration, including the coordinates and dimensions of the region-of-interest for each channel.</div> <br> <div>The dark transfer curve for each experiment and each channel (<code>[experiment]_[channel]_[wavelength]_dtc</code>) gives the data derived from each raw image data, with the following values:</div> <br> <div><code>exposure # exposure duration (seconds)</code></div> <div><code>n_pix # number of pixels in the region of interest</code></div> <div><code>mean # average value of the region of interest</code></div> <div><code>std_t # total standard deviation of the region of interest</code></div> <div><code>std_rs # read+shot-noise standard deviation, copmuted from the difference of the image pair</code></div> <br> <div>This data constitutes <strong>Figures 5 and 8</strong> of the manuscript.</div> <br> <div><strong>Photon Transfer</strong></div> <br> <div>The <code>photon_transfer_curve</code> directory hosts the derived Photon Transfer Curve data (<code>derived_data</code>) and the source region-of-interest illuminated image pairs and associated dark frame image data (<code>raw_data</code>) for each LOROS channel.</div> <br> <div><code>photo_transfer_curve</code></div> <div><code>├── derived_data</code></div> <div><code>│ ├── F*S5L10_0_850_ptc.csv</code></div> <div><code>│ ├── F*S5L10_1_475_ptc.csv</code></div> <div><code>│ ├── F*S5L10_2_400_ptc.csv</code></div> <div><code>│ ├── F*S5L10_3_550_ptc.csv</code></div> <div><code>│ ├── F*S5L10_4_725_ptc.csv</code></div> <div><code>│ ├── F*S5L10_5_950_ptc.csv</code></div> <div><code>│ ├── F*S5L10_6_650_ptc.csv</code></div> <div><code>│ └── F*S5L10_7_550_ptc.csv</code></div> <div><code>└── raw_data</code></div> <div><code>├── 0_850</code></div> <div><code>│ ├── 850_104us_1_calibration.tif</code></div> <div><code>│ ├── 850_104us_2_calibration.tif</code></div> <div><code>│ ├── 850_104us_d_drk.tif</code></div> <div><code>│ ├── 850_105828us_1_calibration.tif</code></div> <div><code>│ ├── ...</code></div> <div><code>├── 1_475</code></div> <div><code>├── 2_400</code></div> <div><code>├── 3_550</code></div> <div><code>├── 4_725</code></div> <div><code>├── 5_950</code></div> <div><code>├── 6_650</code></div> <div><code>├── 7_550</code></div> <div><code>└── camera_config.csv</code></div> <br> <div>The <code>raw_data</code> directory hosts an image pair and dark frame for each exposure time used, and the <code>camera_config.csv</code> spreadsheet gives metadata for the system configuration, including the coordinates and dimensions of the region-of-interest for each channel.</div> <br> <div>The photon transfer curve for each experiment and each channel (<code>[experiment]_[channel]_[wavelength]_ptc</code>) gives the data derived from each raw image data, with the following values across the region-of-interest:</div> <br> <div><code>exposure # exposure duration (seconds)</code></div> <div><code>n_pix # number of pixels in the region of interest</code></div> <div><code>mean # average value (Digital Numbers)</code></div> <div><code>std_t # total standard deviation (Digital Numbers)</code></div> <div><code>std_rs # read+shot-noise standard deviation (Digital Numbers), computed from the difference of the image pair</code></div> <div><code>d_mean # average value of the dark (Digital Numbers)</code></div> <div><code>d_dsnu # Dark Signal Nonuniformity (Digital Numbers)</code></div> <div><code>std_s # Shot Noise (read noise removed) (Digital Numbers)</code></div> <div><code>k_adc # Sensitivity (note this the point-wise sensitivity, rather than fitted) (electrons/Digital Number)</code></div> <div><code>linearity # Linearity Error (point-wise distance to least-squares linear fit) (%)</code></div> <div><code>snr # Signal-to-Noise Ratio, derived from shot-noise (point-wise)</code></div> <div><code>snr_t # Signal-to-Noise Ratio, derived from total noise (point-wise)</code></div> <div><code>e- # Electron count, derived from sensitivity</code></div> <div><code>e-_noise # Electron shot-noise, derived from sensitivity</code></div> <br> <div>This data constitutes <strong>Figures 3, 4, 6, 7 & 9</strong> of the manuscript.</div> <br> <div><strong>Measured Sensor Properties</strong></div> <br> <div>The <code>measured_sensor_properties.csv</code> spreadsheet collects and averages the following metrics over the 4 experiments performed, to give the values for each channel, along with the expected values for LOROS in off-the-shelf configuration, gain-adjusted LOROS, and OROCHI downsampled to 12-bit resolution.</div> <br> <div><code>SNR Max</code></div> <div><code>Dynamic Range (dB)</code></div> <div><code>Dynamic Range (bits)</code></div> <div><code>Sensitivity (e-/DN)</code></div> <div><code>Saturation Capacity (e-)</code></div> <div><code>Saturation Capacity (DN)</code></div> <div><code>Read Noise (e-)</code></div> <div><code>Read Noise (DN)</code></div> <div><code>Nonlinearity (%)</code></div> <div><code>Dark Signal@30°C (e-/s)</code></div> <div><code>Dark Signal@30°C (DN/s)</code></div> <div><code>Bias (e-)</code></div> <div><code>Bias (DN)</code></div> <div><code>DSNU1288 (DN)</code></div> <div><code>DSNU1288 (e-)</code></div> <div><code>PRNU1288 (%)</code></div> <br> <div>This data constitutes <strong>Table 3</strong> of the manuscript.</div> <div> </div> <div> <h2>Dataset D: Dark Spectralon Validation</h2> This dataset contains the raw image and derived data used to demonstrate the ability of LOROS to measure the spectral reflectance of the 5% reflectance Spectralon calibration target (<a href="https://www.labsphere.com/wp-content/uploads/2021/09/SpectralonStandards.pdf" target="_blank" rel="noopener">SCT5</a>).<br> <div>The image data is hosted in the directory:</div> <br> <div>- <code>raw_data</code></div> <br> <div>and the processed data (e.g. reflectance products) are hosted in the directory:</div> <br> <div>- <code>processed_data</code></div> <br> <div><code>D_dark_spectralon_validation</code></div> <div><code>├── processed_data</code></div> <div><code>│ ├── SCT5</code></div> <div><code>│ └── SCT99</code></div> <div><code>├── raw_data</code></div> <div><code>│ ├── SCT5</code></div> <div><code>│ ├── SCT5_dark</code></div> <div><code>│ ├── SCT99</code></div> <div><code>│ └── SCT99_dark</code></div> <div><code>└── README.md</code></div> <br> <div><strong>Raw Data</strong></div> <br> <div>The raw data directory contains images captured of <code>SCT5</code> and <code>SCT99</code> (99% reflectance white Spectralon), and accompanying dark frames, hosted in the <code>SCT5_dark</code> and <code>SCT99_dark</code> frames respectively.</div> <br> <div>For each channel, 25 repeat images have been captured for the illuminated and dark frames.</div> <br> <div><strong>Processed Data</strong></div> <br> <div>The processed SCT99 and SCT5 datasets differ slightly. Both include:</div> <br> <div><code>├── img</code></div> <div><code>├── rfl</code></div> <div><code>└── rois</code></div> <br> <div>directories, with the SCT99 scene also including a <code>cal</code> directory.</div> <br> <div><code>img</code> hosts a set of <code>context</code> figures, showing the regions of interest selected, <code>fits</code> hosts the floating point mean (<code>ave</code>), standard error (<code>err</code>), standard deviation (<code>std</code>) and single-frame (<code>one</code>), all in units of Digital Number, after dark frame subtraction, flat-fielding and linearity correction. <code>uint8</code> hosts the same data rescaled to 8-bit resolution, for quick-view.</div> <br> <div><code>rfl</code> hosts the same set as <code>img</code>, after conversion to units of reflectance against the results of the SCT99 calibration (see §3.5 of the manuscript).</div> <br> <div><code>rois</code> gives plots of the mean and error of the reflectance spectrum of the region of interest, as well as the Signal-to-Noise Ratio, as well as the data for each region-of-interest (<code>roi_data</code>).</div> <br> <div><code>cal</code> also gives context figures for each channel region-of-interest, as converted to units of reflectance coefficients (1/DN/s).</div> </div> </div> </div> </div> </div>
GLIB: image dataset
<p>data/images:</p> <ul> <li><em>data/images/Base</em> : 132 screenshots of game1 & game2 with UI display issues from 466 test reports.</li> <li><em>data/images/Code</em> : 9,412 screenshots of game1 & game2 with UI display issues generated by our Code augmentation method.</li> <li><em>data/images/Normal</em>: 7,750 screenshots of game1 & game2 without UI display issues collected by randomly traversing the game scene.</li> <li><em>data/images/Rule(F)</em> : 7,750 screenshots of game1 & game2 with UI display issues generated by our Rule(F) augmentation method.</li> <li><em>data/images/Rule(R)</em> : 7,750 screenshots of game1 & game2 with UI display issues generated by our Rule(R) augmentation method.</li> <li><em>data/images/testDataSet</em> : 192 screenshots with UI display issues from 466 test reports(exclude game1 & game2).</li> </ul> <p>data/data_csv:</p> <ul> <li><em>data/data_csv/Base</em> : dataset for baseline method.</li> <li><em>data/data_csv/Code</em> : dataset for our Code Augmentation method.</li> <li><em>data/data_csv/Rule(F)</em> : dataset for our Rule(F) Augmentation method.</li> <li><em>data/data_csv/Rule(R)</em> : dataset for our Rule(R) Augmentation method.</li> <li><em>data/data_csv/Code_plus_Rule(F)</em> : dataset for our Code&Rule(F) Augmentation method.</li> <li><em>data/data_csv/Code_plus_Rule(R)</em> : dataset for our Code&Rule(R) Augmentation method.</li> <li><em>data/data_csv/testDataSet</em> : test dataset(normal image and real glitch images from 466 test reports).</li> </ul>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.