Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
105
datasets available to search
ShareScore release 0.9.0
Dataset results
105 results for “Ground truth”
Ground Truth for Alpenwort corpus
<p>The files contain the 7 semantic annotated ground truth documents from the Alpenwort Corpus (SEMOHI project, see http://www.semanticmountain.at), together with gazetteer, IOB-tags (https://spacy.io/api/annotation#iob) and RDF (https://www.w3.org/RDF/) for ground truth.</p>
Ground Truth Data
<div> <div> <div> <div> </div> </div> </div> </div> <div> <div> <div> <div> <div> <div> <p>The repository contains a dataset of labeled RGB and NDVI tree images with health statuses and the related trainined classification model. The dataset images are generated from the original images using geometric transformations from the Albumentations library. The labels folder contains three CSV files, each representing a health status. Health status 1 is labeled as Asymptomatic, 2 as Mild symptoms, and 3 and 4 as Evident symptomatic/compromised. We decided to merge labels 3 and 4 to create a more balanced dataset.</p> <p>The leading number "i_" indicates that the image has been generated through augmentation. This means that the original image has been modified to create this new version. The original unmodified image would be named "DJI_" without the leading number.</p> <p>For example:</p> <p>Original image: DJI_20240525124035_0032_NDVI_0.JPG</p> <p>Augmented images: 1_DJI_20240525124035_0032_NDVI_0.JPG, 2_DJI_20240525124055_0042_RGB_0.JPG, ...</p> <p> </p> <p>The classification model was chosen as the best-performing one among the various created during training. The model is a custom neural network designed to handle both NDVI and RGB images.</p> </div> </div> </div> </div> </div> </div>
Synthetic images of fluorescent spots and ground truth data
<p>Synthetical images of fluorescent spots and ground truth data created with the simcep software.</p>
Simulated data ground truth and results for CALDER
<p>Supporting data for CALDER 2019 publication</p> <ul> <li>branchsim_results.ipynb: code used to load simulated results and generate figures used in paper</li> <li>exome_sim3_groundtruth.pickle: ground truth for simulated datasets, see notebook for usage example</li> <li>exome3_s.pickle: results from methods used for benchmarking on simulated data, see notebook for usage example</li> <li>calder_util.py: supporting functions used in notebook</li> <li>BsimExomeResults.zip: raw result files generated by CALDER, CALDER+pyclone, PhyloWGS, and CITUP on simulated datasets</li> </ul>
Ground truth 3d tetrahedra models
<p>This dataset contains three 3D model files of a tetrahedron, each at a different scale, in .stl format.</p> <p>These models were generated from a 3D model file authored by Anenome at Thingiverse: <a href="https://www.thingiverse.com/thing:942120">https://www.thingiverse.com/thing:942120</a> .</p> <p> </p> <p> </p>
Test dataset with ground truth for segmentation models testing
<p>Test dataset with ground truth for segmentation models testing</p>
Ground Truth for DCASE 2021 Challenge Task 2 Evaluation Dataset
<p><strong>Description</strong></p> <p>This data is the ground truth for the "<a href="https://zenodo.org/record/4884786">evaluation dataset</a>" for the <a href="http://dcase.community/challenge2021/task-unsupervised-detection-of-anomalous-sounds"><strong>DCASE 2021 Challenge Task 2 "Unsupervised Anomalous Sound Detection for Machine Condition Monitoring under Domain Shifted Conditions"</strong></a>. </p> <p>In the task, three datasets have been released: "<a href="http://zenodo.org/record/4562016">development dataset</a>", "<a href="https://zenodo.org/record/4660992">additional training dataset</a>", and "<a href="https://zenodo.org/record/4884786">evaluation dataset</a>". The evaluation dataset was the last of the three released and includes around 200 samples for each machine type, section index, and domain, none of which have a condition label (i.e., normal or anomaly). This ground truth dataset contains the condition labels.</p> <p> </p> <p><strong>Data format</strong></p> <p>The CSV file for each machine type, section index, and domain includes the ground truth data like the following:</p> <p>---------------------------------</p> <p>section_03_source_test_0000.wav,1<br> section_03_source_test_0001.wav,1</p> <p>...</p> <p>section_03_source_test_0198.wav,0<br> section_03_source_test_0199.wav,1</p> <p>---------------------------------</p> <p>The first column shows the name of a wave file. The second column shows the condition label (i.e., 0: normal or 1: anomaly).</p> <p> </p> <p><strong>How to use</strong></p> <p>A script for calculating the AUC, pAUC, precision, recall, and F1 scores for the "evaluation dataset" is available on the Github repository <a href="https://github.com/y-kawagu/dcase2021_task2_evaluator">[URL]</a>. The ground truth data are used by this system. For more information, please see the Github repository.</p> <p> </p> <p><strong>Conditions of use</strong></p> <p>This dataset was created jointly by <strong>Hitachi, Ltd.</strong> and <strong>NTT Corporation</strong> and is available under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0) license.</p> <p> </p> <p><strong>Publication</strong></p> <p>If you use this dataset, please cite <strong>all the following three papers</strong>:</p> <ul> <li>Yohei Kawaguchi, Keisuke Imoto, Yuma Koizumi, Noboru Harada, Daisuke Niizumi, Kota Dohi, Ryo Tanabe, Harsh Purohit, and Takashi Endo, "Description and Discussion on DCASE 2021 Challenge Task 2: Unsupervised Anomalous Sound Detection for Machine Condition Monitoring under Domain Shifted Conditions," in arXiv e-prints: 2106.04492, 2021. [<a href="https://arxiv.org/abs/2106.04492">URL</a>]</li> <li>Noboru Harada, Daisuke Niizumi, Daiki Takeuchi, Yasunori Ohishi, Masahiro Yasuda, Shoichiro Saito, "ToyADMOS2: Another Dataset of Miniature-Machine Operating Sounds for Anomalous Sound Detection under Domain Shift Conditions," in arXiv e-prints: 2106.02369, 2021. [<a href="https://arxiv.org/abs/2106.02369">URL</a>]</li> <li>Ryo Tanabe, Harsh Purohit, Kota Dohi, Takashi Endo, Yuki Nikaido, Toshiki Nakamura, and Yohei Kawaguchi, "MIMII DUE: Sound Dataset for Malfunctioning Industrial Machine Investigation and Inspection with Domain Shifts due to Changes in Operational and Environmental Conditions," in arXiv e-prints: 2105.02702, 2021. [<a href="https://arxiv.org/abs/2105.02702">URL</a>]</li> </ul> <p><br> <strong>Feedback</strong></p> <p>If there is any problem, please contact us:</p> <ul> <li>Yohei Kawaguchi, <a href="mailto:yohei.kawaguchi.xk@hitachi.com">yohei.kawaguchi.xk@hitachi.com</a></li> <li>Daisuke Niizumi, <a href="mailto:daisuke.niizumi.dt@hco.ntt.co.jp">daisuke.niizumi.dt@hco.ntt.co.jp</a></li> <li>Keisuke Imoto, <a href="mailto:keisuke.imoto@ieee.org">keisuke.imoto@ieee.org</a></li> </ul> <p> </p>
Ground truths for mesmerize-core tests.
<p>Ground truths for mesmerize-core tests (https://github.com/nel-lab/mesmerize-core)</p> <p>This version updates the CNMF correlation image to match the motion correction correlation image, for PR #340.</p>
Ground Truth for ÖNB, Cod. 3891
<p>The Ground Truth was produced by the participants of the HTR Winter School 2022 in the Late Latin Group (more information: <a href="https://www.oeaw.ac.at/imafo/veranstaltungen/detail/introduction-into-handwritten-text-recognition">https://www.oeaw.ac.at/imafo/veranstaltungen/detail/introduction-into-handwritten-text-recognition</a>).</p> <p>The Ground Truth includes the following folios: 1-3r, 6-8, 11r, 27 and is still work in progress. We are adding more pages soon. If you find any errors we kindly ask you to contact Jan Odstrčilík (<a href="mailto:jan.odstrcilik@oeaw.ac.at">jan.odstrcilik@oeaw.ac.at</a>).</p> <p>The Supervisors of the Late Latin Group: Jan Odstrčilík PhD, Austrian Acadamy of Sciences, Daniela Mairhofer PhD, Princeton University, Tobias Hodel PhD, University of Bern.</p>
Ground truth annotations for boiling bubble detection and measurement in microgravity
<p>This is a dataset of ground truth annotations for benchmark data provided in A. Sielaff, D. Mangini, O. Kabov, M. Raza, A. Garivalis, M. Zupančič, S. Dehaeck, S. Evgenidis, C. Jacobs, D. Van Hoof, O. Oikonomidou, X. Zabulis, P. Karamaounas, A. Bender, F. Ronshin, M. Schinnerl,</p> <p>J. Sebilleau, C. Colin, P. Di Marco, T. Karapantsios, I. Golobič, A. Rednikov, P. Colinet, P. Stephan, L. Tadrist, The multiscale boiling investigation on-board the international space station:</p> <p>An overview, Applied Thermal Engineering 205 (2022) 117932. doi:10.1016/j.applthermaleng.2021.117932.</p> <p> </p> <p>The annotations regard the 15 image sequences provided in the benchmark data and denoted as D1-D15.</p> <p>The annotators were asked to localize the contact points and points on the bubble boundary so an adequate contour identification is provided, according to the judgement of the expert. The annotators were two multiphase dynamics experts (RO, SE) and one image processing expert (ICS). The annotators used custom-made software to pinpoint samples upon contour locations in the images carefully, using magnification, undo, and editing facilities. The experts annotated the contact points and multiple points on the contour of the bubble until they were satisfied with the result.</p> <p>The annotations were collected for the first bubble of each sequence. For each bubble, 20 frames were sampled in chronological order and in equidistant temporal steps and annotated. All experts annotated data sets D1-D15. The rest were annotated by ICS after learning annotation insights from the multiphase dynamics experts.</p> <p>The format of the dataset is as follows. A directory is dedicated to each bubble annotation. The directory name notes the number of the dataset and the annotator id. Each directory contains 20 text files and 20, corresponding, images. Each text file contains a list with the 2D coordinates of one bubble annotation. The first coordinate marks the left contact point and the last coordinate marks the right contact point. These coordinates refer to a corresponding image contained in the same directory. Text files and image files are corresponded through their file names, which contain the frame number. The frame number refers to the image sequence. Images are in lossless PNG format.</p>
Detection of the Fire Drill anti-pattern: 15 real-world projects with ground truth, issue-tracking data, source code density, models and code
<p>This package contains artifacts for <strong>15</strong> real-world software projects. The data is supposed to aid the detection of the presence of the Fire Drill anti-pattern. We include original data, ground truth, code (experimental setups and models), and notebooks. The data supports two distinct methods of detecting the AP: a) through issue-tracking data, and b) through the underlying source code. This version of the dataset corresponds to <strong>v8</strong> of the <a href="https://arxiv.org/abs/2104.15090v8">technical report</a> and the <a href="https://github.com/MrShoenel/anti-pattern-models/releases/tag/arxiv-v8">GitHub repository</a>. The package includes the following:</p> <p>Original data:</p> <ul> <li>For each project, its <strong>original</strong> artifacts (e.g., wikis, meeting minutes, mentor's notes, etc.)</li> <li>Evaluation of raters' notes by the assessor</li> </ul> <p>Fire Drill in issue-tracking data:</p> <ul> <li><strong>Ground truth</strong> for whether and how strong each project exhibits the Fire Drill AP, on a scale from [0,10]. This was determined by two individual raters, who also reached a consensus.</li> <li>Coefficients for indicators for the first method, per project.</li> <li>Detailed issue-tracing data for each project: what occurred and when.</li> <li>Time logs for each project.</li> </ul> <p>Fire Drill in source-code data:</p> <ul> <li><strong>Four</strong> technical reports that document the developed method of how to translate a description into a detectable pattern, and to use the pattern to detect the presence and to score it (similar to the rating). Also includes a report for how activities were assigned to individual commits.</li> <li>Source code density data (metrics) for each commit in each of the nine projects as a separate dataset.</li> <li>Code: a snapshot of the repository that holds all code, models, notebooks, and pre-computed results, for utmost reproducibility (the code is written in R).</li> </ul>
Person detection using UWB and Monocular camera (With LiDAR ground-truth)
<p>This dataset was record using the ROS2-foxy framework and can be utilized with:</p> <pre><code>ros2 bag play square_test_with_gt</code></pre> <table> <tbody> <tr> <td>Name of ROS2 topic</td> <td>Type of ROS2 topic</td> <td>Information</td> </tr> <tr> <td>/Detections</td> <td>vision_msgs/msg/Detection2DArray</td> <td>This topic includes person detections from the monocular camera that is performing the Deep Learning object detection</td> </tr> <tr> <td>/GT_POINT</td> <td>geometry_msgs/msg/PointStamped</td> <td>Contains the PointStamped message obtained from the LiDAR person detection for ground truth purposes</td> </tr> <tr> <td>/distance_data_array</td> <td>itrci_hardware/msg/RadioRangeDataArray</td> <td>This topic has person detections from the 3 UWB Anchors relative to the person TAG (Note that is in custom ros2 message itrci_hardware)</td> </tr> <tr> <td>/tf</td> <td>tf2_msgs/msg/TFMessage</td> <td>base_link and odom tf (robot is static)</td> </tr> <tr> <td>/tf_static</td> <td>tf2_msgs/msg/TFMessage</td> <td>Contains tf information of LiDAR cameras and anchors relative to the robot base_link</td> </tr> </tbody> </table>
Person detection using UWB and Monocular camera (With LiDAR ground-truth) 0.7m/s
<p>This dataset was record using the ROS2-foxy framework and can be utilized with:</p> <pre><code>ros2 bag play square_test_with_gt_2 </code></pre> <table> <tbody> <tr> <td>Name of ROS2 topic</td> <td>Type of ROS2 topic</td> <td>Information</td> </tr> <tr> <td>/Detections</td> <td>vision_msgs/msg/Detection2DArray</td> <td>This topic includes person detections from the monocular camera that is performing the Deep Learning object detection</td> </tr> <tr> <td>/GT_POINT</td> <td>geometry_msgs/msg/PointStamped</td> <td>Contains the PointStamped message obtained from the LiDAR person detection for ground truth purposes</td> </tr> <tr> <td>/distance_data_array</td> <td>itrci_hardware/msg/RadioRangeDataArray</td> <td>This topic has person detections from the 3 UWB Anchors relative to the person TAG (Note that is in custom ros2 message itrci_hardware)</td> </tr> <tr> <td>/tf</td> <td>tf2_msgs/msg/TFMessage</td> <td>base_link and odom tf (robot is static)</td> </tr> <tr> <td>/tf_static</td> <td>tf2_msgs/msg/TFMessage</td> <td>Contains tf information of LiDAR cameras and anchors relative to the robot base_link</td> </tr> </tbody> </table>
OCR model for lexical lists in Chinese-IPA Glossing, Ground Truth
<p>The ground truth dataset for the OCR model consisted of jpg, pdf, and xml files. The training process was conducted using Transkribus, employing a PyLaia model constructed using ground truth data derived from lexical lists encompassing ten literary works that document Burmish languages. These languages include Achang, Bola, Chashan, Langsu, Leqi, and Zaiwa. The lexical lists utilized for the training phase were predominantly composed in both Chinese characters and International Phonetic Alphabet (IPA) symbols.</p> <p>The training was done on 311 pages and validation on 34 pages of ten lexical lists of Burmish languages on Transkribus with the default PyLaia model:</p> <p>(1) Achang<br> • adapted by Hill & Cooper (2020) from Dai & Cui (1985)<br> (2) Bola<br> • adapted by Hill & Cooper (2020) from He & Chen (2004)<br> (3) Bola<br> • adapted by Hill & Cooper (2020) from Dai et al. (2007)<br> (4) Chashan<br> • adapted by Hill & Cooper (2020) from Dai et al. (2010)<br> (5) Langsu<br> • adapted by Hill & Cooper (2020) from He & Chen (2004)<br> (6) Langsu<br> • adapted by Hill & Cooper (2020) from Dai (2005)<br> (7) Leqi<br> • adapted by Hill & Cooper (2020) from He & Chen (2004)<br> (8) Leqi<br> • adapted by Hill & Cooper (2020) from Dai & Li (2006)<br> (9) Leqi<br> • adapted by Hill & Cooper (2020) from Dai & Jie (2007)<br> (10) Zaiwa<br> • adapted by Hill & Cooper (2020) from Xu & Xu (1984)</p> <p>The material utilized for assessing the performance of the trained models is enclosed within the dataset, comprising a lexical list of the Tujia language as documented by Tian in 1986. The primary objective of the trained model was the recognition of printed lexical lists employing Chinese-IPA glossing.</p>
Ground truth labels for "BRAVE-NET: Fully Automated Arterial Brain Vessel Segmentation In Patients with Cerebrovascular Disease"
<p>Manual, voxel-wise segmentation ground truth labels for 20 healthy volunteers (4 from each age group) from the publicly available MIDAS data collection website under:<br> https://www.insight-journal.org/midas/community/view/21‌</p> <p> </p>
AHPC - JMBD segmentation ground truth
<p>GT segmentation for the JMBD series from AHPC corpus.</p> <p>Images are under request by PARES</p>
Data from: A dataset of stereoscopic images and ground-truth disparity mimicking human fixations in peripersonal space
[No abstract entered]
ground truths for mesmerize-core tests
<p>ground truths for mesmerize-core tests</p>
3D ground truth annotations of cleared whole mouse brain nuclei imaged with a mesoSPIM system
<p>Please see Achard et al. 2024 for Data Card and other information.</p>
animal soup sample data, ground truth dataset, and pre-trained models
<p>- sample data and ground truth files for animal soup tests</p> <p>- pre-trained models</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.