Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
105
datasets available to search
ShareScore release 0.9.0
Dataset results
105 results for “Ground Truth”
Urban material ground truth data for the 2007 HyMap hyperspectral image of Munich
<p><span>This dataset entails a spectral library file (.sli file with matching .hdr text file) with 12028 labeled spectra derived from the 4m resolution airborne hyperspectral HyMap image of Munich (Germany) that was acquired during the summer of 2007 (June 17 and 25 2007). The labeled image spectra are retrieved from pixels of the HyMap dataset that has been processed to level 2A surface reflectance in 119 bands ranging between the wavelengths of 455 nm and 2496 nm. The preprocessing performed on this image data is explained in Heldens et al. (2008) and Heiden et al. (2012). See the "Related works" section of this data publication.</span></p> <p><span>The ground truth (GT) data have been used in previous research (again, see the "Related works" section) and they were likewise used for the remote sensing-based mapping experiments with a generic urban spectral library performed in the frame of the GENLIB research project. The data set contains reflectance spectra of typical urban surface materials and their spectral variations.</span></p> <p><span>The spectra included in this dataset were sampled from the above mentioned HyMap image by (1) using the methodology described in Jilge et al. (2017) and (2) through the delineation of manually digitized regions of interest. The image spectra are <span> </span>labeled based on the method mentioned above and using ancillary reference data, already published urban spectral libraries, terrain knowledge and some field work. The header of the spectral library contains the various labels that were added to the image spectra. These labels cover:</span></p> <ul> <li><span>EAGLE Land Cover Component (LCC) from the EAGLE matrix version 3.1. Visit the </span><span><a href="https://land.copernicus.eu/en/eagle" target="_blank" rel="noopener"><span>website of the EAGLE framework</span></a></span><span> for more information.</span></li> <li><span>Material Groups (MG).</span></li> <li><span>Artificial Material Types (AMT).</span></li> </ul> <p><span>The value domains of these spectrum attributes are described in the look-up table included as a CSV-file in this data publication.</span></p> <p><span>While considerable efforts have been made to safeguard the accuracy of these data, they are published as is, without any warranty or support. Use at your own discretion.</span></p>
Ground Truth and Automated Classification from Copernicus Sentinel-2 Imagery
<p>Ground-Truth and Sentinel2 imagery classification of <em>Trees Outside Forest</em> in an agroforestry landscape in Umbria, Italy.</p> <p>Location: Alfina plains, Castelgiorgio area, Umbria, Italy. Reference system: EPSG:32632 (WGS84, UTM zone 32 North) Extent: West 740609 — East 750828, South 4726490 — North 4737250</p> <p>Dataset format: geopackage, a single file <strong>data.gpkg</strong> containing 9 vector layers (in alphabetical order):</p> <ol> <li>Areas — Areas of interest, 2 polygons</li> <li>Classification — Automated classification from Sentinel2 imagery, 11781 polygons</li> <li>Hedgerows1 — Ground truth, hedgerows of Area1, 148 lines</li> <li>Hedgerows2 — Ground truth, hedgerows of Area2, 135 lines</li> <li>Sentinel2 — Sentinel2 scenes footprint, one polygon</li> <li>Trees1 — Ground truth, isolated trees of Area1, 55 points</li> <li>Trees2 — Ground truth, isolated trees of Area2, 64 points</li> <li>Woods1 — Ground truth, small forest patches of Area1, 33 polygons</li> <li>Woods2 — Ground truth, small forest patches of Area2, 37 polygons</li> </ol> <p>Accompanying map: <strong>map.qgz</strong>, Qgis 3.6 format. The geopackage dataset is supposed to be stored in the same directory of the map (relative path = ./)</p> <p>Dataset description and metadata: <strong>meta.pdf</strong> </p> <p> </p>
Live-cell STED dataset of mitochondria containing ground truth and corresponding low intensity noisy images
<p>The dataset was acquired as part of the manuscript "Denoising diffusion models for high-resolution microscopy image restoration". The dataset contains ground truth and low intensity STED images of mitochondria acquired in live U2-OS cells stably expressing TOM20 coupled to the dead mutant of HaloTag7 which was made fluorescent by using the exchangeable ligand Hy4 bound to the fluorophore SiR. </p>
Urban material ground truth data for the 2015 APEX hyperspectral image of Brussels
<p>This dataset entails a spectral library file (.sli file with matching .hdr text file) with 1350 georeferenced and labeled spectra derived from the 2m resolution airborne hyperspectral APEX image of Brussels (Belgium) that was acquired during the summer of 2015. The labeled spectra included in this dataset describe level 2A surface reflectance profiles ranging between 450 and 2431 nm. The original APEX image files can be downloaded via the <a href="https://belair.vito.be/en/belair-data" target="_blank" rel="noopener">Belair website</a>, and the preprocessing performed on this image data is explained in Sterckx et al. (2016) and Vreys et al. (2016). See the "Related works" section of this data publication.</p> <p>The main purpose of this dataset is to provide Ground Truth (GT) data for remote sensing-based mapping experiments with a generic urban spectral library, performed in the frame of the GENLIB research project. The content of this dataset hence focuses on the optical reflectance/absorption behaviour of urban surface materials and their variations.</p> <p>The spectra included in this dataset were manually sampled from the above mentioned APEX image and labeled using ancillary reference data (very high-resolution aerial imagery, Google Street View, LiDAR ...), already published urban spectral libraries, terrain knowledge and some field work. The header of the spectral library contains the various labels that were added to these spectra. These labels cover:</p> <ul> <li>EAGLE Land Cover Component (LCC) from the EAGLE matrix version 3.1. Visit the <a href="https://land.copernicus.eu/en/eagle" target="_blank" rel="noopener">website of the EAGLE framework</a> for more information.</li> <li>Material Groups (MG).</li> <li>Artificial Material Types (AMT).</li> <li>Artificial Material Coating or Fabrication (AMCF).</li> <li>Artificial Material Forms (AMF).</li> <li>Latitude (degrees, WGS84).</li> <li>Longitude (degrees, WGS84).</li> </ul> <p>The value domains of these spectrum attributes are described in the look-up table included as a CSV-file in this data publication.</p> <p>While considerable efforts have been made to safeguard the accuracy of these data, they are published as is, without any warranty or support. Use at your own discretion.</p>
Audio Commons Ground Truth Data for deliverables D4.4, D4.10 and D4.12
<p>This dataset contains the ground truth data used to evaluate the musical <strong>pitch</strong>, <strong>tempo</strong> and <strong>key </strong>estimation algorithms developed during the AudioCommons H2020 EU project and which are part of the <a href="https://www.audiocommons.org/2018/07/15/audio-commons-audio-extractor.html">Audio Commons Audio Extractor tool</a>. It also includes ground truth information for the <strong>single-event<em>ness</em> </strong>audio descriptor also developed for the same tool.</p> <p>This ground truth data has been used to generate the following documents:</p> <ul> <li><strong>Deliverable D4.4</strong>: Evaluation report on the first prototype tool for the automatic semantic description of music samples</li> <li><strong>Deliverable D4.10</strong>: Evaluation report on the second prototype tool for the automatic semantic description of music samples</li> <li><strong>Deliverable D4.12</strong>: Release of tool for the automatic semantic description of music samples</li> </ul> <p>All these documents are available in the <a href="https://www.audiocommons.org/materials/">materials section </a>of the AudioCommons website.</p> <p>All ground truth data in this repository is provided in the form of CSV files. Each CSV file corresponds to one of the individual datasets used in one or more evaluation tasks of the aforementioned deliverables. This repository <strong>does not include the audio files</strong> of each individual dataset, but includes references to the audio files. The following paragraphs describe the structure of the CSV files and give some notes about how to obtain the audio files in case these would be needed.</p> <p><br> <strong>Structure of the CSV files</strong></p> <p>All CSV files in this repository (with the sole exception of <em>SINGLE EVENT - Ground Truth.csv</em>) feature the following 5 columns:</p> <ol> <li><strong>Audio reference</strong>: reference to the corresponding audio file. This will either be a string withe the <strong>filename</strong>, or the <strong>Freesound ID </strong>(for one dataset based on Freesound content). See below for details about how to obtain those files. </li> <li><strong>Audio reference type</strong>: will be one of <em>Filename</em> or <em>Freesound ID</em>, and specifies how the previous column should be interpreted. </li> <li><strong>Key annotation</strong>: tonality information as a string with the form "RootNote minor/major". Audio files with no ground truth annotation for tonality are left blank. Ground truth annotations are parsed from the original data source as described in the text of deliverables D4.4 and D4.10.</li> <li><strong>Tempo annotation</strong>: tempo information as an integer representing beats per minute. Audio files with no ground truth annotation for tempo are left blank. Ground truth annotations are parsed from the original data source as described in the text of deliverables D4.4 and D4.10. Note that integer values are used here because we only have tempo annotations for <em>music loops</em> which typically only feature integer tempo values.</li> <li><strong>Pitch annotation</strong>: pitch information as an integer representing the MIDI note number corresponding to annotated pitch's frequency. Audio files with no ground truth pitch for tempo are left blank. Ground truth annotations are parsed from the original data source as described in the text of deliverables D4.4 and D4.10.</li> </ol> <p>The remaining CSV file, <em>SINGLE EVENT - Ground Truth.csv</em>, has only the following 2 columns:</p> <ul> <li><strong>Freesound ID</strong>: sound ID used in Freesound to identify the audio clip.</li> <li><strong>Single Event: </strong>boolean indicating whether the corresponding sound is considered to be a single event or not. Single event annotations were collected by the authors of the deliverables as described in deliverable D4.10.</li> </ul> <p> </p> <p><strong>How to get the audio data</strong></p> <p>In this section we provide some notes about how to obtain the audio files corresponding to the ground truth annotations provided here. Note that due to licensing restrictions we are not allowed to re-distribute the audio data corresponding to most of these ground truth annotations.</p> <ul> <li><strong>Apple Loops (APPL)</strong>: This dataset includes some of the music loops included in Apple's music software such as Logic or GarageBand. Access to these loops requires owning a license for the software. Detailed instructions about how to set up this dataset are <a href="https://github.com/ffont/ismir2016/blob/master/docs/create_dataset.md#appl">provided here</a>. </li> <li><strong>Carlos Vaquero Instruments Dataset (CVAQ)</strong>: This dataset includes single instrument recordings carried out by <a href="https://www.linkedin.com/in/carlosvaquero/">Carlos Vaquero</a> as part of this <a href="http://mtg.upf.edu/node/2609">master thesis</a>. Sounds are available as Freesound packs and can be downloaded at this page: https://freesound.org/people/Carlos_Vaquero/packs</li> <li><strong>Freesound Loops 4k (FSL4)</strong>: This dataset set includes a selection of music loops taken from Freesound. Detailed instructions about how to set up this dataset are <a href="https://github.com/ffont/ismir2016/blob/master/docs/create_dataset.md#instructions-for-setting-up-datasets">provided here</a>.</li> <li><strong>Giant Steps Key Dataset (GSKY)</strong>: This dataset includes a selection of previews from Beatport annotated by key. Audio and original annotations <a href="https://github.com/GiantSteps/giantsteps-key-dataset">available here</a>.</li> <li><strong>Good-sounds Dataset (GSND)</strong>: This dataset contains monophonic recordings of instrument samples. Full description, original annotations and audio are <a href="https://zenodo.org/record/820937#.XEYMiy2ZN25">available here</a>.</li> <li><strong>University of IOWA Musical Instrument Samples (IOWA)</strong>: This dataset was created by the Electronic Music Studios of the University of IOWA and contains recordings of instrument samples. The dataset is available upon request by <a href="http://theremin.music.uiowa.edu/MIS.html">visiting this website</a>.</li> <li><strong>Mixcraft Loops (MIXL)</strong>: This dataset includes some of the music loops included in Acoustica's Mixcraft music software. Access to these loops requires owning a license for the software. Detailed instructions about how to set up this dataset are <a href="https://github.com/ffont/ismir2016/blob/master/docs/create_dataset.md#mixl">provided here</a>.</li> <li><strong>NSynth Dataset Test and Validation sets (NSYT and NSYV)</strong>: NSynth is a large-scale and high-quality dataset of annotated musical notes built with synthesized sounds by Google's Magenta team. Full dataset description including original annotations and audio files is <a href="https://magenta.tensorflow.org/datasets/nsynth">available here</a>.</li> <li><strong>Philarmonia Orchestra Sound Samples Dataset (PHIL)</strong>: This includes thousands of free, downloadable sound samples specially recorded by Philharmonia Orchestra players. Audio files are freely downloadable from the <a href="http://www.philharmonia.co.uk/explore/sound_samples">philarmonia orchestra website</a>.</li> <li><strong>Freesound Single Events Dataset (SINGLE EVENT)</strong>: This includes a selection of Freesound audio clips representing audio signals containing either a single audio <em>event</em> or multiple ones. Original audio files can be retrieved by downloading individual audio clips from Freesound using the ID identifier provided in the CSV file. A similar procedure to that described <a href="https://github.com/ffont/ismir2016/blob/master/docs/create_dataset.md#getting-fsl4-by-downloading-content-from-freesound">here</a> could be followed.</li> </ul>
The VAROS Synthetic Underwater Data Set: Towards realistic multi-sensor underwater data with ground truth
<p>Underwater visual perception requires being able to deal with bad and rapidly varying illumination and with reduced visibility due to water turbidity. The verification of such algorithms is crucial for safe and efficient underwater exploration and intervention operations. Ground truth data play an important role in evaluating vision algorithms. However, obtaining ground truth from real underwater environments is in general very hard, if possible at all. In a synthetic underwater 3D environment, however, (nearly) all parameters are known and controllable, and ground truth data can be absolutely accurate in terms of geometry. In this paper, we present the VAROS environment, our approach to generating highly realistic underwater video and auxiliary sensor data with precise ground truth, built around the Blender modeling and rendering environment. VAROS allows for physically realistic motion of the simulated underwater (UW) vehicle including moving illumination. Pose sequences are created by first defining way-points for the simulated underwater vehicle which are expanded into a smooth vehicle course sampled at IMU data rate (200Hz). This expansion uses a vehicle dynamics model and a discrete-time controller algorithm that simulates the sequential following of the way-points. The scenes are rendered using the raytracing method, which generates realistic images, integrating direct light, and indirect volumetric scattering. The VAROS dataset version 1 provides images, inertial measurement unit (IMU) and depth gauge data, as well as ground truth poses, depth images and surface normal images.</p>
Ground-truthing of satellite imagery to track harmful algal blooms in Pigeon Lake, Alberta, Canada 2017-2022
This data was collected to create a calibrated model that would enable the use of satellite imagery to track harmful algal blooms by using chlorophyll a estimates as a proxy for cyanobacteria in the lake. Samples from Pigeon Lake were collected on the same day that the Sentinel-2 satellite would pass over the lake. These samples were analyzed for different algal pigments and enumerated to genus level to ensure that the satellite imagery was of cyanobacteria rather than different algal groups. An algorithm was developed which we termed the three band index (TBI) that best matched with the cholorophyll a from the in situ samples. This model was used on satellite imagery from 2017-2022 of Pigeon Lake to get chlorophyll a estimates for every 20 x 20 pixel of each image of the lake. This pixel data was used to determine different bloom metrics like the intensity, the area (extent) and severity.
Vegetation and invertebrate communities in 500 plots in the Duplin and Dean Creek watersheds: ground truth data for matching hyperspectral imagery
We measured characteristics of vegetation (Aster tenuifolius, Batis maritima, Borrichia frutescens, Distichlis spicata, Iva frutescens, Juncus roemerianus, Limonium carolinianum, Salicornia biglovii, Salicornia virginica, Spartina alterniflora, Spartina patens, Sporobolus virginicus), soil (salinity, proportion organic and proportion water) and densities of common gastropods and bivalves in 500 plots in the Duplin and Dean Creek watersheds on Sapelo Island on June 20-26, 2006. Plot locations were determined using a high precision hand-held GPS. These data were used to help ground-truth hyperspectral aerial images collected at the same time by Dr. John Schalles.
Ground Truth for DCASE 2020 Challenge Task 2 Evaluation Dataset
<p><strong>Description</strong></p> <p>This data is the ground truth for the "<a href="https://zenodo.org/record/3841772">evaluation dataset</a>" for the <strong>DCASE 2020 Challenge Task 2 "Unsupervised Detection of Anomalous Sounds for Machine Condition Monitoring" </strong><a href="http://dcase.community/challenge2020/task-unsupervised-detection-of-anomalous-sounds">[task description]</a>. </p> <p>In the task, three datasets have been released: "<a href="http://zenodo.org/record/3678171">development dataset</a>", "<a href="https://zenodo.org/record/3727685">additional training dataset</a>", and "<a href="https://zenodo.org/record/3841772">evaluation dataset</a>". The evaluation dataset was the last of the three released and includes around 400 samples for each Machine Type and Machine ID used in the evaluation dataset, none of which have any condition label (i.e., normal or anomaly). This ground truth data contains the condition labels.</p> <p> </p> <p><strong>Data format</strong></p> <p>The ground truth data is a CSV file like the following:</p> <p>---------------------------------</p> <p>fan<br> id_01_00000000.wav,normal_id_01_00000098.wav,0<br> id_01_00000001.wav,anomaly_id_01_00000064.wav,1<br> ...</p> <p>id_05_00000456.wav,anomaly_id_05_00000033.wav,1<br> id_05_00000457.wav,normal_id_05_00000049.wav,0<br> pump<br> id_01_00000000.wav,anomaly_id_01_00000049.wav,1<br> id_01_00000001.wav,anomaly_id_01_00000039.wav,1<br> ...</p> <p>id_05_00000346.wav,anomaly_id_05_00000052.wav,1<br> id_05_00000347.wav,anomaly_id_05_00000080.wav,1<br> slider<br> id_01_00000000.wav,anomaly_id_01_00000035.wav,1<br> id_01_00000001.wav,anomaly_id_01_00000176.wav,1<br> ...</p> <p>---------------------------------</p> <p>"Fan", "pump", "slider", etc mean "Machine Type" names. The lines following a Machine Type correspond to pairs of a wave file in the Machine Type and a condition label. The first column shows the name of a wave file. The second column shows the original name of the wave file, but this can be ignored by users. The third column shows the condition label (i.e., 0: normal or 1: anomaly).</p> <p> </p> <p><strong>How to use</strong></p> <p>A system for calculating AUC and pAUC scores for the "evaluation dataset" is available on the Github repository <a href="https://github.com/y-kawagu/dcase2020_task2_evaluator">[URL]</a>. The ground truth data is used by this system. For more information, please see the Github repository.</p> <p> </p> <p><strong>Conditions of use</strong></p> <p>This dataset was created jointly by <strong>NTT Corporation</strong> and <strong>Hitachi, Ltd.</strong> and is available under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0) license.</p> <p> </p> <p><strong>Publication</strong></p> <p>If you use this dataset, please cite <strong>all the following three papers</strong>:</p> <p>Yuma Koizumi, Shoichiro Saito, Noboru Harada, Hisashi Uematsu, and Keisuke Imoto, "ToyADMOS: A Dataset of Miniature-Machine Operating Sounds for Anomalous Sound Detection," in Proc. of IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA), 2019. <a href="https://ieeexplore.ieee.org/document/8937164">[pdf]</a></p> <p>Harsh Purohit, Ryo Tanabe, Kenji Ichige, Takashi Endo, Yuki Nikaido, Kaori Suefusa, and Yohei Kawaguchi, “MIMII Dataset: Sound Dataset for Malfunctioning Industrial Machine Investigation and Inspection,” in Proc. 4th Workshop on Detection and Classification of Acoustic Scenes and Events (DCASE), 2019. <a href="http://dcase.community/documents/workshop2019/proceedings/DCASE2019Workshop_Purohit_21.pdf">[pdf]</a></p> <p>Yuma Koizumi, Yohei Kawaguchi, Keisuke Imoto, Toshiki Nakamura, Yuki Nikaido, Ryo Tanabe, Harsh Purohit, Kaori Suefusa, Takashi Endo, Masahiro Yasuda, and Noboru Harada, "Description and Discussion on DCASE2020 Challenge Task2: Unsupervised Anomalous Sound Detection for Machine Condition Monitoring<em>,"</em> in Proc. 5th Workshop on Detection and Classification of Acoustic Scenes and Events (DCASE), 2020. <a href="https://dcase.community/documents/workshop2020/proceedings/DCASE2020Workshop_Koizumi_3.pdf">[pdf]</a></p> <p><br> <strong>Feedback</strong></p> <p>If there is any problem, please contact us:</p> <ul> <li>Yuma Koizumi, <a href="mailto:koizumi.yuma@ieee.org">koizumi.yuma@ieee.org</a></li> <li>Yohei Kawaguchi, <a href="mailto:yohei.kawaguchi.xk@hitachi.com">yohei.kawaguchi.xk@hitachi.com</a></li> <li>Keisuke Imoto, <a href="mailto:keisuke.imoto@ieee.org">keisuke.imoto@ieee.org</a></li> </ul>
Fluorescently-labelled zebrafish pronephroi + ground truth classes (normal/cystic) + trained CNN model
<p>This upload contains :</p> <p>- <strong>images.zip: </strong> microscope images of fluorescently-labelled pronephroi in larvae of the <em>Tg(wt1b:EGFP)</em> transgenic zebrafish line showing 2 morphologies (normal vs cystic) upon injection with Co-Mo or ift172-MO, respectively. Images were obtained using an ACQUIFER Imaging Machine widefield high content screening microscope.</p> <p>Reference: </p> <p>Pandey, G., Westhoff, J., Schaefer, F. and Gehrig, J. (2019). <strong>A Smart Imaging Workflow for Organ-Specific Screening in a Cystic Kidney Zebrafish Disease Model</strong>. International Journal of Molecular Sciences <em>20</em>, 1290, doi:<a href="https://doi.org/10.3390/ijms20061290">10.3390/ijms20061290</a>.</p> <p> </p> <p>- <strong>Annotations-***.csv : </strong>Tables containing ground-truth category classes (normal vs cystic) for the images in the zip file.</p> <p>The tables contain columns with the image filename, folder and category.</p> <p>Note : <strong>the Folder column should be updated with the root folder directory once downloaded on your machine.</strong></p> <p>These files were generated with the Fiji plugin <em>single-class (button)</em> from the <em>Qualitative-Annotations</em> update site.</p> <p>The 2 files contain the same information, they only differ in the formatting of the category, the <em>singleColumn </em>file has a single category column while the <em>multiColumn</em> has 2 columns (normal/cystic) with 0/1 encoding.</p> <p>The choice of category encoding solely depends on how the table is used, i.e. in which training workflow, home-made script or software.</p> <p>- <strong>trainedModel.zip : </strong>This archive contains 2 files: <strong>(1) </strong>a h5 file corresponding to a trained deep-learning model to classify the images of the dataset in the 2 categories (normal vs cystic), and <strong>(2)</strong> a text file containing the class names. Both files are necessary to predict the category of new images similar to the one in the dataset, for instance using the published KNIME workflows.</p>
A Public Ground-Truth Dataset for Handwritten Circuit Diagram Images
<p><strong>CGHD</strong></p> <p>This dataset contains images of hand-drawn electrical circuit diagrams as well as accompanying annotation and segmentation ground-truth files. It is intended to train (e.g. ANN) models for extracting electrical graphs from raster graphics.</p> <p><strong>Content</strong></p> <ul> <li><strong>3.269</strong> Annotated Raw Images<br> <ul> <li>31 Main Drafters <ul> <li>12 Circuits per Drafter</li> <li>2 Drawings per Circuit</li> <li>4 Photos per Drawing</li> </ul> </li> <li>Additional Circuit Images provided by TU Dresden (from Real-World Examinations, Drafter 0)</li> <li>Additional Circuit Images provided by RPTU Kaiserslautern-Landau (Drafter -1)</li> <li><strong>248.020 </strong>Bounding Box Annotations</li> <li><strong>40.711</strong> Rotation Annotations</li> <li><strong>1.437</strong> Mirror Annotations</li> <li><strong>85.417</strong> Text String Annotations (equals <strong>93.74%</strong> completeness)<br> <ul> <li><strong>289.850</strong> Text Characters</li> <li><strong>98</strong> Character Types (Upper/Lower Case Latin, Numbers, Special Characters)</li> </ul> </li> </ul> </li> <li><strong>320</strong> Binary Segmentation Maps<br> <ul> <li>Strokes vs. Background</li> <li>Accompanying Polygon Annotation Files</li> <li><strong>22.929</strong> Polygon Annotations</li> </ul> </li> <li><strong>59 </strong>Object Classes</li> <li><strong>Scripts</strong> for Data Loading, Statistics, Consistency Check and Training Preparation</li> </ul>
Ground truth and raw hyperspectral files of olive trees for plant stress detection
<p>This dataset contains raw hyperspectral images from Cubert S-185 collected on 13 May 2021 from an olive field in Halkidiki, Northern Greece. Included is also a matrix containing the id of each recorded olive tree (the samples) that also appears in the hyperspectral images. QGIS (ver.3.28.0) software plugin 'zonal statistics multiband' was used to compute zonal statistics for each of the 138 spectral bands available for each sample. Accompanying each sample is also the ground truthing data recorded, which addresses the present stress of 3 stressors (<i>Verticillium dahliae, Pleospora herbarum </i>and 'other stressors').</p>
Ground-Truthed Data Set of Zenon Papyri for Handwritten Text Recognition
<p>Diplomatic transcription of papyri found in the Zenon archive [see <a href="https://en.wikipedia.org/wiki/Zenon_of_Kaunos">en.wikipedia.org/wiki/Zenon_of_Kaunos</a>]</p> <p>Manually prepared as PageXML with Transkribus within <a href="http://d-scribes.philhist.unibas.ch/">D-Scribes</a> project.</p> <p> </p>
UAV-based monocular SLAM video datasets in vineyards with RTK ground truth
<p>The dataset provides a UAV-based monocular visual SLAM data, designed to evaluate the potential of using monocular visual SLAM in vineyards. It includes videos in ".mp4" format collected by UAV, and "xlsx" tables which include latitude, longitude, height, speed in x, y and z, comjpass, pitch, roll. The ".xlsx" tables were measured by RTK and can be used as ground truth of UAV trajectory and pose.</p> <p>This dataset can be combined with other datasets to enable a comprehensive view of the vineyards:</p> <p>Vélez S, Ariza-Sentís M, Valente J. EscaYard: Precision viticulture multimodal dataset of vineyards affected by Esca disease consisting of geotagged smartphone images, phytosanitary status, UAV 3D point clouds and Orthomosaics. Data in Brief. 2024 Jun 1;54:110497. <a title="Persistent link using digital object identifier" href="https://doi.org/10.1016/j.dib.2024.110497" target="_blank" rel="noreferrer noopener"><span>https://doi.org/10.1016/j.dib.2024.110497</span></a></p> <p><span>Ariza-Sentís M, Wang K, Cao Z, Vélez S, Valente J. GrapeMOTS: UAV vineyard dataset with MOTS grape bunch annotations recorded from multiple perspectives for enhanced object detection and tracking. Data in Brief. 2024 Jun 1;54:110432. <a title="Persistent link using digital object identifier" href="https://doi.org/10.1016/j.dib.2024.110432" target="_blank" rel="noreferrer noopener">https://doi.org/10.1016/j.dib.2024.110432</a></span></p> <p> </p> <p> </p>
Ground truth recordings for validation of spike sorting algorithms
<p><strong>Ground-truth recordings for validation of spike sorting algorithms</strong><br> </p> <p>This datasets is composed of simultaneous loose patch recordings of Ganglion Cells in mice retina, combined with dense extra-cellular recordings (252 channels). The details of the dataset can be found here <a href="https://elifesciences.org/articles/34518">https://elifesciences.org/articles/34518</a></p> <p><strong>Probe layout</strong></p> <p>The probe layout can be found as mea_256.prb. This is a 16x16 Multi Electrode Array with 30um spacing. Only 252 channels are extra-cellular signals, and the 4 corners are devoted to triggers/sync/juxta.</p> <p><strong>Struture of the data</strong></p> <p>In this dataset, you will find several individual recordings, at max 5min long each (but please do not hesitate to contact us if interested by longer recordings). The extra-cellular data are saved as 16bits unsigned integer, with a variable offset at the beginning of the file. The value of this offset is given, for every datafile, in the additional text file (padding value (see following for more details)). The files have already been filtered with a Butterworth filter of order 3 with a cut-off frequency at 100Hz</p> <p><strong>Structure of a given dataset</strong></p> <p>Please read carefully the following to understand how to load and perform spike sorting with the data. In every .tar.gz file, you will find:</p> <ul> <li> a jpg image, displaying a small chunk of the juxta-cellular signal (top left), with detected peaks and threshold. The extra-cellular spike triggered waveform, across all channels, for the juxta-spike times (top right). In the bottom, you can see the juxta-cellular spikes, for all the detected triggers (left), and on the right the voltage on the channel where the Spike Triggered Average of the extra-cellular waveform is peaking the most.</li> <li>a file .juxta.raw, as float32, with the juxta-cellular trace at 20kHz, no data offset</li> <li>a file .raw, as uint16, with the extra-cellular signals recorded for 256 channels at a sampling rate of 20kHZ. In fact, only 252 channels are extra-cellular signals, the 4 corners of the arrays are devoted to juxta-cellular and sync signals (see probe layout mea_256.prb)</li> <li>a file .triggers.npy containing the spike times of the juxta-cellular spikes, detected using a threshold of k.MAD. The exact value of k can vary on a per dataset basis, and is written in the .txt file (threshold)</li> <li>a .txt file describing some information for a given dataset, such as the threshold value used to detect the spikes, the channel in the raw file where the juxta-cellular signal is located, the minimal value of the peak for the STA (and on which channel it is located), and the header size to read the raw data</li> <li>a .params file, if you want to analyze the data with SpyKING CIRCUS</li> </ul> <p><strong>How to load the raw data in numpy</strong></p> <pre><code class="language-python">#Using the offset value from the txt file, we can load the data with memmap arrays data=numpy.memmap('mydata.raw', dtype='uint16', offset=offset, mode='r') data=data.reshape(len(data)//256, 256) #Then for example, to display the first second of channel 0 one_channel = data[:20000, 0].astype('float32') #If we want to center data around 0 one_channel -= 2**15 - 1 #And if we want to display data in micro volt, we must use the gain factor of 0.1042 provided in the header one_channel *= 0.1042</code></pre> <p> </p>
MB2017: Artificial spiking neural data with ground truth
<p>This dataset comprises simulated extracellular spiking neural signals, for which the activity of the active neurons is known. These data can thus be used to test spike sorting algorithms.</p> <p>This dataset has been generated for the study reported in:</p> <p>Bernert M, Yvert B (2018) An attention-based spiking neural network for unsupervised spike-sorting. International Journal of Neural Systems, https://doi.org/10.1142/S0129065718500594</p> <p>Please cite this paper as a reference.</p>
Simulated dMRI images and ground truth of random fiber phantoms in various configurations
<p>This archive contains simulated dMRI images of random fiber phantoms in various configurations created with Fiberfox and other tools available in MITK Diffusion (<a href="http://mitk.org/wiki/DiffusionImaging">http://mitk.org/wiki/DiffusionImaging</a>). RandomFibers_Example.png illustrates one of the random fiber configurations used for these phantoms.</p> <p>If you are using any of these datasets or the tools used to generate them, please don't forget to cite the dataset itself as well as other relevant publications.</p> <p>Each subfolder contains the following elements:<br> The simulated dMRI image with b-values and gradient directions: dwi.nii.gz, dwi.bvals, dwi.bvecs<br> The fibers used for simulation: AllBundles.fib (binary vtk format)<br> parameters.ffp: Fiberfox simulation parameters<br> parameters.ffp.bvals: b-value file for Fiberfox simulation<br> parameters.ffp.bvecs: gradient vector file for Fiberfox simulation<br> parameters.ffp_VOLUME1.nii.gz: fiber compartment volume fraction map for Fiberfox simulation<br> The logfile detailing all steps of the generation process of the respective phantom: LOGFILE.json</p> <p>bundles: folder containing the individual fiber bundles (binary vtk format .fib)<br> centroids: folder containing the centerlines of each bundle<br> masks: folder containing the binary envelope of each bundle<br> peaks: folder containing the principal fiber direction image (peaks) of each bundle</p> <p>Each subfolder contains the fibers and dMRI simulations with the following fiber specifications:<br> Phantom 1:<br> - Number of bundles: 25<br> - Fiber density: 250 streamlines per cm²<br> - Bundle curvature: 0-30 in degree<br> - Bundle start radius: 5-15 in mm</p> <p>Phantom 2:<br> - Number of bundles: 25<br> - Fiber density: 250 streamlines per cm²<br> - Bundle curvature: 0-30 in degree<br> - Bundle start radius: 15-30 in mm</p> <p>Phantom 3:<br> - Number of bundles: 25<br> - Fiber density: 250 streamlines per cm²<br> - Bundle curvature: 30-60 in degree<br> - Bundle start radius: 5-15 in mm</p> <p>Phantom 4:<br> - Number of bundles: 25<br> - Fiber density: 250 streamlines per cm²<br> - Bundle curvature: 30-60 in degree<br> - Bundle start radius: 15-30 in mm</p> <p>Phantom 5:<br> - Number of bundles: 25<br> - Fiber density: 50-500 streamlines per cm²<br> - Bundle curvature: 0-30 in degree<br> - Bundle start radius: 5-15 in mm</p> <p>Phantom 6:<br> - Number of bundles: 25<br> - Fiber density: 50-500 streamlines per cm²<br> - Bundle curvature: 0-30 in degree<br> - Bundle start radius: 15-30 in mm</p> <p>Phantom 7:<br> - Number of bundles: 25<br> - Fiber density: 50-500 streamlines per cm²<br> - Bundle curvature: 30-60 in degree<br> - Bundle start radius: 5-15 in mm</p> <p>Phantom 8:<br> - Number of bundles: 25<br> - Fiber density: 50-500 streamlines per cm²<br> - Bundle curvature: 30-60 in degree<br> - Bundle start radius: 15-30 in mm</p> <p>Phantom 9:<br> - Number of bundles: 50<br> - Fiber density: 250 streamlines per cm²<br> - Bundle curvature: 0-30 in degree<br> - Bundle start radius: 5-15 in mm</p> <p>Phantom 10:<br> - Number of bundles: 50<br> - Fiber density: 250 streamlines per cm²<br> - Bundle curvature: 0-30 in degree<br> - Bundle start radius: 15-30 in mm</p> <p>Phantom 11:<br> - Number of bundles: 50<br> - Fiber density: 250 streamlines per cm²<br> - Bundle curvature: 30-60 in degree<br> - Bundle start radius: 5-15 in mm</p> <p>Phantom 12:<br> - Number of bundles: 50<br> - Fiber density: 250 streamlines per cm²<br> - Bundle curvature: 30-60 in degree<br> - Bundle start radius: 15-30 in mm</p> <p>Phantom 13:<br> - Number of bundles: 50<br> - Fiber density: 50-500 streamlines per cm²<br> - Bundle curvature: 0-30 in degree<br> - Bundle start radius: 5-15 in mm</p> <p>Phantom 14:<br> - Number of bundles: 50<br> - Fiber density: 50-500 streamlines per cm²<br> - Bundle curvature: 0-30 in degree<br> - Bundle start radius: 15-30 in mm</p> <p>Phantom 15:<br> - Number of bundles: 50<br> - Fiber density: 50-500 streamlines per cm²<br> - Bundle curvature: 30-60 in degree<br> - Bundle start radius: 5-15 in mm</p> <p>Phantom 16:<br> - Number of bundles: 50<br> - Fiber density: 50-500 streamlines per cm²<br> - Bundle curvature: 30-60 in degree<br> - Bundle start radius: 15-30 in mm</p>
Layout Ground Truth for Historical Commentaries
<p>This release contains the public domain portion of the dataset used in the paper <em>Page Layout Analysis of Text-heavy Historical Documents: a Comparison of Textual and Visual Approaches</em>.</p>
The e-NDP project : collaborative digital edition of the Chapter registers of Notre-Dame of Paris (1326-1504). Ground-truth for handwriting text recognition (HTR) on late medieval manuscripts.
<p>The <a href="https://endp.hypotheses.org/">e-NDP project</a>, funded by the ANR, is led by the <a href="https://lamop.hypotheses.org/6870">LaMOP</a> (Julie Claustre and Darwin Smith).</p> <p>The project's partners are the Archives nationales, the Bibliothèque nationale de France (Department of Manuscripts, Bibliothèque de l'Arsenal), the École nationale des chartes and the Bibliothèque Mazarine.</p> <p>The e-NDP project aims at renewing our knowledge on <strong>Notre-Dame de Paris cathedral</strong> through the creation of a collaborative digital edition of the registers of its Chapter (1326-1504, <em>AN LL 105-128</em>), the community of 51 canons meeting three times a week on set days to take all administrative, financial and practical decisions pertaining to the cathedral, its estate and the society living in its cloister. This corpus has never been the object of a comprehensive study to understand the workings and history of this urban enclave and powerful community. The collaborative digital edition is based on a process of<strong> handwriting text recognition (HTR)</strong>, tested and supervised by scholars, researchers and engineers combining expertise in Medieval history, paleography, philology and digital humanities. The edition shall allow a better insight into the Chapter’s administration, into its economical and political power within Paris, and the relationships it maintained with other institutions in the city.</p> <p> </p> <p><strong>Section 1 : The e-NDP ground-truth dataset for Handwriting text recognition.</strong></p> <p>The full e-NDP corpus kept today in the French National Archives and was entirely digitized and described in its <a href="https://www.siv.archives-nationales.culture.gouv.fr/siv/rechercheconsultation/consultation/ir/consultationIR.action?formCaller=GENERALISTE&irId=FRAN_IR_059635">catalog</a> in 2022.</p> <p>The first major goal of the e-NDP projet is to propose a first automatic transcription of the 14k pages composing the 26 chapter registers. To achieve this goal representative samples from each one of the volumes were selected and transcribed in order to train a specialized HTR model able to propose a high quality automatic transcription. The collected ground-truth released on this repository currently has <strong>512 pages from the 26 registers</strong> of the cathedral chapter preserved in the National Archives (LL105 - LL128, <strong>1326-1504</strong>). The transcriptions were manually completed in <strong>two rounds</strong> by a group of 12 contributors, historians and paleographers, over the course of 2021-2022 using <a href="https://escriptorium.paris.inria.fr/">eScriptorium </a>as annotation environment. </p> <p> </p> <p><strong>Ground-truth features :</strong></p> <p><br> <em>Number of hands </em>: according to our estimates no fewer than 18 main hands were involved in the writing of the registers during the medieval period. </p> <p><em>Language</em> : More than 98% of the content of the registers was written in Latin, the rest in French. The exact percentage is hard to estimate because the vernacular language is often used in formulae, notes and comments. It is rare to find entire pages or blocks written in French. </p> <p><em>Script family</em> : The registers were written using a Cursive script (ca. late XIIIe - XVIe).</p> <p><em>Documental typology</em> : The volumes containing the chapter conclusions were conceived to serve as memorial records, but above all as documents for regular use and consultation in the daily practice of administration and management. In diplomatics the notion of "documentary manuscripts" is used to describe this kind of sources also by opposition to books and litterary or normative manuscripts.</p> <table align="center"> <caption><strong>Ground truth statistics</strong></caption> <tbody> <tr> <th>Text units</th> <th>Count</th> </tr> <tr> <td>Pages</td> <td>512</td> </tr> <tr> <td>Annotated regions (see section 2)</td> <td>2448</td> </tr> <tr> <td>Lines of text</td> <td>34231</td> </tr> <tr> <td>Tokens</td> <td>205083</td> </tr> <tr> <td>Characters</td> <td>3320407</td> </tr> </tbody> </table> <p> </p> <p><strong>Rules of transcription :</strong></p> <ul> <li>The abbreviations have been resolved, both those by suspension (<code>facimꝰ</code> ---> <code>facimus</code>) and by contraction (<code>dñi</code> --> <code>domini</code>). Likewise, those using conventional signs (<code>⁊</code> --> <code>et</code> ; <code>ꝓ</code> --> <code>pro</code>) have been resolved. </li> <li>The named entities (names of persons, places and institutions) have been <code>capitalized</code>. The beginning of a block of text as well as the original capitals used by the notary are also capitalized.</li> <li>The consonantal <code>i</code> and <code>u</code> characters have been transcribed as <code>j</code> and <code>v</code> in both French and Latin.</li> <li>The punctuation marks used in the text: <code>.</code> and <code>/</code> have been transcribed, but the transcription has not been standardized with modern punctuation.</li> <li>Corrections and words that appear cancelled in the manuscript have been transcribed surrounded by the sign <code>$</code> at the beginning and at the end.</li> <li>More specific transcription rules can be found into the file <code>transcription_guidelines.pdf</code></li> </ul> <p> </p> <p><strong>Section 2. e-NDP Layout Segmentation.</strong></p> <p>Layout segmentation is a compulsory step before HTR recognition in order to distinguish sections and regions inside a document. This process intend to separate interdependant page zones to produce a recognition in a section-sequence order and not in a line-sequence order which mix textual and peri-textual content.</p> <p>The regions of 364 pages (see <code>GT-layout_list</code>) of the e-NDP corpus were annotated using a 5 sections vocabulary (see <code>endp_layout_regions</code>) in order to describe the page distribution in all the 26 volumes :</p> <ol> <li><em>Block</em> : All the central text blocks, that normally corresponds to the main content called "conclusions" in registers.</li> <li><em>Liste</em> : List of names of the canons who were present during the meeting. Normally located before the <em>conclusions</em>.</li> <li><em>Entrée</em> : Marginal notes or entries to inform about the content of <em>conclusions</em>.</li> <li><em>Date</em> : Paragraph contending the date. Normally at the head of a <em>conclusion</em>, but separate of the main body.</li> <li><em>Numérotation</em> : Page numbers in roman or arabic. Usually appear in the top corners of the pages.</li> </ol> <table align="center"> <caption><strong>Layout GT statistics</strong></caption> <tbody> <tr> <th>Region</th> <th>Count</th> </tr> <tr> <td>block</td> <td>833</td> </tr> <tr> <td>liste</td> <td>431</td> </tr> <tr> <td>date</td> <td>448</td> </tr> <tr> <td>entrée</td> <td>205</td> </tr> <tr> <td>numérotation</td> <td>531</td> </tr> </tbody> </table> <p> </p> <p><strong>Section 3. The e-NDP HTR modeling.</strong></p> <p>The e-NDP project has progressively trained several HTR models adapted to work on late medieval cursive in order to accelerate the production of ground truth. Currently the best model delivers an average <strong>CER (Character error ratio) of 9.7%</strong> in handwriting recognition on the 26 registers (see <code>endp_learning_curve</code>) and can serve as generalist model for other manuscripts of the same period and similar script family. These models and their training implementation details can be found in the project's github <a href="https://github.com/chartes/e-NDP_HTR">repository</a>. </p> <p>Additionally, the automatic HTR transcriptions of the 26 registers (14k pages, 4.5M tokens) enriched with lexical and semantical information has been the subject of a first <a href="https://nosketch-engine.lamop.fr/#dashboard?corpname=endp">online publication</a> using the NoSketch engine that allows advanced data mining based on the combination of data, metadata and NLP features. </p> <p> </p> <p><strong>Section 4. Dataset content.</strong></p> <p>This zip dataset contains :</p> <p>- <code>HTR_ground_truth</code> : Two folders containing the jpg / jpeg images and their curated transcriptions in PAGE XML format.</p> <p>- <code>images_docs</code> : 4 files illustrating the different phases of the project (list of GT for layout segmentation, layout ontologie, transcription guideline and HTR evaluation curves)</p>
DUDE competition train - validation - test splits ground truth
<p>This JSON file contains the ground truth annotations for the train and validation set of the DUDE competition (https://rrc.cvc.uab.es/?ch=23&com=tasks) of ICDAR 2023 (https://icdar2023.org/).</p> <p> </p> <p><strong>V1.0.7 release</strong>: 41454 annotations for 4974 documents (train-validation-test)</p> <pre>DatasetDict({ train: Dataset({ features: ['docId', 'questionId', 'question', 'answers', 'answers_page_bounding_boxes', 'answers_variants', 'answer_type', 'data_split', 'document', 'OCR'], num_rows: 23728 }) val: Dataset({ features: ['docId', 'questionId', 'question', 'answers', 'answers_page_bounding_boxes', 'answers_variants', 'answer_type', 'data_split', 'document', 'OCR'], num_rows: 6315 }) test: Dataset({ features: ['docId', 'questionId', 'question', 'answers', 'answers_page_bounding_boxes', 'answers_variants', 'answer_type', 'data_split', 'document', 'OCR'], num_rows: 11402 }) }) ++update on answer_type +++formatting change to answers_variants ++++stricter check on answer_variants & rename annotations file <strong>+ blind test set (no ground truth answers provided) </strong>++ removed duplicates from test set: </pre> <blockquote> <p> "92bd5c758bda9bdceb5f67c17009207b_ac6964cbdf483e765b6668e27b3d0bc4",</p> <p> "6ee71a16d4e4d1dbd7c1f569a92d4e08_549f2a163f8ff3e9f0293cf59fdd98bc",</p> <p> "e6f3855472231a7ca6aada2f8e85fe5a_827c03a72f2552c722f2c872fd7f74c3",</p> <p> "e3eecd7cca5de11f1d17cd94ae6a8d77_6300df64e4cf6ba0600ac81278f68de2",</p> <p> "107b4037df8127a92ee4b6ae9b5df8fb_d7a60e7a9fc0b27487ea39cd7f56f98e",</p> <p> "300cc3900080064d308983f958141232_6a7cf1aad908d58a75ab8e02ddc856f4",</p> <p> "fdd3308efacddb88d4aa6e2073f481d4_138cb868ecc804a63cc7a4502c0009b2",</p> <p> "1f7de256ff1743d329a8402ba0d132e7_95b6e8758533a9817b9f20a958e7b776",</p> <p> "4f399b8c526ffb6a2fd585a18d4ed5ec_51097231bc327c26c59a4fd8d3ff3069",</p> </blockquote> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.