Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

193

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

193 results for “labeled data”

Learn how ShareScore rates datasets ↗
zenodo52/100

Labeled Time Series Data of Force/Torque for Monitoring Assembly Processes with a Delta Robot

<p>This dataset comprises 524 recordings of 6-dimensional time series data, capturing forces in three directions and torques in three directions during the assembly of small car model wheels. The data was collected using an equidistant sampling method with a sampling period of 0.004 seconds. Each time series represents the process of assembling one wheel, specifically the placement of a tire onto a rim, and includes a label indicating whether the assembly was successful (OK). The wheels were assembled in batches of four, and the recordings were obtained over six different days. The labels of recordings from two (days 3 and 4) of the six days are invalid as described in [1].&nbsp; The labels presented in this data set are only binary (they do not describe the reason of the failure). The labels of recordings from days 5 and 6 are created by human while the other labels came from a convolutional neural network based computer vision classifier and can be inaccurate as described in section 5.4 of [1].&nbsp; &nbsp;</p> <h4>Dataset Structure:</h4> <ul> <li><strong>File:</strong> <code>ForceTorqueTimeSeries.csv</code> <ul> <li><strong>Columns:</strong> <ul> <li><code>idx (1-524)</code>: Index of the recording corresponding to the assembly of one wheel.</li> <li><code>label (true/false)</code>: Indicates whether the assembly was successful (TRUE = product is OK).</li> <li><code>meas_id (1-6)</code>: Identifier for the day on which the recording was made (refer to Table 2.1 in [1]).</li> <li><code>force_x</code>: X-component of the force measured by the sensor mounted on the delta robot's end effector.</li> <li><code>force_y</code>: Y-component of the force.</li> <li><code>force_z</code>: Z-component of the force.</li> <li><code>torque_x</code>: X-component of the torque.</li> <li><code>torque_y</code>: Y-component of the torque.</li> <li><code>torque_z</code>: Z-component of the torque.</li> </ul> </li> </ul> </li> </ul> <h4>Additional Files:</h4> <ul> <li><strong><code>IMG_3351.MOV</code>:</strong> A video demonstrating the assembly process for one batch of four wheels.</li> <li><strong><code>F3-BP-2024-Trna-Ales-Ales Trna - 2024 - Anomaly detection in robotic assembly process using force and torque sensors.pdf</code>:</strong> Bachelor thesis [1] detailing the dataset and preliminary experiments on fault detection.</li> <li><strong><code>F3-BP-2024-Hanzlik-Vojtech-Anomaly_Detection_Bachelors_Thesis.pdf</code>:</strong> Bachelor thesis [2] describing the data acquisition process.</li> </ul> <h3>References:</h3> <ol> <li>Trna, A. (2024). <em>Anomaly detection in robotic assembly process using force and torque sensors</em> [Bachelor&rsquo;s thesis, Czech Technical University in Prague].</li> <li>Hanzlik, V. (2024). <em>Edge AI integration for anomaly detection in assembly using Delta robot</em> [Bachelor&rsquo;s thesis, Czech Technical University in Prague].</li> </ol>

opencc-by-4.0Sep 2024View details →
zenodo48/100

Labelled magnetic reconnection simulation data set

<p>Numerical simulations have been performed on Marconi at CINECA (Italy) under the ISCRA initiative.&nbsp;The corresponding data can be found at:&nbsp;<a href="https://doi.org/10.5281/zenodo.3935887">https://doi.org/10.5281/zenodo.3935887</a></p>

opencc-by-4.0Jun 2020View details →
zenodo48/100

Data from 'Tracability of Forest Reproductive Material with the quality label 'Plant van Hier': A DNA database with genetic profiles of native autochthonous tree and shrub species of Flanders, Belgium'

<h2>Background</h2> <p>Indigenous trees and shrubs play an important role in multifunctional forest management. They form a significant part of the biodiversity in our forests. Forest reproductive material (FRM) of autochthonous Flemish origin is sold under the quality label &lsquo;Plant van Hier&rsquo;, a certification mark of the Agency for Nature and Forests. To ensure the provenance of the seedlings, we developed a DNA-database of genetic profiles of potential parent trees, using species-specific genetic markers. This database enables the traceability of FRM of the &lsquo;Plant van Hier&rsquo; label throughout the entire production chain; from seed harvesting and cultivation to planting by the end user.</p> <p>This database contains the genetic profiles of almost all possible parent trees present within 27 Flemish autochthonous seed orchards of eight ecologically important tree and shrub species: <em>Carpinus betulus</em>, <em>Corylus avellana</em>, <em>Frangula alnus</em>, <em>Populus tremula</em>, <em>Sorbus aucuparia</em>, <em>Tilia cordata</em>, <em>Tilia platyphyllos,</em> and <em>Ulmus laevis</em>. The profiles were established using microsatellite markers (11 to 24 markers per species).&nbsp;&nbsp;New genetic markers were developed for&nbsp;<em>Carpinus betulus</em> and <em>Ulmus laevis</em>. PCR products were run on an ABI 3500 Genetic Analyser (Thermo Fisher Scientific).</p> <h2>Files</h2> <p>The files will be updated when new genotypes are added to the seed orchards. The current data files contain data from genotypes collected in the period 2018-2023.&nbsp;</p> <h3>Species_genotypes</h3> <p>These files contain the genetic fingerprints of the parent trees of autochthonous Flemish seed orchards. Missing data is indicated as &lsquo;MD&rsquo;. For <em>Carpinus betulus</em>, an octoploid species, the allelic phenotype is given instead of the genotype as the number of times that an allele occurs on a specific locus is not known.</p> <p>The next metadata is additionally given:<br>- Species: the Latin name of the species<br>- Seed_orchard: the name of the seed orchard in which the genotypes are located<br>- Code_seed_orchard: the code of the seed orchard in which the genotypes are located as given in the Register of Flemish Forest Reproductive Material (&lsquo;Register bosbouwkundig uitgangsmateriaal&rsquo;; inbo.be)<br>- Genotype: the fieldname given to the genotype<br>- Origin: the location where the genotype was collected in Flanders, Belgium. Genotypes were collected from natural stands which are assumed to have an autochthonous origin. When the specific location is unknown, the location &lsquo;Flanders&rsquo; is given.&nbsp;<br>- Year_sampled: the year in which the genotypes were sampled in the respective seed orchard for genetic analysis.</p> <h3>Species_binsets</h3> <p>These files contain the binsets and allele names that are used to score the alleles of the genotypes in the programme Geneious Prime 2019.3.2 (<a href="https://www.geneious.com">https://www.geneious.com</a>). For <em>Tilia platyphyllos </em>and <em>Tilia cordata</em>, the same binsets were used.</p>

opencc-by-4.0Nov 2024View details →
zenodo48/100

Raw planetary images and boulder labels data (as shapefiles) collected during the BOULDERING Marie Skłodowska-Curie Global fellowship

<p>This database contains 64 large images of craters on the lunar and martian surfaces and 3 images of boulder fields on Earth (see manuscript <a href="https://agupubs.onlinelibrary.wiley.com/doi/full/10.1029/2023JE008013">https://agupubs.onlinelibrary.wiley.com/doi/full/10.1029/2023JE008013</a> for more information on those terrestrial locations). The data was collected during the BOULDERING Marie Skłodowska-Curie Global fellowship between October 2021 and 2024.</p> <p>For each image, the boulder outlines within specific tiles within the image were carefully mapped in QGIS. More information about the labelling procedure can be found in the following manuscript (<a href="https://agupubs.onlinelibrary.wiley.com/doi/full/10.1029/2023JE008013">https://agupubs.onlinelibrary.wiley.com/doi/full/10.1029/2023JE008013</a>). This dataset differs from the previous dataset included along with the manuscript&nbsp;<a href="https://zenodo.org/records/8171052">https://zenodo.org/records/8171052</a>, as it contains more mapped images, especially of boulder populations around young impact structures on the Moon (cold spots).&nbsp;</p> <p>For each location, you will find a raster with a .tif format, and three shapefiles:</p> <ul> <li> <p>a boulder-mapping file, which is the manually digitized outline of boulders.</p> </li> <li> <p>a tiles-completely-mapped file, which depicts the patches/tiles/windows on which the boulder mapping has been conducted.</p> </li> <li> <p>a global-tiles file, which shows all of the image patches/tiles/windows (pick the term you are the most familiar with) within a raster.</p> </li> </ul> <p>In addition you will find .pkl (which stands for pickle), which contains some information about the patches/tiles/windows if you would need to clip those windows out from the original raster. You can find more information in the way we process this raw data into a format which can be ingested in a deep learning model (see <a href="https://zenodo.org/records/14250874" target="_blank" rel="noopener">https://zenodo.org/records/14250874</a>) in the two following github repositories (<a href="https://github.com/astroNils/YOLOv8-BeyondEarth" target="_blank" rel="noopener">https://github.com/astroNils/YOLOv8-BeyondEarth</a> and&nbsp;<a href="https://github.com/astroNils/MLtools/tree/main" target="_blank" rel="noopener">https://github.com/astroNils/MLtools</a>). If you don't plan in adding more training data, you can directly used the pre-processed database (see <a href="https://zenodo.org/records/14250874" target="_blank" rel="noopener">https://zenodo.org/records/14250874</a>).</p> <p>There are multiple locations/images per planetary body. Cold spots are located on the Moon, but they are saved in a folder of their own.&nbsp;</p> <p>Note that the cold spots boulder mapping shapefiles are partially manually mapped, and partially originating from predictions made from a deep learning model (which explains the outline of boulders are predicted within one pixel).</p> <p><strong>How to cite:</strong></p> <p>Please refer to the "how to cite" section of the readme file of <a href="https://github.com/astroNils/YOLOv8-BeyondEarth" target="_blank" rel="noopener">https://github.com/astroNils/YOLOv8-BeyondEarth.</a></p> <p><strong>Structure:</strong></p> <pre><code>. └── raw_data/ ├── coldspots/ │ └── image_name/ │ ├── shp/ │ │ ├── &lt;image_name&gt;-tiles-completely-mapped.shp │ │ ├── &lt;image_name&gt;-boulder-mapping.shp │ │ └── &lt;image_name&gt;-global-tiles.shp │ └── raster/ │ └── &lt;image_name&gt;.tif ├── earth/ │ └── image_name/ │ ├── shp/ │ │ ├── &lt;image_name&gt;-tiles-completely-mapped.shp │ │ ├── &lt;image_name&gt;-boulder-mapping.shp │ │ └── &lt;image_name&gt;-global-tiles.shp │ └── raster/ │ └── &lt;image_name&gt;.tif ├── mars/ │ └── image_name/ │ ├── shp/ │ │ ├── &lt;image_name&gt;-tiles-completely-mapped.shp │ │ ├── &lt;image_name&gt;-boulder-mapping.shp │ │ └── &lt;image_name&gt;-global-tiles.shp │ └── raster/ │ └── &lt;image_name&gt;.tif └── moon/ └── image_name/ ├── shp/ │ │ ├── &lt;image_name&gt;-tiles-completely-mapped.shp │ │ ├── &lt;image_name&gt;-boulder-mapping.shp │ │ └── &lt;image_name&gt;-global-tiles.shp └── raster/ └── &lt;image_name&gt;.tif</code></pre>

opencc-by-4.0Nov 2024View details →
zenodo44/100

Quantitative Content Analysis Data for Hand Labeling Road Surface Conditions in New York State Department of Transportation Camera Images

<p><strong>Foundational Codebook and Data:&nbsp;</strong></p> <p>Traffic camera images from the New York State Department of Transportation (511ny.org) are used to create a hand-labeled dataset of images classified into to one of six road surface conditions: 1) severe snow, 2) snow, 3) wet, 4) dry, 5) poor visibility, or 6) obstructed. Six labelers (authors Sutter, Wirz, Przybylo, Cains, Radford, and Evans) went through a series of four labeling trials where reliability across all six labelers were assessed using the Krippendorff&rsquo;s alpha (KA) metric (Krippendorff, 2007). The online tool by Dr. Freelon (Freelon, 2013; Freelon, 2010) was used to calculate reliability metrics after each trial, and the group achieved inter-coder reliability with KA of 0.888 on the 4th trial. This process is known as quantitative content analysis, and three pieces of data used in this process are shared, including: 1) a PDF of the codebook which serves as a set of rules for labeling images, 2) images from each of the four labeling trials, including the use of New York State Mesonet weather observation data (Brotzge et al., 2020), and 3) an Excel spreadsheet including the calculated inter-coder reliability (ICR) metrics and other summaries used to asses reliability after each trial. The data are included in NYSDOT_quantitative_content_analysis.zip.</p> <p>The broader purpose of this work is that the six human labelers, after achieving inter-coder reliability,&nbsp;can then label large sets of images independently, each contributing to the creation of larger labeled dataset&nbsp;used for&nbsp;training supervised machine learning models to predict road surface conditions from camera images. The xCITE lab&nbsp;(xCITE, 2023) is used to store&nbsp;camera images from 511ny.org, and the lab provides computing resources for training machine learning models.</p> <p><strong>Obstructed Class Variation: </strong></p> <p>There are many applications for labeling roadside camera images, and as a variation of the foundational codebook, an addendum codebook provides another version of labeling the obstructed class. Specifically, this variation prioritizes labeling an image as &ldquo;obstructed&rdquo; only in extreme circumstances where there is a camera- or image- specific problem that prevents the assessment of any road surfaces. For labelers who want to use this version of the obstructed class (in this document) and also the other five weather-related classes (in the foundational codebook), the guidance is to use both documents in tandem, making sure to use the obstructed rules/definitions in this document while disregarding the obstructed rules/definitions in the foundational codebook. Alternatively, this codebook may be used alone in applications where the goal is to solely classify obstructed vs not obstructed.&nbsp;To ensure reliability and quality of this variation, quantitative content analysis was conducted on this addendum codebook, just as it was for the foundational codebook. Two labelers were tested with a sample of 30 images and achieved inter-coder reliability with Krippendorff's Alpha of 0.934 after one trial. The data, including the addendum codebook and labeling trial data (images and results) are included in ObstructedVariation_quantitative_content_analysis.zip.</p> <p>This material is based upon work supported by the U.S. National Science Foundation under Grant No. RISE-2019758.</p>

opencc-by-4.0Sep 2023View details →
zenodo44/100

Raw EEG Data for: Learning from Label Proportions in Brain-Computer Interfaces

<p>If you prefer to use the preprocessed and epoched data, please refer to: https://zenodo.org/record/192684</p> <p>Note that this repository ontains only the visual paradigm with the N=13 subjects recorded at 31 EEG channels, as described in the above link. We copied the relevant section of the description below:</p> <blockquote> <p>This data repository contains raw EEG of an EEG experiment utilizing visual event-related potentials (ERPs) with N=13 healthy subjects.</p> <p>The dataset is used and described in the following journal article:</p> <p><em>H&uuml;bner, D., Verhoeven, T., Schmid, K., M&uuml;ller, K. R., Tangermann, M., &amp; Kindermans, P. J. (2017). Learning from label proportions in brain-computer interfaces: online unsupervised learning with guarantees. PloS one, 12(4), e0175856.</em></p> <p><strong>Please cite the above article when using the data.</strong></p> <p>The data set with N=13 subjects is different to ordinary ERP datasets in the sense that the train of stimuli to spell one character (68) is divided into repetitions of two interleaved sequences with length 8 and 18, respectively. We added &#39;#&#39; symbols to the spelling matrix which should never be attended by the subject and hence, are non-targets by definition. The first, shorter sequence, now highlights only ordinary characters, while the second sequence also highlights &#39;#&#39; -- visual blank symbols. By construction, sequence 1 has a higher target ratio than sequence 2. These known, but different target and non-target proportions are then used to reconstruct the target and non-target class means. This approach which does not need explicit class labels is termed Learning from Label Proportions (LLP). It can be used to decode brain signals without prior calibration session. More details can be found in the article.</p> <p>In another study, the above data set was used to simulate a new unsupervised mixture approach which combines the mean estimation of the unsupervised expectation-maximization algorithm by Kindermans et al. (2012, PLoS One) with the means obtained with the LLP approach. This leads to an unsupervised solution for which the performance is as good as in the supervised scenario. Please find more details in the following article:</p> <p><em>Verhoeven, T., H&uuml;bner, D., Tangermann, M., M&uuml;ller, K. R., Dambre, J., &amp; Kindermans, P. J. (2017). Improving zero-training brain-computer interfaces by mixing model estimators. Journal of neural engineering, 14(3), 036021.</em></p> </blockquote> <p>The data was recorded with BrainVision recorder. A new file was recorded for every group of 7 characters. The .eeg file contains the RAW EEG data in the format as described in the .vhdr file. Events / stimuli markers are provided in the .vmrk files. Note that there is a wrapper available to use this data in MOABB here: TODO INSERT LINK</p> <p>The subjects had the task to spell a specific sentence with 63 letters. In the online experiment, this was repeated 3 times and each time the online unsupervised classifier was reset at the start of the sentence.</p>

opencc-by-4.0Jan 2022View details →
zenodo44/100

A benchmark dataset of herbarium specimen images with label data: Summary

<p>This landing page contains a CSV file compiling all data associated with herbarium specimens that are part of this dataset, as they could be found on GBIF, JACQ or FinBIF. A CSV file with and without Darwin Core extension data is available, as some CSV readers have trouble with the JSON format that is used for those extensions.</p> <p>In addition, DOI&#39;s of the individual specimens uploaded to Zenodo and direct links to the different files (JPEG, TIFF, JSON, PNG) are also included. Index of these added&nbsp;variables:</p> <p>- persistentID: Persistent Identifier of the collection specimen. Data uploaded as part of this dataset will not be kept in sync with changes at the collection&#39;s repository. Hence, this URI will always point to the most up to date information known about the herbarium specimen.</p> <p>- jpegURL, tiffURL, jsonURL: URL&#39;s pointing straight to the respective image and data files themselves, to facilitate (selective) batch downloads.</p> <p>- pngSegAllURL and pngSegSelURL: Segmented overlays of the herbarium specimens indicating the location of different labels and reference material on the sheet (&quot;All&quot;) and their content (&quot;Sel&quot;). More information can be found in the paper (in prep) associated with this data publication and the individual depositions themselves.</p> <p>- DOI: The DOI of the deposition of images and data of these specimens on Zenodo. DOI&#39;s point to the most up-to-date version of these depositions at the time of the publication of this CSV file. As a rule, this CSV file will be updated should any changes happen to any of the depositions.</p> <p>- jpegURL2, tiffURL2: A few herbarium sheets had labels on the back and consisted therefore of two scans. As a rule, the label scans are in this category.</p>

opencc-zeroNov 2018View details →
zenodo44/100

Mars orbital image (HiRISE) labeled data set

<p>This data set contains 3820 landmarks that were extracted from 168 HiRISE images. The landmarks were detected in HiRISE browse images. For each landmark, we cropped a square bounding box the included the full extent of the landmark plus a 30-pixel margin to left, right, top, and bottom. Each cropped image was then resized to 227x227 pixels.</p> <p><strong>Contents</strong>:</p> <ul> <li>map-proj/: Directory containing individual cropped landmark images</li> <li>labels-map-proj.txt: Class labels (ids) for each landmark image</li> <li>landmark_mp.py: Python dictionary that maps class ids to semantic names</li> </ul> <p><strong>Attribution</strong>:</p> <p>If you use this data set in your own work, please cite this DOI: 10.5281/zenodo.1048301</p> <p>Please also cite this paper, which provides additional details about the data set.</p> <p>Kiri L. Wagstaff, You Lu, Alice Stanboli, Kevin Grimes, Thamme Gowda, and Jordan Padams. &quot;Deep Mars: CNN Classification of Mars Imagery for the PDS Imaging Atlas.&quot; <em>Proceedings of the Thirtieth Annual Conference on Innovative Applications of Artificial Intelligence</em>, 2018.</p> <p>&nbsp;</p>

opencc-by-sa-4.0Nov 2017View details →
zenodo44/100

Mars Target Encyclopedia - LPSC abstracts labeled data set

<p>This data set contains annotated text versions of 2-page abstracts published at the Lunar and Planetary Science Conference in 2015 and 2016.</p> <p>The original PDF abstracts are available at:</p> <ul> <li>https://www.hou.usra.edu/meetings/lpsc2015/programAbstracts/view/</li> <li>https://www.hou.usra.edu/meetings/lpsc2016/programAbstracts/view/</li> </ul> <p>The text files in this archive were extracted using the Apache Tika PDF parsing tool.  The text is provided here so that the annotations can be viewed.  The text content remains copyright of the original abstract authors.</p> <p>The annotations (entities and relations) are provided in the format used by the brat annotation tool.  To view the annotations in a web-based graphical form, install the brat tool (http://brat.nlplab.org/).  These annotations were generated using brat v1.3.  The annotation files are also human-readable and can be parsed in to be used directly in code.</p> <p><strong>Contents</strong>:</p> <ul> <li>lpsc15/: 62 abstracts</li> <li>lpsc16/: 55 abstracts</li> </ul> <p>Each directory contains a .txt and .ann file for each abstract.  The .ann file is in brat standoff format (http://brat.nlplab.org/standoff.html).</p> <p>Additional .conf files are provided to generate color highlighting and keyboard shortcuts.  These are used by the brat tool.</p> <p><strong>Attribution</strong>:</p> <p>If you use this data set in your own work, please cite this DOI:</p> <p>10.5281/zenodo.1048419</p> <p>Please also cite this paper, which provides additional details about the data set.</p> <p>Kiri L. Wagstaff, Raymond Francis, Thamme Gowda, You Lu, Ellen Riloff, Karanjeet Singh, and Nina Lanza. "Mars Target Encyclopedia: Rock and Soil Composition Extracted from the Literature."  <em>Proceedings of the Thirtieth Annual Conference on Innovative Applications of Artificial Intelligence</em>, 2018.</p>

opencc-by-sa-4.0Nov 2017View details →
zenodo44/100

Mars surface image (Curiosity rover) labeled data set

<p>This data set consists of 6691 images spanning 24 classes that were collected by the Mars Science Laboratory (MSL, Curosity) rover by three instruments (Mastcam Right eye, Mastcam Left eye, and MAHLI).&nbsp; These images are the &quot;browse&quot; version of each original data product, not full resolution.&nbsp; They are roughly 256x256 pixels each.</p> <p>We divided the MSL images into train, validation, and test data sets according to their sol (Martian day) of acquisition.&nbsp; This strategy was chosen to model how the system will be used operationally with an image archive that grows over time.&nbsp; The images were collected from sols 3 to 1060 (August 2012 to July 2015).&nbsp; The exact train/validation/test splits are given in individual files.&nbsp; Full-size images can be obtained from the PDS at https://pds-imaging.jpl.nasa.gov/search/ .</p> <p><strong>Contents</strong>:</p> <ul> <li>calibrated/: Directory containing calibrated MSL images</li> <li>train-calibrated-shuffled.txt: Training labels (images in shuffled order)</li> <li>val-calibrated-shuffled.txt: Validation labels</li> <li>test-calibrated-shuffled.txt: Test labels</li> <li>msl_synset_words-indexed.txt: Mapping from class IDs to class names</li> </ul> <p><strong>Attribution</strong>:</p> <p>If you use this data set in your own work, please cite this DOI:</p> <p>10.5281/zenodo.1049137</p> <p>Please also cite this paper, which provides additional details about the data set.</p> <p>Kiri L. Wagstaff, You Lu, Alice Stanboli, Kevin Grimes, Thamme Gowda, and Jordan Padams. &quot;Deep Mars: CNN Classification of Mars Imagery for the PDS Imaging Atlas.&quot; <em>Proceedings of the Thirtieth Annual Conference on Innovative Applications of Artificial Intelligence</em>, 2018.</p>

opencc-by-sa-4.0Nov 2017View details →
zenodo44/100

Post-hoc labeling of arbitrary EEG recordings for data-efficient evaluation of neural decoding methods

<p>EEG signals&nbsp;recorded from seven healthy subjects. On average,&nbsp;Seventy-three minutes of EEG data&nbsp;were recorded&nbsp;from 31 electrodes placed according to the extended 10-20 system. Signals are used in the paradigm-agnostic post-hoc labeled dataset generation framework for benchmarking of oscillatory neural decoding methods.</p>

opencc-by-4.0Nov 2017View details →
zenodo44/100

Eye image data with gaze labels recorded using custom video-oculography hardware at 120Hz

<p>The repository of eye image data with corresponding gaze labels collected from 40 subjects. The preview contains a collage of random image samples, one per subject.&nbsp;</p> <p>All recorded subjects gave informed consent under an experimental protocol approved by the Institutional Research Board of Texas State University (approval code 2018044) and their data were anonymized prior to public release.</p> <p>The data were recorded using the custom video-oculography (VOG) desktop hardware setup at 120Hz. The full description of this eye-tracking system's capabilities is provided at https://doi.org/10.48550/arXiv.1904.07361.</p> <p>This VOG set contains recordings of the random oblique saccades task. It is comprised of 174 on-screen fixation targets that densely cover the range of &plusmn;20.51&deg; horizontally and &plusmn;16.7&deg; vertically (in degrees of visual angle). More detail on the presented stimuli can be found at https://doi.org/10.1145/3379156.3391370.</p> <p>The data were also used in Dmytro Katrychuk's Ph.D. thesis "Generating Realistic Eye Images to Evaluate Photosensor Oculography Eye-Tracking for Portable Headsets" (https://hdl.handle.net/10877/19437); with the release for public use in the upcoming publication "An appearance-based gaze estimation as a benchmark for eye image data generation methods" accepted to MDPI Journal of Applied Sciences.&nbsp;</p> <p>Each .zip archive represents a recording from one subject, which includes:</p> <ul> <li>Video of the close eye capture in ".avi" format</li> <li>Calibration data in ".xml" format</li> <li>Gaze data in ".tsv" format</li> <li>On-screen target stimulus position in ".tsv" format</li> </ul> <p>The "src.zip" provides a Python script to unpack each ".avi" video recording to a set of ".png" images. The direct playback of ".avi"s may require special codecs and is not supported.&nbsp;</p> <p>Any additional code will be uploaded to https://github.com/dkatrychuk/psog-eval-diss2023</p> <p>The authors can be contacted at their corresponding emails: Dmytro Katrychuk - d_k139@txstate.edu; Oleg Komogortsev - ok@txstate.edu.</p>

opencc-by-4.0Oct 2024View details →
zenodo44/100

Detecting coarse beach sediment using remotely sensed imagery at the FRF, Duck, NC, USA: Labeled images, deep learning model, testing data, and predictions.

<p>This data record contains 5 zip files all used to build and use a semantic segmentation model to operate on beach imagery taken at the Field Research Facility (FRF) in Duck, North Carolina, USA. &nbsp;All data is from 2015-2021</p> <p>The `training_data.zip` contains all data used to train the ML model. All images come from the north facing (c1) camera. This zip file includes: a list of classes used to label the imagery, and folders of 107 images, 107 sparse annotations (doodles), 107 labels, and 107 overlays. All labeling was done with the open-source labeling tool &lsquo;Doodler (Buscombe et al., 2021).</p> <p>The `model.zip` file contains the ML model, and associated metadata. This includes: a JSON model configuration file, a figure showing model training statistics, an `.npz` file of model training output, a list of training and validation files, the model as an h5 file and in the Tensorflow &lsquo;saved model&rsquo; format. &nbsp;All modeling was done with Segmentation Gym (Buscombe &amp; Goldstein 2022).</p> <p>The `test_data_c6.zip` file contains all data from the south facing (c6) camera to test the ML model. This includes: a list of classes used to label the imagery, and folders of 10 images, 10 sparse annotations (doodles), 10 labels, and 10 overlays. &nbsp;All labeling was done with the open-source labeling tool &lsquo;Doodler (Buscombe et al., 2021). Testing the model with this data was done with codes in: https://github.com/ebgoldstein/FRF_GrainSize</p> <p>The `test_data_c1.zip` file contains all data from the north facing (c1) camera to test the ML model. This includes: a list of classes used to label the imagery, and folders of 10 images, 10 sparse annotations (doodles), 10 labels, and 10 overlays. &nbsp;All labeling was done with an open-source labeling tool &lsquo;Doodler (Buscombe et al., 2021). Testing the model with this data was done with codes in: https://github.com/ebgoldstein/FRF_GrainSize</p> <p>The `predictions.zip` file contains 4418 images from the north facing (c1) camera that were run through the trained segmentation model as well as the resulting output (presented as side-by-side image and overlays). These images were created using codes in Segmentation Gym (Buscombe &amp; Goldstein 2022).</p>

opencc-by-4.0Jan 2023View details →
zenodo44/100

Data on soil compounds, respiration and incorporation of 13C-labeled substrate

<p>Root exudation increases the concentration of readily available carbon (C) compounds in its immediate environment. This creates &lsquo;hotspots&rsquo; of microbial activity characterized by accelerated soil organic matter turnover with direct implications for nutrient availability for plants. However, we still lack a deeper understanding of the microbial metabolic processes that occur in the immediate vicinity of the roots during and after a root exudation event. Even though theoretical concepts have been developed, the direct consequences of root exudation on microbial metabolism and nutrient availability have never been measured in their immediate environment in intact soil.</p> <p>Here, we used reverse microdialysis to simulate root exudation by releasing a <sup>13</sup>C-labelled mix of low-molecular-weight organic C compounds at discrete, mm-sized locations in undisturbed soil in combination with <sup>13</sup>C stable isotope tracing. This approach allowed us to investigate the fine-scale temporal and spatial response of microbial metabolism and soil chemistry to root exudation at the mm-scale, and to trace microbial respiration and uptake of exuded compounds.</p> <p>Our results show that a 9-hour simulated root exudation pulse leads to i) a large local respiration event and ii) alteration of the temporal dynamics of soil metabolites over the following twelve days right at the spot of exudate release. Notably, we observed an approximately threefold increase in ammonium concentrations twelve hours after the pulse and increased nitrate concentrations five days after the pulse. We also observed an increase of various short-chain fatty acids, such as acetate, propionate and formate over the following days, indicating altered microbial metabolic pathways and activity. Phospholipid and neutral lipid fatty acids (PLFAs and NLFAs) of all major microbial groups were significantly enriched in <sup>13</sup>C within a radius of 5 mm around the microdialysis probes, but not beyond. The highest relative <sup>13</sup>C enrichment was observed in fungal NLFAs, indicating that a significant proportion of the exuded compounds had been incorporated into fungal storage compounds.</p> <p>Our findings indicate that the punctual release of low-molecular weight organic C compounds into intact soil significantly changes microbial metabolism and activity in its immediate surroundings, which lead to enhanced mineralisation of native organic nitrogen (N). Our observations emphasise the versatility of microbial metabolic pathways that underlie the response of soil microbes to rapidly altered C availability. They furthermore demonstrate the effectiveness of this response, as triggered by root exudation pulses, to increase nutrient availability for plants around the root.</p>

opencc-by-4.0Jul 2023View details →
zenodo44/100

Input and output data (images + boulder labels, model setup, model weights and more) for the manuscript "Automatic characterization of boulders on planetary surfaces from high-resolution satellite images"

<p><strong>File 1:</strong> raw_data_BOULDERING.zip</p> <p><strong>Size:</strong> 8.8 GB</p> <p><strong>Summary: </strong>It contains all of the rasters (planetary images) and labeled boulders (raw data):</p> <ul> <li> <p>a boulder-mapping file, which is the manually digitized outline of boulders.</p> </li> <li> <p>a ROM file (stands for Region of Mapping), which depicts the image patches on which the boulder mapping has been conducted.</p> </li> <li> <p>a global-tiles file, which shows all of the image patches within a raster.</p> </li> </ul> <p>There are multiple locations/images per planetary body.</p> <p><strong>Structure:</strong></p> <pre>. └── raw_data/ ├── earth/ │ └── image_name/ │ &nbsp; ├── shp/ │ &nbsp; │ ├── &lt;image_name&gt;-ROM.shp │ &nbsp; │ ├── &lt;image_name&gt;-boulder-mapping.shp │ &nbsp; │ └── &lt;image_name&gt;-global-tiles.shp │ &nbsp; └── raster/ │ &nbsp; &nbsp; └── &lt;image_name&gt;.tif ├── mars/ │ └── image_name/ │ &nbsp; ├── shp/ │ &nbsp; │ ├── &lt;image_name&gt;-ROM.shp │ &nbsp; │ ├── &lt;image_name&gt;-boulder-mapping.shp │ &nbsp; │ └── &lt;image_name&gt;-global-tiles.shp │ &nbsp; └── raster/ │ &nbsp; &nbsp; └── &lt;image_name&gt;.tif └── moon/ &nbsp; └── image_name/ &nbsp; &nbsp; ├── shp/ &nbsp; &nbsp; │ ├── &lt;image_name&gt;-ROM.shp &nbsp; &nbsp; │ ├── &lt;image_name&gt;-boulder-mapping.shp &nbsp; &nbsp; │ └── &lt;image_name&gt;-global-tiles.shp &nbsp; &nbsp; └── raster/ &nbsp; &nbsp; &nbsp; └── &lt;image_name&gt;.tif</pre> <p>&nbsp;</p> <p><strong>File 2:</strong> best_model.zip</p> <p><strong>Size:</strong> 624.7 MB</p> <p><strong>Summary:</strong></p> <p>This zip file contains all of the inputs and outputs required/obtained from the training of the BoulderNet Mask R-CNN model (model setup, augmentation pipeline, model weights, log during training, logged metrics):</p> <ul> <li> <p>augmentation_pipeline.json (required as inputs for the training of the algorithm to apply augmentations). See <a href="https://github.com/astroNils">https://github.com/astroNils</a> and the MLtools repository for more information.</p> </li> </ul> <ul> <li> <p>Base-RCNN-FPN.yaml (base model setup file).</p> </li> <li> <p>config.yaml (complete model setup file, merge of the base and Mars-Moon-Earth setup file).</p> </li> <li> <p>Mars-MoonEarth-v050...yaml (model setup file).</p> </li> <li> <p>log.txt (log during training of the algorithm).</p> </li> <li> <p>model_0055999.pth (model weights at second last saving step)</p> </li> <li> <p>model_0063999.pth (model weights at last saving step)</p> </li> </ul> <p>We advice the use of model weights model_0055999.pth (to avoid slight overfitting).</p> <p><strong>File 3:</strong> Apr2023-Mars-Moon-Earth-mask-5px.zip (pre-processed input images)</p> <p><strong>Size:</strong> 252.8 MB</p> <p><strong>Summary:</strong></p> <p>This zip files contains the input data (images and boulder outlines) for the train, validation and test datasets. See <a href="https://github.com/astroNils">https://github.com/astroNils</a> and the MLtools repository for more information in how-to-use the different files.</p> <ul> <li> <p>The json folder contains json files that can be given as input (as a custom dataset) to the Detectron2 platform. The only differences between the two files is how the bounding boxes around masks have been generated. We advised to use &quot;Apr2023-Mars-Moon-Earth-mask-5px.json&quot;.</p> </li> <li> <p>The pkl folder and pickle file includes some informations about the 950 image patches in our boulder dataset.</p> </li> <li> <p>The pre-processing folder contains all of the training, validation and test image patches and corresponding shapefiles.</p> </li> <li> <p>The shapefile folder is actually empty (it should not be there!).</p> </li> </ul> <p><strong>Structure:</strong></p> <pre>. └── preprocessed_inputs/ &nbsp; ├── json &nbsp; ├── pkl &nbsp; ├── preprocessing/ &nbsp; │ &nbsp; ├── train/ &nbsp; │ &nbsp; │ &nbsp; ├── images &nbsp; │ &nbsp; │ &nbsp; └── labels &nbsp; │ &nbsp; ├── validation/ &nbsp; │ &nbsp; │ &nbsp; ├── images &nbsp; │ &nbsp; │ &nbsp; └── labels &nbsp; │ &nbsp; └── test/ &nbsp; │ &nbsp; &nbsp; &nbsp; ├── images &nbsp; │ &nbsp; &nbsp; &nbsp; └── labels &nbsp; └── shp</pre> <p>&nbsp;</p>

opencc-by-4.0Jul 2023View details →
zenodo44/100

Human Inner Ear Anatomy: Labeled Volume CT Data of Inner Ear Fluid Space and Anatomical Landmarks

<p>The provided dataset comprises 43 instances of temporal bone volume CT scans. The scans were performed on human cadaveric specimen with a resulting isotropic voxel size of <span class="math-tex">\(99 \times 99 \times 99 \, \, \mathrm{\mu m}^3\)</span>. Voxel-wise image labels of the fluid space of the bony labyrinth, subdivided in the three semantic classes cochlear volume, vestibular volume and semicircular canal volume are provided. In addition, each dataset contains JSON-like descriptor data defining the voxel coordinates of the anatomical landmarks: (1) apex of the cochlea, (2) oval window and (3) round window. The dataset can be used to train and evaluate algorithmic machine learning models for automated innear ear analysis in the context of the supervised learning paradigm.</p> <p>&nbsp;</p> <p><strong>Usage Notes</strong></p> <p>The datasets are formatted in the HDF5 format developed by the <a href="https://www.hdfgroup.org/solutions/hdf5/">HDF5 Group</a>. We utilized and thus recommend the usage of Python bindings <a href="https://www.h5py.org/">pyHDF</a> to handle the datasets.</p> <p>The flat-panel volume CT raw data, labels and landmarks are saved in the HDF5-internal file structure using the respective group and datasets:</p> <pre><code>raw/raw-0 label/label-0 landmark/landmark-0 landmark/landmark-1 landmark/landmark-2</code></pre> <p>Array raw and label data can be read from the file by indexing into an opened h5py file handle, for example as numpy.ndarray. Further metadata is contained in the attribute dictionaries of the raw and label datasets.</p> <p>Landmark coordinate data is available as an attribute dict and contains the coordinate system (LPS or RAS), IJK voxel coordinates and label information. The helicotrema or cochlea top is globally saved in landmark 0, the oval window in landmark 1 and the round window in landmark 2. Read as a Python dictionary, exemplary landmark information for a dataset may reads as follows:</p> <pre><code class="language-python">{'coordsys': 'LPS', 'id': 1, 'ijk_position': array([181, 188, 100]), 'label': 'CochleaTop', 'orientation': array([-1., -0., -0., -0., -1., -0., 0., 0., 1.]), 'xyz_position': array([ 44.21109689, -139.38058589, -183.48249736])}</code></pre> <p>&nbsp;</p> <pre><code class="language-python">{'coordsys': 'LPS', 'id': 2, 'ijk_position': array([222, 182, 145]), 'label': 'OvalWindow', 'orientation': array([-1., -0., -0., -0., -1., -0., 0., 0., 1.]), 'xyz_position': array([ 48.27890112, -139.95991131, -179.04103763])}</code></pre> <p>&nbsp;</p> <pre><code class="language-python">{'coordsys': 'LPS', 'id': 3, 'ijk_position': array([223, 209, 147]), 'label': 'RoundWindow', 'orientation': array([-1., -0., -0., -0., -1., -0., 0., 0., 1.]), 'xyz_position': array([ 48.33120126, -137.27135678, -178.8665465 ])}</code></pre> <p>&nbsp;</p>

opencc-by-4.0Aug 2023View details →
zenodo40/100

Table S3. List of Locustella sound recordings included in bioacoustic analysis surrounding description of the Taliabu Grasshopper-Warbler. The table provides information on sound library sources and sampling localities of recordings as well as raw data on all 11 bioacoustic parameters measured (see Supplementary Materials section SM3 for more details on parameters). Recordings whose source is labeled as "private recording" were obtained by colleagues and are available upon demand from the corresponding author.

<p>supplement to&nbsp;Rheindt, Frank E., Prawiradilaga, Dewi M., Ashari, Hidayat, Suparno, Gwee, Chyi Yin, Lee, Geraldine W. X., Wu, Meng Yue, Ng, Nathaniel S. R. (2020): A lost world in Wallacea: Description of a montane archipelagic avifauna. Science 367: 167-170, DOI: 10.1126/science.aax2146</p>

opencc-by-4.0Jan 2020View details →
zenodo40/100

cigKast: A data of 3D synthetic seismic volumes with labeled paleokarsts for deep-learning-based paleokarst interpretation

<p>cigKarst is a dataset created by the <a href="http://cig.ustc.edu.cn/">Computational Interpretation Group (CIG)</a> for the deep-learning-based peleokarst interpretation in 3D seismic images, <a href="http://cig.ustc.edu.cn/xinming/list.htm" target="_blank" rel="noopener">Xinming Wu</a> is the main contributor to the dataset.</p> <p>This dataset contains 120 pairs of synthetic 3D seismic images and the corresponding label images with the ground truth of the paleokarst systems simulated in the seismic images. More detail of building this dataset is discussed in the paper published at the journal of JGR Solid Earth:</p> <p><strong>Wu, X.</strong>, S. Yan, J. Qi, and H. Zeng, 2020, Deep learning for characterizing paleokarst collapse features in 3D seismic images.&nbsp;<strong>JGR, Solid Earth</strong>, Vol. 125(9), 1-23, e2020JB019685.&nbsp;<a href="http://cig.ustc.edu.cn/_upload/tpl/05/cd/1485/template1485/papers/wu2020karst.pdf">[PDF]</a>. doi: 10.1029/2020JB019685</p> <p>Below are some brief description of the dataset:</p> <p>1) The "seismic.zip" contains 120 3D seismic images, each image is with the dimension of 256X256X256;</p> <p>&nbsp;2) The "karst.zip" contains 120 3D label images of the karsts. Each label image is with the same dimension of 256X256X256. The values in a label image are set with ones in the karst areas while zeros elsewhere, which is why the compressed label images in the karst.zip is much smaller than the&nbsp;seismic images compressed in the seismic.zip</p>

opencc-by-4.0Nov 2020View details →
zenodo40/100

Training data for MaxQuant and Msstats label-free analysis in Galaxy

<p>The files serve as input and intermediate results for a MaxQuant and Msstats training on skin cancer tissues (<a href="https://doi.org/10.1016/j.matbio.2017.11.004">https://doi.org/10.1016/j.matbio.2017.11.004</a>) in the Galaxy training network (https://training.galaxyproject.org).</p> <p>Input files: human FASTA database for Maxquant. Annotation file and comparison matrix file for Msstats.</p> <p>Intermediate result files: MaxQuant protein groups, evidence and PTXQC.</p>

opencc-by-4.0Feb 2021View details →
zenodo40/100

LSD4WSD : An Open Dataset for Wet Snow Detection with SAR Data and Physical Labelling

<p><strong>LSD4WSD V2.0</strong></p><p><strong>L</strong>earning <strong>S</strong>AR <strong>D</strong>ataset for <strong>W</strong>et <strong>S</strong>now <strong>D</strong>etection - Full Analysis Version.&nbsp;</p><p>The aim of this dataset is to provide a basis for automatic learning to detect wet snow. It is based on Sentinel-1 SAR GRD satellite images acquired between August 2020 and August 2021 over the French Alps. The new version of this dataset is no longer simply restricted to a classification task, and provides a set of metadata for each sample.</p><p>Modification and improvements of the version 2.0.0 :</p><ul><li><i>Number of massif:</i> add 7 new massif to cover the all Sentinel-1 images (cf `info.pdf`).</li><li><i>Acquisition:</i> add images of the descending pass in addition to those originally used in the ascending pass.</li><li><i>Sample: </i>reduction in the size of the samples considered to 15 by 15 to facilitate evaluation at the central pixel.</li><li><i>Sample: </i>increased density of extracted windows, with a distance of approximately 500 meters between the centers of the windows.</li><li><i>Sample:</i> removal of the pre-processing involving the use of logarithms.</li><li><i>Sample:</i> removal of the pre-processing involving the normalisation.</li><li><i>Labels:</i> new structure for the labels part: dictionary with keys: `topography`, `metadata` and `physics`.</li><li><i>Labels:</i> `physics`: addition of direct information from the CROCUS model for 3 simulations: Liquid Water Content, snow height and minimum snowpack temperature.</li><li><i>Labels:</i> `topography`: information on the slope, altitude and average orientation of the sample.</li><li><i>Labels:</i> `metadata` : information on the date of the sample, the mountain massif and the run (ascending or descending).</li><li><i>Dataset</i>: removal of the train/test split*</li></ul><p>*We leave it up to the user to use the Group Kfold method to validate the models using the alpine massif information.</p><p>Finally, it consists of 2467516 samples of size 15 by 15 by 9. For each sample, the 9 metadata are provided, using in particular the <a href="https://www.umr-cnrm.fr/spip.php?article265&amp;lang=en">Crocus</a> physical model:</p><ul><li>topography:<ul><li>elevation (meters) (average),</li><li>orientation (degrees) (average),</li><li>slope (degrees) (average),</li></ul></li><li>metadata:<ul><li>name of the alpine massif,</li><li>date of acquisition,</li><li>type of acquisition (ascending/descending),</li></ul></li><li>physics<ul><li>Liquid Water Content (km/m2),</li><li>snow height (m),</li><li>minimum snowpack temperature (Celsius degree).</li></ul></li></ul><p>The 9 channels are in the following order:</p><ul><li>Sentinel-1 polarimetric channels: VV, VH and the combination C: VV/VH in linear,</li><li>Topographical features: altitude, orientation, slope</li><li>Polarimetric ratio with a reference summer image: VV/VVref, VH/VHref, C/Cref**</li></ul><p>** The reference image selected is that of August 9th 2020, as a reference image without snow (cf. <a href="https://ieeexplore.ieee.org/document/842004">Nagler&amp;al</a>)</p><p>An overview of the distribution and a summary of the sample statistics can be found in the file info.pdf.</p><p>The data is stored in .hdf5 format with gzip compression. We provide a python script to read and request the data. The script is dataset_load.py. It is based on the h5py, numpy and pandas libraries. It allows to select a part or the whole dataset using requests on the metadata. The script is documented and can be used as described in the README.md file</p><p>The processing chain is available at the following <a href="https://github.com/Matthieu-Gallet/LSD4WSD-dataset"><strong>Github</strong></a> address.</p><p>The authors would like to acknowledge the support from the National Centre for Space Studies (CNES) in providing computing facilities and access to SAR images via the PEPS platform.</p><p>The authors would like to deeply thank Mathieu Fructus for running the Crocus simulations.</p><p><strong>Erratum :</strong></p><p>In the dataloader file, the name of the "aquisition" column must be added twice, see the correction below.:</p><blockquote><p>dtst_ld = Dataset_loader(path_dataset,shuffle=False,descrp=["date","massif","aquisition","aquisition","elevation","slope","orientation","tmin","hsnow","tel",],)&nbsp;</p></blockquote><p>If you have any comments, questions or suggestions, please contact the authors:&nbsp;</p><ul><li>matthieu.gallet@univ-smb.fr</li><li>fatima.karbou@meteo.fr</li><li>abdourrahmane.atto@univ-smb.fr</li><li>emmanuel.trouve@univ-smb.fr</li></ul>

opencc-by-4.0Dec 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record