Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

4,612

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

4,612 results for “labeling”

Learn how ShareScore rates datasets ↗
zenodo44/100

Automatic labelling of HeLa "Kyoto" cells using Deep Learning tools

<p><strong>Name</strong>: Automatic labelling&nbsp;of HeLa &ldquo;Kyoto&rdquo; cells using Deep Learning tools</p> <p><strong>Data type</strong>: Microscopy images from the dataset &ldquo;<strong>HeLa &ldquo;Kyoto&rdquo; cells&nbsp;under the scope</strong>&rdquo;, Brightfield (BF), Digital Phase Contrast (DPC, either &ldquo;raw&rdquo; or &ldquo;square-rooted&rdquo;), Tubulin and H2B fluorescent channel, paired with their corresponding nuclei or cell/cyto label images.</p> <p><strong>Labels images</strong>: Labels images were generated using the script <em>&ldquo;prepare_trainingDataset_cellpose.ijm</em>&rdquo;.</p> <p>Briefly, for 5 defined time-points (1,10,50,100,150), channels of interest were duplicated, resaved and :</p> <p>-&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; nuclei label images were obtained using <a href="https://github.com/stardist/stardist">StarDist</a> on H2B channel</p> <p>-&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; cell label images were obtained using <a href="https://github.com/MouseLand/cellpose">Cellpose</a> on Tubulin and H2B channels</p> <p>A quick visual inspection of the resulting label images concluded that they were satisfying enough, despite certainly not being perfect.</p> <p>Notes :</p> <p>-&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; This labelling&nbsp;strategy:</p> <p>o&nbsp;&nbsp; will not produce 100% accurate labels, but they might be more reproducible than labels generated by humans and are (definitely) much faster to obtain.</p> <p>o&nbsp;&nbsp; is <strong>NOT a recommended way of generating labels images</strong>, but for educational purposes.</p> <p>-&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; The fluorescent channels are part of the dataset to ease the process of review of the labels and are NOT used for training. We generated the labels from the fluorescent channels to later predict labels from the BF or DPC channels only. As such, the fluorescent channels should not be &ldquo;reused&rdquo; with our labels during training.</p> <p><strong>File format</strong>: .tif (16-bit)</p> <p><strong>Image size</strong>: 540x540 (Pixel size: 0.299 nm)</p> <p>&nbsp;</p> <p><strong><em>NOTE</em></strong>: This dataset uses&nbsp;the &ldquo;HeLa &ldquo;Kyoto&rdquo; cells&nbsp;under the scope&rdquo; &nbsp;dataset (<a href="https://doi.org/10.5281/zenodo.6139958">https://doi.org/10.5281/zenodo.6139958</a>) to automatically generate annotations</p> <p><strong><em>NOTE</em></strong>: This dataset was used to train cellpose models in the following Zenodo entry&nbsp;<a href="https://doi.org/10.5281/zenodo.6140111">https://doi.org/10.5281/zenodo.6140111</a></p>

opencc-by-4.0Feb 2022View details →
zenodo44/100

Paddy Rice Mapping Learning Material(Sentinel-1 & labeling) in South Korea

<p>This dataset includes time series Sentinel-1 images and paddy rice labeling in South Korea for ML/DL model training. It consists of&nbsp;7,762 training patches and 5,180 validation patches for each patch consists of 256 x 256 pixels.&nbsp;The dataset is saved in hdf5 format&nbsp;separated into training/valdation data, image/labeling, and part number which can be accessed by key: {tr/va}_{im/lb}_{0~4}.</p> <p>According to the phonological stage of paddy rice, the Sentinel-1 images were acquired through 8-time steps&nbsp;from May 10 to October 20 in 20 days&rsquo; interval. In order for the images to capture similar features of rice invariant to more or less difference of growth, minimum and maximum value composite were used at transplanting season and ripening season each.&nbsp;The acquisition year for each patch varies from 2017 to 2019 since it was matched to that of labeling source.</p> <p>The paddy rice labeling is a rasterized version of farm map produced by Korean Ministry of Agriculture, Food and Rural Affairs(MAFRA). The original source data was produced by visual interpreted by high-resolution satellite images and aerial photos referring the other national GIS data and it is accessible through the national open data platform (<a href="http://data.nsdi.go.kr/dataset/20210707ds00001">http://data.nsdi.go.kr/dataset/20210707ds00001</a>). As the data is distributed in a vector format, it was converted to 10 m x 10 m raster format which is compatible to the Sentinel-1, and used for labeling the images.</p>

opencc-by-4.0Mar 2022View details →
zenodo44/100

A benchmark dataset of herbarium specimen images with label data: Summary

<p>This landing page contains a CSV file compiling all data associated with herbarium specimens that are part of this dataset, as they could be found on GBIF, JACQ or FinBIF. A CSV file with and without Darwin Core extension data is available, as some CSV readers have trouble with the JSON format that is used for those extensions.</p> <p>In addition, DOI&#39;s of the individual specimens uploaded to Zenodo and direct links to the different files (JPEG, TIFF, JSON, PNG) are also included. Index of these added&nbsp;variables:</p> <p>- persistentID: Persistent Identifier of the collection specimen. Data uploaded as part of this dataset will not be kept in sync with changes at the collection&#39;s repository. Hence, this URI will always point to the most up to date information known about the herbarium specimen.</p> <p>- jpegURL, tiffURL, jsonURL: URL&#39;s pointing straight to the respective image and data files themselves, to facilitate (selective) batch downloads.</p> <p>- pngSegAllURL and pngSegSelURL: Segmented overlays of the herbarium specimens indicating the location of different labels and reference material on the sheet (&quot;All&quot;) and their content (&quot;Sel&quot;). More information can be found in the paper (in prep) associated with this data publication and the individual depositions themselves.</p> <p>- DOI: The DOI of the deposition of images and data of these specimens on Zenodo. DOI&#39;s point to the most up-to-date version of these depositions at the time of the publication of this CSV file. As a rule, this CSV file will be updated should any changes happen to any of the depositions.</p> <p>- jpegURL2, tiffURL2: A few herbarium sheets had labels on the back and consisted therefore of two scans. As a rule, the label scans are in this category.</p>

opencc-zeroNov 2018View details →
zenodo44/100

Labeled songs of domestic canary M1-2016-spring (Serinus canaria)

<p><strong>Labeled songs of domestic canary M1-2016-spring (Serinus canaria)</strong></p> <p><em>J. Giraudon*<sup>123</sup>, N. Trouvain*<sup>123</sup>, A. Cazala<sup>4</sup>, C. Del Negro<sup>4</sup>, X. Hinaut<sup>123</sup></em></p> <p><sup>1</sup> Inria Bordeaux Sud-Ouest, France</p> <p><sup>2</sup> LaBRI, Bordeaux INP, CNRS, UMR 5800, France</p> <p><sup>3</sup> Institut des Maladies Neurog&eacute;g&eacute;n&eacute;ratives, Universit&eacute; de Bordeaux, CNRS, UMR 5293, France</p> <p><sup>4 </sup>Paris-Saclay University,&nbsp;UMR 9197&nbsp;CNRS, Paris-Saclay Institute of Neuroscience, France&nbsp;</p> <p><em>* these authors participated equally to this work.</em></p> <p><strong>General information</strong></p> <p>This dataset contains ~3h of labeled songs (459 songs) of one male canary (called M1) recorded between May 24th and June 15th 2016. Songs were recorded in a sound-isolation chamber using a RODE M3 microphone, an external sound card for microphone amplification (M-Audio Fast Track Ultra 8R), and the software Sound Analysis Pro 2011 (SAP). SAP parameters were set with conservative thresholds (software threshold to 4-6) in order to record the initiation of canary&#39;s songs which can be low in volume.</p> <p>Songs were hand labelled by one human expert using Audacity. They were then checked and corrected by another human expert assisted by an automated program based on recurrent neural networks (see References).</p> <p><strong>Dataset description</strong></p> <p>Canary songs are labeled using 27 different identified syllable classes + 1 &quot;call&quot; class identifying simple off-song calls + 1 &quot;TRASH&quot; class for irrelevant sounds (very rare vocalizations or non-bird sounds) + 1 &quot;SIL&quot; class for silence between vocalizations. Songs are annotated at the phrase level: a phrase consists of a repetition of a single syllable type and each phrase type is assigned a label.</p> <p>Annotations are provided in CSV format in the &quot;M1-2016-spring_csv_annotations.zip&quot; archive. There is one file per song, containing:</p> <ul> <li>a &quot;wave&quot; column indicating the song&#39;s audio filename;</li> <li>&quot;start&quot; and &quot;end&quot; columns indicating the temporal delimitation of the label from the begining of the song, in seconds;</li> <li>a &quot;syll&quot; column indicating the labels.</li> </ul> <p>Annotations are also provided in <a href="https://manual.audacityteam.org/man/importing_and_exporting_labels.html">Audacity TXT&nbsp;format</a>&nbsp;in the &quot;M1-2016-spring_audacity_annotations.zip&quot; archive. There is one file per song, containing&nbsp;three tabulation-separated&nbsp;columns. The first two column indicates the temporal delimitation (start and end) of the phrase from the begining of the song. The thrid one contains&nbsp;the associated label. Annotations filenames match corresponding song audio filename.</p> <p>Songs are provided in WAV format (44kHz sampling rate) in the &quot;M1-2016-spring_audio.zip&quot; archive. There is one file per song: audio filenames match corresponding annotation filenames.</p> <p><strong>References</strong></p> <p>This dataset was used in:</p> <p>N. Trouvain, X. Hinaut (2021) Canary Song Decoder: Transduction and Implicit Segmentation with ESNs and LTSMs. HAL preprint <a href="https://hal.inria.fr/hal-03203374">&lang;hal-03203374&rang;</a></p>

opencc-by-4.0May 2021View details →
zenodo44/100

An annotated high-content fluorescence microscopy dataset with Hoechst 33342-stained nuclei and manually labelled outlines

<p>Here we present a benchmarking dataset of fluorescence microscopy images with Hoechst 33342-stained nuclei together with annotations of nuclei, nuclear fragments and micronuclei. Images were randomly selected from an RNA interference screen with a modified U2OS osteosarcoma cell line, acquired on a Thermo Fischer CX7 high-content imaging system at 20x magnification. Labelling was performed by a single annotator and reviewed by a biomedical expert.</p> <p>The dataset contains 50 images showing over 2000 labelled nuclear objects in total, which is sufficiently large to train well-performing neural networks for instance or semantic segmentation. It is pre-split into training, development and test set, each in a zip file. The dataset should be referred to as Aitslab_bioimaging1. A brief article describing the dataset is also available (Arvidsson M, Kazemi Rashed S, Aits S. <a href="https://doi.org/10.1016/j.dib.2022.108769">10.1016/j.dib.2022.108769</a> )</p> <p><strong>Dataset description:</strong></p> <p>Fluorescence microscopy images: original .C01 files and files converted to 8-bit .png format (Grayscale)</p> <p>Annotations: 24-bit .png format (RGB)</p> <p>Script used to convert C01 to png images:&nbsp;C01_to_png.py file with python code and readme.md file with instructions to run it</p>

opencc-by-4.0Jun 2022View details →
zenodo44/100

Urban Sound & Sight (Urbansas) - Labeled set

<p><strong>Urban Sound &amp; Sight (Urbansas):&nbsp;</strong></p> <p>Version 1.0, May 2022</p> <p><strong>Created by</strong><br> Magdalena Fuentes (1, 2), Bea Steers (1, 2), Pablo Zinemanas (3), Mart&iacute;n Rocamora (4), Luca Bondi (5), Julia Wilkins (1, 2), Qianyi Shi (2), Yao Hou (2), Samarjit Das (5), Xavier Serra (3), Juan Pablo Bello (1, 2)<br> 1. Music and Audio Research Lab, New York University<br> 2. Center for Urban Science and Progress, New York University<br> 3. Universitat Pompeu Fabra, Barcelona, Spain<br> 4. Universidad de la Rep&uacute;blica, Montevideo, Uruguay<br> 5. Bosch Research, Pittsburgh, PA, USA</p> <p><strong>Publication</strong></p> <p>If using this data in academic work, please cite the following paper, which presented this dataset:<br> M. Fuentes, B. Steers, P. Zinemanas, M. Rocamora, L. Bondi, J. Wilkins, Q. Shi, Y. Hou, S. Das, X. Serra, J. Bello. &ldquo;Urban Sound &amp; Sight: Dataset and Benchmark for Audio-Visual Urban Scene Understanding&rdquo;. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2022.</p> <p><strong>Description</strong></p> <p>Urbansas is a dataset for the development and evaluation of machine listening systems for audiovisual spatial urban understanding. One of the main challenges to this field of study is a lack of realistic, labeled data to train and evaluate models on their ability to localize using a combination of audio and video.<br> We set four main goals for creating this dataset:&nbsp;<br> 1. To compile a set of real-field audio-visual recordings;<br> 2. The recordings should be stereo to allow exploring sound localization in the wild;<br> 3. The compilation should be varied in terms of scenes and recording conditions to be meaningful for training and evaluation of machine learning models;<br> 4. The labeled collection should be accompanied by a bigger unlabeled collection with similar characteristics to allow exploring self-supervised learning in urban contexts.<br> Audiovisual data<br> We have compiled and manually annotated Urbansas from two publicly available datasets, plus the addition of unreleased material. The public datasets are the TAU Urban Audio-Visual Scenes 2021 Development dataset (street-traffic subset) and the Montevideo Audio-Visual Dataset (MAVD):</p> <p><br> Wang, Shanshan, et al. &quot;A curated dataset of urban scenes for audio-visual scene analysis.&quot; ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2021.</p> <p>Zinemanas, Pablo, Pablo Cancela, and Mart&iacute;n Rocamora. &quot;MAVD: A dataset for sound event detection in urban environments.&quot; Detection and Classification of Acoustic Scenes and Events, DCASE 2019, New York, NY, USA, 25&ndash;26 oct, page 263--267 (2019).</p> <p><br> The TAU dataset consists of 10-second segments of audio and video from different scenes across European cities, traffic being one of the scenes. Only the scenes labeled as traffic were included in Urbansas. MAVD is an audio-visual traffic dataset curated in different locations of Montevideo, Uruguay, with annotations of vehicles and vehicle components sounds (e.g. engine, brakes) for sound event detection. Besides the published datasets, we include a total of 9.5 hours of unpublished material recorded in Montevideo, with the same recording devices of MAVD but including new locations and scenes.</p> <p>Recordings for TAU were acquired using a GoPro Hero 5 (30fps, 1280x720) and a Soundman OKM II Klassik/studio A3 electret binaural in-ear microphone with a Zoom F8 audio recorder (48kHz, 24 bits, stereo). Recordings for MAVD were collected using a GoPro Hero 3 (24fps, 1920x1080) and a SONY PCM-D50 recorder (48kHz, 24 bits, stereo).&nbsp;</p> <p>When compiled in Urbansas, it includes 15 hours of stereo audio and video, stored in separate 10 second MPEG4 (1280x720, 24fps) and WAV (48kHz, 24 bit, 2 channel) files. Both released video datasets are already anonymized to obscure people and license plates, the unpublished MAVD data was anonymized similarly using this anonymizer. We also distribute the 2fps video used for producing the annotations.</p> <p>The audio and video files both share the same filename stem, meaning that they can be associated after removing the parent directory and extension.</p> <p>MAVD:<br> video/&lt;location_id&gt;_&lt;mavd_clip_id&gt;_&lt;clip_split_id&gt;.mp4<br> audio/&lt;location_id&gt;_&lt;mavd_clip_id&gt;_&lt;clip_split_id&gt;.wav</p> <p>TAU:<br> video/&lt;location_id&gt;_&lt;tau_clip_id&gt;.mp4<br> audio/&lt;location_id&gt;_&lt;tau_clip_id&gt;.wav</p> <p><br> where location_id in both cases includes the city and an ID number.</p> <p><br> &nbsp; &nbsp; &nbsp; city &amp; &nbsp;places &amp; &nbsp;clips &amp; &nbsp;mins &amp; &nbsp;frames &amp; &nbsp;labeled mins &nbsp; &nbsp;\\<br> Montevideo &amp; &nbsp; &nbsp; &nbsp; 8 &amp; &nbsp; 4085 &amp; &nbsp; 681 &amp; &nbsp;980400 &amp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;92 \\<br> &nbsp;Stockholm &amp; &nbsp; &nbsp; &nbsp; 3 &amp; &nbsp; &nbsp; 91 &amp; &nbsp; &nbsp;15 &amp; &nbsp; 21840 &amp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 2 \\<br> &nbsp;Barcelona &amp; &nbsp; &nbsp; &nbsp; 4 &amp; &nbsp; &nbsp;144 &amp; &nbsp; &nbsp;24 &amp; &nbsp; 34560 &amp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;24 \\<br> &nbsp; Helsinki &amp; &nbsp; &nbsp; &nbsp; 4 &amp; &nbsp; &nbsp;144 &amp; &nbsp; &nbsp;24 &amp; &nbsp; 34560 &amp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;16 \\<br> &nbsp; &nbsp; Lisbon &amp; &nbsp; &nbsp; &nbsp; 4 &amp; &nbsp; &nbsp;144 &amp; &nbsp; &nbsp;24 &amp; &nbsp; 34560 &amp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;19 \\<br> &nbsp; &nbsp; &nbsp; Lyon &amp; &nbsp; &nbsp; &nbsp; 4 &amp; &nbsp; &nbsp;144 &amp; &nbsp; &nbsp;24 &amp; &nbsp; 34560 &amp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 6 \\<br> &nbsp; &nbsp; &nbsp;Paris &amp; &nbsp; &nbsp; &nbsp; 4 &amp; &nbsp; &nbsp;144 &amp; &nbsp; &nbsp;24 &amp; &nbsp; 34560 &amp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 2 \\<br> &nbsp; &nbsp; Prague &amp; &nbsp; &nbsp; &nbsp; 4 &amp; &nbsp; &nbsp;144 &amp; &nbsp; &nbsp;24 &amp; &nbsp; 34560 &amp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 2 \\<br> &nbsp; &nbsp; Vienna &amp; &nbsp; &nbsp; &nbsp; 4 &amp; &nbsp; &nbsp;144 &amp; &nbsp; &nbsp;24 &amp; &nbsp; 34560 &amp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 6 \\<br> &nbsp; &nbsp; London &amp; &nbsp; &nbsp; &nbsp; 5 &amp; &nbsp; &nbsp;144 &amp; &nbsp; &nbsp;24 &amp; &nbsp; 34560 &amp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 4 \\<br> &nbsp; &nbsp; &nbsp;Milan &amp; &nbsp; &nbsp; &nbsp; 6 &amp; &nbsp; &nbsp;144 &amp; &nbsp; &nbsp;24 &amp; &nbsp; 34560 &amp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 6 \\<br> \midrule<br> &nbsp; &nbsp; &nbsp;Total &amp; &nbsp; &nbsp; &nbsp;50 &amp; &nbsp; 5472 &amp; &nbsp; 912 &amp; 1.3M &amp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 180 \\</p> <p><br> <strong>Annotations</strong></p> <p><br> Of the 15 hours of audio and video, 3 hours of data (1.5 hours TAU, 1.5 hours MAVD) are manually annotated by our team both in audio and image, along with 12 hours of unlabeled data (2.5 hours TAU, 9.5 hours of unpublished material) for the benefit of unsupervised models. The distribution of clips across locations was selected to maximize variance across different scenes. The annotations were collected at 2 frames per second (FPS) as it provided a balance between temporal granularity and clip coverage.</p> <p>The annotation data is contained in video_annotations.csv and audio_annotations.csv.&nbsp;</p> <p><strong>Video Annotations</strong></p> <p>Each row in the video annotations represents a single object in a single frame of the video. The annotation schema is as follows:</p> <ul> <li>frame_id: The index of the frame within the clip the annotation is associated with. This index is 0-based and goes up to 19 (assuming 10-second clips with annotations at 2 FPS)</li> <li>track_id: The ID of the detected instance that identifies the same object across different frames. These IDs are guaranteed to be unique within a clip.</li> <li>x, y, w, h: The top-left corner and width and height of the object&rsquo;s bounding box in the video. The values are given in absolute coordinates with respect to the image size (1280x720).&nbsp;</li> <li>class_id: The index of the class corresponding to: [0, 1, 2, 3, -1] &mdash; see label for the index mapping. The -1 value corresponds to the case where there are no events, but still clip-level annotations, like night and city. When operating on bounding boxes, class_id of -1 should be filtered.</li> <li>label: The label text. This is equivalent to LABELS[class_id], where LABELS=[car, bus, motorbike, truck, -1]. The label -1 has the same role as above.</li> <li>visibility: The visibility of the object. This is 1 unless the object becomes obstructed, where it changes to 0.</li> <li>filename: The file ID of the associated file. This is the file&rsquo;s path minus the parent directory and extension.</li> <li>city: The city where the clip was collected in.</li> <li>location_id: The specific name of the location. This may include an integer ID following the city name for cases where there are multiple collection points.</li> <li>time: The time (in seconds) of the annotation, relative to the start of the file. Equivalent to frame_id / fps .</li> <li>night: Whether the clip takes place during the day or at night. This value is singular per clip.</li> <li>subset: Which data source the data originally belongs to (TAU or MAVD).</li> </ul> <p><strong>Audio Annotations</strong></p> <p>Each row represents a single object instance, along with the time range that it exists within the clip. The annotation schema is as follows:</p> <ul> <li>filename: The file ID odd the associated audio file. See filename above.&nbsp;</li> <li>class_id, label: See above. Audio has an additional class_id of 4 (label=offscreen) which indicates an off-screen vehicle - meaning a vehicle that is heard but not seen. A class_id of -1 indicates a clip-level annotation for a clip that has no object annotations (an empty scene).</li> <li>non_identifiable_vehicle_sound: True if the region contains the sound of vehicles where individual instances cannot be uniquely identified.&nbsp;</li> <li>start, end: The start and end times (in seconds) of the annotation relative to the file.&nbsp;</li> </ul> <p><strong>Conditions of use</strong></p> <p>Dataset created by Magdalena Fuentes, Bea Steers, Pablo Zinemanas, Mart&iacute;n Rocamora, Luca Bondi, Julia Wilkins, Qianyi Shi, Yao Hou, Samarjit Das, Xavier Serra, and Juan Pablo Bello.</p> <p>The Urbansas dataset is offered free of charge under the following terms:</p> <ul> <li>Urbansas annotations are release under the CC BY 4.0 license</li> <li>Urbansas video and audio replicates the original sources licenses: <ul> <li>&nbsp; &nbsp;MAVD subset is released under &nbsp;CC BY 4.0&nbsp;</li> <li>&nbsp; &nbsp;TAU subset is released under a Non-Commercial license</li> </ul> </li> </ul> <p><strong>Feedback</strong></p> <p>Please help us improve Urbansas by sending your feedback to:</p> <ul> <li>Magdalena Fuentes: mfuentes@nyu.edu</li> <li>Bea Steers: bsteers@nyu.edu&nbsp;</li> </ul> <p>In case of a problem, please include as many details as possible.</p> <p><strong>Acknowledgments</strong></p> <p>This work was partially supported by the National Science Foundation award 1955357 and Bosch RTC.</p>

opencc-by-4.0May 2022View details →
zenodo44/100

Historic manuscript page images with noisy labels

<p>Images of digitised manuscript pages sourced from <a href="https://iiif.biblissima.fr/collections/">https://iiif.biblissima.fr/collections/</a>. This dataset aims to facilitate experiments using existing data/metadata to train computer vision models. In particular, using &#39;noisy&#39; labels in some capacity.</p> <p>Each image is taken from a page of a manuscript listed on <a href="https://iiif.biblissima.fr/collections/">https://iiif.biblissima.fr/collections/</a>. Each example includes the labels included in the IIIF manifests for these images. The data includes the following columns:</p> <ul> <li>image: an IIIF URL for the image</li> <li>manifest_url: A URL for the IIIF manifest for the image</li> <li>license: for each image</li> <li>label: the text found in the manifest &#39;label&#39; field.</li> <li>attribution: which institution the image comes from</li> </ul> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Jun 2022View details →
zenodo44/100

Dataset of Open-Source Software Developers Labeled by their Experience Level and Associated with their Software Metrics

<p>This dataset contains 703&nbsp;anonymized developers extracted from 17 open-source projects from GitHub. Projects were chosen because they use:</p> <ul> <li>the Java programming language</li> <li>the <a href="https://spring.io/projects/spring-framework">Spring framework</a></li> <li><a href="https://maven.apache.org/">Maven</a> / <a href="https://gradle.org/">Gradle</a> build tools</li> </ul> <p>For all these developers, 23 software metrics were calculated for each project to which they contribute. These metrics are either calculated by analyzing the source code or relative to project management metadata. Each of these developers then have been manually annotated. To do this, developers have been searched&nbsp; for in professionnal social media such as:</p> <ul> <li><a href="https://www.linkedin.com/">Linkedin</a></li> <li><a href="https://twitter.com/">Twitter</a></li> <li><a href="https://github.com/">Github</a></li> </ul> <p><strong>This dataset is published in the following journal article: </strong></p> <p><strong>Dataset of Open-Source Software Developers Labeled by their Experience Level in the Project and their Associated Software Metrics, Q. Perez, C. Urtado and </strong><strong>S. Vauttier, Data In Brief, </strong></p> <p><a href="https://www.sciencedirect.com/science/article/pii/S2352340922010459">https://www.sciencedirect.com/science/article/pii/S2352340922010459</a></p>

opencc-by-4.0Aug 2022View details →
zenodo44/100

Mars Target Encyclopedia - Labeled LPSC abstracts for four Mars missions

<p>This data set contains annotated text versions of 1635 two-page abstracts published at the Lunar and Planetary Science Conference from 1998&nbsp;to 2020 of relevance to four Mars missions.&nbsp; The annotations were generated using named entity recognition and relation extraction provided by the MTE processing pipeline (available at&nbsp;https://github.com/wkiri/MTE), followed by manual review.&nbsp; Annotated entities include Element, Mineral, Property, and Target.&nbsp; Annotated relations include <strong>Contains</strong>(Target, Element | Mineral) and <strong>HasProperty</strong>(Target, Property).&nbsp; The extracted&nbsp;information (without full texts) is also available as a database (stored in .csv files) at&nbsp;https://pds-geosciences.wustl.edu/missions/mte/mte.htm .&nbsp;The complete annotated texts are provided here as a resource for further research and experimentation on&nbsp;information extraction methods.&nbsp; For more information about the Mars Target Encyclopedia and these annotations, please see:</p> <ul> <li>&quot;<a href="https://www.hou.usra.edu/meetings/lpsc2022/pdf/1231.pdf">Targets from the Spirit Mars Exploration Rover in the Mars Target Encyclopedia</a>&quot;,&nbsp;Kiri L. Wagstaff, Raymond Francis, Matthew Golombek, Steven Lu, Ellen Riloff, Leslie Tamppari, Yuan Zhuang, and Thomas Stein.<br> <em>53rd Lunar and Planetary Science Conference</em>, Abstract #1231, March 2022.</li> <li>&quot;<a href="https://www.hou.usra.edu/meetings/lpsc2021/pdf/1278.pdf">The Mars Target Encyclopedia Now Includes Mars Pathfinder and Mars Phoenix Targets</a>&quot;,&nbsp;Kiri L. Wagstaff, Raymond Francis, Matthew Golombek, Steven Lu, Ellen Riloff, Leslie Tamppari, and Thomas C. Stein.<br> <em>52nd Lunar and Planetary Science Conference</em>, Abstract #1278, March 2021.</li> </ul> <p>The original PDF abstracts are available at:&nbsp;</p> <ul> <li>For years prior to 2000:&nbsp; https://www.lpi.usra.edu/meetings/LPSC${two-digit-year}/pdf/${id}.pdf</li> <li>For year 2000:&nbsp; https://www.lpi.usra.edu/meetings/LPSC${four-digit-year}/pdf/${id}.pdf</li> <li>For years 2001-2017 (note lower-case lpsc):&nbsp; https://www.lpi.usra.edu/meetings/lpsc${four-digit-year}/pdf/${id}.pdf</li> <li>For years 2018-2020:&nbsp; https://www.hou.usra.edu/meetings/lpsc${four-digit-year}/pdf/${id}.pdf</li> </ul> <p>where ${id} is a four-digit abstract number, starting with 1001 (if available).</p> <p>The text files provided in this archive were extracted from the PDF files using the Apache Tika PDF parsing tool.&nbsp; They are named as ${four-digit-year}_${id}.txt.&nbsp; The text is provided here so that the annotations can be viewed in context.&nbsp; The text content remains copyright of the original abstract authors.</p> <p>The annotations (entities and relations) are provided in the format used by the brat annotation tool.&nbsp; They are named as ${four-digit-year}_${id}.ann. To view the annotations in a web-based graphical form, install the brat tool (http://brat.nlplab.org/).&nbsp; These annotations were generated using brat v1.3.&nbsp; The annotation files are also human-readable and can be parsed in to be used directly in code.&nbsp; If the .ann file is empty, then there are no relevant annotations for the associated text file.</p> <p><strong>Contents</strong>:</p> <ul> <li>mpf.zip: 591 abstracts relating to the Mars Pathfinder mission (1998-2020)</li> <li>mer-a.zip: 397 abstracts relating to the MER-A (Spirit) rover mission (2004-2020)</li> <li>mer-b.zip: 256 abstracts relating to the MER-B (Opportunity) rover mission (2005-2020)</li> <li>phx.zip: 391 abstracts relating to the Mars Phoenix Lander mission (2009-2020)</li> </ul> <p>Each directory contains a .txt and .ann file for each abstract.&nbsp; The .ann file is in brat standoff format (http://brat.nlplab.org/standoff.html).&nbsp; Additional .conf files are provided to generate color highlighting and keyboard shortcuts.&nbsp; These are used by the brat tool.</p> <p>Note: the same abstract may appear in more than one mission directory, if it discusses targets from more than one mission.&nbsp; It will have a different .ann file for each such appearance.&nbsp; Within each directory, a&nbsp;&quot;Target&quot; annotation is understood to refer to a target of the relevant mission.</p> <p><strong>Attribution</strong>:</p> <p>If you use this data set in your own work, please cite it as follows:</p> <p>Kiri L. Wagstaff, Raymond Francis, Matthew Golombek, Leslie Tamppari, and Steven Lu. (2022).&nbsp;Mars Target Encyclopedia - Labeled LPSC abstracts for four Mars missions&nbsp;(1.0.0.0) [Data set]. Zenodo. DOI: 10.5281/zenodo.7066107</p>

opencc-by-4.0Sep 2022View details →
zenodo44/100

Mars orbital image (HiRISE) labeled data set

<p>This data set contains 3820 landmarks that were extracted from 168 HiRISE images. The landmarks were detected in HiRISE browse images. For each landmark, we cropped a square bounding box the included the full extent of the landmark plus a 30-pixel margin to left, right, top, and bottom. Each cropped image was then resized to 227x227 pixels.</p> <p><strong>Contents</strong>:</p> <ul> <li>map-proj/: Directory containing individual cropped landmark images</li> <li>labels-map-proj.txt: Class labels (ids) for each landmark image</li> <li>landmark_mp.py: Python dictionary that maps class ids to semantic names</li> </ul> <p><strong>Attribution</strong>:</p> <p>If you use this data set in your own work, please cite this DOI: 10.5281/zenodo.1048301</p> <p>Please also cite this paper, which provides additional details about the data set.</p> <p>Kiri L. Wagstaff, You Lu, Alice Stanboli, Kevin Grimes, Thamme Gowda, and Jordan Padams. &quot;Deep Mars: CNN Classification of Mars Imagery for the PDS Imaging Atlas.&quot; <em>Proceedings of the Thirtieth Annual Conference on Innovative Applications of Artificial Intelligence</em>, 2018.</p> <p>&nbsp;</p>

opencc-by-sa-4.0Nov 2017View details →
zenodo44/100

Mars Target Encyclopedia - LPSC abstracts labeled data set

<p>This data set contains annotated text versions of 2-page abstracts published at the Lunar and Planetary Science Conference in 2015 and 2016.</p> <p>The original PDF abstracts are available at:</p> <ul> <li>https://www.hou.usra.edu/meetings/lpsc2015/programAbstracts/view/</li> <li>https://www.hou.usra.edu/meetings/lpsc2016/programAbstracts/view/</li> </ul> <p>The text files in this archive were extracted using the Apache Tika PDF parsing tool.  The text is provided here so that the annotations can be viewed.  The text content remains copyright of the original abstract authors.</p> <p>The annotations (entities and relations) are provided in the format used by the brat annotation tool.  To view the annotations in a web-based graphical form, install the brat tool (http://brat.nlplab.org/).  These annotations were generated using brat v1.3.  The annotation files are also human-readable and can be parsed in to be used directly in code.</p> <p><strong>Contents</strong>:</p> <ul> <li>lpsc15/: 62 abstracts</li> <li>lpsc16/: 55 abstracts</li> </ul> <p>Each directory contains a .txt and .ann file for each abstract.  The .ann file is in brat standoff format (http://brat.nlplab.org/standoff.html).</p> <p>Additional .conf files are provided to generate color highlighting and keyboard shortcuts.  These are used by the brat tool.</p> <p><strong>Attribution</strong>:</p> <p>If you use this data set in your own work, please cite this DOI:</p> <p>10.5281/zenodo.1048419</p> <p>Please also cite this paper, which provides additional details about the data set.</p> <p>Kiri L. Wagstaff, Raymond Francis, Thamme Gowda, You Lu, Ellen Riloff, Karanjeet Singh, and Nina Lanza. "Mars Target Encyclopedia: Rock and Soil Composition Extracted from the Literature."  <em>Proceedings of the Thirtieth Annual Conference on Innovative Applications of Artificial Intelligence</em>, 2018.</p>

opencc-by-sa-4.0Nov 2017View details →
zenodo44/100

Mars surface image (Curiosity rover) labeled data set

<p>This data set consists of 6691 images spanning 24 classes that were collected by the Mars Science Laboratory (MSL, Curosity) rover by three instruments (Mastcam Right eye, Mastcam Left eye, and MAHLI).&nbsp; These images are the &quot;browse&quot; version of each original data product, not full resolution.&nbsp; They are roughly 256x256 pixels each.</p> <p>We divided the MSL images into train, validation, and test data sets according to their sol (Martian day) of acquisition.&nbsp; This strategy was chosen to model how the system will be used operationally with an image archive that grows over time.&nbsp; The images were collected from sols 3 to 1060 (August 2012 to July 2015).&nbsp; The exact train/validation/test splits are given in individual files.&nbsp; Full-size images can be obtained from the PDS at https://pds-imaging.jpl.nasa.gov/search/ .</p> <p><strong>Contents</strong>:</p> <ul> <li>calibrated/: Directory containing calibrated MSL images</li> <li>train-calibrated-shuffled.txt: Training labels (images in shuffled order)</li> <li>val-calibrated-shuffled.txt: Validation labels</li> <li>test-calibrated-shuffled.txt: Test labels</li> <li>msl_synset_words-indexed.txt: Mapping from class IDs to class names</li> </ul> <p><strong>Attribution</strong>:</p> <p>If you use this data set in your own work, please cite this DOI:</p> <p>10.5281/zenodo.1049137</p> <p>Please also cite this paper, which provides additional details about the data set.</p> <p>Kiri L. Wagstaff, You Lu, Alice Stanboli, Kevin Grimes, Thamme Gowda, and Jordan Padams. &quot;Deep Mars: CNN Classification of Mars Imagery for the PDS Imaging Atlas.&quot; <em>Proceedings of the Thirtieth Annual Conference on Innovative Applications of Artificial Intelligence</em>, 2018.</p>

opencc-by-sa-4.0Nov 2017View details →
zenodo44/100

Post-hoc labeling of arbitrary EEG recordings for data-efficient evaluation of neural decoding methods

<p>EEG signals&nbsp;recorded from seven healthy subjects. On average,&nbsp;Seventy-three minutes of EEG data&nbsp;were recorded&nbsp;from 31 electrodes placed according to the extended 10-20 system. Signals are used in the paradigm-agnostic post-hoc labeled dataset generation framework for benchmarking of oscillatory neural decoding methods.</p>

opencc-by-4.0Nov 2017View details →
zenodo44/100

Measure While Drilling (MWD) dataset with rock type labels for 15 Norwegian hard rock tunnels

<p>The dataset is presented in the paper:&nbsp;</p> <p><em>Building and analysing a labelled Measure While Drilling dataset from 15 hard rock tunnels in Norway</em>, by&nbsp;T.F. Hansen, Z. Liu, J. Torressen</p> <p>The paper has a preprint on SSRN: <a href="http://dx.doi.org/10.2139/ssrn.4729646" target="_blank" rel="noopener">http://dx.doi.org/10.2139/ssrn.4729646</a>&nbsp;and is under review in a peer-reviewed journal.</p> <p>The dataset is utilised in a machine learning analysis in the paper:</p> <p><em>Predicting rock type from MWD tunnel data using a reproducible ML-modelling process</em>, by T.F. Hansen, Z. Liu, J. Torressen</p> <p>The paper is published in the journal <em>Tunnelling and Underground Space Technology</em>:&nbsp;</p> <p><a href="https://doi.org/10.1016/j.tust.2024.105843">https://doi.org/10.1016/j.tust.2024.105843</a></p> <p>&nbsp;</p> <p><strong>Description of the dataset:</strong></p> <p>Measure While Drilling (MWD) is a technique in rock drilling, mainly used in drill and blast tunnelling, where data about the rock mass is registered by sensors while drilling. The extensive and geologically diversified dataset contains corresponding MWD-data and rock mass mappings for 5205 blasting rounds from 15 hard rock tunnels in Norway. MWD-data are presented as tabular data. 10 different rocktypes are the corresponding labels.</p> <p>Four files are given:</p> <ul> <li>A csv-file of the training dataset - with outliers removed</li> <li>A csv-file of the testing dataset (split train/test 0.75/0.25) - with outliers removed</li> <li>A csv-file with the full unsplitted dataset, cleaned and with outliers removed</li> <li>A csv-file with the raw dataset, before cleaning, processing and outlier removal</li> </ul> <p>The author gratefully acknowledge the tunnel software/hardware company Bever Control, which have facilitated data from the clients Bane NOR, Statens Vegvesen, Nye Veier, and the contractor AF-Gruppen.</p> <p>&nbsp;</p> <p><strong>NOTE:</strong> The dataset is only available for research, no commercial use.</p>

opencc-by-4.0May 2024View details →
zenodo44/100

OME-Zarr 3D hiPSCs with 3D labels & 3D measurements, 2x2 field of views

<p>These are 2 small OME-Zarr files of the data from&nbsp;<a href="../records/7057076">10.5281/zenodo.7057076</a>.</p> <p>The images have been processed using <a href="https://fractal-analytics-platform.github.io/">Fractal</a>, the workflow is attached as a json file. It ran with fractal-server==2.3.6, fractal-client==2.0.1, fractal-web==1.4.0 and fractal-tasks-core==1.2.1.</p> <p>Both Zarr files are Zip-compressed to allow easier upload &amp; download from Zenodo.&nbsp;</p> <p>20200812-CardiomyocyteDifferentiation14-Cycle1.zarr contains 3 3D channels, a nuclear segmentation produced by&nbsp;<a href="https://cellpose.readthedocs.io/en/latest/">cellpose</a> as labels and 4 tables: A ROI table for the whole well, a ROI table for the 4 field of views, a masking ROI table for the nuclear segmentation, as well as measurements performed with <a href="https://github.com/haesleinhuepf/napari-skimage-regionprops">napari-skimage-regionprops</a>.</p> <p>20200812-CardiomyocyteDifferentiation14-Cycle1_mip.zarr contains the same 3 channels, but as maximum intensity projections. It contains nuclear segmentation through cellpose, as well as 3 more labels generated by napari workflows (different thresholds, less accurate segmentations). It also contains 7 tables: The region of interests like in the 3D data, as well as measurements performed with <a href="https://github.com/haesleinhuepf/napari-skimage-regionprops">napari-skimage-regionprops</a>.</p> <p>The tables are stored in the OME-Zarr file according to the <a href="https://fractal-analytics-platform.github.io/fractal-tasks-core/tables/">Fractal table specification</a>&nbsp;spec in AnnData.</p> <p>The 3 channels are:</p> <p>- 0: DAPI, nuclear stain</p> <p>- 1: nanog, antibody staining with&nbsp;Bio-Techne AG, AF1997-SP, Lot&nbsp;KKJ0617121 for the stemness marker nanog</p> <p>- 2: Lamin B1, antibody staining with Abcam, ab16048, Lot GR3244890-2 for the nuclear envelope marker Lamin B1</p>

opencc-by-4.0Feb 2023View details →
zenodo44/100

Cholec80-Boxes: Bounding-Box Labels for Surgical Tools in Five Cholecystectomy Videos

<p>&nbsp;</p> <p>The dataset is descriped in a pending publication titled "Cholec80-Boxes: Bounding-box Labeling Data for Surgical Tools in Cholecystectomy Images". The dataset was used in the following studies titled:</p> <ul> <li>"Surgical tool classification &amp; localisation using attention and multi-feature fusion deep learning approach".</li> <li>"Laparoscopic video analysis using temporal, attention, and multi-feature fusion based-approaches".</li> <li>"Analysing attention convolutional neural network for surgical tool localisation: A feasibility study".</li> </ul> <p>The dataset consists of cholecystectomy images and bounding-box labels for surgical tools. These images were extracted from five videos of the Cholec80 dataset (Twinanda et al., 2016) at a rate of 1 Hz. The images are stored in '.png' format with a resolution of 854*480 pixels. Each video&rsquo;s images are organized in a separate folder. The labeling data are stored in a CSV file, which contains the region of interest (ROI) labels for each surgical tool visible in the extracted images. Additionally, the CSV file provides information about each labeled image. Table 1 presents a content description of the 'ROI_Labels.csv' file.</p> <p><strong>Table 1:</strong> Description of 'ROI_Labels.csv' file.</p> <table> <tbody> <tr> <td><strong>Column Name</strong></td> <td><strong>Description</strong></td> <td><strong>Type</strong></td> </tr> <tr> <td><em>Surgery_num</em></td> <td>Procedure number in the Cholec80 dataset from which the image was extracted.</td> <td>Integer</td> </tr> <tr> <td><em>Dir</em></td> <td>Directory of the image folder.</td> <td>String</td> </tr> <tr> <td><em>FrameName</em></td> <td>Image name in the format '<em>Video_SS_fffff.png', </em>where&nbsp;<em>SS is the Surgery_num and fffff is the frame number in the video.</em></td> <td>String</td> </tr> <tr> <td><em>NumBBox_inFrame</em></td> <td>The bounding-box number in the image.</td> <td>Integer</td> </tr> <tr> <td><em>ToolName</em></td> <td>Name of the surgical tool.</td> <td>String</td> </tr> <tr> <td><em>BBox</em>_<em>X</em></td> <td>X-coordinate of the top-left corner.</td> <td>Integer</td> </tr> <tr> <td><em>BBox_Y</em></td> <td>Y-coordinate of the top-left corner.</td> <td>Integer</td> </tr> <tr> <td><em>BBox_Width</em></td> <td>Bounding box&nbsp;width.</td> <td>Integer</td> </tr> <tr> <td><em>BBox_Height</em></td> <td>Bounding box height.</td> <td>Integer</td> </tr> </tbody> </table> <p>&nbsp;</p> <p><strong>Citing This Dataset:</strong></p> <p>When using this dataset, please cite the following publications:</p> <ul> <li>Jalal, N. A., Alshirbaji, T. A., Docherty, P. D., Arabian, H., Laufer, B., Krueger-Ziolek, S., Neumuth, T. &amp; Moeller, K. (2023). Laparoscopic video analysis using temporal, attention, and multi-feature fusion based-approaches. <em>Sensors</em>,&nbsp;<em>23</em>(4), 1958.<br><br></li> <li>Jalal, N. A., Alshirbaji, T. A., Docherty, P. D., Arabian, H., Neumuth, T., &amp; M&ouml;ller, K. (2023). Surgical tool classification &amp; localisation using attention and multi-feature fusion deep learning approach. IFAC-PapersOnLine, 56(2), 5626-5631.</li> <li> <p>Abdulbaki Alshirbaji, T., Arabian, H., Jalal, N. A., Battistel, A., Docherty, P. D., Neumuth, T., &amp; Moeller, K. &nbsp;Cholec80-Boxes: Bounding-box labeling data for surgical tools in cholecystectomy images. (<em>to be submitted</em>).&nbsp;</p> </li> <li>Twinanda, A. P., Shehata, S., Mutter, D., Marescaux, J., De Mathelin, M., &amp; Padoy, N. (2016). Endonet: a deep architecture for recognition tasks on laparoscopic videos.&nbsp;<em>IEEE transactions on medical imaging</em>,&nbsp;<em>36</em>(1), 86-97.</li> </ul>

opencc-by-nc-sa-4.0Sep 2022View details →
zenodo44/100

Revealing real-time 3D in vivo pathogen dynamics in plants by label-free optical coherence tomography

<p>This repository contains all data and code underlying the publication: J. de Wit et al. "<em>Revealing real-time 3D in vivo pathogen dynamics in plants by label-free optical coherence tomography</em>" in Nature Communications (2024) (https://doi.org/10.1038/s41467-024-52594-x)</p> <p><strong>--------------Code description------------------</strong></p> <p>The set of scripts is largely organized around the figures. For each (sub)figure, also from supplementary materials, that involves data and plotting, there is a script that generates the plot from data that can be found in the different zip files that are present in the Zenodo repository under https://doi.org/10.5281/zenodo.11428245.</p> <p>The scripts use the data that is contained in the ZIP folders. The ZIP folders are organized by experiment (Experiment 1, including contrast optimization; Experiment 2), one for the other data (OtherData, the validation for with Trypan blue, and the Arabidopsis, Radish and nematode) and one as a smaller dataset to explain the method on a single B-scan (Example_Bscan_dynamicOCT).</p> <p>IMPORTANT: The folder where the ZIP files are unzipped should be put in the file '<em>basepath.txt</em>', such that the data can be automatically loaded.</p> <p>Besides the figures that mention 'MakeFig...' there are a few more scripts:</p> <ul> <li><em>pointcloud_generation_experiment1.py</em>: this file makes the point clouds from the dynamic OCT images as described in Fig 2b. The resulting data is saved as maximum intensity projections and axial sums(forming the basis for Fig S4, S6 and S7) and as voxel counts (forming the basis of Fig.2c and Fig S5)</li> <li><em>pointcloud_generation_timelapses.py</em>: this file does the segmentation for experiment 2 and saves the maximum intensity projections and axial sums of the different stages in the segmentation (forming the basis of Fig3a,d,e and FigS9a,b), and saves the point clouds of the data. These point clouds were refined manually in CloudCompare as described in methods. These segmented point clouds are contained in the data zip folder of experiment 2.</li> <li><em>StatisticalTests.R</em>: This R file calculates the statistical tests for Fig.2cd and Fig.S5d. Here the path is not automatically updated, and should be manually set. The input file is contained in "Experiment1/SegmentationData/segmentationdata_samples.csv" and the output of the file is "D:/DataZenodo/Experiment1/SegmentationData/data_combined_Rstats_output.csv"</li> <li><em>example_dynamic_Bscan.py</em>: This script gives an example of the dynamic OCT processing as proposed in this paper. First it shows the process from an OCT interference spectrum to a B-scan. Then it loads 100 B-scans and applies dynamic OCT, including normalization with histograms. Finally it gives a dynamic B-scan and plots this against the average normal OCT image. This script can be used with only the zip folder "Example_Bscan_dynamicOCT", which reduces the amount of data needed to download/unzip.</li> </ul> <p>The list of other script files to load the data and generate the figures (guiding to the path of uncropped figures) is:</p> <ul> <li><em>MakeFig1bce_Fig2e.py</em></li> <li><em>MakeFig1d.py</em></li> <li><em>MakeFig1agraphs_FigureS1.py</em></li> <li><em>MakeFig3acde_S9ab.py</em></li> <li><em>MakeFigS2_determine_dynamic_range_experiment1.py</em></li> <li><em>MakeFigS8.py</em></li> <li><em>MakeFigureS4-S6-S7.py</em></li> <li><em>MakeFigureS5.py</em></li> <li><em>MakeHistFig2b_makeFigS3b-e.py</em></li> <li><em>MakePlotsFig2ab.py</em></li> <li><em>MakePlotsFig2cd.py</em></li> </ul> <p>Code was all run in Python 3 using Anaconda Spyder.</p> <p>Moreover, the zip file with the code contains the folder '<em>figures</em>' with all subfigures. Some of them are automatically saved from the scripts, others (like photos, icons, but also the Trypan blue microscopy figure) are added in the respective folder. The are logically organized by figure number.</p> <p><strong>--------------Dataset Description-----------------</strong></p> <p>As mentioned above, the data is organized in four zip folders for both experiments, the other data (validation with Trypan blue, other plant-pathogens) and one for the dynamic OCT B-scan example. The data contain the following:</p> <p><strong>Experiment 1:&nbsp;</strong></p> <ul> <li>DynamicOCTimages whose subfolders (organized by date) contain a folder per volume dataset in experiment 1 with a z-stack of .tif files that form the imaged volume. The lateral sampling is 3 um and the axial sampling is 1.37 um.&nbsp;</li> <li>ContrastOptimization: This folder contains&nbsp; <ul> <li><em>Bscans_with_segmentation</em>: segmented B-scans for contrast optimization (Fig S3)</li> <li><em>Bscan_figS1_fig1</em>: The B-scans and segementation for Figure S1.</li> <li><em>histogramdata_dynamicrange</em>: The histograms, bins and deducted reference data for determining the dynamic range per color channel for experiment 1 (Fig S2)</li> <li><em>logcompressed_3value_dOCT_example</em>: An example data stack for obtaining histograms (see script MakeFigS2_determine_dynamic_range_experiment1.py)</li> <li><em>overlaps_threshold-100-98-95-92-90-85-80-75-70-65-60-55-50-45perc_red-1_2_blue_-3_0_green1_filt.npy</em>: A file with intermediate data for the contrast optimization, which can also be generated with the script "<em>MakeHistFig2b_makeFigS3b-e.py</em>"</li> </ul> </li> <li>SegmentationData: This folder contains&nbsp; <ul> <li><em>MIP_segmentation_stages</em>: maximum intensityp projections and axial sums for all images at different stages in the segmentation (basis for Fig S4,6,7)</li> <li><em>processed_masks and StackMasks</em>: the manually obtained masks (segmented in StackMasks, made into masks in the folder 'processed_masks') for filtering out stomata, veins and artefacts.</li> <li><em>Unmasked_axialsum_th34_formanualsegmentation</em>: This folder contains the images of Fig.S4 and were used for the segmentation (we addes a small offset, such that the in segmentation we could set it to 0 and have a unique mask).&nbsp;</li> <li><em>overview_samples_bremiayn.csv</em>: A dataframe with the data for all the samples in experiment 1 that is used as input for the segmentation. It also contains the result of the manual check whether it has infection (Fig2c, left).</li> <li><em>segmentationdata_samples.csv</em>: This supplements the file of overview_samples_bremiayn.csv with the results from the segmentation and is output to script "<em>pointcloud_generation_experiment1.py</em>". It forms the basis of Fig.2a-c, and FigS5, as well as for the R-script to do the statistical testing.</li> <li><em>qPCR_dOCT.csv</em>: This script contains the qPCR data and is input to Fig2d.&nbsp;</li> </ul> </li> </ul> <p><strong>Experiment 2:</strong></p> <ul> <li><em>DynamicOCTimages</em>: This contains the z-stacks of .tif files of the volumes for experiment 2 (and one extra, where a z-slice is used in Fig.1b, bottom). Sampling step size is here again 3 um in lateral direction and 1.37 um in axial direction.</li> <li><em>.npy files </em>with the histograms (with same bins as Experiment 1), maxvalues and reference values for the dynamic range calculation.</li> <li><em>segmentation_data</em>: this folder contains: <ul> <li><em>quantification_volume_disc160_33_10.csv</em> and <em>quantification_volume_disc160_33_10.xlsx</em>: data from the manually segmented point clouds that form the basis of Fig.3c.</li> <li><em>timelapse_sampleoverview.csv</em>: overview of the samples that is used as input in the file "<em>pointcloud_generation_timelapses.py</em>"</li> <li><em>pointclouds</em>: Folder with segmented point clouds for the three leaf discs. These files could &nbsp;be loaded in CloudCompare.</li> <li><em>overviewMIPs</em>: folder with overview maximum intensity projections for the different steps in segmentation, which also forms the input of Fig.3a, FigS9ab.</li> <li>rawpointclouds: folder with the automatically generated point clouds from file&nbsp;<em>pointcloud_generation_timelapses.py&nbsp;</em>which were imported into CloudCompare as the basis for the segmented point clouds.</li> </ul> </li> </ul> <p><strong>OtherData:</strong></p> <p>This folder contains the z-stacks of dynamic OCT tif images for Arabidopsis (here both a normal contrast and one that has been increased to only contain the original 0-180 range); nematodes, radish (called radijs_test_PP_py_0002), spores for Fig1c (SporesImaging) and the dynamic OCT image of Fig1d.&nbsp;</p> <p><strong>example_Bscan_dynamicOCT:</strong></p> <p>This folder contains data to run the script example_dynamic_Bscan.py to show the dynamic OCT imaging process from raw OCT spectra.</p> <ul> <li><em>raw_spectra_exampleframe:</em> contains interference spectra, a reference spectrum and interpolation grid to show how to get from a raw OCT spectrum to a normal single B-scan.</li> <li><em>abs_images:</em> contains 100 subsequent B-scans that can be used to generate a dynamic OCT image as done in example_dynamic_Bscan.py</li> </ul> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Aug 2024View details →
zenodo44/100

Eye image data with gaze labels recorded using custom video-oculography hardware at 120Hz

<p>The repository of eye image data with corresponding gaze labels collected from 40 subjects. The preview contains a collage of random image samples, one per subject.&nbsp;</p> <p>All recorded subjects gave informed consent under an experimental protocol approved by the Institutional Research Board of Texas State University (approval code 2018044) and their data were anonymized prior to public release.</p> <p>The data were recorded using the custom video-oculography (VOG) desktop hardware setup at 120Hz. The full description of this eye-tracking system's capabilities is provided at https://doi.org/10.48550/arXiv.1904.07361.</p> <p>This VOG set contains recordings of the random oblique saccades task. It is comprised of 174 on-screen fixation targets that densely cover the range of &plusmn;20.51&deg; horizontally and &plusmn;16.7&deg; vertically (in degrees of visual angle). More detail on the presented stimuli can be found at https://doi.org/10.1145/3379156.3391370.</p> <p>The data were also used in Dmytro Katrychuk's Ph.D. thesis "Generating Realistic Eye Images to Evaluate Photosensor Oculography Eye-Tracking for Portable Headsets" (https://hdl.handle.net/10877/19437); with the release for public use in the upcoming publication "An appearance-based gaze estimation as a benchmark for eye image data generation methods" accepted to MDPI Journal of Applied Sciences.&nbsp;</p> <p>Each .zip archive represents a recording from one subject, which includes:</p> <ul> <li>Video of the close eye capture in ".avi" format</li> <li>Calibration data in ".xml" format</li> <li>Gaze data in ".tsv" format</li> <li>On-screen target stimulus position in ".tsv" format</li> </ul> <p>The "src.zip" provides a Python script to unpack each ".avi" video recording to a set of ".png" images. The direct playback of ".avi"s may require special codecs and is not supported.&nbsp;</p> <p>Any additional code will be uploaded to https://github.com/dkatrychuk/psog-eval-diss2023</p> <p>The authors can be contacted at their corresponding emails: Dmytro Katrychuk - d_k139@txstate.edu; Oleg Komogortsev - ok@txstate.edu.</p>

opencc-by-4.0Oct 2024View details →
zenodo44/100

Webinar: Potential and challenges of certification and development of new labels

<p>The webinar will address the vital topic of new certifications and labels for the renewable, and especially bio-based, economy.</p> <p>Certification and labelling play a crucial role in empowering consumers to make informed and sustainable purchasing decisions by impeding greenwashing and instead providing credible and reliable sustainability information. Furthermore, they facilitate the tracking and traceability of biological and other renewable feedstock throughout value chains, fostering transparency and accountability among stakeholders of the industry. However, there are a lot of challenges for current certification and labelling schemes, since they are often not sufficiently laid-out for bio-based and other renewable products and their value-chains.</p> <p>In this webinar, REDcert and T&Uuml;V AUSTRIA will show their sytems for sustainability certification and will present new or soon to be published label and certification schemes (LCS) that are relevant for the bio-based economy. The speakers will talk about challenges on the way to the final label, and which gaps will be closed.</p> <p>Join us on November the 5th and listen to the two presentations by REDcert and T&Uuml;V AUSTRIA. Together, we will explore their innovative new LCS that will support shaping a strong circular EU (bio)economy.</p> <p>Speakers:</p> <p>Phillipe Dewolfs (T&Uuml;V AUSTRIA): From standardisation to communication &ndash; the role of a certification body<br>Simon Schwarzwald (REDcert): Certification of sustainable products in the REDcert&sup2; scheme</p>

opencc-by-4.0Nov 2024View details →
zenodo44/100

FeM dataset – An iron ore labeled images dataset for segmentation training and testing

<p>This dataset is composed of 81 pairs of correlated images. Each pair contains one image of an iron ore sample acquired through reflected light microscopy (RGB, 24-bit), and the corresponding binary reference image (8-bit), in which the pixels are labeled as belonging to one of two classes: ore (0) or embedding resin (255).</p> <p>The sample came from an itabiritic iron ore concentrate from Quadril&aacute;tero Ferr&iacute;fero (Brazil) mainly composed of hematite and quartz, with little magnetite and goethite. It was classified by size and concentrated with a dense liquid. Then, the fraction -149+105 &mu;m with density greater than 3.2 was cold mounted with epoxy resin and subsequently ground and polished.</p> <p>Correlative microscopy was employed for image acquisition. Thus, 81 fields were imaged on a reflected light microscope with a 10&times; (NA 0.20) objective lens and on a scanning electron microscope (SEM). In sequence, they were registered, resulting in images of 999&times;756 pixels with a resolution of 1.05 &micro;m/pixel. Finally, the images from SEM were thresholded to generate the reference images.</p> <p>Further description of this sample and its imaging procedure can be found in the work by Gomes and Paciornik (2012).</p> <p>This dataset was created for developing and testing deep learning models on semantic segmentation tasks. The paper of Filippo et al. (2021) presented a variant of the DeepLabv3+ model that reached mean values of 91.43% and 93.13% for overall accuracy and F1 score, respectively, for 5 rounds of experiments (training and testing), each with a different, random initialization of network weights.</p> <p>For further questions and suggestions, please do not hesitate to contact us.</p> <p>&nbsp;</p> <p><strong>Contact email</strong>: ogomes@gmail.com</p> <p>&nbsp;</p> <p>If you use this dataset in your own work, please cite this DOI: 10.5281/zenodo.5014700</p> <p>&nbsp;</p> <p>Please also cite this paper, which provides additional details about the dataset:</p> <p>Michel Pedro Filippo, Ot&aacute;vio da Fonseca Martins Gomes, Gilson Alexandre Ostwald Pedro da Costa, Guilherme Lucio Abelha Mota. <em>Deep learning semantic segmentation of opaque and non-opaque minerals from epoxy resin in reflected light microscopy images</em>. <strong>Minerals Engineering</strong>, Volume 170, 2021, 107007, https://doi.org/10.1016/j.mineng.2021.107007.</p> <p>&nbsp;</p>

opencc-by-4.0Jun 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record