Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
4,612
datasets available to search
ShareScore release 0.7.1
Dataset results
4,612 results for “labels”
Calorie-labeled food cues
Open the record for dataset details and reuse information.
Labelled acoustic dataset of roding Eurasian Woodcock (Scolopax rusticola)
<p>This dataset contains manually labelled audio data of roding Eurasian Woodcock (<em>Scolopax rusticola</em>). </p> <h2><strong>Description</strong></h2> <p>Bioacoustic surveys of roding Eurasian Woodcock were conducted in Baden-Württemberg, Germany in May and June in 2020 and 2021. The audio data of this collection was used for the evaluation of BirdNET as a means for the automated analysis of large quantities of audio data. The original dataset consisted of 12.236 minutes of recording, which were reviewed manually. Each call element of a male roding Woodcock (i.e. croak, whistle, chasing male) was annotated. Individual call elements were subsequently clustered into so called roding events, which are ecologically more meaningful. BirdNET was then tested against this manually labelled dataset.</p> <p>The dataset uploaded to zenodo contains:</p> <ul> <li>audio data of 2545 woodcock call element selections with a duration of 145 minutes</li> <li>audio data of 782 aggregated woodcock roding events with a duration of 115 minutes </li> <li>selection tables for call elements and roding events</li> <li>associated metadata</li> </ul> <p>Audio information in between roding events (i.e. non woodcock audio) ist not included due to data privacy reasons (see below). </p> <h3>Selections</h3> <p>Woodcock call elements were manually selected/annotated in Raven Pro with bounding boxes. For this dataset, all selections with a duration of less than 3 seconds were extended symmetrically until 3 seconds were reached. This may result in overlapping selections in the case of croaks that are directly followed by a whistle. Signals at the beginning or end of these selections may thus be included twice.</p> <h3>Roding events</h3> <p>A roding event was defined as a continuous series of Woodcock call elements with a maximum gap of six seconds between consecutive elements. Each event can be interpreted as a roding bird that passes by the recording location, similar to a typical woodcock roding survey conducted by a human observer. Roding events were not created with the extended 3 seconds clips described above, but with the original bounding box selections drawn in Raven Pro.</p> <h3>Audio files</h3> <ul> <li>selections.zip: each wav-file contains a single selections. Filenames correspond to the column selec in the table <em>selections.csv</em></li> <li>events.zip: each wav-file contains a single roding event, typically consisting of multiple call elements (croaks and/or whistles). In the case of faint signals of distant birds, roding events may consist of a single call element only. Filenames correspond to the column <em>event.id</em> in the table<em> events.csv</em>.</li> </ul> <h2><strong>Data collection</strong></h2> <p>All wav-files in this dataset originate from audio files that were recorded with autonomous recording units of the type AudioMoth. ARUs were housed in custom made waterproof casings (See details and files for 3D-printing: https://www.thingiverse.com/thing:6428228). ARUs were programmed to record continuously for 2 hours during dusk and were placed at edges of forest clearings. The devices were mounted to tree trunks at a height of approximately 1.5m above ground. </p> <h2><strong>Metadata files</strong></h2> <table> <tbody> <tr> <td><strong>filename</strong></td> <td><strong>content</strong></td> </tr> <tr> <td>sites.csv</td> <td> <p>contains locations of the recording sites. Since exact recording locations can not be made public, only recording sites (= cells of the 1km² UTM-grid) are provided. CRS: EPSG - 25832, ETRS89 / UTM 32N </p> <p>Data source of the underlying ETRS89 UTM 32N grid: https://gdz.bkg.bund.de/index.php/default/digitale-geodaten/nicht-administrative-gebietseinheiten/geographische-gitter-fur-deutschland-in-utm-projektion-geogitter-national.html</p> <p><strong>columns</strong></p> <p>site.id = unique id of recording sites,</p> <p>cellcode = official cellcode of the 1km²-UTM-grid</p> <p>elevation = mean elevation a.s.l.</p> <p>x.centroid = x-coordinate of centroid (EPSG: 25832)</p> <p>y.centroid = y-coordinate of centroid (EPSG: 25832)</p> <p>wkt.geometry = polygon geometry of the grid cell</p> </td> </tr> <tr> <td>arus.csv</td> <td> <p>metadata of the recording hardware</p> <p> </p> <p><strong>columns</strong></p> <p>aru.id = unique id of recording device</p> <p>type = recorder type</p> <p>manufacturer = manufacturer of recording hardware</p> <p>hardware.version = hardware version of the recording device</p> <p>acquisition.date = date the device was purchased (for reasons of microphone degradation)</p> </td> </tr> <tr> <td>deploys.csv</td> <td> <p>information on recorder deployment, includes aru settings, location, recording times </p> <p> </p> <p><strong>columns</strong></p> <p>deploy.id = unique id of recorder deployment</p> <p>aru.id = unique id of deployed aru</p> <p>start.date = date the aru was deployed in the field (YYYY-MM-DD)</p> <p>end.date = date the aru was collected (YYYY-MM-DD)</p> <p>firmware = firmware version used in this deployment</p> <p>rec.periods = number of daily recording periods (corresponds to start.rec1, start.rec2 ...)</p> <p>sample.rate = sample rate in kHz</p> <p>gain = gain setting</p> <p>sleep.duration = duration off stand-by phases in seconds, when set on a sleep/record-cycle</p> <p>rec.duration = duration of each recording in seconds, when set on a sleep/record-cycle</p> <p>start.rec1 = start of first recording period (UTC, hh:mm:ss)</p> <p>end.rec1 = end of first recording period (UTC, hh:mm:ss)</p> <p>start.rec2 = start of secondrecording period (UTC, hh:mm:ss)</p> <p>end.rec2 = end of second recording period (UTC, hh:mm:ss)</p> <p>site.id = unique id of recording site</p> </td> </tr> <tr> <td>recordings.csv</td> <td> <p>metadata of the audio files from which the roding events originate</p> <p> </p> <p> <strong>columns</strong></p> <p>recording.id = unique id of the recording</p> <p>deploy.id = unique id of aru deployment, during which the recording was made</p> <p>date = date on which the recording was made (YYYY-MM-DD)</p> <p>time = time of day at which the recording started (UTC, hh:mm:ss)</p> <p>duration = duration in seconds</p> <p>sampler.rate = sample rate in kHz</p> <p>channels = number of channels</p> <p>bits = bit depth</p> <p>samples = number of audio samples</p> <p>gain = gain setting of the aru</p> <p>voltage = battery voltage of the aru during recording</p> <p>temperature = ambient temperature during recording </p> <p>reviewer = anonymous id of staff who reviewed the file and annotated calls</p> <p> </p> </td> </tr> <tr> <td>selections.csv</td> <td> <p>manually labelled woodcock call elements (i.e. croaks, whistles, chases). Short selections were extended to 3 seconds by symmetrically adding time before and after the original selection. In the format of raven pro selection tables.</p> <p> </p> <p> <strong>columns</strong></p> <p>selec = unique id of the selection. Corresponds to the filename of the wav-files in the archive <em>selections.zip</em></p> <p><em>deploy.id = unique id of the aru deployment during which the roding event was recorded</em></p> <p>channel = audio channel</p> <p>start = start of the event in seconds from the start of the recording</p> <p>end = end of the event in seconds from the start of the recording</p> <p>bottom.freq = bottom frequency of the annotation bounding box</p> <p>top.frequency = top frequency of the annotation bounding box</p> <p>species.code = species code as used by BirdNET</p> <p>common.name = English common name as used by BirdNET</p> <p>annotation = contains annotations of call elements that are pooled in the roding event. Thus typcally equal to the number of annotated call element </p> <p>recording.id = id of the recording this roding eventoriginates from</p> </td> </tr> <tr> <td>events.csv</td> <td> <p>aggregated roding events consisting of contiuous sequences of manually labelled call elements. In the format of raven pro selection tables</p> <p> </p> <p> <strong>columns</strong></p> <p>event.id = unique id of roding event. Corresponds to the filename of the wav-files in the archive <em>events.zip </em></p> <p>channel = audio channel</p> <p>start = start of the event in seconds from the start of the recording</p> <p>end = end of the event in seconds from the start of the recording</p> <p>bottom.freq = bottom frequency of the annotation bounding box</p> <p>top.frequency = top frequency of the annotation bounding box</p> <p>species.code = species code as used by BirdNET</p> <p>common.name = English common name as used by BirdNET</p> <p>annotation = contains annotations of call elements that are pooled in the roding event. Thus typcally equal to the number of annotated call element </p> <p>recording.id = id of the recording this roding eventoriginates from</p> <p>deploy.id = unique id of the aru deployment during which the roding event was recorded</p> </td> </tr> <tr> <td>removed_audio_files.txt</td> <td>selection ids and event ids of audio files that were deleted because they included voices. Their metadata is still included in the files described above</td> </tr> </tbody> </table> <p> </p> <h2><strong>Data privacy</strong></h2> <p>Selections and roding events were checked for human voices and audio information was removed, in case it contained any. Audio segments that did not contain woodcock calls were not completely checked for human voices and can thus not be made available.</p>
Pre-processed (in Detectron2 and YOLO format) planetary images and boulder labels collected during the BOULDERING Marie Skłodowska-Curie Global fellowship
<p>This database contains 4976 planetary images of boulder fields located on Earth, Mars and Moon. The data was collected during the BOULDERING Marie Skłodowska-Curie Global fellowship between October 2021 and 2024. The data was already splitted into train, validation and test datasets, but feel free to re-organize the labels at your convenience. </p> <p>For each image, all of the boulder outlines within the image were carefully mapped in QGIS. More information about the labelling procedure can be found in the following manuscript (<a href="https://agupubs.onlinelibrary.wiley.com/doi/full/10.1029/2023JE008013">https://agupubs.onlinelibrary.wiley.com/doi/full/10.1029/2023JE008013</a>). This dataset differs from the previous dataset included along with the manuscript <a href="https://zenodo.org/records/8171052">https://zenodo.org/records/8171052</a>, as it contains more mapped images, especially of boulder populations around young impact structures on the Moon (cold spots). In addition, the boulder outlines were also pre-processed so that it can be ingested directly in YOLOv8.</p> <p>A description of what is what is given in the README.txt file (in addition in how to load the custom datasets in Detectron2 and YOLO). Most of the other files are mostly self-explanatory. Please see previous dataset or manuscript for more information. If you want to have more information about specific lunar and martian planetary images, the IDs of the images are still available in the name of the file. Use this ID to find more information (e.g., M121118602_00875_image.png, ID M121118602 ca be used on https://pilot.wr.usgs.gov/). I will also upload the raw data from which this pre-processed dataset was generated (see <a href="https://zenodo.org/records/14250970">https://zenodo.org/records/14250970</a>).</p> <p>Thanks to this database, you can easily train a Detectron2 Mask R-CNN or YOLO instance segmentation models to automatically detect boulders. </p> <p><strong>How to cite:</strong></p> <p>Please refer to the "how to cite" section of the readme file of <a href="https://github.com/astroNils/YOLOv8-BeyondEarth" target="_blank" rel="noopener">https://github.com/astroNils/YOLOv8-BeyondEarth.</a></p> <p><strong>Structure:</strong></p> <pre><code>. └── boulder2024/ ├── jupyter-notebooks/ │ └── REGISTERING_BOULDER_DATASET_IN_DETECTRON2.ipynb ├── test/ │ └── images/ │ ├── <image_name>_image.png │ ├── ... │ └── labels/ │ ├── <image_name>_image.txt │ ├── ... ├── train/ │ └── images/ │ ├── <image_name>_image.png │ ├── ... │ └── labels/ │ ├── <image_name>_image.txt │ ├── ... ├── validation/ │ └── images/ │ ├── <image_name>_image.png │ ├── ... │ └── labels/ │ ├── <image_name>_image.txt │ ├── ... ├── detectron2_inst_seg_boulder_dataset.json ├── README.txt ├── yolo_inst_seg_boulder_dataset.yaml</code></pre> <p> </p> <pre><code>detectron2_inst_seg_boulder_dataset.json</code></pre> <p>is a json file containing the masks as expected by Detectron2 (see <a href="https://detectron2.readthedocs.io/en/latest/tutorials/datasets.html">https://detectron2.readthedocs.io/en/latest/tutorials/datasets.html</a> for more information on the format). In order to use this custom dataset, you need to register the dataset before using it in the training. There is an example how to do that in the jupyter-notebooks folder. You need to have detectron2, and all of its depedencies installed. </p> <pre><code>yolo_inst_seg_boulder_dataset.yaml</code></pre> <p>can be used as it is, however you need to update the paths in the .yaml file, to the test, train and validation folders. More information about the YOLO format can be found here (<a href="https://docs.ultralytics.com/datasets/segment/">https://docs.ultralytics.com/datasets/segment/</a>).</p>
Labels for Emergency Response Imagery from Hurricane Barry, Delta, Dorian, Florence, Ida, Isaias, Laura, Michael, Sally, Zeta, and Tropical Storm Gordon
<p>The csv files contain human-generated labels for Emergency Response Imagery collected by US National Oceanic and Atmospheric Administration (NOAA) after Hurricane Barry, Delta, Dorian, Florence, Ida, Isaias, Laura, Michael, Sally, Zeta, and Tropical Storm Gordon. All authors contributed to labeling the imagery. All labeling was done with an open-source labeling tool (Rafique et al., 2020).</p> <p>All csv files provide the userID (the ID of the anonymous labeler), the NOAA flight, the NOAA image, and 6 labels — allWater (if the image was all water), devType (if the image had buildings/development), washoverType (if the image had washover deposits), dmgType (if the image showed damage to built environment), impactType (if the labeler could identify the coastal impact, using the Storm Impact Scale from Sallenger, 2000), and terrainType (the type of physical environment).</p> <p>Images labeled here correspond to multiple NOAA flights — all listed in the csv file for each jpeg image. These jpeg images can be downloaded directly from NOAA (https://storms.ngs.noaa.gov/) or using Moretz et al. (2020a, 2020b).</p> <p>There are three csv files:</p> <p>ReleaseData_10172022.csv has 10,237 labels for 4250 images. These labels were generated by coastal scientists. The csv also contains the Latitude and Longitude of the image center (from NOAA).</p> <p>ReleaseDataQuads.csv has 400 labels for 100 images. These labels were generated by coastal scientists. The images labeled in this set correspond to original NOAA images that have been split into quadrants. Splitting images was done with ImageMagick. The command used to split the images was:</p> <p>`magick mogrify -crop 2x2@ +repage -path ../quadrants *.jpg`</p> <p>The naming convention corresponds to the image quarter — the *-0.jpg is upper left, *-1.jpg is upper right, *-2.jpg is lower left, and *-3.jpg is the lower right.</p> <p>ReleaseDataNCE.csv has 400 labels for 100 images. These images were labeled by non-coastal scientists. Note that the 100 images were also labeled by coastal scientists — those labels can be found in ReleaseData_v3.csv.</p> <p>There is another companion dataset to this, with slightly different labels (Goldstein et al., 2020).</p> <p>A zip file of images is also provided for demonstration purposes (images.zip). These are resized copies made with imagemagick, with the longest dimension set at 2000 pixels ( `mogrify -resize 2000x2000`). For full size images, please download the jpegs directly from NOAA.</p>
STAR4BBS D1.3 Report impact and contribution SCS and Labels_Appendix II dataset
<p>This dataset contains the full coding sheet for the systematic mapping that formed STAR4BBS deliverable D1.3 (Appendix II). The systematic mapping exercise reviewed literature on the impact of and contribution to GHG emissions reductions of existing sustainability systems and certification schemes (SCS) and B2B labels used within the bioeconomy. A coding sheet in the context of a systematic map is a structured tool used to extract and record specific data from studies being reviewed, ensuring consistency and accuracy in data collection. It forms the basis of the analysis and is included for transparency.</p>
Labeled Time Series Data of Force/Torque for Monitoring Assembly Processes with a Delta Robot
<p>This dataset comprises 524 recordings of 6-dimensional time series data, capturing forces in three directions and torques in three directions during the assembly of small car model wheels. The data was collected using an equidistant sampling method with a sampling period of 0.004 seconds. Each time series represents the process of assembling one wheel, specifically the placement of a tire onto a rim, and includes a label indicating whether the assembly was successful (OK). The wheels were assembled in batches of four, and the recordings were obtained over six different days. The labels of recordings from two (days 3 and 4) of the six days are invalid as described in [1]. The labels presented in this data set are only binary (they do not describe the reason of the failure). The labels of recordings from days 5 and 6 are created by human while the other labels came from a convolutional neural network based computer vision classifier and can be inaccurate as described in section 5.4 of [1]. </p> <h4>Dataset Structure:</h4> <ul> <li><strong>File:</strong> <code>ForceTorqueTimeSeries.csv</code> <ul> <li><strong>Columns:</strong> <ul> <li><code>idx (1-524)</code>: Index of the recording corresponding to the assembly of one wheel.</li> <li><code>label (true/false)</code>: Indicates whether the assembly was successful (TRUE = product is OK).</li> <li><code>meas_id (1-6)</code>: Identifier for the day on which the recording was made (refer to Table 2.1 in [1]).</li> <li><code>force_x</code>: X-component of the force measured by the sensor mounted on the delta robot's end effector.</li> <li><code>force_y</code>: Y-component of the force.</li> <li><code>force_z</code>: Z-component of the force.</li> <li><code>torque_x</code>: X-component of the torque.</li> <li><code>torque_y</code>: Y-component of the torque.</li> <li><code>torque_z</code>: Z-component of the torque.</li> </ul> </li> </ul> </li> </ul> <h4>Additional Files:</h4> <ul> <li><strong><code>IMG_3351.MOV</code>:</strong> A video demonstrating the assembly process for one batch of four wheels.</li> <li><strong><code>F3-BP-2024-Trna-Ales-Ales Trna - 2024 - Anomaly detection in robotic assembly process using force and torque sensors.pdf</code>:</strong> Bachelor thesis [1] detailing the dataset and preliminary experiments on fault detection.</li> <li><strong><code>F3-BP-2024-Hanzlik-Vojtech-Anomaly_Detection_Bachelors_Thesis.pdf</code>:</strong> Bachelor thesis [2] describing the data acquisition process.</li> </ul> <h3>References:</h3> <ol> <li>Trna, A. (2024). <em>Anomaly detection in robotic assembly process using force and torque sensors</em> [Bachelor’s thesis, Czech Technical University in Prague].</li> <li>Hanzlik, V. (2024). <em>Edge AI integration for anomaly detection in assembly using Delta robot</em> [Bachelor’s thesis, Czech Technical University in Prague].</li> </ol>
Ovenbird song recordings from Alberta (Canada) with individual labels and spatial locations, 2015-2016
This dataset includes spatially localized and individually identified Ovenbird songs. We used automated species detection and acoustic localization to localize Ovenbird singing events from microphone arrays in Alberta, Canada (2015-2016). We then hand-annotated songs to individuals based on acoustic characteristics. This dataset includes the manual annotations and annotations from automated individual identification approaches. This data publication pertains to the manuscript [in prep] by Lapp et al on Ovenbird individual identification and provides further details on the study and the individual identification approach.
Labelled magnetic reconnection simulation data set
<p>Numerical simulations have been performed on Marconi at CINECA (Italy) under the ISCRA initiative. The corresponding data can be found at: <a href="https://doi.org/10.5281/zenodo.3935887">https://doi.org/10.5281/zenodo.3935887</a></p>
Synchrotron diffraction images for the 2.9 Å crystal structure of L-Selenomethionine labeled human GDAP1
<p>Dataset collected at DLS, I04 beamline 16.5.2019. L-SeMet substituted crystals collected with SAD-method.</p> <ul> <li>Flux: 1.32e+11</li> <li>Ω Start: 0.0°</li> <li>Ω Osc: 0.10°</li> <li>Ω Overlap: 0°</li> <li>No. Images: 3600</li> <li>Resolution: 2.90Å</li> <li>Wavelength: 0.9790Å</li> <li>Exposure: 0.040s</li> <li>Transmission: 100.00%</li> <li>Beam size: 63x50μm</li> <li>Type: SAD</li> <li>Comment: X,Y,Z (-561,302,301), Aperture: Large</li> </ul> <p> </p>
Sidescan Sonar Substrate, Depth, and Shadow Image-Label-Pairs
<p>Substrate, depth, and shadow image-label pairs used to train side scan sonar segmentation models v1.0 implemented in PINGMapper v2.0.</p><p> </p><p>Images were labeled with <a href="https://github.com/Doodleverse/dash_doodler">Doodler</a> and <a href="https://www.makesense.ai/">Make Sense.</a></p>
Labeled high-resolution orthoimagery time-series of an alluvial river corridor; Elwha River, Washington, USA.
<h2>Labeled high-resolution orthoimagery time-series of an alluvial river corridor; Elwha River, Washington, USA.</h2><h4>Daniel Buscombe, Marda Science LLC</h4><p>There are two datasets in this data release:</p><p>1. <strong>Model training dataset</strong>. A manually (or semi-manually) labeled image dataset that was used to train and evaluate a machine (deep) learning model designed to identify subaerial accumulations of large wood, alluvial sediment, water, and vegetation in orthoimagery of alluvial river corridors in forested catchments. </p><p>2. <strong>Model output dataset</strong>. A labeled image dataset that uses the aforementioned model to estimate subaerial accumulations of large wood, alluvial sediment, water, and vegetation in a larger orthoimagery dataset of alluvial river corridors in forested catchments. </p><p>All of these label data are derived from raw gridded data that originate from the U.S. Geological Survey (<i>Ritchie et al., 2018</i>). That dataset consists of 14 orthoimages of the Middle Reach (MR, in between the former Aldwell and Mills reservoirs) and 14 corresponding Lower Reach (LR, downstream of the former Mills reservoir) of the Elwha River, Washington, collected between the period 2012-04-07 and 2017-09-22. That orthoimagery was generated using SfM photogrammetry (following <i>Over et al., 2021</i>) using a photographic camera mounted to an aircraft wing. The imagery capture channel change as it evolved under a ~20 Mt sediment pulse initiated by the removal of the two dams. The two reaches are the ~8 km long Middle Reach (MR) and the lower-gradient ~7 km long Lower Reach (LR). </p><p>The orthoimagery have been labeled (pixelwise, either manually or by an automated process) according to the following classes (inter class in the label data in parentheses):</p><p>1. vegetation / other (0)</p><p>2. water (1)</p><p>3. sediment (2)</p><p>4. large wood (3)</p><h3>1. Model training dataset.</h3><p>Imagery was labeled using a combination of the open-source software Doodler (<i>Buscombe et al., 2021</i>; <a href="https://github.com/Doodleverse/dash_doodler">https://github.com/Doodleverse/dash_doodler</a>) and hand-digitization using QGIS at 1:300 scale, rasterizeing the polygons, and gridded and clipped in the same way as all other gridded data. Doodler facilitates relatively labor-free dense multiclass labeling of natural imagery, enabling relatively rapid training dataset creation. The final training dataset consists of 4382 images and corresponding labels, each 1024 x 1024 pixels and representing just over 5% of the total data set. The training data are sampled approximately equally in time and in space among both reaches. All training and validation samples purposefully included all four label classes, to avoid model training and evaluation problems associated with class imbalance (<i>Buscombe and Goldstein, 2022</i>). </p><p>Data are provided in geoTIFF format. The imagery and label grids (imagery) are reprojected to be co-located in the NAD83(2011) / UTM zone 10N projection, and to consist of 0.125 x 0.125m pixels.</p><p>Pixel-wise labels measurements such as these facilitate development and evaluation of image segmentation, image classification, object-based image-analysis (OBIA), and object-in-image detection models, and numerous potential other machine learning models for the general purposes of river corridor classification, description, enumeration, inventory, and process or state quantification. For example this dataset may serve in transfer learning contexts for application in different river or coastal environments or for different tasks or class ontologies.</p><h4>Files:</h4><p>1. Labels_used_for_model_training_Buscombe_Labeled_high_resolution_orthoimagery_time_series_of_an_alluvial_river_corridor_Elwha_River_Washington_USA.zip, 63 MB, label tiffs</p><p>2. Model_<i>training_</i> images1of4.zip, 1.5 GB, imagery tiffs</p><p>3. Model_<i>training_</i> images2of4.zip, 1.5 GB, imagery tiffs</p><p>4. Model_<i>training_</i> images3of4.zip, 1.7 GB, imagery tiffs</p><p>5. Model_<i>training_</i> images4of4.zip, 1.6 GB, imagery tiffs</p><h3>2. Model output dataset.</h3><p>Imagery was labeled using a deep-learning based semantic segmentation model (<i>Buscombe, 2023</i>) trained specifically for the task using the Segmentation Gym (<i>Buscombe and Goldstein, 2022</i>) modeling suite. We use the software package Segmentation Gym (<i>Buscombe and Goldstein, 2022</i>) to fine-tune a Segformer (<i>Xie et al., 2021</i>) deep learning model for semantic image segmentation. We take the instance (i.e. model architecture and trained weights) of the model of <i>Xie et al. (2021)</i>, itself fine-tuned on ADE20k dataset (<i>Zhou et al., 2019</i>) at resolution 512x512 pixels, and fine-tune it on our 1024x1024 pixel training data consisting of 4-class label images.</p><p>The spatial extent of the imagery in the MR is [455157.2494695878122002,5316532.9804129302501678 : 457076.1244695878122002,5323771.7304129302501678] (NAD83(2011) / UTM zone 10N). Imagery width is 15351 pixels and imagery height is 57910 pixels. The spatial extent of the imagery in the LR is [457704.9227139975992031,5326631.3750646486878395 : 459241.6727139975992031,5333311.0000646486878395] (NAD83(2011) / UTM zone 10N). Imagery width is 12294 pixels and imagery height is 53437 pixels. Data are provided in Cloud-Optimzed geoTIFF (COG) format. The imagery and label grids (imagery) are reprojected to be co-located in the NAD83(2011) / UTM zone 10N projection, and to consist of 0.125 x 0.125m pixels. All grids have been clipped to the union of extents of active channel margins during the period of interest.</p><p>Reach-wide pixel-wise measurements such as these facilitate comparison of wood and sediment storage at any scale or location. These data may be useful for studying the morphodynamics of wood-sediment interactions in other geomorphically complex channels, wood storage in channels, the role of wood in ecosystems and conservation or restoration efforts. </p><h4>Files:</h4><p>1. Elwha_MR_labels_Buscombe_Labeled_high_resolution_orthoimagery_time_series_of_an_alluvial_river_corridor_Elwha_River_Washington_USA.zip, 9.67 MB, label COGs from Elwha River Middle Reach (MR)</p><p>2. Elwha<i>MR_ imagery_ part1_ of</i>_<i> </i>2.zip, 566 MB, imagery COGs from Elwha River Middle Reach (MR)</p><p>3. Elwha<i>MR_ imagery_ part2_ of</i>_<i> </i>2.zip, 618 MB, imagery COGs from Elwha River Middle Reach (MR)</p><p>3. Elwha_LR_labels_Buscombe_Labeled_high_resolution_orthoimagery_time_series_of_an_alluvial_river_corridor_Elwha_River_Washington_USA.zip, 10.96 MB, label COGs from Elwha River Lower Reach (LR)</p><p>4. ElwhaL<i>R_ imagery_ part1_ of</i>_<i> </i>2.zip, 622 MB, imagery COGs from Elwha River Middle Reach (MR)</p><p>5. ElwhaL<i>R_ imagery_ part2_ of</i>_<i> </i>2.zip, 617 MB, imagery COGs from Elwha River Middle Reach (MR)<br> </p><p>This dataset was created using open-source tools of the Doodleverse, a software ecosystem for geoscientific image segmentation, by Daniel Buscombe (<a href="https://github.com/dbuscombe-usgs">https://github.com/dbuscombe-usgs</a>) and Evan Goldstein (<a href="https://github.com/ebgoldstein">https://github.com/ebgoldstein</a>). Thanks to the contributors of the Doodleverse!. Thanks especially Sharon Fitzpatrick (<a href="https://github.com/2320sharon">https://github.com/2320sharon</a>) and Jaycee Favela for contributing labels. </p><h3>References</h3><p>• Buscombe, D. (2023). <strong>Doodleverse/Segmentation Gym SegFormer models for 4-class (other, water, sediment, wood) segmentation of RGB aerial orthomosaic imagery (v1.0)</strong> [Data set]. Zenodo. <a href="https://doi.org/10.5281/zenodo.8172858">https://doi.org/10.5281/zenodo.8172858</a></p><p>• Buscombe, D., Goldstein, E. B., Sherwood, C. R., Bodine, C., Brown, J. A., Favela, J., et al. (2021).<strong> Human-in-the-loop segmentation of Earth surface imagery</strong>. Earth and Space Science, 9, e2021EA002085. <a href="https://doi.org/10.1029/2021EA002085">https://doi.org/10.1029/2021EA002085</a></p><p>• Buscombe, D., & Goldstein, E. B. (2022). <strong>A reproducible and reusable pipeline for segmentation of geoscientific imagery.</strong> Earth and Space Science, 9, e2022EA002332. <a href="https://doi.org/10.1029/2022EA002332">https://doi.org/10.1029/2022EA002332</a> See: <a href="https://github.com/Doodleverse/segmentation_gym">https://github.com/Doodleverse/segmentation_gym</a></p><p>• Over, J.R., Ritchie, A.C., Kranenburg, C.J., Brown, J.A., Buscombe, D., Noble, T., Sherwood, C.R., Warrick, J.A., and Wernette, P.A., 2021, <strong>Processing coastal imagery with Agisoft Metashape Professional Edition, version 1.6—Structure from motion workflow documentation</strong>: U.S. Geological Survey Open-File Report 2021–1039, 46 p., <a href="https://doi.org/10.3133/ofr20211039">https://doi.org/10.3133/ofr20211039</a>.</p><p>• Ritchie, A.C., Curran, C.A., Magirl, C.S., Bountry, J.A., Hilldale, R.C., Randle, T.J., and Duda, J.J., 2018, <strong>Data in support of 5-year sediment budget and morphodynamic analysis of Elwha River following dam removals</strong>: U.S. Geological Survey data release, <a href="https://doi.org/10.5066/F7PG1QWC">https://doi.org/10.5066/F7PG1QWC</a>.</p><p>• Xie, E., Wang, W., Yu, Z., Anandkumar, A., Alvarez, J.M. and Luo, P., 2021. <strong>SegFormer: Simple and efficient design for semantic segmentation with transformers</strong>. Advances in Neural Information Processing Systems, 34, pp.12077-12090.</p><p>• Zhou, B., Zhao, H., Puig, X., Xiao, T., Fidler, S., Barriuso, A. and Torralba, A., 2019. <strong>Semantic understanding of scenes through the ade20k dataset</strong>. International Journal of Computer Vision, 127, pp.302-321.</p><p><br> </p>
SCLabels: Labelled rectified RGB images from the Spanish CoastSnap network
<h1>Training dataset</h1> <p><span>The SCLabels dataset is intended to be used in the exploring and development of Artificial Intelligence (AI) applications aimed at the automation of the shoreline extraction process from rectified images. SCLabels includes rectified RGB images from the Spanish CoastSnap network and their corresponding masks, together with a metadata file and a README file. RGB images encompass variable geographic locations, fields of view, beach types and degrees of occupation, tidal regimes, meteoceanic and lightning conditions, and a variety of environmental characteristics. Masks account for dense pixel labels including 5 categories: i) No data; ii) Not classified; iii) Landwards; iv) Seawards; and v) Shoreline. In the metadata file, images are linked to their corresponding masks, and information about the geographic location of each image, capture characteristics and image source, shoreline position and other auxiliary data are provided. The README file enhances the explainability and comprehension of the dataset, elaborating on the context and contents, and providing detailed explanations of the metadata, potential limitations, technical aspects of the image processing and annotation stages, usage recommendations, and related works. </span></p> <h1>Technical details</h1> <p>The SCLabels dataset version 1.0.0 is packaged in a compressed file (SCLabels_v1.0.0.zip). A total of 1717 RGB images are shared in JPG format, corresponding masks in PNG format, a metadata file in JSON format, and the README file in PDF format.</p> <h2>Data preprocessing</h2> <p><span>To generate the SCLabels masks, rectified RGB images and their corresponding shorelines were used. RGB images were cropped to the minimum and maximum alongshore pixel coordinates of the shoreline (vertical axis) plus 10 additional pixels above and below to preserve contextual information. A grayscale image was then derived from each cropped RGB image for subsequent pixel labelling. First, a binary mask was derived, marking "NoData'' for black and white padded pixels resulting from the registration and rectification steps. Subsequently, the shoreline was densified, ensuring at least one pixel per row was assigned the "Shoreline" label. Next, "Landwards" and "Seawards" labels were assigned to the right and left of the shoreline. Pixels left unlabelled were categorised as "NotClassified". Finally, masks’ values were reclassified to align with the predefined labels, and the grayscale masks were exported. For additional information, please consult the README file. </span></p> <h2>Data splitting</h2> <p><span>Data splitting requirements may vary depending on the chosen AI approach (e.g., splitting by entire images, image patches, or image rows). Researchers should use a consistent data splitting method and document the approach and splits used in publications. This transparency enables reproducible results and facilitates comparisons between studies.</span></p> <h2>Classes, labels and annotations</h2> <p><span>The SCLabels dataset includes one mask per rectified RGB image, sharing the same width and height. These masks are in greyscale and PNG format, and consist of five different labels:</span></p> <table> <tbody> <tr> <td><strong> Mask value</strong></td> <td><strong> Label</strong></td> <td><strong> Description</strong></td> </tr> <tr> <td>0</td> <td>NoData</td> <td>High probability of being black or white padded pixels, used to pad non-rectangular images within the image registration and rectification processes</td> </tr> <tr> <td>25</td> <td>NotClassified</td> <td>Not labeled pixels</td> </tr> <tr> <td>75</td> <td>Landwards</td> <td>All pixels that are towards the landside with respect to the shoreline (row-wise), excluding “NoData” ones</td> </tr> <tr> <td>150</td> <td>Seawards</td> <td>All pixels that are towards the seaside with respect to the shoreline (row-wise), excluding “NoData” ones</td> </tr> <tr> <td>255</td> <td>Shoreline</td> <td>Pixels intersected by the mapped shoreline densified to cover one pixel per row, at least</td> </tr> </tbody> </table> <h2>Parameters</h2> <p><span>RGB values or any transformation in the colour space can be used as parameters.</span><span> </span></p> <h2>Data sources</h2> <p><span>In the CoastSnap initiative, citizens capture images (oblique smartphone photos) from fixed CoastSnap stations and share them with the scientific managers. Images are subjected to a quality control process, spatially registered to a designated target image, and rectified (georeferencing). The shoreline is subsequently digitised from each rectified image.</span><span> </span></p> <h2>Data quality</h2> <p><span>All images included have been supervised by CSs’ scientific managers. However, citizen scientists take images by smartphones (different camera quality) at irregular intervals across various sites with varying weather and illumination conditions. Users of SCLabels dataset must be aware of this variance. </span></p> <h2>Image resolution</h2> <p><span>The resolution of the images depends on the CoastSnap station and the length of the shoreline, ranging from 241x188 pixels to 801x796 pixels.</span></p> <h2>Spatial coverage</h2> <p><span>The SCLabels dataset version 1.0.0 contains data from five Spanish CoastSnap stations, including sandy beaches in the northwest (</span><span>agrelo</span><span>), the Cíes Islands (</span><span>cies</span><span>), the south (</span><span>cadiz</span><span>), and the Balearic Islands (</span><span>samarador </span><span>and </span><span>arenaldentem</span><span>).</span></p> <table> <tbody> <tr> <td><strong> CoastSnap station</strong></td> <td><strong> Longitude</strong></td> <td><strong> Latitude</strong></td> </tr> <tr> <td><em>agrelo</em></td> <td>-8.772</td> <td>42.331</td> </tr> <tr> <td><em>cies</em></td> <td>-8.900</td> <td>42.226</td> </tr> <tr> <td><em>cadiz</em></td> <td>-6.288</td> <td>36.522</td> </tr> <tr> <td><em>samarador</em></td> <td>3.185</td> <td>39.350</td> </tr> <tr> <td><em>arenaldentem</em></td> <td>2.974</td> <td>39.353</td> </tr> </tbody> </table> <h2>Contact information</h2> <p><span>For further technical inquiries or additional information about the annotated dataset, please contact jsoriano@socib.es.</span></p>
Data from 'Tracability of Forest Reproductive Material with the quality label 'Plant van Hier': A DNA database with genetic profiles of native autochthonous tree and shrub species of Flanders, Belgium'
<h2>Background</h2> <p>Indigenous trees and shrubs play an important role in multifunctional forest management. They form a significant part of the biodiversity in our forests. Forest reproductive material (FRM) of autochthonous Flemish origin is sold under the quality label ‘Plant van Hier’, a certification mark of the Agency for Nature and Forests. To ensure the provenance of the seedlings, we developed a DNA-database of genetic profiles of potential parent trees, using species-specific genetic markers. This database enables the traceability of FRM of the ‘Plant van Hier’ label throughout the entire production chain; from seed harvesting and cultivation to planting by the end user.</p> <p>This database contains the genetic profiles of almost all possible parent trees present within 27 Flemish autochthonous seed orchards of eight ecologically important tree and shrub species: <em>Carpinus betulus</em>, <em>Corylus avellana</em>, <em>Frangula alnus</em>, <em>Populus tremula</em>, <em>Sorbus aucuparia</em>, <em>Tilia cordata</em>, <em>Tilia platyphyllos,</em> and <em>Ulmus laevis</em>. The profiles were established using microsatellite markers (11 to 24 markers per species). New genetic markers were developed for <em>Carpinus betulus</em> and <em>Ulmus laevis</em>. PCR products were run on an ABI 3500 Genetic Analyser (Thermo Fisher Scientific).</p> <h2>Files</h2> <p>The files will be updated when new genotypes are added to the seed orchards. The current data files contain data from genotypes collected in the period 2018-2023. </p> <h3>Species_genotypes</h3> <p>These files contain the genetic fingerprints of the parent trees of autochthonous Flemish seed orchards. Missing data is indicated as ‘MD’. For <em>Carpinus betulus</em>, an octoploid species, the allelic phenotype is given instead of the genotype as the number of times that an allele occurs on a specific locus is not known.</p> <p>The next metadata is additionally given:<br>- Species: the Latin name of the species<br>- Seed_orchard: the name of the seed orchard in which the genotypes are located<br>- Code_seed_orchard: the code of the seed orchard in which the genotypes are located as given in the Register of Flemish Forest Reproductive Material (‘Register bosbouwkundig uitgangsmateriaal’; inbo.be)<br>- Genotype: the fieldname given to the genotype<br>- Origin: the location where the genotype was collected in Flanders, Belgium. Genotypes were collected from natural stands which are assumed to have an autochthonous origin. When the specific location is unknown, the location ‘Flanders’ is given. <br>- Year_sampled: the year in which the genotypes were sampled in the respective seed orchard for genetic analysis.</p> <h3>Species_binsets</h3> <p>These files contain the binsets and allele names that are used to score the alleles of the genotypes in the programme Geneious Prime 2019.3.2 (<a href="https://www.geneious.com">https://www.geneious.com</a>). For <em>Tilia platyphyllos </em>and <em>Tilia cordata</em>, the same binsets were used.</p>
Raw planetary images and boulder labels data (as shapefiles) collected during the BOULDERING Marie Skłodowska-Curie Global fellowship
<p>This database contains 64 large images of craters on the lunar and martian surfaces and 3 images of boulder fields on Earth (see manuscript <a href="https://agupubs.onlinelibrary.wiley.com/doi/full/10.1029/2023JE008013">https://agupubs.onlinelibrary.wiley.com/doi/full/10.1029/2023JE008013</a> for more information on those terrestrial locations). The data was collected during the BOULDERING Marie Skłodowska-Curie Global fellowship between October 2021 and 2024.</p> <p>For each image, the boulder outlines within specific tiles within the image were carefully mapped in QGIS. More information about the labelling procedure can be found in the following manuscript (<a href="https://agupubs.onlinelibrary.wiley.com/doi/full/10.1029/2023JE008013">https://agupubs.onlinelibrary.wiley.com/doi/full/10.1029/2023JE008013</a>). This dataset differs from the previous dataset included along with the manuscript <a href="https://zenodo.org/records/8171052">https://zenodo.org/records/8171052</a>, as it contains more mapped images, especially of boulder populations around young impact structures on the Moon (cold spots). </p> <p>For each location, you will find a raster with a .tif format, and three shapefiles:</p> <ul> <li> <p>a boulder-mapping file, which is the manually digitized outline of boulders.</p> </li> <li> <p>a tiles-completely-mapped file, which depicts the patches/tiles/windows on which the boulder mapping has been conducted.</p> </li> <li> <p>a global-tiles file, which shows all of the image patches/tiles/windows (pick the term you are the most familiar with) within a raster.</p> </li> </ul> <p>In addition you will find .pkl (which stands for pickle), which contains some information about the patches/tiles/windows if you would need to clip those windows out from the original raster. You can find more information in the way we process this raw data into a format which can be ingested in a deep learning model (see <a href="https://zenodo.org/records/14250874" target="_blank" rel="noopener">https://zenodo.org/records/14250874</a>) in the two following github repositories (<a href="https://github.com/astroNils/YOLOv8-BeyondEarth" target="_blank" rel="noopener">https://github.com/astroNils/YOLOv8-BeyondEarth</a> and <a href="https://github.com/astroNils/MLtools/tree/main" target="_blank" rel="noopener">https://github.com/astroNils/MLtools</a>). If you don't plan in adding more training data, you can directly used the pre-processed database (see <a href="https://zenodo.org/records/14250874" target="_blank" rel="noopener">https://zenodo.org/records/14250874</a>).</p> <p>There are multiple locations/images per planetary body. Cold spots are located on the Moon, but they are saved in a folder of their own. </p> <p>Note that the cold spots boulder mapping shapefiles are partially manually mapped, and partially originating from predictions made from a deep learning model (which explains the outline of boulders are predicted within one pixel).</p> <p><strong>How to cite:</strong></p> <p>Please refer to the "how to cite" section of the readme file of <a href="https://github.com/astroNils/YOLOv8-BeyondEarth" target="_blank" rel="noopener">https://github.com/astroNils/YOLOv8-BeyondEarth.</a></p> <p><strong>Structure:</strong></p> <pre><code>. └── raw_data/ ├── coldspots/ │ └── image_name/ │ ├── shp/ │ │ ├── <image_name>-tiles-completely-mapped.shp │ │ ├── <image_name>-boulder-mapping.shp │ │ └── <image_name>-global-tiles.shp │ └── raster/ │ └── <image_name>.tif ├── earth/ │ └── image_name/ │ ├── shp/ │ │ ├── <image_name>-tiles-completely-mapped.shp │ │ ├── <image_name>-boulder-mapping.shp │ │ └── <image_name>-global-tiles.shp │ └── raster/ │ └── <image_name>.tif ├── mars/ │ └── image_name/ │ ├── shp/ │ │ ├── <image_name>-tiles-completely-mapped.shp │ │ ├── <image_name>-boulder-mapping.shp │ │ └── <image_name>-global-tiles.shp │ └── raster/ │ └── <image_name>.tif └── moon/ └── image_name/ ├── shp/ │ │ ├── <image_name>-tiles-completely-mapped.shp │ │ ├── <image_name>-boulder-mapping.shp │ │ └── <image_name>-global-tiles.shp └── raster/ └── <image_name>.tif</code></pre>
BWILD: Beach seagrass Wrack Identification Labelled Dataset
<h1>Training dataset</h1> <p>BWILD is a dataset tailored to train Artificial Intelligence applications to automate beach seagrass wrack detection in RGB images. It includes oblique RGB images captured by SIRENA beach video-monitoring systems, along with corresponding annotations, auxiliary data and a README file. BWILD encompasses data from two microtidal sandy beaches in the Balearic Islands, Spain. The dataset consists of images with varying fields of view (9 cameras), beach wrack abundance, degrees of occupation, and diverse meteoceanic and lighting conditions. The annotations categorise image pixels into five classes: i) Landwards, ii) Seawards, iii) Diffuse wrack, iv) Intermediate wrack, and v) Dense wrack.</p> <h1>Technical details</h1> <p>The BWILD version 1.1.0 is packaged in a compressed file (BWILD_v1.1.0.zip). A total of 3286 RGB images are shared in PNG format, corresponding annotations and masks in various formats (PNG, XML, JSON,TXT), and the README file in PDF format.</p> <h2>Data preprocessing</h2> <p>The BWILD dataset utilizes snapshot images from two SIRENA beach video-monitoring systems. To facilitate annotation while maintaining a diverse range of scenarios, the original 1280x960 pixel images were cropped to smaller regions, with a uniform resolution of 640x480 pixels. A subset of images was carefully curated to minimize annotation workload while ensuring representation of various time periods, distances to camera, and environmental conditions. Image selection involved filtering for quality, clustering for diversity, and prioritizing scenes containing beach seagrass wracks. Further details are available in the README file. </p> <h2>Data splitting</h2> <p>Data splitting requirements may vary depending on the chosen Artificial Intelligence approach (e.g., splitting by entire images or by image patches). Researchers should use a consistent method and document the approach and splits used in publications, enabling reproducible results and facilitating comparisons between studies. </p> <h2>Classes, labels and annotations</h2> <p>The BWILD dataset has been labelled manually using the 'Computer Vision Annotation Tool' (CVAT), categorising pixels into five labels of interest using polygon annotations.</p> <table> <tbody> <tr> <td><strong> Label</strong></td> <td><strong> Description</strong></td> </tr> <tr> <td>landwards</td> <td>Pixels that are towards the landside with respect to the shoreline</td> </tr> <tr> <td>seawards</td> <td>Pixels that are towards the seaside with respect to the shoreline</td> </tr> <tr> <td>diffuse wrack</td> <td>Pixels that potentially resembled beach wracks based on colour and shape, yet the annotator could not confirm this with certainty, were denoted as ‘diffuse wrack’</td> </tr> <tr> <td>Intermediate wrack</td> <td>Pixels with low-density beach wracks or mixed beach wracks and sand surfaces</td> </tr> <tr> <td>Dense wrack</td> <td>Pixels with high-density beach wracks</td> </tr> </tbody> </table> <p>Annotations were exported from CVAT in four different formats: (i) CVAT for images (XML); (ii) Segmentation Mask 1.0 (PNG); (iii) COCO (JSON); (iv) Ultralytics YOLO Segmentation 1.0 (TXT). These diverse annotation formats can be used for various applications including object detection and segmentation, and simplify the interaction with the dataset, making it more user-friendly. Further details are available in the README file. </p> <h2>Parameters</h2> <p>RGB values or any transformation in the colour space can be used as parameters.</p> <h2>Data sources</h2> <p>A SIRENA system consists of a set of RGB cameras mounted at the top of buildings on the beachfront. These cameras take oblique pictures of the beach, with overlapping sights, at 7.5 FPS during the first 10 minutes of each hour in daylight hours. From these pictures, different products are generated, including snapshots, which correspond to the frame of the video at the 5th minute. In the Balearic Islands, SIRENA stations are managed by the Balearic Islands Coastal Observing and Forecasting System (SOCIB), and are mounted at the top of hotels located in front of the coastline. The present dataset includes snapshots from the SIRENA systems operating since 2011 at Cala Millor (5 cameras) and Son Bou (4 cameras) beaches, located in Mallorca and Menorca islands (Balearic Islands, Spain), respectively. All latest and historical SIRENA images are available at the Beamon app viewer (https://apps.socib.es/beamon). </p> <h2>Data quality</h2> <p>All images included in BWILD have been supervised by the authors of the dataset. However, variable presence of beach segrass wracks across different beach segments and seasons impose a variable distribution of images across different SIRENA stations and cameras. Users of BWILD dataset must be aware of this variance. Further details are available in the README file. </p> <h2>Image resolution</h2> <p>The resolution of the images in BWILD is of 640x480 pixels.</p> <h2>Spatial coverage</h2> <p>The BWILD version 1.1.0 contains data from two SIRENA beach video-monitoring stations, encompassing two microtidal sandy beaches in the Balearic Islands, Spain. These are: Cala Millor (<em>clm</em>) and Son Bou (<em>snb</em>). </p> <table> <tbody> <tr> <td><strong>SIRENA station</strong></td> <td><strong> Longitude</strong></td> <td><strong> Latitude</strong></td> </tr> <tr> <td><em>clm</em></td> <td>3.383</td> <td>39.596</td> </tr> <tr> <td><em>snb</em></td> <td>4.077</td> <td>39.898</td> </tr> </tbody> </table> <h2>Contact information</h2> <p>For further technical inquiries or additional information about the annotated dataset, please contact jsoriano@socib.es.</p>
Labeled Images at OBSEA for Object Detection Algorithms
<p>Images from OBSEA underwater cameras labeled with marine species to train AI-based Object Detection algorithms.</p>
Dataset: Label-free detection of methicillin resistance in Staphylococcus aureus using different Raman-spectroscopy approaches
<p>This is the dataset accompanying the submission of the manuscript: Label-free detection of methicillin resistance in Staphylococcus aureus using different Raman-spectroscopy approaches in the journal Microbiology Spectrum.</p> <p>The data description is the following:</p> <p>Strains<br> 16859MRSA= Strain AUSTR-07-16859 MRSA<br> 16859MSSA= Strain AUSTR-07-16859 MSSA<br> CC8MRSA= Strain 08V15773<br> CC8MSSA= Strain MRSA2010-174<br> AUSTR05MRSA= Strain AUSTR-05-15441 MRSA<br> AUSTR05MSSA= Strain AUSTR-05-15441 MSSA<br> CC361MRSA= Strain UAE-Abu Dhabi-020<br> CC361MSSA= Strain UAE-Dubai-80-MS 1368.9/09</p> <p>Datasets<br> UVRR: UV-Resonance Raman with 244 nm excitation on bulk samples, calibration standard Polystyrene, measurements were time series of 10 consecutive spectra, for each strain and batch 25 time series were collected from 3 different slides<br> 532nm: Single cell analysis with 532nm excitation, calibration standard 4AAP, one spectrum per bacterial cell was collected<br> 785nm: Bulk analysis of bacterial colonies using 785 nm excitation and a Raman fibre probe, calibration standard 4AAP, bulk analysis, individual spectra of colonies were collected</p> <p>Data structure is in the metadata files.<br> Individual spectra are in the folders sorted by the date they were measured.</p>
LigPCDS: Labeled Dataset of X-ray Protein Ligand Images in 3D Point Cloud and Validated Deep Learning Models
<p>The difference electron density from X-ray protein crystallography was used to create the first dataset of labeled ligand images in 3D point clouds, named <strong>LigPCDS</strong>. The dataset contain 244,226 entries of free organic ligands containing 3D representations labeled with two major labeling approaches: SP-based and AtomSymbol-based.</p> <p> </p> <p>The data from free organic molecules (non-covalent ligands) was retrieved from the Research Collaboratory for Structural Bioinformatics Protein Data Bank (RCSB PDB) in december 2019 with resolutions ranging from 1.5 to 2.2 Å. The ligand images (blobs) were interpolated from their calculated difference electron density map in a 3D grid-like bounding box, around their atomic positions, and stored in point clouds. These ligand grid representations were further processed to retrive the final ligands representation in 3D point clouds using a mask of the shape of the ligand. A grid spacing of 0.5 Å gave the best results. The density value of the grid points was used as feature. The labeling approach used the structure of the ligands to propose vocabularies of chemical classes based on the chemical atoms themselves and their cyclic substructures. These structure annotations were applied pointwise to the ligand 3D representations using an atomic sphere model. Four proposed vocabularies were validated by successfully training good performance deep learning models for the semantic segmentation of a stratified dataset from LigPCDS, using 78902 entries.</p> <p>The four validated deep learning models are: (i) the LigandRegion, composed by generic atoms of any type; (ii) the AtomCycle, composed by generic atoms outside cycles and generic cycles; (iii) the AtomC347CA56, composed by generic atoms outside cycles, not aromatic cycles of size 3 to 7 and aromatic cycles of size 5 and 6; and (iv) the AtomSymbolGroups, composed by the atoms symbols with groupings. The mean accuracy of these models in their cross-validation was between 49.7% <span lang="EN-GB">[-19.4,20.</span><span lang="EN-GB">2]</span> and 77.4% <span lang="EN-GB">[-11.7,12.1]</span> in terms of Intersection over Union (mIoU) metric and between 62.4% <span lang="EN-GB">[-18.8,19.</span><span lang="EN-GB">7]</span> and 87.0% <span lang="EN-GB">[-8.4,8.8]</span> in F1-score (mF1), confidence interval between squared brackets. The models i, ii and iii and the used labeled representations in 3D point cloud are contained in the SP-based record; and model iv and its used labeled representations are contained in the AtomSymbol-based record.</p> <p>The dataset and validated models may be used to tackle problems regarding known and unknown ligand building to drug discovery and fragment screening pipelines. </p> <p>The code used to create and validated the LigPCDS is available at the following repository: https://github.com/danielatrivella/np3_ligand</p> <p>This repository also contains the NP³ Blob Label application for ligand building using the validated deep learning models from LigPCDS.</p>
Record Label & Music Publishing Turnover in Europe
<p>Imputed and forecasted values of the recording and music publishing industry from the <a href="https://appsso.eurostat.ec.europa.eu/nui/show.do?dataset=sbs_na_1a_se_r2&lang=en">Annual detailed enterprise statistics for services (NACE Rev. 2 H-N and S95)</a> Eurostat folder.</p>
DRM-labeled HIV-1 protease sequence dataset
<p>An HIV-1 protease dataset with labeled DRMs derived from the Stanford HIV Drug Resistance Database<sup>1,2</sup> is provided. It is in <em>fasta</em> format with major protease drug resistance mutations (as defined by Wensing et al.<sup>3</sup>) provided in the sequence name section (following the ">" symbol) as a comma-separated list. The dataset was used in the following paper: "Ahmed A., de Souza D. R., Link R. W., Nonnemacher M. R., Wigdahl B., Dampier W. Design of a SHERLOCK-based low resource screening assay for HIV-1 drug resistance, in preparation, 2021. "</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.