Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
118
datasets available to search
ShareScore release 0.9.0
Dataset results
118 results for “Object detection”
Variable Message Signal annotated images for object detection
<p><strong>If you use this dataset, please cite this paper: <em>Puertas, E.; De-Las-Heras, G.; Sánchez-Soriano, J.; Fernández-Andrés, J. Dataset: Variable Message Signal Annotated Images for Object Detection. Data 2022, 7, 41. https://doi.org/10.3390/data7040041</em></strong></p> <p>This dataset consists of Spanish road images taken from inside a vehicle, as well as annotations in XML files in PASCAL VOC format that indicate the location of Variable Message Signals within them. Also, a CSV file is attached with information regarding the geographic position, the folder where the image is located, and the text in Spanish. This can be used to train supervised learning computer vision algorithms, such as convolutional neural networks. Throughout this work, the process followed to obtain the dataset, image acquisition, and labeling, and its specifications are detailed. The dataset is constituted of <strong>1216</strong> instances, <strong>888</strong> positives, and <strong>328</strong> negatives, in <strong>1152</strong> jpg images with a resolution of 1280x720 pixels. These are divided into <strong>576</strong> real images and <strong>576</strong> images created from the data-augmentation technique. The purpose of this dataset is to help in road computer vision research since there is not one specifically for VMSs.</p> <p>The folder structure of the dataset is as follows:</p> <ul> <li>vms_dataset/ <ul> <li>data.csv</li> <li>real_images/ <ul> <li>imgs/</li> <li>annotations/</li> </ul> </li> <li>data-augmentation/ <ul> <li>imgs/</li> <li>annotations/</li> </ul> </li> </ul> </li> </ul> <p>In which:</p> <ul> <li><strong>data.csv:</strong> Each row contains the following information separated by commas (,): image_name, x_min, y_min, x_max, y_max, class_name, lat, long, folder, text.</li> <li><strong>real_images:</strong> Images extracted directly from the videos.</li> <li><strong>data-augmentation:</strong> Images created using data-augmentation</li> <li><strong>imgs:</strong> Image files in .jpg format.</li> <li><strong>annotations:</strong> Annotation files in .xml format.</li> </ul>
artificial dataset of small engine parts for object detection-segmentation
<p>This dataset contains images and masks of parts used in engine assembly. The images were generated artificially from CAD models using gazebo simulator. The dataset consists of three classes: Large Bolt, Small Bolt and Rocker Arm. Annotations are represented as mask images with the same name as corresponding RGB images. 1080 images for each class, 3240 images total. Resolution of each image: 640x480 pixels. CAD models for each part are also included.</p>
KITTI 3D Object Detection
<p>The dataset downloaded from official KITTI website was used in eGAC3D's evaluation. This dataset includes raw images (data_object_label_2.zip) and corresponding ground-truth (data_object_image_2.zip). </p>
DataSet for UAV-based Untrained Small Object Detection using Distance Metric Method
<p>DataSet for UAV-based Untrained Small Object Detection using Distance Metric Method</p>
Corpus Nummorum - Object Detection Coin Dataset
<p>This Object Detection dataset is a collection of ancient coin images from three different sources: the <a href="https://www.corpus-nummorum.eu/">Corpus Nummorum (CN)</a> project, the <a href="https://ikmk.smb.museum/home?lang=en">Münzkabinett Berlin</a> and the <a href="https://www.bnf.fr/fr/departement-monnaies-medailles-antiques">Bibliothèque nationale de France, Département des Monnaies, médailles et antiques</a>. It covers Greek and Roman coins from ancient Thrace, Moesia Inferior, Troad and Mysia. This is a selection of the coins published on the <a href="https://www.corpus-nummorum.eu/search/coins">CN</a> portal (due to copyrights). </p> <p>This dataset contains 506 different classes with about 179.000 coin images (approx. 29.000 unique coins). The classes come from four different categories: persons, objects, animals and plants. The coin images were assigned to the classes using our <a href="https://github.com/Frankfurt-BigDataLab/NLP-on-multilingual-coin-datasets">NLP pipeline</a>. For this purpose, our Named Entity Recognition and Relation Extraction were performed on every coin's description (separated into obverse and reverse). Each coin image assigned to this description was then <span>copied to the folder of the predicted classes</span>. A coin image can therefore also be assigned to different classes. The file name contains both the coin id and the coin type of the <a href="https://www.corpus-nummorum.eu/">CN database</a>. Whether the image belongs to a coin obverse or reverse can be recognized by the suffix obv or rev. An "sources" csv file holds the sources for every image. Due to copyrights the image size is limited to 299*299 pixels. However, this should be sufficient for most ML approaches.</p> <p>Due to the numerically different occurrences of the individual entities, the data set is not balanced. In addition, a class can contain very different representations of the same entity. Therefore, some classes can be difficult to train. Unfortunately, we cannot provide any annotations for the data set.</p> <p>During the summer semester 2024, we held the "Data Challenge" event at our Department of Computer Science at the Goethe-University. Our students could choose between the Object Detection dataset and a Natural Language dataset as their challenge. One team opted for the Object Detection challenge. We gave them this dataset with the task to use to try out their own ideas. Here are their results:</p> <p><a href="https://github.com/PatrickMelnic/DataChallenges_ObjectDetection/">Multilabel Classification as Backbone for Object Detection</a></p> <p> </p> <p>Now we would like to invite you to try out your own ideas and models on our coin data.</p> <p>If you have any questions or suggestions, please, feel free to contact us. </p>
Dangerous Items Dataset for 5-Class Object Detection (Pascal VOC annotation)
<p>This repository contains the data from the manuscript "Dangerous Items Detection in Surveillance Camera Images Using Faster R-CNN". It contains 4,000 images representing 5 classes of objects: baseball bat, gun, knife, machete and rifle (<a href="https://drive.google.com/file/d/1aG30Hpupctnp8ywhAprz313jmYbFnZk3/view?usp=drive_link" target="_blank" rel="noopener">https://drive.google.com/file/d/1aG30Hpupctnp8ywhAprz313jmYbFnZk3/view?usp=drive_link</a>). All images were scaled so that the smaller side is no shorter than 600 pixels and the larger one is no longer than 1,000 pixels. The full set was randomly divided into training and testing parts (75% and 25% of the full set respectively). As a result, the training part contains 3,000 images, and the testing part – 1,000 images. Both parts are balanced, that is, they contain similar number of objects to be detected. Images were annotated using bounding boxes in the Pascal VOC format, according to which the description about each image is included in the corresponding XML file.</p> <p>Using this dataset please cite:<br>Omiotek, Z. (2025). Dangerous items’ detection in surveillance camera images using Faster R‑CNN. Przegląd Elektrotechniczny, 101(5), 156-168. https://www.red.pe.org.pl/articles/2025/5/36.pdf</p>
Popular Animals Dataset for 6-Class Object Detection (Pascal VOC annotation)
<p>This dataset contains images used in the monograph titled <em>Zastosowanie wybranych metod uczenia głębokiego w wizji komputerowej</em> (Application of Selected Deep Learning Methods in Computer Vision) to build the Faster R-CNN model. The full collection consists of 600 image files showing 6 classes of objects: <em>kot</em> (cat), <em>krowa</em> (cow), <em>pies</em> (dog), <em>koń</em> (horse), <em>człowiek</em> (human), <em>owca</em> (sheep) (<a title="Popular Animals Dataset" href="https://drive.google.com/file/d/1g3O1Qq2YqmCb8WSMDHaoFJPoHe6yujlQ/view?usp=drive_link" target="_blank" rel="noopener">https://drive.google.com/file/d/1g3O1Qq2YqmCb8WSMDHaoFJPoHe6yujlQ/view?usp=drive_link</a>). All images were scaled so that the smaller side is no shorter than 600 pixels and the larger side is no longer than 1000 pixels. The set was randomly divided into a training part (75% of the full set) and a test part (25% of the full set). As a result, the training part contains 450 files, and the test part - 150. Both parts are balanced - they contain a similar number of detected objects. The images are labeled with bounding boxes in the Pacal VOC format, according to which the description of each file is contained in an XML file with the same name. The dataset can be used to build models for object detection.</p>
BadODD: Bangladeshi Autonomous Vehicle Object Detection Dataset
<p>The dataset covers the following 9 districts in Bangladesh: Sylhet, Dhaka, Rajshahi, Mymensingh, Maowa, Chittagong, Sirajganj, Sherpur, and Khulna. Participants will encounter a wide range of road types, including towns, expressways, highways, and village roads. This diversity in locations aims to challenge algorithms to perform well across various driving contexts commonly encountered on Bangladesh roads.</p> <p> </p>
Deep-sea observatories images labeled by citizen for object detection algorithms
<div>All information is available from the original publication page: <a href="https://doi.org/10.17882/101899">https://doi.org/10.17882/101899</a>.</div>
RBOD: An annotated satellite imagery dataset for automated river barrier object detection
<h3>Introduction:</h3> <p>Millions of river barriers have been constructed worldwide for flood control, hydropower generation, and agricultural irrigation. The lack of comprehensive records on the locations and types of river barriers, particularly small barriers such as weirs, limits our ability to assess their societal and environmental impacts. Integrating satellite imagery with object detection algorithms holds promise for the automatic identification of river barriers on a global scale. However, achieving this objective requires high-quality image datasets for algorithm training and testing. Hence, this study presents a large-scale dataset named the River Barrier Object Detection (RBOD), making it the first publicly available dataset specifically for river barrier object detection.</p> <p>The RBOD dataset comprises 4,872 high-resolution satellite images and 11,741 meticulously annotated oriented bounding boxes (OBBs). In this dataset, river barriers can be classified into five classes: dams, groynes, locks, sluices, and weirs. The effectiveness of the RBOD dataset was validated using five typical object detection algorithms, namely YOLOV8-OBB, Oriented R-CNN, Rotated Faster R-CNN, R3Det, and Rotated RetinaNet, which provide performance benchmarks for future applications.</p> <h3><span>Usage Notes:</span></h3> <p>The RBOD dataset consists of three folders (namely, '<em>images</em>', '<em>labels_voc</em>', and '<em>labels_yolo</em>') and a .txt file named '<em>class</em>':</p> <p>·'<em>images</em>' folder - contains 4872 satellite images (.jpg).</p> <p>·'<em>labels_voc</em>' folder - contains 11,741 .xml files for annotations in PASCAL VOC format. In these .xml files, the position of OBB is represented as (cx, cy, width, height, angle), where 'cx' and 'cy' denote the center coordinates, 'width' and 'height' are the lengths along the x- and y-axes, and 'angle' is the clockwise rotation angle relative to the x-axis.</p> <p>·'<em>labels_yolo</em>' folder - contains 11,741 .txt files for annotations in YOLO format. In these .txt files, the OBB is represented as (class_index, x1, y1, x2, y2, x3, y3, x4, y4), where ‘class_index’ denotes the target category, and ‘x1, y1, x2, y2, x3, y3, x4, y4’ are the normalized coordinates of the four corners of the bounding box.</p> <p>·'<em>class</em>' txt file - record the classifications of river barriers and their indices, which correspond to the ‘class_index’.</p> <p>Note that each folder splits into three subfolders: train (70%), test (20%), and val (10%).</p>
Object detection for graphical user interface: old fashioned or deep learning or a combination? - Model&Datasets
<p>This repo contains the datasets, trained models, and data splitting in ESEC/FSE 2020 "Object detection for graphical user interface: old fashioned or deep learning or a combination?" paper.</p>
VISIONE Feature Repository for VBS: Multi-Modal Features and Detected Objects from V3C1+V3C2 Dataset
<p>This repository contains a diverse set of features extracted from the V3C1+V3C2 dataset, sourced from the Vimeo Creative Commons Collection. These features were utilized in the VISIONE system [Amato et al. 2023, Amato et al. 2022] during the latest editions of the Video Browser Showdown (VBS) competition (<a href="https://www.videobrowsershowdown.org/">https://www.videobrowsershowdown.org/</a>).</p> <p>The original V3C1+V3C2 dataset, provided by NIST, can be downloaded using the instructions provided at <a href="https://videobrowsershowdown.org/about-vbs/existing-data-and-tools/">https://videobrowsershowdown.org/about-vbs/existing-data-and-tools/</a>.</p> <p>It comprises 7,235 video files, amounting for 2,300h of video content and encompassing 2,508,113 predefined video segments.</p> <p>We subdivided the predefined video segments longer than 10 seconds into multiple segments, with each segment spanning no longer than 16 seconds. As a result, we obtained a total of 2,648,219 segments. For each segment, we extracted one frame, specifically the middle one, and computed several features, which are described in detail below.</p> <p>This repository is released under a Creative Commons Attribution license. If you use it in any form for your work, please cite the following paper:</p> <blockquote> <pre>@inproceedings{amato2023visione, title={VISIONE at Video Browser Showdown 2023}, author={Amato, Giuseppe and Bolettieri, Paolo and Carrara, Fabio and Falchi, Fabrizio and Gennaro, Claudio and Messina, Nicola and Vadicamo, Lucia and Vairo, Claudio}, booktitle={International Conference on Multimedia Modeling}, pages={615--621}, year={2023}, organization={Springer} }</pre> </blockquote> <p> </p> <p>This repository comprises the following files:</p> <ul> <li><strong><em>msb.tar.gz </em></strong> contains tab-separated files (.tsv) for each video. Each tsv file reports, for each video segment, the timestamp and frame number marking the start/end of the video segment, along with the timestamp of the extracted middle frame and the associated identifier ("id_visione"). </li> <li><em><strong>extract-keyframes-from-msb.tar.gz</strong></em> contains a Python script designed to extract the middle frame of each video segment from the MSB files. To run the script successfully, please ensure that you have the original V3C videos available.</li> <li><strong><em>features-aladin.tar.gz<sup>†</sup></em> </strong>contains <a href="https://github.com/mesnico/ALADIN">ALADIN</a> [Messina N. et al. 2022] features extracted for all the segment's middle frames. </li> <li><em><strong>features-clip-laion.tar.gz<sup>†</sup></strong></em> contains <a href="https://huggingface.co/laion/CLIP-ViT-H-14-laion2B-s32B-b79K">CLIP ViT-H/14 - LAION-2B </a>[Schuhmann et al. 2022] features extracted for all the segment's middle frames.</li> <li><em><strong>features-clip-openai.tar.gz<sup>†</sup> </strong></em>contains <a href="https://huggingface.co/openai/clip-vit-large-patch14">CLIP ViT-L/14</a> [Radford et al. 2021] features extracted for all the segment's middle frames. </li> <li><em><strong>features-clip2video.tar.gz<sup>†</sup> </strong></em>contains <a href="https://github.com/CryhanFang/CLIP2Video">CLIP2Video</a> [Fang H. et al. 2021] extracted for all the video segments. <strong> </strong>In particular 1) we concatenate consecutive short segments so to create segments at least 3 seconds long; 2) we downsample the obtained segments to 2.5 fps; 3) we feed the network with the first min(36, n) frames, where n is the number of frames of the segment. Notice that the minimum processed length consists of 7 frames, given that the segment is no shorter than 3s. </li> <li><em><strong>objects-frcnn-oiv4.tar.gz<sup>*</sup> </strong></em>contains the objects detected using <a href="http://tfhub.dev/google/faster_rcnn/openimages_v4/inception_resnet_v2/1">Faster R-CNN+Inception ResNet</a> (trained on the Open Images V4 [Kuznetsova et al. 2020]). </li> <li><em><strong>objects-mrcnn-lvis.tar.gz<sup>*</sup></strong></em> contains the objects detected using Mask R-CNN [He et al. 2017] (trained on LVIS).</li> <li><em><strong>objects-vfnet64-coco.tar.gz<sup>*</sup></strong></em> contains the objects detected using VfNet [Zhang et al. 2021] (trained on COCO dataset).</li> </ul> <p>*Please be sure to use the <strong>v2 version </strong>of this repository, since v1 feature files may contain inconsistencies that have now been corrected</p> <p><em><strong>*Note on the object annotations:</strong></em> Within an object archive, there is a jsonl file for each video, where each row contains a record of a video segment (the <em>"_id"</em> corresponds to the <em>"id_visione"</em> used in the msb.tar.gz) . Additionally, there are three arrays representing the objects detected, the corresponding scores, and the bounding boxes. The format of these arrays is as follows:</p> <ul> <li><em>"object_class_names"</em>: vector with the class name of each detected object.</li> <li><em>"object_scores"</em>: scores corresponding to each detected object.</li> <li><em>"object_boxes_yxyx"</em>: bounding boxes of the detected objects in the format <em>(ymin, xmin, ymax, xmax).</em></li> </ul> <p> </p> <p><em><strong><sup>†</sup>Note on the cross-modal features: </strong></em>The extracted multi-modal features (ALADIN, CLIPs, CLIP2Video) enable internal searches within the V3C1+V3C2 dataset using the query-by-image approach (features can be compared with the dot product). However, to perform searches based on free text, the text needs to be transformed into the joint embedding space according to the specific network being used. Please be aware that t<strong>he service for transforming text into features is not provided within this repository and should be developed independently using the original feature repositories linked above.</strong></p> <p>We have plans to release the code in the future, allowing the reproduction of the VISIONE system, including the instantiation of all the services to transform text into cross-modal features. However, this work is still in progress, and the code is not currently available.</p> <p> </p> <p><strong>References:</strong></p> <p>[Amato et al. 2023] Amato, G.et al., 2023, January. VISIONE at Video Browser Showdown 2023. In International Conference on Multimedia Modeling (pp. 615-621). Cham: Springer International Publishing.</p> <p>[Amato et al. 2022] Amato, G. et al. (2022). VISIONE at Video Browser Showdown 2022. In: , et al. MultiMedia Modeling. MMM 2022. Lecture Notes in Computer Science, vol 13142. Springer, Cham. </p> <p>[Fang H. et al. 2021] Fang H. et al., 2021. Clip2video: Mastering video-text retrieval via image clip. arXiv preprint arXiv:2106.11097.</p> <p>[He et al. 2017] He, K., Gkioxari, G., Dollár, P. and Girshick, R., 2017. Mask r-cnn. In Proceedings of the IEEE international conference on computer vision (pp. 2961-2969).</p> <p>[Kuznetsova et al. 2020] Kuznetsova, A., Rom, H., Alldrin, N., Uijlings, J., Krasin, I., Pont-Tuset, J., Kamali, S., Popov, S., Malloci, M., Kolesnikov, A. and Duerig, T., 2020. The open images dataset v4. International Journal of Computer Vision, 128(7), pp.1956-1981.</p> <p>[Lin et al. 2014] Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P. and Zitnick, C.L., 2014, September. Microsoft coco: Common objects in context. In European conference on computer vision (pp. 740-755). Springer, Cham.</p> <p>[Messina et al. 2022] Messina N. et al., 2022, September. Aladin: distilling fine-grained alignment scores for efficient image-text matching and retrieval. In Proceedings of the 19th International Conference on Content-based Multimedia Indexing (pp. 64-70).</p> <p>[Radford et al. 2021] Radford A. et al., 2021, July. Learning transferable visual models from natural language supervision. In International conference on machine learning (pp. 8748-8763). PMLR.</p> <p>[Schuhmann et al. 2022] Schuhmann C. et al., 2022. Laion-5b: An open large-scale dataset for training next generation image-text models. Advances in Neural Information Processing Systems, 35, pp.25278-25294.</p> <p>[Zhang et al. 2021] Zhang, H., Wang, Y., Dayoub, F. and Sunderhauf, N., 2021. Varifocalnet: An iou-aware dense object detector. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 8514-8523).</p>
Infrared Salient Object Detection (ISOD)
<p>It is a Thermal Infrared Image dataset that focuses on salient objects.</p>
Objective Quality of Life Detection Validation
ClinicalTrials.gov study NCT04121793. IPD Sharing: NO. Countries: 1. Publications: 1.
Pilot Study for the Early Detection of Chronic Kidney Disease, Non-Dialysis Objective (NDO).
ClinicalTrials.gov study NCT06447038. IPD Sharing: UNDECIDED. Countries: 1. Publications: 1.
LiaoSteve/Trash-Dataset-and-object-detection: Trash-Dataset-and-object-detection
<p>Create my trash dataset and use hexacopter to detect trash pollution with YOLOV3.</p> <pre><code>Trash dataset : 1. bottle 2. bag 3. plastic_bag YOLOV3 weight : 1. anchor.txt 2. classes.txt 3. weight.h5 </code></pre>
The SoccerSum Dataset for Automated Detection, Segmentation, and Tracking of Objects on the Soccer Pitch
<p>The SoccerSum dataset was curated by capturing and annotating soccer videos from the Norwegian Eliteserien league. This collection spans three years, covering 2021, 2022, and 2023. It comprises 750 frames from 41 unique sequences, with 4 to 40 frames per sequence, carefully chosen to represent a diverse selection of scenarios encountered in professional soccer games.</p>
Dataset of the Floating Objects Detection Notebook
<p>This dataset contains the data used in the notebook "Detecting floating objects using Deep Learning and Sentinel-2 imagery", published in the ocean modelling section of The Environmental Data Science Book.</p>
Classical Center-Surround Receptive Fields Facilitate Novel Object Detection in Retinal Bipolar Cells
<p>Database and Figures in Igor format</p>
auto_dataset_for_object_detection
<p>auto_dataset_for_object_detection</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.