Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
118
datasets available to search
ShareScore release 0.7.1
Dataset results
118 results for “object detection”
Smartbay Marine Species Object Detection Training dataset
<h1>Training dataset</h1> <p>The SmartBay Observatory in Galway Bay is an important contribution by Ireland to the growing global network of real-time data capture systems deployed within the ocean – technology giving us new insights into the ocean which we have not had before.</p> <p>The observatory was installed on the seafloor 1.5km off the coast of Spiddal, County Galway, Ireland . The observatory uses cameras, probes and sensors to permit continuous and remote live underwater monitoring. This observatory equipment allows ocean researchers unique real-time access to monitor ongoing changes in the marine environment. Data relating to the marine environment at the site is transferred in real-time from the SmartBay Observatory through a fibre optic telecommunications cable to the Marine Institute headquarters and onwards onto the internet. The data includes a live video stream, the depth of the observatory node, the sea temperature and salinity, and estimates of the chlorophyll and turbidity levels in the water which give an indication of the volume of phytoplankton and other particles, such as sediment, in the water.</p> <p>The Smartbay Marine Species Object Detection training Dataset is an initial Bounding Box Annotated image dataset used in attempting to Train a YOLOv8 Object Detection Model to classify the Marine Fauna observed in the Smartbay Observatory Video footage using species names.</p> <p>The imagery used in this training dataset consists of image frame captures from the <a href="https://smartbay.marine.ie">Smartbay</a> video Archive files, CC-BY imagery from the <a href="https://www.minka-sdg.org">www.minka-sdg.org</a> website and images taken by Eva Cullen in the "<a href="https://nationalaquarium.ie/">Galway Atlantaquaria</a>" Aquarium in Galway, Ireland.</p> <p>The imagery were annotated using CVAT, collated on <a href="https://www.roboflow.com/">Roboflow</a> and exported in YOLOv8 training dataset format. </p>
Smartbay Marine Types Object Detection Training dataset
<h1>Training Dataset</h1> <p>The SmartBay Observatory in Galway Bay is an important contribution by Ireland to the growing global network of real-time data capture systems deployed within the ocean – technology giving us new insights into the ocean which we have not had before.</p> <p>The observatory was installed on the seafloor 1.5km off the coast of Spiddal, County Galway, Ireland . The observatory uses cameras, probes and sensors to permit continuous and remote live underwater monitoring. This observatory equipment allows ocean researchers unique real-time access to monitor ongoing changes in the marine environment. Data relating to the marine environment at the site is transferred in real-time from the SmartBay Observatory through a fibre optic telecommunications cable to the Marine Institute headquarters and onwards onto the internet. The data includes a live video stream, the depth of the observatory node, the sea temperature and salinity, and estimates of the chlorophyll and turbidity levels in the water which give an indication of the volume of phytoplankton and other particles, such as sediment, in the water.</p> <p>The Smartbay Marine Types Object Detection training Dataset is an initial Bounding Box Annotated image dataset used in attempting to Train a YOLOv8 Object Detection Model to classify the Marine Fauna observed in the Smartbay Observatory Video footage using broad "Marine Type" classes.</p> <p>The imagery used in this training dataset consists of image frame captures from the <a href="https://smartbay.marine.ie">Smartbay</a> video Archive files, CC-BY imagery from the <a href="https://www.minka-sdg.org">www.minka-sdg.org</a> website and images taken by Eva Cullen in the "<a href="https://nationalaquarium.ie/">Galway Atlantaquaria</a>" Aquarium in Galway, Ireland.</p> <p>The imagery were annotated using CVAT, collated on <a href="https://www.roboflow.com/">Roboflow</a> and exported in YOLOv8 trainign dataset format. </p>
Labeled Images at OBSEA for Object Detection Algorithms
<p>Images from OBSEA underwater cameras labeled with marine species to train AI-based Object Detection algorithms.</p>
The Object Detection for Olfactory References (ODOR) Dataset
<p><strong>The Object Detection for Olfactory References (ODOR) Dataset</strong></p> <p>Real-world applications of computer vision in the humanities require algorithms to be robust against artistic abstraction, peripheral objects, and subtle differences between fine-grained target classes. </p> <p>Existing datasets provide instance-level annotations on artworks but are generally biased towards the image centre and limited with regard to detailed object classes. The ODOR dataset fills this gap, offering 38,116 object-level annotations across 4,712 images, spanning an extensive set of 139 fine-grained categories. </p> <p>It has challenging dataset properties, such as a detailed set of categories, dense and overlapping objects, and spatial distribution over the whole image canvas. </p> <p>Inspiring further research on artwork object detection and broader visual cultural heritage studies, the dataset challenges researchers to explore the intersection of object recognition and smell perception.</p> <p><strong>How to use</strong></p> <p>The annotations are provided in COCO JSON format. To represent the two-level hierarchy of the object classes, we make use of the supercategory field in the categories array as defined by COCO. In addition to the object-level annotations, we provide an additional CSV file with image-level metadata, which includes content-related fields, such as Iconclass codes or image descriptions, as well as formal annotations, such as artist, license, or creation year. </p> <p>In addition to a zip containing the dataset images, we provide links to their source collections in the metadata file and a Python script to conveniently download the artwork images (`download_imgs.py`).</p> <p>The mapping between the `images` array of the `annotations.json` and the `metadata.csv` file can be accomplished via the `file_name` attribute of the elements of the `images` array and the unique `File Name` column of the `metadata.csv` file, respectively.</p>
Nephrops (Nephrops norvegicus) Burrow object detection simple training dataset from Irish Underwater TV surveys
<div> <div> <div> <div> <h1>Training dataset</h1> <p>Norway prawns (<em>Nephrops norvegicus</em>), also known as the Dublin Bay prawn, are common around the Irish coast. They are found in distinct sandy/muddy areas where the sediment is suitable for them to construct their burrows. <em>Nephrops </em>spend a great deal of time in their burrows and their emergence from these is related to time of year, light intensity and tidal strength. The Irish <em>Nephrops </em>fishery is extremely valuable with landings recently worth around €55m at first sale, supporting an important Irish fishing industry. </p> <p><em>Nephrops</em> are managed in Functional Units (FUs). The Marine Institute has conducted under water television surveys since 2002 to independently estimate abundance, distribution and stock sizes of <em>Nephrops</em> <em>norvegicus </em>for:</p> <ul> <li>Irish Sea <em>Nephrops</em> Grounds (FU 14 and 15) in collaboration with <a title="Link to 'Fisheries and Aquatic Ecosystems' work in AFBI Northern Ireland" href="https://www.afbini.gov.uk/area-of-expertise/fisheries-and-aquatic-ecosystems">AFBI</a> an <a title="Link to Cefas (the Centre for Environment, Fisheries, and Aquaculture Science) in the UK" href="https://www.cefas.co.uk/">CEFAS</a>.</li> <li>Porcupine Bank <em>Nephrops</em> Grounds (FU16)</li> <li>Aran, Galway Bay and Slyne Head <em>Nephrops</em> Grounds (FU17)</li> <li>South and South west Ireland <em>Nephrops</em> Grounds (FU19)</li> <li>Labadie, Jones and Cockburn <em>Nephrops</em> Grounds (FU20 and 21)</li> <li>“Smalls” <em>Nephrops</em> Grounds (FU22)</li> </ul> <p>Each year during the summer months, on average 300 stations are surveyed each year, in three survey legs, covering all the FUs in depths from 20 to 650 metres.</p> <p>A high definition camera system is towed over the sea bed for 10 minutes travelling approx. 200m at 0.8 knots on a purpose built sledge. The UWTV survey follows survey protocols available <a title="Link to survey protocols" href="https://doi.org/10.17895/ices.pub.8014">here</a> agreed by International Council for the Exploration of the Sea (ICES) Working Group on <em>Nephrops </em>surveys (WGNEPS). </p> <p>As part of the iMagine project a selection of images from the Underwater TV survey Functional Units were annotated with bounding boxes and labels in YOLOv8 format to train an YOLOv8 Object Detection Models. The training dataset is saved in YOLOv8 format. It is intended to train a YOLOv8 Nephrrops burrow object detection model to assess the utility of an Object Detection model is assisting Prawn Survey work in the semi automated annotation of prawn burrow imagery.</p> </div> </div> </div> </div>
Simplified Object Detection for Manufacturing: Introducing a Low-Resolution Dataset
<p>This dataset was published with the dataset descriptor "Simplified Object Detection for Manufacturing: Introducing a Low-Resolution Dataset".</p> <p>ACKNOWLEDGEMENTS</p> <p>The project ”ZUKIPRO” is funded as part of the ”Future Centers” program by the Federal<br>Ministry of Labour and Social Affairs and the European Union through the European Social<br>Fund Plus (ESF Plus).Roles and Contributions.</p>
Small Object Aerial Person Detection Dataset
<p><strong>Small Object Aerial Person Detection Dataset:</strong></p> <p>The aerial dataset publication comprises a collection of frames captured from unmanned aerial vehicles (UAVs) during flights over the University of Cyprus campus and Civil Defense exercises. The dataset is primarily intended for people detection, with a focus on detecting small objects due to the top-view perspective of the images. The dataset includes annotations generated in popular formats such as YOLO, COCO, and VOC, making it highly versatile and accessible for a wide range of applications. Overall, this aerial dataset publication represents a valuable resource for researchers and practitioners working in the field of computer vision and machine learning, particularly those focused on people detection and related applications.</p> <p> </p> <table> <tbody> <tr> <td>Subset</td> <td>Images</td> <td>People</td> </tr> <tr> <td>Training</td> <td>2092</td> <td>40687</td> </tr> <tr> <td>Validation</td> <td>523</td> <td>10589</td> </tr> <tr> <td>Testing</td> <td>521</td> <td>10432</td> </tr> </tbody> </table> <p> </p> <p>It is advised to further enhance the dataset so that random augmentations are probabilistically applied to each image prior to adding it to the batch for training. Specifically, there are a number of possible transformations such as geometric (rotations, translations, horizontal axis mirroring, cropping, and zooming), as well as image manipulations (illumination changes, color shifting, blurring, sharpening, and shadowing).</p>
Data for publication 'Detection of Artificial Seed-like Objects from UAV Imagery'
<p>This resource contains the datasets supporting the model development as published in the article 'Detection of Artificial Seed-like Objects from UAV Imagery' (https://doi.org/10.3390/rs15061637).</p> <p>In the last two decades, unmanned aerial vehicle (UAV) technology has been widely utilized as an aerial survey method. Recently, a unique system of self-deployable and biodegradable microrobots akin to winged achene seeds was introduced to monitor environmental parameters in the air above the soil interface, which requires geo-localization. This research focuses on detecting these artificial seed-like objects from UAV RGB images in real-time scenarios, employing the object detection algorithm YOLO (You Only Look Once). Three environmental parameters, namely, daylight condition, background type, and flying altitude, were investigated to encompass varying data acquisition situations and their influence on detection accuracy. Artificial seeds were detected using four variants of the YOLO version 5 (YOLOv5) algorithm, which were compared in terms of accuracy and speed. The most accurate model variant was used in combination with slice-aided hyper inference (SAHI) on full resolution images to evaluate the model’s performance. It was found that the YOLOv5n variant had the highest accuracy and fastest inference speed. After model training, the best conditions for detecting artificial seed-like objects were found at a flight altitude of 4 m, on an overcast day, and against a concrete background, obtaining accuracies of 0.91, 0.90, and 0.99, respectively. YOLOv5n outperformed the other models by achieving a mAP0.5 score of 84.6% on the validation set and 83.2% on the test set. This study can be used as a baseline for detecting seed-like objects under the tested conditions in future studies.</p>
Outputs of the Jupyter Notebook - Detecting floating objects using Deep Learning and Sentinel-2 imagery
<p>The dataset contains the outputs of the notebook "Detecting floating objects using Deep Learning and Sentinel-2 imagery" published in the ocean modelling section of The Environmental Data Science Book.</p> <p><strong>Contributions</strong></p> <p><em>Notebook</em></p> <ul> <li> <p>Jamila Mifdal (author), European Space Agency Φ-lab, <a href="https://github.com/jmifdal">@jmifdal</a></p> </li> <li> <p>Raquel Carmo (author), European Space Agency Φ-lab, <a href="https://github.com/raquelcarmo">@raquelcarmo</a></p> </li> <li> <p>Alejandro Coca-Castro (reviewer), The Alan Turing Institute, <a href="https://github.com/acocac">@acocac</a></p> </li> </ul> <p><em>Modelling codebase</em></p> <ul> <li> <p>Jamila Mifdal (author), European Space Agency Φ-lab, <a href="https://github.com/jmifdal">@jmifdal</a></p> </li> <li> <p>Raquel Carmo (author), European Space Agency Φ-lab, <a href="https://github.com/raquelcarmo">@raquelcarmo</a></p> </li> <li> <p>Marc Rußwurm (author), EPFL-ECEO, <a href="https://github.com/MarcCoru">@marccoru</a></p> </li> </ul>
FG-OVD: Fine-grained Open-Vocabulary Object Detection Benchmark Suite
<p>A collection of annotations for PACO images containing free-form fine-grained textual captions of objects, their parts, and their attributes. It also comprises several sets of negative captions that can be used to test and evaluate the fine-grained recognition ability of open-vocabulary models.</p>
AGS_apple_detection - Apple fruit images dataset for full image object detection
<p>This dataset correspond to full apple tree images (623) annotated for the task of object detection with its corresponding annotations in yolo format saved as txt files. The dataset was divided into test, train and validation<br><br>The data was collected in 2017 on 4 different apple varieties using a Samsung sm-a510F cell phone at two different resolutions: 2448 x 3264 px and 3096 x 4128 px in the orchards of Agroscope located in Wallis, Switzerland. </p>
TensorFlow models for CK object detection
<p>Tarball containing the yolo model for the tensorflow object detection program in CK repositories</p> <p> </p>
The Object Detection for Olfactory References (ODOR) Dataset.
<p>Real-world applications of computer vision in the humanities require algorithms to be robust against artistic abstraction, peripheral objects, and subtle differences between fine-grained target classes. Existing datasets provide instance-level<br> annotations on artworks but are generally biased towards the image centre and limited with regard to detailed object classes. The proposed ODOR dataset fills this gap, offering 38,116 object-level annotations across 4,712 images, spanning an extensive set of 139 fine-grained categories. Conducting a statistical analysis, we showcase challenging dataset properties, such as a detailed set of categories, dense and overlapping objects, and spatial distribution over the whole image canvas. Furthermore, we provide an extensive baseline analysis for object detection models and highlight the challenging properties of the dataset through a set of secondary studies. Inspiring further research on artwork object detection and broader visual cultural heritage studies, the dataset challenges researchers to explore the intersection of object recognition and smell perception.</p> <p><strong>How to use</strong></p> <p>The annotations are provided in COCO JSON format. To represent the two-level hierarchy of the object classes, we make use of the supercategory field in the categories array as defined by COCO. In addition to the object-level annotations, we provide an additional CSV file with image-level metadata, which includes content-related fields, such as Iconclass codes [72 , 73]) or image<br> descriptions, as well as formal annotations, such as artist, license, or creation year. For the sake of license compliance, we do not publish the images directly (although most of the images are public domain). Instead, we provide links to their source collections in the metadata file (meta.csv) and a python script to download the artwork images (download_images.py).</p> <p> </p>
Different spectral sensitivities of ON- and OFF-motion pathways enhance the detection of approaching color objects in Drosophila - Processed Data
<p>Processed data and code for plotting figures for the paper:</p><p>"Different spectral sensitivities of ON- and OFF-motion pathways enhance the detection of approaching color objects in Drosophila", by Kit D. Longden, Edward M. Rogers, Aljoscha Nern, Heather Dionne, Michael B. Reiser.</p><p>Data (compressed results folder) and plotting code (compressed src folder) are MATLAB files (see READ_ME for version information and toolboxes). The Source Data excel file also contains the data plotted in the paper figures.</p>
Next-generation 3D object detection and tracking for self-driving vehicles using object velocity
<p>The synthetic dataset was generated using KITTI-like specifications and annotations format. It is comprised by the training and testing sets, that include KITTI standard folders: label_2, image_2 and calib. Furthermore, there is a velodyne file for each of the following use cases:</p><ul><li>Point cloud 1: (x,y,z, (Float)Radial_Velocity): this point cloud has the relative radial velocity as an additional feature for each point. File: velodyne_radial_velocity;</li><li>Point cloud 2: (x,y,z,(Float)Absolute_Speed): in this point cloud, every point has the absolute speed of the object as the additional feature. File: velodyne_abs_speed;</li><li>Point cloud 3: (x,y,z,(Bool)Is_Moving): the additional feature of this point cloud is a Boolean value that is set to 1.0 if the object is moving; contrariwise, it is set to 0.0 for static objects. File: velodyne_is_moving;</li><li>Point cloud 4: (x,y,z,0): no additional feature information. If desired, requires post-processing to convert to (x,y,z) or changing the toolbox point cloud configuration to not consider the additional feature. File: velodyne_xyz;</li></ul><p>Additionally, the detections generated with the OpenPCDet toolbox and Second-IoU model are provided.</p><p>This work was made as part of a master thesis of Informatics Engineering in the University of Aveiro.</p>
VISIONE Feature Repository for VBS: Multi-Modal Features and Detected Objects from VBSLHE Dataset
<p>This repository contains a diverse set of features extracted from the VBSLHE dataset (laparoscopic gynecology) . These features will be utilized in the VISIONE system [Amato et al. 2023, Amato et al. 2022] in the next editions of the Video Browser Showdown (VBS) competition (<a href="https://www.videobrowsershowdown.org/">https://www.videobrowsershowdown.org/</a>). </p> <p>We used a snapshot of the dataset provided by the Medical University of Vienna and Toronto that can be downloaded using the instructions provided at <a href="https://download-dbis.dmi.unibas.ch/mvk/">https://download-dbis.dmi.unibas.ch/mvk/</a>. It comprises 75 video files. We divided each video into video shots with a maximum duration of 5 seconds.</p> <p>This repository is released under a Creative Commons Attribution license. If you use it in any form for your work, please cite the following paper:</p> <blockquote> <p>@inproceedings{amato2023visione, title={VISIONE at Video Browser Showdown 2023}, author={Amato, Giuseppe and Bolettieri, Paolo and Carrara, Fabio and Falchi, Fabrizio and Gennaro, Claudio and Messina, Nicola and Vadicamo, Lucia and Vairo, Claudio}, booktitle={International Conference on Multimedia Modeling}, pages={615--621}, year={2023}, organization={Springer} } </p> </blockquote> <p> </p> <p>This repository (v2) comprises the following files:</p> <ul> <li><em><strong>msb.tar.gz </strong></em> contains tab-separated files (.tsv) for each video. Each tsv file reports, for each video segment, the timestamp and frame number marking the start/end of the video segment, along with the timestamp of the extracted middle frame and the associated identifier ("id_visione").</li> <li><em><strong>extract-keyframes-from-msb.tar.gz</strong></em> contains a Python script designed to extract the middle frame of each video segment from the MSB files. To run the script successfully, please ensure that you have the original VBSLHE videos available.</li> <li><em><strong>features-aladin.tar.gz†</strong></em><strong> </strong>contains <a href="https://github.com/mesnico/ALADIN">ALADIN</a> [Messina N. et al. 2022] features extracted for all the segment's middle frames.</li> <li><em><strong>features-clip-laion.tar.gz†</strong></em> contains <a href="https://huggingface.co/laion/CLIP-ViT-H-14-laion2B-s32B-b79K">CLIP ViT-H/14 - LAION-2B </a>[Schuhmann et al. 2022] features extracted for all the segment's middle frames.</li> <li><em><strong>features-clip-openai.tar.gz† </strong></em>contains <a href="https://huggingface.co/openai/clip-vit-large-patch14">CLIP ViT-L/14</a> [Radford et al. 2021] features extracted for all the segment's middle frames.</li> <li><em><strong>features-clip2video.tar.gz† </strong></em>contains <a href="https://github.com/CryhanFang/CLIP2Video">CLIP2Video</a> [Fang H. et al. 2021] extracted for all the video segments. <strong> </strong></li> <li><em><strong>objects-frcnn-oiv4.tar.gz* </strong></em>contains the objects detected using <a href="http://tfhub.dev/google/faster_rcnn/openimages_v4/inception_resnet_v2/1">Faster R-CNN+Inception ResNet</a> (trained on the Open Images V4 [Kuznetsova et al. 2020]).</li> <li><em><strong>objects-mrcnn-lvis.tar.gz*</strong></em> contains the objects detected using Mask R-CNN [He et al. 2017] (trained on LVIS).</li> <li><em><strong>objects-vfnet64-coco.tar.gz*</strong></em> contains the objects detected using VfNet [Zhang et al. 2021] (trained on COCO dataset).</li> </ul> <p>*Please be sure to use the <strong>v2 version </strong>of this repository, since v1 feature files may contain inconsistencies that have now been corrected</p> <p><em><strong>*Note on the object annotations:</strong></em> Within an object archive, there is a jsonl file for each video, where each row contains a record of a video segment (the <em>"_id"</em> corresponds to the <em>"id_visione"</em> used in the msb.tar.gz) . Additionally, there are three arrays representing the objects detected, the corresponding scores, and the bounding boxes. The format of these arrays is as follows:</p> <ul> <li><em>"object_class_names"</em>: vector with the class name of each detected object.</li> <li><em>"object_scores"</em>: scores corresponding to each detected object.</li> <li><em>"object_boxes_yxyx"</em>: bounding boxes of the detected objects in the format <em>(ymin, xmin, ymax, xmax).</em></li> </ul> <p> </p> <p><em><strong>†Note on the cross-modal features: </strong></em>The extracted multi-modal features (ALADIN, CLIPs, CLIP2Video) enable internal searches within the VBSLHE dataset using the query-by-image approach (features can be compared with the dot product). However, to perform searches based on free text, the text needs to be transformed into the joint embedding space according to the specific network being used (see links above). Please be aware that t<strong>he service for transforming text into features is not provided within this repository and should be developed independently using the original feature repositories linked above.</strong></p> <p>We have plans to release the code in the future, allowing the reproduction of the VISIONE system, including the instantiation of all the services to transform text into cross-modal features. However, this work is still in progress, and the code is not currently available.</p> <p> </p> <p><strong>References:</strong></p> <p>[Amato et al. 2023] Amato, G.et al., 2023, January. VISIONE at Video Browser Showdown 2023. In International Conference on Multimedia Modeling (pp. 615-621). Cham: Springer International Publishing.</p> <p>[Amato et al. 2022] Amato, G. et al. (2022). VISIONE at Video Browser Showdown 2022. In: , et al. MultiMedia Modeling. MMM 2022. Lecture Notes in Computer Science, vol 13142. Springer, Cham. </p> <p>[Fang H. et al. 2021] Fang H. et al., 2021. Clip2video: Mastering video-text retrieval via image clip. arXiv preprint arXiv:2106.11097.</p> <p>[He et al. 2017] He, K., Gkioxari, G., Dollár, P. and Girshick, R., 2017. Mask r-cnn. In Proceedings of the IEEE international conference on computer vision (pp. 2961-2969).</p> <p>[Kuznetsova et al. 2020] Kuznetsova, A., Rom, H., Alldrin, N., Uijlings, J., Krasin, I., Pont-Tuset, J., Kamali, S., Popov, S., Malloci, M., Kolesnikov, A. and Duerig, T., 2020. The open images dataset v4. International Journal of Computer Vision, 128(7), pp.1956-1981.</p> <p>[Lin et al. 2014] Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P. and Zitnick, C.L., 2014, September. Microsoft coco: Common objects in context. In European conference on computer vision (pp. 740-755). Springer, Cham.</p> <p>[Messina et al. 2022] Messina N. et al., 2022, September. Aladin: distilling fine-grained alignment scores for efficient image-text matching and retrieval. In Proceedings of the 19th International Conference on Content-based Multimedia Indexing (pp. 64-70).</p> <p>[Radford et al. 2021] Radford A. et al., 2021, July. Learning transferable visual models from natural language supervision. In International conference on machine learning (pp. 8748-8763). PMLR.</p> <p>[Schuhmann et al. 2022] Schuhmann C. et al., 2022. Laion-5b: An open large-scale dataset for training next generation image-text models. Advances in Neural Information Processing Systems, 35, pp.25278-25294.</p> <p>[Zhang et al. 2021] Zhang, H., Wang, Y., Dayoub, F. and Sunderhauf, N., 2021. Varifocalnet: An iou-aware dense object detector. In Proceedings of the IEEE/CV</p>
VISIONE Feature Repository for VBS: Multi-Modal Features and Detected Objects from MVK Dataset
<p>This repository contains a diverse set of features extracted from the marine video (underwater) dataset (MVK) . These features were utilized in the VISIONE system [Amato et al. 2023, Amato et al. 2022] during the latest editions of the Video Browser Showdown (VBS) competition (<a href="https://www.videobrowsershowdown.org/">https://www.videobrowsershowdown.org/</a>). </p> <p>We used a snapshot of the MVK dataset from 2023, that can be downloaded using the instructions provided at <a href="https://download-dbis.dmi.unibas.ch/mvk/">https://download-dbis.dmi.unibas.ch/mvk/</a>. It comprises 1,372 video files. We divided each video into 1 second segments. </p> <p>This repository is released under a Creative Commons Attribution license. If you use it in any form for your work, please cite the following paper:</p> <blockquote> <pre>@inproceedings{amato2023visione, title={VISIONE at Video Browser Showdown 2023}, author={Amato, Giuseppe and Bolettieri, Paolo and Carrara, Fabio and Falchi, Fabrizio and Gennaro, Claudio and Messina, Nicola and Vadicamo, Lucia and Vairo, Claudio}, booktitle={International Conference on Multimedia Modeling}, pages={615--621}, year={2023}, organization={Springer} }</pre> </blockquote> <p> </p> <p>This repository comprises the following files:</p> <ul> <li><strong><em>msb.tar.gz </em></strong> contains tab-separated files (.tsv) for each video. Each tsv file reports, for each video segment, the timestamp and frame number marking the start/end of the video segment, along with the timestamp of the extracted middle frame and the associated identifier ("id_visione"). </li> <li><em><strong>extract-keyframes-from-msb.tar.gz</strong></em> contains a Python script designed to extract the middle frame of each video segment from the MSB files. To run the script successfully, please ensure that you have the original MVK videos available.</li> <li><strong><em>features-aladin.tar.gz<sup>†</sup></em> </strong>contains <a href="https://github.com/mesnico/ALADIN">ALADIN</a> [Messina N. et al. 2022] features extracted for all the segment's middle frames. </li> <li><em><strong>features-clip-laion.tar.gz<sup>†</sup></strong></em> contains <a href="https://huggingface.co/laion/CLIP-ViT-H-14-laion2B-s32B-b79K">CLIP ViT-H/14 - LAION-2B </a>[Schuhmann et al. 2022] features extracted for all the segment's middle frames.</li> <li><em><strong>features-clip-openai.tar.gz<sup>†</sup> </strong></em>contains <a href="https://huggingface.co/openai/clip-vit-large-patch14">CLIP ViT-L/14</a> [Radford et al. 2021] features extracted for all the segment's middle frames. </li> <li><em><strong>features-clip2video.tar.gz<sup>†</sup> </strong></em>contains <a href="https://github.com/CryhanFang/CLIP2Video">CLIP2Video</a> [Fang H. et al. 2021] extracted for all the 1s video segments. <strong> </strong></li> <li><em><strong>objects-frcnn-oiv4.tar.gz<sup>*</sup> </strong></em>contains the objects detected using <a href="http://tfhub.dev/google/faster_rcnn/openimages_v4/inception_resnet_v2/1">Faster R-CNN+Inception ResNet</a> (trained on the Open Images V4 [Kuznetsova et al. 2020]). </li> <li><em><strong>objects-mrcnn-lvis.tar.gz<sup>*</sup></strong></em> contains the objects detected using Mask R-CNN [He et al. 2017] (trained on LVIS).</li> <li><em><strong>objects-vfnet64-coco.tar.gz<sup>*</sup></strong></em> contains the objects detected using VfNet [Zhang et al. 2021] (trained on COCO dataset).</li> </ul> <p>*Please be sure to use the <strong>v2 version </strong>of this repository, since v1 feature files may contain inconsistencies that have now been corrected</p> <p><em><strong>*Note on the object annotations:</strong></em> Within an object archive, there is a jsonl file for each video, where each row contains a record of a video segment (the <em>"_id"</em> corresponds to the <em>"id_visione"</em> used in the msb.tar.gz) . Additionally, there are three arrays representing the objects detected, the corresponding scores, and the bounding boxes. The format of these arrays is as follows:</p> <ul> <li><em>"object_class_names"</em>: vector with the class name of each detected object.</li> <li><em>"object_scores"</em>: scores corresponding to each detected object.</li> <li><em>"object_boxes_yxyx"</em>: bounding boxes of the detected objects in the format <em>(ymin, xmin, ymax, xmax).</em></li> </ul> <p> </p> <p><em><strong><sup>†</sup>Note on the cross-modal features: </strong></em>The extracted multi-modal features (ALADIN, CLIPs, CLIP2Video) enable internal searches within the MVK dataset using the query-by-image approach (features can be compared with the dot product). However, to perform searches based on free text, the text needs to be transformed into the joint embedding space according to the specific network being used (see links above). Please be aware that t<strong>he service for transforming text into features is not provided within this repository and should be developed independently using the original feature repositories linked above.</strong></p> <p>We have plans to release the code in the future, allowing the reproduction of the VISIONE system, including the instantiation of all the services to transform text into cross-modal features. However, this work is still in progress, and the code is not currently available.</p> <p> </p> <p><strong>References:</strong></p> <p>[Amato et al. 2023] Amato, G.et al., 2023, January. VISIONE at Video Browser Showdown 2023. In International Conference on Multimedia Modeling (pp. 615-621). Cham: Springer International Publishing.</p> <p>[Amato et al. 2022] Amato, G. et al. (2022). VISIONE at Video Browser Showdown 2022. In: , et al. MultiMedia Modeling. MMM 2022. Lecture Notes in Computer Science, vol 13142. Springer, Cham. </p> <p>[Fang H. et al. 2021] Fang H. et al., 2021. Clip2video: Mastering video-text retrieval via image clip. arXiv preprint arXiv:2106.11097.</p> <p>[He et al. 2017] He, K., Gkioxari, G., Dollár, P. and Girshick, R., 2017. Mask r-cnn. In Proceedings of the IEEE international conference on computer vision (pp. 2961-2969).</p> <p>[Kuznetsova et al. 2020] Kuznetsova, A., Rom, H., Alldrin, N., Uijlings, J., Krasin, I., Pont-Tuset, J., Kamali, S., Popov, S., Malloci, M., Kolesnikov, A. and Duerig, T., 2020. The open images dataset v4. International Journal of Computer Vision, 128(7), pp.1956-1981.</p> <p>[Lin et al. 2014] Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P. and Zitnick, C.L., 2014, September. Microsoft coco: Common objects in context. In European conference on computer vision (pp. 740-755). Springer, Cham.</p> <p>[Messina et al. 2022] Messina N. et al., 2022, September. Aladin: distilling fine-grained alignment scores for efficient image-text matching and retrieval. In Proceedings of the 19th International Conference on Content-based Multimedia Indexing (pp. 64-70).</p> <p>[Radford et al. 2021] Radford A. et al., 2021, July. Learning transferable visual models from natural language supervision. In International conference on machine learning (pp. 8748-8763). PMLR.</p> <p>[Schuhmann et al. 2022] Schuhmann C. et al., 2022. Laion-5b: An open large-scale dataset for training next generation image-text models. Advances in Neural Information Processing Systems, 35, pp.25278-25294.</p> <p>[Zhang et al. 2021] Zhang, H., Wang, Y., Dayoub, F. and Sunderhauf, N., 2021. Varifocalnet: An iou-aware dense object detector. In Proceedings of the IEEE/CV</p>
RafanoSet: Dataset of raw, manual and automatically annotated Raphanus Raphanistrum weed images for object detection and segmentation in Heterogenous Agriculture Environment
<p>This dataset is a collection of raw and annotated Multispectral (MS) images acquired in a heterogenous agricultural environment with MicaSense RedEdge-M camera. The spectra particularly Green, Blue, Red, Red Edge and Near Infrared (NIR) were acquired at sub-metre level.. <br><br>The MS images were labelled manually using VIA and automatically using Grounding DINO in combination with Segment Anything Model. The segmentation masks obtained using these two annotation techniqes over as well as the source code to perform necessary image processing operations are provided in the repository. The images are focussed over Horseradish (Raphanus Raphanistrum) infestations in Triticum Aestivum (wheat) crops.</p> <p>The nomenclature of sequecncing and naming images and annotations has been in this format: IMG_<scene number>_<spectral channel number><br><strong>_1</strong>: Blue<br><strong>_2</strong>: Green<br><strong>_3</strong>: Red<br><strong>_4</strong>: Near Infrared<br><strong>_5</strong>: RedEdge<br><br>Example: An image name <strong>IMG_0200_3 </strong>represents the scene number<strong> 200</strong> in <strong>Red channel</strong></p> <p>This dataset 'RafanoSet'is categorized in 6 directories namely 'Raw Images', 'Manual Annotations', 'Automated Annotations', 'Binary Masks - Manual', 'Binary Masks - Automated' and 'Codes'. The sub-directory 'Raw Images' consists of manually acquired 85 images in .PNG format. over 17 different scenes. The sub-directory 'Manual Annotations' consists of annotation file 'region_data' in COCO segmentation format. The sub-directory 'Automated Annotations' consists of 80 automatically annotated images in .JPG format and 80 .XML files in Pascal VOC annotation format.</p> <p>The scientific framework of image acquisition and annotations are explained in the Data in Brief paper which is the course of peer review. This is just a prerequisite to the data article. <br><br>Field experimentation roles:</p> <p>The image acquisition was performed by Mariano Crimaldi, a researcher, on behalf of Department of Agriculture and the hosting institution University of Naples Federico II, Italy.</p> <p>Shubham Rana has been the curator and analyst for the data under the supervision of his PhD supervisor Prof. Salvatore Gerbino. They are affiliated with Department of Engineering, University of Campania 'Luigi Vanvitelli'. </p> <p>Domenico Barretta, Department of Engineering has been associated in consulting and brainstorming role particularly with data validation, annotation management and litmus testing of the datasets.</p>
DeepBacs – Escherichia coli antibiotic phenotyping object detection dataset and YOLOv2 model
<p>Training and test images of <em>E. coli</em> cells treated with different antibiotics for antibiotic phenotyping using YOLOv2 object detection.</p> <p>Additional information can be found on this <a href="https://github.com/HenriquesLab/DeepBacs/wiki">github wiki</a>.</p> <p>Example images show predictions of drug-treated <em>E. coli</em> cells.</p> <p> </p> <p><strong>Training and test dataset</strong></p> <p><strong>Data type</strong>: Paired microscopy images (confocal fluorescence) and manual annotations</p> <p><strong>Microscopy data type</strong>: Confocal fluorescence images of fixed <em>E. coli</em> cells stained for membrane (Nile Red) and DNA (DAPI) paired with annotations in PASCAL VOC format</p> <p><strong>Microscope</strong>: Zeiss LSM710 confocal microscope with a Plan-Apo 63x oil objective (1.4 NA)</p> <p><strong>Cell type</strong>: Chemically fixed <em>E. coli</em> NO34 cells (MreB-sfGFPsw, kindly provided by Zemer Gitai) (untreated or drug-treated);</p> <p><strong>File format</strong>: .png (RGB)</p> <p><strong>Image size</strong>: 400 x 400 px² (Pixel size: 84 nm)</p> <p> </p> <p><strong>YOLOv2 model</strong></p> <p>The YOLOv2 model was generated using the ZeroCostDL4Mic platform (Chamier et al., 2021). It was trained from scratch for 97 epochs on 153 manually annotated images (image dimensions: (400, 400, 3)) with a batch size of 16 and a custom loss function combining MSE and crossentropy losses, using the YOLOv2 ZeroCostDL4Mic notebook (v 1.12) (von Chamier & Laine et al., 2020). Key python packages used include tensorflow (v0.1.12), Keras (v 2.3.1), numpy (v 1.19.5), cuda (v 10.1.243). The training was accelerated using a Tesla P100GPU and data was augmented by a factor of 8 using rotation and flipping.</p> <p>The model weights can be used with the ZeroCostDL4Mic YOLOv2 notebook.</p> <p> </p> <p><strong>Author(s)</strong>: Christoph Spahn<sup>1,2</sup>, Mike Heilemann<sup>1,3</sup></p> <p><strong>Contact email</strong>: christoph.spahn@mpi-marburg.mpg.de</p> <p> </p> <p><strong>Affiliation(s)</strong>: </p> <p>1) Institute of Physical and Theoretical Chemistry, Max-von-Laue Str. 7, Goethe-University Frankfurt, 60439 Frankfurt, Germany</p> <p>2) ORCID: 0000-0001-9886-2263 </p> <p>3) ORCID: 0000-0002-9821-3578</p>
DeepBacs – Escherichia coli growth stage object detection dataset and YOLOv2 model
<p>Training and test images of E. coli cells for object detection and classification using YOLOv2, as well as a trained YOLOv2 model.</p> <p>Additional information can be found on this <a href="https://github.com/HenriquesLab/DeepBacs/wiki">github wiki</a>.</p> <p>The example shows a bright field image of live <em>E. coli</em> cells and the respective annotation for specific growth stages.</p> <p> </p> <p><strong>Training and test dataset</strong></p> <p><strong>Data type</strong>: Paired microscopy images (bright field) and annotations in PASCAL VOC format</p> <p><strong>Microscopy data type</strong>: 2D bright field images recorded at 1 min interval</p> <p><strong>Microscope</strong>: Nikon Eclipse Ti-E equipped with an Apo TIRF 1.49NA 100x oil immersion objective</p> <p><strong>Cell type</strong>: <em>E. coli</em> MG1655 wild type strain (CGSC #6300).</p> <p><strong>File format</strong>: .png (8-bit)</p> <p><strong>Image size</strong>: 256 x 256 px² (158 nm / pixel), 100/15 individual frames (training/test dataset)</p> <p>1024 x 1024 px² (79 nm / pixel), 9 regions of interest with 80 frames @ 1 min time interval (live-cell time series)</p> <p><strong>Image preprocessing</strong>: Raw images were recorded in 16-bit mode (image size 512x512 px² @ 158 nm/px). 256 x 256 px² patches were extracted from individual frames and converted into 8-bit .png images after adjusting brightness and contrast. Annotation was performed online using <em>LabelImg </em>(https://github.com/tzutalin/labelImg).</p> <p> </p> <p><strong>YOLOv2 model</strong></p> <p>The YOLOv2 model was generated using the ZeroCostDL4Mic platform (Chamier et al., 2021). It was trained from scratch for 97 epochs on 100 manually annotated images (image dimensions: (256, 256)) with a batch size of 8 and a custom loss function combining MSE and crossentropy losses, using the YOLOv2 ZeroCostDL4Mic notebook (v 1.12.1). Key python packages used include tensorflow (v 0.1.12), Keras (v 2.3.1), numpy (v 1.19.5), cuda (v 11.0.221). The training was accelerated using a Tesla T4 GPU and data were augmented by a factor of 4 using flipping and rotation.</p> <p>The model weights can be used with the ZeroCostDL4Mic YOLOv2 notebook.</p> <p> </p> <p><strong>Author(s)</strong>: Christoph Spahn<sup>1,2</sup>, Mike Heilemann<sup>1,3</sup></p> <p><strong>Contact email</strong>: christoph.spahn@mpi-marburg.mpg.de</p> <p> </p> <p><strong>Affiliation(s)</strong>: </p> <p>1) Institute of Physical and Theoretical Chemistry, Max-von-Laue Str. 7, Goethe-University Frankfurt, 60439 Frankfurt, Germany</p> <p>2) ORCID: 0000-0001-9886-2263 </p> <p>3) ORCID: 0000-0002-9821-3578</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.