Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
56
datasets available to search
ShareScore release 0.7.1
Dataset results
56 results for “image recognition”
Generating Physically Sound Training Data for Image Recognition of Additively Manufactured Parts Data and Scripts
<p>The repository contains the data corresponding to the Paper "Generating Physically Sound Training Data for Image Recognition of Additively Manufactured Parts".</p> <p>Random30, Random50, Random100, Similiar10, Similar30 and Similar50.zip contain the data sets (obj Files).</p> <p>R30_physical_images.zip and sim50_physical_images.zip contain the photos made from the physical components which are used for the evaluation.</p>
Historical German Children's Playbooks - 6 Digitized Books with Images, OCR-Fulltext, and Named Entity Recognition
<p>The dataset consists of 6 digitized books with 1750 images and OCR-fulltext.</p> <p>Additionally, named entity recognition has been carried out on basis of flair's de-ner model, see https://github.com/flairNLP for details.</p>
Image recognition based on deep learning in Haemonchus contortus motility assays
<p>The repository contains the data associated with the paper `Image recognition based on deep learning in Haemonchus contortus motility assays`. The following folders form part of the repository:</p> <p>- Annotation Data - annotated microscope images used for the training of the Mask R-CNN model. The data are divided into `train` and `val`<br> - Mask R-CNN - contains the trained weights for the Mask R-CNN model<br> - Motility Output - output of the 3 compared algorithms (Wiggle Index, WF-NTP and Mask R-CNN).<br> - Motility Videos - input videos used for the motility detection. The naming convetno is `XXXzYYY.avi`, where `XXX` denotes the motility group and `YYY` the sequence number for the video within a given group</p>
Interpretable Geotechnical Artificial Intelligence (XGeoT-AI) Application to Demystify Image Recognition of Soil Cracks [Datasets]
<p>Here is the test data for the paper "Interpretable Geoscience Artificial Intelligence (XGeoS-AI): Application to Demystify Image Recognition".</p>
Confocal microscopy images for: Surface remodeling and inversion of cell-matrix interactions underlie community recognition and dispersal in Vibrio cholerae biofilms
Open the record for dataset details and reuse information.
Training images for physical stress image face recognition
<p>A proof-of-concept dataset for the initial training for physical stress recognition based on patient's faces. The dataset consists on multiple frames of a training session containing 11 different individuals.</p>
Dataset containing images for training and testing deep learning image recognition models
<p>The dataset contains a training and a test subsets of images. Each image belongs to one out of 10 categories of animals</p>
The Bangladesh Road Traffic Sign Dataset in Real-World Images for Traffic Sign Recognition
Open the record for dataset details and reuse information.
Animal Recognition Using Methods Of Fine-Grained Visual Analysis - Kashtanka Pets (400 Hand-labelled Images - Cats & Dogs, Single Folder)
<p>400 images (200 cats, 200 dogs) hand-labelled by Maria E. with head and body bounding box labels in YOLOv5 format. Images are in a single folder, no separate folders for cats and dogs.</p>
Animal Recognition Using Methods Of Fine-Grained Visual Analysis - Kashtanka Pets (200 Hand-labelled Images, Cats and Dogs, Separate Folders)
<p>400 images (200 cats, 200 dogs) hand-labelled by Maria E. with head and body bounding box labels in YOLOv5 format. Images for cats, for dogs are in a separate folders.</p>
DAVE - Images from Trave for recognition or segmentation
<p>The dataset contains images from a measurement run on the River Trave, with selected objects blurred to ensure privacy. Each image is accompanied by a JSON file that provides the coordinates of the blurred bounding box and the object category.</p> <p>Image data was recorded using a Hik Vision DS-2CD2T47G2-LSU/SL camera.</p> <p> </p> <div> <div> <div> <p>This publication is a result of the research of the Center of Excellence CoSA and funded by the Federal Ministry for Digital and Transport of the Federal Republic of Germany (Id 19F2225C, DAVE).</p> <p> </p> <p>Project website: https://www.th-luebeck.de/cosa/projekt/dave/</p> </div> </div> </div>
PLANT SPECIES RECOGNITION USING LEAF IMAGES AND CONVOLUTIONAL NEURAL NETWORKS (CAAR dataset, version 1)
<p>The CAAR dataset contains leaf images from plants obtained from the Arboreal Collection at Augusto Ribas Agricultural College (CAAR/UEPG). This plant collection is situated in Augusto Ribas College, located at the Ponta Grossa State University, Ponta Grossa, Paraná, Brasil. The images were taken using a smartphone camera with a resolution of 1659 x 2658 pixels and 24 bits of color depth. For each plant, images samples were collected using a white paper sheet background, varying the leaf orientation. The<br>number of plant species used to build the dataset is equal to 35. Data augmentation, using rotation and zoom, were used to increase the data size from 730 to 1986 images in this dataset.</p>
Performance of chemical structure string representations for chemical image recognition using transformers dataset
<p>The datasets contain string representations used for DECIMER short communication paper.</p> <p><strong>ChEMBL dataset:</strong></p> <p>Train and test datasets downloaded from ChEMBL and curated. Contains data with and without stereochemistry. Separated as Canonical and Isomeric.</p> <p>String representations contain SMILES, DeepSMILES, SELFIES and InChIs.</p> <ul> <li>Train dataset: 1.5 Mio molecules</li> <li>Test dataset: ~100K molecules</li> </ul> <p><strong>Pubchem dataset:</strong></p> <p>Train and test datasets downloaded from PubChem and curated. Contains data with and without stereochemistry. Separated as Canonical and Isomeric.</p> <p>String representations contain SMILES, DeepSMILES and SELFIES.</p> <ul> <li>Train dataset: 3 Mio molecules</li> <li>Test dataset: 250K molecules</li> </ul>
The raw images of Laser Confocal Microscopy experiments in the manuscript: Inert Pepper aptamer-mediated endogenous mRNA recognition and imaging in living cells
<p>The <strong>original imaging data</strong> folder contains the raw images of Laser confocal microscopy experiments in the manuscript: Inert Pepper aptamer-mediated endogenous mRNA recognition and imaging in living cells. <a href="https://doi.org/10.1093/nar/gkac368">https://doi.org/10.1093/nar/gkac368</a> </p>
Keypoints Method for Recognition of Ship Wake Components in Sentinel-2 Images by Deep Learning
<p>The dataset used in the study consists of imagery capturing ship wake patterns. It is a manually curated dataset specifically created for the purpose of training and evaluating the wake component detection model. The dataset contains a collection of image chips, each focusing on a specific ship wake instance.</p> <p>The imagery in the dataset is acquired from satellite sensors, specifically on Sentinel-2 satellite imagery. Sentinel-2 provides multispectral data with high spatial resolution, allowing for detailed analysis of ship wake patterns. The dataset includes images captured on B8 spectral band, enabling the exploration of the wake detection model's performance under various spectral conditions. These images have been pre-processed (by scaling+CLAHE) to highlight ocean surface features.</p> <p>Each image chip in the dataset is annotated with keypoint locations representing specific wake components, such as the ship wake vertex, the ending of the turbulent wake, and the ending of Kelvin arms. These annotations serve as ground truth labels for training and evaluating the wake component detection model. </p> <p>Additionally, the dataset includes samples with variations in environmental conditions, such as different sea states, lighting conditions, and wake complexities. This variability allows for a comprehensive evaluation of the model's generalization capability and robustness across diverse scenarios.</p>
Artificial Intelligence in Image Recognition of Pouchoscopies in Patients With Restorative Proctocolectomy
ClinicalTrials.gov study NCT04864587. IPD Sharing: NO. Countries: 1. Publications: 1.
Optimization of Fetal Biometry With 3D Ultrasound and Image Recognition
ClinicalTrials.gov study NCT03812471. IPD Sharing: NO. Countries: 1. Publications: 1.
GAN Generated Images for Facial Expression Recognition systems
<p>Most facial expression recognition (FER) systems rely on machine learning approaches that require large databases (DBs) for effective training. As these are not easily available, a good solution is to augment the DBs with appropriate techniques, which are typically based on either geometric transformation or deep learning based technologies (e.g., Generative Adversarial Networks (GANs)). Whereas the first category of techniques has been fairly adopted in the past, studies that use GAN-based techniques are limited for FER systems. To advance in this respect, we evaluate the impact of the GAN techniques by creating a new DB containing the generated synthetic images. </p> <p>The face images contained in the KDEF DB serve as the basis for creating novel synthetic images by combining the facial features of two images (i.e., Candie Kung and Cristina Saralegui) selected from the YouTube-Faces DB. The novel images differ from each other, in particular concerning the eyes, the nose, and the mouth, whose characteristics are taken from the Candie and Cristina images.</p> <p>The total number of novel synthetic images generated with the GAN is 980 (70 individuals from KDEF DB x 7 emotions x 2 subjects from YouTube-Faces DB).</p> <p>The zip file "GAN_KDEF_Candie" contains the 490 images generated by combining the KDEF images with the Candie Kung image. The zip file "GAN_KDEF_Cristina" contains the 490 images generated by combining the KDEF images with the Cristina Saralegui image. The used image IDs are the same used for the KDEF DB. The synthetic generated images have a resolution of 562x762 pixels.</p> <p> </p> <p><strong>If you make use of this dataset, please consider citing the following publication:</strong></p> <p>Porcu, S., Floris, A., & Atzori, L. (2020). Evaluation of Data Augmentation Techniques for Facial Expression Recognition Systems. Electronics, 9, 1892, doi: 10.3390/electronics9111892, url: https://www.mdpi.com/2079-9292/9/11/1892.</p> <p>BibTex format:</p> <p>@article{porcu2020evaluation, title={Evaluation of Data Augmentation Techniques for Facial Expression Recognition Systems}, author={Porcu, Simone and Floris, Alessandro and Atzori, Luigi}, journal={Electronics}, volume={9}, pages={108781}, year={2020}, number = {11}, article-number = {1892}, publisher={MDPI}, doi={10.3390/electronics9111892} }</p> <p> </p>
Figure 2 from: Greeff M, Caspers M, Kalkman V, Willemse L, Sunderland BD, Bánki O, Hogeweg L (2022) Sharing taxonomic expertise between natural history collections using image recognition. Research Ideas and Outcomes 8: e79187. https://doi.org/10.3897/rio.8.e79187
Figure 2 In the Central Library of Datasets, natural history collection staff will find correctly identified images of their target organisms and download the data for training of an individually customized classifier (photos: Lepidoptera by Entomological Collection of ETH Zürich; Orthoptera by Naturalis Biodiversity Center; Brassicaceae by United Herbaria Z+ZT, ZT-00164967, ZT-00167494, ZT-00171530, CC BY-SA 4.0). The current figure shows a mock-up.
Figure 3 from: Greeff M, Caspers M, Kalkman V, Willemse L, Sunderland BD, Bánki O, Hogeweg L (2022) Sharing taxonomic expertise between natural history collections using image recognition. Research Ideas and Outcomes 8: e79187. https://doi.org/10.3897/rio.8.e79187
Figure 3 Sharing of taxonomic knowledge between institutes. (1) Each algorithm contains two basic components: the feature extractor and the classifier. (2) The Central Library of Datasets allows the user to browse through all available images of collection objects; (3) based on all available images, a regularly updated central feature extractor is created and published; (4) custom made algorithms can relatively easily be created by building a classifier based on a selection of taxa from the central library and combining this with the central feature extractor; (5) newly created algorithms together with their metadata (probability & information on content) are published through a web service in the Central Library of Algorithms (6) and can be used through the Identification web services (API) either for batch processing of images or through a mobile app. Models can be easily extended by other institutions by combining data sources (7).
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.