Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

107

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

107 results for “Computer vision”

Learn how ShareScore rates datasets ↗
zenodo28/100

Example computer vision classification training data derived from British Library 19th Century Books Image collection

<p>Example computer vision classification training data derived from British Library 19th Century Books Image collection</p> <p>This dataset provides training data for image classification for use in a computer vision workshop. The images are derived from &#39;<a href="https://doi.org/10.21250/db17">Digitised Books - Images identified as Embellishments. c. 1510 - c. 1900. JPG&#39;</a>&nbsp;from the year &#39;1839&#39;.</p> <p>Currently, included are four folders containing a variety of images derived from the BL books corpus.</p> <ul> <li>&#39;cv_workshop_exercise_data&#39; include images of: &#39;building&#39;, &#39;people&#39;, &#39;coat of arms&#39;</li> <li>&#39;humancats&#39; contains images of humans and images of cats</li> </ul> <p>The &#39;fashion&#39; and &#39;portraits&#39; folders both contain images of people organised into &#39;female&#39; and &#39;male&#39;. These labels were annotated by a single annotator and these categories may themselves not be meaningful. They are included in the workshop data as a point of discussion about how we should label data both in general and when working with historical data.&nbsp;</p> <p>This data is intended primarily as an educational resource.</p>

openother-pdFeb 2020View details →
zenodo28/100

Transfer Learning for leveraging computer vision in infrastructure maintenance [extracted features]

<p>Dataset containing features extracted from images taken on single case of&nbsp;infrastructure facility for the purpose of training Transfer Learned CNN classifier. It is meant to be used with KrakN framework (https://github.com/MatZar01/KrakN), published with the research paper.</p>

opencc-by-4.0Apr 2020View details →
zenodo28/100

Date Fruit Detection Dataset for Computer Vision-Based Automatic Harvesting.

<p>The "Date Fruit Detection Dataset for Computer Vision-Based Automatic Harvesting" is a collection of videos and images showcasing date fruits from four Moroccan varieties, namely Majhoul, Boufagous, Bouisthami, and Khoult. This dataset is specifically designed to detect and classify date fruits, with the primary goal being the automation of the harvesting process.</p><p>All the images in this dataset were captured in two orchards located in Morocco, with the first orchard situated in the southeast of Errachidia and the second in Tismoumine, Alnif, Tinghir. These images were taken under various natural conditions, encompassing differing lighting, contrast, shadows, and instances where the dates were concealed by bags or hidden beneath palm leaves.</p><p>The dataset was meticulously compiled over the period spanning from June to September 2022, ensuring comprehensive coverage of all four maturity stages of date fruits, which include immature, khalal, rutab, and tamer. The dataset is intended for both object detection and classification purposes, and it includes a YOLO annotation txt file for each image. These annotations have been tailored to precisely recognize not only the date fruit but also to distinguish the specific variety and its maturity stage.</p>

openDec 2022View details →
dryad28/100

Data from: Investigating human repeatability of a computer vision based task to identify meristems on a potato plant (Solanum tuberosum)

<p>Labelled training data in artificial intelligence (AI) is used to teach so-called 'supervised learning models'. However, such data may contain error or bias, which can impact model prediction accuracy. Thus, obtaining accurate training data is of high importance. In applications of AI, such as in classification and detection problems, raw training data is not always made available in published research. Likewise, the process of obtaining labelled data is not always documented well enough to enable reproducibility. This training data set captures a repeatability exercise in AI training data collection for a task that is difficult for humans to perform, delineating a bounding box in a two-dimensional image of a growing apical meristem in potato plants.</p>

opencc-zeroJan 2022View details →
zenodo28/100

Marked crosswalks in US transit-oriented station areas, 2007–2020: A computer vision approach using street view imagery

<p>This is the supplementary dataset for the article &ldquo;Marked crosswalks in US transit-oriented station areas, 2007&ndash;2020: A computer vision approach using street view imagery&rdquo; (<a href="https://journals.sagepub.com/doi/10.1177/23998083221112157">https://journals.sagepub.com/doi/10.1177/23998083221112157</a>) published on Environment and Planning B: Urban Analytics and City Science. The dataset contains comprehensive information of the presence of marked crosswalks (i.e., parallel-line and high-visibility) at each street intersection within a 250-meter buffer of all US TOD stations for each consecutive year between 2007 and 2020. We describe each field of the dataset below. Most of the column station variable names are taken directly from the National TOD Database from the <a href="https://toddata.cnt.org/">Center for Transit-Oriented Development</a>.&nbsp;</p> <ul> <li>intersecton_id: numeric, unique identifier of each intersection.</li> <li>tod_id: numeric, unique identifier of the associated TOD station. We retain all records of intersections that are associated with multiple TOD stations.&nbsp;</li> <li>year: numeric, one year between 2007 and 2020.&nbsp;</li> <li>plain: numeric, the number of presence of parallel-line crosswalks (bounding boxes) detected among the four 90-degree images at each intersection in that particular year. Zero for the years where GSV are not available.&nbsp;</li> <li>zebra: numeric, the number of presence of high-visibility crosswalks (bounding boxes) detected among the four 90-degree images at each intersection in that particular year. Zero for the years where GSV are not available.&nbsp;</li> <li>impute: binary, 1 if imputed; 0 otherwise.&nbsp;</li> <li>lon_intersection: longitude of each intersection obtained from OSM.&nbsp;</li> <li>lat_intersection: latitude of each intersection obtained from OSM.&nbsp;</li> <li>distance_intersection_to_tod: distance of each intersection from TOD station by meters.&nbsp;</li> <li>Agency: operating transit agency of the associated TOD station.&nbsp;</li> <li>Lines (s): transit lines of the associated TOD station.&nbsp;</li> <li>Station Name: station name of the associated TOD station.&nbsp;</li> <li>Year Opened: the year when the associated station opened.&nbsp;</li> <li>lat_tod: latitude of the associated TOD station obtained from the National TOD Database.&nbsp;</li> <li>lon_tod: longitude of the associated TOD station obtained from the Nation TOD Database.&nbsp;</li> <li>Both: numeric, the total number of marked (parallel-line + high-visibility) crosswalks detected among the four 90-degree images at each intersection in that particular year.&nbsp;</li> <li>plain_b: binary, 1 if presence of parallel-line crosswalk; 0 if otherwise.&nbsp;</li> <li>zebra_b: binary, 1 if presence of high-visibility crosswalk; 0 if otherwise.&nbsp;</li> <li>both_b: binary, 1 if presence of marked crosswalk; 0 if otherwise.&nbsp;</li> </ul>

opencc-by-4.0Dec 2021View details →
zenodo28/100

Computer Vision-Based Algorithm for Precise Defect Detection and Classification in Photovoltaic Modules

Open the record for dataset details and reuse information.

opencc-by-4.0Jul 2024View details →
zenodo28/100

Filming the sound: Anomaly Detection on Audio Tape Recordings using Computer Vision Algorithms

<p>This repository makes available the dataset related to the paper:</p> <p>Zafer &Ccedil;ınar, Alessandro Russo, Matteo Spanio, Niccol&ograve; Pretto, and Sergio Canazza, <em>Filming the Sound: Anomaly Detection on Audio Tape Recordings using Computer Vision Algorithms</em>, IAI4CH, Bozen, 2024.</p> <p>The dataset and the experiment are described in the publication above.</p> <p>This repository contains two main directories (<strong>bold</strong> indicates directory names):</p> <ul> <li><strong>video samples</strong>: the actual videos used in the paper's experiment. This folder contains four subdirectories - 3.75 ips, 7.5 ips, 15 ips, and 30 ips - each representing a different playback speed (in inches per second). Within each subdirectory are several MP4 files, recorded on an A810 Studer open reel recorder, documenting the playback of magnetic audio tapes. The files follow the naming convention &ldquo;Xips (Y).mp4,&rdquo; where&nbsp;<em>X</em> represents the tape playback speed and&nbsp;<em>Y</em> is a serial number identifier for each video.</li> <li><strong>irregularities</strong>: the metadata for each video with timestamp and type of irregularity. The folder includes four CSV files - 3.75.csv, 7.5.csv, 15.csv, and 30.csv - corresponding to the playback speeds of the video samples. Each CSV file provides handmade annotations for its respective videos, with three columns: <ul> <li><em>video_id</em>: name of the video file in the format &ldquo;Xips (Y).mp4,&rdquo; where&nbsp;<em>X</em> is the tape speed and <em>Y</em> is the ID number.</li> <li><em>time_label</em>: timestamp indicating the irregularity, formatted as HH:MM:SS.mls.</li> <li><em>irregularity_type</em>: category of the detected anomaly, which may be one of the following: &ldquo;splice,&rdquo; &ldquo;shadow,&rdquo; &ldquo;end-of-tape,&rdquo; or &ldquo;annotation.&rdquo;</li> </ul> </li> </ul>

restrictedcc-by-4.0Nov 2024View details →
zenodo28/100

NOVA: Rendering Virtual Worlds with Humans for Computer Vision Tasks

<p>Today, the cutting edge of computer vision research greatly depends on the availability of large datasets, which are critical for effectively training and testing new methods. Manually annotating visual data, however, is not only a labor-intensive process but also prone to errors. In this study, we present NOVA, a versatile framework to create realistic-looking 3D rendered worlds containing procedurally generated humans with rich pixel-level ground truth annotations. NOVA can simulate various environmental factors such as weather conditions or different times of day, and bring an exceptionally diverse set of humans to life, each having a distinct body shape, gender and age. To demonstrate NOVA&#39;s capabilities, we generate two synthetic datasets for person tracking. The first one includes 108 sequences, each with different levels of difficulty like tracking in crowded scenes or at nighttime and aims for testing the limits of current state-of-the-art trackers. A second dataset of 97 sequences with normal weather conditions is used to show how our synthetic sequences can be utilized to train and boost the performance of deep-learning based trackers. Our results indicate that the synthetic data generated by NOVA represents a good proxy of the real-world and can be exploited for computer vision tasks.</p>

opencc-by-4.0May 2021View details →
zenodo28/100

DeepLandforms: A Deep Learning Computer Vision toolset applied to a prime use case for mapping planetary skylights - ANc

<p>Initial training dataset for&nbsp;<strong>DeepLandforms: A Deep Learning Computer Vision toolset applied to a prime use case for mapping planetary skylights</strong></p>

openNov 2021View details →
zenodo28/100

P in Identification of morphologically cryptic species with computer vision models: wall lizards (Squamata: Lacertidae: Podarcis) as a case study

P. Ʋirescens

opennotspecifiedDec 2022View details →
zenodo28/100

P in Identification of morphologically cryptic species with computer vision models: wall lizards (Squamata: Lacertidae: Podarcis) as a case study

P. Ʋaucheri

opennotspecifiedDec 2022View details →
zenodo28/100

P in Identification of morphologically cryptic species with computer vision models: wall lizards (Squamata: Lacertidae: Podarcis) as a case study

P. tunesiacus

opennotspecifiedDec 2022View details →
zenodo28/100

P in Identification of morphologically cryptic species with computer vision models: wall lizards (Squamata: Lacertidae: Podarcis) as a case study

P. lusitanicus

opennotspecifiedDec 2022View details →
zenodo28/100

P in Identification of morphologically cryptic species with computer vision models: wall lizards (Squamata: Lacertidae: Podarcis) as a case study

P. liolepis

opennotspecifiedDec 2022View details →
zenodo28/100

P in Identification of morphologically cryptic species with computer vision models: wall lizards (Squamata: Lacertidae: Podarcis) as a case study

P. Ʋirescens

opennotspecifiedDec 2022View details →
zenodo28/100

P in Identification of morphologically cryptic species with computer vision models: wall lizards (Squamata: Lacertidae: Podarcis) as a case study

P. Ʋaucheri

opennotspecifiedDec 2022View details →
zenodo28/100

P in Identification of morphologically cryptic species with computer vision models: wall lizards (Squamata: Lacertidae: Podarcis) as a case study

P. liolepis

opennotspecifiedDec 2022View details →
zenodo28/100

P in Identification of morphologically cryptic species with computer vision models: wall lizards (Squamata: Lacertidae: Podarcis) as a case study

P. hispanicus

opennotspecifiedDec 2022View details →
zenodo28/100

P in Identification of morphologically cryptic species with computer vision models: wall lizards (Squamata: Lacertidae: Podarcis) as a case study

P. hispanicus

opennotspecifiedDec 2022View details →
zenodo28/100

P in Identification of morphologically cryptic species with computer vision models: wall lizards (Squamata: Lacertidae: Podarcis) as a case study

P. guadarramae

opennotspecifiedDec 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record