Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

979

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

979 results for “image dataset”

Learn how ShareScore rates datasets ↗
zenodo40/100

Dataset used in "Ocean floor imaging with Distributed Acoustic Sensing and water phases reverberations" by Spica et al. in Geophysical Research Letters

<p>earthquake #1<br> earthquake #2</p>

opencc-by-4.0Jul 2021View details →
zenodo40/100

Dataset on UAV High-resolution Images from Grassland with Broad-leaved Dock (Rumex Obtusifolius)

<p>The dataset consists of orthophotos (build from UAV images) from&nbsp;a&nbsp;grassland field in which several <em>Rumex obtusifolius</em> plants were&nbsp;detected. The field is located in Germany (Kleve). The UAV images were acquired at 10, 15, and 30 meters height. Moreover, the&nbsp;labels/annotations from the <em>Rumex obtusifolius</em> plants in the images&nbsp;are also provided.&nbsp;</p>

opencc-by-4.0Dec 2020View details →
zenodo40/100

RELLISUR: A Real Low-Light Image Super-Resolution Dataset

<p>The RELLISUR dataset contains real low-light low-resolution images paired with normal-light high-resolution reference image counterparts. This dataset aims to fill the gap between low-light image enhancement and low-resolution image enhancement (Super-Resolution (SR)) which is currently only being addressed separately in the literature, even though the visibility of real-world images is often limited by both low-light and low-resolution.&nbsp; The dataset contains 12750 paired images of different resolutions and degrees of low-light illumination, to facilitate learning of deep-learning based models that can perform a direct mapping from degraded images with low visibility to high-quality detail rich images of high resolution. The associated paper can be found here: <a title="https://datasets-benchmarks-proceedings.neurips.cc/paper/2021/file/7ef605fc8dba5425d6965fbd4c8fbe1f-Paper-round2.pdf" href="https://datasets-benchmarks-proceedings.neurips.cc/paper/2021/file/7ef605fc8dba5425d6965fbd4c8fbe1f-Paper-round2.pdf" target="_blank" rel="noopener">https://datasets-benchmarks-proceedings.neurips.cc/paper/2021/file/7ef605fc8dba5425d6965fbd4c8fbe1f-Paper-round2.pdf</a></p>

opencc-by-4.0Aug 2021View details →
zenodo40/100

Dataset used in "Uncertainty-Aware Learning for Improvements in Image Quality of the Canada-France-Hawaii Telescope" (https://arxiv.org/abs/2107.00048)

<p>&#39;x_train.p&#39;, &#39;y_train.p&#39;: pickle files for training split containing&nbsp;50,757 samples</p> <p>&#39;x_val.p&#39;, &#39;y_val.p&#39;: pickle file for validation split containing 5,640 samples</p> <p>&#39;x_test.p&#39;, &#39;y_test.p&#39;: pickle file for test split containing 6,267 samples</p> <p>&#39;feature_names.p&#39;: pickle file containing names of all 119 features</p>

opencc-by-4.0Aug 2021View details →
zenodo40/100

Data Models for Dataset Drift Controls in Machine Learning With Optical Images - Datasets

<p>This dataset accompanies the paper&nbsp;titled</p> <p><em>Data Models for Dataset Drift Controls in Machine Learning with Images</em><br> <br> that appeared in the Transactions on Machine Learning Research<br> <br> <a href="https://openreview.net/forum?id=I4IkGmgFJz">https://openreview.net/forum?id=I4IkGmgFJz</a><br> &nbsp;</p> <pre><code>@article{ oala2023data, title={Data Models for Dataset Drift Controls in Machine Learning With Optical Images}, author={Luis Oala and Marco Aversa and Gabriel Nobis and Kurt Willis and Yoan Neuenschwander and Mich{\`e}le Buck and Christian Matek and Jerome Extermann and Enrico Pomarico and Wojciech Samek and Roderick Murray-Smith and Christoph Clausen and Bruno Sanguinetti}, journal={Transactions on Machine Learning Research}, issn={2835-8856}, year={2023}, url={https://openreview.net/forum?id=I4IkGmgFJz}, note={} }</code></pre> <p>We make available two datasets.</p> <p><strong>Raw-Microscopy:</strong></p> <ul> <li><strong>940 raw bright-field microscopy images</strong> of human blood smear slides for leukocyte classification (microscopy/images/raw_scale100) with corresponding labels (microscopy/labels).</li> <li><strong>5,640 variations measured at six additional different intensities </strong>(microscopy/images/raw_scale001-raw_scale0075)</li> <li><strong>11,280 images of the raw sensor data processed through twelve different pipelines</strong> (microscopy/images/processed_views)</li> </ul> <p><strong>Raw-Drone:</strong></p> <ul> <li><strong>548 raw drone camera images for car segmentation</strong> (drone/images_tiles_256/raw_scale100) with corresponding binary segmentation mask (drone/masks_tiles_256). The images and the masks are cropped from 12 raw drone camera images (drone/images_full/raw_scale100) and 12 masks (drone/masks_full) of size 3648 by 5472.</li> <li><strong>3,288 variations measured at six additional different intensities</strong> (drone/images_tiles_256/raw_scale001-raw_scale075).</li> <li><strong>6,576 images of the raw sensor data processed through twelve different pipelines</strong> (drone/images_tiles_256/processed_views).</li> </ul> <p>Detailed datasheets for the two datasets can be found in the appendices of the TMLR paper.</p> <p>The code repository for this project can be found at&nbsp;<a href="https://github.com/aiaudit-org/raw2logit">https://github.com/aiaudit-org/raw2logit</a></p> <p>&nbsp;</p>

opencc-by-4.0May 2023View details →
zenodo40/100

Dataset - Impact of 3D radiative transfer on airborne NO2 imaging remote sensing over cities with buildings

<p>This dataset was created by Marc Schwaerzel (marc.schwaerzel@empa.ch) and is intended to get along with the Schwaerzel et al. (2021) AMT publication (amt-2020-146) . The data and the data structure is described in the<em> <strong>readme.md </strong></em>text file.</p> <p>The dataset contains:</p> <p>- libRadtran output (radiances and AMFs)</p> <p>- Synthetic SCDs</p>

opencc-by-4.0Sep 2021View details →
zenodo40/100

Lake Malawi Cichlid image dataset

<p>This photo dataset is the raw data used in the work&nbsp;Identification of Cichlid Fishes from Lake Malawi Using Computer Vision (https://doi.org/10.1371/journal.pone.0077686).</p> <p>https://github.com/forcecore/ghoti : The repository of the original work</p> <p>https://github.com/forcecore/ghoti-2021 : Renewed, deep-learning-powered version of the work (as a tutorial)</p> <p>&nbsp;</p> <p>Later, a genetic level study was done on these specimens:&nbsp;https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6764894/<br> The related data is here:&nbsp;https://datadryad.org/stash/dataset/doi:10.5061/dryad.258nm86</p>

opencc-by-4.0Oct 2021View details →
zenodo40/100

Dataset with Agricultural Parcels Markup on Satellite Images

<p>The dataset was created for the development and testing of the algorithm proposed in the paper &quot;Segmentation of agricultural parcel in satellite images based on historical vegetation index data&quot;. However, this markup can be used to test other algorithms and compare their quality.</p> <p>This work employs data from the remote sensing programs Sentinel-2A, Sentinel-2B.</p> <p>The dataset contains the agricultural parcel markup&nbsp;for four regions within Russia and Ukraine, where agriculture is well developed.&nbsp;The areas were chosen so that each of them had other types of terrain in addition to fields: urban area, water surface, swamps, and forests.</p> <p>&nbsp;</p>

opencc-by-4.0Oct 2021View details →
zenodo40/100

OCTA image dataset with label annotation for quality assessment

<p>This dataset is publish by the research &quot;<em>A Deep Learning-based Quality Assessment and Segmentation System with a Large-scale Benchmark Dataset for Optical Coherence Tomographic Angiography Image</em>&quot;</p> <p>Detail:</p> <p>OCTA image dataset with label annotation for quality assessment.&nbsp;sOCTA-3x3-10k: 10,480 3 &times; 3 mm<sup>2</sup> superficial vascular layer OCTA (sOCTA) images divided into three classes; sOCTA-6x6-14k: 14,042 6 &times; 6 mm<sup>2</sup> sOCTA images divided into three classes.&nbsp;</p> <p>GitHub:&nbsp;<a href="https://github.com/shanzha09/COIPS">https://github.com/shanzha09/COIPS</a></p> <p>These datasets are public available, if you use the dataset or our system in your research, please <strong>cite</strong> our paper:&nbsp;<em><code>A Deep Learning-based Quality Assessment and Segmentation System with a Large-scale Benchmark Dataset for Optical Coherence Tomographic Angiography Image</code></em>.</p> <p>arXiv:<a href="https://arxiv.org/abs/2107.10476v1">https://arxiv.org/abs/2107.10476v1</a></p>

opencc-by-4.0Jul 2021View details →
zenodo40/100

Dataset for Adaptive Light-Sheet Fluorescence Microscopy with a Deformable Mirror for Video-Rate Volumetric Imaging

<p>1. Underlying data of figures in the&nbsp;paper&nbsp;</p> <p>2. Background images used to process the experimental data</p> <p>3. image stack of 250 nm beads</p> <p>4. image stack of sunflower pollen grains</p> <p>5. image stacks and videos of Fluo-4 labelled cells</p> <p>6. image stacks and videos of CMO-labelled cells</p> <p>The data is organised according to the figures they are related to in the following publication:</p> <p>&nbsp;</p> <p><a href="https://aip.scitation.org/author/Hong%2C+Wenzhi">Wenzhi Hong</a><em>,&nbsp;</em><a href="https://aip.scitation.org/author/Wright%2C+Terry">Terry Wright</a><em>,&nbsp;</em><a href="https://aip.scitation.org/author/Sparks%2C+Hugh">Hugh Sparks</a><em>,&nbsp;</em><a href="https://aip.scitation.org/author/Dvinskikh%2C+Liuba">Liuba Dvinskikh</a><em>,&nbsp;</em><a href="https://aip.scitation.org/author/MacLeod%2C+Ken">Ken MacLeod</a><em>,&nbsp;</em><a href="https://aip.scitation.org/author/Paterson%2C+Carl">Carl Paterson</a><em>, and&nbsp;</em><a href="https://aip.scitation.org/author/Dunsby%2C+Chris">Chris Dunsby</a>&nbsp;</p> <p>, &quot;Adaptive light-sheet fluorescence microscopy with a deformable mirror for video-rate volumetric imaging&quot;, Appl. Phys. Lett.&nbsp;121, 193703&nbsp;(2022)&nbsp;<a href="https://doi.org/10.1063/5.0125946">https://doi.org/10.1063/5.0125946</a></p>

opencc-by-4.0Oct 2022View details →
zenodo40/100

WHU-OHS: A benchmark dataset for large-scale Hyperspectral Image classification

<p>The WHU-OHS dataset is made up of 42 OHS satellite images acquired from more than 40 different locations in China. The imagery has a spatial resolution of 10 m (nadir) and a swath width of 60 km (nadir). There are 32 spectral channels ranging from the visible to near-infrared range, with an average spectral resolution of 15 nm. We cropped each image into 512 &times; 512 pixels with a stride of 32. There are 4822, 513, and 2460 sub-images in the training, validation, and test sets, respectively.</p> <p>For transferability test, we choose eight pairs of OHS images, and each pair contains one source image (S) and one target image (T):</p> <p>S1: Changchun</p> <p>T1: Jilin</p> <p>S2: Wuxi</p> <p>T2: Shanghai</p> <p>S3: Guangzhou</p> <p>T3: Zhongshan</p> <p>S4: Xining</p> <p>T4: Lanzhou</p> <p>S5: Hetian</p> <p>T5: Kelamayi</p> <p>S6: Anyi</p> <p>T6: Nanchang</p> <p>S7: Changde</p> <p>T7: Changsha</p> <p>S8: Tianjin</p> <p>T8: Tangshan</p> <p>The 26 OHS images except for the eight pairs:</p> <p>O1: Baoding</p> <p>O2: Chongqing</p> <p>O3: Fujin</p> <p>O4: Huainan</p> <p>O5: Huhehaote</p> <p>O6: Jinzhong</p> <p>O7: Luliang</p> <p>O8: Manasi_1</p> <p>O9: Manasi_2</p> <p>O10: Nanmulin</p> <p>O11: Neimenggu</p> <p>O12: Qingdao</p> <p>O13: Qinghuangdao</p> <p>O14: Shawan</p> <p>O15: Shenyang</p> <p>O16: Shuozhou</p> <p>O17: Songpan</p> <p>O18: Taian</p> <p>O19: Tongjiang_1</p> <p>O20: Tongjiang_2</p> <p>O21: Wuzhong</p> <p>O22: Xundian</p> <p>O23: Xuzhou</p> <p>O24: Yidu</p> <p>O25: Zangzu</p> <p>O26: Zhongshan</p> <p>The image patches have been normalized and scaled by 10000 to reduce storage cost. Divide the pixel values by 10000 and then the image patches can be used directly.</p>

opencc-by-4.0Sep 2022View details →
zenodo40/100

UAV RGB imagery dataset captured at nadir and oblique angles over pistachio trees in Spain, including images, GCPs, 3D point cloud and orthomosaic.

<p>The dataset comprises 248 images taken in two flights on 29 July 2021&nbsp;over a pistachio orchard in Spain. In addition, GCPs (ground control points) were collected to improve the photogrammetric process accuracy. The photos were taken using a UAV DJI Phantom Advance quadcopter equipped with a DJI FC6310 RGB 20-megapixel camera. The first flight mission was planned to take nadir images (-90&ordm; gimbal pitch degree), whereas the second flight was scheduled to take oblique images (-60&ordm; gimbal pitch degree), both at 55 metres above the ground. In addition, the images were used to generate a 3D point cloud, DEM and&nbsp;orthomosaic, which were included in the dataset.This dataset is useful for precision agriculture researchers interested in photogrammetric reconstruction.</p>

opencc-by-4.0Dec 2021View details →
zenodo40/100

Dataset Literature Review Digital Forensic and Image Processing

<p>Data ini digunakan untuk membuat penelitian sesuai dengan tinjauan literatur dengan kata kunci &quot;<em>digital forensic</em>&quot; dan &quot;<em>image processing</em>&quot;</p>

opencc-by-4.0Nov 2022View details →
zenodo40/100

Dataset Literature Review Digital Forensic AND Image Processing

<p>Data ini digunakan untuk membuat penelitian berdasarkan tinjauan literatur dengan kata kunci &quot;digital forensic&quot; dan &quot;image processing&quot;&nbsp;</p>

opencc-by-4.0Nov 2022View details →
zenodo40/100

GLAMI-1M: A Multilingual Image-Text Fashion Dataset

<p>We introduce GLAMI-1M: the largest multilingual image-text classification dataset and benchmark. The dataset contains images of fashion products with item descriptions, each in 1 of 13 languages. Categorization into 191 classes has high-quality annotations: all 100k images in the test set and 75% of the 1M training set were human-labeled. The paper presents baselines for image-text classification showing that the dataset presents a challenging fine-grained classification problem: The best scoring EmbraceNet model using both visual and textual features achieves 69.7% accuracy. Experiments with a modified Imagen model show the dataset is also suitable for image generation conditioned on text.&nbsp;The dataset, source code and model checkpoints are published at: https://github.com/glami/glami-1m.</p>

openapache2.0Nov 2022View details →
zenodo40/100

GLAMI-1M: A Multilingual Image-Text Fashion Dataset - 800px

<p>We introduce GLAMI-1M: the largest multilingual image-text classification dataset and benchmark. The dataset contains images of fashion products with item descriptions, each in 1 of 13 languages. Categorization into 191 classes has high-quality annotations: all 100k images in the test set and 75% of the 1M training set were human-labeled. The paper presents baselines for image-text classification showing that the dataset presents a challenging fine-grained classification problem: The best scoring EmbraceNet model using both visual and textual features achieves 69.7% accuracy. Experiments with a modified Imagen model show the dataset is also suitable for image generation conditioned on text.&nbsp;The dataset, source code and model checkpoints are published at: https://github.com/glami/glami-1m</p>

openapache2.0Nov 2022View details →
zenodo40/100

A crowdsourced dataset of aerial images with annotated solar photovoltaic arrays and installation metadata

<p><strong>Summary</strong></p> <p>Photovoltaic (PV) energy generation plays a crucial role in the energy transition. Small-scale, residential PV installations are deployed at an unprecedented pace, and their safe integration into the grid necessitates up-to-date, high-quality information. Overhead imagery is increasingly used to improve the knowledge of residential PV installations with machine learning models capable of automatically mapping these installations. However, these models cannot be reliably transferred from one region or imagery source to another without incurring a decrease in accuracy. To address this issue, known as distribution shift, and foster the development of PV array mapping pipelines, we propose a dataset containing aerial images, segmentation masks, and installation metadata. We provide installation metadata for more than 28000 installations. We provide ground truth segmentation masks for 13000 installations, including 7000 with annotations for two different image providers. Finally, we provide installation metadata that matches the annotation for more than 8000 installations. Dataset applications include end-to-end PV registry construction, robust PV installations mapping, and analysis of crowdsourced datasets.</p> <p>This dataset contains the complete records&nbsp;associated with the article &quot;A crowdsourced dataset of aerial images of solar panels, their segmentation masks, and characteristics&quot;, published in Scientific data. The article is accessible here :&nbsp;<a href="https://www.nature.com/articles/s41597-023-01951-4">https://www.nature.com/articles/s41597-023-01951-4</a> These complete records consist of:</p> <ol> <li>The complete training dataset containing RGB overhead imagery, segmentation masks and metadata of PV installations (folder <strong>bdappv</strong>),</li> <li>The raw crowdsourcing data, and the postprocessed data for replication and validation (folder <strong>data</strong>).</li> </ol> <p><strong>Data records</strong></p> <p>Folders are organized as follows:</p> <ul> <li><strong>bdappv/</strong> Root data folder <ul> <li><strong>google / ign:</strong>&nbsp; One folder for each campaign <ul> <li><strong>img/</strong>: Folder containing all the images presented to the users. This folder contains 28807 images for Google and 17325 images for IGN.</li> <li><strong>mask/</strong>: Folder containing all segmentations masks generated from the polygon annotations of the users. This folder contains 13303 masks for Google and&nbsp;7686 masks for IGN.</li> </ul> </li> <li><em>metadata.csv</em> The <code>.csv</code>&nbsp; file with the installations&#39; metadata.</li> </ul> </li> </ul> <p>&nbsp;</p> <ul> <li><strong>data/ </strong>Root data folder <ul> <li><strong>raw/</strong> Folder containing the raw crowdsourcing data and raw metadata; <ul> <li><em>input-google.json</em>: <code>.json </code>input data data containing all information on images and raw annotators&rsquo; contributions for both phases (clicks and polygons) during the first annotation campaign;</li> <li><em>input-ign.json</em>:<em> </em><code>.json </code>input data containing all information on images and raw annotators&rsquo; contributions for both phases (clicks and polygons) during the second annotation campaign;</li> <li><em>raw-metadata.json</em>: <code>.json </code>output containing the PV systems&rsquo; metadata extracted from the BDPV database before filtering. It can be used to replicate the association between the installations and the segmentation masks, as done in the notebook metadata.</li> </ul> </li> <li><strong>replication/</strong> Folder containing the compiled data used to generate the segmentation masks; <ul> <li><strong>campaign-google/campaign-ign</strong>: One folder for each campaign <ul> <li><em>click-analysis.json</em>: <code>.json </code>output on the click analysis, compiling raw input into a few best-guess locations for the PV arrays. This dataset enables the replication of our annotations,</li> <li><em>polygon-analysis.json</em>: <code>.json </code>output of polygon analysis, compiling raw input into a best-guess polygon for the PV arrays.</li> </ul> </li> </ul> </li> <li><strong>validation/</strong> Folder containing the compiled data used for technical validation. <ul> <li><strong>campaign-google/campaign-ign</strong>: One folder for each campaign <ul> <li><em>click-analysis-thres=1.0.json</em>: <code>.json </code>output of the click analysis with a lowered threshold to analyze the effect of the threshold on image classification, as done in the notebook annotation;</li> <li><em>polygon-analysis-thres=1.0.json</em>: <code>.json </code>output of polygon analysis, with a lowered threshold to analyze the effect of the threshold on polygon annotation, as done in the notebook annotations.</li> </ul> </li> <li><em>metadata.csv</em>: the <code>.csv </code>file of filtered installations&#39; metadata.</li> </ul> </li> </ul> </li> </ul> <p><strong>License</strong></p> <p>We extracted the thumbnails contained in the <strong>google/img/</strong> folder using Google Earth Engine API and we generated the thumbnails contained in the <strong>ign/img</strong><strong>/</strong> folder from high resolution tiles downloaded from the online IGN portal accessible here: <a href="https://geoservices.ign.fr/bdortho">https://geoservices.ign.fr/bdortho</a>. Images provided by Google are subjet to Google&#39;s terms and conditions. Images provided by the IGN are subject to an open license 2.0.</p> <p>Access the terms and conditions of Google images at this URL: <a href="https://www.google.com/intl/en/help/legalnotices_maps/">https://www.google.com/intl/en/help/legalnotices_maps/</a></p> <p>Access the terms and conditions of IGN images at this URL: <a href="https://www.etalab.gouv.fr/wp-content/uploads/2018/11/open-licence.pdf">https://www.etalab.gouv.fr/wp-content/uploads/2018/11/open-licence.pdf</a></p>

opencc-by-4.0Jul 2022View details →
zenodo40/100

Benchmark datasets for detection and identification of insects from camera trap images with deep learning

<p><strong>Insect benchmark datasets for training, validation and test (train1201.zip, val1201.zip and test1201.zip)&nbsp;with time-lapse images as described in paper:</strong></p> <p><a href="https://www.biorxiv.org/content/10.1101/2022.10.25.513484v1">Bjerge K, Alison J, Dyrmann M, Frigaard C.E., Mann H. M. R., H&oslash;ye T.T., Accurate detection and identification of insects from camera trap images with deep learning, bioRxiv:10.1101/2022.10.25.513484v1</a></p> <p>Labels in&nbsp;<strong>YOLO format:&nbsp;<a href="https://github.com/ultralytics/yolov5/issues/2293">ultralytics/yolov5: label format</a></strong></p> <p>The annotated training and validation datasets contains insects of nine different species as listed below:</p> <table> <tbody> <tr> <td>0&nbsp;<em>Coccinellidae septempunctata</em></td> </tr> <tr> <td>1&nbsp;<em>Apis mellifera</em></td> </tr> <tr> <td>2&nbsp;<em>Bombus lapidarius</em></td> </tr> <tr> <td>3&nbsp;<em>Bombus terrestris</em></td> </tr> <tr> <td>4&nbsp;<em>Eupeodes corolla</em></td> </tr> <tr> <td>5&nbsp;<em>Episyrphus balteatus</em></td> </tr> <tr> <td>6&nbsp;<em>Aglais urticae</em></td> </tr> <tr> <td>7&nbsp;<em>Vespula vulgaris</em></td> </tr> <tr> <td>8&nbsp;<em>Eristalis tenax</em></td> </tr> </tbody> </table> <p>The test dataset contains additional classes of insects.</p> <table> <tbody> <tr> <td>9 Non-Bombus Anthophila</td> </tr> <tr> <td>10 Bombus spp.</td> </tr> <tr> <td>11 Syrphidae</td> </tr> <tr> <td>12 Fly spp.</td> </tr> <tr> <td>13 Unclear insect</td> </tr> <tr> <td>14 Mixed animals:<br> &mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;<br> Rhopalocera<br> Non-Anthophila Hymenoptera<br> Non-Syrphidae Diptera<br> Non-Conccinalidae Coleoptera<br> Concinellidae<br> Other animals</td> </tr> </tbody> </table> <p><strong>There are two naming conventions for image (.jpg) and label (.txt) files.</strong></p> <p><em>Background images without insects are named</em>:<br> &ldquo;<strong>X_Seq-YYYYMMDDHHMMSS</strong>-snapshot&rdquo;.<br> E.g.:<br> Background image: 12_13-20190704172200-snapshot.jpg<br> Empty label file: 12_13-20190704172200-snapshot.txt</p> <p><em>Images annotated with insects are named:</em><br> &ldquo;<strong>SZ_IP-MonthDate_C_Seq-YYYYMMDDHHMMSS</strong>&rdquo;.<br> E.g.:<br> Image file: S1_146-Aug23_1_156-20190822133230.jpg<br> Label file: S1_146-Aug23_1_156-20190822133230.txt</p> <p><strong>Abbreviations</strong>:</p> <p><strong>YYYYMMDDHHMMSS&nbsp;</strong>&ndash; Capture timestamp with year, month, date, hour, minutes, and second<br> <strong>Seq</strong>&nbsp;&ndash; Sequence number created by the motion program to separate images<br> <strong>C</strong>&nbsp;&ndash; Identification of two cameras with Id=0 or Id=1 in system identified by&nbsp;<strong>SZ_IP</strong><br> <strong>MonthDate&nbsp;</strong>&ndash; Folder name for where the original image were stored in the system<br> <strong>SZ_IP</strong>&nbsp;&ndash; Identification of five camera systems: S1_123, S2_146, S3_194, S4_199, S5_187 (Two cameras in each system)<br> <strong>X</strong>&nbsp;&ndash; An index number related to a specific camera and folder ensuring unique file names of background images from different camera systems.<br> <br> The important information in a filename is system (<strong>SZ_IP</strong>), camera Id (<strong>C</strong>) and timestamp (<strong>YYYYMMDDHHMMSS</strong>).</p> <p><strong>The three best YOLOv5 models (YOLOv5models.zip)&nbsp;from the paper are available in pytorch format.</strong></p> <p>All models are tested with YOLOv5 release v7.0 (22-11-2022):&nbsp;<a href="https://github.com/ultralytics/yolov5">ultralytics/yolov5: YOLOv5&nbsp;&nbsp;in PyTorch</a></p> <p><strong>insect1201-bestF1-640v5m.pt</strong>: Model no. 6 in Table 2 (F1=0.912)<br> <strong>insect1201-bestF1-1280v5m6.pt</strong>: Model no. 8 in Table 2 (F1=0.925)<br> <strong>insect1201-bestF1-1280v5m6.pt</strong>: Model no. 10 in Table 2 (F1=0.932)</p> <p><strong>insects-1201val.yaml</strong>: YAML file with label names to train YOLOv5</p> <p><strong>trainInsects-1201m.sh</strong>: Linux bash shell script with parameters to train YOLOv5m6<br> <strong>valInsectsF1-1201.sh</strong>: Linux bash shell script with parameters to validated models</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2022View details →
zenodo40/100

The PS-Battles Dataset - an Image Collection for Image Manipulation Detection

<p>The boost of available digital media has led to a significant increase in derivative work. With tools for manipulating objects becoming more and more mature, it can be very difficult to determine whether one piece of media was derived from another one or tampered with. As derivations can be done with malicious intent, there is an urgent need for reliable and easily usable tampering detection methods. However, even media considered semantically untampered by humans might have already undergone compression steps or light post-processing, making automated detection of tampering susceptible to false positives. In this paper, we present the PS-Battles dataset which is gathered from a large community of image manipulation enthusiasts and provides a basis for media derivation and manipulation detection in the visual domain. The dataset consists of 102&#39;028 images grouped into 11&#39;142 subsets, each containing the original image as well as a varying number of manipulated derivatives.</p> <p>Mirror of the git repository:&nbsp;<a href="https://github.com/dbisUnibas/PS-Battles">https://github.com/dbisUnibas/PS-Battles</a></p> <p>Paper on arxiv:&nbsp;<a href="https://arxiv.org/abs/1804.04866">https://arxiv.org/abs/1804.04866</a></p>

opencc-by-4.0Apr 2018View details →
zenodo40/100

Image Dataset for Object Detection of Small Size Construction Tools

<p>&nbsp;This is an image dataset established as input data for object detection model of small-sized construction tools. In the dataset, there are 12 classes of target tools&nbsp; (bucket, cutter, drill, grinder, hammer, knife, saw, shovel, spanner, tacker, trowel, and wrench) which are typically used at indoor construction sites. 25,084 sets of image and the corresponding label data have been established and shared.&nbsp;</p> <p>&nbsp;The diversity of objects in the images of the 12 small tools was considered by photographing tools of various shapes, sizes, and colors. In addition, to improve the model performance, images were also captured with various changes (e.g., image resolution, occlusion, lighting, and background). Among the 25,084 images in the dataset, 6,258 (25%) were obtained from the actual construction site.&nbsp;</p> <p>&nbsp;Object annotations in each image were done by bounding boxes and were saved into a text file. The coordinates of the bounding box have the form of (Class, Center X, Center Y, Width, Height). Class refers to one of 12 construction tool types. Center X and Center Y are the center coordinates of the bounding box for an object from an image when the resolution of the image has min-max normalized. Width and Height are the width and height of the bounding box for an object, respectively, also from the image with the min-max normalized resolution.</p> <p>&nbsp;</p> <p>The peer-reviewed publication for this dataset has now been published in &quot; KSCE Journal of Civil Engineering&quot; a Springer journal as follows:</p> <p><strong>* Lee, K., Jeon, C., and Shin, D. (2023, In press) &quot;Small Tool Image Database and Object Detection Approach for Indoor Construction Site Safety&quot; <em>KSCE Journal of Civil Engineering</em>. DOI: https://doi.org/10.1007/s12205-022-1011-7</strong></p> <p>Please cite this reference when using the dataset.</p>

opencc-by-4.0Dec 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record