Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
979
datasets available to search
ShareScore release 0.9.0
Dataset results
979 results for “image dataset”
PRMI: A dataset of minirhizotron images for diverse plant root study
<p>Understanding a plant's root system architecture (RSA) is crucial for a variety of plant science problem domains including sustainability and climate adaptation. Minirhizotron (MR) technology is a widely-used approach for phenotyping RSA non-destructively by capturing root imagery over time. Precisely segmenting roots from the soil in MR imagery is a critical step in studying RSA features. In this paper, we introduce a large-scale dataset of plant root images captured by MR technology. In total, there are over 72K RGB root images across six different species including cotton, papaya, peanut, sesame, sunflower, and switchgrass in the dataset. The images span a variety of conditions including varied root age, root structures, soil types, and depths under the soil surface. All of the images have been annotated with weak image-level labels indicating whether each image contains roots or not. The image-level labels can be used to support weakly supervised learning in plant root segmentation tasks. In addition, 63K images have been manually annotated to generate pixel-level binary masks indicating whether each pixel corresponds to root or not. These pixel-level binary masks can be used as ground truth for supervised learning in semantic segmentation tasks. By introducing this dataset, we aim to facilitate the automatic segmentation of roots and the research of RSA with deep learning and other image analysis algorithms.</p>
A Dataset of Thermal images of User Interfaces
<p>Recent advancement in sensor technology facilitates having a thermal camera at a lower price. These cameras have many potential applications but can also be used for malicious purposes, such as capturing user interfaces and retrieving user information from heat traces in the thermal images. This dataset is created during an interactive study investigating the threat of thermal attacks on user interfaces. We adapted the following experimental setup during data collection.</p> <ol> <li>2 camera perspectives- FLIR E8-XT camera placed behind the participant and Optris PI 450i camera placed left of the participant. </li> <li>4 types of input devices- i) smartphone, ii) 3 keyboards- a PBT keyboard, an ABS keyboard, and a metal frame keyboard</li> <li>3 types of user input data (text, email address, password)</li> </ol> <p>In summary, we have collected 1152 images from the FLIR camera and another 1152 images from the Optris camera through an interactive study with 32 participants. For each participant, we captured 36 images (9 types of user input, 4 types of input devices). The created dataset can be used to evaluate the deep learning model developed to prevent thermal imaging attacks. Furthermore, the ground truth user input of text, email address, and passwords are structured along with the corresponding image ID so that the advanced data-driven model can be employed to identify user input and investigate the type of user input that can be easily cracked using machine learning techniques.</p>
Low Resolution Thermal Imaging Dataset
<p>The dataset contains low resolution thermal images corresponding to various sign language digits represented by hand and captured using the Omron D6T thermal camera with a resolution of 32x32.</p>
Dataset for Mistic: an open-source multiplexed image t-SNE viewer
<p>This link consists of 10 anonymized non-small cell lung cancer (NSCLC) field of Views (FoVs) to test Mistic.</p> <p><strong>Mistic</strong></p> <p>Understanding the complex ecology of a tumor tissue and the spatio-temporal relationships between its cellular and microenvironment components is becoming a key component of translational research, especially in immune-oncology. The generation and analysis of multiplexed images from patient samples is of paramount importance to facilitate this understanding. In this work, we present Mistic, an open-source multiplexed image t-SNE viewer that enables the simultaneous viewing of multiple 2D images rendered using multiple layout options to provide an overall visual preview of the entire dataset. In particular, the positions of the images can be taken from t-SNE or UMAP coordinates. This grouped view of all the images further aids an exploratory understanding of the specific expression pattern of a given biomarker or collection of biomarkers across all images, helps to identify images expressing a particular phenotype or to select images for subsequent downstream analysis. Currently there is no freely available tool to generate such image t-SNEs.</p> <p><strong>Links</strong></p> <p><br> <a href="https://github.com/MathOnco/Mistic">Mistic code</a></p> <p><a href="https://mistic-rtd.readthedocs.io/">Mistic documentation</a></p> <p><a href="https://www.biorxiv.org/content/10.1101/2021.10.08.463728v1">Paper</a></p> <p> </p>
Evaluating Feature Attribution Methods in the Image Domain: High-Dimensional Datasets
<p>Here you can find the versions of the Places-365 and Caltech-256 datasets used in the paper Evaluating Feature Attribution Methods in the Image Domain.</p>
Plain_Background_Thermal_Imaging_Dataset
<p>The dataset contains the images captured from high resolution thermal camera. The thermal images contains ten different hand gestures captured from random people. We also captured images of both color and gray scale under different environment conditions. Further, different hand orientations are considered for the creation of effective dataset.</p>
DeepCAD-RT dataset: synthetic calcium imaging data
<p>DeepCAD-RT dataset: synthetic calcium imaging data</p>
RIGA+ Dataset for Unsupervised Domain Adaptation in Medical Image Segmentation
<p>Different from the previous combined multi-domain dataset for unsupervised domain adaptation (UDA) in medical image segmentation, this multi-domain fundus image dataset contains annotations made by the same group of ophthalmologists. Hence the annotator bias among different datasets can be mitigated. Therefore, this dataset can provide a relatively fair benchmark for evaluating UDA methods in fundus image segmentation.</p> <p>This dataset is based on the RIGA[1] dataset and MESSIDOR[2] dataset. We appreciate their efforts devoted by the authors of [1] and [2].</p> <p>The six duplicated cases in the RIGA dataset are filtered out according to the <a href="https://www.adcis.net/en/third-party/messidor/">Errata</a>. We also remove the duplicated cases that exist in both the RIGA dataset and the MESSIDOR dataset by hash value matching.</p> <table align="center"> <caption>Details of the RIGA+ dataset</caption> <thead> <tr> <th scope="row">Domain</th> <th scope="col">Dataset</th> <th scope="col"> <p>Labeled Samples</p> <p>(Train+Test)</p> </th> <th scope="col"> <p>Unlabeled</p> <p>Samples</p> </th> </tr> </thead> <tbody> <tr> <th scope="row">Source</th> <td>BinRushed</td> <td>195 (195+0)</td> <td>0</td> </tr> <tr> <th scope="row">Source</th> <td>Magrabia</td> <td>95 (95+0)</td> <td>0</td> </tr> <tr> <th scope="row">Target</th> <td>MESSIDOR-BASE1</td> <td>173 (138+35)</td> <td>227</td> </tr> <tr> <th scope="row">Target</th> <td>MESSIDOR-BASE2</td> <td>148 (118+30)</td> <td>238</td> </tr> <tr> <th scope="row">Target</th> <td>MESSIDOR-BASE3</td> <td>133 (106+27)</td> <td>252</td> </tr> </tbody> </table> <p>[1] Almazroa A, Alodhayb S, Osman E, et al. Retinal fundus images for glaucoma analysis: the RIGA dataset[C]//Medical Imaging 2018: Imaging Informatics for Healthcare, Research, and Applications. International Society for Optics and Photonics, 2018, 10579: 105790B.</p> <p>[2] Decencière E, Zhang X, Cazuguel G, et al. Feedback on a publicly distributed image database: the Messidor database[J]. Image Analysis & Stereology, 2014, 33(3): 231-234.</p> <p>If you find this dataset useful for your research, please consider citing the paper as follows:</p> <pre><code>@inproceedings{hu2022domain, title={Domain Specific Convolution and High Frequency Reconstruction based Unsupervised Domain Adaptation for Medical Image Segmentation}, author={Shishuai Hu and Zehui Liao and Yong Xia}, booktitle={International Conference on Medical Image Computing and Computer-Assisted Intervention}, year={2022}, organization={Springer} }</code></pre> <p> </p>
(SEN12MS) deepNIR: Dataset for generating synthetic NIR images
<p>This dataset contains <strong>SEN12MS </strong>NIR+RGB dataset used in our paper; deepNIR: Dataset for generating synthetic NIR images and improved fruit detection system using deep learning techniques.</p> <p>Please refer to <a href="http://tiny.one/deepNIR">http://tiny.one/deepNIR</a> for more detail.</p>
(capsicum) deepNIR: Dataset for generating synthetic NIR images
<p>This dataset contains <strong>capsicum</strong> NIR+RGB dataset used in our paper; deepNIR: Dataset for generating synthetic NIR images and improved fruit detection system using deep learning techniques.</p> <p>Please refer to <a href="http://tiny.one/deepNIR">http://tiny.one/deepNIR</a> for more detail.</p>
Automatic taxonomic identification based on the Fossil Image Dataset (>415,000 images) and deep convolutional neural networks
<p>The rapid and accurate taxonomic identification of fossils is of great significance in paleontology, biostratigraphy, and other fields. However, taxonomic identification is often labor-intensive and tedious, and the requisition of extensive prior knowledge about a taxonomic group also requires long-term training. Moreover, identification results are often inconsistent across researchers and communities. Accordingly, in this study, we used deep learning to support taxonomic identification. We used web crawlers to collect the Fossil Image Dataset (FID) via the Internet, obtaining 415,339 images belonging to 50 fossil clades. Then we trained three powerful convolutional neural networks on a high-performance workstation. The Inception ResNet v2 architecture achieved an average accuracy of 0.90 in the test dataset when transfer learning was applied. The clades of microfossils and vertebrate fossils exhibited the highest identification accuracies of 0.95 and 0.90, respectively. In contrast, clades of sponges, bryozoans, and trace fossils with various morphologies or with few samples in the dataset exhibited a performance below 0.80. Visual explanation methods further highlighted the discrepancies among different fossil clades and suggested similarities between the identifications made by machine classifiers and taxonomists. Collecting large paleontological datasets from various sources, such as the literature, digitization of dark data, citizen-science data, and public data from the Internet may further enhance deep learning methods and their adoption. Such developments will also possibly lead to image-based systematic taxonomy to be replaced by machine-aided classification in the future. Pioneering studies can include microfossils and some invertebrate fossils. To contribute to this development, we deployed our model on a server for public access at www.ai-fossil.com.</p>
(nirscene) deepNIR: Dataset for generating synthetic NIR images
<p>This dataset contains <strong>nirscene</strong> NIR+RGB dataset used in our paper; deepNIR: Dataset for generating synthetic NIR images and improved fruit detection system using deep learning techniques.</p> <p>Please refer to <a href="http://tiny.one/deepNIR">http://tiny.one/deepNIR</a> for more detail.</p> <p> </p>
An Ordovician to Silurian graptolite specimen image dataset for global correlation and shale gas exploration
<p>A unique high-resolution image dataset consists of key graptolite species used for dating rocks, global correlation, and “gold caliper” for locating shale gas favourable exploration beds (FEBs) in China.</p> <p>All images were taken from 1,550 carefully curated graptolite specimens, taxonomically belong to 113 graptolite species or subspecies. These specimens were collected from 154 representative geological sections of the Ordovician to Silurian sediments of China and published in 1958-2020. All specimens are housed at the Nanjing Institute of Geology and Palaeontology (NIGP), Chinese Academy of Sciences (CAS). Detailed scientific information of every piece of fossil specimen is given in the attached spreadsheet file.</p> <p>My working group spent over two years to complete photographing every specimen using a single-lens reflex camera Nikon D800E with Nikkor 60 mm macro-lens and Leica M125 and M205C microscopes equipped with Leica cameras. Every image is well focused and better shows the morphology of graptolite bodies.</p> <p>In total, we took 40,597 images, including 20,644 camera photos (each with a resolution of 4,912 × 7,360) and 19,953 microscope photos (each with a resolution of 2,720 × 2,048). Photos of low contrast or bad focus were removed from the whole collection. We only kept and selected the photos that show the visual morphology of every specimen and the diagnostic character of each graptolite species that the specimens represent. We selected one image for each specimen as the present final dataset, uploaded to and stored in our cloud server.</p> <p>We incorporated revision suggestions from distinguished palaeontologists to generate the ground-truth labels, providing a taxonomical authority of the dataset. The dataset potentially contributes to a range of scientific activities and provides 1) easy access to high-resolution images of 2951 specimens of 113 graptolite species for teaching and training in palaeontology and geologic survey; 2) Global bio-stratigraphic correlation using graptolites, especially with those bio-zone species; 3) A standard fossil specimen image dataset used in shale gas industry to improve exploration efficiency, and 4) The potential aid of developing image-based automated classification model.</p> <p>Every specimen has two photos, one is original, another shows specimen with a scale bar. Occasionally in some large image the scale bar is embedded and beside the fossil specimen. Example: The file name: ‘9721Cardiograptus_amplus_S.jpg’, ‘9721’ is the specimens number, ‘Cardiograptus_amplus’ means species name is ‘Cardiograptus amplus’, with ‘_S’ means it is a photo with scale bar. In all scale bar, the minimum unit is millimeter.</p> <p>Author and contact:</p> <p>Hong-He Xu</p> <p>Nanjing Institute of Geology and Palaeontology, Chinese Academy of Sciences</p> <p>39 East Beijing Road, Nanjing, 210008</p> <p>China</p> <p>E-mail: hhxu@nigpas.ac.cn</p>
A Dataset Containing Tiny and Low Quality Images for Vehicle Classification
<p>This dataset contains 4800 tiny and low resolution vehicle images. The vehicles in the images are grouped in six classes: Bike, Car, Juggernaut, Minibus, Pickup, and Truck. For each class, there are 800 vehicle images with 100 × 100 pixels and 96 dpi resolution.</p> <p><strong>The peer-reviewed research article for this dataset has been published in MDPI Sensors, and can be accessed here: <a href="https://doi.org/10.3390/s22134740">https://doi.org/10.3390/s22134740</a>. Please cite this when using the dataset.</strong></p>
Dataset_3 for "An integrated imaging sensor for aberration-corrected 3D photography"
<p>...</p>
Dataset_2 for "An integrated imaging sensor for aberration-corrected 3D photography"
<p>...</p>
Dataset_1 for "An integrated imaging sensor for aberration-corrected 3D photography"
<p>...</p>
Dataset with square plots across Sierra Nevada (Spain) where the contours of all juniper shrubs were annotated as polygons using centimetric GPS and very high resolution aerial and satellite RGB images
<p><strong>This dataset is a shapefile of 767 polygons describing the contours of Juniperus communis L. and Juniperus sabina L. shrubs for the year 2021 in rectangular plots across Sierra Nevada. The coordinates of the polygons were obtained from a field work campaign with a differential centimetric GPS, and their contours were drawn manually in QGIS using the Google Earth satellite image for 2020 and the PNOA aerial image for the 2020. </strong></p> <p><strong>This dataset also contains an excel file describing the features of each polygon: the polygon centroid coordinates, the type of species, the sexgender, the morphotype, the damage in the vegetation cover estimated in the field and telematically, certainty of the digitalization with QGIS and also if the differential centimetric GPS used belongs to the University of Granada or the University of Almeria. </strong></p>
CODEX multiplexed imaging cell datasets used for using STELLAR to transfer cell type annotations to other tissues and donors
<p>We performed CODEX (co-detection by indexing) multiplexed imaging on 24 sections of the human intestine from 3 donors (B004, B005, B006) using a panel of 47 oligonucleotide-barcoded antibodies. We also performed CODEX imaging on both human tonsil and Barrett's esophagus (BE) using a panel of 57 oligonucleotide-barcoded antibodies. Subsequently images underwent standard CODEX image processing (tile stitching, drift compensation, cycle concatenation, background subtraction, deconvolution, and determination of best focal plane), single cell segmentation, and column marker z-normalization by tissue. Output of this process were dataframes of 870,000 cells and 220,000 cells respectively with fluorescence values quantified from each marker.</p>
PFRR ASC Image Dataset for Troyer et al. 2022 (Substorm Activity as a Driver of Energetic Pulsating Aurora)
<p>This data set contains movies taken from the Poker Flat Digital All Sky Camera of nights where pulsating aurora occurred at the same time that the Poker Flat Incoherent Scatter Radar was running in an MSWinds (D-region) mode. There are two types of movies here. File names looking like 428-YYYY-MM-DD.mp4 are movies compiled from the 428 nm images. File names looking like YYYY-MM-DD-white.mp3 are movies compiled by creating a grayscale image using the 428, 558, and 630nm images.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.