Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
166
datasets available to search
ShareScore release 0.9.0
Dataset results
166 results for “Image Classification”
Transurethral Ultrasonic Imaging For Detection and Classification of Prostate Cancer
ClinicalTrials.gov study NCT02307552. IPD Sharing: UNDECIDED. Countries: 1. Publications: 1.
Training and test data for: Not getting in too deep: A practical deep learning approach to routine crystallisation image classification
Open the record for dataset details and reuse information.
Data from: Integrating a UAV-derived DEM in object-based image analysis increases habitat classification accuracy on coral reefs
Open the record for dataset details and reuse information.
Mapping built infrastructure in semi-arid systems using data integration and open-source approaches for image classification
Open the record for dataset details and reuse information.
Urban forest classification of the central Arizona-Phoenix area from a Landiscor image, year 1997
The dataset is the result of applying parameters developed for the object-oriented classification of high resolution, true-color aerial photography (Walker and Briggs in press). The source image used is the true-color Landiscor aerial photography (3 m) collected over Phoenix in the spring of 1997.
Continental Splitted Non-IID Image Classification Dataset
<p>A non-IID dataset with images from Flickr and labels from Open Image Dataset. The dataset is split into parts from three continents, North America, Europe, and Asia, between which the data turned out to be non-IID, with same objects looking different. The images in the dataset were cropped from original images with the help of bounding boxes in Open Image Dataset.</p>
Deep Brain Image Contrast Classification
<p>Deep Brain Image Contrast Classification</p>
Deep Brain Image Contrast Classification
<p>Deep Brain Image Contrast Classification</p>
MIHIC: A multiplex IHC histopathological image classification dataset for lung cancer immune microenvironment quantification
<p>A cohort of 47 TMA sections from 114 patients was collected from Liaoning cancer hospital \& Institute, where each TMA section has the size of 188,416$\times$110,080 pixels (i.e., 42660.87um$\times$24924.15um) at 40$\times$ magnification. TMA sections contain different number of tissue cores, ranging from 28 to 48. After excluding poor quality TMA sections with tissue folding, missing or contamination, there are totally 114 patients. Each patient has tissue cores with 12 different IHC stains, including CD3, CD20, CD34, CD38, CD68, CDK4, cyclin-D1, D2-40, FAP, Ki67, P53, and SMA. Two pathologists have manually labeled clear tissue regions (i.e., without controversy) in TMA sections based on visual examination via Qupath software, where six tissue types including Alveoli, Immune cells, Nerosis, Other, Stroma, Tumor were annotated. Besides the annotated six tissue types, we added one more Background type.</p> <p>To build histological classification models, we split 309,698 image patches in MIHIC dataset into three sets: training, validation and test. Note that image patches extracted from the same annotated tissue region are distributed into the same set, which avoids data leakage during classification model optimization. According to the number of extracted ROIs, train, val and test accounted for 64\%, 16\% and 20\%.</p> <h1>if you use this dataset, please cite:</h1> <pre>@article{wang2024mihic, title={MIHIC: a multiplex IHC histopathological image classification dataset for lung cancer immune microenvironment quantification}, author={Wang, Ranran and Qiu, Yusong and Wang, Tong and Wang, Mingkang and Jin, Shan and Cong, Fengyu and Zhang, Yong and Xu, Hongming}, journal={Frontiers in Immunology}, volume={15}, year={2024}, publisher={Frontiers Media SA} }</pre>
multi class dataset_Dipper Throated Optimization with Deep Convolutional Neural Network-based Crop Classification on Remote Sensing Image Analysis
Open the record for dataset details and reuse information.
MosqVision-3K: A Balanced Multi-Source Dataset of 3,000 Annotated Images for Culex, Anopheles, and Aedes Mosquito Species Classification
<p><strong>Comprehensive Mosquito Species Image Dataset for Machine Learning</strong><br><strong>Description:</strong><br>This dataset is a meticulously curated collection of high-quality images featuring three major mosquito species: <strong>Culex</strong>, <strong>Anopheles</strong>, and <strong>Aedes</strong>. These species are significant vectors for transmitting vector-borne diseases such as malaria, dengue, and Zika. The dataset has been compiled to support research and development in entomology, vector-borne disease control, and image recognition.<br>With <strong>3,000 images in total</strong>, the dataset is structured to ensure a balanced representation of the three species, each having <strong>1,000 images</strong>. Images were sourced from four reputable platforms, including <strong>MosquitoAlert.com</strong>, <strong>Mendeley Data</strong>, <strong>IEEE DataPort</strong>, and the <strong>Dryad Digital Repository</strong>. These sources ensure a comprehensive and diverse representation of mosquito appearances, including variations in morphology, lighting conditions, and orientations.</p> <h2>The dataset is organized into directories for each species, making it easy to integrate into machine learning workflows for tasks like species identification and classification. The collection also includes metadata and annotations to enhance usability.</h2> <p><strong>Key Features:</strong></p> <ul> <li><strong>Species Represented:</strong> <ul> <li><em>Culex</em></li> <li><em>Anopheles</em></li> <li><em>Aedes</em></li> </ul> </li> <li><strong>Total Images:</strong> 3,000 (1,000 images per species)</li> <li><strong>Image Sources:</strong> <ul> <li><strong>MosquitoAlert.com</strong> (1,234 images)</li> <li><strong>Mendeley Data</strong> (876 images)</li> <li><strong>IEEE DataPort</strong> (748 images)</li> <li><strong>Dryad Digital Repository</strong> (600 images)</li> </ul> </li> </ul> <h2>- <strong>Image Annotations:</strong> Metadata and species labels are included for enhanced usability.</h2> <p><strong>Applications:</strong><br>This dataset is ideal for a variety of applications, including:</p> <ul> <li>Training machine learning models for mosquito species identification.</li> <li>Developing computer vision algorithms for pest control and public health.</li> </ul> <h2>- Enhancing vector control strategies to mitigate disease spread.</h2> <p><strong>Data Structure:</strong><br>The dataset is organized as follows:</p> <pre><code>Mosquito_Dataset/ ├── Anopheles/ │ ├── img_001<span>.jpg</span> │ ├── img_002<span>.jpg</span> │ └── ... ├── Aedes/ │ ├── img_001<span>.jpg</span> │ ├── img_002<span>.jpg</span> │ └── ... └── Culex/ ├── img_001<span>.jpg</span> ├── img_002<span>.jpg</span> └── ... </code></pre> <h2> </h2> <h2><strong>Acknowledgments:</strong><br>We acknowledge the following data sources for their contributions:</h2> <blockquote> <ul> <li>MosquitoAlert.com</li> <li>Mendeley Data</li> <li>IEEE DataPort</li> <li> Dryad Digital Repository</li> </ul> </blockquote>
Advances in Leukemia Detection and Classification: A Systematic Review of AI and Image Processing Techniques
Open the record for dataset details and reuse information.
Image-based taxonomic classification of bulk biodiversity samples using deep learning and domain adaptation
<p>Complex bulk samples of insects from biodiversity surveys present a challenge for taxonomic identification, which could be overcome by high-throughput imaging combined with machine learning for rapid classification of specimens. These procedures require that taxonomic labels from an existing source data set are used for model training and prediction of an unknown target sample. However, such transfer learning may be problematic for the study of new samples not previously encountered in an image set, e.g. from unexplored ecosystems, and require methods of domain adaptation that reduce the differences in the feature distribution of the source and target domains (training and test sets). We assessed the efficiency of domain adaptation for family-level classification of bulk samples of Coleoptera, as a critical first step in the characterisation of biodiversity samples. Neural network models trained with images from a global database of Coleoptera were applied to a biodiversity sample from understudied forests in Cyprus as the target. Within-dataset classification accuracy reached 98% and depended on the number and quality of training images and on dataset complexity. The accuracy of between-datasets predictions (across disparate source-target pairs that do not share any species or genera) was at most 82% and depended greatly on the standardisation of the imaging procedure. Algorithms for domain adaptation significantly improved the prediction performance of models trained by non-standardised, low-quality images. Our findings demonstrate that existing databases can be used to train models and successfully classify images from unexplored biota, but the imaging conditions and classification algorithms need careful consideration.</p>
FIGURE. Images of representative members of tribe Phyllantheae. (A) Flowers of Nellica maderaspatensis. (B) Pistillate flowers of Cathetus gracilis, note the unique disc covering the ovary. (C) Pistillate and staminate flowers of Cathetus glaucophyllus. (D) Staminate flowers of Nymphanthus glaucescens. (E) Fruits of Kirganelia muelleriana. (F) Fruits of Lysiandra subcrenulata. (G) Phylloclade with flowers of Phyllanthus angustifolius. (H) Flowering branchlet of Phyllanthus incrustatus, note the ornamentation on the axes. (I) Habit of Moeroris tenella. (J) Fruiting branch of Dendrophyllanthus tenuirhachis. (K) Fruits of Cicca profusa. (L) Fruiting branch of Emblica officinalis. (M) Flowering plant of Emblica urinaria. (N) Pistillate flower of Breynia disticha. (O) Staminate flower of Breynia disticha. (P) flower of Glochidion dunnianum. (Q) Staminate of Glochidion lanceolarium. (R) Dehisced capsule of Glochidion sp. showing seeds covered with a red sarcotesta. Photos: A & F by J.J. Bruhl; B & P by M.S. Nuraliev; C by T. Williams; E & K by C. Jongkind; H by B. Falcón; J by R.-Y. Yu; D, G, I, L, M, N, O, Q & R by R.W.Bouman. in A revised phylogenetic classification of tribe Phyllantheae (Phyllanthaceae)
FIGURE. Images of representative members of tribe Phyllantheae. (A) Flowers of Nellica maderaspatensis. (B) Pistillate flowers of Cathetus gracilis, note the unique disc covering the ovary. (C) Pistillate and staminate flowers of Cathetus glaucophyllus. (D) Staminate flowers of Nymphanthus glaucescens. (E) Fruits of Kirganelia muelleriana. (F) Fruits of Lysiandra subcrenulata. (G) Phylloclade with flowers of Phyllanthus angustifolius. (H) Flowering branchlet of Phyllanthus incrustatus, note the ornamentation on the axes. (I) Habit of Moeroris tenella. (J) Fruiting branch of Dendrophyllanthus tenuirhachis. (K) Fruits of Cicca profusa. (L) Fruiting branch of Emblica officinalis. (M) Flowering plant of Emblica urinaria. (N) Pistillate flower of Breynia disticha. (O) Staminate flower of Breynia disticha. (P) flower of Glochidion dunnianum. (Q) Staminate of Glochidion lanceolarium. (R) Dehisced capsule of Glochidion sp. showing seeds covered with a red sarcotesta. Photos: A & F by J.J. Bruhl; B & P by M.S. Nuraliev; C by T. Williams; E & K by C. Jongkind; H by B. Falcón; J by R.-Y. Yu; D, G, I, L, M, N, O, Q & R by R.W.Bouman.
A Dataset Containing Tiny and Low Quality Images for Vehicle Classification
<p>This dataset contains 4800 tiny and low resolution vehicle images. The vehicles in the images are grouped in six classes: Bike, Car, Juggernaut, Minibus, Pickup, and Truck. For each class, there are 800 vehicle images with 100 × 100 pixels and 96 dpi resolution.</p> <p><strong>The peer-reviewed research article for this dataset has been published in MDPI Sensors, and can be accessed here: <a href="https://doi.org/10.3390/s22134740">https://doi.org/10.3390/s22134740</a>. Please cite this when using the dataset.</strong></p>
Colorectal Cancer Histology Image Tiles for Tissue Multi-class Classification
<p><strong>Content</strong></p> <p>The present dataset is linked to a research aimed at discovering the best normalization pipeline and classification model for colorectal cancer multi-class tissue classification.<br> The 15,856 histological image tiles are completely anomized and are extracted from 10 formalin-fized paraffine-embedded samples of patients affected by colorectal cancer.</p> <p>The materials are inside the following zip file:</p> <p>“CRC_Tiles_IRCCS_ISTITUTO_TUMORI_BARI.zip”: a zipped folder containing tiles (n=15,856) annotated by a pathologist, grouped in 6 subdirectories, each of them representing a class. Tiles are of size 224 x 224 px, taken at a resolution of 0.5 μm/px.<br> <br> <strong>Ethical Statement</strong></p> <p>The study has been funded by “Tecnopolo per la Medicina di Precisione (CUP B84I18000540002)”. The institutional Ethic Committee approved the study (Prot n. 780/CE).</p> <p><br> <strong>Related Datasets and Works</strong><br> <br> For further details concerning the aforementioned dataset, refer to the papers below. <br> Please cite the following articles if you need this dataset for your research.</p> <p>Altini N. et al. (2021) Multi-class Tissue Classification in Colorectal Cancer with Handcrafted and Deep Features. In: Huang DS., Jo KH., Li J., Gribova V., Bevilacqua V. (eds) Intelligent Computing Theories and Application. ICIC 2021. Lecture Notes in Computer Science, vol 12836. Springer, Cham. <br> https://doi.org/10.1007/978-3-030-84522-3_42</p> <p>Altini, N., Marvulli, T. M., Zito, F. A., Caputo, M., Tommasi, S., Azzariti, A., ... & Bevilacqua, V. (2023). The Role of Unpaired Image-to-Image Translation for Stain Color Normalization in Colorectal Cancer Histology Classification. <em>Computer Methods and Programs in Biomedicine</em>, 107511. <br> <a href="https://doi.org/10.1016/j.cmpb.2023.107511">https://doi.org/10.1016/j.cmpb.2023.107511</a></p> <p>Please also consider the dataset offered in our previous work:</p> <p>Altini N. et al. (2021). Pathologist's Annotated Image Tiles for Multi-Class Tissue Classification in Colorectal Cancer (v1.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.4785131</p>
GradDA – A novel dataset for investigating domain shifts in image classification
<p>A domain shift occurs when the testing data is drawn from a distribution different from that of the training dataset. This shift presents a significant challenge and may compromise the performance of machine learning models, which leads to poor generalization. Over the past years, various models have been developed and evaluated on benchmark datasets such as VisDA, Office-Home and DomainNet. These datasets consist of discrete domains with different object classes. However, a notable limitation when addressing the domain shift is the absence of data samples where the exact same object exists in both domains. </p> <p>We propose a new dataset designed to address this challenge. In particular, we introduce a domain shift from a purely synthetic style (grey object on white background) to a more realistic appearance (object with texture against a realistic background) with differential modifications, which enables the representation of the same object in both synthetic and real domains, consequently facilitating the analysis of a transition between the two domains. The dataset comprises five distinct classes (Airplane, Bicycle, Bus, Car, Train), with multiple objects per class. Additionally, each object is depicted from 20 different perspectives, resulting in a total of 101 images per perspective that captures the transition from pure synthetic to a more real-world-like domain. This dataset offers a unique opportunity to investigate the impact of domain shift on model performance in classification tasks, as it focuses solely on domain changes without other interfering effects. It is the objective of our work to trigger new discussions about the domain shift problem, and how it can be tackled with alternative data driven model designs.</p>
Contrastive Learning for Fine-Grained Ship Classification in Remote Sensing Images
<p>Dataset for Contrastive Learning for Fine-Grained Ship Classification in Remote Sensing Images from https://github.com/WindVChen/Push-and-Pull-Network?tab=readme-ov-file</p>
Tongue image dataset for Tri-Dhat classification in traditional Thai medicine
<p>Traditional Thai medicine (TTM) is an increasingly popular treatment option. Tongue diagnosis is a highly efficient method for determining overall health, as practiced by TTM practitioners. However, the diagnosis naturally varies depending on the practitioner's expertise. In this work, we propose tongue image analysis with raw pixels using artificial intelligence (AI) to support TTM diagnoses. The target classification of Tri-Dhat consists of three classes: Vata, Pitta, and Kapha. We utilized our own organized, genuine datasets collected from our university TTM hospital. Class balancing and data augmentation were conducted, and we present analysis approaches and experimental designs. Transfer learning techniques for various pretrained deep learning models were developed. We used two-tailed paired t-tests and single-factor ANOVA for performance comparisons. Our work demonstrated that the DenseNet121 and Xception models provided the most significant results with cropped image datasets, including DSLR-taken and mobile-taken images. Notably, model ensemble evaluations yielded the highest average predictions, achieving a precision of 0.94, an F1 score of 0.96, an accuracy of 0.96, a sensitivity of 0.96, and a specificity of 0.97, supported by a p-value of 0.0003 from ANOVA. We suggest that our methods could be effectively deployed in real-world scenarios to aid TTM practitioners in their diagnoses.</p>
Histological images for MSI vs. MSS classification in gastrointestinal cancer, FFPE samples
<p>This repository contains 411,890 unique image patches derived from histological images of colorectal cancer and gastric cancer patients in the TCGA cohort (original whole slide SVS images are freely available at https://portal.gdc.cancer.gov/). All images in this repository are derived from formalin-fixed paraffin-embedded (FFPE) diagnostic slides ("DX" at the GDC data portal). This is explained well in this blog: http://www.andrewjanowczyk.com/download-tcga-digital-pathology-images-ffpe/</p> <p><strong>Preprocessing</strong></p> <p>All SVS slides were preprocessed as follows</p> <p>1. automatic detection of tumor</p> <p>2. resizing to 224 px x 224 px at a resolution of 0.5 µm/px</p> <p>4. color normalization with the Macenko method (Macenko et al., 2009, http://wwwx.cs.unc.edu/~mn/sites/default/files/macenko2009.pdf)</p> <p>5. assignment of patients to either "MSS" (microsatellite stable) or "MSIMUT" (microsatellite instable or highly mutated)</p> <p>6. randomization of patients to training and testing sets (~70% and ~30%). Randomization was done on a patient level rather than on a slide or tile level</p> <p>7. equilibration of training sets by undersampling (removing excess tiles in MSS class in a random way)</p> <p><strong>File description</strong></p> <p>1. STAD_TRAIN_MSS - training images (~70% of all patients) for gastric (stomach) cancer TCGA patients with MSS (microsatellite stable) tumors, 50285 unique image patches; FFPE samples</p> <p>2. STAD_TRAIN_MSIMUT - training images ( (~70% of all patients) for gastric (stomach) cancer TCGA patients with MSI (microsatellite instable) or highly mutated tumors, 50285 unique image patches; FFPE samples</p> <p>3. STAD_TEST_MSS - test images (~30% of all patients) for gastric (stomach) cancer TCGA patients with MSS (microsatellite stable) tumors, 90104 unique image patches; FFPE samples</p> <p>4. STAD_TEST_MSIMUT - test images ( ~30% of all patients) for gastric (stomach) cancer TCGA patients with MSI (microsatellite instable) or highly mutated tumors, 27904 unique image patches; FFPE samples</p> <p>5. CRC_DX_TEST_MSIMUT - test images (~30% of all patients) for colorectal cancer TCGA patients with MSI (microsatellite instable) or highly mutated tumors, 29335 unique image patches; FFPE samples</p> <p>6. CRC_DX_TEST_MSS - test images (~30% of all patients) for colorectal cancer TCGA patients with MSS (microsatellite stable) tumors, 70569 unique image patches; FFPE samples</p> <p>7. CRC_DX_TRAIN_MSIMUT - training images (~70% of all patients) for colorectal cancer TCGA patients with MSI (microsatellite instable) or highly mutated tumors, 46704 unique image patches; FFPE samples</p> <p>8. CRC_DX_TRAIN_MSS - training images (~70% of all patients) for colorectal cancer TCGA patients with MSS (microsatellite stable) tumors, 46704 unique image patches; FFPE samples</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.