Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
166
datasets available to search
ShareScore release 0.9.0
Dataset results
166 results for “Image Classification”
Example computer vision classification training data derived from British Library 19th Century Books Image collection
<p>Example computer vision classification training data derived from British Library 19th Century Books Image collection</p> <p>This dataset provides training data for image classification for use in a computer vision workshop. The images are derived from '<a href="https://doi.org/10.21250/db17">Digitised Books - Images identified as Embellishments. c. 1510 - c. 1900. JPG'</a> from the year '1839'.</p> <p>Currently, included are four folders containing a variety of images derived from the BL books corpus.</p> <ul> <li>'cv_workshop_exercise_data' include images of: 'building', 'people', 'coat of arms'</li> <li>'humancats' contains images of humans and images of cats</li> </ul> <p>The 'fashion' and 'portraits' folders both contain images of people organised into 'female' and 'male'. These labels were annotated by a single annotator and these categories may themselves not be meaningful. They are included in the workshop data as a point of discussion about how we should label data both in general and when working with historical data. </p> <p>This data is intended primarily as an educational resource.</p>
Data for "Current challenges for unseen-epitope TCR interaction prediction and a new perspective derived from image classification" (ImRex)
<p>Repository containing the different experiments described in the manuscript titled: "Current challenges for epitope-agnostic TCR interaction prediction and a new perspective derived from image classification".</p> <p>Publication DOI: TBA</p> <p>Originally appeared as a preprint on bioRxiv: <a href="https://doi.org/10.1101/2019.12.18.880146">https://doi.org/10.1101/2019.12.18.880146</a>.</p> <p>Contains:</p> <ul> <li>Trained model files (.h5)</li> <li>Associated train and validation datasets for each model.</li> <li>Learning curves and evaluation metrics.</li> <li>Log files with training and data arguments (full training scripts are available in GitHub repository).</li> <li>Comparisons between different models.</li> <li>Complete raw and processed datasets (also available in the associated GitHub repository).</li> </ul> <p><strong>Please refer to the associated GitHub repository (<a href="https://github.com/pmoris/ImRex">https://github.com/pmoris/ImRex</a>) for more information on the directory structure and contents, as well as the scripts that generated these output files.</strong></p> <p><strong>Contents:</strong></p> <ul> <li> <p><code>data.zip</code>: Contains raw and preprocessed datasets. READMEs in subdirectory describe the data sources and preprocessing steps. Please refer to the associated GitHub repository for the specific scripts that generated these files. Note that the full training and test sets (i.e. containing both positive and negative examples) are stored separately for each model/CV iteration in the <code>models</code> archives.</p> </li> <li> <p><code>models-main.zip</code>: contains the trained models and evaluation metrics for the main different experiments described in the bash and pbs scripts in <code>./src/scripts/hpc_scripts</code>. Log files for the experiments outlined here can be found in <code>./src/scripts/hpc_scripts</code>.</p> </li> <li> <p><code>models-full.zip</code>: contains models that were trained on the complete VDJdb dataset without cross-validation, filtered on human TRB data, no 10x data and restricted to 10-20 (CDR3) or 8-11 (epitope) amino acid residues, with negatives that were generated by shuffling (i.e. sampling an negative epitope for each positive CDR3 sequence). One set of models uses downsampling to reduce the most abundant epitopes down to 400 pairs each, the other one does not use any downsampling. These models were also used for evaluating on the external Adaptive dataset, as outlined in <code>./src/scripts/evaluate/evaluate_adaptive.sh</code>, and the TRA subset of sequences (<code>./src/scripts/evaluate/evaluate_tra.sh</code>).</p> </li> <li> <p><code>models-decoyfit.zip</code>: contains models that were trained on true data, but evaluated on data where epitopes were replaced by decoys.</p> </li> <li> <p><code>models-padded-epitoperatio.zip</code>: contains a quick test of trained models (padded/interaction map) that use a different type of negative shuffling, see docstrings in <code>./src/processing/negative_sampler.py</code> for more info.</p> </li> <li> <p><code>models-repeat-local.zip</code>: contains a number of repeated runs from <code>models-main</code>, used to estimate variability in model performance for multiple identical runs.</p> </li> <li> <p><code>comparisons.zip</code>: contains comparison directories, each consisting of two or more model output directories, that contrast the performance metrics of the models. These outputs were generated by using the <code>./src/scripts/evaluate/visualize.py</code> script, or by using the oneliners in <code>./src/scripts/evaluate/visualise.sh</code>, which can operate on the entire comparisons directory at once.</p> </li> </ul> <p><strong>Note that any file paths described here are in reference to the associated GitHub repository (<a href="https://github.com/pmoris/ImRex">https://github.com/pmoris/ImRex</a>).</strong></p> <p><strong>Overview of different experiments:</strong></p> <ul> <li>Two main architectures were compared: the interaction map (or <code>padded</code>) CNN and a dual input CNN based on NetTCR (<code>nettcr</code>).</li> <li>Two different cross-validation strategies were used: a 5x repeated 5-fold CV (<code>repeated5fold</code>) and an epitope-grouped CV (<code>epitope_grouped</code>).</li> <li>The different dataset subsets are labelled as follows. Check the Makefile's <code>preprocess-vdjdb-aug-2019</code> command (and the underlying script <code>./src/scripts/preprocessing/preprocess_vdjdb.py</code>) for a more thorough overview of the different filtering options. <ul> <li><code>mhci</code>: only MHCI class presented epitopes.</li> <li><code>trb</code>: only TRB CDR3 sequences.</li> <li><code>tra</code>: only TRA CDR3 sequences.</li> <li><code>tratrb</code>: both types of CDR3 sequences.</li> <li><code>down</code>: moderate downsampling of most abundant epitopes to 1000 pairs.</li> <li><code>down400</code>: strong downsampling of most abundant epitopes to 400 pairs.</li> <li><code>decoy</code>: decoy epitope data.</li> <li><code>reg001</code>: regularization factor 0.01 (only for padded/interaction type models, fixed value)</li> </ul> </li> <li>Two different methods of generating negative TCR-epitope pairs were used: shuffling of positive pairs, i.e. sampling a single epitope from the positive pairs for each CDR3 sequence (<code>shuffle</code>), and sampling CDR3s from a reference repertoire (<code>negref</code>).</li> <li>The batch size is labelled as <code>b32</code> = a batch size of 32.</li> <li>The learning rate was always 0.0001 (<code>lre4</code>) or 0.001 (<code>lre3</code>).</li> </ul> <p> </p>
Remote sensing image classification dataset
Open the record for dataset details and reuse information.
Dipper Throated Optimization with Deep Convolutional Neural Network-based Crop Classification on Remote Sensing Image Analysis
Open the record for dataset details and reuse information.
Combining Natural Language and Images for Garbage Classification: A Public Benchmark
Open the record for dataset details and reuse information.
Image classification in Galaxy with fruit 360 dataset
<p>Credit: 'Fruit recognition from images using deep learning' by H. Muresan and M. Oltean (<a href="https://arxiv.org/abs/1712.00580">https://arxiv.org/abs/1712.00580</a>)<br> <br> Fruit 360 is a dataset with 90380 images of 131 fruits and vegetables (<a href="https://www.kaggle.com/moltean/fruits">https://www.kaggle.com/moltean/fruits</a>). Images are 100 pixel by 100 pixel and are RGB (color) images (3 values for each pixel). This dataset is a subset of Fruit 360 dataset, containing only 10 fruits/vegetables (Strawberry, Apple_Red_Delicious, Pepper_Green, Corn, Banana, Tomato_1, Potato_White, Pineapple, Orange, and Peach). We selected a subset of fruits/vegetables, so the dataset size is smaller and the neural network can be trained faster.</p> <p> </p> <p>The utilities used to create the dataset, along with step by step instructions, can be found here: https://github.com/kxk302/fruit_dataset_utilities<br> <br> First, we created feature vectors for each image. Each image is 100 pixel by pixel and are RGB (color) images (3 values for each pixel). Hence, each image can be represented by 30,000 values (100 X 100 X 3). Second, we selected a subset of 10 fruits/vegetables images (training and test dataset sizes go from 7 GB and 2.5 GB for 131 fruits/vegetables to 500 MB and 177 MB for 10 fruits/vegetables, respectively). Third, we created separate files for feature vectors and labels. Finally, we mapped the labels for the 10 selected fruits/vegetables to a range of 0 to 9.</p> <p> </p>
A Deep Learning-Based Approach for Efficient Detection and Classification of Local Ca²⁺ Release Events in Full-Frame Confocal Imaging
<p>U-Net implementation and project code: <a href="https://github.com/dottipr/sparks_project">https://github.com/dottipr/sparks_project</a></p> <p>GUI's github repository: <a href="https://github.com/r-janicek/xytCalciumSignalsDetection">https://github.com/r-janicek/xytCalciumSignalsDetection</a></p>
Detection, quantification and classification of ripened tomatoes: a comparative analysis of image processing and machine learning
<p>This is an open dataset.</p>
Detection, quantification and classification of ripened tomatoes: a comparative analysis of image processing and machine learning
<p>In this study, specifically for the detection of ripe/unripe tomatoes with/without defects in the crop field, two distinct methods are described and compared from captured images by a camera mounted on a mobile robot. One is a machine learning approach, known as 'Cascaded Object Detector' (COD) and the other is a composition of traditional customised methods, individually known as 'Colour Transformation': 'Colour Segmentation' and 'Circular Hough Transformation'. The (Viola-Jones) COD generates 'histogram of oriented gradient' (HOG) features to detect tomatoes. For ripeness checking, the RGB mean is calculated with a set of rules. However, for traditional methods, colour thresholding is applied to detect tomatoes either from natural or solid background and RGB colour is adjusted to identify ripened tomatoes. This algorithm is shown to be optimally feasible for any micro-controller based miniature electronic devices in terms of its run time complexity of <i>O</i>(<i>n</i><sup>3</sup>) for a traditional method in best and average cases. Comparisons show that the accuracy of the machine learning method is 95%, better than that of the Colour Segmentation Method using MATLAB.</p>
Data from MODIS images classification
<p>This data is related to MODIS images classification in Ghana (Guinea-savannah and Forest-savannah)</p>
IMAging With Opto-acoustics to downgradE BI-RADS claSsificaTion Relative tO Other Diagnostic Methodologies (MAESTRO)
ClinicalTrials.gov study NCT02364388. IPD Sharing: Not stated. Countries: 1. Publications: 0.
Detection, quantification and classification of ripened tomatoes: a comparative analysis of image processing and machine learning
Open the record for dataset details and reuse information.
Data from: Cross-sectional study of patients with axial spondyloarthritis fulfilling imaging arm of ASAS classification criteria: baseline clinical characteristics and subset differences in a single centre cohort
Open the record for dataset details and reuse information.
BOREAS TE-18 Landsat TM Maximum Likelihood Classification Image of the SSA
A Landsat-5 TM image from 06-Aug-1990 was used to derive this classification. The objective of this classification is to provide the BOREAS investigators with a data product that characterizes the land cover of the SSA. A standard supervised maximum likelihood approach was used to produce this classification. Companion files include example thumbnail images that may be viewed using a convenient viewer utility.
BOREAS TE-18 Landsat TM Maximum Likelihood Classification Image of the NSA
The objective of this classification is to provide the BOREAS investigators with a data product that characterizes the land cover of the NSA. A Landsat-5 TM image from 20-Aug-1988 was used to derive this classification. A standard supervised maximum likelihood approach was used to produce this classification. Companion files include example thumbnail images that may be viewed using a convenient viewer utility.
BOREAS TE-18 Landsat TM Physical Classification Image of the NSA
The objective of this classification is to provide the BOREAS investigators with a data product that characterizes the land cover of the NSA. A Landsat-5 TM image from 21-Jun-1995 was used to derive the classification. A technique was implemented that uses reflectances of various land cover types along with a geometric optical canopy model to produce spectral trajectories. These trajectories are used as training data to classify the image into the different land cover classes. Companion files include example thumbnail images that may be viewed and using a convenient viewer utility.
BOREAS TE-18 Landsat TM Physical Classification Image of the SSA
The objective of this classification is to provide the BOREAS investigators with a data product that characterizes the land cover of the SSA. A Landsat-5 TM image from 02-Sep-1994 was used to derive the classification. A technique was implemented that uses reflectances of various land cover types along with a geometric optical canopy model to produce spectral trajectories. These trajectories are used as training data to classify the image into the different land cover classes. These data are provided in a binary image file format. Companion files include example thumbnail images that may be viewed and the image data files downloaded using a convenient viewer utility.
Image datasets for jammer classification in GNSS
<p>Set of spectrogram binary images corresponding to GNSS signal with and without interference in a variety of scenarios and signal parameters. It also includes two Matlab functions to perform the analisys. </p>
Datasets corresponding to publication: AIDeveloper: deep learning image classification in life science and beyond
<p>Datasets and videos corresponding to publication:<br> AIDeveloper: deep learning image classification in life science and beyond</p>
Parasitic Egg Detection and Classification in Microscopic Images
<div> <p>Parasitic infections have been recognised as one of the most significant causes of illnesses by WHO. Most infected persons shed cysts or eggs in their living environment, and unwittingly cause transmission of parasites to other individuals. Diagnosis of intestinal parasites is usually based on direct examination in the laboratory, of which capacity is obviously limited. Targeting to automate routine faecal examination for parasitic diseases, this challenge aims to gather experts in the field to develop robust automated methods to detect and classify eggs of parasitic worms in a variety of microscopic images. Participants will work with a large-scale dataset, containing 11 types of parasitic eggs from faecal smear samples. They are the main interest because of causing major diseases and illness in developing countries. We open to any techniques used for parasitic egg recognition, ranging from conventional approaches based on statistical models to deep learning techniques. Finally, the organisers expect a new collaboration come out from the challenge.</p> </div> <h4>Instructions:</h4> <div> <p>Datasets contain 11 parasitic egg types. Each category has 1,000 images.</p> <ul> <li>category_id 0: Ascaris lumbricoides</li> <li>category_id 1: Capillaria philippinensis</li> <li>category_id 2: Enterobius vermicularis</li> <li>category_id 3: Fasciolopsis buski</li> <li>category_id 4: Hookworm egg</li> <li>category_id 5: Hymenolepis diminuta</li> <li>category_id 6: Hymenolepis nana</li> <li>category_id 7: Opisthorchis viverrine</li> <li>category_id 8: Paragonimus spp</li> <li>category_id 9: Taenia spp. egg</li> <li>category_id 10: Trichuris trichiura</li> </ul> <p>Please visit the Challenge Homepage (https://icip2022challenge.piclab.ai/).</p> <p>Results must be submitted to the Leaderboard at the Challenge Homepage (https://icip2022challenge.piclab.ai/submission/).</p> <p>Please cite our paper for the usage after the competition: <a href="https://arxiv.org/abs/2208.06063" target="_blank" rel="noopener">N. Anantrasirichai, T. H. Chalidabhongse, D. Palasuwan, K. Naruenatthanaset, T. Kobchaisawat, N. Nunthanasup, K. Boonpeng, X. Ma and A. Achim, "ICIP 2022 Challenge on Parasitic Egg Detection and Classification in Microscopic Images: Dataset, Methods and Results," IEEE ICIP2022</a>.</p> </div>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.