Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
28
datasets available to search
ShareScore release 0.9.0
Dataset results
28 results for “supervised classification”
Dataset for Insect Lidar Supervised Classification
<p>This repository contains the data used for our paper Detection of Insects in Class-imbalanced Lidar Field Measurements, which was published in and presented at the 2021 IEEE Machine Learning for Signal Processing conference.</p> <p>The code is archived on Zenodo at <a href="https://doi.org/10.5281/zenodo.5504408">https://doi.org/10.5281/zenodo.5504408</a>. The code can also be found on Github at <a href="https://github.com/BMW-lab-MSU/insect-lidar-supervised-classification/tree/mlsp-2021">https://github.com/BMW-lab-MSU/insect-lidar-supervised-classification/tree/mlsp-2021</a>.</p>
Multi-granular Software Classification using File-Level Weak Supervision - Data
<p>Data for our paper: Multi-granular Software Classification using File-Level Weak Supervision</p>
Data from: Phenotype classification of zebrafish embryos by supervised learning
Open the record for dataset details and reuse information.
Self-supervised retinal thickness prediction enables deep learning from unlabeled data to boost classification of diabetic retinopathy
<p><strong>This data repository contains the OCT images and binary annotations for segmentation of retinal tissue using deep learning. To use, please refer to the Github repository </strong><a href="https://github.com/theislab/DeepRT">https://github.com/theislab/DeepRT</a>.</p> <p> </p> <p><strong>#######</strong></p> <p><strong>Access to large, annotated samples represents a considerable challenge for training accurate deep-learning models in medical imaging. While current leading-edge transfer learning from pre-trained models can help with cases lacking data, it limits design choices, and generally results in the use of unnecessarily large models. We propose a novel, self-supervised training scheme for obtaining high-quality, pre-trained networks from unlabeled, cross-modal medical imaging data, which will allow for creating accurate and efficient models. We demonstrate this by accurately predicting optical coherence tomography (OCT)-based retinal thickness measurements from simple infrared (IR) fundus images. Subsequently, learned representations outperformed advanced classifiers on a separate diabetic retinopathy classification task in a scenario of scarce training data. Our cross-modal, three-staged scheme effectively replaced 26,343 diabetic retinopathy annotations with 1,009 semantic segmentations on OCT and reached the same classification accuracy using only 25% of fundus images, without any drawbacks, since OCT is not required for predictions. We expect this concept will also apply to other multimodal clinical data-imaging, health records, and genomics data, and be applicable to corresponding sample-starved learning problems.</strong></p> <p><strong>#######</strong></p>
Self Supervised Cloud Classification Data Supplement
<p>Supplementary data for: "Self Supervised Cloud Classification" Geiss et al. 2023.</p> <p>Includes trained neural networks and a copy of the image chips extracted from the "SGFF" dataset that were used for evaluating the trained networks.</p> <p>Update 9-16-2025: Uploaded the correct 'modis_encoder' file. Previously the 'modis_projector' file was mistakenly uploaded as both the encoder and projector.</p>
Data for "LAND COVER CLASSIFICATION FROM A MAPPING PERSPECTIVE: PIXELWISE SUPERVISION IN THE DEEP LEARNING ERA"
<p><strong>Contents</strong></p> <ul> <li>clc_maps.zip contains the dataset.</li> <li>LICENSE.txt describes the usage terms of the maps.</li> </ul> <p>The maps contained in clc_maps.zip follow the naming convention of BigEarthNet [1], i.e. each sample of BigEarthNet has a corresponding pixel-level label map in the dataset.</p> <p>[1] G. Sumbul, M. Charfuelan, B. Demir, and V. Markl, “Bigearthnet: A large-scale benchmark archive for remote sensing image understanding,” in IGARSS 2019-2019 IEEE International Geoscience and Remote Sensing Symposium. IEEE, 2019, pp. 5901–5904.</p> <p><strong>Description</strong></p> <p>The original shape file (<a href="https://land.copernicus.eu/pan-european/corine-land-cover/clc2018?tab=download">link</a>) was altered by reprojecting the shape file onto the coordinate reference system (CRS) of the respective BigEarthNet sample images to ensure pixel synchronicity. Afterwards, the shapes present in the sample CRS are rasterized by burning a linearly increasing class index which replaces the textual CLC nomenclature. The class IDs and their corresponding class names are presented in the following section. </p> <p><strong>Classes</strong></p> <p>Class ID - Corine Land Cover 2018 class name<br> 1 - Continuous urban fabric<br> 2 - Discontinuous urban fabric<br> 3 - Industrial or commercial units<br> 4 - Road and rail networks and associated land<br> 5 - Port areas<br> 6 - Airports<br> 7 - Mineral extraction sites<br> 8 - Dump sites<br> 9 - Construction sites<br> 10 - Green urban areas<br> 11 - Sport and leisure facilities<br> 12 - Non-irrigated arable land<br> 13 - Permanently irrigated land<br> 14 - Rice fields<br> 15 - Vineyards<br> 16 - Fruit trees and berry plantations<br> 17 - Olive groves<br> 18 - Pastures<br> 19 - Annual crops associated with permanent crops<br> 20 - Complex cultivation patterns<br> 21 - Land principally occupied by agriculture, with significant areas of natural vegetation<br> 22 - Agro-forestry areas<br> 23 - Broad-leaved forest<br> 24 - Coniferous forest<br> 25 - Mixed forest<br> 26 - Natural grasslands<br> 27 - Moors and heathland<br> 28 - Sclerophyllous vegetation<br> 29 - Transitional woodland-shrub<br> 30 - Beaches, dunes, sands<br> 31 - Bare rocks<br> 32 - Sparsely vegetated areas<br> 33 - Burnt areas<br> 34 - Glaciers and perpetual snow<br> 35 - Inland marshes<br> 36 - Peat bogs<br> 37 - Salt marshes<br> 38 - Salines<br> 39 - Intertidal flats<br> 40 - Water courses<br> 41- Water bodies<br> 42 - Coastal lagoons<br> 43 - Estuaries<br> 44 - Sea and ocean<br> 48 - NODATA<br> 49 - UNCLASSIFIED LAND SURFACE<br> 50 - UNCLASSIFIED WATER BODIES </p> <p>More details about the CLC classes and conventions can be found in the CLC nomenclature guide (<a href="https://land.copernicus.eu/user-corner/technical-library/corine-land-cover-nomenclature-guidelines/html">Link</a>).</p> <p><strong>Attribution</strong></p> <p>If you find this work useful please consider citing:</p> <p>Wilhelm, T.; Koßmann, D. LAND COVER CLASSIFICATION FROM A MAPPING PERSPECTIVE: PIXELWISE SUPERVISION IN THE DEEP LEARNING ERA. In Proceedings of the IGARSS 2021—2021 IEEE International Geoscience and Remote Sensing Symposium, Brussels, Belgium, 12 – 16 July 2021; to appear.</p> <p><strong>License</strong></p> <p>The generated maps are based on data from the Copernicus program, which are subject to the terms described here:<br> <a href="https://land.copernicus.eu/pan-european/corine-land-cover/clc2018?tab=metadata">https://land.copernicus.eu/pan-european/corine-land-cover/clc2018?tab=metadata</a></p>
SECANT: a biology-guided semi-supervised method for clustering, classification, and annotation of single-cell multi-omics
GEO Series GSE168264. Homo sapiens. 4 samples. Type: Expression profiling by high throughput sequencing; Other.
MULTI-LABEL ASRS DATASET CLASSIFICATION USING SEMI-SUPERVISED SUBSPACE CLUSTERING
MULTI-LABEL ASRS DATASET CLASSIFICATION USING SEMI-SUPERVISED SUBSPACE CLUSTERING MOHAMMAD SALIM AHMED, LATIFUR KHAN, NIKUNJ OZA, AND MANDAVA RAJESWARI Abstract. There has been a lot of research targeting text classification. Many of them focus on a particular characteristic of text data - multi-labelity. This arises due to the fact that a document may be associated with multiple classes at the same time. The consequence of such a characteristic is the low performance of traditional binary or multi-class classification techniques on multi-label text data. In this paper, we propose a text classification technique that considers this characteristic and provides very good performance. Our multi-label text classification approach is an extension of our previously formulated [3] multi-class text classification approach called SISC (Semi-supervised Impurity based Subspace Clustering). We call this new classification model as SISC-ML(SISC Multi-Label). Empirical evaluation on real world multi-label NASA ASRS (Aviation Safety Reporting System) data set reveals that our approach outperforms state-of-theart text classification as well as subspace clustering algorithms.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.