Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
101
datasets available to search
ShareScore release 0.9.0
Dataset results
101 results for “Supervised learning”
Raw Data for Publication "Using Supervised Learning to Classify Metadata of Research Data by Field of Study"
<p>Automated classification of metadata of research data by their field of study can be used in scientometric research, by repository service providers, and in the context of research data aggregation services. To evaluate different machine learning approaches, data from the DataCite index have been downloaded in May 2019 with a GeRDI harvester (filtering out any metadata without a qualified subject, i.e. a subject with either a subjectName or a subjectURI) . These is the resulting raw data set.</p>
Audio-visual self-supervised learning
<p>CVPR 2021 tutorial</p>
Data from: Phenotype classification of zebrafish embryos by supervised learning
Zebrafish is increasingly used to assess biological properties of chemical substances and thus is becoming a specific tool for toxicological and pharmacological studies. The effects of chemical substances on embryo survival and development are generally evaluated manually through microscopic observation by an expert and documented by several typical photographs. Here, we present a methodology to automatically classify brightfield images of wildtype zebrafish embryos according to their defects by using an image analysis approach based on supervised machine learning. We show that, compared to manual classification, automatic classification results in 90 to 100% agreement with consensus voting of biological experts in nine out of eleven considered defects in 3 days old zebrafish larvae. Automation of the analysis and classification of zebrafish embryo pictures reduces the workload and time required for the biological expert and increases the reproducibility and objectivity of this classification.
Semi-supervised learning for sensitive open modification spectral library searching - dataset PXD009476
<p>This is the analysis results of ANN-SoLo + Rescoring as an integrated module. ANN-SoLo spectral library search engine is a tool for efficient open modification searching. ANN-SoLo uses a cascade search strategy to optimally identify both unmodified and modified peptides: in the first stage a standard search is performed to identify unmodified peptides, after which the remaining unidentified spectra are submitted to second stage during which an open search is performed to additionally identify modified peptides. We have augmented this approach by natively integrating PSM rescoring into ANN-SoLo using the mokapot Python framework for semi-supervised learning for peptide detection.</p> <p>The repository includes the results and search database used to analyze a human glycoproteomics dataset that was acquired from human kidney tissue, serum, and T cells to study O-linked glycosylation. Samples were trypsin-digested and separated into 24 fractions after enrichment of intact glycopeptides and release of O-linked glycopeptides. Next, the data was acquired on a Fusion Lumos mass spectrometer with an Easy-nLC 1200 system or a Q-Exactive HF mass spectrometer with a Waters NanoAcquity UPLC. From this dataset, four raw files derived from kidney tissue samples were retrieved from PRIDE (project PXD009476) and converted to MGF files using ThermoRawFileParser (version 1.7.2).</p> <p>The four files are available both in RAW and MGF format below. </p> <p> </p> <ul> <li>Arab, Issar, William E. Fondrie, Kris Laukens, and Wout Bittremieux. "<strong>Semisupervised Machine Learning for Sensitive Open Modification Spectral Library Searching</strong>." Journal of Proteome Research 22, no. 2 (2023): 585-593. <a title="DOI URL" href="https://doi.org/10.1021/acs.jproteome.2c00616">doi.org/10.1021/acs.jproteome.2c00616</a></li> </ul>
A Pilot Trial of Remotely-Supervised Transcranial Direct Current Stimulation (RS-tDCS) to Enhance Motor Learning in Progressive Multiple Sclerosis (MS)
ClinicalTrials.gov study NCT03499314. IPD Sharing: Not stated. Countries: 1. Publications: 0.
Data from: Phenotype classification of zebrafish embryos by supervised learning
Open the record for dataset details and reuse information.
Data from: ZeitZeiger: supervised learning for high-dimensional data from an oscillatory system
Open the record for dataset details and reuse information.
Solo: doublet identification via semi-supervised deep learning
GEO Series GSE140262. Mus musculus. 4 samples. Type: Expression profiling by high throughput sequencing.
Self-supervised learning for predicting transcriptomic groups on whole slides images in intrahepatic cholangiocarcinoma
GEO Series GSE244807. Homo sapiens. 246 samples. Type: Expression profiling by high throughput sequencing.
Self-supervised retinal thickness prediction enables deep learning from unlabeled data to boost classification of diabetic retinopathy
<p><strong>This data repository contains the OCT images and binary annotations for segmentation of retinal tissue using deep learning. To use, please refer to the Github repository </strong><a href="https://github.com/theislab/DeepRT">https://github.com/theislab/DeepRT</a>.</p> <p> </p> <p><strong>#######</strong></p> <p><strong>Access to large, annotated samples represents a considerable challenge for training accurate deep-learning models in medical imaging. While current leading-edge transfer learning from pre-trained models can help with cases lacking data, it limits design choices, and generally results in the use of unnecessarily large models. We propose a novel, self-supervised training scheme for obtaining high-quality, pre-trained networks from unlabeled, cross-modal medical imaging data, which will allow for creating accurate and efficient models. We demonstrate this by accurately predicting optical coherence tomography (OCT)-based retinal thickness measurements from simple infrared (IR) fundus images. Subsequently, learned representations outperformed advanced classifiers on a separate diabetic retinopathy classification task in a scenario of scarce training data. Our cross-modal, three-staged scheme effectively replaced 26,343 diabetic retinopathy annotations with 1,009 semantic segmentations on OCT and reached the same classification accuracy using only 25% of fundus images, without any drawbacks, since OCT is not required for predictions. We expect this concept will also apply to other multimodal clinical data-imaging, health records, and genomics data, and be applicable to corresponding sample-starved learning problems.</strong></p> <p><strong>#######</strong></p>
m-sized Training and Evaluation Data for Publication "Using Supervised Learning to Classify Metadata of Research Data by Field of Study"
<p>Automated classification of metadata of research data by their discipline(s) of research can be used<br> in scientometric research, by repository service providers, and in the context of research data aggregation<br> services. Openly available metadata of the DataCite index for research data were used to compile a large<br> training and evaluation set comprised of 609,524 records. This is the cleaned and vectorized version with a feature selection of medium size</p>
Data for "LAND COVER CLASSIFICATION FROM A MAPPING PERSPECTIVE: PIXELWISE SUPERVISION IN THE DEEP LEARNING ERA"
<p><strong>Contents</strong></p> <ul> <li>clc_maps.zip contains the dataset.</li> <li>LICENSE.txt describes the usage terms of the maps.</li> </ul> <p>The maps contained in clc_maps.zip follow the naming convention of BigEarthNet [1], i.e. each sample of BigEarthNet has a corresponding pixel-level label map in the dataset.</p> <p>[1] G. Sumbul, M. Charfuelan, B. Demir, and V. Markl, “Bigearthnet: A large-scale benchmark archive for remote sensing image understanding,” in IGARSS 2019-2019 IEEE International Geoscience and Remote Sensing Symposium. IEEE, 2019, pp. 5901–5904.</p> <p><strong>Description</strong></p> <p>The original shape file (<a href="https://land.copernicus.eu/pan-european/corine-land-cover/clc2018?tab=download">link</a>) was altered by reprojecting the shape file onto the coordinate reference system (CRS) of the respective BigEarthNet sample images to ensure pixel synchronicity. Afterwards, the shapes present in the sample CRS are rasterized by burning a linearly increasing class index which replaces the textual CLC nomenclature. The class IDs and their corresponding class names are presented in the following section. </p> <p><strong>Classes</strong></p> <p>Class ID - Corine Land Cover 2018 class name<br> 1 - Continuous urban fabric<br> 2 - Discontinuous urban fabric<br> 3 - Industrial or commercial units<br> 4 - Road and rail networks and associated land<br> 5 - Port areas<br> 6 - Airports<br> 7 - Mineral extraction sites<br> 8 - Dump sites<br> 9 - Construction sites<br> 10 - Green urban areas<br> 11 - Sport and leisure facilities<br> 12 - Non-irrigated arable land<br> 13 - Permanently irrigated land<br> 14 - Rice fields<br> 15 - Vineyards<br> 16 - Fruit trees and berry plantations<br> 17 - Olive groves<br> 18 - Pastures<br> 19 - Annual crops associated with permanent crops<br> 20 - Complex cultivation patterns<br> 21 - Land principally occupied by agriculture, with significant areas of natural vegetation<br> 22 - Agro-forestry areas<br> 23 - Broad-leaved forest<br> 24 - Coniferous forest<br> 25 - Mixed forest<br> 26 - Natural grasslands<br> 27 - Moors and heathland<br> 28 - Sclerophyllous vegetation<br> 29 - Transitional woodland-shrub<br> 30 - Beaches, dunes, sands<br> 31 - Bare rocks<br> 32 - Sparsely vegetated areas<br> 33 - Burnt areas<br> 34 - Glaciers and perpetual snow<br> 35 - Inland marshes<br> 36 - Peat bogs<br> 37 - Salt marshes<br> 38 - Salines<br> 39 - Intertidal flats<br> 40 - Water courses<br> 41- Water bodies<br> 42 - Coastal lagoons<br> 43 - Estuaries<br> 44 - Sea and ocean<br> 48 - NODATA<br> 49 - UNCLASSIFIED LAND SURFACE<br> 50 - UNCLASSIFIED WATER BODIES </p> <p>More details about the CLC classes and conventions can be found in the CLC nomenclature guide (<a href="https://land.copernicus.eu/user-corner/technical-library/corine-land-cover-nomenclature-guidelines/html">Link</a>).</p> <p><strong>Attribution</strong></p> <p>If you find this work useful please consider citing:</p> <p>Wilhelm, T.; Koßmann, D. LAND COVER CLASSIFICATION FROM A MAPPING PERSPECTIVE: PIXELWISE SUPERVISION IN THE DEEP LEARNING ERA. In Proceedings of the IGARSS 2021—2021 IEEE International Geoscience and Remote Sensing Symposium, Brussels, Belgium, 12 – 16 July 2021; to appear.</p> <p><strong>License</strong></p> <p>The generated maps are based on data from the Copernicus program, which are subject to the terms described here:<br> <a href="https://land.copernicus.eu/pan-european/corine-land-cover/clc2018?tab=metadata">https://land.copernicus.eu/pan-european/corine-land-cover/clc2018?tab=metadata</a></p>
Synthetic dataset for dual-perspective self-supervised learning
<p>The synthetic Ca datasets for training and testing, including training dataset with bidirectional collinear scan (N<sub>y</sub> = 2N<sub>x</sub>) for MP-SSL, training dataset with normal scan (N<sub>y</sub> = N<sub>x</sub>) for TP-SSL, and testing data (N<sub>y</sub> = N<sub>x</sub>)</p> <p>If you use these data simulated using our modified <a href="https://doi.org/10.1016/j.jneumeth.2021.109173">NAOMi</a> model, please cite the corresponding work:</p> <p><a href="https://doi.org/10.1186/s43074-023-00117-0"><strong><span>B. Shen</span></strong><span>, C. Luo, W. Pang, Y. Jiang, W. Wu, R. Hu, J. Qu, B. Gu, L. Liu. Surmounting photon limits and motion artifacts for biological dynamics imaging via dual-perspective self-supervised learning. PhotoniX 5, 1 (2024). </span></a></p>
Experimental dataset for dual-perspective self-supervised learning
<p>Experimental data for training and testing, including astrocyte data, rapid hemodynamic data (larger vessels), vascular data (smaller vessels), zebrafish cardiac data.</p> <p>If you use these data acquired using our imaging system, please cite the corresponding work:</p> <p><a href="https://doi.org/10.1186/s43074-023-00117-0"><span><span><span> </span></span></span><strong><span>B. Shen</span></strong><span>, C. Luo, W. Pang, Y. Jiang, W. Wu, R. Hu, J. Qu, B. Gu, L. Liu. Surmounting photon limits and motion artifacts for biological dynamics imaging via dual-perspective self-supervised learning. PhotoniX 5, 1 (2024). </span></a></p> <p> </p>
Systematic Evaluation of Psychotherapy Training: Supervision, Student Learning, and Patient Outcomes
ClinicalTrials.gov study NCT06105502. IPD Sharing: NO. Countries: 1. Publications: 0.
Evaluation via Supervised Machine Learning of the Broiler Pectoralis Major and Liver Transcriptome in Association with the Muscle Myopathy Wooden Breast
GEO Series GSE144000. Gallus gallus. 35 samples. Type: Expression profiling by high throughput sequencing.
AI-Based Self-Supervised Learning Model Using Non-Contrast Breast MRI for Early Screening and Clinical Utility Evaluation
ClinicalTrials.gov study NCT07205276. IPD Sharing: YES. Countries: 0. Publications: 0.
SPATIALLY ADAPTIVE SEMI-SUPERVISED LEARNING WITH GAUSSIAN PROCESSES FOR HYPERSPECTRAL DATA ANALYSIS
SPATIALLY ADAPTIVE SEMI-SUPERVISED LEARNING WITH GAUSSIAN PROCESSES FOR HYPERSPECTRAL DATA ANALYSIS GOO JUN * AND JOYDEEP GHOSH* Abstract. A semi-supervised learning algorithm for the classification of hyperspectral data, Gaussian process expectation maximization (GP-EM), is proposed. Model parameters for each land cover class is first estimated by a supervised algorithm using Gaussian process regressions to find spatially adaptive parameters, and the estimated parameters are then used to initialize a spatially adaptive mixture-of-Gaussians model. The mixture model is updated by expectationmaximization iterations using the unlabeled data, and the spatially adaptive parameters for unlabeled instances are obtained by Gaussian process regressions with soft assignments. Two sets of hyperspectral data taken from the Botswana area by the NASA EO-1 satellite are used for experiments. Empirical evaluations show that the proposed framework performs significantly better than baseline algorithms that do not use spatial information, and the results are also better than any previously reported results by other algorithms on the same data.
Self-Supervised Learning Cell Image Dataset of Master Thesis "Enhancing Cell Instance Segmentation in 3D Microscopy using Self-Supervised ViTs"
<p>This is the self-supervised learning cell image dataset of master thesis "Enhancing Cell Instance Segmentation in 3D Microscopy using Self-Supervised ViTs". We gather images from datasets such as the LIVECell dataset, the EVICAN dataset, as well as datasets available on Image Data Resource (https://idr.openmicroscopy.org/) and the Broad Bioimage Benchmark Collection (https://bbbc.broadinstitute.org/). Only images with sizes larger than 512x512 are collected. For datasets containing more than 1000 images, we randomly select 1000 images. Otherwise, we retain all images in the dataset. </p> <p> </p> <p>After download, please put all compressed folders of subdatasets in the "image" folder under the root directory.</p>
Prediction of illness remission in patients with Obsessive-Compulsive Disorder with supervised machine learning
<p>Prediction of illness remission in patients with Obsessive-Compulsive<br> Disorder with supervised machine learning</p> <p> </p> <p>Introduction: The course of OCD differs widely among OCD patients, varying from chronic symptoms to full<br> remission. No tools for individual prediction of OCD remission are currently available. This study aimed to<br> develop a machine learning algorithm to predict OCD remission after two years, using solely predictors easily<br> accessible in the daily clinical routine.<br> Methods: Subjects were recruited in a longitudinal multi-center study (NOCDA). Gradient boosted decision<br> trees were used as supervised machine learning technique. The training of the algorithm was performed with 227<br> predictors and 213 observations collected in a single clinical center. Hyper-parameter optimization was performed<br> with cross-validation and a Bayesian optimization strategy. The predictive performance of the algorithm<br> was subsequently tested in an independent sample of 215 observations collected in five different centers.<br> Between-center differences were investigated with a bootstrap resampling approach.<br> Results: The average predictive performance of the algorithm in the test centers resulted in an AUROC of<br> 0.7820, a sensitivity of 73.42%, and a specificity of 71.45%. Results also showed a significant between-center<br> variation in the predictive performance. The most important predictors resulted related to OCD severity, OCD<br> chronic course, use of psychotropic medications, and better global functioning.<br> Limitations: All recruiting centers followed the same assessment protocol and are in The Netherlands. Moreover,<br> the sample of the data recruited in some of the test centers was limited in size.<br> Discussion: The algorithm demonstrated a moderate average predictive performance, and future studies will<br> focus on increasing the stability of the predictive performance across clinical settings.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.