Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
101
datasets available to search
ShareScore release 0.9.0
Dataset results
101 results for “Supervised learning”
Combination of whole genome sequencing and Supervised Machine Learning provides unambiguous identification of enterohemorrhagic Escherichia coli in raw milk
<p>These dataset are used in the "rename_list_of_groups.ipynb" notebook</p>
Supplementary material for the publication: "Combining unsupervised and supervised learning in microscopy enables defect analysis of a full 4H-SiC wafer"
<p><span><span>This dataset contains postprocessed data for the publication „<span>Combining unsupervised and supervised learning in microscopy enables defect analysis of a full 4H-SiC wafer</span>“.</span></span></p>
Valentwin: Using Self-Supervised Contrastive Learning on Language Model for Schema Matching Datasets
<div>ValenTwin is a schema matching framework that uses self-supervised contrastive learning to train the model, uses the model to generate embeddings of table columns, then uses different similarity measures to match the column embeddings.</div> <div> </div> <div> <div>We provide two types of zip files for the datasets:<br>1. `data.zip` contains the raw data files, the ground truth files, the sampled data (n=[100, 200, 300, 400, 500] used in the experiments, as well as the contrastive data used to train the model.<br>2. `data-raw.zip` contains only the raw data files and the ground truth files. You can sample the data and generate the contrastive dataset yourself by following step 1 and 2 in the `How to Run` section. <br>Download and unzip one of the zip files to the `data` folder.</div> </div>
Shear Sonic Prediction Using Supervised Machine Learning: Case Study Talang Akar Formation
<p>This material has presented on 2nd International Conference on Advanced Research in Engineering and Technology in October 25, 2023.</p>
Predicting the failure of dental implants using supervised learning techniques
<p>A total of 747 fixtures from patients who completed their prosthodontics treatments. The dependent variable is dental implant failure; a total of 20 independent variables were collected, including age, gender, factors of missing, systemic disease, tobacco smoking, alcohol consumption, betel nut chewing, department of surgeon, surgeon experience, location of implant, bone density, ridge augmentation, Maxillary sinus augmentation, implant system, fixture length, fixture width, types of prosthesis, angle of abutment, and prosthesis fixation.</p>
images for self-supervised learning
<p>Real images with reduced resolution with associated masks generated from data obtained from High Frequency Receiver onboard Van Allen Probes.</p> <p>Sythetic images with associated masks generated based on statistics and radomlization. </p> <p>Inputs for the self-supervised contrastive learning software published at https://github.com/Yi-JiunSu/SSL-Contrastive</p>
STS-Tooth: A multi-modal dental dataset for semi-supervised deep learning image segmentation
<p>In response to the increasing prevalence of dental diseases, dental health, a vital aspect of human well-being, warrants greater attention. Panoramic X-ray images (PXI) and Cone Beam Computed Tomography (CBCT) are key tools for dentists in diagnosing and treating dental conditions. Additionally, deep learning for tooth segmentation can focus on relevant treatment information and localize lesions. However, the scarcity of publicly available PXI and CBCT datasets hampers their use in tooth segmentation tasks. Therefore, this paper presents a multimodal dataset for semi-supervised deep learning in dental PXI and CBCT, named STS-2D-Tooth and STS-3D-Tooth. STS-2D-Tooth includes 4,000 images and 900 masks, categorized by age into children and adults. Moreover, we have collected CBCTs providing more detailed and three-dimensional information, resulting in the STS-3D-Tooth dataset comprising 148,400 unlabeled scans and 8,800 masks. To our knowledge, this is the first multimodal dataset combining dental PXI and CBCT, and it is the largest tooth segmentation dataset, a significant step forward for the advancement of tooth segmentation.</p>
Predicting Equatorial Spread F at JICAMARCA Sector via Supervised Machine Learning
<p>Dataset used for ESF prediction model</p> <p> </p>
Data for "Supervised machine learning methods for crystal structure prediction of the binary Cs-Te system"
<p>Crystal structures, high-throughput calculations and trained machine learning models presented in the paper "Supervised machine learning methods for crystal structure prediction of the binary Cs-Te system".</p> <ul> <li><em>crystal_datasets </em>contains the input/output data sets of crystal structures for high-throughput calculations and ML models.</li> <li><em>aiida_ht_calculations </em>contains the data regarding the high-throughput DFT calculations.</li> <li><em>ml_models</em> contains the trained ML models.</li> </ul> <p>Eeach zip-archive contains a jupyter-notebook examplifying how the data can be accessed and reused.</p>
Weakly-Supervised Learning Significantly Reduces the Number of Labels Required for Intracranial Hemorrhage Detection on Head CT
<p>Modern machine learning pipelines, in particular those based on deep learning (DL) models, require large amounts of labeled data. For classification problems, the most common learning paradigm consists of presenting labeled examples during training, thus providing \emph{strong supervision} by directly presenting examples from the different classes, e.g. positive and negative samples. As a result, the adequate training of these models demands the curation of large datasets with high-quality labels. This constitutes a major obstacle for the development of DL models in radiology---in particular for cross-sectional imaging (e.g., computed tomography [CT] scans)---where labels must come from manual annotations by expert radiologists at the image or slice-level (as opposed to the examination level, such as could be obtained using natural language processing of radiology reports). <br> This work studies the question of what kind of labels should be collected for the problem of intracranial hemorrhage detection in brain CT. We investigate whether image-level annotations should be preferred to examination-level ones. By framing this task as a Multiple Instance Learning (MIL) problem, and employing modern attention-based DL architectures, we analyze the degree to which different levels of supervision improve the detection performance. We find that strong supervision (learning with local image-level annotations) and weak supervision (learning with only global examination-level labels) achieve comparable performance in both examination- and image-level hemorrhage detection, as well as in hemorrhage localization at the pixel-level (explainability). Furthermore, we study this behavior as a function of the number of labels available during training. Our results suggest that local labels may not be necessary at all, drastically reducing the time and cost involved in collecting and curating datasets.</p>
Supplementary to "Generalizable biomarker prediction from cancer pathology slides with self-supervised deep learning - a retrospective multicentric study"
<p>High-resolution images of heatmaps. Top tiles and GradCam of figure 5</p>
Unlabeled Sentinel 2 time series dataset (training, T30TXT): Self-supervised Spatio-Temporal Representation Learning of Satellite Image Time Series
<p> <strong> T30TXT unlabeled S2 dataset </strong></p> <p>This is a part of the unlabeled Sentinel 2 (S2) L2A dataset composed of patch time series acquired over France used to pretrain U-BARN. For further details, see section IV.A of the pre-print article "Self-Supervised Spatio-Temporal Representation Learning Of Satellite Image Time Series" available <a href="https://hal.science/hal-04084839">here</a>. Each patch is constituted of the 10 bands [B2,B3,B4,B5,B6,B7,B8,B8A,B11,B12] and the three masks ['CLM_R1', 'EDG_R1', 'SAT_R1']. The global dataset is composed of two disjoint datasets: training (9 tiles) and validation dataset (4 tiles).</p> <p>In this repo,<strong> only data from the S2 tile T30TXT</strong> are available. To download the full pretraining dataset, see : <a href="https://doi.org/10.5281/zenodo.7891924">10.5281/zenodo.7891924</a></p> <table> <caption><strong>Global unlabeled dataset description</strong></caption> <tbody> <tr> <td>Dataset name</td> <td>S2 tiles</td> <td>ROI size</td> <td>Temporal extent</td> </tr> <tr> <td>Train</td> <td> <p>T30TXT,T30TYQ,T30TYS,T30UVU,</p> <p>T31TDJ,T31TDL,T31TFN,T31TGJ,T31UEP</p> </td> <td>1024*1024</td> <td>2018-2020</td> </tr> <tr> <td>Val</td> <td>T30TYR,T30UWU,T31TEK,T31UER</td> <td>256*256</td> <td>2016-2019</td> </tr> </tbody> </table> <p> </p>
Unlabeled Sentinel 2 time series dataset (training, T30TYS): Self-Supervised Spatio-Temporal Representation Learning of Satellite Image Time Series
<p>This is a part of the unlabeled Sentinel 2 (S2) L2A dataset composed of patch time series acquired over France used to pretrain U-BARN. For further details, see section IV.A of the pre-print article "Self-Supervised Spatio-Temporal Representation Learning Of Satellite Image Time Series" available <a href="https://hal.science/hal-04084839">here</a>. Each patch is constituted of the 10 bands [B2,B3,B4,B5,B6,B7,B8,B8A,B11,B12] and the three masks ['CLM_R1', 'EDG_R1', 'SAT_R1']. The global dataset is composed of two disjoint datasets: training (9 tiles) and validation dataset (4 tiles).</p> <p>In this repo,<strong> only data from the S2 tile T30TYS</strong> are available. To download the full pretraining dataset, see: <a href="https://doi.org/10.5281/zenodo.7891924">10.5281/zenodo.7891924</a></p> <table> <caption><strong>Global unlabeled dataset description</strong></caption> <tbody> <tr> <td>Dataset name</td> <td>S2 tiles</td> <td>ROI size</td> <td>Temporal extent</td> </tr> <tr> <td>Train</td> <td> <p>T30TXT,T30TYQ,T30TYS,T30UVU,</p> <p>T31TDJ,T31TDL,T31TFN,T31TGJ,T31UEP</p> </td> <td>1024*1024</td> <td>2018-2020</td> </tr> <tr> <td>Val</td> <td>T30TYR,T30UWU,T31TEK,T31UER</td> <td>256*256</td> <td>2016-2019</td> </tr> </tbody> </table>
Dataset for "Extraction of stratigraphic exposures on visible images using a supervised machine learning technique"
<p>This is the dataset used in a research paper "Extraction of stratigraphic exposures on visible images using a supervised machine learning technique".</p> <p>The content is</p> <ul> <li>Augmented images used in the U-Net training (aug_images.zip) <ul> <li>train/*.png: augmented original images (14,219 files)</li> <li>train_masks/*.png: augmented hand-masked images (14,219 files).</li> </ul> </li> </ul> <p>Note that original images include images obtained using <em>google-image-download</em>, a Python script published on GitHub (<a href="https://github.com/Joeclinton1/google-images-download/tree/patch-1">https://github.com/Joeclinton1/google-images-download/tree/patch-1</a>, Copyright © 2015-2019 Hardik Vasa). The whole images we obtained by <em>google-image-download</em> were labeled as noncommercial reuse with modification.</p> <p>For more details, please refer to a research paper "Extraction of stratigraphic exposures on visible images using a supervised machine learning technique".</p> <p>Correspondence: Rina Noguchi (r-noguchi@env.sc.niigata-u.ac.jp)</p>
Structure-based self-supervised learning enables ultrafast prediction of stability changes upon mutation at the protein universe scale
<p>Pythia computed all single mutations of <em>E.coli</em> proteome, high quality high quality of Swiss-Prot structures and thermophilic proteins used in analysis.</p>
Improving Sub-pixel Accuracy in Ultrasound Localization Microscopy Using Supervised and Self-supervised Deep Learning
<p>These are the original PSFs used for generating the training, validation, and evaluation dataset for the paper "Improving Sub-pixel Accuracy in Ultrasound Localization Microscopy Using Supervised and Self-supervised Deep Learning".</p>
Community-Based Care for Minority Adolescents With ADHD: Improving Fidelity With Machine Learning-Assisted Supervision and Fidelity Feedback.
ClinicalTrials.gov study NCT05135065. IPD Sharing: Not stated. Countries: 0. Publications: 1.
Systematic review of validation of supervised machine learning models in accelerometer-based animal behaviour classification literature
Open the record for dataset details and reuse information.
Data from: ZeitZeiger: supervised learning for high-dimensional data from an oscillatory system
Numerous biological systems oscillate over time or space. Despite these oscillators' importance, data from an oscillatory system is problematic for existing methods of regularized supervised learning. We present ZeitZeiger, a method to predict a periodic variable (e.g. time of day) from a high-dimensional observation. ZeitZeiger learns a sparse representation of the variation associated with the periodic variable in the training observations, then uses maximum-likelihood to make a prediction for a test observation. We applied ZeitZeiger to a comprehensive dataset of genome-wide gene expression from the mammalian circadian oscillator. Using the expression of 13 genes, ZeitZeiger predicted circadian time (internal time of day) in each of 12 mouse organs to within ∼1 h, resulting in a multi-organ predictor of circadian time. Compared to the state-of-the-art approach, ZeitZeiger was faster, more accurate and used fewer genes. We then validated the multi-organ predictor on 20 additional datasets comprising nearly 800 samples. Our results suggest that ZeitZeiger not only makes accurate predictions, but also gives insight into the behavior and structure of the oscillator from which the data originated. As our ability to collect high-dimensional data from various biological oscillators increases, ZeitZeiger should enhance efforts to convert these data to knowledge.
The impacts of active and self-supervised learning on efficient annotation of single-cell expression data - source data
<p>Source data used to create all figures in the manuscript.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.