Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

101

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

101 results for “Supervised learning”

Learn how ShareScore rates datasets ↗
zenodo32/100

Combination of whole genome sequencing and Supervised Machine Learning provides unambiguous identification of enterohemorrhagic Escherichia coli in raw milk

<p>These dataset are used in the &quot;rename_list_of_groups.ipynb&quot; notebook</p>

opencc-by-4.0Sep 2022View details →
zenodo32/100

Supplementary material for the publication: "Combining unsupervised and supervised learning in microscopy enables defect analysis of a full 4H-SiC wafer"

<p><span><span>This dataset contains postprocessed data for the publication &bdquo;<span>Combining unsupervised and supervised learning in microscopy enables defect analysis of a full 4H-SiC wafer</span>&ldquo;.</span></span></p>

opencc-by-4.0May 2024View details →
zenodo32/100

Valentwin: Using Self-Supervised Contrastive Learning on Language Model for Schema Matching Datasets

<div>ValenTwin is a schema matching framework that uses self-supervised contrastive learning to train the model,&nbsp;uses the model to generate embeddings of table columns, then uses different similarity measures to match the column embeddings.</div> <div>&nbsp;</div> <div> <div>We provide two types of zip files for the datasets:<br>1. `data.zip` contains the raw data files, the ground truth files, the sampled data (n=[100, 200, 300, 400, 500] used in the experiments, as well as the contrastive data used to train the model.<br>2. `data-raw.zip` contains only the raw data files and the ground truth files. You can sample the data and generate the contrastive dataset yourself by following step 1 and 2 in the `How to Run` section. <br>Download and unzip one of the zip files to the `data` folder.</div> </div>

opencc-by-4.0May 2024View details →
zenodo32/100

Shear Sonic Prediction Using Supervised Machine Learning: Case Study Talang Akar Formation

<p>This material has presented on 2nd International Conference on Advanced Research in Engineering and Technology in October 25, 2023.</p>

opencc-by-4.0Jun 2024View details →
zenodo32/100

Predicting the failure of dental implants using supervised learning techniques

<p>A total of 747 fixtures from patients who completed their prosthodontics treatments.&nbsp;The dependent variable is dental implant failure;&nbsp;a total of 20 independent variables were collected, including age, gender, factors of missing, systemic disease, tobacco smoking, alcohol consumption, betel nut chewing, department of surgeon, surgeon experience, location of implant, bone density, ridge augmentation, Maxillary sinus augmentation, implant system, fixture length, fixture width, types of prosthesis, angle of abutment, and prosthesis fixation.</p>

opencc-by-nc-nd-4.0Apr 2018View details →
zenodo32/100

images for self-supervised learning

<p>Real images with reduced resolution with associated masks generated from data obtained from High Frequency Receiver onboard Van Allen Probes.</p> <p>Sythetic images with associated masks generated based on statistics and radomlization.&nbsp;</p> <p>Inputs for the self-supervised contrastive learning software published at https://github.com/Yi-JiunSu/SSL-Contrastive</p>

opencc-by-4.0Aug 2024View details →
zenodo32/100

STS-Tooth: A multi-modal dental dataset for semi-supervised deep learning image segmentation

<p>In response to the increasing prevalence of dental diseases, dental health, a vital aspect of human well-being, warrants greater attention. Panoramic X-ray images (PXI) and Cone Beam Computed Tomography (CBCT) are key tools for dentists in diagnosing and treating dental conditions. Additionally, deep learning for tooth segmentation can focus on relevant treatment information and localize lesions. However, the scarcity of publicly available PXI and CBCT datasets hampers their use in tooth segmentation tasks. Therefore, this paper presents a multimodal dataset for semi-supervised deep learning in dental PXI and CBCT, named STS-2D-Tooth and STS-3D-Tooth. STS-2D-Tooth includes 4,000 images and 900 masks, categorized by age into children and adults. Moreover, we have collected CBCTs providing more detailed and three-dimensional information, resulting in the STS-3D-Tooth dataset comprising 148,400 unlabeled scans and 8,800 masks. To our knowledge, this is the first multimodal dataset combining dental PXI and CBCT, and it is the largest tooth segmentation dataset, a significant step forward for the advancement of tooth segmentation.</p>

opencc-by-4.0Sep 2024View details →
zenodo32/100

Predicting Equatorial Spread F at JICAMARCA Sector via Supervised Machine Learning

<p>Dataset used for ESF prediction model</p> <p>&nbsp;</p>

opencc-by-4.0Aug 2024View details →
zenodo32/100

Data for "Supervised machine learning methods for crystal structure prediction of the binary Cs-Te system"

<p>Crystal structures, high-throughput calculations and trained machine learning models presented in the paper "Supervised machine learning methods for crystal structure prediction of the binary Cs-Te system".</p> <ul> <li><em>crystal_datasets&nbsp;</em>contains the input/output data sets of crystal structures for high-throughput calculations and ML models.</li> <li><em>aiida_ht_calculations&nbsp;</em>contains the data regarding the high-throughput DFT calculations.</li> <li><em>ml_models</em> contains the trained ML models.</li> </ul> <p>Eeach zip-archive contains a jupyter-notebook examplifying how the data can be accessed and reused.</p>

opencc-by-4.0Nov 2024View details →
zenodo32/100

Weakly-Supervised Learning Significantly Reduces the Number of Labels Required for Intracranial Hemorrhage Detection on Head CT

<p>Modern machine learning pipelines, in particular those based on deep learning (DL) models, require large amounts of labeled data. For classification problems, the most common learning paradigm consists of presenting labeled examples during training, thus providing \emph{strong supervision} by directly presenting examples from the different classes, e.g. positive and negative samples. As a result, the adequate training of these models demands the curation of large datasets with high-quality labels. This constitutes a major obstacle for the development of DL models in radiology---in particular for cross-sectional imaging (e.g., computed tomography [CT] scans)---where labels must come from manual annotations by expert radiologists at the image or slice-level (as opposed to the examination level, such as could be obtained using natural language processing of radiology reports).&nbsp;<br> This work studies the question of what kind of labels&nbsp;should be collected for the problem of intracranial hemorrhage detection in brain CT. We investigate whether image-level annotations should be preferred to examination-level ones. By framing this task as a Multiple Instance Learning (MIL) problem, and employing modern attention-based DL architectures, we analyze the degree to which different levels of supervision improve the detection performance. We find that strong supervision (learning with local image-level annotations) and weak supervision (learning with only global examination-level labels) achieve comparable performance in both examination- and image-level hemorrhage detection, as well as in hemorrhage localization at the pixel-level (explainability). Furthermore, we study this behavior as a function of the number of labels available during training. Our results suggest that local labels may not be necessary at all, drastically reducing the time and cost involved in collecting and curating datasets.</p>

opencc-by-4.0Nov 2022View details →
zenodo32/100

Supplementary to "Generalizable biomarker prediction from cancer pathology slides with self-supervised deep learning - a retrospective multicentric study"

<p>High-resolution images of&nbsp;heatmaps. Top tiles and GradCam of figure 5</p>

opencc-by-4.0Dec 2022View details →
zenodo32/100

Unlabeled Sentinel 2 time series dataset (training, T30TXT): Self-supervised Spatio-Temporal Representation Learning of Satellite Image Time Series

<p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;<strong> T30TXT unlabeled S2 dataset </strong></p> <p>This is a part of the unlabeled Sentinel 2 (S2) L2A dataset composed of patch time series acquired over France used to pretrain U-BARN. For further details, see section IV.A of the pre-print article &quot;Self-Supervised Spatio-Temporal Representation Learning Of Satellite Image Time Series&quot; available <a href="https://hal.science/hal-04084839">here</a>.&nbsp; Each patch is constituted of the 10 bands&nbsp; [B2,B3,B4,B5,B6,B7,B8,B8A,B11,B12] and the three masks [&#39;CLM_R1&#39;, &#39;EDG_R1&#39;, &#39;SAT_R1&#39;]. The global dataset is composed of two disjoint datasets: training (9 tiles) and validation dataset (4 tiles).</p> <p>In this repo,<strong> only data from the S2 tile T30TXT</strong> are available. To download the full pretraining dataset, see : <a href="https://doi.org/10.5281/zenodo.7891924">10.5281/zenodo.7891924</a></p> <table> <caption><strong>Global unlabeled dataset description</strong></caption> <tbody> <tr> <td>Dataset name</td> <td>S2 tiles</td> <td>ROI size</td> <td>Temporal extent</td> </tr> <tr> <td>Train</td> <td> <p>T30TXT,T30TYQ,T30TYS,T30UVU,</p> <p>T31TDJ,T31TDL,T31TFN,T31TGJ,T31UEP</p> </td> <td>1024*1024</td> <td>2018-2020</td> </tr> <tr> <td>Val</td> <td>T30TYR,T30UWU,T31TEK,T31UER</td> <td>256*256</td> <td>2016-2019</td> </tr> </tbody> </table> <p>&nbsp;</p>

openApr 2023View details →
zenodo32/100

Unlabeled Sentinel 2 time series dataset (training, T30TYS): Self-Supervised Spatio-Temporal Representation Learning of Satellite Image Time Series

<p>This is a part of the unlabeled Sentinel 2 (S2) L2A dataset composed of patch time series acquired over France used to pretrain U-BARN. For further details, see section IV.A of the pre-print article &quot;Self-Supervised Spatio-Temporal Representation Learning Of Satellite Image Time Series&quot; available <a href="https://hal.science/hal-04084839">here</a>.&nbsp; Each patch is constituted of the 10 bands&nbsp; [B2,B3,B4,B5,B6,B7,B8,B8A,B11,B12] and the three masks [&#39;CLM_R1&#39;, &#39;EDG_R1&#39;, &#39;SAT_R1&#39;]. The global dataset is composed of two disjoint datasets: training (9 tiles) and validation dataset (4 tiles).</p> <p>In this repo,<strong> only data from the S2 tile T30TYS</strong> are available. To download the full pretraining dataset, see: <a href="https://doi.org/10.5281/zenodo.7891924">10.5281/zenodo.7891924</a></p> <table> <caption><strong>Global unlabeled dataset description</strong></caption> <tbody> <tr> <td>Dataset name</td> <td>S2 tiles</td> <td>ROI size</td> <td>Temporal extent</td> </tr> <tr> <td>Train</td> <td> <p>T30TXT,T30TYQ,T30TYS,T30UVU,</p> <p>T31TDJ,T31TDL,T31TFN,T31TGJ,T31UEP</p> </td> <td>1024*1024</td> <td>2018-2020</td> </tr> <tr> <td>Val</td> <td>T30TYR,T30UWU,T31TEK,T31UER</td> <td>256*256</td> <td>2016-2019</td> </tr> </tbody> </table>

opencc-by-4.0Apr 2023View details →
zenodo32/100

Dataset for "Extraction of stratigraphic exposures on visible images using a supervised machine learning technique"

<p>This is the dataset used in a research paper &quot;Extraction of stratigraphic exposures on visible images using a supervised machine learning technique&quot;.</p> <p>The content&nbsp;is</p> <ul> <li>Augmented images used in the U-Net training (aug_images.zip) <ul> <li>train/*.png: augmented original images (14,219 files)</li> <li>train_masks/*.png: augmented hand-masked images (14,219 files).</li> </ul> </li> </ul> <p>Note that original images include&nbsp;images obtained using <em>google-image-download</em>, a Python script published on GitHub (<a href="https://github.com/Joeclinton1/google-images-download/tree/patch-1">https://github.com/Joeclinton1/google-images-download/tree/patch-1</a>, Copyright &copy; 2015-2019 Hardik Vasa).&nbsp;The whole images we obtained by <em>google-image-download</em> were labeled as noncommercial reuse with modification.</p> <p>For more details, please refer to a research paper &quot;Extraction of stratigraphic exposures on visible images using a supervised machine learning technique&quot;.</p> <p>Correspondence: Rina Noguchi (r-noguchi@env.sc.niigata-u.ac.jp)</p>

opencc-by-4.0Jul 2023View details →
zenodo32/100

Structure-based self-supervised learning enables ultrafast prediction of stability changes upon mutation at the protein universe scale

<p>Pythia computed all single mutations of <em>E.coli</em> proteome, high quality high quality of Swiss-Prot structures and thermophilic proteins used in analysis.</p>

opencc-by-4.0Aug 2023View details →
zenodo32/100

Improving Sub-pixel Accuracy in Ultrasound Localization Microscopy Using Supervised and Self-supervised Deep Learning

<p>These are&nbsp;the original PSFs used for generating the training, validation, and evaluation dataset for the paper &quot;Improving Sub-pixel Accuracy in Ultrasound Localization Microscopy Using Supervised and Self-supervised Deep Learning&quot;.</p>

opencc-by-4.0Aug 2023View details →
ClinicalTrials.gov32/100

Community-Based Care for Minority Adolescents With ADHD: Improving Fidelity With Machine Learning-Assisted Supervision and Fidelity Feedback.

ClinicalTrials.gov study NCT05135065. IPD Sharing: Not stated. Countries: 0. Publications: 1.

restrictedIPD-UNDECIDEDFeb 2026View details →
dryad32/100

Systematic review of validation of supervised machine learning models in accelerometer-based animal behaviour classification literature

Open the record for dataset details and reuse information.

publicJun 2025View details →
dryad28/100

Data from: ZeitZeiger: supervised learning for high-dimensional data from an oscillatory system

Numerous biological systems oscillate over time or space. Despite these oscillators' importance, data from an oscillatory system is problematic for existing methods of regularized supervised learning. We present ZeitZeiger, a method to predict a periodic variable (e.g. time of day) from a high-dimensional observation. ZeitZeiger learns a sparse representation of the variation associated with the periodic variable in the training observations, then uses maximum-likelihood to make a prediction for a test observation. We applied ZeitZeiger to a comprehensive dataset of genome-wide gene expression from the mammalian circadian oscillator. Using the expression of 13 genes, ZeitZeiger predicted circadian time (internal time of day) in each of 12 mouse organs to within ∼1 h, resulting in a multi-organ predictor of circadian time. Compared to the state-of-the-art approach, ZeitZeiger was faster, more accurate and used fewer genes. We then validated the multi-organ predictor on 20 additional datasets comprising nearly 800 samples. Our results suggest that ZeitZeiger not only makes accurate predictions, but also gives insight into the behavior and structure of the oscillator from which the data originated. As our ability to collect high-dimensional data from various biological oscillators increases, ZeitZeiger should enhance efforts to convert these data to knowledge.

opencc-zeroDec 2015View details →
zenodo28/100

The impacts of active and self-supervised learning on efficient annotation of single-cell expression data - source data

<p>Source data used to create all figures in the manuscript.</p>

opencc-by-4.0Dec 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record