Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

101

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

101 results for “Supervised learning”

Learn how ShareScore rates datasets ↗
zenodo36/100

EnGRaiN : A Supervised Ensemble Learning Method for Recovery of Large-scale Gene Regulatory Networks

<p>EnGRaiN is a supervised machine learning method to construct ensemble networks. To benefit from the typical accuracy advantages of supervised learning methods while taking into account the impossibility of knowing true networks for training, we devised a method that uses small training datasets of true positives and true negatives among gene pairs.</p> <p>The datasets used to evaluate the performance of EnGaiN include (i) simulated datasets generated from Yeast networks and (ii) A. thaliana gene expression datasets.</p>

opencc-by-4.0Dec 2021View details →
zenodo36/100

scPretrain: Multi-task self-supervised learning for cell type classification

<p>The dataset and code for paper, scPretrain: Multi-task self-supervised learning for cell type classification.</p>

opencc-by-4.0Nov 2020View details →
zenodo36/100

AbdomenCT-1K: Weakly Supervised Learning Benchmark

<p>This is the dataset of AbdomenCT-1K: Weakly Supervised Learning Benchmark.</p> <p>Related paper: <a href="https://ieeexplore.ieee.org/document/9497733/">https://ieeexplore.ieee.org/document/9497733/</a></p> <p>Benchmark homepage: https://abdomenct-1k-weaklysupervisedlearning.grand-challenge.org/</p>

opencc-by-4.0Jul 2021View details →
zenodo36/100

AbdomenCT-1K: Fully Supervised Learning Benchmark

<p>This is the dataset of AbdomenCT-1K: Fully Supervised Learning Benchmark.</p> <p>Related paper: https://ieeexplore.ieee.org/document/9497733/</p> <p>Benchmark homepage: https://abdomenct-1k-fully-supervised-learning.grand-challenge.org/</p> <p>&nbsp;</p>

opencc-by-4.0Jul 2021View details →
zenodo36/100

AbdomenCT-1K: Semi-supervised Learning Benchmark (Subtask 1)

<p>This is the Subtask 1 dataset of AbdomenCT-1K: Semi-supervised Learning Benchmark.</p> <p>Related paper: https://ieeexplore.ieee.org/document/9497733/</p> <p>Benchmark homepage: https://abdomenct-1k-semi-supervised-learning.grand-challenge.org/Home/</p> <p>&nbsp;</p>

opencc-by-4.0Jul 2021View details →
zenodo36/100

AbdomenCT-1K: Semi-supervised Learning Benchmark (Subtask 2 Part 1)

<p>This is the Subtask 2 (Part 1) dataset of AbdomenCT-1K: Semi-supervised Learning Benchmark.</p> <p>Related paper: https://ieeexplore.ieee.org/document/9497733/</p> <p>Benchmark homepage: https://abdomenct-1k-semi-supervised-learning.grand-challenge.org/Home/</p> <p>&nbsp;</p>

opencc-by-4.0Jul 2021View details →
zenodo36/100

AbdomenCT-1K: Semi-supervised Learning Benchmark (Subtask 2 Part 2)

<p>This is the Subtask 2 (Part 2) dataset of AbdomenCT-1K: Semi-supervised Learning Benchmark.</p> <p>Related paper: https://ieeexplore.ieee.org/document/9497733/</p> <p>Benchmark homepage: https://abdomenct-1k-semi-supervised-learning.grand-challenge.org/Home/</p> <p>&nbsp;</p>

opencc-by-4.0Jul 2021View details →
dryad36/100

Data from: Correcting a bias in the computation of behavioral time budgets that are based on supervised learning

<p>Supervised learning of behavioral modes from body-acceleration data has become a widely used research tool in Behavioral Ecology over the past decade. One of the primary usages of this tool is to estimate behavioral time budgets from the distribution of behaviors as predicted by the model. These serve as the key parameters to test predictions about the variation in animal behavior. In this paper we show that the widespread computation of behavioral time budgets is biased, due to ignoring the classification model confusion probabilities. Next, we introduce <em>the confusion matrix correction for time budgets</em> -- a simple correction method for adjusting the computed time budgets based on the model's confusion matrix. Finally, we show that the proposed correction is able to eliminate the bias, both theoretically and empirically in a series of data simulations on body acceleration data of a fossorial rodent species (Damaraland mole-rat, <em>Fukomys damarensis</em>). Our paper provides a simple implementation of <em>the confusion matrix correction for time budgets</em>, and we encourage researchers to use it to improve accuracy of behavioral time budget calculations.</p>

opencc-zeroMar 2022View details →
zenodo36/100

Multi-task self-supervised learning for wearables - human activity recognition

<p>Datasets used to train and evaluated the self-supervised-learning model</p>

opencc-by-4.0May 2022View details →
zenodo36/100

Using photodiodes and supervised Machine Learning for automatic classification of weld defects in laser welding of thin foils copper-to-steel battery tabs

<p>In this folder, excel files are stored with the results of signal processing that supported findings in the following paper:</p> <p>&quot;Using photodiodes and supervised Machine Learning for automatic classification of weld defects in laser welding of thin foils copper-to-steell battery tabs&quot;.</p> <p>Matlab scripts and orginal signals will be uploaded soon with more detailed description.</p> <p>&nbsp;</p>

opencc-by-4.0Jun 2022View details →
zenodo36/100

Supervised learning is an accurate method for network-based gene classification - Data

<p>This file contains the data that was used in the paper titled &quot;Supervised learning is an accurate method for network-based gene classification&quot; (https://doi.org/10.1093/bioinformatics/btaa150). Some data was excluded if the license was not permissive enough.</p>

opencc-by-nc-sa-4.0Jul 2019View details →
zenodo36/100

l-sized Training and Evaluation Data for Publication "Using Supervised Learning to Classify Metadata of Research Data by Field of Study"

<p>Automated classification of metadata of research data by their discipline(s) of research can be used in scientometric research, by repository service providers, and in the context of research data aggregation services. Openly available metadata of the DataCite index for research data were used to compile a large training and evaluation set comprised of 609,524 records. This is the cleaned and vectorized version with a feature selection of large size.</p>

opencc-by-4.0Oct 2019View details →
zenodo36/100

s-sized Training and Evaluation Data for Publication "Using Supervised Learning to Classify Metadata of Research Data by Field of Study"

<p>Automated classification of metadata of research data by their discipline(s) of research can be used in scientometric research, by repository service providers, and in the context of research data aggregation services. Openly available metadata of the DataCite index for research data were used to compile a large training and evaluation set comprised of 609,524 records. This is the cleaned and vectorized version with a feature selection of small size.</p>

opencc-by-4.0Oct 2019View details →
zenodo36/100

Dataset for supervised learning with a deep neural network to assess azimuthal localisation in sound field synthesis

<p>Dataset for supervised learning with a deep neural network to assess azimuthal localisation in sound field synthesis.<br> Released as part of the Master Thesis &#39;An Auditory Model for Azimuthal Localisation in Sound Field Synthesis&#39;.</p> <p>This database is calculated from the data of listening experiments.</p>

opencc-by-4.0Nov 2019View details →
dryad36/100

Data from: Behaviour-specific spatiotemporal patterns of habitat use by sea turtles revealed using biologging and supervised machine learning

<ol> <li>Conservation of threatened species and anthropogenic threat mitigation commonly rely on spatially managed areas selected according to habitat preference. Since the impact of threats can be behaviour-specific, such information could be incorporated into spatial management to improve conservation outcomes. However, collecting spatially explicit behavioural data is challenging.</li> <li>Using multi-sensor biologging tags containing high-resolution movement sensors (e.g., accelerometer, magnetometer, GPS) and animal-borne video cameras, combined with supervised machine learning, we developed a method to automatically identify and geolocate typically ambiguous behaviours for the poorly understood flatback turtle <em>Natator depressus</em>. Subsequently, we evaluated behaviour-specific spatiotemporal patterns of habitat use.</li> <li>Boosted regression trees successfully identified the presence of foraging and resting in 7074 dives (AUC &gt; 0.9), using dive features representing characteristics of locomotory activity, body posture, and three-dimensional dive paths validated by ancillary video data. Foraging was characterised by dives with longer duration, variable depth, tortuous bottom phases; resting was characterised by dives with decreased locomotory activity and longer duration bottom phases.</li> <li>Foraging and resting showed minimal spatial segregation based on 50% and 95% utilisation distributions. Expected diel patterns of behaviour-specific habitat use were superseded by the extreme tides at the near-shore study site. Turtles rested in areas close to the subtidal and intertidal boundary within larger overlapping foraging areas, allowing efficient access to intertidal food resources upon inundation at high tides when foraging was ~25% more likely.</li> <li> <em>Synthesis and applications:</em><span> Using supervised machine learning and biologging tools, we show the potential for dynamic spatial management of flatback turtles to mitigate behaviour-specific threats by prioritising protection of important locations at pertinent times. Although results are a species-specific response to a super-tidal environment</span>, our approach can be generalised to a broad range of taxa and study systems, facilitating a conceptual advance in spatial management.</li> </ol>

opencc-zeroMay 2023View details →
dryad36/100

Data from: Correcting a bias in the computation of behavioral time budgets that are based on supervised learning

Open the record for dataset details and reuse information.

publicMar 2022View details →
dryad36/100

Data from: Behaviour-specific spatiotemporal patterns of habitat use by sea turtles revealed using biologging and supervised machine learning

Open the record for dataset details and reuse information.

publicMay 2023View details →
zenodo32/100

Learning Lenient Parsing & Typing via Indirect Supervision

<p><strong>Validation Set:</strong> We have shared 3 CSV files containing human-annotated validation sets of our paper (<em>Validation Data.zip</em>).</p> <p><strong>AST and Student Code Correction:</strong>&nbsp; For generating AST and code correction, please check the two files, AST.py and Top1.py ( in <em>AST &amp; Top-1.zip</em> ). In AST.py, we present the output of different parts of the program with an example. Please read that one before Top1.py. We follow the implementation of<a href="https://github.com/Lsdefine/attention-is-all-you-need-keras?fbclid=IwAR2UCy9NvD_PLHFVqM6B1VqvMVZzWRIS25BTG4nJudEhs3684RSOpGBhOyk"> https://github.com/Lsdefine/attention-is-all-you-need-keras</a>. Please check the remaining code in the above link. We made a minor correction in the dataloader.py to use two separate vocabulary cutoffs for input and output. dataloader1.py, transformer1.py, etc. are an exact replication of dataloader.py and transformer.py. Since we are using two models, we did it that way to avoid any conflict.</p> <p><strong>TypeFix:</strong> Check the code in <em>TypeFix.zip.</em></p> <p><em>We have also published data set for FragFix and BlockFix.</em></p>

opencc-by-4.0Aug 2019View details →
zenodo32/100

[Data] Self-Supervised Bayesian Representation Learning of Acoustic Emissions from Laser Powder Bed Fusion Process for In-situ Monitoring

<div> <div> <div> <p>Different Laser Powder Bed Fusion (LPBF) process spaces were deliberately introduced by employing two distinct 316L stainless steel powder distributions (with particle sizes &gt;45 &mu;m and &lt; 45 &mu;m) and processing them with two sets of laser parameters, resulting in the creation of four datasets [D1, D2, D3, and D4]. These datasets encompass LoF pores, conduction mode, and keyhole formations, each associated with three LPBF regimes denoted as D1, D2, D3, and D4. The experiments utilized a Sisma MYSINT 100 commercial LPBF printer and an airborne AE sensor system with a flat frequency response ranging from 0 to 150 kHz.&nbsp;Validation of the ground truths for the three laser regimes across the four datasets, representing distinct process spaces, was accomplished through the confirmation of cross-sectional images. In the course of fabricating a cube using a powder bed and laser, data acquisition from an AE sensor was triggered when the optical intensity reached a threshold of 0.5 V for each scan length. The photodiode trigger gain was adjusted to saturate at 5 V, and the ensuing continuous-time window, where the optical signal remained at 5 V for 12.5 ms, was calculated and segmented to generate the dataset.&nbsp;Irrespective of the specific regime (Lack of Fusion, Conduction, and Keyhole) or the cube being fabricated (with two powder distributions), the signals obtained during this process were then segmented into a 12.5 ms window comprising 5000 data points. To eliminate any noise, an offline application of a low-pass Butterworth filter with a 150 kHz cut-off frequency was employed, aligned with the frequency response specification of the AE sensor. Each dataset has two files against it [raw/groundtruth label].</p> </div> </div> </div>

opencc-by-4.0Nov 2023View details →
zenodo32/100

Self-Supervised Learning for Avian Diversity Monitoring SC22

<p><strong>Clusters&nbsp;</strong>The clusterization generated from the output of the pre-trained backbone</p> <p><strong>Morton Spectrograms</strong> The original spectrograms with which we trained the model to check the clusterization&nbsp;</p> <p><strong>Pretrained Model</strong> The pre-trained model</p> <p><strong>Single Image Attentional Maps</strong>&nbsp;The attentional maps and masks generated for a single image</p> <p><strong>features attentional&nbsp;maps and names</strong>&nbsp;Features, attentional maps and names of all the Spectrogram Images</p>

opencc-by-4.0Aug 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record