Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
20
datasets available to search
ShareScore release 0.7.1
Dataset results
20 results for “supervised machine learning”
Datasets for "Mapping Lunar Swirls with Machine Learning: The Application of Unsupervised and Supervised Classification Algorithms in Reiner Gamma and Mare Ingenii"
<p>Final surface reflectance data at 2.6 m/pixel resolution with floating point values are available as GeoTiff and ASCII text files. Definition files for the K-Means and MLC algorithms in classifying swirl units are also available as ASCII text files. See README file for further details.</p> <p>Data used in the research article:</p> <p>Chuang, F.C., M.D. Richardson, J.R. Weirich, A.A. Sickafoose, and D.L. Domingue, 2022. Mapping Lunar Swirls with Machine Learning: The Application of Unsupervised and Supervised Image Classification Algorithms in Reiner Gamma and Mare Ingenii. The Planetary Science Journal, 3:231. doi://10.3847/PSJ/ac8f43</p> <p> </p>
Data Set for 'Self-Supervised Machine Learning for Live Cell Imagery Segmentation'
<p><strong>Self-supervised machine learning code and data for segmenting live cell imagery (Matlab)</strong></p> <p><em>Running the Code</em></p> <p>SSL_Demo_2.m : main program for self-supervised machine learning segmentation</p> <p>SSL_Declumping_2.m : main program for declumping application (applied to output of SSL_Demo_2.m)</p> <p>This Matlab code is designed to be used with time-resolved live cell microscopy images (tiffs) for the automated segmentation of cells from background.</p> <p>It is recommended you first run this code with its accompanying demo data (included in this package), keeping the current directory structure.</p> <p>Simply open SSL_Demo_2.m or SSL_Declumping_2.m in Matlab and hit Run.</p> <p><em>Code Methodology</em></p> <p>The principle of self-supervised machine learning is that you simply load your images and Run - no parameter tuning needed, no training imagery required.</p> <p>Run from start to finish, the SSL_Demo_2.m code uses consecutive pairs of images to generate training data of 'cells' and 'background' via dynamic feature vectors based on optical flow (unsupervised). These self-labeled pixels are then used to generate static feature vectors (entropy, gradient), which in turn are used to train a classifier model. The training data is updated every image in order to automatically adapt to temporal changes in cell morphologies or background illumination.</p> <p>The code was tested for high fidelity segmentation using five different modes of light microscopy: transmitted light, DIC, phase contrast, fluorescence and interference reflection microscopy.</p> <p>Six different cell lines were imaged to cover a range of morphologies and phenotypic dynamics using three cameras of differing resolutions.</p> <p>The associated manuscript for this work can be found here (although the latest version is under peer review as of this writing): </p> <p><a href="https://www.biorxiv.org/content/10.1101/2021.01.07.425773v1">https://www.biorxiv.org/content/10.1101/2021.01.07.425773v1</a></p> <p>This code was tested on Matlab v2020a and v2021a using commercially available laptop computers running the Windows 10 operating system.</p>
Assessing predictive performance of supervised machine learning algorithms for a diamond pricing model
<p>The diamond is 58 times harder than any other mineral in the world, and its elegance as a jewel has long been appreciated. Forecasting diamond prices is challenging due to nonlinearity in important features such as carat, cut, clarity, table, and depth. Against this backdrop, the study conducted a comparative analysis of the performance of multiple supervised machine learning models (regressors and classifiers) in predicting diamond prices. Eight supervised machine learning algorithms were evaluated in this work including Multiple Linear Regression, Linear Discriminant Analysis, eXtreme Gradient Boosting, Random Forest, k-Nearest Neighbors, Support Vector Machines, Boosted Regression and Classification Trees, and Multi-Layer Perceptron. The analysis is based on data preprocessing, exploratory data analysis (EDA), training the aforementioned models, assessing their accuracy, and interpreting their results. Based on the performance metrics values and analysis, it was discovered that eXtreme Gradient Boosting was the most optimal algorithm in both classification and regression, with a R<sup>2</sup> score of 97.45% and an Accuracy value of 74.28%. As a result, eXtreme Gradient Boosting was recommended as the optimal regressor and classifier for forecasting the price of a diamond specimen.</p>
A Supervised Machine-Learning Prediction of Textile's Antimicrobial Capacity Coated with Nanomaterials
<p>The dataset contains P-Chem properties of NMs and experimental conditions for assessing the antimicrobial properties of inorganic and organic NMs using machine learning tools.</p>
Image for the dataset in "Extraction of stratigraphic exposures on visible images using a supervised machine learning technique"
<p>This is the original and hand-masked image for the dataset used in a research paper "Extraction of stratigraphic exposures on visible images using a supervised machine learning technique".</p> <p>The content is</p> <ul> <li>Original images with hand-masked images (original_images_NOGUCHIandShoji.zip) <ul> <li>training/* : original images used for the training dataset generation (60 files)</li> <li>training_masks/* : hand-masked images for the training dataset generation (60 files)</li> <li>validation/* : original images used for the training dataset generation (10 files)</li> <li>validation_masks/* : hand-masked images for the validation dataset generation (10 files)</li> <li>test/* : original images used as the test data (5 files)</li> <li>test_masks/* : hand-masked images used as the test data (5 files).</li> </ul> </li> </ul> <p>Note that original images include images obtained using <em>google-image-download</em>, a Python script published on GitHub (<a href="https://github.com/Joeclinton1/google-images-download/tree/patch-1">https://github.com/Joeclinton1/google-images-download/tree/patch-1</a>, Copyright © 2015-2019 Hardik Vasa). The whole images we obtained by <em>google-image-download</em> were labeled as noncommercial reuse with modification.</p> <p>For more details, please refer to a research paper "Extraction of stratigraphic exposures on visible images using a supervised machine learning technique".</p> <p>Correspondence: Rina Noguchi (r-noguchi@env.sc.niigata-u.ac.jp)</p>
A data-driven supervised machine learning approach to estimating global ambient air pollution concentrations with associated prediction intervals
Open the record for dataset details and reuse information.
Assessing predictive performance of supervised machine learning algorithms for a diamond pricing model
Open the record for dataset details and reuse information.
Using photodiodes and supervised Machine Learning for automatic classification of weld defects in laser welding of thin foils copper-to-steel battery tabs
<p>In this folder, excel files are stored with the results of signal processing that supported findings in the following paper:</p> <p>"Using photodiodes and supervised Machine Learning for automatic classification of weld defects in laser welding of thin foils copper-to-steell battery tabs".</p> <p>Matlab scripts and orginal signals will be uploaded soon with more detailed description.</p> <p> </p>
Data from: Behaviour-specific spatiotemporal patterns of habitat use by sea turtles revealed using biologging and supervised machine learning
<ol> <li>Conservation of threatened species and anthropogenic threat mitigation commonly rely on spatially managed areas selected according to habitat preference. Since the impact of threats can be behaviour-specific, such information could be incorporated into spatial management to improve conservation outcomes. However, collecting spatially explicit behavioural data is challenging.</li> <li>Using multi-sensor biologging tags containing high-resolution movement sensors (e.g., accelerometer, magnetometer, GPS) and animal-borne video cameras, combined with supervised machine learning, we developed a method to automatically identify and geolocate typically ambiguous behaviours for the poorly understood flatback turtle <em>Natator depressus</em>. Subsequently, we evaluated behaviour-specific spatiotemporal patterns of habitat use.</li> <li>Boosted regression trees successfully identified the presence of foraging and resting in 7074 dives (AUC > 0.9), using dive features representing characteristics of locomotory activity, body posture, and three-dimensional dive paths validated by ancillary video data. Foraging was characterised by dives with longer duration, variable depth, tortuous bottom phases; resting was characterised by dives with decreased locomotory activity and longer duration bottom phases.</li> <li>Foraging and resting showed minimal spatial segregation based on 50% and 95% utilisation distributions. Expected diel patterns of behaviour-specific habitat use were superseded by the extreme tides at the near-shore study site. Turtles rested in areas close to the subtidal and intertidal boundary within larger overlapping foraging areas, allowing efficient access to intertidal food resources upon inundation at high tides when foraging was ~25% more likely.</li> <li> <em>Synthesis and applications:</em><span> Using supervised machine learning and biologging tools, we show the potential for dynamic spatial management of flatback turtles to mitigate behaviour-specific threats by prioritising protection of important locations at pertinent times. Although results are a species-specific response to a super-tidal environment</span>, our approach can be generalised to a broad range of taxa and study systems, facilitating a conceptual advance in spatial management.</li> </ol>
Data from: Behaviour-specific spatiotemporal patterns of habitat use by sea turtles revealed using biologging and supervised machine learning
Open the record for dataset details and reuse information.
Combination of whole genome sequencing and Supervised Machine Learning provides unambiguous identification of enterohemorrhagic Escherichia coli in raw milk
<p>These dataset are used in the "rename_list_of_groups.ipynb" notebook</p>
Shear Sonic Prediction Using Supervised Machine Learning: Case Study Talang Akar Formation
<p>This material has presented on 2nd International Conference on Advanced Research in Engineering and Technology in October 25, 2023.</p>
Predicting Equatorial Spread F at JICAMARCA Sector via Supervised Machine Learning
<p>Dataset used for ESF prediction model</p> <p> </p>
Data for "Supervised machine learning methods for crystal structure prediction of the binary Cs-Te system"
<p>Crystal structures, high-throughput calculations and trained machine learning models presented in the paper "Supervised machine learning methods for crystal structure prediction of the binary Cs-Te system".</p> <ul> <li><em>crystal_datasets </em>contains the input/output data sets of crystal structures for high-throughput calculations and ML models.</li> <li><em>aiida_ht_calculations </em>contains the data regarding the high-throughput DFT calculations.</li> <li><em>ml_models</em> contains the trained ML models.</li> </ul> <p>Eeach zip-archive contains a jupyter-notebook examplifying how the data can be accessed and reused.</p>
Dataset for "Extraction of stratigraphic exposures on visible images using a supervised machine learning technique"
<p>This is the dataset used in a research paper "Extraction of stratigraphic exposures on visible images using a supervised machine learning technique".</p> <p>The content is</p> <ul> <li>Augmented images used in the U-Net training (aug_images.zip) <ul> <li>train/*.png: augmented original images (14,219 files)</li> <li>train_masks/*.png: augmented hand-masked images (14,219 files).</li> </ul> </li> </ul> <p>Note that original images include images obtained using <em>google-image-download</em>, a Python script published on GitHub (<a href="https://github.com/Joeclinton1/google-images-download/tree/patch-1">https://github.com/Joeclinton1/google-images-download/tree/patch-1</a>, Copyright © 2015-2019 Hardik Vasa). The whole images we obtained by <em>google-image-download</em> were labeled as noncommercial reuse with modification.</p> <p>For more details, please refer to a research paper "Extraction of stratigraphic exposures on visible images using a supervised machine learning technique".</p> <p>Correspondence: Rina Noguchi (r-noguchi@env.sc.niigata-u.ac.jp)</p>
Community-Based Care for Minority Adolescents With ADHD: Improving Fidelity With Machine Learning-Assisted Supervision and Fidelity Feedback.
ClinicalTrials.gov study NCT05135065. IPD Sharing: Not stated. Countries: 0. Publications: 1.
Systematic review of validation of supervised machine learning models in accelerometer-based animal behaviour classification literature
Open the record for dataset details and reuse information.
Evaluation via Supervised Machine Learning of the Broiler Pectoralis Major and Liver Transcriptome in Association with the Muscle Myopathy Wooden Breast
GEO Series GSE144000. Gallus gallus. 35 samples. Type: Expression profiling by high throughput sequencing.
Prediction of illness remission in patients with Obsessive-Compulsive Disorder with supervised machine learning
<p>Prediction of illness remission in patients with Obsessive-Compulsive<br> Disorder with supervised machine learning</p> <p> </p> <p>Introduction: The course of OCD differs widely among OCD patients, varying from chronic symptoms to full<br> remission. No tools for individual prediction of OCD remission are currently available. This study aimed to<br> develop a machine learning algorithm to predict OCD remission after two years, using solely predictors easily<br> accessible in the daily clinical routine.<br> Methods: Subjects were recruited in a longitudinal multi-center study (NOCDA). Gradient boosted decision<br> trees were used as supervised machine learning technique. The training of the algorithm was performed with 227<br> predictors and 213 observations collected in a single clinical center. Hyper-parameter optimization was performed<br> with cross-validation and a Bayesian optimization strategy. The predictive performance of the algorithm<br> was subsequently tested in an independent sample of 215 observations collected in five different centers.<br> Between-center differences were investigated with a bootstrap resampling approach.<br> Results: The average predictive performance of the algorithm in the test centers resulted in an AUROC of<br> 0.7820, a sensitivity of 73.42%, and a specificity of 71.45%. Results also showed a significant between-center<br> variation in the predictive performance. The most important predictors resulted related to OCD severity, OCD<br> chronic course, use of psychotropic medications, and better global functioning.<br> Limitations: All recruiting centers followed the same assessment protocol and are in The Netherlands. Moreover,<br> the sample of the data recruited in some of the test centers was limited in size.<br> Discussion: The algorithm demonstrated a moderate average predictive performance, and future studies will<br> focus on increasing the stability of the predictive performance across clinical settings.</p>
Database of antimicrobial resistant bacterial taxa for supervised machine learning
<p>It is a database of antimicrobial resistant bacterial taxa for supervised machine learning.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.