Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
78
datasets available to search
ShareScore release 0.9.0
Dataset results
78 results for “Machine Learning Classification”
IMBALANCED MACHINE LEARNING CLASSIFICATION MODELS FOR REMOVAL BIOSIMILAR DRUGS AND INCREASED ACTIVITY IN PATIENTS WITH RHEUMATIC DISEASES
<p>Objective: Predict long-term disease worsening and the removal of biosimilar medication in patients with rheumatic diseases.</p><p>Methodology: Observational, retrospective, and descriptive study. Review of a database of patients with immune-mediated inflammatory rheumatic diseases. Disease worsening and removing biosimilars are imbalanced variables, that require using imbalanced machine learning models selected based on their superior f1-scores and great accuracy. Previously, we selected the most important variables using mutual information tests.</p><p>Results: The best imbalanced machine learning models to predict disease worsening and the removal of the biosimilar obtained f1-scores of 0.52 and 0.63, respectively. Both models are decision trees. In the first one, two important factors are switching of biosimilar and age, and in the second, the relevant variables are optimization and the value of the initial CRP. </p><p>Conclusions: Biosimilar drugs do not always work well for rheumatic diseases. We obtained two imbalanced machine learning models to detect those cases, where the drug should be removed or where the activity of the disease increases from low to high. Our decision trees use variables, such as age or switching, not considered in previous studies.</p>
Machine Learning Classification Workflow and Datasets for Ionospheric VLF Data Exclusion
<p><span>This data includes the pre-processed dataset, along with a novel workflow that utilizes the PyCaret library and a post-processing workflow. The code and data serve educational purposes in the interdisciplinary field of machine learning and ionospheric physics science, as well as being useful to other researchers for diverse objectives. </span></p> <p><span><span>Acknowledgements:</span></span></p> <p><span><span>The WALDO database (<a href="https://waldo.world"><span>https://waldo.world</span></a>, accessed on October 1, 2023) provides VLF data. It is run collaboratively by the University of Colorado Denver and the Georgia Institute of Technology, utilizing data gathered from Stanford University and those two institutions. It has been made possible by numerous grants from the Department of Defense, NASA, and the NSF.<span> </span></span></span></p> <p> </p> <p> </p>
Using photodiodes and supervised Machine Learning for automatic classification of weld defects in laser welding of thin foils copper-to-steel battery tabs
<p>In this folder, excel files are stored with the results of signal processing that supported findings in the following paper:</p> <p>"Using photodiodes and supervised Machine Learning for automatic classification of weld defects in laser welding of thin foils copper-to-steell battery tabs".</p> <p>Matlab scripts and orginal signals will be uploaded soon with more detailed description.</p> <p> </p>
Machine learning-based pulse wave analysis for classification of circle of Willis topology: an in silico study with 30,618 virtual subjects (database: Complete CoW)
<p>This repository contains the dataset for the complete CoW described in the article with the same name. MATLAB and Python codes for post-processing the dataset and the code for training and testing all machine learning models using the open-source library TensorFlow 2.12, the Keras application programming interface, and the Scikit-learn Python package can be found in here (<a href="https://zenodo.org/records/12519322" target="_blank" rel="noopener">https://zenodo.org/records/12519322</a>).</p>
Data from: OoCount: A machine-learning based approach to mouse ovarian follicle counting and classification
<p>The number and distribution of ovarian follicles in each growth stage provides a reliable readout of ovarian health and function. Leveraging techniques for three-dimensional (3D) imaging of ovaries in toto has the potential to uncover total, accurate ovarian follicle counts. However, because of the size and holistic nature of these images, counting oocytes is time consuming and difficult. The advent of deep-learning algorithms has allowed for the rapid development of ultra-fast, automated methods to analyze microscopy images. In recent years, these pipelines have become more user-friendly and accessible to non-specialists. We used these tools to create OoCount, a high-throughput, open-source method for automatic oocyte segmentation and classification from fluorescent 3D microscopy images of whole mouse ovaries using a deep-learning convolutional neural network (CNN) based approach. We developed a fast clearing and spinning disk confocal-based imaging protocol to obtain 3D images of whole mount perinatal and adult mouse ovaries. Then, fluorescently labeled oocytes from 3D images of ovaries were manually annotated to develop a machine learning training dataset. This dataset was used to train a CNN to automatically label all oocytes in the ovary. In a second phase, we trained another CNN to classify labeled oocytes and sort them into growth stages. Using OoCount, we can obtain accurate counts of oocytes in each growth stage in the perinatal and adult ovary, improving our ability to study ovarian function and fertility. Here, we provide an end-to-end protocol for developing high quality 3D images of the perinatal and adult mouse ovary, obtaining follicle counts and stages, and how to customize OoCount to fit images produced in any lab.</p>
Training data for "Machine learning: classification and regression"
<p>The data provided here are part of a Galaxy Training Network tutorial for "Machine learning: classification and regression".</p>
Machine learning classification of archaea and bacteria identifies novel predictive genomic features
<p>Dataset used for the classification analysis of archaea and bacteria based on 77 genomic features calculated using GBRAP (GenBank Retrieving, Analyzing and Parsing) tool (Vischioni, C. et al. Gbrap: a tool to retrieve, parse and analyze genbank files of viral and bacterial species. bioRxiv 2021–09 (2021)).</p>
Evaluating Machine Learning Models for Supernova Gravitational Wave Signal Classification
<p>This dataset contains gravitational wave (GW) data used in our research work <a href="https://doi.org/10.1088/2632-2153/ada33a" target="_blank" rel="noopener">Abylkairov et al. (2024)</a>. The first 10,000 columns represent the gravitational wave strain <em>D · h</em> [cm] for the corresponding time values ranging from -993 ms to 6.9 ms, with a step size of 0.1 ms. The zero time refers to the time of core bounce. Each row within these first 10,000 columns corresponds to 864 different gravitational wave signals.</p> <p>Columns 10,001 to 10,005 contain the following additional parameters:</p> <ul> <li><strong>T/|W|</strong>: The rotational parameter.</li> <li><strong>GR_or_GREP</strong>: Binary indicator for the signal type, where 0 denotes GR and 1 denotes GREP.</li> <li><strong>EOS</strong>: The equation of state (EOS) model, where 0 corresponds to SFHo, 1 to LS220, 2 to HSDD2, and 3 to GShenFSU2.1.</li> <li><strong>f_peak</strong>: The peak frequency [Hz].</li> <li><strong>D Delta h</strong>: <em>D · ∆h</em> [cm].</li> </ul> <p>For each row (representing a single gravitational wave signal), these parameters provide information about the signal's rotational parameter, type (GR or GREP), EOS model, peak frequency, and <em>D · ∆h</em>.</p> <p><strong>Note</strong>: In the f_peak calculation procedure, we truncated the GW signal at 4.5 ms after the end of the core bounce (see <a href="https://doi.org/10.1103/PhysRevD.95.063019">Richers et al. (2017)</a> for details).</p>
Additional files for Horvath et al., 2024. Detection and classification of long terminal repeat sequences in plant LTR-retrotransposons and their analysis using explainable machine learning.
<p>Additional data for Horvath et al., 2024 (source code freeze, models, data, supplementary figures, tables and files(.</p>
Generating a Labeled Dataset to Train Machine Learning Algorithms for Lithological Classification of Drill Cuttings
<p>This dataset contains 16,700 fully labeled SEM images of rock chips isolated from 14 thin sections of drill cutting samples. These samples come from a low-permeability reservoir in western Canada.</p>
ECG and EEG stress features for: ECG and EEG based detection and multilevel classification of stress using machine learning for specified genders: A preliminary study
<p>Mental health, especially stress, plays a crucial role in the quality of life. During different phases (luteal and follicular phases) of the menstrual cycle, women may exhibit different responses to stress from men. This, therefore, may have an impact on stress detection and classification accuracy of machine learning models that genders are not taken into account. However, this has never been investigated before. In addition, only a handful of stress detection devices are scientifically validated. To this end, this work proposes stress detection and multilevel stress classification models for unspecified and specified genders through ECG and EEG signals. Models for stress detection are achieved through developing and evaluating multiple individual classifiers. On the other hand, stacking technique is employed to obtain models for multilevel stress classification. ECG and EEG features extracted from 40 subjects (21 females and 19 males) were used to train and validate the models. In the low&high combined stress condition, RBF-SVM and kNN yielded the highest average classification accuracy for females (79.81%) and males (73.77%), respectively. Combining ECG and EEG, the average classification accuracy increased to at least 87.58% (male, high stress) and up to 92.70% (female, high stress). For multilevel stress classification from ECG and EEG, the accuracy for females was 62.60% and for males was 71.57%. This study shows that the difference in genders influences the classification performance for both the detection and multilevel classification of stress. The developed models can be used for both personal (through ECG) and clinical (through ECG and EEG) stress monitoring with and without taking genders into account.</p>
Data and code example for the article: "Massively parallel hybrid quantum-classical machine learning for kernelized time-series classification"
<p>Data needed to reproduce the figures of <a href="https://arxiv.org/abs/2305.05881">https://arxiv.org/abs/2305.05881</a> and a simple code example of a quantum-convex-classical neural network used to train a sine versus cosine classification problem.</p>
ECG and EEG stress features for: ECG and EEG based detection and multilevel classification of stress using machine learning for specified genders: A preliminary study
Open the record for dataset details and reuse information.
Data from: OoCount: A machine-learning based approach to mouse ovarian follicle counting and classification
Open the record for dataset details and reuse information.
Identifying galaxies, quasars and stars with machine learning: a new catalogue of classifications for 111 million SDSS sources without spectra
<p>The Paper: <a href="https://arxiv.org/abs/1909.10963">https://arxiv.org/abs/1909.10963</a></p> <p>Abstract: We used 3.1 million spectroscopically labelled sources from the Sloan Digital Sky Survey (SDSS) to train an optimised random forest classifier using photometry from the SDSS and the Widefield Infrared Survey Explorer (WISE). We applied this machine learning model to 111 million previously unlabelled sources from the SDSS photometric catalogue which did not have existing spectroscopic observations. Our new catalogue contains 50.4 million galaxies, 2.1 million quasars, and 58.8 million stars. We provide individual classification probabilities for each source, with 6.7 million galaxies (13%), 0.33 million quasars (15%), and 41.3 million stars (70%) having classification probabilities greater than 0.99; and 35.1 million galaxies (70%), 0.72 million quasars (34%), and 54.7 million stars (93%) having classification probabilities greater than 0.9. Precision, Recall, and F1 score were determined as a function of selected features and magnitude error. We investigate the effect of class imbalance on our machine learning model and discuss the implications of transfer learning for populations of sources at fainter magnitudes than the training set. We used a non-linear dimension reduction technique (Uniform Manifold Approximation and Projection: UMAP) in unsupervised, semi-supervised, and fully-supervised schemes to visualise the separation of galaxies, quasars, and stars in a two-dimensional space. When applying this algorithm to the 111 million sources without spectra, it is in strong agreement with the class labels applied by our random forest model.</p> <p>When using this dataset, please reference our paper via the journal (<a href="https://arxiv.org/abs/1909.10963">https://arxiv.org/abs/1909.10963</a>) and this DOI (10.5281/zenodo.3459293). If you make use of our scripts please reference our Github repository DOI (10.5281/zenodo.3855160).</p> <p>File descriptions:</p> <p>All of these files are Pandas Dataframes, saved as pickle files. df_spec_classprobs.pkl contains the spectroscopically observed sources used for training and testing. This has been cleaned, and has the results of the random forest classifier added as additional columns (sources used for training have NaNs in the class_pred column). SDSS-ML-all contains the 111 million photometrically observed sources, with our class labels and probabilities added. SDSS-ML-galaxies/quasars/stars is the same file broken up by assigned class for convenience.</p>
Data from: Learning to see the wood for the trees: machine learning, decision trees and the classification of isolated theropod teeth
Taxonomic identification of fossils based on morphometric data traditionally relies on the use of standard linear models to classify such data. Machine learning and decision trees offer powerful alternative approaches to this problem but are not widely used in palaeontology. Here, we apply these techniques to published morphometric data of isolated theropod teeth in order to explore their utility in tackling taxonomic problems. We chose two published datasets consisting of 886 teeth from 14 taxa and 3020 teeth from 17 taxa, respectively, each with five morphometric variables per tooth. We also explored the effects that missing data have on the final classification accuracy. Our results suggest that machine learning and decision trees yield superior classification results over a wide range of data permutations, with decision trees achieving accuracies of 96% in classifying test data in some cases. Missing data or attempts to generate synthetic data to overcome missing data seriously degrade all classifiers predictive accuracy. The results of our analyses also indicate that using ensemble classifiers combining different classification techniques and the examination of posterior probabilities is a useful aid in checking final class assignments. The application of such techniques to isolated theropod teeth demonstrate that simple morphometric data can be used to yield statistically robust taxonomic classifications and that lower classification accuracy is more likely to reflect preservational limitations of the data or poor application of the methods.
Data and code for comparison of different machine learning methods and dimensionality reduction for classification astrocytoma and glioblastoma tissues by mass spectra
<p>This upload contains all replication material for "Comparison of different machine learning methods and dimensionality reduction for classification astrocytoma and glioblastoma tissues by mass spectra" (forthcoming).</p> <p><strong>Authors:</strong> E.S. Zhvansky, A.A. Sorokin, V.A. Shurkhay, V.A. Eliferov, D.S. Bormotov, D.G. Ivanov, D.S. Zavorotnyuk, A.A. Potapov.</p> <p><strong>Code and data are located within data_and_code.zip.</strong> Code is written in Python 3.7.7 using Jupyter Notebook, MATLAB R2019b, and Python 3.5.2.</p> <p>Please find the readme.txt for code using and the code to replicate the main findings of the paper described below:</p> <ul> <li>venn_diagramm.py for Venn diagram figures.</li> <li>SSM.m for SSM calculation and visualization.</li> <li>DR_ML.ipynb for dimensionality reduction and machine learning algorithms comparing on the datasets.</li> </ul> <p> </p>
Supplementary Information: CHAPTER 3 - Classification of genomic features of plant-associated bacteria using machine learning
<p>Appendix A- List of all bacterial genomes used in orthologous genes clustering in the feature extraction step and in the further steps to build and test classifiers’ models. The list includes the isolation source information and the related category for the genome classification and features selection purposes.</p> <p>Appendix B - Distribution of genomes by phylum, family, and genus among the categories defined according to bacteria lifestyle association.</p> <p>Appendix C - Enriched orthogroups by genus according to each enrichment test (Material and Methods). Values for each test are "Y" (enriched), "N" (not enriched), or "Untested" (clusters were untested when there was insufficient phylogenetic signal, they were too small or were found in all genomes).</p> <p>Appendix D - Classification performance of random forest and logistic regression techniques applied to genus-specific datasets of genomic features (orthogroups) using both matrices from gene count number and presence/absence values. Sensitivity is a measure of how well a test identifies true positives; Specificity: is a measure how well a test or model avoids false positives; Positive Predictive Value (Pos. Pred. Value): The probability that a positive prediction is correct; Negative Predictive Value (Neg. Pred. Value): The probability that a negative prediction is correct; Precision: The accuracy of positive predictions; Recall (Sensitivity): The ability to find all relevant cases; F1 Score: A combined measure of precision and recall; Prevalence: The proportion of positive cases in the total; Detection Rate: The proportion of true positive cases identified; Detection Prevalence: The proportion of positive predictions; Balanced Accuracy: An average of sensitivity and specificity; Area Under the Curve (AUC): The overall performance of the model in distinguishing between positive and negative cases.</p> <p>Appendix E - Orthogroups assigned with predicted COGs as an important feature for classifying plant-associated genomes. COG categories: A - RNA processing and modification; B - Chromatin structure and dynamics; C - Energy production and conversion; D - Cell cycle control, cell division, chromosome partitioning; E - Amino acid transport and metabolism; F - Nucleotide transport and metabolism; G - Carbohydrate transport and metabolism; H - Coenzyme transport and metabolism; I - Lipid transport and metabolism; J - Translation, ribosomal structure and biogenesis; K - Transcription; L - Replication, recombination and repair; M - Cell wall/membrane/envelope biogenesis; N - Cell motility; O - Posttranslational modification, protein turnover, chaperones; P - Inorganic ion transport and metabolism; Q - Secondary metabolites biosynthesis, transport and catabolism; R - General function prediction only; S - Function unknown; T - Signal transduction mechanisms; U - Intracellular trafficking, secretion, and vesicular transport; V - Defense mechanisms; W - Extracellular structures; X - Mobilome: prophages, transposons; Y - Nuclear structure; Z - Cytoskeleton.</p>
K-Nearest-Neighbor algorithm to predict the survival time and classification of various stages of Oral Cancer: A machine learning approach
<p>This project predicts the survival time of a cancer patient in terms of the number of days and also classifies the dataset into various stages of cancer</p> <p>This is executed on the SPYDER platform using Python 3.7 on Anaconda Navigator.</p> <p>The dataset includes the oral cancer patient's record of 4 countries</p> <p> </p>
Machine learning-based pulse wave analysis for classification of circle of Willis topology: an in silico study with 30,618 virtual subjects (database: Missing ACoA)
<p>This repository contains the dataset for the Missing ACoA described in the article with the same name. MATLAB and Python codes for post-processing the dataset and the code for training and testing all machine learning models using the open-source library TensorFlow 2.12, the Keras application programming interface, and the Scikit-learn Python package can be found in here (<a href="https://zenodo.org/records/12519322" target="_blank" rel="noopener">https://zenodo.org/records/12519322</a>).</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.