Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,943
datasets available to search
ShareScore release 0.9.0
Dataset results
1,943 results for “machine learning”
Dataset related to article "Multicentric evaluation of a machine learning model to streamline the radiotherapy patient specific quality assurance process"
<p> This record contains raw data related to article “Multicentric evaluation of a machine learning model to streamline the radiotherapy patient specific quality assurance process"</p> <p><em>Purpose:</em> Patient-specific quality assurance (PSQA) is performed to ensure that modulated treatment plans can be delivered as intended, but constitutes a substantial workload that could slow down the radiotherapy process and delay the start of clinical treatments. In this study, we investigated a machine learning (ML) tree-based ensemble model to predict the gamma passing rate (GPR) for volumetric modulated arc therapy (VMAT) plans.</p> <p><em>Materials and Methods:</em> 5622 VMAT plans from multiple treatment sites were selected from a database of Institution 1 and the ML model trained using 19 metrics. PSQA analyses were performed automatically using criteria 3%/1 mm (global normalization, absolute dose, 10% threshold) and 95% action limit. Model’s performance was evaluated on an out-of-sample test set of Institution 1 and on two independent sets of measurements collected at Institution 2 and Institution 3. Mean absolute error (MAE), as well as the model’s sensitivity and specificity, were computed.</p> <p><em>Results:</em> The model obtained a MAE of 2.33%, 2.54% and 3.91% for the three Institutions, with a specificity of 0.90, 0.90 and 0.68, and a sensitivity of 0.61, 0.25, and 0.55, respectively. Small positive median values of the residuals (i.e., the difference between measurements and predictions) were observed for each Institution (0.95%, 1.66%, and 3.42%). Thus, the model’s predictions were, on average, close to the real values and provided a conservative estimation of the GPR.</p> <p><em>Conclusions</em>: ML models can be integrated into clinical practice to streamline the radiotherapy workflow, but they should be center-specific or thoroughly verified within centers before clinical use.</p> <p> </p>
Exploring Machine Learning-Based Methods for anomalies detection: Evidence from cryptocurrencies returns
<p>The data consists of 4500 observation for each of the 6 cryptocurrencies.</p>
Deciphering Complex Antibiotic Resistance Patterns in Helicobacter pylori Through Whole Genome Sequencing and Machine Learning
<p>Helicobacter pylori affects billions of people worldwide. Despite the availability of different antibiotics, emerging resistance of H. pylori renders antibiotic treatment ineffective. Next generation sequencing provides a powerful technology to investigate the genotype-phenotype connection for H. pylori. However, the prediction of antibiotic resistance using whole genome sequencing data remains a formidable challenge. Here we conducted a comprehensive investigation into the antibiotic resistance profiles of H. pylori strains against five distinct antibiotics, alongside assessing clinical treatment outcomes for Amoxicillin and Clarithromycin combination therapy. Concurrently, we performed whole-genome sequencing on a collection of H. pylori isolates. We rigorously evaluated the potential for predicting antibiotic resistance through univariate statistical tests, multivariate unsupervised and supervised machine learning. Our study contributes valuable insights towards enhancing precision and effectiveness in antibiotic treatment strategies for H. pylori infections with the application of whole-genome sequencing for H. pylori.</p>
Automated Generation of Complex Bugs in the Machine Learning Era
<p>Artifacts for reproducing BugFarm.</p>
Multi-omics and machine learning reveal context-specific gene regulatory activities of PML-RARA in Acute Promyelocytic Leukemia [APL PBMCs ATAC-seq]
GEO Series GSE215101. Homo sapiens. 2 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.
MRI radiomics-based machine-learning classification of bone chondrosarcoma
<p><strong>Purpose: </strong>To evaluate the diagnostic performance of machine learning for discrimination between low-grade and high-grade cartilaginous bone tumors based on radiomic parameters extracted from unenhanced magnetic resonance imaging (MRI).</p> <p><strong>Methods: </strong>We retrospectively enrolled 58 patients with histologically-proven low-grade/atypical cartilaginous tumor of the appendicular skeleton (n = 26) or higher-grade chondrosarcoma (n = 32, including 16 appendicular and 16 axial lesions). They were randomly divided into training (n = 42) and test (n = 16) groups for model tuning and testing, respectively. All tumors were manually segmented on T1-weighted and T2-weighted images by drawing bidimensional regions of interest, which were used for first order and texture feature extraction. A Random Forest wrapper was employed for feature selection. The resulting dataset was used to train a locally weighted ensemble classifier (AdaboostM1). Its performance was assessed via 10-fold cross-validation on the training data and then on the previously unseen test set. Thereafter, an experienced musculoskeletal radiologist blinded to histological and radiomic data qualitatively evaluated the cartilaginous tumors in the test group.</p> <p><strong>Results: </strong>After feature selection, the dataset was reduced to 4 features extracted from T1-weighted images. AdaboostM1 correctly classified 85.7 % and 75 % of the lesions in the training and test groups, respectively. The corresponding areas under the receiver operating characteristic curve were 0.85 and 0.78. The radiologist correctly graded 81.3 % of the lesions. There was no significant difference in performance between the radiologist and machine learning classifier (P = 0.453).</p> <p><strong>Conclusions: </strong>Our machine learning approach showed good diagnostic performance for classification of low-to-high grade cartilaginous bone tumors and could prove a valuable aid in preoperative tumor characterization.</p>
Data set from Machine learning to predict in-hospital mortality in covid-19 patients using computed tomography-derived pulmonary and vascular features
<p>Data set from Machine learning to predict in-hospital mortality in covid-19 patients using computed tomography-derived pulmonary and vascular features</p>
Data set from Machine learning to predict in-hospital mortality in covid-19 patients using computed tomography-derived pulmonary and vascular features
<p>Data Set from the study Machine learning to predict in-hospital mortality in covid-19 patients using computed tomography-derived pulmonary and vascular features</p>
Identifying lumbar fragility fractures: a comparison of traditional machine learning and Deep Learning
<p>The class-0 folder contains ROIs belonging to a non-fracture group.</p> <p>The Class-1 folder contains ROIs belonging to the fractured group Features folder.</p>
Machine Learning Based Digital Twin in Manufacturing: A Bibliometric Analysis and Evolutionary Overview
<p>The data files include bibliometric file. data extracted from bibliometric data.</p>
Revisiting Machine Learning based Test Case Prioritization for Continuous Integration
<p>This repository contains a replication package for a research paper submitted to the 45th International Conference on Software Engineering (https://conf.researchr.org/home/icse-2023). We provide our code, data, and result for the ease of replicating our experiments.</p> <p><strong>Code</strong></p> <ul> <li>In the <strong>collect_data</strong> subdirectory, scripts for constructing TCP datasets are provided. We do dependency analysis using Understand (https://www.scitools.com/), so please download the related tools in advance.</li> <li>In the<strong> rl</strong> subdirectory, we provide the python implementation for algorithms RL, COLEMAN, PPO2-PO, ACER-PA, PPO1-LI.</li> <li>In the <strong>supervised_learning</strong> subdirectory, we provide implementations for MART, RankNet, RankBoost, CA, L-MART, which mainly rely on Ranklib (https://sourceforge.net/p/lemur/wiki/RankLib/.). We also provide implementation for DeepOrder.</li> </ul> <p><strong>Data</strong></p> <ul> <li>The <strong>origin</strong> subdirectory contains the original datasets collected from github using our scripts, including 11 projects.</li> <li>The <strong>smote</strong> subdirectory contains the datasets pre-processed by SMOTE.</li> </ul> <p><strong>Result</strong></p> <ul> <li>Results for <strong>RQ1</strong>, <strong>RQ2</strong>, <strong>RQ3</strong>, and <strong>threats to validity</strong> are provided in the corresponding subdirectories. Scripts for plotting figures are also provided.</li> </ul> <p><strong>Reference</strong></p> <p>We adopt code from previous work</p> <p>Learning-to-Rank vs Ranking-to-Learn: Strategies for Regression Testing in Continuous Integration (https://dl.acm.org/doi/abs/10.1145/3377811.3380369) Github repository: https://github.com/icse20/RT-CI</p> <p>Reinforcement Learning for Test Case Prioritization (https://ieeexplore.ieee.org/abstract/document/9394799) Github repository: https://github.com/moji1/tp_rl</p> <p>DeepOrder: Deep Learning for Test Case Prioritization in Continuous Integration Testing (https://ieeexplore.ieee.org/abstract/document/9609187) Github repository: https://github.com/AizazSharif/DeepOrder-ICSME21</p> <p>A Multi-Armed Bandit Approach for Test Case Prioritization in Continuous Integration Environments (https://ieeexplore.ieee.org/abstract/document/9086053) Github repository: https://github.com/jacksonpradolima/coleman4hcs</p>
Particle characterization by analyzing light scattering signals with a machine learning approach.
<p>This container includes the measurement data and trained machine learning models associated with the publication: "Particle Characterization by Analyzing Light Scattering Signals Using a Machine Learning Approach."</p> <p>There are three types of data, each marked with a specific prefix:<br>- <strong>Data_</strong><br>- <strong>Pred Data_</strong><br>- <strong>Test Data_</strong></p> <p>The files with the prefix <strong>Data_</strong> contain the data used to train the machine learning model.</p> <p>The files with the prefix <strong>Test Data_</strong> contain the data used to test the machine learning model.</p> <p>The files with the prefix <strong>Pred Data_</strong> contain the results generated after the test data was applied to the machine learning model.</p> <p>Additionally, there are pre-trained machine learning models with the prefix <strong>Model_</strong>.</p> <p>Moreover, a Wolfram Mathematica script is included for training and testing the machine learning models. The script <strong>Script Wolfram Mathematica</strong> is added in <strong>.nb</strong> und in <strong>.pdf</strong> formats.</p> <p><strong>The use of the data is permitted only for academic purposes and not for any commercial purposes.</strong></p>
Dataset related to article: "Machine learning to predict mortality after rehabilitation among patients with severe stroke"
<p>We provide the raw data used for the following article:</p> <ul> <li>Scrutinio D, Ricciardi C, Donisi L, Losavio E, Battista P, Guida P, Cesarelli M, Pagano G, D'Addio G.<br> <em>Machine learning to predict mortality after rehabilitation among patients with severe stroke.</em> "Sci Rep." 2020 Nov 18;10(1):20127.<br> doi: 10.1038/s41598-020-77243-3. PMID: 33208913; PMCID: PMC7674405.</li> </ul> <p> </p> <p><strong>Abstract:</strong> Stroke is among the leading causes of death and disability worldwide. Approximately 20–25% of stroke survivors present severe disability, which is associated with increased mortality risk. Prognostication is inherent in the process of clinical decision-making. Machine learning (ML) methods have gained increasing popularity in the setting of biomedical research. The aim of this study was twofold: assessing the performance of ML tree-based algorithms for predicting three-year mortality model in 1207 stroke patients with severe disability who completed rehabilitation and comparing the performance of ML algorithms to that of a standard logistic regression. The logistic regression model achieved an area under the Receiver Operating Characteristics curve (AUC) of 0.745 and was well calibrated. At the optimal risk threshold, the model had an accuracy of 75.7%, a positive predictive value (PPV) of 33.9%, and a negative predictive value (NPV) of 91.0%. The ML algorithm outperformed the logistic regression model through the implementation of synthetic minority oversampling technique and the Random Forests, achieving an AUC of 0.928 and an accuracy of 86.3%. The PPV was 84.6% and the NPV 87.5%. This study introduced a step forward in the creation of standardisable tools for predicting health outcomes in individuals affected by stroke.</p>
Calibration of Polyvinylidene fluoride (PVDF) stress gauges under high-impact dynamic compression by machine learning
<p><br> The shared materials are for calibration of PVDF stress gauges under high-impact dynamic compression by machine learning. Details will be found in the submitted paper to JAP, with a DOI to be updated in the next version. </p> <p>‘data.csv’: experimental data used for machine learning<br> ‘training.jpynb’: notebook to train the model <br> ‘model.joblib’: one example of the trained model </p> <p>The model is used for calibrating the PVDF stress gauge from 0.3 to 10 GPa, best appropriate with remnant polarization from 6.7 to 8.3 μC/cm2 and active sensing thickness from 20 to 30 μm. <br> </p>
Bacteria-Specific Features Selection for Enhanced Antimicrobial Peptide Activity Predictions Using Machine-Learning Methods
<p>We developed a new computational approach that allowed us to train several supervised machine-learning models using a specific set of data associated with peptides targeting E. coli bacteria. LASSO regression and Support Vector Machine techniques have been utilized to select, among more than 1500 physio-chemical descriptors, the most important features that can be used to classify a peptide as antimicrobial or ineffective against E. coli. We then performed the classification of active versus inactive AMPs using the Support Vector classifiers, Logistic Regression, and Random Forest methods. This computational study allows us to make recommendations of how to design more efficient anti-bacterial drug therapies.</p>
Octopus: A Novel Approach for Health Data Masking and Retrieving using Physically Unclonable Function and Machine Learning
<p>The health equipment is used to keep track of significant health indicators, automate health interventions, and analyze health indicators. People have begun using mobile applications to track health characteristics and medical demands because all devices are linked to high-speed internet and phones. Such a combination of smart devices, the internet, and mobile applications expands the usage of remote health monitoring through the Internet of Medical Things (IoMT). The accessibility and unpredictable aspects of IoMT create massive security and confidentiality threats in IoMT systems. In this proposed paper - Octopus, Physically Unclonable Functions (PUFs) have been used to provide privacy to the healthcare device by masking the data, and machine learning (ML) techniques are used to retrieve the health data back and reduce security breaches on networks. This technique has exhibited 99.45% accuracy, which proves that this technique could be used to secure health data with masking.</p>
MolToxPred: Small molecule toxicity prediction using machine learning approach
<p>SMILES data of training, test set, and another dataset with 180 molecules of the external validation set. Data has been curated from various sources as mentioned in the manuscript.</p>
Machine_Learning_to_Hardware_for_Instrumentation
<p>Data used in "Exploring_Machine_Learning_to_Hardware_for_Instrumentation" paper, which is under review by IOP Machine learning: Science and technology journal</p>
Machine Learning in Hypertension Detection: A Study on World Hypertension Day Data
<p>Many modifiable and non-modifiable risk factors have been associated with hypertension. However, current screening programs are still failing in identifying individuals at higher risk of hypertension. Given the major impact of high blood pressure on cardiovascular events and mortality, there is an urgent need to find new strategies to improve hypertension detection. We aimed to explore whether a machine learning (ML) algorithm can help identifying individuals predictors of hypertension. We analysed the data set generated by the questionnaires administered during the World Hypertension Day from 2015 to 2019. A total of 20206 individuals have been included for analysis. We tested five ML algorithms, exploiting different balancing techniques. Moreover, we computed the performance of the medical protocol currently adopted in the screening programs. Results show that a gain of sensitivity reflects in a loss of specificity, bringing to a scenario where there is not an algorithm and a configuration which properly outperforms against the others. However, Random Forest provides interesting performances (0.818 sensitivity - 0.629 specificity) compared with medical protocols (0.906 sensitivity - 0.230 specificity). Detection of hypertension at a population level still remains challenging and a machine learning approach could help in making screening programs more precise and cost effective, when based on accurate data collection. More studies are needed to identify new features to be acquired and to further improve the performances of ML models.</p>
Predicting SARS-CoV-2 exposure using T-cell repertoire sequencing and machine learning
<p>The dataset contains processed T-cell receptor repertoire sequencing data from >1200 individuals of different sex and age, originally published in [1]. Note that only samples with good sequencing coverage are published (>10^5 reads per file). </p> <p>TODO</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.