Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,943
datasets available to search
ShareScore release 0.9.0
Dataset results
1,943 results for “machine learning”
An Integrated Approach for enhanced SMAP Soil Moisture Retrieval: Multi-Source Data Fusion and Data-Driven Machine Learning
<p><span>Accurate satellite-based soil moisture (SM) retrieval is essential for hydrometeorological and agroecological applications, yet traditional physical models for L-band SM retrieval are hindered by uncertainties stemming from inaccuracies in prior parameters. This work combines multi-source data fusion and a physically-guided machine learning framework to develop a Soil Moisture Active Passive (SMAP) SM retrieval model (Fusion-LightGBM, F-LGB) that bypasses the need for static prior parameters, resulting in a new SM product. The retrieval benchmark is a new seamless SM data constructed by combining Triple Collection correlation coefficients (TC-R) and the Maximized-R method, which demonstrates superior temporal correlation on 20 International Soil Moisture Network (ISMN) <em>in-situ</em> networks compared to existing SM data, including ECMWF Reanalysis v5-Land (ERA5-Land), SMAP Level 4 (SMAP L4), and Global Land Data Assimilation System (GLDAS) Noah. The machine learning model incorporates input variables that represent the Tau-Omega model’s radiative transfer process, including brightness temperature, vegetation optical depth, soil temperature, and an external variable for precipitation. In the 2015-2020 validation set, F-LGB demonstrated the highest correlation (mean R = 0.72, significantly surpassing the second-best SMAP-INRAE-BORDEAUX (SMAP-IB) SM and deep neural network (DNN) SM at 0.67) and the lowest ubRMSE (mean value of 0.052 m<sup>3</sup>/m<sup>3</sup>, better than 0.055 m<sup>3</sup>/m<sup>3</sup> for both DNN and SMAP-IB). F-LGB performed well across diverse land covers, vegetation densities, and climates, with SHAP analysis showing H-polarized brightness temperature as crucial, especially in areas with low to moderate vegetation. This new machine learning-based SMAP SM product may improve global satellite-based SM estimation capabilities.</span></p>
Data for "Supervised machine learning methods for crystal structure prediction of the binary Cs-Te system"
<p>Crystal structures, high-throughput calculations and trained machine learning models presented in the paper "Supervised machine learning methods for crystal structure prediction of the binary Cs-Te system".</p> <ul> <li><em>crystal_datasets </em>contains the input/output data sets of crystal structures for high-throughput calculations and ML models.</li> <li><em>aiida_ht_calculations </em>contains the data regarding the high-throughput DFT calculations.</li> <li><em>ml_models</em> contains the trained ML models.</li> </ul> <p>Eeach zip-archive contains a jupyter-notebook examplifying how the data can be accessed and reused.</p>
Prediction model of in-hospital mortality in intensive care unit patients with heart failure: machine learning-based, retrospective analysis of the MIMIC-III database
<p><b>Objective:</b> The predictors of in-hospital mortality for intensive care units (ICU)-admitted HF patients remain poorly characterized.We aimed to develop and validate a prediction model for all-cause in-hospital mortality among ICU-admitted HF patients.</p> <p><b>Design: </b>A retrospective cohort study.</p> <p><b>Setting and Participants: </b>Data were extracted from the MIMIC-III database. Data on 1,177 heart failure patients were analysed.</p> <p><strong>Methods</strong>: Patients meeting the inclusion criteria were identified from the MIMIC-III database and randomly divided into derivation and validation groups. Independent risk factors for in-hospital mortality were screened using XGBoost and LASSO regression models in the derivation sample. Multivariable logistic regression analysis was used to build prediction models. Discrimination, calibration, and clinical usefulness of the predicting model were assessed using the C-index, calibration plot, and decision curve analysis. After pairwise comparison, the best performing model was chosen to build a nomogram according to the regression coefficients.</p> <p><b>Results:</b> Among the 1,177 admissions, in-hospital mortality was 13.52%. In both groups, the XGBoost, LASSO regression, and GWTG-HF risk score models showed acceptable discrimination. The XGBoost and LASSO regression models also showed good calibration. In pairwise comparison, the prediction effectiveness was higher with the XGBoost and LASSO regression models than with the GWTG-HF risk score model (P<0.05). The XGBoost model was chosen as our final model for its more concise and wider net benefit threshold probability range and was presented as the nomogram.</p> <p><b>Conclusions</b><b>:</b> Our nomogram enabled good prediction of in-hospital mortality in ICU-admitted HF patients, which may help clinical decision-making for such patients.</p>
Comparison and assessment of different object-based classifications using machine learning algorithms and UAVs multispectral imagery in the framework of precision agriculture
<p>Supplementary material of the paper</p>
A Hybrid Approach to Atmospheric Modeling that Combines Machine Learning with a Physics-Based Numerical Model
<p>Data used to generate the figures in "A Hybrid Approach to Atmospheric Modeling that Combines Machine Learning with a Physics-Based Numerical Model" 2021. The zip files contains the hybrid forecasts and regridded ERA5 data used to verify the forecasts as well as the SPEEDY-LLR and ML-only runs. </p>
Machine Learning-Assisted Sampling of SERS Substrates Improves Data Collection Efficiency: raw data and code
<p>Raw datasets and media accompanying the manuscript: <strong>Machine Learning-Assisted Sampling of SERS Substrates Improves Data Collection Efficiency</strong>: data, published in <em>Applied Spectroscopy </em>in 2021</p>
List of documents and patterns identified by multivocal literature review of software engineering patterns for machine learning applications
We performed a multivocal literature review of both academic and gray literature to collect software engineering for machine learning (ML) application systems and software design. For the academic literature, we chose Engineering Village. For the gray literature, we used a Google search on August 16, 2019. We retrieved 32 scholarly documents and 48 gray literature documents. We vetted whether each document should be included in our review using the following criteria: Documents written in English addressing concrete software-engineering patterns or practices to design ML application systems and software should be included. Documents focusing on design of ML techniques and algorithms should be excluded. This process identified 19 scholarly documents and 19 gray documents. Although 69 patterns related to the design of ML application systems were initially identified, 33 remained after the vetting process. Finally, industrial ML developers reviewed the 33 candidates from the viewpoint of practical usefulness. They identified only 15 ML patterns.
Clinical Categorization Algorithm (Clical) and Machine-Learning Approach (Srf-clical) to Predict Clinical Benefit to Immunotherapy in Metastatic Melanoma Patients: Real-world Evidence from Istituto Nazionale Tumori Irccs Fondazione Pascale, Napoli, Italy.
<p>Raw-data related to a manuscript submitted to "Cancers" journal - MDPI - https://www.mdpi.com/journal/cancers</p> <p><strong>Manuscript Title</strong>: Clinical Categorization Algorithm (Clical) and Machine-Learning Approach (Srf-clical) to Predict Clinical Benefit to Immunotherapy in Metastatic Melanoma Patients: Real-world Evidence from Istituto Nazionale Tumori Irccs Fondazione Pascale, Napoli, Italy.</p> <p><strong>Authors:</strong> Gabriele Madonna1,#, Giuseppe V. Masucci2,3,#, Mariaelena Capone1, Domenico Mallardo1, Antonio Maria Grimaldi1, Ester Simeone1, Vito Vanella1, Lucia Festino1, Marco Palla1, Luigi Scarpato1, Marilena Tuffanelli1, Grazia D’angelo1, Lisa Villabona2, Isabelle Krakowski2,4, Hanna Eriksson2,3, Felipe Simao5, Rolf Lewensohn2,3, Paolo Antonio Ascierto1,+</p> <p><strong>Affiliations</strong>:</p> <p>1 Cancer Immunotherapy and Development Therapeutics Unit, Istituto Nazionale Tumori IRCCS Fondazione "G. Pascale", Napoli, Italy</p> <p>2 Theme Cancer, Karolinska University Hospital, Stockholm, Sweden</p> <p>3 Department of Oncology-Pathology, Karolinska Institutet, Stockholm, Sweden</p> <p>4 Theme Inflammation, Karolinska University Hospital Stockholm, Sweden</p> <p>5 Genevia technologies OY, Tampere, Finland</p> <p># these authors equally contributed</p> <p>+ Corresponding author</p> <p><strong>Abstract of submitted Manuscript:</strong> The real-life application of immune checkpoint inhibitors (ICI) may yield different outcomes compared to the benefit presented in clinical trials. For this reason, there is a need to define the group of patients that may benefit from treatment. We retrospectively investigated 578 metastatic melanoma patients treated with ICI at Istituto Nazionale Tumori IRCCS Fondazione “G. Pascale” of Napoli Italy (INT-NA). To compare patients’ clinical variables (age, Lactate Dehydrogenase (LDH), Neutrophil-Lymphocyte Ratio (NLR), eosinophil, BRAF status, previous treatment) and their predictive and prognostic power in a comprehensive non-hierarchical way, a Clinical Categorization Algorithm (CLICAL) was defined and validated by the application of machine learning, Survival Random Forest (SRF-CLICAL). The comprehensive analysis of the clinical parameters by log risk-based algorithms convened into predictive signatures that could identify groups of patients with great benefit or not, regardless of the ICI received. From a real-life retrospective analysis of metastatic melanoma patients, we generated and validated an algorithm based on machine learning that could assist with the clinical decision of whether or not to apply ICI therapy by defining five signatures of predictability with a 95% accuracy.</p> <p><strong>Funding: </strong>This research was funded by Italian Ministry of Health (IT-MOH) through “Ricerca Corrente”, grants number M2-2. Additional funding [N#184093) from the Stockholm Cancer Society and King Gustav V’s Jubilee foundation Stockholm.</p>
DSM Water Level: An UAV photogrammetry dataset for determination of river surface level using machine learning
<p>Orthophotos and digital surface models (DSMs) obtained using UAV photogrammetry can be used to determine the water surface level of a river. However, this task is difficult due to disturbances of the water surface on DSMs caused by limitations of photogrammetric algorithms. Machine Learning can be used to correct these disturbances as well as to extract a single water surface elevation value. The presented dataset contains raw photogrammetric orthophotos and DSMs of areas representing parts of a small river and the corresponding DSMs with corrected water surface disturbances. Also a single ground truth value of mean water surface level for each DSM sample is provided. This allows the dataset to be used for supervised training of a neural network performing a denoising or regression task.</p> <p>Acknowledgement: some of the samples were extracted from photogrammetric data acquired by Bandini et. al (https://doi.org/10.5281/zenodo.3519888)</p>
Data archive for paper "Copula-based synthetic data augmentation for machine-learning emulators"
<p><strong>Overview</strong></p> <p>This is the data archive for paper "<a href="https://doi.org/10.5194/gmd-14-5205-2021">Copula-based synthetic data augmentation for machine-learning emulators</a>". It contains the paper’s data archive with model outputs (see <code>results</code> folder) and the Singularity image for (optionally) re-running experiments.</p> <p>For the Python tool used to generate synthetic data, please refer to <a href="https://github.com/dmey/synthia">Synthia</a>.</p> <p><strong>Requirements</strong></p> <ul> <li><a href="https://sylabs.io/singularity/">Singularity</a> >= 3</li> <li><a href="https://en.wikipedia.org/wiki/Portable_Batch_System">Portable Batch System</a> (PBS) job scheduler*</li> <li>Today's high-performance computer (e.g. ~ 32 CPUs @ 2 500 MHz with 64 GB of RAM )</li> </ul> <p>*Although PBS in not a strict requirement, it is required to run all helper scripts as included in this repository. Please note that depending on your specific system settings and resource availability, you may need to modify PBS parameters at the top of submit scripts stored in the <code>hpc</code> directory (e.g. <code>#PBS -lwalltime=72:00:00</code>).</p> <p><strong>Usage</strong></p> <p>To reproduce the results from the experiments described in the paper, first fit all copula models to the reduced NWP-SAF dataset with:</p> <pre><code>qsub hpc/fit.sh</code></pre> <p>then, to generate synthetic data, run all machine learning model configurations, and compute the relevant statistics use:</p> <pre><code>qsub hpc/stats.sh qsub hpc/ml_control.sh qsub hpc/ml_synth.sh</code></pre> <p>Finally, to plot all artifacts included in the paper use:</p> <pre><code>qsub hpc/plot.sh</code></pre> <p><strong>Licence</strong></p> <p>Code released under <a href="./LICENSE.txt">MIT license</a>. Data from the reduced NWP-SAF dataset released under <a href="./data/LICENSE.txt">CC BY 4.0</a>.</p>
Air Quality Forecasts Improved by Combining Data Assimilation and Machine Learning with Satellite AOD
<p>Input data for random forest model. </p> <p> </p> <p>1) UM_RDAPS.egg file: It provides analysis and forecast products four times a day (00, 06, 12, 18 UTC) in 12 km x 12 km spatial resolution. In this study, analysis products were only considered as the input variables (i.e., 2m temperature and dew-point temperature, relative humidity (RH), maximum wind speed, visibility at height above the ground, planetary boundary layer height (PBLH), and surface pressure). The accumulated maximum wind speed during 1, 3, 5, 7 days were also used in this study.</p> <p>2) data_1.zip file: GOCI Aerosol product, MODIS Land cover, MODIS NDVI, Population density, Road density, SRTM_DEM. </p> <p> </p> <p>The detailed information of input variables is written in the supporting information of the paper.</p> <p> </p> <p> </p> <p> </p>
The impact of the cross-docked poses on the performance of machine learning classifier for protein-ligand binding pose prediction
<p>Datasets, features, and some representative scripts utilized in the paper "The impact of the cross-docked poses on the performance of machine learning classifier for protein-ligand binding pose prediction".</p>
Dataset used in "Machine-learning Kondo physics using variational autoencoders"
<p>Spectral functions from the single-impurity Anderson model, generated on a log-linear frequency mesh using the numerical renormalization group algorithm. Used in both https://doi.org/10.1103/PhysRevB.103.245118 (where a detailed description of dataset generation can be found) and https://arxiv.org/abs/2107.08013.</p>
Additional File 1 and Notebook Code for "Machine learning approaches for hospital acquired pressure injuries: a retrospective study of electronic medical records"
<p>Supplementary materials (Pressure_Injuries_Additional_File_1_final_double_blind.pdf) and Jupyter Notebook code (HAPI_Prediction_Script.pdf) developed as supplement for study "Machine learning approaches for hospital acquired pressure injuries: a retrospective study of electronic medical records".</p>
Data-Error Scaling Laws in Machine Learning on Combinatorial Mutation-prone Sets: Proteins and Small Molecules
<div> </div> <h3>Data</h3> <p> This folder contains the raw data used during this work. `out_seq_total.txt` contains information on the sequences used (mutations, number of mutations, etc.). `output_energies_total.txt` contains the response variables, which include: a) unrelaxed EvoEF energies (peptides); b) relaxed EvoEF energies (peptides, `*_repaired.txt`); and c) solvation energies (molecules). 3D structures are provided in `.xyz` format in the subfolder `XYZ` (molecules).</p> <div> <div><strong>GB1 dataset</strong></div> <br> <div>The GB1 dataset was <strong>not</strong> generated by us (https://doi.org/10.48550/arXiv.2405.05167). If you use the GB1 dataset, please cite the original paper:</div> <br> <div>Wu, N. C., Dai, L., Olson, C. A., Lloyd-Smith, J. O., & Sun, R. (2016). <em>Adaptation in protein fitness landscapes is facilitated by indirect paths. </em><strong>eLife</strong>, 5:e16965. doi:10.7554/eLife.16965</div> </div> <h3>Results</h3> <p>This folder contains the results (outputs) of the ML models trained using the provided scripts (see github repository). Such results are incuded in the form of `.npy` files. To load the files please include the option `allow_pickle=True`.</p> <p> Each `.npy` file contains the following keys:<br> * `initial_parameters`: script inputs.<br> * `d_encoder`: encoder used (not always included).<br> * `ns_train`: number of training points used for the LCs (rounded, integers).<br> * `ns_train_float`: number of training points used for the LCs (not rounded, float).<br> * `ns_train_norm`: number of training points used for the LCs (normalized, float).<br> * `res`: test MAEs.<br> * `res_tot`: (train,validation,test) MAEs.<br> * `res_tot_mut`: (train,validation,test) MAEs sorted by mutation number.<br> * `l_opt`: optimal kernel length used during the test.<br> * `ls`: kernel lengths used for grid search.<br> * `idx_seeds`: indices used to reshuffle the data. If one want to rebild the initial order use `np.argsort(idx_seeds)`.<br> * `alpha_opt`: optimal regression parameters used to calculate the test error. To sort the data use `alpha_opt[i][ii][np.argsort(idx_seeds[ii,0:arg_train_max].astype(int)[:ns_train[i]]`. Where `i` is the replicate number (0-99) and ii is the idex in the LC.<br> * `valid_errs`: validation error (MAE) calculated for each point in the hyperparameter (kernel scale) optimisation.<br> * `test_errs`: test error (MAE) calculated for each point in the hyperparameter (kernel scale) optimisation.</p>
All data support the published articel "Loop-optimization of Trichoderma reesei endoglucanases for balancing the activity–stability trade-off through cross-strategy between machine learning and the B-factor analysis"
<p><em>Trichoderma reesei</em> endoglucanases (EGs) have limited industrial applications due to its low thermostability and activity. Here, we aimed to improve the thermostability of EGs from<em> T.reesei</em> without reducing its activity counteracting the activity-stability trade-off. A cross-strategy combination of machine learning and B-factor analysis was used to predict beneficial amino acid substitution in EG loop optimization. Experimental validation showed single-site mutated EG concomitantly improved enzymatic activity and thermal properties by 17.21%–18.06% and 49.85%–62.90%, respectively, compared with wild-type EGs. Furthermore, the mechanism explained mutant variants had lower RMSD values and a more stable overall structure than the wild type. According to this study, EGs loop optimization is crucial for balancing the activity-stability trade-off, which may provide new insights into how loop region function interacts with enzymatic characteristics. Moreover, the cross-strategy between machine learning and B-factor analysis improved superior enzyme activity-stability performance, which integrated structure-dependent and sequence-dependent information.</p>
Data to support "Physics-based representations for machine learning properties of chemical reactions
<p>4 datasets of reaction data: </p> <p>1. SN2-20 dataset adapted from https://iopscience.iop.org/article/10.1088/2632-2153/aba822/meta</p> <p>2. Proparg-21-TS dataset from https://pubs.rsc.org/en/content/articlehtml/2021/sc/d1sc00482d</p> <p>3. GDB7-22-TS dataset from https://www.nature.com/articles/s41597-020-0460-4</p> <p>4. Our Hydroform-22-TS dataset of 2,350 structures of reactant and product structures and associated barriers</p> <p>In all cases, there are xyz files of reactant(s) and product(s) structures, and a csv file of associated properties (reaction energies for the first case, e.e. values for the second, and barriers for the third and fourth).</p> <p>For example usage see https://github.com/lcmd-epfl/b2r2-reaction-rep</p>
Supplementary Materials for 'Spatiotemporal Prediction of Air Quality Using Machine Learning Techniques'
<p>This package includes supplementary materials used to implement air quality prediction in the city of Madrid. It consists of two main subdirectories: Data and Code. The Data directory contains Raw-Data (air quality, meteorological and traffic data from the period of January-June 2019 and January-June 2020, and the location of air quality and meteorological monitoring stations and traffic measurement points of the city of Madrid) and Processed-Data (the output after raws data has gone through the workflow to meet the requirements corresponding to the implementation of the proposed forecasting approaches). The Code directory contains Process Raw Data, Chapter4-ConvLSTM, Chapter5-BiConvLSTM, and Chapter6-A3T_GCN, which provides the procedure for constructing and implementing the proposed approaches.</p>
Dataset for puplication: Machine Learning in Automated Monitoring of Metabolic Changes Accompanying the Differentiation of Adipose Tissue-Derived Human Mesenchymal Stem cells employing 1H-1H TOCSY NMR
<p>Data set used in the publication: <strong>Machine Learning in Automated Monitoring of Metabolic Changes Accompanying the Differentiation of Adipose Tissue-Derived Human Mesenchymal Stem cells employing <sup>1</sup>H-<sup>1</sup>H TOCSY NMR. </strong></p> <p>Abstract: In this work, the dynamic evolution of adipose tissue-derived human MSCs (AT-derived hMSCs) after fourteen days of cultivation, adiobocytes and osteocytes differentiation has been inspected based on 2D NMR TOCSY using machine learning techniques. Multi-class classification in addition to novelty detection of metabolites was established based on the profile of a control hMSCs sample at four days cultivation and successively detect the absence and the abundance of metabolites in differentiated MSCs following a set of <sup>1</sup>H-<sup>1</sup>H TOCSY profiles. The uploaded files are:</p> <p>File: metabolites_names.xlsx contain the names of the used metabolites.</p> <p>File: metabolites.xlsx</p> <p> column 1: metabolite abbreviation</p> <p>column 2: 2D NMR TOCSY horizontal and vertical frequencies of metabolites in the control group at 4 days cultivation (Ct d4)</p> <p>column 3: 2D NMR TOCSY horizontal and vertical frequencies of metabolites found after 14 days of cultivation (Ct d14)</p> <p>column 4: 2D NMR TOCSY horizontal and vertical frequencies of metabolites found after 14 days of differentiation into adipocytes (AT d14)</p> <p>column 5: 2D NMR TOCSY horizontal and vertical frequencies of metabolites found after14 days of differentiation into osteocytes (OS d14)</p> <p>column 6: 2D NMR TOCSY horizontal and vertical standard frequencies of all metabolites measured at broadband high resolution 600.13 MHz NMR</p>
Machine learning reveals a taxonomy of bacterial sensors in Earth ecosystems
<p>Datasets and FASTA files used for the manuscript titled: "Machine learning reveals a taxonomy of bacterial sensors in Earth ecosystems." Description.ipynb contains descriptions and first 5 lines of each file.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.