Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,943
datasets available to search
ShareScore release 0.9.0
Dataset results
1,943 results for “machine learning”
Data for "Machine Learning Scoring Functions for Drug Discovery from Experimental and Computer-generated Protein-Ligand Structures: Towards Per-target Scoring Functions"
<p>Data used in "<em>Machine Learning Scoring Functions for Drug Discovery from Experimental and Computer-generated Protein-Ligand Structures: Towards Per-target Scoring Functions</em>"<br> by F. Pellicani, D. Dal Ben, A. Perali, S. Pilati</p> <p>If you use these data or the python script for your research or other activities, please cite the corresponding journal article.</p> <p> </p> <p>====================</p> <p>Uncompressing the zipped file <em>DataSFUnicam.zip</em> provies the following files and folders:</p> <p><br> <strong>DataSFUnicam/</strong></p> <p> </p> <p> ExperimentalDataPDBFiles/<br> <em>This folder contains 2408 .pdb files of experimental complex structures. The files are named with a univocal code corresponding to the protein-ligand complex.</em></p> <p> </p> <p> ExperimentalDataXLSXFile.xlsx<br> <em>This Excel file reports the experimental protein-ligand chemical information. In the sheet named “Foglio1”, the first column contains the univocal code of the protein-ligand complex, the second column contains the experimentally measured pK_d.</em></p> <p> </p> <p> SyntheticDataPDBFiles/<br> <em>This folder contains the .pdb files of the synthetic complex structures. The .pdb files are grouped in 17 folders according to just as many target proteins. The folders are named after the corresponding protein. Each folder contains the .pdb files for the best position of each protein-ligand pair according to the MOE docking score. The files are named with a univocal code.</em></p> <p> </p> <p> SyntheticDataXLSXFiles/<br> <em> The folder contains 17 Excel files with the chemical information of the synthetic protein-ligand complexes. The files are named after the corresponding target protein. In the sheet named “Foglio1” of each .xlsx file, the first column contains a univocal code of the protein-ligand complex in each conformation, the second column contains an auxiliary numerical code corresponding to the protein-ligand pair, the third column contains the experimentally measured pK_i, and the fourth column contains the docking score provided by the MOE software.</em></p> <p>====================</p> <p>USER GUIDE FOR THE PYTHON SCRIPT</p> <p>Download and uncompress the zipped file "<em>SFUnicam.zip</em>" with a command like "<em>unzip SFUnicam.zip</em>". </p> <p>The following file structure is created:</p> <p><em>SFUnicam/</em></p> <p> <em>ComplexToBePredictedFolder/4ey5_30.pdb <br> MaxAssMatrix.npy<br> my_model<br> devStndSynt.npy<br> mediaSynt.npy<br> UnicamSF13prot.py<br> README.txt</em><br> <br> The subfolder "<em>ComplexToBePredictedFolder/</em>" contains the example PDB file "<em>4ey5_30.pdb</em>".</p> <p>-) To execute the script "<em>UnicamSF13prot.py</em>", Python 3 should be installed with the following libraries and sublibraries:<br> <em>Keras:<br> Regularizers<br> Sequential (keras.models)<br> Conv1D, Dense, MaxPooling1D, GlobalMaxPooling1D, GlobalAveragePooling1D, AveragePooling1D (keras.layers)<br> Adam (keras.optimizers)<br> Numpy</em><br> <em>Tensorflow</em></p> <p>Operation:<br> -) Copy the .pdb file related to the protein-ligand complex whose affinity is to be predicted in the subfolder “<em>ComplexToBePredictedFolder/</em>”.<br> -) Make sure the following files are in the same folder where the python script is:<br> <em>MaxAssMatrix.npy<br> mediaSynt.npy<br> devStndSynt.npy<br> my_model</em><br> -) Run the code using Python 3 with a command like "<em>python3.x UnicamSF13prot.py</em>".<br> -) Enter the name of the protein-ligand PDB file whose affinity is to be predicted (excluding the extension ".pdb").<br> -) Read the predicted affinity from screen.<br> </p> <p> </p>
Machine Learning Potentials for Metal-Organic Frameworks using an Incremental Learning Approach: Workflow and Data
<p>This repository contains input files, workflow scripts, and output datasets and interatomic potentials for a diverse set of metal-organic frameworks, as discussed in this <a href="https://chemrxiv.org/engage/chemrxiv/article-details/6363dbf718a8ccae675d2ac8">preprint paper</a>. In addition, we provide the scripts to compute the extended Hessian and subsequently the elastic constants using automatic differentiation.</p>
Curated Dataset of Association Constants Between a Cyclodextrin and a Guest for Machine Learning
<p>Determining the association constant between a cyclodextrin and a guest molecule is an important task for various applications in various industrial and academical fields. However, such a task is time consuming, tedious and requires samples of both molecules. A significant number of association constants and relevant data is available from the literature. The availability of data makes the use of machine learning techniques to predict association constants possible. However, such data is mainly available from tables in articles or appendices. It is necessary to make them available in a computer friendly format and to curate them. Furthermore, the raw data need to be enriched with physicochemical information about each molecule and when such information does not allow to discriminate molecules, some additional data is needed. We present a dataset built from data gathered from the literature.</p>
Dataset for machine learning guided prediction of the yield strength and hardness of multi-principal element alloys
<p>These data sets were used to develop machine-learning models to predict yield strength and hardness of multi-principal element alloys. We mainly collected the alloys and their mechanical properties from different published works. A list of references is provided at the end of each data set.</p>
Health, socioeconomic and genetic predictors of COVID-19 vaccination uptake: a nationwide machine-learning study
<p>Reduced participation in COVID-19 vaccination programs is a key societal concern. Understanding factors associated with vaccination uptake can help in planning effective immunization programs. We considered 2,890 health, socioeconomic, familial, and demographic factors measured on the entire Finnish population aged 30 to 80 (N=3,192,505) and genome-wide information for a subset of 273,765 individuals. Risk factors were further classified into 12 thematic categories and a machine learning model was trained for each category. The main outcome was uptaking the first COVID-19 vaccination dose by 31.10.2021, which has occurred for 90.3% of the individuals.</p> <p>The strongest predictor category was labor income in 2019 (AUC evaluated in a separate test set = 0.710, 95% CI: 0.708-0.712), while drug purchase history, including 376 drug classes, achieved a similar prediction performance (AUC = 0.706, 95% CI: 0.704-0.708). Higher relative risks of being unvaccinated were observed for some mental health diagnoses (e.g. dissocial personality disorder, OR=1.26, 95% CI : 1.24-1.27) and when considering vaccination status of first-degree relatives (OR=1.31, 95% CI:1.31-1.32 for unvaccinated mothers)</p> <p>We derived a prediction model for vaccination uptake by combining all the predictors and achieved good discrimination (AUC = 0.801, 95% CI: 0.799-0.803). The 1% of individuals with the highest risk of not vaccinating according to the model predictions had an average observed vaccination rate of only 18.8%.</p> <p>We identified 8 genetic loci associated with vaccination uptake and derived a polygenic score, which was a weak predictor of vaccination status in an independent subset (AUC=0.612, 95% CI: 0.601-0.623). Genetic effects were replicated in an additional 145,615 individuals from Estonia (genetic correlation=0.80, 95% CI: 0.66-0.95) and, similarly to data from Finland, correlated with mental health and propensity to participate in scientific studies. Individuals at higher genetic risk for severe COVID-19 were less likely to get vaccinated (OR=1.03, 95% CI: 1.02-1.05).</p> <p>Our results, while highlighting the importance of harmonized nationwide information, not limited to health, suggest that individuals at higher risk of suffering the worst consequences of COVID-19 are also those less likely to uptake COVID-19 vaccination. The results can support evidence-informed actions for COVID-19 and other areas of national immunization programs.</p>
Data for paper 'Machine Learning Force Fields for Molecular Liquids: Ethylene Carbonate / Ethyl Methyl Carbonate Binary Solvent'
<p>This data is supplied in conjunction with the paper:</p> <p>Magdău, I. B., Arismendi-Arrieta, D. J., Smith, H. E., Grey, C. P., Hermansson, K., and Csányi, G. NPJ Computational Materials, accepted. (2023). "Machine Learning Force Field for Molecular Liquids: Ethylene Carbonate / Ethyl Methyl Carbonate Binary Solvent."</p> <p>The archive contains the final EC:EMC training data and test sets (Volume Scans, Intra/Inter splits), final GAP potential and the MD trajectories described in the paper.</p> <p>The data is accompanied by a Jupyter Notebook: HowTo.ipynb (also compiled as *.pdf and *.html) which explains in detail the structure of the data and how to interact with it. The Notebook also demonstrates how to create volume scans, intra/inter splits and analyze configurations and MD trajectories.</p>
Open data for "Predicting dominant terrestrial biomes at a Global Scale: Assessments of machine learning algorithms, climate variables indexing, and extreme climate"
<p>______________________________________________________<br> This page contains public-domain data required to reconstruct simulation results in the manuscript "Predicting dominant terrestrial biomes at a Global Scale: Assessments of machine learning algorithms, climate variables indexing, and extreme climate," submitted by the following author.</p> <p>Author: Hisashi SATO (JAMSTEC) <br> email : hsatoscb_(at)_gmail.com</p> <p>______________________________________________________<br> 1. Folder "Code"<br> Detailed descriptions are available on the code. </p> <p>1-1. MachineLearningComparison.R<br> Machine learning programs using random forest (RF), naive Bayes classifier (NV), and support vector machine (SVM) algorithms.</p> <p>1-2. Analyse_MapSimilarity.R<br> Calculate coincidences of simulated potential natural vegetation (PNV) maps simulated by different models.</p> <p>1-3. Visualize_VCE.R<br> Generating VCE (Visualize Climate Image) for training CNN models.</p> <p>1-4. Visualize_Maps.R<br> Visualizing global PNV maps.</p> <p>1-5. Visualize_ClimateHistgrams.R<br> Visualizing histograms of climate datasets.</p> <p>______________________________________________________<br> 2. Folder "Input"</p> <p>2-1. Unified_BIOCLIM_WorldClim.csv<br> Input data for the current climate.<br> This file contains the following variables.<br> lon Longitude at the center of the grid<br> lat Latitude at the center of the grid<br> bio1~19 Average climate indices from BIOCLIM (AveI)<br> CDD~WSDI Extreme climate indices (CEI)<br> c1~c16 Fraction of PNV from MODIS data<br> tavg01~tavg12 Monthly mean air temperature from January to December (Ave)<br> prec01~prec12 Monthly precipitation from January to December (Ave)</p> <p>2-2. Unified_BIOCLIM_WorldClimFutureRCP85.csv<br> Input data for future climate (@RCP8.5)<br> Including variables are the same as Unified_BIOCLIM_WorldClim.csv</p> <p>2-3. BIOCLIM_RefNo.csv<br> This CSV file contains the following information for each grid.<br> lat: Latitude at the center of the grid<br> lon: Longitude at the center of the grid<br> latNo: Latitude number corresponding to the image file name<br> lonNo: Longitude number corresponding to the image file name<br> lineNo: No use. Don't mind.<br> vegNo: Most dominant PNV based on the Unified_BIOCLIM_WorldClim.csv</p> <p>______________________________________________________<br> 3. Folder "Output"</p> <p>3-1. PNV_sim<br> 3-2. PNV_sim_RCP85.csv<br> Current and future PNV maps from various models. These files are the main output files from the code MachineLearningComparison.R. For PNV maps from CNN models (m4p1~6) were supplemented. Detailed methods to build CNN models, please refer to the following manuscript.<br> Sato, H. & T. Ise (2022). "Predicting global terrestrial biomes with the LeNet convolutional neural network." Geoscientific Model Development 15(7): 3121-3132.</p> <p>Labels indicate combinations of machine-learning-algorithm and dataset for training the model. For example, In case of "m1p1", that column shows the simulation result of models trained with randomForest (RF) algorithm and Ave dataset.<br> m1: randomForest (RF)<br> m2: Support vector machine (SVM)<br> m3: Naive Bayes (NB)<br> m4: Convolutional Neural Network (CNN), which is NOT analysed in this code<br> p1: Ave<br> p2: Ave + CEI<br> p3: Ave + CEIpart<br> p4: AveI <br> p5: AveI + CEI<br> p6: AveI + CEIpart</p>
Numerical analysis and machine learning techniques on the behavior of FRP confined circular reinforced concrete columns
<p>This study presents a comprehensive nonlinear finite element study on the behavior of circular fibre reinforced polymer (FRP) confined reinforced and plain concrete columns under concentric loads. For this investigation, 65 test models with a combination of spiral hoop reinforced concrete, concrete with longitudinal and circular hoop reinforcements, and FRP confined plain concrete were designed . Four different machine learning (ML) techniques were developed to predict the ultimate axial load and strain at the tensile rupture of FRP. The accuracy of the proposed finite element model (FEM) was verified by comparing it with the existing experimental test results. The impact of unconfined concrete strength, hoop reinforcement ratio, thickness of FRP, and spiral hoop spacing on the confinement effectiveness, load-carrying capacity, and ductility behavior of circular FRP confined concrete columns were demonstrated. The parametric analysis found that axial load capacity of FRP-confined concrete columns increased when unconfined concrete strength increased, while low-strength confined concrete achieved a larger strength improvement ratio than high-grade concrete. The investigation also revealed that the thickness of the confining FRP has a significant impact on the confinement effectiveness of hoop reinforcement. The correlation between the FEM and experimental tests yielded 99.60% R<sup>2</sup> for ultimate axial load and 93.40% R<sup>2</sup> for ultimate strain. Extra tree regressor (ETR) and gradient boosting of ML yielded accurate predictions of ultimate axial load and strain at the tensile rupture of FRP compared to other approaches, but ETR has the best comprehensive prediction performance using the comprehensive ranking system. Overall, ETR can be applied in the ultimate axial load and strain prediction of circular FRP confined reinforced and plain concrete columns under concentric loads, conserving resources, time, and cost through laboratory testing.</p> <p> </p> <p> </p> <p> </p> <p> </p>
Dataset for "Polarimetric retrieval of raindrop size distribution: double-moment normalization approach and machine learning techniques"
<p>This is the dataset for the publication [1]. This dataset contains:<br>1) Training dataset (train_data.pkl)<br>2) Independent validation dataset (test_data.pkl)</p> <p>This dataset was also used for the publication [2]. The training and validation datasets were merged in [2].</p> <p> </p> <p>References</p> <p>[1] Shin, K., Kim, K., Song, J. J., & Lee, G. (2024). Polarimetric Retrieval of Raindrop Size Distribution: Double‐Moment Normalization Approach and Machine Learning Techniques. <em>Geophysical Research Letters</em>, <em>51</em>(1), e2023GL106057. <a href="https://doi.org/10.1029/2023GL106057">https://doi.org/10.1029/2023GL106057</a></p> <p>[2] Tapiador, F. J., Shin, K., Leganés, L. J., Lim, K.-S., Juárez, G., Bang, W., et al. (2025). A Physically Consistent Particle Size Distribution Modelling of the Microphysics of Precipitation for Weather and Climate Models. <em>Submitted</em>.</p>
Bugs in machine learning-based systems: a faultload benchmark
<p>Replication package of paper "<strong>Bugs in Machine Learning-based Systems: A Faultload Benchmark</strong>".<br> This repository is faultload benchmark of ML-based systems namely <strong>defect4ML</strong>.</p>
Dataset A of article "A machine learning approach for classifying Mediterranean obsidian of archaeometric interest by SEM investigation"
<p>Dataset A used for the manuscript titled "A machine learning approach for classifying Mediterranean obsidian of archaeometric interest by SEM investigation".</p><p>This dataset is a txt file, tab separated, encoded in UTF-8, with newline line terminator.</p><p>A compressed zip file is also provided, which, expanded, will create a directory containing 890 files. The files come in couples with the same name and two different extensions: spc (with SEM-EDS spectroscopy data) and tif (with SEM-BSE images). The SEM-EDS data refers to the same sample area shown in the SEM-BSE image going under the same name.</p>
Machine Learning in Guiding rTMS Treatment for GWI-Related Headaches and Body Pain
ClinicalTrials.gov study NCT07325513. IPD Sharing: NO. Countries: 0. Publications: 0.
Evaluation of Pulmonary Complications in Liver Transplantation Patients Based on Machine Learning
ClinicalTrials.gov study NCT06534840. IPD Sharing: UNDECIDED. Countries: 1. Publications: 0.
Predicting HIF-2α Levels in Clear Cell Kidney Cancer Using Machine Learning
ClinicalTrials.gov study NCT07332923. IPD Sharing: UNDECIDED. Countries: 1. Publications: 0.
Construction of Perioperative Medical Data Platform and Its Typical Practice to Predict Postoperative Acute Moderate to Severe Pain With Machine Learning Models
ClinicalTrials.gov study NCT05569460. IPD Sharing: NO. Countries: 1. Publications: 0.
Use of Machine-learning Algorithms, Biomarkers and Measures of Quality of Life to Personalize Medical Management of Liver and Heart Transplant Recipients
ClinicalTrials.gov study NCT06774768. IPD Sharing: Not stated. Countries: 1. Publications: 0.
Machine Learning for Estimating Cardiorespiratory Fitness in Patients With Obesity
ClinicalTrials.gov study NCT07011108. IPD Sharing: NO. Countries: 0. Publications: 0.
Pre-operative Characteristics for Prediction of Supraglottic Airway Failure Using Machine Learning (ERICA)
ClinicalTrials.gov study NCT06617403. IPD Sharing: Not stated. Countries: 1. Publications: 0.
A Machine Learning Predictive Model for Sepsis
ClinicalTrials.gov study NCT04771429. IPD Sharing: UNDECIDED. Countries: 1. Publications: 0.
Prospective Validation of the Model Predicting Postoperative Delirium Occurrence With Machine Learning-based Analysis of Intraoperative Biological Signals During Anesthesia in Cardiac Surgery
ClinicalTrials.gov study NCT05320965. IPD Sharing: UNDECIDED. Countries: 1. Publications: 0.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.