Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,943
datasets available to search
ShareScore release 0.9.0
Dataset results
1,943 results for “machine learning”
Identification of Kidney Cell Types in scRNA-seq and snRNA-seq Data Using Machine Learning Algorithms
<p>Metadata files for our final analysis as part of the "<span><span>Identification of Kidney Cell Types in scRNA-seq and snRNA-seq Data Using Machine Learning Algorithms" study.</span></span></p>
Data for "Online Learning of Entrainment Closures in a Hybrid Machine Learning Parameterization"
<p>Calibration data and visualization tools for "Online Learning of Entrainment Closures in a Hybrid Machine Learning Parameterization".</p>
Descriptor and Graph-based Molecular Representations in Prediction of Copolymer Properties using Machine Learning
<p>This dataset accompanies a study that investigates the use of machine learning (ML) approaches for predicting seven different physical properties of 140 binary copolymers. Two computational methods were employed: a random forest (RF) model based on molecular descriptors and a Graph Neural Network (GNN) using 2D polymer graphs. These methods were applied in both single- and multi-task settings to explore the strengths of each approach in capturing various polymer properties.</p> <p>The dataset includes two files:</p> <ol> <li> <p><strong>Dataset.xlsx</strong>: Contains the following information for each of the 140 copolymers:</p> <ul> <li>Polymer names</li> <li>SMILES notation of the monomers</li> <li>Fraction of monomers in each copolymer</li> <li>Simulated values (calculated using molecular dynamics simulation) and experimental values of various physical properties, including density, specific heat capacity at constant pressure ad volume ,radius of gyration, linear expansion coefficient, volume expansion coefficient, and bulk modulus .</li> </ul> </li> <li> <p><strong>Descriptors.xlsx</strong>: Provides the molecular descriptors calculated using PaDEL-Descriptor software, which were used as input for the RF model to predict polymer properties.</p> </li> </ol> <p>This work provides insight into the comparative strengths of descriptor- and graph-based representations in machine learning models for predicting material properties. It also highlights the importance of selecting appropriate molecular representations based on the nature of the properties being predicted.</p>
DFT datasets for training machine-learning potential to model Cl-doped lithium borosilicate glasses using DeePMD
<h2><strong>Li diffusion in oxygen-chlorine mixed anion borosilicate glasses using</strong></h2> <h2><strong>a machine-learning simulation</strong></h2> <h5>Shingo Urata, Noriyoshi Kayaba</h5> <ul> <li>DFT_Data_for_Cl-doped_LBSCl_glass.zip inlucudes atom configurations, energies, forces, box size, atom types, atom kinds, and virial in coord.raw, energy.raw, force.raw, type.raw, type_map.raw, and virial.raw, respectively. </li> <li>All DFT data were evaluated using PBE with a cutoff energy of 600 eV by VASP.</li> <li>The other detasets are available from https://doi.org/10.5281/zenodo.10577559</li> <li>LBSCl_DMD_model.pb is force field developed using DeePMD-kit.</li> <li>LBSCl_DMD_model_c.pb.zip is the compressed version of LBSCl_DMD_model.pb.</li> </ul>
Dataset for the paper 'Predicting mechanical properties of polycrystalline nanopillars by interpretable machine learning'
<div> <div>This dataset contains the data produced for the above paper. The dataset consists of:</div> <div> </div> <div>- input nanopillars before deformation (molecular_dynamics/nanopillars)</div> <div>- stress-strain curves acquired by deforming the nanopillars (molecular_dynamics/stress_strain_curves)</div> <div>- weights of the CNNs trained to predict mechanical properties of the nanopillars (machine_learning/train_CNN)</div> <div>- Grad-CAM fields of the predictions (machine_learning/train_CNN)</div> <br> <div>Codes used for creating and analyzing the dataset are available at https://github.com/tekoivisto/nanopillar-ML</div> </div>
Supplemental materials for "Boosting Barlow Twins reduced order modeling for machine learning-based surrogate models in multiphase flow problems"
<p>Supplemental materials for "Boosting Barlow Twins reduced order modeling for machine learning-based surrogate models in multiphase flow problems" in Water Resources Research. Detailed information is available in readme.md.</p>
Linked Papers With Code: The Latest in Machine Learning as an RDF Knowledge Graph
<p><strong>Linked Papers With Code (LPWC)</strong> is an <strong>RDF knowledge graph </strong>that comprehensively models the research field of <strong>machine learning</strong>. It contains information about almost 400,000 machine learning <strong>publications</strong>, including the <strong>tasks</strong> addressed, the <strong>datasets</strong> utilized, the <strong>methods</strong> implemented, and the <strong>evaluations</strong> conducted, along with their <strong>results</strong>. The data set is based on <strong>Papers With Code</strong> and licensed under the CC BY-SA 4.0 license. Furthermore, we provide <strong>knowledge graph embeddings</strong> for entities and relations represented in LPWC.</p><p>More information can be found at <a href="https://linkedpaperswithcode.com/"><strong>https://linkedpaperswithcode.com/</strong></a> and in the ISWC'23 publication <a href="https://linkedpaperswithcode.com/"><strong>"</strong></a><a href="https://aifb.kit.edu/web/Inproceedings3993"><strong>Linked Papers With Code: The Latest in Machine Learning as an RDF Knowledge Graph".</strong></a></p>
Predicting Equatorial Spread F at JICAMARCA Sector via Supervised Machine Learning
<p>Dataset used for ESF prediction model</p> <p> </p>
Incorporating Hourly Convective Cloud data into Tropical Cyclone Rapid Intensification Forecasting with Machine Learning
<p>These models were developed to predict both the probability of RI and the binary RI classification for tropical cyclones. They are a weighted average of probabilities derived from logistic regression, random forest, decision tree, and extremely randomized tree algorithms within the standard Python scikit-learn package. The uploaded files include the hyperparameters and weights for each individual machine learning model.</p>
Supplementary initial and final MD configurations for the manuscript: General-purpose machine-learned potential for 16 elemental metals and their alloys
<p>Extended XYZ files for the initial and final configurations of the molecular dynamics simulations from the article 'General-purpose machine-learned potential for 16 elemental metals and their alloys' (https://arxiv.org/abs/2311.04732).</p> <p><span>The Supplementary Data is contained in the folder named:<br>1) Polycrystalline-MoTaVW: <span> </span>The plasticity MD simulations in multi-principal element alloys.</span></p> <p><span>2) MoTaVW-radiation: The primary radiation damage MD simulations in multi-principal element alloys.</span></p> <p><span>3) Goldene: Comparisons between UNEP-v1 and EAM models in MD simulations.</span></p> <p><span>4) NiAlMo: Comparisons between UNEP-v1 and EAM models in MCMD simulations.<br>5) AlCrCuNiV: Comparisons between UNEP-v1 and EAM models in MCMD simulations.</span></p>
Source Data for the manuscript: General-purpose machine-learned potential for 16 elemental metals and their alloys
<p>***Source Data***<br>This folder contains multiple .txt files that provide the source data for the figures and tables presented in the paper: "General-purpose machine-learned potential for 16 elemental metals and their alloys."</p> <p>The source data are organized in the following folders and files:</p> <p>1) Fig2<br>2) Fig3<br>3) Fig4<br>4) Fig5<br>5) Fig6<br>6) FigS1<br>7) FigS2<br>8) FigS3<br>9) FigS4-6-pure<br>10) FigS7-9-binary<br>11) FigS10-12-ternany<br>12) FigS13-15-quaternary<br>13) FigS16-17-quinary<br>14) FigS18-20<br>15) FigS21<br>16) FigS22<br>17) FigS23<br>18) FigS26<br>19) Table1-Element-atoms-GPU-Speed.txt<br>20) Table-S1-DFT-EAM-UNEP-Elastic.txt<br>21) Table-S2-DFT-EAM-UNEP-Mono-vacanc.txt<br>22) Table-s3-Surface-100-110-111-DFT-EAM-UNEP.txt<br>23) Table-S4-Melting-EAM-UNEP-Exp.txt</p>
Supporting data and code for the published paper: Machine Learning Nonadiabatic Dynamics: Eliminating Phase Freedom of Nonadiabatic Couplings with the State-Interaction State-Averaged Spin-Restricted Ensemble-Referenced Kohn–Sham Approach
Open the record for dataset details and reuse information.
A COMPREHENSIVE STUDY OF MACHINE LEARNING APPROACHES FOR CUSTOMER SENTIMENT ANALYSIS IN BANKING SECTOR
<p>This study explores the application of sentiment analysis in the banking sector, focusing on customer feedback to enhance service quality and customer experiences. We collected a comprehensive dataset of approximately 100,000 entries from diverse sources, including customer satisfaction surveys, social media platforms, and direct feedback. A robust preprocessing pipeline was employed to address challenges associated with unstructured data, informal language, and mixed sentiments. We evaluated several machine learning and natural language processing models, including Logistic Regression, Naive Bayes, Support Vector Machine (SVM), Random Forest, Long Short-Term Memory (LSTM), and BERT (Bidirectional Encoder Representations from Transformers), using metrics such as accuracy, precision, recall, F1 score, AUC-ROC, and training time. The results revealed that advanced models, particularly BERT, achieved superior performance with an accuracy of 88% and an F1 score of 0.86, demonstrating an exceptional ability to capture nuanced sentiments. This study underscores the importance of employing sophisticated sentiment analysis techniques in banking to derive actionable insights from customer feedback. The findings suggest that leveraging advanced models can significantly improve service quality and customer satisfaction, while also presenting avenues for future research into real-time sentiment analysis and its integration with customer relationship management systems.</p>
ADVANCEMENTS IN AIRLINE SECURITY: EVALUATING MACHINE LEARNING MODELS FOR THREAT DETECTION
<p>This study assessed the performance of four machine learning algorithms—Decision Tree (DT), Random Forest (RF), Support Vector Machine (SVM), and Neural Network (NN)—for predicting airline security threats using a dataset of 100,000 entries with 30 features. The models were evaluated based on accuracy, precision, recall, F1-Score, and AUC-ROC. The Neural Network achieved the highest performance, with an accuracy of 88%, precision of 86%, recall of 85%, F1-Score of 85.5%, and AUC-ROC of 0.90, demonstrating superior capability in capturing complex, non-linear patterns. The Random Forest model followed, with an accuracy of 85%, precision of 83%, recall of 82%, F1-Score of 82.5%, and AUC-ROC of 0.87, offering a robust and generalizable solution. The SVM model attained an accuracy of 81%, precision of 80%, recall of 78%, F1-Score of 79%, and AUC-ROC of 0.84, showing effective binary classification but with higher computational costs. The Decision Tree model, while interpretable, had the lowest performance metrics: accuracy of 78%, precision of 76%, recall of 72%, F1-Score of 74%, and AUC-ROC of 0.79. The results indicate that Neural Networks and Random Forests are the most effective models for airline security threat detection, with Neural Networks providing the highest overall accuracy and AUC-ROC.</p>
Reproduction Package: Using machine learning techniques to mitigate confidentiality violations
<p>Reproduction package for the masters thesis "Using machine learning techniques to mitigate confidentiality violations"</p>
Power grid attack detection and state estimation with machine learning
<p>Detecting attacks and estimating states of power grids from partial observations with machine learning. A manuscript submitted to PRX Energy.</p>
Data related to the publication "structure and transport properties of LiTFSI-based deep eutectic electrolytes from machine-learned interatomic potential simulations"
<p>Reference training and test datasets, trained ML potential models, and input scripts for the training (Allegro) and MD simulations (LAMMPS).</p>
DYNAMIC PRICING IN FINANCIAL TECHNOLOGY: EVALUATING MACHINE LEARNING SOLUTIONS FOR MARKET ADAPTABILITY
<p>The rapid advancement of technology has transformed the financial services sector, leading to the rise of fintech companies that leverage cutting-edge tools such as artificial intelligence (AI) and machine learning (ML) to offer innovative solutions. One area where fintech is particularly impactful is dynamic pricing, which involves adjusting prices in real-time based on market conditions, user behavior, and external factors. The ability to optimize pricing in response to fluctuating conditions is critical for maximizing profitability, improving customer satisfaction, and maintaining competitiveness. In this context, machine learning algorithms provide a powerful framework for making data-driven pricing decisions by learning from historical data and predicting future trends.</p>
[Supplementary Data] PowerModel-AI: A First On-the-fly Machine-Learning Predictor for AC Power Flow Solutions.
<h1><strong>Abstract</strong></h1> <p>The real-time creation of machine-learning models via active or on-the-fly learning has attracted considerable interest across various scientific and engineering disciplines. These algorithms enable machines to autonomously build models while remaining operational. Through a series of query strategies, the machine can evaluate whether newly encountered data fall outside the scope of the existing training set. In this study, we introduce <em>PowerModel-AI</em>, an end-to-end machine learning software designed to accurately predict AC power flow solutions. We present detailed justifications for our model design choices and demonstrate that selecting the right input features effectively captures the load flow decoupling inherent in power flow equations. Our approach incorporates on-the-fly learning, where power flow calculations are initiated only when the machine detects a need to improve the dataset in regions where the model's performance is sub-optimal, based on specific criteria. Otherwise, the existing model is used for power flow predictions. This study includes analyses of five Texas A&M synthetic power grid cases, encompassing the 14-, 30-, 37-, 200-, and 500-bus systems. The training and test datasets were generated using <em>PowerModel.jl</em>, an open-source power flow solver/optimizer developed at Los Alamos National Laboratory, NM, USA.</p> <div> <h1><strong>Overview</strong></h1> <p>This dataset, provided as supplementary material for the above-referenced study, includes a comprehensive collection of images (plots) from the study’s analyses, along with Jupyter notebooks containing Python scripts used for the training, validation, and testing phases of PowerModel-AI. Additionally, it includes all training and external test data used in this work, generated via LANL-based open-source power flow solver, PowerModels.jl.</p> <p>The primary objective of this dataset is to ensure full reproducibility of the study’s analyses and facilitate critical examination by the scientific community, thereby maximizing the overall impact of the work.</p> <h1>Directory Structure</h1> <p>The hierarchy of folders and file organization of the dataset is illustrated in the chart below. The directory contains a README.md file which contains the information provided here and a requirements.txt that contains all libraries necessary to run the python scripts or Jupyter notebooks in this directory. A brief description of the folders and what they contain are provided below: </p> </div> <div> <strong>.</strong></div> <div>├── <strong>Models/</strong></div> <div>│ ├── _PM_AI_Models/</div> <div>│ │ ├── Type1/</div> <div>│ │ ├── Type2/</div> <div>│ │ └── Type3/</div> <div>│ │</div> <div>│ ├── _PM_AI_Module/</div> <div>│ │ ├── __init__.py</div> <div>│ │ └── PM_Methods.py</div> <div>│ │</div> <div>│ ├── _PM_JL_Data/</div> <div>│ ├── C1_Model/</div> <div>│ ├── C2_Model/</div> <div>│ ├── C3_Model/</div> <div>│ ├── M1_Model/</div> <div>│ ├── M2_Model/ </div> <div>│ └── M3_Model/</div> <div>│</div> <div>├── <strong>NodeSensitivityAnalysis/</strong></div> <div>│ ├── Sensitivity_Plots/</div> <div>│ ├── Sensitivity_PM_JL_Data/</div> <div>│ ├── get_sensitivity_PMJL_data.py</div> <div>│ └── NodeSensitivityAnalysis.ipynb</div> <div>│</div> <div>├── <strong>PlotsForOnTheFlyAnalysis/</strong></div> <div>│ ├── BaseModel_A/</div> <div>│ ├── BaseModel_B/ </div> <div>│ └── BaseModel_C/</div> <div>│ </div> <div>├── <strong>README.md</strong></div> <div>└── <strong>requirements.txt</strong></div> <p> </p> <h2>Files Description </h2> <h3><strong>1. </strong><strong>Models/</strong></h3> <p><strong><em>_PM_AI_Models/</em></strong> contains the ML models discussed in the manuscript. Type1, Type2, and Type3 refers to the C and M models with numbers "1", "2," and "3". Each Type folder contains individual subfolders for the synthetic grids discussed.</p> <p><strong> <em>_PM_AI_Module/</em></strong> contains a python script that has all the functions used in model training and analysis. It is imported in the Jupyter notebooks in the C and M subfolders in this directory.</p> <p><strong><em>_PM_JL_Data/</em> </strong>contains the following subfolders:</p> <p>a) <strong> </strong><em>_</em><strong><em><strong>G</strong>eneratePowerModelData/</em></strong> has a python script (<u>get_PowerModelJLData.py)</u> that is used to parse <u>PowerModels.jl</u> to compute AC power flow solutions for different power demand configurations. It also contains a subfolder, <strong><em>BusData_MATLAB/</em></strong>, that has all the synthetic grids used in the study and in MATLAB format.</p> <p>b) It also contains other subfolders (not shown in the chart above) that contains AC power flow solutions generated using LANL’s PowerModels.jl, for each synthetic grids and other grid-related data.</p> <p>The folders starting with C and M are the control and candidate models (more details in manuscript). Each folder contains Jupyter notebooks (for each power grid) that has python algorithms used in model training and testing, as well as functions to analyze and plot the prediction performance of the models. It accesses (if already trained) or stores (if newly trained) the models in the <strong><em>_PM_AI_Models/</em></strong> directory. Additionally, they contain 2 subfolders (not shown in the chart above): <strong>AbsoluteErrorPlots/</strong> and <strong>PredictionPlots/</strong> were the results (plots) from the analyses are stored. Some of these results are shown in the manuscript (<em>Figures 4</em>,<em> 5</em>, <em>6</em>, and <em>7</em>).</p> <h3><strong>2. </strong><strong>NodeSensitivityAnalysis/</strong></h3> <p>This folder contains the tools used for the node dependency analysis in Section 3.1.1 of the manuscript.</p> <p>It contains a python script called get_sensitivity_PMJL_data.py, which has similar operation like the <u>get_PowerModelJLData.py</u> script but only compute changes for one bus at a time (see details in manuscript). There is also a Jupyter notebook called <u>NodeSensitivityAnalysis.ipynb </u>that analyzes the bus node dependencies and produces the plots that are shown in Figure 3 in the manuscript. It contains 2 additional sub-folders:</p> <p>a) <strong> <em>Sensitivity_PM_JL_Data/</em></strong> where PowerModels.jl generated AC power flow solution data are stored, and</p> <p>b) <em><strong>Sensitivity_Plots/</strong> </em>where the results from <u>NodeSensitivityAnalysis.ipynb</u> are stored.</p> <h3><strong>3. </strong><strong>PlotsForOnTheFlyAnalysis/</strong></h3> <p>This directory contains only the results for the discussion in Section 3.2 in the manuscript, which is the on-the-fly implementation of PowerModel-AI. The on-the-fly algorithm will be provided and distributed separately in the PowerModel-AI package, which will be publicly available through LANL’s <a href="https://github.com/lanl-ansi" target="_blank" rel="noopener">The Advanced Network Science Initiative</a> (Github). It contains 3 subfolders with similar names but ends with "A", "B," and "C" which correspond to the designations shown and discussed in <em>Figure 8</em> of the manuscript.</p> <h1>Summary</h1> <h3><strong> </strong><strong>Models/</strong></h3> <p><strong>1. _PM_AI_Models/: </strong></p> <p>Contains the machine learning models discussed in the manuscript. The Type1, Type2, and Type3 folders refer to C and M models labeled "1", "2," and "3". Each type folder includes subfolders for the corresponding synthetic grids analyzed.</p> <div><strong>2. _PM_AI_Module/: </strong> </div> <div>Contains Python scripts for model training and analysis. PM_Methods.py is the script that defines functions for model training and analysis, which are imported in the Jupyter notebooks in the C and M subfolders. </div> <div> </div> <div><strong>3. _PM_JL_Data/: </strong></div> <div>-_GeneratePowerModelData/: Contains a Python script (get_PowerModelJLData.py) used to parse PowerModels.jl and compute AC power flow solutions for various power demand configurations. This folder also includes BusData_MATLAB/, which holds the synthetic grid data in MATLAB format. </div> <div>- Other Subfolders: Contain PowerModels.jl AC power flow solutions for each synthetic grid, as well as related data.</div> <div> </div> <div><strong>4. C1_Model/, C2_Model/, C3_Model/, M1_Model/, M2_Model/, M3_Model/: </strong></div> <div>These folders contain the control and candidate models (refer to manuscript details). Each folder includes pre-run Jupyter notebooks for model training and analysis, as well as two subfolders: </div> <div> - <em>AbsoluteErrorPlots/</em>: Contains saved analysis results for each grid.</div> <div> <strong> </strong>- <em>PredictionPlots/</em>: Stores model prediction results. </div> <div>Some results are shown in <em>Figures 4, 5, 6,</em> and <em>7</em> of the manuscript and can be reproduced using the notebooks.</div> <h3>NodeSensitivityAnalysis/</h3> <div><strong>1. Node Dependency Analysis Tools: </strong> </div> <div> This folder contains scripts and data used for the node dependency analysis in Section 3.1.1 of the manuscript.</div> <div> </div> <div><strong>2. Scripts: </strong></div> <div> - get_sensitivity_PMJL_data.py: Computes power flow changes for individual buses (refer to the manuscript for details). </div> <div> - NodeSensitivityAnalysis.ipynb: Analyzes dependencies and generates plots for Figure 3 in the manuscript. </div> <div> </div> <div><strong>3. Subfolders:</strong> </div> <div> - Sensitivity_PM_JL_Data/: Stores PowerModels.jl data for sensitivity analysis. </div> <div> - Sensitivity_Plots/: Contains the generated results from the Jupyter notebook.</div> <h3>PlotsForOnTheFlyAnalysis/</h3> <div><strong>On-the-Fly Learning Results: </strong>Contains the results for the on-the-fly learning analysis discussed in Section 3.2 of the manuscript. The subfolders BaseModel_A/, BaseModel_B/, and BaseModel_C/ correspond to the designations in <em>Figure 8</em> of the manuscript. These subfolders contain plots generated during analysis.</div> <h3>Additional Files</h3> <div>- README.md: This file.</div> <div>- requirements.txt: Contains a list of necessary Python libraries required to run the scripts and Jupyter notebooks.</div> <h3>Notes:</h3> <div>- The PowerModel-AI code is publicly available on GitHub.</div> <div>- The supplementary materials include additional Jupyter notebooks and results for each power grid analysis. </div> <div>- The generated results and predictions in this repository are consistent with those discussed in the manuscript.</div>
Using Machine Learning With Supplementary NC Code To Predict Machining Energy - Excel Documents
<p>The Excel Files Housed within this DOI represent the raw data collected during machining each of the test parts, and the excel documents made which prevent model summaries for each model created., during the execution of the, "Using Machine Learning With Supplementary NC Code to Predict Machining Energy. These files were created by Samuel D. Stencel, a Graduate Research Assistant and Purdue University.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.