Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,943
datasets available to search
ShareScore release 0.9.0
Dataset results
1,943 results for “machine learning”
Supplementary Material for the paper "Developers' Perception Matters: Machine Learning to Detect Developer-sensitive Smells"
<p>Supplementary material for the paper "Daniel Oliveira; Wesley K. G. Assunção; Alessandro Garcia; Baldoino Fonseca; and Márcio Ribeiro. Developers' Perception Matters: Machine Learning to Detect Developer-sensitive Smells. In: Empirical Software Engineering. 2022. Springer."</p>
Datasets from "Electrostatic Embedding of Machine Learning Potentials"
<p>Data required to reproduce results in "Electrostatic Embedding of Machine Learning Potentials" <a href="https://doi.org/10.26434/chemrxiv-2022-rknwt">article</a>. See <a href="https://github.com/emedio/embedding">https://github.com/emedio/embedding</a> for details.</p> <ul> <li>QM7_B3LYP_cc-pVTZ.tgz - outputs of single point B3LYP/cc-pVTZ calculations of structures in <a href="http://quantum-machine.org/data/qm7.mat">QM7 dataset</a> with ORCA 5. Include molecular dipolar polarizabilities.</li> <li>QM7_B3LYP_cc-pVTZ_horton.tgz - MBIS partitioning of the B3LYP/cc-pVTZ densities with <a href="https://github.com/theochem/horton">Horton 2.1.0</a>.</li> <li>mpro_xyz.tgz - coordinates of the ligand and surrounding point charges from 100 snapshots of SARS-CoV-2 Mpro complex with PF-00835231.</li> <li>mpro_*.tgz - DFT and semiempirical single point calculations with ORCA 5 for the coordinates from mpro_xyz.tgz, <em>in vacuo </em>and in presence of point charges.</li> <li>mlmm.mat - learned parameters and SOAP feature vectors of reference atomic environments</li> </ul>
Ship tracks detected using machine learning algorithm
<p>The filtered, vector ship tracks detected using the linked machine learning algorithm and derived from the linked segmentation masks. Each dataset contains the date and other related data for each shiptrack polygon. The `_geo` dataset contains the polygons on a lat/lon coordinate system while the other provides the polygons on the MODIS swath (pixel) indices.</p>
CASM: A long-term Consistent Artificial-intelligence based Soil Moisture dataset based on machine learning and remote sensing
<p>Paper to cite: Skulovich, O., Gentine, P. A Long-term Consistent Artificial Intelligence and Remote Sensing-based Soil Moisture Dataset. <em>Sci Data</em> 10, 154 (2023). https://doi.org/10.1038/s41597-023-02053-x</p> <p> </p> <p>The Consistent Artificial Intelligence (AI)-based Soil Moisture (CASM) dataset is a global, consistent, and long-term, remote sensing soil moisture (SM) dataset created using machine learning. It is based on the NASA Soil Moisture Active Passive (SMAP) satellite mission SM data as a target and is aimed at extrapolating SMAP-like quality SM data back in time with previous satellite microwave platforms. Machine learning approach, such as neural network (NN) has the advantage of being both nonlinear, and state-dependent, and naturally imposing a global distribution matching between the source and the target data. Utilizing this, the new CASM dataset was created using high-quality SMAP SM as a target and Soil Moisture and Ocean Salinity (SMOS) or Advanced Microwave Scanning Radiometer - Earth Observing System (AMSR-E/2) brightness temperature as a source, which allowed extrapolating SM data 13 years back from before SMAP mission launch. CASM represents SM in the top soil layer, defined on a global 25 km EASE-2 grid and covers 2002-2020 with a 3-day temporal resolution. The resulting dataset exhibits excellent spatial and temporal homogeneity, without compromising the interannual variability, and is in excellent agreement with the SMAP data (with a mean correlation of 0.97 between the SMAP and CASM SM for the period when the two overlap). Moreover, the input and target datasets were divided into seasonal cycle and residuals, with the NN trained on the residuals. This approach ensures that the high performance does not mask a simple seasonal cycle matching but rather exemplifies the skill targeted at predicting extremes; with the NN achieving a correlation of 0.75 on the test data for the residuals. Comparison to 367 global in-situ SM monitoring sites shows a SMAP-like median correlation of 0.66 between station SM and CASM SM from the corresponding grid cell. Additionally, the SM product uncertainty was assessed, and both aleatoric and epistemic uncertainties were estimated and included in the dataset. Mean epistemic uncertainty, related to the NN model structure, ranges from 0.007 m<sup>3</sup>/m<sup>3</sup> to 0.014 m<sup>3</sup>/m<sup>3</sup> and on average is close to a desired SM product stability threshold of 0.01 m<sup>3</sup>/m<sup>3</sup> per year. Aleatoric uncertainty, defined as input noise propagated through the system, depends on the introduced level of noise. With 10% noise applied to the residuals, the resulting mean standard deviation of the model outputs rises from 0.005 to 0.007 m<sup>3</sup>/m<sup>3</sup>. </p>
On the Application of Machine Learning Models to Assess and Predict Software Reusability
<p>This is the dataset, results and notebook for the submission into Maltesque 2022 conference.</p>
Network Theme: The potential of machine learning and AI for blood based investigations - Professor Jeremy Frey (University of Southampton)
<p>This video is the sixth talk from our Future Blood Testing Network Plus Launch that took place on the 23/11/2021.</p> <p>Network Theme: The potential of machine learning and AI for blood based investigations - Professor Jeremy Frey (University of Southampton)</p> <p>Bio: Prof Jeremy Frey Professor of Physical Chemistry, Head of Computational Systems Chemistry, University of Southampton (UoS). He is PI of AI for Scientific Discovery Network+, and co_I on the Internet of Food Things Digital Economy Network+ and has had considerable involvement in the UK e-Science and Digital Economy programmes for many years (e.g., PI of the Digital Economy IT as a Utility Network+. He is a strong proponent of interdisciplinary research and the use of digital technology and ideas to enhance methods of scientific research & development. His own research involves activities across the physical land life sciences, from the application of novel mathematical analysis (e.g., Topological Data Analysis), laser spectroscopy and imagining techniques to chemical and biological problems, with the development of sensors and imagining systems such as the novel soft x-ray microscope. In parallel he works on the integration of these techniques with full provenance environment into laboratory systems using semantic web technologies.</p> <p>Further details on this event can be found at: https://futurebloodtesting.org/event/23-11-21-future-blood-testing-network-launch/</p> <p>This video is an output from the Future Blood Testing Network which is funded by EPSRC under Grant Number EP/W000652/1</p> <p>YouTube Link: https://youtu.be/eNORwfMy5cE</p>
The Application of Machine Learning for Classification on Blood Pressure Variability. A New Approach for an Old Idea - Professor Kelvin Tsoi (The Chinese University of Hong Kong, School of Public Health and Primary Care)
<p>This video is the eighth talk from our Future Blood Testing Network Plus Launch that took place on the 23/11/2021.</p> <p>The Application of Machine Learning for Classification on Blood Pressure Variability. A New Approach for an Old Idea - Professor Kelvin Tsoi (The Chinese University of Hong Kong, School of Public Health and Primary Care)</p> <p>Bio: Professor Kelvin Tsoi is an Epidemiologist specialized in Digital Health. His research interests focus on digital innovation in chronic disease management, including mobile and telecare application for hypertension management, technological implementation and social engagement for cognitive screening, artificial intelligent application on electronic health records. He also works as the traditional epidemiologist on evidence-based medicine and population cohort studies. He obtained his Bachler Degree from Department of Statistics and Doctor of Philosophy from School of Public Health in the Chinese University of Hong Kong. He further received post-doctoral training in the Division of Gastroenterology and Hepatology, Department of Medicine and Therapeutics. He was also appointed as a Director of CUHK JC Bowel Cancer Education Centre to promote colorectal cancer screening. In 2011, he worked as a research scientist in Hospital Authority. He led projects covering a wide range of service areas on chronic diseases, such as service demand projection for schizophrenia and dementia. The experience of database management enhanced his understanding of the HA database structures. In 2013, he was invited to join the interdisciplinary team for Big Data research and worked closely with a team of engineers and data scientists. Currently, Professor Tsoi is an Associate Professor in JC School of Public Health and Primary Care, SH big Data Decision Analytics Research Centre and JC Institute of Ageing.</p> <p>Further details on this event can be found at: https://futurebloodtesting.org/event/23-11-21-future-blood-testing-network-launch/</p> <p>This video is an output from the Future Blood Testing Network which is funded by EPSRC under Grant Number EP/W000652/1</p> <p>YouTube Link: https://youtu.be/liLVKA-JHiI</p>
Soil information on a regional scale: Two machine learning based approaches for predicting saturated hydraulic conductivity
<p><strong>Version 1.0 - This version is the final revised one.</strong></p> <p>This is the dataset accompanying the paper: Zeitfogel et al., Soil information on a regional scale: Two machine learning based approaches for predicting saturated hydraulic conductivity, published at Geoderma, 2023 (https://doi.org/10.1016/j.geoderma.2023.116418).</p> <p>Soil property and Ksat maps for Austria. The digital soil maps were generated based on a Machine Learning and PTF-based approach (indirect approach) and a pure Machine Learning based approach (direct approach). By downloading the datasets, you agree that we nor the provider of the used source datasets cannot be liable for the data provided.</p> <p>This study was funded by the Austrian Federal Ministry of Agriculture, Regions and Tourism (Project InfCapAT), the Austrian Academy of Science (Project RechAUT) and the Austrian Science Fund project P 31213.</p> <p> </p>
Machine Learning Enabled Multi-Radio Access Technology Selection in 5G Networks
<p>In this paper, we present a machine learning algorithm for effective RAT selection in 5G networks by considering the geo-location (latitude and longitude) of the user as well as the received signal strength intensity (RSSI) from the base station as basic parameters, real live data from a 5G network base-station were collated, divided into training and testing data-sets, the training data-sets (input) were used to train models of supervised machine learning classification algorithm: Decision Tree (DT), Extra Tree (XTREE), Random Forest (RF), Gradient Boosting (GB), and eXtreme Gradient Boosting (XGBoost); these trained models are further tested with input test data-sets to predict/select the appropriate RAT (4G/5G) as labelled output. Evaluation of results showed a measure of accuracy of our chosen model of RAT selection; (XGBoost) at optimal level 93.86\%, which was further cross validated at 92.9\% when compared with other algorithms for its effectiveness on future data and mitigation ability on over-fitting and under-fitting issues, hence recommended for planning and optimization purposes in similar urban/dense-urban environment to assist in maintaining the rapidly increasing demand of network connections and devices.</p>
Research data for "Exploring the configurational space of amorphous graphene with machine-learned atomic energies"
<p>This dataset supports the paper: "Exploring the configurational space of amorphous graphene with machine-learned atomic energies" (<a href="https://doi.org/10.1039/D2SC04326B">https://doi.org/10.1039/D2SC04326B</a>).</p> <p>Trajectory data for the 200-atom structures (Fig. 3) and the final configurations for the 612-atom structures as well as the GAP-17-optimised 610-atom structure from Toh et al are provided (Fig. 4). Additionally, the structures used for data analysis in Fig. 5 are given.</p> <p>The files are in extended xyz (.xyz) format and contain the raw data for coordinates, forces, and atomic energies (labelled 'c_1'). The files also contain the atomic energies relative to pristine graphene, labelled "Energy_per_atom", and the locally averaged energy relative to pristine graphene, labelled "NN_Energy_per_atom". Topological information is included at the end of the .xyz file for the 612-atom structures ('fig_4'/) and for the structures in 'fig_5/'.</p> <p>All raw atomic energies were computed using LAMMPS default settings and were output with six significant figures, with the exception of the Toh et al. structure (for which ASE was used, outputting a higher number of significant figures). </p> <p>The data can be read using, for example, the Atomic Simulation Environment (ASE), or visualised using Ovito.</p> <p> </p>
Assessing predictive performance of supervised machine learning algorithms for a diamond pricing model
<p>The diamond is 58 times harder than any other mineral in the world, and its elegance as a jewel has long been appreciated. Forecasting diamond prices is challenging due to nonlinearity in important features such as carat, cut, clarity, table, and depth. Against this backdrop, the study conducted a comparative analysis of the performance of multiple supervised machine learning models (regressors and classifiers) in predicting diamond prices. Eight supervised machine learning algorithms were evaluated in this work including Multiple Linear Regression, Linear Discriminant Analysis, eXtreme Gradient Boosting, Random Forest, k-Nearest Neighbors, Support Vector Machines, Boosted Regression and Classification Trees, and Multi-Layer Perceptron. The analysis is based on data preprocessing, exploratory data analysis (EDA), training the aforementioned models, assessing their accuracy, and interpreting their results. Based on the performance metrics values and analysis, it was discovered that eXtreme Gradient Boosting was the most optimal algorithm in both classification and regression, with a R<sup>2</sup> score of 97.45% and an Accuracy value of 74.28%. As a result, eXtreme Gradient Boosting was recommended as the optimal regressor and classifier for forecasting the price of a diamond specimen.</p>
Machine Learning Process
<p><span>Image illustrating the essential steps for applying machine learning (ML) to your data.</span></p> <p><span>The learning process of an ML model starts with </span><strong><span>data selection</span></strong><span>, which involves collecting the necessary data from various places like databases, online repositories, or real-time systems. The second step is the </span><strong><span>preprocessing</span></strong><span> or </span><strong><span>data preparation</span></strong><span>. It includes cleaning the data to remove errors or inconsistencies and handling missing data. After that, the specific attributes from the </span><strong><span>structured data</span></strong><span> </span><span>are selected</span><span> </span><span>that will</span><span> help the ML algorithm learn. Then, the data </span><span>is split</span><span> into </span><strong><span>training and testing sets</span></strong><span>. The training set </span><span>is used</span><span> to build and train the ML model, while the testing set </span><span>is used</span><span> to evaluate its performance. Then, the ML model is chosen based on the problem </span><span>and</span><span> the training data is fed into the model to create the </span><strong><span>candidate model</span></strong><span>. Once satisfied with the model's performance, </span><span>the final ML model can be deployed</span><span>.</span></p>
Supplementary data: A machine learning approach for dynamical modelling of Al distributions in zeolites via 23Na/27Al solid-state NMR
<p><strong>Content:</strong></p> <p>This dataset provides supplementary data to "A machine learning approach for dynamical modelling of Al distributions in zeolites via 23Na/27Al solid-state NMR". It contains trained Neural Network Potentials (NNP), energy and force data used for accuracy evaluation of the NNPs. Energy and forces are stored as ASE trajectory files (traj), readable by the <a href="https://wiki.fysik.dtu.dk/ase/index.html">Atomic Simulation Environment </a>(ASE). In addition, this repository contains the generated training database with DFT (SCAN+D3(BJ)) energies and forces as SchNetPack1.0 database (SiAlOHNa.db) file readable by ASE and <a href="https://github.com/atomistic-machine-learning/schnetpack/tree/schnetpack1.0">SchNetPack version 1.0</a>. Also, the structure files used to calculate NMR properties are involved.</p> <ul> <li>"nnps.zip" - (pytorch) NNP model files (compatible with <a href="https://github.com/atomistic-machine-learning/schnetpack/tree/schnetpack1.0">SchNetPack version 1.0</a>)</li> <li>"SiAlOHNa.db" - DFT (SCAN+D3(BJ)) training database as SchNetPack1.0 database file readable by ASE and <a href="https://github.com/atomistic-machine-learning/schnetpack/tree/schnetpack1.0">SchNetPack version 1.0</a></li> <li>"error_stats.zip" - traj files storing energies/forces at the DFT (SCAN+D3(BJ)) and NNP level for all test simulations to calcuate energy/force errors</li> <li>"Structures_CHA17.zip" - the structures files of CHA(17). </li> </ul>
Cloud to Thing Continuum based Sports Monitoring System using Machine Learning and Deep Learning Model
<p><span>Sports monitoring and analysis have seen significant advancements with the integration of cloud computing and continuum paradigms, facilitated by machine learning and deep learning techniques. In this study, we present a novel approach for sports monitoring that seamlessly transitions from traditional cloud-based architectures to a continuum paradigm, enabling real-time analysis and insights into player performance and team dynamics. Leveraging machine learning and deep learning algorithms, our framework offers enhanced capabilities for player tracking, action recognition, and performance evaluation in various sports scenarios. This research proposes a Cloud-to-Thing Continuum based Sports Monitoring System utilizing Machine Learning (ML) and Deep Learning (DL) models. The system integrates data acquisition, preprocessing, feature extraction, cloud-based processing, continuum paradigm integration, and decision-making stages. It leverages innovative techniques such as Improved Mask R-CNN for pose estimation, hybrid metaheuristic algorithms with Generative Adversarial Network (GAN) for classification, and fuzzy decision-making Based on the integrated analysis, decisions are made regarding player performance, team strategies, and tactical adjustments. The continuum approach ensures a balance between centralized cloud processing and distributed edge processing, optimizing resource utilization and reducing latency. Through this system, real-time analysis of sports events is achieved, enabling immediate feedback for time-sensitive applications.</span></p>
Supplementary data files for "Machine Learning Tools for Baraminology"
Open the record for dataset details and reuse information.
SurrogateLIB: An extendable library of mixed-integer programs with embedded machine learning predictors
<p>We constructed a set of Mixed-Integer Programming (MIP) instances with embedded Machine Learning (ML) predictors. The generators for these instances are available at: <a href="https://github.com/Opt-Mucca/PySCIPOpt-ML">https://github.com/Opt-Mucca/PySCIPOpt-ML</a>. The full paper which introduces SurrogateLIB and the larger Python framework PySCIPOpt-ML is available at: <a href="https://arxiv.org/abs/2312.08074">https://arxiv.org/abs/2312.08074</a></p>
Supplementary data for global distribution of mercury in foliage predicted by machine learning
<p>Global distribution of foliar mercury concentrations and pools with a spatial resolution of 0.25 latitude by 0.25 longitude, predicted by machine learning.</p>
The compiled 8-year dataset (2012-2019) consisting of weekly river water quality indicators (CODMn, DO, NH3-N and PH ) in majors 10 sub-basin of Yangtze river based on imputation of machine learning
<p>Water quality is significantly affected by global climate change and human activities, with diverse critical factors shaping its state in rivers and lakes. In the study, we utilized four indicators to characterize water quality: the physical water quality parameters included dissolved oxygen (DO, mg/L) and PH, while the chemical water quality parameters encompassed chemical oxygen demand (CODMn, mg/L) and ammonia nitrogen (NH3-N, mg/L). This study establishes weekly water quality models for typical 10 sub-basins along the Yangtze River using machine learning methods, which incorporate the impacts of hydro-meteorological and anthropogenic factors.These 10 sub-basins represent the principal tributaries of the Yangtze River basin and include Dongting Lake, the upper Han River, the lower Han River, the Jialing River, the Jinsha River, the Li River, the Min River, Poyang Lake, the Xiang River, and the Yuan River. This data collection was performed by National Environmental Monitoring Centre (http://www.cnemc.cn/sssj/szzdjczb/index_1.shtml). The water quality indicators discussed in this study are assessed in accordance with the national standard GB 3838-2002. Please refer to the paper for details.</p>
Unlabeled AnuraSet: A dataset for leveraging unlabeled data in machine learning models for passive acoustic monitoring
<p>The Unlabeled AnuraSet (U-AnuraSet) is an extension of the original AnuraSet dataset. It consists of soundscape recordings from passive acoustic monitoring conducted in Brazil. The recording sites are identical to those in the original AnuraSet. Each site comprises 2,666 one-minute raw audio files of unlabeled data. The U-AnuraSet is publicly available to encourage machine learning researchers to explore innovative methods for leveraging unlabeled data in the training of models aimed at solving problems such as anuran call identification.</p> <p>If you find the Unlabeled AnuraSet useful for your research, please consider citing it as follows:</p> <p>Cañas, J.S., Toro-Gómez, M.P., Sugai, L.S.M., et al. A dataset for benchmarking Neotropical anuran calls identification in passive acoustic monitoring. Sci Data 10, 771 (2023). https://doi.org/10.1038/s41597-023-02666-2</p>
Integrated Machine Learning model in Early Urban Flooding Warning System - Data
<p>AI_DATA.npy - Inundation data (mm) generated from MIKE+ model that has been converted to numpy array</p> <p>INDEX.npy - The index where inundation is > 0 </p> <p>source.tif - Source tif image for creating map from ML models</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.