Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
54
datasets available to search
ShareScore release 0.9.0
Dataset results
54 results for “machine learning systems”
Machine learning reveals dynamic controls of soil nitrous oxide (N2O) emissions from diverse long-term cropping systems
Open the record for dataset details and reuse information.
Data from: SLICE-MSI: A machine learning interface for system suitability testing of mass spectrometry imaging platforms
Open the record for dataset details and reuse information.
Understanding the Nature of System-Related Issues in Machine Learning Frameworks: An Exploratory Study
<p>Modern systems are built using development frameworks. The infrastructure provided by these frameworks have a major impact on how the developed system executes, how configurations are managed, how it is tested, and how and where it is deployed. Machine learning (ML) systems have revolutionized multiple industries and come with different kinds of frameworks. Naturally, the issues that manifest in such systems may differ as well---as may the behavior of developers correcting those issues. We are interested in characterizing the types of system-related issues---issues impacting performance, memory and resource usage, and other quality attributes---that emerge in machine learning frameworks, and how they differ from those in traditional frameworks. To this end, we have conducted a large-scale exploratory study analyzing real-world system-related issues from 10 popular machine learning frameworks.<br> <br> Our findings offer a number of interesting observations, with implications for the development of machine learning systems, including differences in the frequency of occurrence of certain issue types, observations regarding the impact of debate and time on issue correction, and differences in the specialization of developers. We hope that this exploratory study will enable developers to improve their expectations, plan for risk, and allocate resources accordingly when making use of the tools provided by these frameworks to develop ML-based systems.</p>
Dynamical ising dataset for the paper Machine learning stochastic differential equations for the evolution of order parameters of classical many-body systems in and out of equilibrium
<p>This dataset provide the evolution in time for the magnetizaion in the 2D Ising model evolved with Gluber dynamics for a lattice of size 64 x 64.</p>
Machine Learning for predicting chaotic systems – Data
<p>The data used in our article "Machine Learning for Predicting Chaotic Systems" - <a href="https://arxiv.org/abs/2407.20158">https://arxiv.org/abs/2407.20158</a></p> <p>DeebDbDysts*.zip contain the Dysts database, DeebDbLorenz*.zip the DeebLorenz database (with DeebDbLorenzBig*.zip being the "extension" dataset for Lorenz63std with different time series lengths).</p> <p>The observation and truth data of the Dysts database originates from <a href="https://github.com/williamgilpin/dysts">https://github.com/williamgilpin/dysts</a> (we converted the data format from json to csv).</p> <p>For DeebLorenz, we used the R package <a href="https://github.com/chroetz/DEEBdata">DEEBdata</a> to create it.</p>
An automated system for inspecting rock faces and detecting potential rock falls using machine learning
<p>Rockfall is a hazard in mountainous areas threatening infrastructure and human lives. Rockfall hazards are often mitigated by manual inspections using pry bars. The inspector must access the rock face, hit the rock surface, detect, and remove the loose rocks. This method is very labor demanding, unsafe, and challenging. This research presents a method that automatize the inspection of rock blocks that are prone to rockfall events. A robot is developed to replace the manual hammer tap process and collect the sound data remotely; subsequently, the sound signal is used to identify different types of the discontinuity in rocks in controlled laboratory environment. Machine learning is used to train the method to discriminate between intact rock and rock that may be prone to fall. This methodology was successfully applied to laboratory tests on rock. Finally, the research involves the implementation of this system in field to understand the potential and limitations of the proposing system in automatizing the rock inspections. This research enables the inspectors to collect data remotely, detect loose rocks, and save data for future references.</p>
Code and extensive data for training neural networks for radiation, used in "Implementation of a machine-learned gas optics parameterization in the ECMWF Integrated Forecasting System: RRTMGP-NN 2.0""
<p>Data and code used in a paper submitted to JAMES titled :<em> Implementation of a machine-learned gas optics parameterization in the ECMWF Integrated Forecasting System</em></p> <p>1) The files <strong>ml_training_*.7z</strong> contain extensive datasets (in NetCDF format) for training neural network versions of the RRTMGP gas optics scheme as described in the paper. The datasets are read by <a href="https://github.com/peterukk/rte-rrtmgp-nn/blob/main/examples/rrtmgp-nn-training/ml_train.py">ml_train.py.</a></p> <p>2) The ML datasets were in turn generated using the input profiles (in NetCDF format) inside <strong>inputs_to_RRTMGP.zip </strong>by running the Fortran programs <code>rrtmgp_sw_gendata_rfmipstyle.F90 and rrtmgp_lw_gendata_rfmipstyle.F90 </code>in <em>rte-rrtmgp-nn/examples/rrtmgp-nn-training</em>, which call the RRTMGP gas optics scheme, The input profiles contain <strong>millions of columns, hundreds of perturbation experiments (including hypercube-sampled gas concentrations), are derived from several different data sources (including CAMS reanalysis, GCM, and CKDMIP-MMM), and span present-day, preindustrial, and future atmospheric conditions.</strong> They could be used to generate training data for developing emulators of the full RTE+RRTMGP radiation scheme, not just gas optics (see nn_dev on the <a href="https://github.com/peterukk/rte-rrtmgp-nn">RTE+RRTMGP-NN repository on Github</a>, used in a previous paper where different emulation methods were compared)</p> <p>3) The Fortran and Python code used for data generation and NN training are found in<a href="https://github.com/peterukk/rte-rrtmgp-nn/tree/main/examples/rrtmgp-nn-training"> <em>rte-rrtmgp-nn/examples/rrtmgp-nn-training</em> </a>on the main branch on Github; <strong>an archived version is also included here </strong>(<strong>rte-rrtmgp-nn-2.0.zip</strong>). See the readme in the above sub-directory for further information.</p> <p> </p>
Semantic Web resources and Machine Learning systems - Knowledge Graph (SWeMLS-KG)
<p>This resource is part of our submission to ESWC 2023 resource track, which includes:</p> <p>Datasets:<br> - Folder "pattern" - a set of SWeMLS patterns represented based on OPMW and P-Plan ontology,<br> - Folder "shapes" - a set of SHACL constraints to check the conformance of SWeML Systems against SWeMLS patterns as well as a set of SHACL-AF rules to generate links between system components,<br> - File "swemls-ontology.ttl" - an ontology to represent Semantic Web resources and Machine Learning systems (SWeMLS),<br> - File "swemls-instances.ttl" - a set of triples representing the extracted metadata from 476 SWeML systems and papers,<br> - File "swemls-kg.ttl" - an integrated and validated KG containing all above files, including enrichment from SHACL-AF rules using "swemls-toolkit" [2].</p> <p>These resources are produced based on the result of the Systematic Mapping Study (SMS) reported in [1]. The latest SNAPSHOT-version of the resource can be accessed through our resource landing page: <a href="https://w3id.org/semsys/sites/swemls-kg/">https://w3id.org/semsys/sites/swemls-kg/</a></p> <p>[1] Breit, A., Waltersdorfer, L., Ekaputra, J.F., Sabou, M., Ekelhart, A., Iana, A., Paulheim, H., Portisch, J., Revenko, A., Ten Teije, A., van Harmelen, F.: Combining Machine Learning and Semantic Web -A Systematic Mapping Study (under review). ACM CSUR (2022)<br> [2] Source code of swemls-toolkit is available at: https://github.com/semanticsystems/swemls-toolkit</p>
Replication Package for Identifying Self-Admitted Technical Debt in Issue Tracking Systems using Machine Learning
<p>This dataset includes pre-trained word embeddings and a weighted file that can be used to identify self-admitted technical debt (SATD) from issue tracking systems.</p>
Data from: Machine learning improves predictions of agricultural nitrous oxide (N2O) emissions from intensively managed cropping systems
<p><span>The potent greenhouse gas nitrous oxide (N</span><sub><span>2</span></sub><span>O) is accumulating in the atmosphere at unprecedented rates largely due to agricultural intensification, and cultivated soils contribute ~60% of the agricultural flux. Empirical models of N</span><sub><span>2</span></sub><span>O fluxes for intensively managed cropping systems are confounded by highly variable fluxes and limited </span><span><span>geographic coverage;</span></span><span> process-based biogeochemical models are rarely able to predict daily to monthly emissions with > 20% accuracy even with site-specific calibration. Here we show the promise for machine learning (ML) to significantly improve field-level flux predictions, especially when coupled with a cropping systems model to simulate unmeasured </span><span><span>soil</span></span><span> parameters. We used sub-daily N</span><sub><span>2</span></sub><span>O flux data from six years of automated flux chambers installed in a continuous corn rotation at a site in the upper U.S. Midwest (~3000 sub-daily flux observations), supplemented with weekly to biweekly manual chamber measurements (~1100 daily fluxes), to train an ML model that explained 65-89% of daily flux variance with very few input variables –soil moisture, days after fertilization, soil texture, air temperature, soil carbon, precipitation, and N fertilizer rate. When applied to a long-term test site not used to train the model, the model explained 38% of the variation observed in weekly to biweekly manual chamber measurements from corn, and 51% upon coupling the ML model with a cropping systems model that predicted daily soil N availability. </span><span><span>This represents a 2-3 times improvement over conventional process-based models and with substantially fewer input requirements.</span></span><span> This coupled approach </span><span><span>offers promise</span></span><span> for better predictions of agricultural N</span><sub><span>2</span></sub><span>O emissions and thus more precise global models and more effective </span><span><span>agricultural mitigation interventions.</span></span></p>
Demonstration of Portable Performance of Scientific Machine Learning on High Performance Computing Systems
<p>With the largest datasets to date and a diverse set of discoveries to be made, the current generation of scientific analyses are well poised to utilize artificial intelligence (AI) and machine learning (ML) on high performance computing (HPC) resources. Like never before, these workflows can be written in one portable language, python, which thanks to highly-optimized ML libraries achieves excellent cross-platform performance with little to no intervention by the user. In this demonstration, we explore the performance of several scientific AI/ML applications across leading HPC resources and highlight best practices for portable performance.</p>
Ripple: A Long-Sighted Self-Adaptation Approach to Retrain Machine-Learning-Enabled Systems
<p>Data files required to reproduce the results of paper "Ripple: A Long-Sighted Self-Adaptation Approach to Retrain Machine-Learning-Enabled Systems" submitted to ICSME 2025</p>
Data for "Supervised machine learning methods for crystal structure prediction of the binary Cs-Te system"
<p>Crystal structures, high-throughput calculations and trained machine learning models presented in the paper "Supervised machine learning methods for crystal structure prediction of the binary Cs-Te system".</p> <ul> <li><em>crystal_datasets </em>contains the input/output data sets of crystal structures for high-throughput calculations and ML models.</li> <li><em>aiida_ht_calculations </em>contains the data regarding the high-throughput DFT calculations.</li> <li><em>ml_models</em> contains the trained ML models.</li> </ul> <p>Eeach zip-archive contains a jupyter-notebook examplifying how the data can be accessed and reused.</p>
Tricycle Accident Prevention and Control System using Machine Learning Techniques
<p>Tricyle accident prevention and control system using a feed forward neural network.</p>
Use of Machine Learning Techniques for Serial Assessment of Systemic Inflammatory Markers in Breast Cancer Patients
ClinicalTrials.gov study NCT06447532. IPD Sharing: NO. Countries: 8. Publications: 3.
Effects of a Machine Learning-based Lower Limb Exercise Training System for Knee Pain
ClinicalTrials.gov study NCT05173064. IPD Sharing: NO. Countries: 1. Publications: 14.
Data from: Machine learning improves predictions of agricultural nitrous oxide (N2O) emissions from intensively managed cropping systems
Open the record for dataset details and reuse information.
Black box attack on machine learning assisted wide area monitoring and protection systems
Open the record for dataset details and reuse information.
An Empirical Study of Refactorings and Technical Debt in Machine Learning Systems
<p>Machine Learning (ML), including Deep Learning (DL), systems, i.e., those with ML capabilities, are pervasive in today's data-driven society. Such systems are complex; they are comprised of ML models and many subsystems that support learning processes. As with other complex systems, ML systems are prone to classic technical debt issues, especially when such systems are long-lived, but they also exhibit debt specific to these systems. Unfortunately, there is a gap of knowledge in how ML systems actually evolve and are maintained. In this paper, we fill this gap by studying refactorings, i.e., source-to-source semantics-preserving program transformations, performed in real-world, open-source software, and the technical debt issues they alleviate. We analyzed 26 projects, consisting of 4.2 MLOC, along with 327 manually examined code patches. The results indicate that developers refactor these systems for various reasons, both specific and tangential to ML; some refactorings correspond to established technical debt categories. In contrast, others do not, and code duplication is a major cross-cutting theme that particularly involved ML configuration and model code, which was also the most refactored. We also introduce 14 and 7 new ML-specific refactorings and technical debt categories, respectively, and put forth several recommendations, best practices, and anti-patterns. The results can potentially assist practitioners, tool developers, and educators in facilitating long-term ML system usefulness.</p>
A Comparison of Machine-Learning Assisted Optical and Thermal Camera Systems for Beehive Activity Counting
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.