Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
3,202
datasets available to search
ShareScore release 0.9.0
Dataset results
3,202 results for “maintenance”
Differential brain mechanisms of selection and maintenance of information during working memory (MEG data)
Open the record for dataset details and reuse information.
Metadata for the urbisphere-Paris campaign during 2022-2024: fieldwork maintenance log [L1]
<p>Machine-readable, formatted and redacted electronic fieldwork logs from the urbisphere-Paris observation campaign conducted between 2022-09-05 and 2024-07-22 in Paris, France. Provided in text format with comma separated columns (.csv) and in Microsoft Excel (.xlsx) format.</p> <p>The fieldwork logs are created from raw google form data submitted by campaign managers, scientists, technicinas and students. The formatting process is detailed in https://github.com/Urban-Meteorology-Reading/urbisphere-paris-fieldwork-log-format. The GitHub output has then been manually edited and adjusted.</p> <p>Contains maintenance information for the following observational sites operated as part of the urbisphere Paris campaign 2022 - 2024:</p> <table> <tbody> <tr> <td>PAARBO</td> <td>Paris – Arboretum de Vallée-aux-Loups </td> </tr> <tr> <td>PAAUNA</td> <td>Paris – Aunay-sous-Auneau</td> </tr> <tr> <td>PABOBI</td> <td>Paris – Bobigny</td> </tr> <tr> <td>PABONN</td> <td>Paris – Bonniel</td> </tr> <tr> <td>PABPAC</td> <td>Paris – Balloon Parc Andre Citroën</td> </tr> <tr> <td>PACHAM</td> <td>Paris – Chamant</td> </tr> <tr> <td>PACHAN</td> <td>Paris – Changis-sur-Marne </td> </tr> <tr> <td>PACHEM</td> <td>Paris – Chemin Vert Bobigny</td> </tr> <tr> <td>PACOMP</td> <td>Paris – Compiègne</td> </tr> <tr> <td>PACOUR</td> <td>Paris – Courdimanche-sur-Essonne</td> </tr> <tr> <td>PACRET</td> <td>Paris – Créteil</td> </tr> <tr> <td>PADENF</td> <td>Paris – Denfert Rocherau </td> </tr> <tr> <td>PADROU</td> <td>Paris – Droue Sur Drouette</td> </tr> <tr> <td>PAHOTE</td> <td>Paris – Hôtel de Ville</td> </tr> <tr> <td>PAJUSS</td> <td>Paris – Jussieu </td> </tr> <tr> <td>PALUPD</td> <td>Paris – Université Paris Diderot (LISA Platform)</td> </tr> <tr> <td>PAMEUD</td> <td>Paris – Meudon</td> </tr> <tr> <td>PANANG</td> <td>Paris – Nangis</td> </tr> <tr> <td>PANATI</td> <td>Paris – Rue Nationale</td> </tr> <tr> <td>PAPRUN</td> <td>Paris – Prunay-le-Temple</td> </tr> <tr> <td>PAROIS</td> <td>Paris – Roissy</td> </tr> <tr> <td>PAROMA</td> <td>Paris – Romainville</td> </tr> <tr> <td>PASIRT</td> <td>Paris – SIRTA Observatory Palaiseau</td> </tr> <tr> <td>PASTFE</td> <td>Paris – Saint Félix</td> </tr> <tr> <td>PAWYDT</td> <td>Paris – Wy-dit-Joli-Village</td> </tr> </tbody> </table> <p> </p>
Predictive maintenance of the Baghouse [CAO1]
<p> </p> <p>The baghouse filter consists of a collector which removes dust, mainly filler content in dry aggregates during drying process in drum. Baghouse performance is heavily depended on inlet and outlet gas temperature and flow speed as well as opacity climatic conditions and pressure drop, in the bag house (temperature and humidity, the recipe of asphalt). The final user of this asset is the plant operator (EIFFAGE), which is interested in the predictive maintenance of the baghouse.</p> <p>In the context of the developed Cognitive Solution (CS), two prediction models will be developed:</p> <ul> <li>The first model predicts if the baghouse is working properly by attempting to identify any abnormal behavior.</li> <li>The second model predict the remaining useful life of a component of baghouse (days or hours).</li> </ul> <p>The former is utilized for generating a set of alarms based on the process measurements while the latter is utilized for providing an estimation of the saturation level of the filters as well as a value of the evolution of the saturation level of the filters.</p>
Respectable Standards of Living: The Alternative Lens of Maintenance Costs, Britain 1270-1860
<p>Data set and code book. Replication materials for paper accepted in Economic History Review, April 2024.</p> <p>Abstract </p> <p><span>This paper argues that in all societies there is considerable agreement about what goods and services are needed to provide a decent living, and that this standard can be measured by the expense involved in maintaining people of good standing.<span> </span><span> </span>Maintenance costs include two components of living costs that are neglected in conventional approaches.<span> </span>First, in contrast to the usual focus on a fixed basket of commodities, maintenance costs capture changes in the composition and quality of the goods required for a respectable lifestyle.<span> </span>Second, unlike the conventional accounting they include the costs of the household services required to turn the basket commodities into livings. Ignored in the conventional methodology, the inclusion of these costs represents a core innovation. More than 4600 observations, drawn mainly from primary sources, trace levels and trends in maintenance costs for Britain, 1270-1860. <span> </span>These can be compared with established cost of living indicators to offer a complementary perspective on real consumption that accommodates aspirational goods and the input of household labour.<span> </span>The struggle to support families at respectable standards emerges as driving industriousness and motivating prudence among a class that played a major role in economic development.<span> </span><span> </span></span></p> <p> </p>
TCM: Benchmark Datasets for Predictive Maintenance in Steel Manufacturing
<h1>Anomaly-TCM</h1> <p>Predictive Maintenance (PdM) is a strategy that uses advanced data analytics to predict equipment failures and maintain industrial machinery in good condition. Its goals are to minimize downtime, reduce operational costs, and ensure product quality. PdM methods are applicable across various industries, including steel manufacturing.</p> <p>In steel production, cold rolling is a critical process that reduces the thickness of hot-rolled steel. Developing PdM methods for tandem cold mills (TCM) can significantly improve production efficiency. However, researchers often rely on real manufacturing data, which is typically unavailable, unlabeled, and noisy, making it difficult to validate and compare methods.</p> <p>To overcome this, we created synthetic datasets for the cold rolling process to identify anomalies based on physical principles. These datasets were generated using a mathematical model of a 5-stand TCM, calculating key process parameters like rolling force, torque, speed, tension, gap, thickness reduction, and motor power. We introduced anomalies related to specific failures in the process.</p> <p>We produced six diverse datasets, each with varying complexity, to enable benchmarking of machine learning-based PdM methods for the cold rolling process. Four different types of anomalies were introduced, which are related to a physics-based deviations in the process:</p> <ol> <li>Anomaly in reduction scheme</li> <li>Anomaly in work roll (increased work roll friction)</li> <li>Anomaly in bearing (increased motor torque)</li> <li>Anomaly in electric motor (decrease efficiency)</li> </ol> <p> The details of the datasets are provided below.</p> <table> <tbody> <tr> <td><strong>Dataset</strong></td> <td><strong>Observations</strong></td> <td><strong>Anomalies</strong></td> <td><strong>Share of Anomalies</strong></td> <td><strong>Features</strong></td> <td><strong>Anomaly Types</strong></td> <td><strong>Products</strong></td> <td><strong>Data Drift</strong></td> </tr> <tr> <td>tcm5_dataset_1</td> <td>20009</td> <td>1045</td> <td>5.2%</td> <td>51</td> <td>1</td> <td>4</td> <td>FALSE</td> </tr> <tr> <td>tcm5_dataset_2</td> <td>20001</td> <td>1035</td> <td>5.2%</td> <td>51</td> <td>1</td> <td>20</td> <td>FALSE</td> </tr> <tr> <td>tcm5_dataset_3</td> <td>20003</td> <td>981</td> <td>4.9%</td> <td>51</td> <td>4 (16)</td> <td>4</td> <td>FALSE</td> </tr> <tr> <td>tcm5_dataset_4</td> <td>20001</td> <td>925</td> <td>4.6%</td> <td>51</td> <td>4 (16)</td> <td>20</td> <td>FALSE</td> </tr> <tr> <td>tcm5_dataset_5</td> <td>20005</td> <td>1031</td> <td>5.2%</td> <td>51</td> <td>4 (16)</td> <td>5</td> <td>TRUE</td> </tr> <tr> <td>tcm5_dataset_6</td> <td>20008</td> <td>954</td> <td>4.8%</td> <td>51</td> <td>4 (16)</td> <td>25</td> <td>TRUE</td> </tr> </tbody> </table> <p> </p> <p>Each dataset is generated as a data stream, meaning the observations follow a chronological order, represented by increasing work roll mileage (which is reset after a predefined threshold). The table below provides details about the features and labels present in the datasets. Several features are recorded for each rolling stand, totaling 51 features. Apart from the anomaly related to reduction, the other anomalies are specific to individual stands, resulting in 16 anomaly labels in total.</p> <table> <tbody> <tr> <td><strong>Feature</strong></td> <td><strong>Suffixes</strong></td> <td><strong>Unit</strong></td> <td><strong>Description</strong></td> </tr> <tr> <td>thickness_entry</td> <td>-</td> <td>mm</td> <td>steel entry thickness</td> </tr> <tr> <td>thickness_exit</td> <td>-</td> <td>mm</td> <td>steel exit thickness</td> </tr> <tr> <td>width</td> <td>-</td> <td>mm</td> <td>steel width</td> </tr> <tr> <td>ys_entry</td> <td>-</td> <td>MPa</td> <td>steel entry yield strength</td> </tr> <tr> <td>ys_exit</td> <td>-</td> <td>MPa</td> <td>steel exit yield strength</td> </tr> <tr> <td>work_roll_diam</td> <td>1 to 5</td> <td>mm</td> <td>work roll diamaeter (stands 1 to 5)</td> </tr> <tr> <td>work_roll_mileage</td> <td>1 to 5</td> <td>km</td> <td>work roll mileage (stands 1 to 5)</td> </tr> <tr> <td>reduction</td> <td>1 to 5</td> <td>-</td> <td>thickness reduction (stands 1 to 5)</td> </tr> <tr> <td>tension</td> <td>0 to 5</td> <td>N</td> <td>interstand tension (0 is tension before stand 1, 1-5 refer to tension after stands 1-5)</td> </tr> <tr> <td>roll_speed</td> <td>1 to 5</td> <td>NaN</td> <td>linear work roll speed (stands 1 to 5)</td> </tr> <tr> <td>force</td> <td>1 to 5</td> <td>N</td> <td>rolling force (stands 1 to 5)</td> </tr> <tr> <td>torque</td> <td>1 to 5</td> <td>Nm</td> <td>rolling torque (stands 1 to 5)</td> </tr> <tr> <td>gap</td> <td>1 to 5</td> <td>mm</td> <td>stand gap (stands 1 to 5)</td> </tr> <tr> <td>motor_power</td> <td>1 to 5</td> <td>kW</td> <td>electric motor power (stands 1 to 5)</td> </tr> <tr> <td>Anomaly_Reduction</td> <td>-</td> <td>-</td> <td>(label) anomaly in reduction scheme</td> </tr> <tr> <td>Anomaly_Electric</td> <td>1 to 5</td> <td>-</td> <td>(label) anomaly in electric motor (stands 1 to 5)</td> </tr> <tr> <td>Anomaly_Bearing</td> <td>1 to 5</td> <td>-</td> <td>(label) anomaly in stand bearing (stands 1 to 5)</td> </tr> <tr> <td>Anomaly_WorkRoll</td> <td>1 to 5</td> <td>-</td> <td>(label) anomaly in work roll friction (stands 1 to 5)</td> </tr> </tbody> </table>
Dataset for "Too Simple? Notions of Task Complexity used in Maintenance-based Studies of Programming Tools"
<p>This dataset contains the data to replicate the findings in the publication "Too Simple? Notions of Task Complexity used in Maintenance-based Studies of Programming Tools". The dataset includes the bibliographies and intermediate analysis tables for the analyzed literature. Further, it contains the experiment materials providing context for the task discussed in Section V of the publication.</p> <p>The artifact contains the following files:</p> <ul> <li>icpc-from-survey.bib: The bibliographic entries of all publications selected from the corpus of the paper "Confounding parameters on program comprehension: a literature survey" by Janet Siegmund and Jana Schumann (https://doi.org/10.1007/s10664-014-9318-8)</li> <li>icpc-manually-added.bib: The bibliographic entries of manually added publications.</li> <li>icpc-phase-2-and-3.[csv|xlsx]: The intermediate analysis tables for phases 2 and 3 as described in the paper. The table also contains columns transferred from the original dataset of the Siegmund and Schumann study (marked as "transferred")</li> <li>experiment-materials.zip: Includes the materials for the experiment used in Section V <ul> <li>introduction-english.[docx|rtf]: The experiment protocol we read to participants, including the introduction to the system architecture.</li> <li>JumpODrom.package.zip: The code of the used system.</li> <li>task-16-description.txt: The textual description of the task discussed in Section V.</li> <li>task-16.patch: The change applied to the system code prior to the task, which causes the faulty behavior outlined in the task description.</li> </ul> </li> </ul> <p> </p>
Eyebright species maintenance (scripts and data accompanying Becher et al., Plant Communications)
<p><strong>This gzipped TAR ball contains data and scripts related to the study on Fair Isle eyebrights by Hannes Becher, Max R. Brown, Gavin Powell, Chris Metherell, Nick J. Riddiford, and Alex D. Twyford, submitted to Plant Communications.</strong></p> <p>Data: genome assembly of<em> Euphrasia arctica</em>, variant call files of the "tetraploid" and "conserved" sets of scaffolds, per-individual k-mer spectra, mapping depths, etc.</p> <p>Scripts: R scripts for the analysis of plant trait data, heterozygosity, ect.; an ipython notebook for the analysis of variant data, and a Mathematica notebook with the derivation of the formulae used to fit pop gen parameters to k-mer spectra.</p>
Publication and Maintenance of Relational Data in Enterprise Knowledge Graphs Created (Files used in the experiments)
<p>This dataset contains two files created for the experiments presented in the article: Publication and Maintenance of RDB2RDF Views Externally Materialized in Enterprise Knowledge Graphs.</p> <p><strong>mapR2RML_MusicBrainz_completo.txt</strong><strong>:</strong> We created the R2RML mapping for translating MBD data into the Music Ontology vocabulary, which is used for publishing the LMB view. The LMB view was materialized using the D2RQ tool. It took 67 minutes to materialize the view with approximately 41.1 GB of NTriples. We also provided SPARQL endpoint for querying LMB View.</p> <p><strong>TriggersAndProcedures.txt</strong>: We created the triggers, procedures, and class in java to implement the rules required to compute and publish the changesets.</p> <p><strong>relationalViewDefinition.pdf</strong>: This document gives details about the process of creating the relational views used in the experiments.</p>
Joint Optimization of Production and Maintenance for Cost-effective Manufacturing and Demand Response Participation Dataset - Machine Breakdown Event
<p>Using the previous dataset at <<a href="https://zenodo.org/record/4106746">https://zenodo.org/record/4106746</a>> an announcement of a machine breakdown event was simulated on Friday at 6:00, describing that machine MAQ119 could breakdown at any moment, detected using a predictive maintenance system. Accordingly, the proposed scheduler imposed a machine available frames constraint, during its event, of 0 usable frames, thus removing the machine from production. The proposed genetic algorithm was executed for 1 hour at period 769 (Friday at 7:00) until the remainder of the schedule’s time window. Also, the predefined optimization weights were 1 for total cost and 0 for machine occupancy deviation.</p> <p> </p> <p>File Description:</p> <ul> <li>Input_JSON_Machine_Breakdown_Optimization - JSON input data for the machine breakdown event</li> <li>Output_JSON_Machine_Breakdown_Optimization - JSON output data for the machine breakdown event</li> <li>Output_Statistics_Machine_Breakdown_Optimization - Excel output machine breakdown event statistics</li> </ul>
Joint Optimization of Production and Maintenance for Cost-effective Manufacturing and Demand Response Participation Dataset - Maintenance Optimization
<p>Using the previous dataset at <<a href="https://zenodo.org/record/4106746">https://zenodo.org/record/4106746</a>> a maintenance optimization scenario was formulated to validate the scheduler's ability to schedule tasks as well as maintenance activities while also minimizing the total costs. Accordingly, it was considered an optimization weight of 1 for the total cost and 0 for machine occupancy deviation, as well as a 2-hour execution time for the genetic algorithm. Each maintenance activity, for every machine, has a duration of 6 hours and 10 minutes, with a labor cost of 3,22 EUR/hour during the stipulated maintenance hours and a monetary penalty, that doubles the cost (i.e., 6.44 EUR/hour) if done out of maintenance hours.</p> <p> </p> <p>File Description:</p> <ul> <li>Input_JSON_Maintenance_Optimization - JSON input data for the maintenance optimization</li> <li>Output_JSON_Maintenance_Optimization - JSON output data for the maintenance optimization</li> <li>Output_Statistics_Maintenance_Optimization - Excel output maintenance optimization statistics</li> </ul>
Joint Optimization of Production and Maintenance for Cost-effective Manufacturing and Demand Response Participation Dataset - Total Cost and Machine Occupancy Deviation Optimization
<p>Using the previous dataset at <<a href="https://zenodo.org/record/4106746">https://zenodo.org/record/4106746</a>> a total cost and machine occupancy deviation optimization scenario was formulated that aims to demonstrate how the proposed scheduler is able to balance tasks between machines while also reducing overall costs. For this scenario, it was considered an optimization weight of 0.5 for both the total costs and machine occupancy deviation objectives and the genetic algorithm was executed for 2 hours.</p> <p> </p> <p>File Description:</p> <ul> <li>Input_JSON_Total_Cost_Machine_Occupancy_Deviation_Optimization - JSON input data for the total cost and machine occupancy deviation optimization</li> <li>Output_JSON_Total_Cost_Machine_Occupancy_Deviation_Optimization - JSON output data for the total cost and machine occupancy deviation optimization</li> <li>Output_Statistics_Total_Cost_Machine_Occupancy_Deviation_Optimization - Excel output total cost and machine occupancy deviation optimization statistics</li> </ul>
Joint Optimization of Production and Maintenance for Cost-effective Manufacturing and Demand Response Participation Dataset - Energy Cost Optimization with Energy Selling
<p>Using the previous dataset at <<a href="https://zenodo.org/record/4106746">https://zenodo.org/record/4106746</a>> an energy cost optimization considering the presence of an energy buyer is proposed to validate the scheduler’s ability to maximize profits while also minimizing energy costs. The scenario considers an added sales value corresponding to 50% of the buying. For this scenario, the genetic algorithm was executed for 2 hours, with 1 and 0 for the optimization weights total cost and machine occupancy deviation, respectively.</p> <p> </p> <p>File Description:</p> <ul> <li>Input_JSON_Energy_Cost_Energy_Selling_Optimization - JSON input data for the energy cost optimization with energy selling</li> <li>Output_JSON_Energy_Cost_Energy_Selling_Optimization - JSON output data for the energy cost optimization with energy selling</li> <li>Output_Statistics_Energy_Cost_Energy_Selling_Optimization - Excel output energy cost optimization with energy selling statistics</li> </ul>
Joint Optimization of Production and Maintenance for Cost-effective Manufacturing and Demand Response Participation Dataset - Joint Optimization of Production and Maintenance
<p>Using the previous datasets at <<a href="https://zenodo.org/record/4106746">https://zenodo.org/record/4106746</a>>, <<a href="https://zenodo.org/record/7055698">https://zenodo.org/record/7055698</a>>, <<a href="https://zenodo.org/record/7055580">https://zenodo.org/record/7055580</a>>, and <<a href="https://zenodo.org/record/7055573">https://zenodo.org/record/7055573</a>> a joint optimization of production and maintenance scenario was formulated which aims at combining all the features from the cited scenarios. For energy selling, it was considered an added sales value corresponding to 50% of the buying. Regarding maintenance activities, it was simulated an announcement of a maintenance activity for MAQ118 from Monday at 07:00 (i.e., period 1) to Monday at 17:00 (i.e., period 120), and another for MAQ120 which can be done at any time. These maintenance activities take 6 hours and 10 minutes to complete and have an associated labor cost of 3,22 EUR/hour in maintenance hours, and a double cost penalty (i.e., 6.44 EUR/hour) if done out of maintenance hours. The scenario was executed in 2 hours, with the corresponding optimization weights of 0.8 and 0.2 for the total cost and machine occupancy deviation, respectively.</p> <p> </p> <p>File Description:</p> <ul> <li>Input_JSON_Joint_Optimization_Production_Maintenance - JSON input data for the joint optimization of production and maintenance</li> <li>Output_JSON_Joint_Optimization_Production_Maintenance - JSON output data for the joint optimization of production and maintenance</li> <li>Output_Statistics_Joint_Optimization_Production_Maintenance - Excel output joint optimization of production and maintenance statistics</li> </ul>
Maintenance of Wakefulness Test (MWT) recordings
<p>Each file contains a MWT trial (first trial after noon) recording of a patient. The data contains occipital EEG and EOG data. All signals were bandpass filtered between 0.5-45 Hz.</p> <p>In each file, the data is structured as the following:</p> <ul> <li>fs: sampling rate.</li> <li>eeg_O1: EEG channel O1-M2 where M2 is the mastoid electrode on the opposite side.</li> <li>eeg_O2: EEG channel O2-M1 where M1 is the mastoid electrode on the opposite side.</li> <li>E1 and E2: EOG channels for left and right eye, both referenced to M1.</li> <li>labels_O1 and labels_O2: arrays with expert scoring (0-wake, 1-MSE, 2-MSEc, 3-ED, according to the BERN scoring criteria published in Hertig-Godeschalk et al. doi:10.1093/sleep/zsz163.); length of the arrays is the same as for other signals, i.e. there is a label per sample.</li> <li>prec: amount of signal samples per label, in this case it is 1. variables prec and half_prec were not used.</li> <li>num_Labels: length of the signal in samples.</li> </ul> <p>Further descriptions, details, and outcomes can be found in the related studies. The published studies which are based on this data and address the borderland between wakefulness and sleep, i.e. microsleep episodes, are listed under related/alternative identifiers.</p>
Maintenance of Convectively Coupled Kelvin waves: Relative Importance of Internal Thermodynamic Feedback and External Momentum Forcing (Code and Data)
<p>This is the dataset and code for generating all figures for the journal article named "Maintenance of Convectively Coupled Kelvin Waves: Relative Importance of Internal Thermodynamic Feedback and External Momentum Forcing," The article was written by Mu-Ting Chien and Daehyun Kim and submitted to Geophysical Research Letters in 2024.</p>
Carbon emission and lifecycle costs supporting digital twins for managing railway maintenance and resilience
<p>The development of railway construction increases the system complexity, which results in difficulty in management with traditional methods. Building Information Modelling (BIM) as an interoperable concept is benefits via whole life-cycle assessment (LCA) of the project, and it has been widely adopted in architecture, construction, and engineering (ACE) fields. This dataset of lifecycle cost and carbon footprint supports the digital twins for managing railway maintenance and resilience.</p>
The monthly operating costs (vehicle, fuel, and maintenance) of each compared vehicle, Tesla 3 (283 HP), and Infiniti Q50 (300 HP).
<p>We compared two vehicles with similar horsepower, Tesla 3 (283 HP), and Infiniti Q50 (300 HP). The monthly operating costs of each compared vehicle were:</p> <ul> <li> <p>for the model, Tesla 3 electric vehicle was US$426.10/month, including purchase and depreciation US$333/month, fuel (electricity) US$27.8/month, maintenance US$65.3/month.</p> </li> <li> <p>for the model, Infiniti Q50, the internal combustion engine car was US$583.7/month, including purchase and depreciation US$321/month, fuel US$166/month, maintenance US$96.7/month.</p> </li> </ul>
Maintenance of Plan Libraries for Case-based Planning
<p>Case-based planning is an approach to planning where previous planning experience provides guidance to solving new problems. Such a guidance can be extremely useful, or even necessary, when the new problem is very hard to solve, or the stored previous experience is highly valuable, because, e.g., it was provided or validated by human experts, and the system should try to reuse it as much as possible. <br> <br> To do so, a case-based planning system stores in a library previous planning experience in the form of already encountered problems and their solutions. </p> <p>The quality of such a plan library critically influences the performance of the planner, and therefore it needs to be carefully designed and created. For this reason, it is also important to update the library during the lifetime of the system, as the type of problems being addressed may evolve or differ from the ones the library was originally designed for. Moreover, like in general case-based reasoning, the library needs to be maintained at a manageable size, otherwise the computational cost of querying it grows excessively, making the entire approach ineffective.</p>
1151 commits with software maintenance activity labels (corrective,perfective,adaptive)
<p>Data format: CSV</p> <p>Separator character: '#'</p> <p><strong>This dataset contains 1151 commits manually labeled with maintenance activities ("c" for corrective, "p" for perfective, "a" for adaptive)</strong> according to the definition by Mockus et al. in <em>"Mockus, A. and Votta, L.G., 2000, October. Identifying Reasons for Software Changes using Historic Databases. In icsm (pp. 120-130)"</em>.</p> <p>In addition, this dataset also contains <strong>further information (features) extracted from the commits</strong>:</p> <ol> <li>The <strong>source code changes</strong> performed by the commit author as part of a given commit (statement added, statement removed, etc.) <ul> <li>The source code change taxonomy is detailed in <em>"Fluri, B. and Gall, H.C., 2006, June. Classifying change types for qualifying change couplings. In Program Comprehension, 2006. ICPC 2006. 14th IEEE International Conference on (pp. 35-45). IEEE."</em></li> </ul> </li> <li>A binary indication (1/0) whether a given commit contains any of the <strong>keywords from a pre-computed </strong>(according to a word frequency analysis)<strong> set of keywords</strong> <strong>indicative of each maintenance activity</strong>.</li> </ol> <p>The dataset consists of commits sampled from the following open source projects:</p> <ol> <li>RxJava</li> <li>hbase</li> <li>elasticsearch</li> <li>intellij-community</li> <li>hadoop</li> <li>drools</li> <li>kotlin</li> <li>restlet-framework-java</li> <li>orientdb</li> <li>camel</li> <li>spring-framework </li> </ol> <p>This dataset is a supporting material for the paper <strong>"Boosting Automatic Commit Classification Into Maintenance Activities By Utilizing Source Code Changes", to appear in PROMISE 2017.</strong></p>
Machine Learning-Based Bridge Maintenance Optimization Model for Maximizing Performance within Available Annual Budgets
<p>Effective maintenance planning for bridges is crucial for maintaining their performance, safety, and minimizing maintenance costs. Timely implementation of interventions can improve the performance of bridges and avoid the need for costly interventions. However, bridge maintenance is often delayed due to inadequate planning and budget allocation, as well as resource constraints such as funding. With availability of historical condition data of bridges in databases such as the National Bridge Inventory (NBI) and National Bridge Elements (NBE), there is an opportunity to use data-driven methods to predict deterioration of bridge elements and optimize their maintenance interventions to maximize performance of bridges. This paper presents the development of a novel system that uses Machine Learning (ML) techniques to predict condition of concrete bridge elements and binary linear programming optimization method to identify the optimal selection of maintenance interventions and their timing to maximize the performance of bridges while complying with available annual budgets. Four ML methods are explored: decision tree, random forest, gradient boosting, and support vector machines. The results of the ML evaluation show that, while the values of the predictive performance metrics varied for different elements, random forest method had the best performance for all elements. A case study of a concrete bridge is analyzed to evaluate the performance of the system and demonstrate its new capabilities. The case study results show that the developed model identifies optimal maintenance interventions for various annual budgets over a 50-year study period. The primary contributions of this research to the body of knowledge are: (1) development of a novel system that integrates machine learning techniques and linear programming for predicting bridge element conditions and optimizing maintenance interventions; (2) modeling and predicting the deterioration of bridge elements based on health index metric; and (3) generating long-term maintenance plans for each of bridge elements to maximize the performance of bridges within available annual budgets. The present system is expected to support decision makers, such as highway agencies, in allocating limited financial resources for bridge maintenance more efficiently and cost-effectively.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.