Skip to main content
zenodoopen

Process Behavior Corpus and Benchmarking Datasets

<p>A corpus of process behaviors and benchmarking datasets for semantics-aware process mining tasks.</p> <p>Files:</p> <ul> <li><strong>process_behavior_corpus.csv</strong>: the text corpus, which contains the behavior allowed by process models as sequences of activities (column <em>string_traces)</em>.</li> <li><strong>T_SAD.csv:</strong> A benchmark dataset generated from the corpus to assess the following task: Given a trace &sigma;, decide if &sigma; is a valid execution of the underlying process or not, without knowing the behavior allowed in the process.<br>Each row contains a trace (column <em>trace</em>) with a corresponding label (column&nbsp;<em>anomalous</em>) indicating whether the trace represents a valid execution of the underlying process. The set of activities that can occur in the process are also given (column&nbsp;<em>unique_activities</em>).</li> <li><strong>A_SAD.csv:</strong> A benchmark dataset generated from the corpus to assess the following task: Given an eventually-follows relation ef = a ≺ b of<br>a trace &sigma;, decide if ef represents a valid execution order of the two activities a and b that are executed in a process&nbsp;or not, without knowing the behavior allowed in the process.<br>Each row contains an eventually-follows relation (column <em>eventually_follows</em>) with a corresponding label (column&nbsp;<em>out_of_order</em>) indicating wether the two activities of the relation were executed in an invalid order (TRUE) or in a valid order (FALSE) according to the underlying process (model). The set of activities that can occur in the process are also given (column <em>unique_activities</em>).</li> <li><strong>S_NAP.csv:</strong> A benchmark dataset generated from the corpus to assess the following task: Given an event log L and a prefix p_k of length k, with 1 &lt; k, predict the next activity a_k+1<br>Each row contains a trace prefix (column <em>prefix</em>) with a corresponding next activity (column <em>next</em>) indicating the activity that should be performed next after the last activity of the prefix &nbsp;according to the trace from which the prefix was generated. The set of activities that can occur in the process are also given (column <em>unique_activities</em>).</li> <li><strong>S-PMD.csv:</strong> A benchmark dataset generated from the corpus to assess the following tasks: <ul> <li>Given a set of possible activities&nbsp;(column&nbsp;<em>unique_activities</em>), generate a difectly follows graph (column <em>dfg</em>) that captures the trace semantics of the process model.&nbsp;</li> <li>Given a set of possible activities (column&nbsp;<em>unique_activities</em>), generate a simple process tree (column&nbsp;<em>pt</em>)&nbsp;that captures the trace semantics of the process model.</li> </ul> </li> </ul> <p>Reference and legal info:</p> <p>The corpus and the benchmark datasets are generated using the SAP-SAM dataset:</p> <p>Kampik, T., Warmuth, C., Sola, D., Sch&auml;fer, B., Axworthy, L., Ivarsson, E., Ouda, K., &amp; Eickhoff, D. (2022). SAP Signavio Academic Models (0.5.1) [Data set]. Zenodo.&nbsp;<a href="https://doi.org/10.5281/zenodo.7012043" target="_blank" rel="noopener">https://doi.org/10.5281/zenodo.7012043</a><br><br>The SAP-SAM dataset is published with a specific license (see "Rights"), which, therefore, also applies to the data published in this record.</p> <p><strong>THE DATASETS AND ASSOCIATED EVALUATION EXPERIMENTS ARE DESCRIBED IN <a href="https://arxiv.org/pdf/2407.02310">THIS</a> PAPER.</strong></p> <p><strong>IN&nbsp;<a href="https://github.com/a-rebmann/llms4pm">THIS</a> REPOSITORY YOU FIND THE CODE AND RAW RESULTS OF EVALUATION EXPERIMENTS USING VARIOUS OPEN SOUCE LLMs TO SOLVE THE TASKS</strong></p>

ShareScore

28/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
4
Harmonization
4
Access
12
Reuse readiness
8
Engagement
0