Process Behavior Corpus and Benchmarking Datasets
<p>A corpus of process behaviors and benchmarking datasets for semantics-aware process mining tasks.</p> <p>Files:</p> <ul> <li><strong>process_behavior_corpus.csv</strong>: the text corpus, which contains the behavior allowed by process models as sequences of activities (column <em>string_traces)</em>.</li> <li><strong>T_SAD.csv:</strong> A benchmark dataset generated from the corpus to assess the following task: Given a trace σ, decide if σ is a valid execution of the underlying process or not, without knowing the behavior allowed in the process.<br>Each row contains a trace (column <em>trace</em>) with a corresponding label (column <em>anomalous</em>) indicating whether the trace represents a valid execution of the underlying process. The set of activities that can occur in the process are also given (column <em>unique_activities</em>).</li> <li><strong>A_SAD.csv:</strong> A benchmark dataset generated from the corpus to assess the following task: Given an eventually-follows relation ef = a ≺ b of<br>a trace σ, decide if ef represents a valid execution order of the two activities a and b that are executed in a process or not, without knowing the behavior allowed in the process.<br>Each row contains an eventually-follows relation (column <em>eventually_follows</em>) with a corresponding label (column <em>out_of_order</em>) indicating wether the two activities of the relation were executed in an invalid order (TRUE) or in a valid order (FALSE) according to the underlying process (model). The set of activities that can occur in the process are also given (column <em>unique_activities</em>).</li> <li><strong>S_NAP.csv:</strong> A benchmark dataset generated from the corpus to assess the following task: Given an event log L and a prefix p_k of length k, with 1 < k, predict the next activity a_k+1<br>Each row contains a trace prefix (column <em>prefix</em>) with a corresponding next activity (column <em>next</em>) indicating the activity that should be performed next after the last activity of the prefix according to the trace from which the prefix was generated. The set of activities that can occur in the process are also given (column <em>unique_activities</em>).</li> <li><strong>S-PMD.csv:</strong> A benchmark dataset generated from the corpus to assess the following tasks: <ul> <li>Given a set of possible activities (column <em>unique_activities</em>), generate a difectly follows graph (column <em>dfg</em>) that captures the trace semantics of the process model. </li> <li>Given a set of possible activities (column <em>unique_activities</em>), generate a simple process tree (column <em>pt</em>) that captures the trace semantics of the process model.</li> </ul> </li> </ul> <p>Reference and legal info:</p> <p>The corpus and the benchmark datasets are generated using the SAP-SAM dataset:</p> <p>Kampik, T., Warmuth, C., Sola, D., Schäfer, B., Axworthy, L., Ivarsson, E., Ouda, K., & Eickhoff, D. (2022). SAP Signavio Academic Models (0.5.1) [Data set]. Zenodo. <a href="https://doi.org/10.5281/zenodo.7012043" target="_blank" rel="noopener">https://doi.org/10.5281/zenodo.7012043</a><br><br>The SAP-SAM dataset is published with a specific license (see "Rights"), which, therefore, also applies to the data published in this record.</p> <p><strong>THE DATASETS AND ASSOCIATED EVALUATION EXPERIMENTS ARE DESCRIBED IN <a href="https://arxiv.org/pdf/2407.02310">THIS</a> PAPER.</strong></p> <p><strong>IN <a href="https://github.com/a-rebmann/llms4pm">THIS</a> REPOSITORY YOU FIND THE CODE AND RAW RESULTS OF EVALUATION EXPERIMENTS USING VARIOUS OPEN SOUCE LLMs TO SOLVE THE TASKS</strong></p>
ShareScore
28/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 4
- Harmonization
- 4
- Access
- 12
- Reuse readiness
- 8
- Engagement
- 0