Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
12
datasets available to search
ShareScore release 0.7.1
Dataset results
12 results for “bpm”
WONDERBREAD: A Benchmark + Dataset for Business Process Management (BPM) Tasks
<p><strong>Paper:</strong> <a href="https://arxiv.org/abs/2406.13264">WONDERBREAD: A Benchmark for Evaluating Multimodal Foundation Models on Business Process Management Tasks</a></p> <h2><strong>Background</strong></h2> <p>The <em>WONDERBREAD</em> dataset contains <strong>2,928 human demonstrations</strong> of <strong>598 web navigation workflows</strong> across <strong>6 types of BPM tasks</strong>. These tasks measure the ability of a model to generate accurate documentation, assist in knowledge transfer, and improve the efficiency of workflows.</p> <p>Please see our website for more details: <a href="https://wonderbread.stanford.edu/">https://wonderbread.stanford.edu/</a></p> <h2><strong>Quick Start</strong></h2> <p>To start, download <strong>debug_demos.zip</strong> (1 GB). It contains a subset of <strong>24 demonstrations</strong> which can give you a sense of how the dataset is structured.</p> <p>To reproduce the paper, download <strong>gold_demos.zip</strong> (33 GB). It contains <strong>724 demonstrations</strong> corresponding to the 162 "Gold" tasks which were used for all the evaluations in the original paper.</p> <p>To obtain the full dataset, download <strong>demos.zip</strong> (133 GB). This contains all <strong>2,928 demonstrations</strong> and can be used for training, fine-tuning, and evaluating models.</p> <h2><strong>Dataset Structure</strong></h2> <p>The dataset contains several files, defined below.</p> <ol> <li><strong>Raw Data</strong><em> (useful for training/fine-tuning/evaluation)</em> <ol> <li><strong>debug_demos.zip </strong>(1 GB)<strong> -- </strong>a subset of only 24 demonstrations taken from the full dataset. Useful to get a sense of the dataset and for debugging.</li> <li><strong>gold_demos.zip</strong> (22 GB) -- a subset of only 724 demonstrations corresopnding to the 162 "Gold" tasks. This is the dataset that was used for all evaluations in the original <em>WONDERBREAD</em> paper.</li> <li><strong>demos.zip</strong> (133 GB) -- all 2,928 demonstrations across 598 tasks. Useful for training your own models.</li> </ol> </li> <li><strong>Modality-Specific Subsets of Raw Data </strong><em>(useful for specific types of training/fine-tuning/evaluation)</em><br> <ol> <li><strong>All Demos</strong> <ol> <li><strong>demos_sop_only.zip</strong> (4 MB)-- only the SOP <code>.txt</code> files for all 2,928 demonstrations</li> <li><strong>demos_sop_and_trace_only.zip</strong> (770 MB)-- only the SOP <code>.txt</code> files and action trace <code>.json</code> files for all 2,928 demonstrations</li> <li><strong>demos_sop_and_trace_and_screenshots_only.zip</strong> (22 GB)-- only the SOP <code>.txt</code> files and action trace <code>.json</code> files and screenshot images for all 2,928 demonstrations</li> </ol> </li> <li><strong>"Gold" Demos</strong> <ol> <li><strong>gold_demos_sop_only.zip</strong> (1 MB)-- only the SOP <code>.txt</code> files for the 724 demonstrations in the "Gold" tasks.</li> <li><strong>gold_demos_sop_and_trace_only.zip</strong> (190 MB) -- only the SOP <code>.txt</code> files and action trace <code>.json</code> files for the 724 demonstrations in the "Gold" tasks.</li> <li><strong>gold_demos_sop_and_trace_and_screenshots_only.zip </strong>(6 GB) -- only the SOP <code>.txt</code> files and action trace <code>.json</code> files and screenshot images for the 724 demonstrations in the "Gold" tasks</li> </ol> </li> </ol> </li> <li><strong>Evaluation</strong><em> (useful for evaluation)</em> <ol> <li><strong>qa_dataset.csv -- </strong>contains all 120 questions and ground truth answers used in the "Knowlege Transfer" evaluation.<strong><br></strong></li> <li><strong>df_rankings.csv -- </strong>contains the rankings of all "Gold" tasks used in the "SOP Ranking" evaluation.<strong><br></strong></li> </ol> </li> <li><strong>Metadata</strong><em> (can be safely ignored)</em> <ol> <li><strong>Process Mining Task Demonstrations.xlsx --</strong> maps human annotators to specific demonstrations; also contains "Gold" task rankings used in the "SOP Ranking" evaluation.</li> <li><strong>metadata.json -- </strong>maps Google Drive URLs to Google Drive Folder IDs to demonstration names</li> <li><strong>df_valid.csv -- </strong>tracks assets associated with each demonstration</li> </ol> </li> </ol>
WP4 - Traitement données - BPM - simulation réalité virtuelle
<p>Ces données ont été collectés dans le cadre du projet de recherche PAsCAL entre février et avril 2021. Il s'agit des résultats bruts de fréquence cardiaque (battements par minutes; BPM) pour les 11 participants en situation de handicap moteur qui ont réalisé l'expérience WP4 sur la plateforme de réalité virtuelle de UBFC.</p>
BPM_ethics
<p>Dataset containing the explicit occurrences of terms associated with ethical or moral aspects in the publications of the BPM, ICPM and S-BPM ONE conferences, from their first editions to the 2022 editions</p>
BPM Synthetic UI Logs Collection
<p>This data package described in the <em>BPM Demos&Resources</em> publication entitled: "<em>BPM Hub: An Open Collection of UI Logs</em>", consists of synthetic UI logs along with corresponding screenshots. The UI logs closely resemble real-world use cases within the administrative domain. They exhibit varying levels of complexity, measured by the number of activities, process variants, and visual features that influence the outcome of decision points. For its generation, the <a href="https://canela.lsi.us.es/bpmloggenerator/">BPM Log Generator tool</a> has been used, which requires the following initial generation configuration:</p> <p><strong>Initial Generation Configuration</strong></p> <ul> <li>Seed log: Includes a single instance for each process variant and their associated screenshots.</li> <li>Variability configuration: <ul> <li>Case-level: Refers to variations in the content that can be introduced or modified by the user, such as variations in the text inputs, selectable options, checkboxes, etc.</li> <li>Scenario-level: Refers to varying the GUI (Graphical User Interface) components related to the look and feel of the different applications appearing in the process screenshots.</li> </ul> </li> </ul> <p><strong>Data Package Contents</strong></p> <p>The data package comprises three distinct processes, P1, P2, P3, for which their initial configuration is provided, i.e., a tuple of <SeedLog, Case-level variability conf., Scenario-level variability conf.>. They are characterized by the following:</p> <p>P1. Client Creation</p> <ul> <li>Activities: 5</li> <li>Variants: 2</li> <li>Decision point: Revolves around the presence of an attachment in the reception of an email.</li> </ul> <p>P2. Client Deletion. User's presence in the system</p> <ul> <li>Activities: 7</li> <li>Variants: 2</li> <li>Decision point: Based on the result of the user's search in the Customer Management System (CRM), represented by a checkbox.</li> </ul> <p>P3. Client Deletion. Validation of customer payments</p> <ul> <li>Activities: 7</li> <li>Variants: 4</li> <li>Decision: Involves two conditions: <ol> <li>The presence of an attachment justifying the payment of the invoices in the email.</li> <li>The existence of pending invoices in the user CRM profile.</li> </ol> </li> </ul> <p>These problems depict processes with a single decision point, without cycles, and executed sequentially to ensure a non-interleaved execution pattern. Particularly, P3 shows higher complexity as its decision point is determined by two visual characteristics.</p> <p><strong>Generation of UI Logs</strong></p> <p>For each problem, case-level variations have been applied to generate logs with different sizes in the range of {10, 25, 50, 100} events. In cases where the log exceeds the desired size, the last instance is removed to maintain completeness. Each log size has its associated balanced and unbalanced log. Balanced logs have an approximately equal distribution of instances across variants, while unbalanced logs have a frequency difference of more than 20% between the most frequent and least frequent variants.</p> <p><strong>Scenarios</strong></p> <p>To ensure the reliability of the obtained results, 30 scenarios are generated for each tuple <Problem, LogSize, Balanced?>. These scenarios exhibit slight variations at the scenario-level, particularly in the look and feel and user interface of the applications depicted in the screenshots. Each scenario consists of UI logs that correspond to specific problems categorized by log size (10, 25, 50, 100) and balanced? (Balanced, Unbalanced). Folders containing UI logs and their corresponding screenshots are organized in folders named as follows: sc{scenarioId}_size_{LogSize}_{Balanced?}.</p> <p><strong>Additional Artefacts</strong></p> <p>In addition, each problem includes two more artefacts:</p> <ul> <li>initial_generation_configuration folder: Holds the data needed for problem data generation using the [5] tool.</li> <li>decision.json file: Specifies the condition driving the decision made at the decision point.</li> </ul> <p><strong>decision.json</strong></p> <p>The decision.json acts as a testing oracle, serving as a label for validating mined data. It contains two main sections: "UICompos" and "decision". The "UICompos" section includes a key for each activity related to the decision, storing key-value pairs that represent the UI components involved, along with their bounding box coordinates. The "decision" section defines the condition for a case to match a specific variant based on the mentioned UI components.</p>
BPM 14, feature 574, selection
Fragment naczynia sitowatego z obiektu 574 na stanowisku Bożepole Małe 14, woj. pomorskie badanego przez archeobaltica.pl. Potsherd (I don't know English name for this type) from excavations on Bożepole Małe 14 archaeological site located in Pomerania, northern Poland. Source: Objaverse 1.0 / Sketchfab
A Characterization of Ambiguity in BPM
<p>Dataset and results, and participants demographics for the evaluation of the paper "A Characterization of Ambiguity in BPM".</p> <p>File "Evaluation_tasks_anonymized.pdf": dataset and tasks used for the evaluation.</p> <p>File "Evaluation_results_anonymized.pdf": results of the evaluation.</p> <p>File "Participant_demographics_anonymized.pdf": demographics of the participants of the evaluation.</p>
Systematic Literature Review on Meta-Models in BPM: Material
<p>Material related to the Systematic Literature Review on Meta-Models in Business Process Management</p>
Bottom-Up BPM methods
Open the record for dataset details and reuse information.
Validation Study of WITHINGS BPM Core for the Detection of Atrial Fibrillation
ClinicalTrials.gov study NCT04464499. IPD Sharing: NO. Countries: 1. Publications: 0.
Clinical Validation of the Blood Pressure Measuring Device Withings BPM Pro 2 (WIHYP-GP)
ClinicalTrials.gov study NCT06957847. IPD Sharing: NO. Countries: 1. Publications: 0.
BPM Cuff: 15-24cm, 20-34cm, 30-44cm, 40-48cm, 22-42cm
ClinicalTrials.gov study NCT02021994. IPD Sharing: Not stated. Countries: 0. Publications: 0.
Transcriptomic profile of amiR-bpm and Col-0
GEO Series GSE131037. Arabidopsis thaliana. 3 samples. Type: Expression profiling by array.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.