Synthetic Datasets from the Article titled Privacy-preserving Ground-truth Data for Evaluating Additive Feature Attribution in Regression Models with Additive CBR and CQV
<p>Synthetic datasets were generated as benchmarks capturing the intrinsic characteristics of original data to investigate the performance of additive feature attribution methods for regression tasks. The synthetic datasets were generated based on 2, 6 and 8 clusters formed with the original data. The 6-cluster dataset was used for primary analysis and the other two were used for sensitivity analysis.</p><p>The synthetic dataset was generated from the original data acquired from <a href="https://www.eurocontrol.int/dashboard/rnd-data-archive">Aviation Data for Research Repository</a>, which was collected and processed by <a href="https://www.eurocontrol.int/">EUROCONTROL</a> from the Enhanced Tactical Flow Management System (ETFMS) flight data messages containing all flights in Europe throughout the year 2019, from May to October. The original dataset consisted of fundamental details of the flights, flight status, preceding flight legs, ATFM regulations, weather conditions, calendar information, etc. </p><p>A brief description of the columns in the synthetic data files is presented in the file 'data_description.pdf' and a more detailed discussion on features can be found in the works of Koolen and Coliban [1] and Dalmau et al. [2].</p><p> </p><p><strong>References</strong><br>[1] H. Koolen and I. Coliban, <a href="https://www.eurocontrol.int/sites/default/files/2020-06/flight-progress-msg-update-230620.pdf">Flight Progress Messages Document</a>, EUROCONTROL, Brussels, Belgium, Tech. Rep., 2020.<br>[2] R. Dalmau, F. Ballerini, H. Naessens, S. Belkoura, and S. Wangnick, <a href="https://www.sciencedirect.com/science/article/pii/S0969699721000739">An Explainable Machine Learning Approach to Improve Take-off Time Predictions</a>, Journal of Air Transport Management, vol. 95, p. 102 090, Aug. 2021. doi: 10.1016/j.jairtraman.2021.102090.</p><p><br> </p>
ShareScore
20/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 8
- Harmonization
- 4
- Access
- 8
- Reuse readiness
- 0
- Engagement
- 0