Experimental Data Set for the study "Exploratory Landscape Analysis is Strongly Sensitive to the Sampling Strategy"
<p>This are the feature values used in the study "Exploratory Landscape Analysis is Strongly Sensitive to the Sampling Strategy".</p> <p>The dataset regroups feature values for every "cheap" features available in the R package <em>flacco </em>and are computed using 5 sampling strategies and in dimension <span class="math-tex">\($d=5$\)</span>:</p> <ol> <li>Random: the classical Mersenne-Twister algorithm;</li> <li>Randu: a random number generator that is notoriously bad;</li> <li>LHS: a centered Latin Hypercube Design;</li> <li>iLHS: an improved Latin Hypercube Design;</li> <li>Sobol: points extracted from a Sobol' low-discrepancy sequence.</li> </ol> <p>The csv file <em>features_summury_dim_5_ppsn.csv </em>regroups 100 values for every features whereas <em>features_summury_dim_5_ppsn_median.csv </em>regroups for every feature the median of the 100 values.</p> <p>In the folder <em>PPSN_feature_plots</em> are the histograms of feature values on the 24 COCO functions for 3 sampling strategies: Random, LHS and Sobol.</p> <p>The Python file <em>sampling_ppsn.py</em> is the code used to generate the sample points from which the feature values are computed.</p> <p>The file <em>stats50_knn_dt.csv</em> provide the raw data of median and IQR (inter quartile interval) for the heatmaps and boxplots available in the paper.</p> <p>Finally, the files <em>results_classif_knn100.csv</em> (resp. dt) provide the accuracy of 100 classifications for every settings.</p> <p> </p>
ShareScore
28/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 4
- Harmonization
- 4
- Access
- 16
- Reuse readiness
- 0
- Engagement
- 4