Skip to main content
zenodoopen

Exploring Design Smells for Smell-Based Defect Prediction

<p>The archived file datasets.zip includes the datasets used for supporting the conclusions in the article <em>Exploring Design Smells for Smell-Based Defect Prediction.</em></p> <p>In this paper, we answer two research questions:</p> <p><strong>RQ1.</strong> Do Design code smells contribute to the performance of defect prediction models trained with Traditional code smells?</p> <p><strong>RQ2. </strong>How do the different categories of Design smells impact the performance of the defect prediction models?</p> <p>Therefore, after extracting the archived file documents, you will find two sub-directories, respectively named &quot;RQ1&quot; and &quot;RQ2&quot;. They include the results obtained for each one of the research questions, thus supporting our conclusions.</p> <p>(You will also find a README.pdf file with these same instructions regarding the datasets.)</p> <p>Inside &quot;RQ1,&quot; you will find two directories, respectively named &quot;configuration_1&quot; and &quot;configuration_2&quot;. They represent the different configurations for the experiments. <strong>&quot;configuration_1&quot;</strong> contains the datasets with results for the ten classifiers configurations with the highest scores and <strong>&quot;configuration_2&quot; </strong>contains the datasets with the results classifier configuration with the overall best results - Support Vector Machine with C=0.1. Furthermore, within each directory, there are three sub-directories, respectively named &quot;designite,&quot; &quot;designite_traditional,&quot; and &quot;traditional.&quot; These have the datasets for each of the considered smell sets in our study. Inside &quot;RQ2,&quot; you will find four directories. Each corresponds to a category from the design smells for the dataset &quot;designite_traditional.&quot; These datasets were build from the same configuration as &quot;configuration_2&quot;.</p> <p>Then, within every directory, there are 97 sub-directories representing the 97 projects analyzed in this study.</p> <p>Every project folder follows the same structure, which we define as follows.</p> <ul> <li>The &quot;dataset&quot; directory contains the original training and testing dataset used.</li> <li>The &quot;oversamples&quot; directory contains the training dataset after oversampling for each of the feature selection approaches.</li> <li>The &quot;score_summary&quot; directory contains all classifier configurations considered, not only the 10 with the highest scores.</li> <li>The &quot;scores.csv&quot; file contains all the scores for the main classifier configurations studied in the particular experiment.</li> <li>The &quot;selected_features&quot; directory contains the selected features&#39; information and the selected features dataset for each feature_selection method.</li> <li>The &quot;selected_testing_X&quot; directory contains the testing datasets.</li> <li>The &quot;top_scores_summary&quot; directory contains the classifier configurations and hyper-parameter scores for the top 10 highest scores.</li> </ul>

ShareScore

40/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
8
Harmonization
4
Access
16
Reuse readiness
8
Engagement
4

Topics