Exploring Design Smells for Smell-Based Defect Prediction
<p>The archived file datasets.zip includes the datasets used for supporting the conclusions in the article <em>Exploring Design Smells for Smell-Based Defect Prediction.</em></p> <p>In this paper, we answer two research questions:</p> <p><strong>RQ1.</strong> Do Design code smells contribute to the performance of defect prediction models trained with Traditional code smells?</p> <p><strong>RQ2. </strong>How do the different categories of Design smells impact the performance of the defect prediction models?</p> <p>Therefore, after extracting the archived file documents, you will find two sub-directories, respectively named "RQ1" and "RQ2". They include the results obtained for each one of the research questions, thus supporting our conclusions.</p> <p>(You will also find a README.pdf file with these same instructions regarding the datasets.)</p> <p>Inside "RQ1," you will find two directories, respectively named "configuration_1" and "configuration_2". They represent the different configurations for the experiments. <strong>"configuration_1"</strong> contains the datasets with results for the ten classifiers configurations with the highest scores and <strong>"configuration_2" </strong>contains the datasets with the results classifier configuration with the overall best results - Support Vector Machine with C=0.1. Furthermore, within each directory, there are three sub-directories, respectively named "designite," "designite_traditional," and "traditional." These have the datasets for each of the considered smell sets in our study. Inside "RQ2," you will find four directories. Each corresponds to a category from the design smells for the dataset "designite_traditional." These datasets were build from the same configuration as "configuration_2".</p> <p>Then, within every directory, there are 97 sub-directories representing the 97 projects analyzed in this study.</p> <p>Every project folder follows the same structure, which we define as follows.</p> <ul> <li>The "dataset" directory contains the original training and testing dataset used.</li> <li>The "oversamples" directory contains the training dataset after oversampling for each of the feature selection approaches.</li> <li>The "score_summary" directory contains all classifier configurations considered, not only the 10 with the highest scores.</li> <li>The "scores.csv" file contains all the scores for the main classifier configurations studied in the particular experiment.</li> <li>The "selected_features" directory contains the selected features' information and the selected features dataset for each feature_selection method.</li> <li>The "selected_testing_X" directory contains the testing datasets.</li> <li>The "top_scores_summary" directory contains the classifier configurations and hyper-parameter scores for the top 10 highest scores.</li> </ul>
ShareScore
40/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 8
- Harmonization
- 4
- Access
- 16
- Reuse readiness
- 8
- Engagement
- 4