Skip to main content
zenodoopen

The dataset of the paper titled "Context-Aware Code Change Embedding for Better Patch Correctness Assessment"

<p>The dataset of the paper titled &quot;Context-Aware Code Change Embedding for Better Patch Correctness Assessment&quot;.</p> <p>This is the online repository of the paper &quot;Context-Aware Code Change Embedding for Better Patch Correctness Assessment&quot; under review by ASE2021. We release the source code of Cache, the patches used in our evaluation, as well as the experiment results.</p> <ul> <li> <p>Patches: Two&nbsp;patch benchmarks included in our study.</p> <ul> <li>Small: The 1,183 deduplicated patches from Tian&#39;s <a href="https://conf.researchr.org/details/ase-2020/ase-2020-papers/3/Evaluating-Representation-Learning-of-Code-Changes-for-Predicting-Patch-Correctness-i">ASE20</a> paper and Wang&#39;s <a href="https://conf.researchr.org/details/ase-2020/ase-2020-papers/54/Automated-Patch-Correctness-Assessment-How-Far-are-We-">ASE20</a> paper.</li> <li>Large:&nbsp;The patches collected by ourselves, which is consist of totally 49,694 patches from&nbsp;<a href="https://dl.acm.org/doi/10.1145/3338906.3338911">RepairThemAll</a>&nbsp;and&nbsp;<a href="https://dl.acm.org/doi/10.1145/3379597.3387491">ManySStuBs</a>.</li> </ul> </li> <li> <p>Results</p> <ul> <li> <p>RQ1: The detailed result files in RQ1, which are named by the format of <em><code>[model]_[classifier].csv</code></em>. For example, the file named <em><code>BERT_DT.csv</code></em> in the folder <em><code>Small</code></em> means that this file is the result of patches from <strong>Small</strong> dataset&nbsp;embedded by <strong>BERT</strong> and classified by <strong>Decision Tree</strong>.</p> <ul> <li>Small: The detailed result files on Small&nbsp;dataset.</li> <li>Large: The detailed result files on Large&nbsp;dataset.</li> <li> <p>Cross: The detailed result files of representation learning techniques when training on Large&nbsp;dataset and testing on Small&nbsp;dataset.</p> </li> </ul> </li> <li> <p>RQ2: The detailed result files in RQ2.</p> <ul> <li>Wang_Cache.csv: The detailed result of Cache on the dataset from Wang&#39;s <a href="https://conf.researchr.org/details/ase-2020/ase-2020-papers/54/Automated-Patch-Correctness-Assessment-How-Far-are-We-">ASE20</a> paper.</li> <li> <p>ODS_Cache.csv: The datailed result of Cache on the dataset from Xiong&#39;s <a href="https://dl.acm.org/doi/10.1145/3180155.3180182">ICSE18</a> paper. We directly compare against the results reported by the authors of ODS on 139 patches from Xiong&#39;s paper since the data and source code of ODS is unavailable.</p> </li> <li> <p>Table_5_Effectiveness_APCA.xlsx: The detailed version of&nbsp;<strong>Table 5</strong>&nbsp;in the paper.</p> </li> <li> <p>Table_6_Effectiveness_ODS.xlsx: The detailed version of&nbsp;<strong>Table 6</strong>&nbsp;in the paper.</p> </li> </ul> </li> </ul> </li> <li>Source: The source code and lib for running Cache. Guidance for replicating our study is available at <em><code>source/Readme.md</code></em>. We will build a homepage for Cache on GitHub upon acceptance.</li> </ul>

ShareScore

32/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
4
Harmonization
4
Access
16
Reuse readiness
8
Engagement
0