Skip to main content
zenodoopen

DIRE: A Neural Approach to Decompiled Identifier Naming

<p>This dataset is released as a companion to the paper &quot;DIRE: A Neural Approach to Decompiled Identifier Naming&quot;, appearing in the proceedings of the&nbsp;34th IEEE/ACM International Conference on Automated Software Engineering (ASE 2019).</p> <p>It contains information generated by decompiling 3,195,962 functions found in 164,632 unique binaries generated from C code scraped from GitHub. For practicality, the dataset is partitioned into 16 archives by the first hexadecimal digit of the SHA-256 hash of the binary used to generate it. Each of the 16 archives contains approximately 10,000&nbsp;JSONL files, named according to a binary&#39;s hash. Each JSONL file consists of a single JSON object per-line corresponding to a single function in the decompiled binary.</p> <p>Archives are provided in both GZIP and BZIP2 format.</p> <p>See the README file for more information.</p>

ShareScore

32/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
4
Harmonization
4
Access
16
Reuse readiness
8
Engagement
0