Skip to main content
zenodoopen

A Systematic Evaluation of Large Language Models of Code

<p>These are datasets for the paper:</p> <p>&quot;A Systematic Evaluation of Large Language Models of Code&quot;</p> <p><a href="https://arxiv.org/pdf/2202.13169.pdf">https://arxiv.org/pdf/2202.13169.pdf</a></p> <p>The code is available at:&nbsp;<a href="https://github.com/VHellendoorn/Code-LMs">https://github.com/VHellendoorn/Code-LMs</a></p> <p>&nbsp;</p> <p>The file &quot;<a href="https://zenodo.org/record/6338015/files/unseen_test_sets.tar.gz">unseen_test_sets.tar.gz</a>&quot; contains test sets of ~100 files in each of 12 programming languages.</p> <p>These files are not included in The Pile, and thus models such as GPT-Neo, GPT-J, GPT-NeoX were not trained on them.</p> <p>In the paper, we use these test sets to compare a variety of language models of code including OpenAI&#39;s Codex, GPT-J, GPT-Neo, GPT-NeoX-20B, and CodeParrot and our PolyCoder model.</p> <p>&nbsp;</p> <p>The file &quot;<a href="https://zenodo.org/record/6341643/files/index.zip?download=1">index.zip</a>&quot; includes an index of the&nbsp;<strong>training set</strong>&nbsp;file paths and commit SHAs.</p> <p>&nbsp;</p> <p>The other files, such as &quot;<a href="https://zenodo.org/record/6344914/files/2-7B-150K.tar">2-7B-150K.tar</a>&quot;, are trained model checkpoints, as explained at&nbsp;<a href="https://github.com/VHellendoorn/Code-LMs">https://github.com/VHellendoorn/Code-LMs</a>&nbsp;.</p>

ShareScore

32/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
4
Harmonization
4
Access
16
Reuse readiness
8
Engagement
0