Skip to main content
zenodoopen

Source Code Embeddings

<p>A set of six pretrained fastText models for semantic representations of source code.&nbsp;</p> <p>Each of the models has been&nbsp;trained on high-quality GitHub repositories where the primary language is one of Java, Python, C++, C#, C, PHP. For collecting training data 13.144 repositories were cloned, 2.402.790.348 lines of code were read out of&nbsp;944,467,560&nbsp;files and preprocessed, to finally produce a total of 944.467.560 tokens of clean training data.&nbsp;</p> <p>For further details refer to the following paper:&nbsp;</p> <p>Efstathiou, V., &nbsp;Spinellis, D., 2019. &quot;Semantic Source Code Models Using Identifier Embeddings&quot;. In <em>16th International Conference on Mining Software Repositories:&nbsp;Data Showcase Track. MSR&#39;19.&nbsp;</em></p>

ShareScore

40/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
8
Harmonization
4
Access
16
Reuse readiness
8
Engagement
4

Topics