Skip to main content
zenodoopen

Early Slavic word embeddings

<p>Word embeddings&nbsp;trained on the lemmatised TOROT Treebank, using Word2Vec and the following parameters:</p> <pre><code>sg = True min_count = &lt;1,3,5&gt; window = &lt;3,5&gt; vector_size = &lt;100,200,300&gt; epochs = 5</code></pre> <p>One model was trained for each combination&nbsp;of the parameters&nbsp;enclosed in angled brackets (&lt; &gt;).&nbsp;</p> <p>The release contains both the full models (.model) and the plain vector files (_vectors.txt). The models are named according to the parameters they were trained with.</p> <p>Note that these are the result of very preliminary experiments and no systematic evaluation of their quality was carried out, so use with caution.</p>

ShareScore

36/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
8
Harmonization
4
Access
16
Reuse readiness
8
Engagement
0

Topics