Early Slavic word embeddings
<p>Word embeddings trained on the lemmatised TOROT Treebank, using Word2Vec and the following parameters:</p> <pre><code>sg = True min_count = <1,3,5> window = <3,5> vector_size = <100,200,300> epochs = 5</code></pre> <p>One model was trained for each combination of the parameters enclosed in angled brackets (< >). </p> <p>The release contains both the full models (.model) and the plain vector files (_vectors.txt). The models are named according to the parameters they were trained with.</p> <p>Note that these are the result of very preliminary experiments and no systematic evaluation of their quality was carried out, so use with caution.</p>
ShareScore
36/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 8
- Harmonization
- 4
- Access
- 16
- Reuse readiness
- 8
- Engagement
- 0