Skip to main content
zenodoopen

German CBOW FastText embeddings with min count 250

<p>FastText embeddings built from Common Crawl german dataset</p> <table> <caption>Parameters</caption> <thead> <tr> <th scope="col">Parameters</th> <th scope="col">Value(s)</th> </tr> </thead> <tbody> <tr> <td>Dimensions</td> <td>256 and 384</td> </tr> <tr> <td>Context window</td> <td>5</td> </tr> <tr> <td>Negative sampled</td> <td>10</td> </tr> <tr> <td>Epochs</td> <td>1</td> </tr> <tr> <td>Number of buckets</td> <td>131072 or 262144</td> </tr> <tr> <td>Min n</td> <td>3</td> </tr> <tr> <td>Max n</td> <td>6</td> </tr> </tbody> </table>

ShareScore

44/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
8
Harmonization
8
Access
16
Reuse readiness
8
Engagement
4

Topics