Skip to main content
zenodoopen

Italian GloVe models

<p>Italian GloVe models trained from scratch&nbsp;on a dataset composed of:</p> <p>- <strong>wiki</strong>: a&nbsp;dump of Italian Wikipedia (as of December 15, 2022), comprising 25,548,651 sentences and 526,640,982 words (3.2 GB of raw text);<br> - <strong>webz</strong>: a&nbsp;dataset of Italian news (159,226 documents) from the webz.io platform, crawled in October 2015, containing 44,041,823 sentences and 44,544,385 words (244 MB);<br> - a dataset of 5,510 Italian news articles from the newspaper ModenaToday&nbsp;(<strong>MT</strong>) or 15,115 documents from the Italian version of Reuters (<strong>RCV2</strong>).</p> <p><strong>glv_wiki_wbz_mt_20_epochs.zip</strong>: GloVe model trained on the&nbsp;dataset consisting of wiki, webz, and MT for 20 epochs</p> <p><strong>glv_wiki_wbz_mt_50_epochs.zip</strong>: GloVe model trained on the&nbsp;dataset consisting of wiki, webz, and MT for 50 epochs</p> <p><strong>glv_wiki_wbz_reut_20_epochs.zip</strong>: GloVe model trained on the&nbsp;dataset consisting of wiki, webz, and RCV2 for 20 epochs</p> <p><strong>glv_wiki_wbz_reut_50_epochs.zip</strong>: GloVe model trained on the&nbsp;dataset consisting of wiki, webz, and RCV2 for 50 epochs</p>

ShareScore

36/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
8
Harmonization
4
Access
16
Reuse readiness
8
Engagement
0

Topics