Italian GloVe models
<p>Italian GloVe models trained from scratch on a dataset composed of:</p> <p>- <strong>wiki</strong>: a dump of Italian Wikipedia (as of December 15, 2022), comprising 25,548,651 sentences and 526,640,982 words (3.2 GB of raw text);<br> - <strong>webz</strong>: a dataset of Italian news (159,226 documents) from the webz.io platform, crawled in October 2015, containing 44,041,823 sentences and 44,544,385 words (244 MB);<br> - a dataset of 5,510 Italian news articles from the newspaper ModenaToday (<strong>MT</strong>) or 15,115 documents from the Italian version of Reuters (<strong>RCV2</strong>).</p> <p><strong>glv_wiki_wbz_mt_20_epochs.zip</strong>: GloVe model trained on the dataset consisting of wiki, webz, and MT for 20 epochs</p> <p><strong>glv_wiki_wbz_mt_50_epochs.zip</strong>: GloVe model trained on the dataset consisting of wiki, webz, and MT for 50 epochs</p> <p><strong>glv_wiki_wbz_reut_20_epochs.zip</strong>: GloVe model trained on the dataset consisting of wiki, webz, and RCV2 for 20 epochs</p> <p><strong>glv_wiki_wbz_reut_50_epochs.zip</strong>: GloVe model trained on the dataset consisting of wiki, webz, and RCV2 for 50 epochs</p>
ShareScore
36/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 8
- Harmonization
- 4
- Access
- 16
- Reuse readiness
- 8
- Engagement
- 0