Skip to main content
zenodoopen

Spanish Skip-Gram Word Embeddings in FastText

<p>These Spanish word embeddings in FastText have been generated from the largest corpus ever made in Spanish till&nbsp;date. The corpus has more than 2TB of high-quality text,&nbsp;compiled from the different web crawlings done by the National Library of Spain from 2009 to 2019.&nbsp;</p> <p>These are the&nbsp;SKIP-GRAM&nbsp;embeddings, for the CBOW embeddings see:&nbsp;https://zenodo.org/record/5044988</p> <p><strong>Citation</strong></p> <pre><code>@article{gutierrezfandino2022, author = {Asier Gutiérrez-Fandiño and Jordi Armengol-Estapé and Marc Pàmies and Joan Llop-Palao and Joaquin Silveira-Ocampo and Casimiro Pio Carrino and Carme Armentano-Oller and Carlos Rodriguez-Penagos and Aitor Gonzalez-Agirre and Marta Villegas}, title = {MarIA: Spanish Language Models}, journal = {Procesamiento del Lenguaje Natural}, volume = {68}, number = {0}, year = {2022}, issn = {1989-7553}, url = {http://journal.sepln.org/sepln/ojs/ojs/index.php/pln/article/view/6405}, pages = {39--60} }</code></pre> <p><strong>Copyright </strong></p> <p>Copyright (c) 2021 Secretar&iacute;a de Estado de Digitalizaci&oacute;n e Inteligencia Artificial</p>

ShareScore

40/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
8
Harmonization
4
Access
16
Reuse readiness
8
Engagement
4

Topics