Skip to main content
zenodoopen

Spanish Biomedical Crawled Corpus

<p>The largest Spanish biomedical and heath corpus to date gathered from a massive Spanish health domain crawler over more than 3,000 URLs were downloaded and preprocessed. All the collected data have been preprocessed to produce the CoWeSe (Corpus Web Salud Espa&ntilde;ol) resource, a large-scale and high-quality corpus intended for biomedical and health NLP in Spanish.</p> <p>Enlarged version with less restrictive document and sentence deduplication.</p> <p><strong>Citation</strong></p> <p>If you use this resource in your work, please cite our paper:</p> <pre>@misc{carrino2021spanish, title={Spanish Biomedical Crawled Corpus: A Large, Diverse Dataset for Spanish Biomedical Language Models}, author={Casimiro Pio Carrino and Jordi Armengol-Estap&eacute; and Ona de Gibert Bonet and Asier Guti&eacute;rrez-Fandi&ntilde;o and Aitor Gonzalez-Agirre and Martin Krallinger and Marta Villegas}, year={2021}, eprint={2109.07765}, archivePrefix={arXiv}, primaryClass={cs.CL} } </pre> <p>Copyright (c) 2022 Secretar&iacute;a de Estado de Digitalizaci&oacute;n e Inteligencia Artificial</p> <p>&nbsp;</p>

ShareScore

44/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
8
Harmonization
4
Access
20
Reuse readiness
8
Engagement
4

Topics