Skip to main content
zenodoopen

TDX Thesis Spanish Corpus

<p>The TDX Thesis Spanish Corpus is a 246-million-token corpus of Spanish clean text extracted from scientific thesis of the domain <a href="http://tdx.cat">tdx.cat</a>, which contains open thesis published by Catalan universities. The corpus has been preprocessed and deduplicated using the <a href="https://github.com/TeMU-BSC/corpus-cleaner-acl">Corpus-Cleaner</a> pipeline.</p> <p>It consists of 248.676.517 tokens, 8.156.059 sentences and 9.790. Documents are separated by single new lines.</p> <p>We license the actual packaging of these data under a <a href="https://creativecommons.org/licenses/by/4.0/">Attribution 4.0 International License</a>.</p> <p>Copyright by Secretar&iacute;a de Estado de Digitalizaci&oacute;n e Inteligencia Artificial (SEDIA) (2022)</p>

ShareScore

40/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
4
Harmonization
4
Access
20
Reuse readiness
8
Engagement
4