TDX Thesis Spanish Corpus
<p>The TDX Thesis Spanish Corpus is a 246-million-token corpus of Spanish clean text extracted from scientific thesis of the domain <a href="http://tdx.cat">tdx.cat</a>, which contains open thesis published by Catalan universities. The corpus has been preprocessed and deduplicated using the <a href="https://github.com/TeMU-BSC/corpus-cleaner-acl">Corpus-Cleaner</a> pipeline.</p> <p>It consists of 248.676.517 tokens, 8.156.059 sentences and 9.790. Documents are separated by single new lines.</p> <p>We license the actual packaging of these data under a <a href="https://creativecommons.org/licenses/by/4.0/">Attribution 4.0 International License</a>.</p> <p>Copyright by Secretaría de Estado de Digitalización e Inteligencia Artificial (SEDIA) (2022)</p>
ShareScore
40/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 4
- Harmonization
- 4
- Access
- 20
- Reuse readiness
- 8
- Engagement
- 4