Skip to main content
zenodoopen

The Annotated Corpus of Classical Tibetan (ACTib) - Version 2.0 (Segmented & POS-tagged)

<p>This corpus consisting of &gt;185 million tokens is a segmented and part-of-speech tagged version of</p> <p>Wallman, Jeff, Rowinski, Zach, Ngawang Trinley, Tomlinson, Chris, &amp; Keutzer, Kurt. (2017). Collection of Tibetan etexts compiled by the Buddhist Digital Resource Center [Data set]. Zenodo. http://doi.org/10.5281/zenodo.821218</p> <p>using the training data of</p> <p>Hill, Nathan W., &amp; Garrett, Edward. (2017). A part-of-speech (POS) tagged corpus of Classical Tibetan [Data set]. Zenodo. http://doi.org/10.5281/zenodo.574878</p> <p>The code for segmenting and POS tagging any Tibetan file can be found on GitHub.</p> <p>This Version 2 of ACTib is based on the same XML files as ACTib Version 1 (http://doi.org/10.5281/zenodo.823707), but contains both segmented and POS-tagged files and is improved in a number of ways, although post-processing was still done automatically and no manual correction was involved. For details of this improved annotation method see:</p> <p>Meelen, Marieke, Roux, &Eacute;lie &amp; Hill, Nathan (forthcoming). &#39;Optimisation of the largest annotated Tibetan corpus combining rule-based, memory-based &amp; deep-learning methods&#39; in <em>TALLIP.</em></p>

ShareScore

28/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
8
Harmonization
4
Access
16
Reuse readiness
0
Engagement
0

Topics