Skip to main content
zenodoopen

Namgyal Manuscript Collection Datasets

<p>These are the official datasets created for the <em>Tibetan Manuscript Project Vienna</em>&nbsp;(<em>TMPV</em>) in the years 2023 and 2024. These datasets contain:</p> <ul> <li>OCR datasets (line image - line label pairs) created from the PageXML annotations</li> <li>PageXML (Transkribus) annotations in Unicode and Wylie</li> <li>PageXML Layout annotations (lines, images, captions, margins) used for image segmentation training</li> <li>OCR models (PyTorch checkpoints and ONNX model files)</li> </ul>

ShareScore

32/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
4
Harmonization
4
Access
16
Reuse readiness
8
Engagement
0