Skip to main content
zenodoopen

Drinov Orthography for Post-OCR Correction dataset

<p>The Drinov Orthography for Post-OCR Correction (DOPOC) dataset was created by annotating a historical newspaper collection provided by the <a href="https://digital.libplovdiv.com/en">National Library "Ivan Vazov"</a> (NLIV) in Plovdiv, Bulgaria. We consider printed versions of these documents, which we manually annotate and align at the character level in the same format as the one from the ICDAR 2019 post-OCR correction competition.</p>

ShareScore

44/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
12
Harmonization
4
Access
16
Reuse readiness
8
Engagement
4

Topics