zenodoopen
Drinov Orthography for Post-OCR Correction dataset
<p>The Drinov Orthography for Post-OCR Correction (DOPOC) dataset was created by annotating a historical newspaper collection provided by the <a href="https://digital.libplovdiv.com/en">National Library "Ivan Vazov"</a> (NLIV) in Plovdiv, Bulgaria. We consider printed versions of these documents, which we manually annotate and align at the character level in the same format as the one from the ICDAR 2019 post-OCR correction competition.</p>
ShareScore
44/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 12
- Harmonization
- 4
- Access
- 16
- Reuse readiness
- 8
- Engagement
- 4