Skip to main content
zenodoopen

TexBiG Dataset for Analysing Complex Document Layouts in the Digital Humanities

<p>This is the dataset for the paper&nbsp;&quot;A Dataset for Analysing Complex Document Layouts in the Digital Humanities and its Evaluation with Krippendorff &rsquo;s Alpha&quot; in its second version, containing an update of the test images (without annotations) from the paper &quot;Drawing the Same Bounding Box Twice? Coping Noisy Annotations in Object Detection with Repeated Labels&quot;. Organization of the dataset is also updated to make it easier to use.</p> <p>TexBiG (from the German Text-Bild-Gef&uuml;ge, meaning Text-Image-Structure) is a document layout analysis dataset for historical documents in the late 19th and early 20th century. The dataset provides instance segmentation (bounding boxes and polygons/masks) annotations for 19 different classes with more then 52.000 instances. The added test images can be used to make submission on the leaderboard on <a href="https://eval.ai/web/challenges/challenge-page/2078/overview">EvalAI</a>.&nbsp;</p> <p>The <a href="https://zenodo.org/record/6885144/files/Annotations_Guideline.pdf?download=1">annotation guideline</a> can be found in the first of the dataset.</p>

ShareScore

40/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
8
Harmonization
4
Access
16
Reuse readiness
8
Engagement
4

Topics