Skip to main content
zenodoopen

Greetings From! Historical Postcards Address Transcription Dataset

<p>This dataset provides both Ground Truth (GT) and Handwritten Text Recognition (HTR) transcriptions of historical postcard addresses, stemming from a project to extract address information from historical picture postcards from Belgium, France, Germany, Luxembourg, the Netherlands, and the UK. The dataset encapsulates the back of 500 historically significant postcards.</p><p>The research associated with this dataset will be presented at <strong>Computational Humanities Research Conference, December 6--8, 2023, Paris, France.</strong></p><p><strong>Scope and Content:</strong></p><ul><li><strong>HTR Material</strong>: Handwritten Text Recognition outputs for 500 postcards.</li><li><strong>GT Material</strong>: Ground Truth transcriptions created by human transcribers for the same set of 500 postcards.</li></ul><p><strong>File Structure and Formats:</strong></p><p><i>For both HTR and GT Material, the following files are provided:</i></p><ul><li><strong>JPEG Images</strong>: Scanned or digitized images of the postcards.</li><li><strong>.txt</strong>: Plain text transcriptions of the postcards.</li><li><strong>_tei.xml</strong>: Transcriptions rendered in the TEI XML format.</li><li><strong>.pdf</strong>: PDF presentation of the postcards along with their transcriptions.</li><li><strong>mets.xml</strong>: METS (Metadata Encoding and Transmission Standard) schema for the data.</li><li><strong>page</strong> folder: XML files for individual images, offering metadata and structural information.</li><li><strong>metadata.xml</strong>: metadata concerning the dataset.</li><li><strong>GT_addresses_GPT4.json &amp; HTR_addresses_GPT4.json</strong>: JSON files detailing individual address data for each postcard in structured format.</li></ul><p><strong>Annotation and Transcription:</strong></p><ul><li><strong>GT</strong>: Ground Truth data was annotated by human transcribers who examined both the images of the postcards and the outputs of the HTR system. Transcribers made corrections according to predefined conventions: using <strong>#</strong> for illegible characters, <strong>*</strong> at the start of lines without address information (e.g., person's name), and starting a line with <strong>@</strong> for irrelevant lines.</li><li><strong>HTR</strong>: The HTR versions emerged from state-of-the-art HTR systems (<a href="https://readcoop.eu/introducing-transkribus-super-models-get-access-to-the-text-titan-i/">Transkribus Text Titan I</a>). The <strong>.json</strong> files hold precise address details derived from the main data, which were processed using OpenAI's GPT-4 Large Language Model.</li></ul>

ShareScore

40/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
4
Harmonization
4
Access
20
Reuse readiness
8
Engagement
4