Skip to main content
zenodoopen

LION Dataset

<p>The presented dataset consists of 198 digitised pages, from eight notepads of the manuscript collection of Swedish author Astrid Lindgren (1907 - 2002), written in the Swedish stenography system <em>Melin</em>. It covers portions of the first six chapters of her novel <em>The Brothers Lionheart</em> (1973), inspiring the name of the dataset, as well as excerpts from three other works by Lindgren.</p> <p>LION is the first of its kind in several regards. Firstly, it is the first dataset containing a portion of Astrid Lindgren&#39;s original drafts and handwriting. Secondly, it is the first to present text, written in the Swedish stenography system Melin. Finally, it is the first publicly available dataset, covering a substantial amount of handwritten lines in any kind of stenographic system.</p> <p>The repository contains page and line images, as well as annotations, in the form of bounding boxes and transliterations. Furthermore, data splittings (train, validation, test) are proposed.</p> <p>&nbsp;</p> <p><strong>This work is licensed under the Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License. Please also see the included LICENSE file or http://creativecommons.org/licenses/by-nc-nd/4.0/ </strong></p>

ShareScore

28/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
8
Harmonization
4
Access
8
Reuse readiness
8
Engagement
0

Topics