Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

20

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

20 results for “Handwritten Text Recognition”

Learn how ShareScore rates datasets ↗
zenodo44/100

Transkribus - Handwritten Text Recognition for Premodern Documents (SIMS 2020 Lightning Talk)

<p>Transkribus is a platform for text recognition and can be used via the Transkribus Expert Software (available after registration: transkribus.eu). Through Transkribus different tools for document analysis and text recognition can be directly applied. The intro demonstrates very briefly how Transkribus can help with regards to premodern documents especially since a variety of pre-trained models are already available: for Latin (prints and handwriting), for early modern vernaculars in French, Dutch, English, and German. For more information go to transkribus.eu.</p> <p>Presented as a Schoenberg Symposium 2020 Lightning Talk</p>

opencc-by-4.0Nov 2020View details →
zenodo44/100

Ground-Truthed Data Set of Zenon Papyri for Handwritten Text Recognition

<p>Diplomatic transcription of papyri found in the Zenon archive [see <a href="https://en.wikipedia.org/wiki/Zenon_of_Kaunos">en.wikipedia.org/wiki/Zenon_of_Kaunos</a>]</p> <p>Manually prepared as PageXML with Transkribus within <a href="http://d-scribes.philhist.unibas.ch/">D-Scribes</a> project.</p> <p>&nbsp;</p>

opencc-by-4.0May 2022View details →
zenodo44/100

A Brazilian Portuguese Dataset for Offline Handwritten Text Recognition (BRESSAY)

<p>The BRESSAY dataset comprises images of handwritten essays in Brazilian Portuguese, which present a series of challenges to optical recognition models. These images were sourced from multiple online platforms, limiting our ability to standardize the capture process. Due to these varied sources and the lack of a uniform collection method, the dataset provides a realistic reflection of real-world conditions. Each essay is unique, contributed by different writers, and addresses a specific content topic. Furthermore, the constraints placed on the writers often lead to various handwriting scenarios, including hard-to-read words, connected words, noise, overwriting, and struck-through texts.</p> <h3><strong>Technical Details</strong></h3> <p>The BRESSAY dataset represents a comprehensive collection of handwritten essays in Brazilian Portuguese, offering detailed insights into various handwriting scenarios. It covers a total of 1,000 pages, each contributed by a unique writer, resulting in 1,000 distinct handwriting styles. This aspect of the dataset adds a layer of diversity, which is further emphasized by the total of 4,214 paragraphs, 30,090 lines, and 416,826 words. Regarding unique tokens, we have 41,318 unique words, and 107 unique characters.</p> <h3><strong>Data Structure</strong></h3> <p>The dataset is organized as follows:</p> <ul> <li>data/: Main folder containing segmented essay images <ul> <li>lines/: Images of individual lines <ul> <li>&nbsp; PNG files: Line images</li> <li>&nbsp; TXT files: Transcriptions of lines</li> </ul> </li> <li>pages/: Full page essay images <ul> <li>&nbsp; PNG files: Page images</li> <li>&nbsp; TXT files: Transcriptions of pages</li> </ul> </li> <li>paragraphs/: Images of paragraphs <ul> <li>&nbsp; PNG files: Paragraph images</li> <li>&nbsp; TXT files: Transcriptions of paragraphs</li> </ul> </li> <li>words/: Images of individual words <ul> <li>&nbsp; PNG files: Word images</li> <li>&nbsp; TXT files: Transcriptions of words</li> </ul> </li> </ul> </li> <li>sets/: Contains partition files <ul> <li>test.txt: Names of images in the test set</li> <li>validation.txt: Names of images in the validation set</li> <li>training.txt: Names of images in the training set</li> </ul> </li> </ul> <h3><strong>Dataset Usage and Annotations</strong></h3> <p>Each name in test.txt, validation.txt and training.txt represents the name of the page and all its content (words, lines, paragraphs) must be in the respective partition.</p> <p>Annotations used in the dataset:</p> <ul> <li>&nbsp; <code>##@@???@@##</code>: Superscript text that has become unidentifiable and unreadable.</li> <li>&nbsp; <code>$$@@???@@$$</code>: Subscript text that has become unidentifiable and unreadable.</li> <li>&nbsp; <code>@@???@@</code>: Text that cannot be read or identified due to its illegibility.</li> <li>&nbsp; <code>##--xxx--##</code>: Text that has been added as a superscript and subsequently crossed out, rendering it illegible.</li> <li>&nbsp; <code>$$--xxx--$$</code>: Text that has been added as a subscript and subsequently crossed out, rendering it illegible.</li> <li>&nbsp; <code>--xxx--</code>: Text that has been crossed out in a way that makes it unreadable.</li> <li>&nbsp; <code>##--text--##</code>: Text that has been added as a superscript and subsequently crossed out, but remains legible.</li> <li>&nbsp; <code>$$--text--$$</code>: Text that has been added as a subscript and subsequently crossed out, but remains legible.</li> <li>&nbsp; <code>##text##</code>: Text added as a superscript in the line, typically as a correction or additional note.</li> <li>&nbsp; <code>$$text$$</code>: Text added as a subscript in the line, typically as a correction or additional note.</li> <li>&nbsp; <code>--text--</code>: Text that has been crossed out but remains readable.</li> </ul>

opencc-by-4.0May 2024View details →
zenodo40/100

ICDAR 2015 Competition HTRtS: Handwritten Text Recognition on the tranScriptorium Dataset

<p>This dataset comprises the dataset used for the ICDAR 2015 Competition on  Handwritten Text Recognition on the tranScriptorium Dataset. The handwritten images for this contest were drawn from the English “Bentham collection” dataset used in the TRAN SCRIPTORIUM project. The selected data has been written by several hands and entails significant variabilities and difficulties regarding the quality of text images, writing styles and crossed-out text. This contest is clearly more difficult than the the first edition both for training and for testing. A portion of the training dataset and the full test dataset were provided in the form of carefully segmented line images, along with the corresponding transcripts. Another portion of the training dataset was provided as raw images and their corresponding transcripts at region level.<br>  </p> <p>ICDAR 2015 competition HTRtS: handwritten text recognition on the tranScriptorium dataset<br> JA Sánchez, AH Toselli, V Romero, E Vidal.  In International Conference on Document Analysis and Recognition (ICDAR), pp. 1166-1170, 2015.</p>

opencc-by-4.0Jan 2017View details →
zenodo40/100

Train-B dataset for ICDAR2017 Competition on Handwritten Text Recognition on the READ Dataset (ICDAR2017 HTR). Batch 1 and Batch 2.

<p>Train-B Dataset.   Dataset of pages without any layout or text line information. The corresponding transcripts are provided at page level with line breaks. It has 10k pages, though for convenience it is divided into two 5k page batches. This information is provided in PAGE format. </p> <p>This dataset is complementary to this other dataset:</p> <p>https://zenodo.org/record/439807#.WOIBZ3WLSkA</p> <p>More information at:</p> <p>https://scriptnet.iit.demokritos.gr/competitions/~icdar2017htr/</p> <p> </p>

opencc-by-4.0Apr 2017View details →
zenodo40/100

Train-A dataset for ICDAR2017 Competition on Handwritten Text Recognition on the READ Dataset (ICDAR2017 HTR)

<p>Train-A Dataset of pages with manually revised baselines and the corresponding transcripts associated to them. This batch is small, 50 pages. Please, keep in mind that only the baselines have been manually corrected, The polygons associated to each line have not been manually reviewed. </p> <p>This dataset is complementary to this other dataset:</p> <p>https://zenodo.org/record/439811#.WOIF9HWLSkA</p> <p>More information at:</p> <p>https://scriptnet.iit.demokritos.gr/competitions/~icdar2017htr/</p>

opencc-by-4.0Apr 2017View details →
zenodo40/100

NorHand v3 / Dataset for Handwritten Text Recognition in Norwegian

<p>This dataset comprises Norwegian letter and diary documents from 19th and early 20th century. It can be used to train Handwritten Text Recognition (HTR) models.</p>

opencc-by-4.0Dec 2023View details →
zenodo40/100

Oficio de Hipotecas de Girona. A dataset of Spanish notarial deeds (18th Century) for Handwritten Text Recognition and Layout Analysis of historical documents.

<p>This dataset is a subset of 596 documents from the&nbsp;<em>Registre d&#39;Hipoteques de Girona</em> of 1769 collection, guarded by the <a href="http://xac.gencat.cat/ca/llista_arxius_comarcals/girones/"><em>Arxiu Hist&ograve;ric de Girona</em></a>. This collection, is composed by hundreds of thousands of notarial deeds from the XVIII-XIX century (1768-1862). Sales, redemption of censuses, inheritance and matrimonial chapters are among the most common documentary&nbsp;typologies in the collection.</p> <p>This dataset is composed of more than 23700 text lines&nbsp;written by a single hand, covering more that 50 different topics (documentary typologies) and a vocabulary of more than 2400 different words. The documents are transcribed using the so-called&nbsp;diplomatic criteria. Additionally, transcripts were tagged with&nbsp;<br> extra enriching/complementary information (e.g. expansion of the&nbsp;abbreviations, hyphen marks, etc.). Along with the transcripts &nbsp;the layout of the document is detected and recorded. Pages have&nbsp;been labeled using six different layout regions.</p> <p>The images along with their respective ground-truth was compiled in PAGE compliant XML format<br> by the <a href="http://www2.udg.edu/tabid/11296/Default.aspx"><em>Centre de Recerca d&#39;Hist&ograve;ria Rural</em></a>&nbsp;and the HTR group of the <a href="https://www.prhlt.upv.es">Pattern Recognition and Human Language Technologies Research Center</a>.</p>

opencc-by-nc-4.0Jul 2018View details →
zenodo36/100

Test-B1 ICDAR2017 Competition on Handwritten Text Recognition on the READ Dataset (ICDAR2017 HTR)

<ul> <li><strong>Test-B1</strong>:  a batch of page images annotated with the geometry of regions where to detect text line and recognize.</li> </ul>

opencc-by-4.0Jun 2017View details →
zenodo36/100

Test-B2 ICDAR2017 Competition on Handwritten Text Recognition on the READ Dataset (ICDAR2017 HTR)

<p><strong>Test-B2</strong>:  a batch of page images annotated with the geometry of regions where to detect text line and recognize.</p>

opencc-by-4.0Jun 2017View details →
zenodo36/100

Dataset for ICDAR2017 Competition on Handwritten Text Recognition on the READ Dataset (ICDAR2017 HTR)

<p><strong>Train-A:</strong> Dataset of pages with manually revised baselines and the corresponding transcripts associated to them. This batch is small, 50 pages. Please, keep in mind that only the baselines have been manually corrected, The polygons associated to each line have not been manually reviewed.</p> <p><strong>Train-B:</strong> Dataset of pages without any layout or text line information. The corresponding transcripts are provided at page level with line breaks. It has 10k pages, though for convenience it is divided into two 5k page batches. This information is provided in PAGE format.</p> <p><strong>Test A:</strong> Dataset of pages with manually revised baselines. This batch has 65 pages. The polygons associated to each line have not been manually reviewed.</p> <p><strong>Test-B1:</strong> The same dataset of pages of the Test A, but annotated only with the geometry of regions. Text line information is not provided.                                                   </p> <p><strong>Test-B2:</strong> Dataset of page images annotated with the geometry of regions where to detect text line and recognize. It has 57 pages.</p> <p><strong>Baseline.tgz:</strong> Baseline system trained using the first 40 pages of Train-A. The system is based on the deep learning toolkit to transcribe handwritten text images called Laia.</p> <p>More information at:</p> <p>https://scriptnet.iit.demokritos.gr/competitions/~icdar2017htr/</p> <p> </p>

opencc-by-4.0Jul 2017View details →
zenodo36/100

Appendix A - The Implications of Handwritten Text Recognition for Accessing the Past at Scale

<p>List of works identified through a Grounded Theory Method (GTM) of the current and near future implications of Handwritten Text Recognition (HTR) on the historical method and wider information environment. The findings of this data collection are provided in 'The Implications of Handwritten Text Recognition for Accessing the Past at Scale'</p>

opencc-by-4.0Nov 2023View details →
zenodo36/100

Appendix A - Understanding the application of Handwritten Text Recognition technology in heritage contexts

<p>This appendix lists all the categorised works mentioning the HTR software Transkribus, used in&nbsp;&#39;Understanding the application of Handwritten Text Recognition technology in heritage contexts: a systematic review of Transkribus in published research&#39;</p>

opencc-by-4.0Jan 2022View details →
zenodo36/100

ICDAR 2015 Competition HTRtS: Handwritten Text Recognition on the tranScriptorium Dataset Rerelease

<p>A new release of the dataset used in the ICDAR 2015 HTR competition in which all Page XML files are based on the same 2013-07-15 schema. It only contains page level images, Page XML files for train and test (including the ground truth transcripts for the test and train batch 1) and plain text files for train batch 2 that have the page level ground truth transcripts. The original version of this dataset can be found at http://doi.org/10.5281/zenodo.248733<br> &nbsp;</p>

opencc-by-4.0Jan 2018View details →
zenodo36/100

Handwritten Text Recognition Ground Truth Set: StABS Ratsbücher O10, Urfehdenbuch X

<p>Ground Truth for &quot;Urfehdenbuch X der Stadt Basel (1563-1569)&quot; at Staatsarchiv Basel-Stadt (StABS).</p> <p>Images and text aligned, using text-to-image (provided within Transkribus, <a href="https://www.readcoop.eu">www.readcoop.eu</a>).</p> <p>ALTO and Page XML are available for the text alignment.</p> <p>TEI to txt on page basis by Peter D&auml;ngeli.</p> <p>Derived from Transcription/TEI file: Urfehdenbuch X der Stadt Basel (1563-1569), in: Die Urfehdeb&uuml;cher der Stadt Basel &ndash; digitale Edition, hg. v. Susanna Burghartz, Sonia Calvi und Georg Vogeler Basel/Graz 2016. (zuletzt ver&auml;ndert am 31.1.2017): <a href="http://hdl.handle.net/11471/1010.2.1">hdl:11471/1010.2.1</a>.</p> <p>Image source: http://dokumente.stabs.ch/view/2010/Ratsbuecher_O_10/</p> <p>CC-BY-NC-SA: The license is inherited from the project &quot;Urfehdeb&uuml;cher der Stadt Basel &ndash; digitale Edition&quot;: http://gams.uni-graz.at/o:ufbas.1563.</p>

opencc-by-nc-sa-4.0Aug 2021View details →
zenodo36/100

The Belfort dataset: Handwritten Text Recognition from Crowdsourced Annotations

<p>This dataset includes minutes of Belfort municipal council drawn up between 1790 and 1946. Documents include deliberations, lists of councillors, convocations, and agendas.</p> <p>The dataset includes 24,105 text-line images that were automatically detected from pages. Up to 4 transcriptions are available for each line image: two from humans, and two from automatic models.</p> <p>We would like to thank the <em>Archives municipales de la ville de Belfort, France</em> for giving us access to these documents.</p>

opencc-by-4.0Jun 2023View details →
zenodo32/100

Test A for the ICDAR2017 Competition on Handwritten Text Recognition on the READ Dataset (ICDAR2017 HTR)

<p><strong>Test A. </strong>A batch of page images annotated with baselines.</p>

opencc-by-4.0Jun 2017View details →
zenodo28/100

M-POPP datasets: Datasets for full page text recognition and information extraction from French handwritten and printed marriage records

<h1><strong>M-POPP datasets</strong></h1> <p>This repository contains 2 datasets created within the <strong>EXO-POPP project</strong> (<a href="https://exopopp.hypotheses.org/">Optical EXtraction of handwritten named entities for marriage records of the POPulation of Paris</a>) for the task of text recognition and information extraction. These datasets have been published in <a href="https://arxiv.org/abs/2404.19329"><code>End-to-end information extraction in handwritten documents: Understanding Paris marriage records from 1880 to 1940</code></a> <code>[1]</code>at ICDAR 2024.</p> <p><strong>This version contains the labels for Handwritten Text Recognition and Handwritten Text Recognition + Information Extraction as used in our new paper "<a href="https://hal.science/hal-04555188">DANIEL: A fast Document Attention Network for Information Extraction and Labelling of handwritten documents</a>" [3].</strong></p> <p><strong>This version makes corrections to the handwritten dataset. </strong>More precisely, it corrects a few errors in transcription annotations and named entities.</p> <p><strong>The printed dataset is unchanged compared to version 2.</strong></p> <p><strong>The performances of the models described in [1] and [3] are detailled in the Leaderboard section.</strong></p> <h2><strong>General information</strong></h2> <p>The <strong>EXO-POPP project</strong> aims to establish a comprehensive database comprising 300,000 marriage records from Paris and its suburbs, spanning the years 1880 to 1940, which are preserved in over 130,000 scans of double pages. Each marriage record may encompass up to 118 distinct types of information that require extraction from plain text. The M-POPP corpus (which stands for Marriage records of the POPulation of Paris) is the corpus on which the EXO-POPP project focuses. This corpus was built by gathering the marriage records of Paris and its suburb regions (Hauts- de-Seine, Seine-Saint-Denis, Val-de-Marne).</p> <p>The M-POPP corpus are a subset of the M-POPP database with annotations for full-page text recognition and named entity recognition/information extraction from both handwritten and printed documents. The first dataset comprises handwritten marriage records, while the second dataset consists of typewritten marriage records. It should be noted that even in typewritten marriage records, some handwritten information occurs, especially concerning the names of the spouses, and notes in the margin.<br>The dataset contains single-page images obtained from the original scans of double pages via page segmentation.</p> <p>The structure of the files is the following:</p> <ul> <li>handwritten:&nbsp;<em>the handwritten dataset</em><br> <ul> <li>images: <em>images of the dataset divided following the split used in [1]</em><br> <ul> <li>train</li> <li>valid</li> <li>test</li> </ul> </li> <li>labels:&nbsp;<em>labels for joint handwritten text recognition and information extraction for each encoding tested in [1]</em></li> </ul> </li> <li>printed: <em>the printed dataset</em><br> <ul> <li>images:&nbsp;<em>images of the dataset divided following the split used in [1]</em><br> <ul> <li>train</li> <li>valid</li> <li>test</li> </ul> </li> <li>labels:&nbsp;<em>labels for joint handwritten text recognition and information extraction for each encoding tested in [1]</em></li> </ul> </li> <li>encoding-2-to-encoding-5.json:&nbsp;<em>a JSON file giving the correspondence between the symbols of encoding 2 and encoding 5.</em></li> </ul> <p>&nbsp;</p> <p>Table 1: Details on the split of the handwritten dataset.</p> <table> <tbody> <tr> <td>&nbsp;</td> <td>Train</td> <td>Validation</td> <td>Test</td> </tr> <tr> <td>Pages</td> <td>250</td> <td>32</td> <td>32</td> </tr> <tr> <td>Acts</td> <td>344</td> <td>51</td> <td>53</td> </tr> <tr> <td>Named entities</td> <td>16727</td> <td>2223</td> <td>2517</td> </tr> </tbody> </table> <p>&nbsp;</p> <p>Table 2: Details on the split of the printed dataset.</p> <table> <tbody> <tr> <td>&nbsp;</td> <td>Train</td> <td>Validation</td> <td>Test</td> </tr> <tr> <td>Pages</td> <td>116</td> <td>14</td> <td>13</td> </tr> <tr> <td>Acts</td> <td>363</td> <td>43</td> <td>30</td> </tr> <tr> <td>Named entities</td> <td>22036</td> <td>2559</td> <td>2405</td> </tr> </tbody> </table> <p>&nbsp;</p> <p>Table 3: Average annotation statistics per act for the two M-POPP datasets.</p> <table> <tbody> <tr> <td>Dataset</td> <td># of characters</td> <td># of words</td> <td># of named entities</td> </tr> <tr> <td>Handwritten</td> <td>1519</td> <td>231</td> <td>48</td> </tr> <tr> <td>Printed</td> <td>1328</td> <td>200</td> <td>60</td> </tr> </tbody> </table> <p>&nbsp;</p> <h3><strong>Document structure Annotation</strong></h3> <p>We employ the procedure applied in [2], which involves adding opening and closing tags to the character set for each text block we want to recognize.<br>In total, we define four types of text blocks.</p> <ul> <li>Block A is located in the margin and contains the last names of the married couple, possibly with their first names and the date of the marriage.</li> <li>Block B is the body of the text. Block B is the one that contains most of the information to be extracted.</li> <li>Block C is optional and corresponds to marginal notes used in various cases, such as the mention of a divorce or a correction made to the act.</li> <li>Block D corresponds to a set containing a block A and a block B, optionally with one or more blocks C.</li> </ul> <p>&nbsp;</p> <h3><strong>Information Extraction annotation</strong></h3> <p>The dataset contains 118 information categories. As explained in the paper, we broke down the named entities into sub-elements pertaining to 4 hierarchical levels, which reduces the total number of categories to 23 instead of 118. Notice that level 1, 2, and 3 categories do not encode named entities but rather the relations that may occur between some lower level categories for example: (day, birth, husband) encodes the fact that the annotated piece of text is the date of birth of the husband.&nbsp;</p> <p>For these datasets, we chose to represent these hierarchical elements with emojis. For instance, the information <em>first name</em> is represented by the emoji 💬.<br>The meaning of each emoji can be found in Table 4. To determine the best way to encode named entities in the ground truth, we compared in [1] 5 types of encoding. To illustrate these encodings, let&rsquo;s take for instance&nbsp;<em>Louis Alexandre MOUDEL</em> that we define as the father of the bride, where <em>Louis Alexandre</em> are his two first names, and <em>Moudel</em> is his last name.&nbsp;</p> <p>1) Single separate tags before each word: In this approach, each level of information is indicated by a dedicated tag, and the tags are placed before the word they encode information for. With this encoding, the ground truth for the example would be:</p> <p>💬👴👰Louis &nbsp; 💬👴👰Alexandre&nbsp; 🗨️👴👰MOUDEL</p> <p>2) Single separate tags after each word: Similar to the previous approach, except here the tags are placed after the word. With this encoding the previous example becomes:</p> <p>Louis👰👴💬&nbsp; Alexandre👰👴💬&nbsp; MOUDEL👰👴🗨️</p> <p>3) Open &amp; close separate tags: Here, each word presenting information to be extracted is surrounded by one or more opening and closing tags, where each tag encodes a level of information. So the example would be as:</p> <p>&lt;👰&gt; &lt;👴&gt; &lt;💬&gt; Louis &lt;\💬&gt; &lt;\👴&gt; &lt;\👰&gt;<br>&lt;👰&gt; &lt;👴&gt; &lt;💬&gt; Alexandre &lt;\💬&gt; &lt;\👴&gt; &lt;\👰&gt;<br>&lt;👰&gt; &lt;👴&gt; &lt;🗨️&gt; MOUDEL &lt;\🗨️&gt; &lt;\👴&gt; &lt;\👰&gt;</p> <p>4) Nested open &amp; close separate tags: Similar to the previous approach, but this time a tag is closed only when the encoded information is no longer the same for that level of information. We can see in the example below that the tags for wife and father are only used twice.</p> <p>&lt;👰&gt; &lt;👴&gt; &lt;💬&gt; Louis Alexandre &lt;\💬&gt; &lt;🗨️&gt; MOUDEL &lt;\🗨️&gt;</p> <p>5) Single combined tags after each word: In the last approach, one tag encodes all the hierarchical levels constituting information. The tags are located after the word they encode information for.&nbsp;</p> <p>Louis&lt;wife_father_first_name&gt;&nbsp; Alexandre&lt;wife_father_first_name&gt;&nbsp; MOUDEL&lt;wife_father_family_name&gt;</p> <p>NB: In the labels file of encoding 5, the information are still encoded with emojis but the chosen emojis do not have a semantic meaning due to the number of information categories to be represented. The correspondence between the symbols of encoding 2 and encoding 5 can be found in the file&nbsp;<em>encoding-2-to-encoding-5.json</em>.<em><br></em></p> <p>&nbsp;</p> <p>Table 4: Details of the hierarchical breakdown of named entities. Each tag is placed in the corresponding hierarchical level and associated with the emoji representing it.</p> <table> <tbody> <tr> <td>Level</td> <td>Tags</td> <td>&nbsp;</td> <td>&nbsp;</td> <td>&nbsp;</td> </tr> <tr> <td>1</td> <td>Administrative 📖</td> <td> <pre>Husband<code> 👨</code></pre> </td> <td>Wife 👰</td> <td>Witness 🥸</td> </tr> <tr> <td>2</td> <td>Father 👴</td> <td>Mother 👵</td> <td>Ex-husband 💔</td> <td>&nbsp;</td> </tr> <tr> <td>3</td> <td>Birth 🏥</td> <td>Residence 🏠</td> <td>&nbsp;</td> <td>&nbsp;</td> </tr> <tr> <td>4</td> <td>First name 💬</td> <td>Family name 🗨️</td> <td>Age ⌛</td> <td>Occupation 🔧</td> </tr> <tr> <td>5</td> <td>Street number 🔟</td> <td>Street type 🛣</td> <td>Street name 🔠</td> <td>City 🌆</td> </tr> <tr> <td>&nbsp;</td> <td>Department 🗺</td> <td>Country 🗺</td> <td>Day 🌞</td> <td>Month 📅</td> </tr> <tr> <td>&nbsp;</td> <td>Year 🗓</td> <td>Hour ⏰</td> <td>Minute ⏱</td> <td>&nbsp;</td> </tr> </tbody> </table> <p>&nbsp;</p> <h2><strong>Leaderboard</strong></h2> <h3><strong>Results on M-POPP handwritten</strong></h3> <p><strong>HTR<br></strong></p> <p>The following table contains the current leaderboard of M-POPP v3 for HTR on the handwritten dataset.</p> <p>In this configuration, layout block C is not considered.</p> <p>These results for DAN NER are given using the named entity encoding format 5 described above.</p> <p>HTR stands for Handwritten Text Recognition and HTR+IE for combined Handwritten Text Recognition and Information Extraction.</p> <p>Metrics are expressed in percentages.</p> <table> <tbody> <tr> <td>Method</td> <td>CER</td> <td>WER</td> <td>LOER</td> <td>mAP CER</td> </tr> <tr> <td>DAN - HTR [1]</td> <td>7.21</td> <td>16.42</td> <td>5.35</td> <td>83.03</td> </tr> <tr> <td>DAN NER - HTR + IE [1]</td> <td>6.52</td> <td>14.80</td> <td>3.79</td> <td>86.29</td> </tr> <tr> <td>DANIEL - HTR [3]</td> <td>5.72</td> <td>14.08</td> <td>1.34</td> <td>89.28</td> </tr> </tbody> </table> <p>&nbsp;</p> <p><strong>NER</strong></p> <p>The following table contains the current leaderboard of M-POPP v3 for NER on the handwritten dataset.</p> <p>In this configuration, layout block C is not considered.</p> <p>These results are given using the named entity encoding format 5 described above.</p> <p>Metrics are expressed in percentages.</p> <table> <tbody> <tr> <td>Method</td> <td>F1</td> </tr> <tr> <td>DAN NER [1]</td> <td>76.37</td> </tr> <tr> <td>DANIEL [3]</td> <td>76.37</td> </tr> </tbody> </table> <h3>&nbsp;</h3> <h3><strong>Results on M-POPP printed</strong></h3> <p><strong>HTR</strong></p> <p>The following table contains the current leaderboard of this version for TR on the printed dataset.</p> <p>The labels, and therefore the results, are the same as in version 2.</p> <p>In this configuration, only layout block B is considered.</p> <p>These results for DAN NER are given using the named entity encoding format 5 described above.</p> <p>TR stands for Text Recognition and TR+IE for combined Text Recognition and Information Extraction.</p> <p>Metrics are expressed in percentages.</p> <p><strong>NB:</strong> The model DANIEL from [3] was not evaluated on the printed dataset.</p> <table> <tbody> <tr> <td>Method</td> <td>CER</td> <td>WER</td> </tr> <tr> <td>DAN - TR [1]</td> <td>0.88</td> <td>3.17</td> </tr> <tr> <td>DAN NER - TR + IE [1]</td> <td>1.54</td> <td>3.55</td> </tr> </tbody> </table> <p>&nbsp;</p> <p><strong>NER</strong></p> <p>The following table contains the current leaderboard of this version for NER on the printed dataset.</p> <p>The labels, and therefore the results, are the same as in version 2.</p> <p>In this configuration, only layout block B is considered.</p> <p>These results are given using the named entity encoding format 5 described above.</p> <p>Metrics are expressed in percentages.</p> <p><strong>NB:</strong> The model DANIEL from [3] was not evaluated on the printed dataset.</p> <table> <tbody> <tr> <td>Method</td> <td>F1</td> </tr> <tr> <td>DAN NER [1]</td> <td>93.04</td> </tr> </tbody> </table> <h2>&nbsp;</h2> <h2><strong>Citation Request</strong></h2> <p>If you publish material based on this database, we request you to include a reference to the paper&nbsp;<code><a href="https://arxiv.org/abs/2404.19329">T. Constum, L. Preel, T. Paquet, P. Tranouez, S. Br&eacute;e, End-to-end information extraction in handwritten documents: Understanding Paris marriage records from 1880 to 1940, International Conference on Document Analysis and Recognition (ICDAR), Athens, Greece, 2024</a>.</code></p> <p>&nbsp;</p> <h2><strong>Bibliography</strong></h2> <p><a href="https://arxiv.org/abs/2404.19329">1: T. Constum, L. Preel, T. Paquet, P. Tranouez, S. Br&eacute;e: End-to-end information extraction in handwritten documents: Understanding Paris marriage records from 1880 to 1940, International Conference on Document Analysis and Recognition (ICDAR), Athens, Greece, 2024.</a></p> <p>2: D.Coquenet, C. Chatelain, T. Paquet: DAN: a Segmentation-free Document Attention Network for Handwritten Document Recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence pp. 1&ndash;17 (2023).</p> <p><a href="https://hal.science/hal-04555188/">3: T. Constum, T. Paquet, P. Tranouez: DANIEL: A fast Document Attention Network for Information Extraction and Labelling of handwritten documents, preprint, 2024&nbsp;</a></p>

opencc-by-4.0Apr 2024View details →
zenodo24/100

NorHand / Dataset for Handwritten Text Recognition in Norwegian

<p>The dataset comprises Norwegian letter and diary line images and text from 19th and early 20th century.</p>

opencc-by-4.0May 2022View details →
zenodo20/100

Dataset and scripts used in paper "Exploiting Existing Modern Transcripts for Historical Handwritten Text Recognition"

<p>This package contains the dataset and scripts used in the research paper titled "Exploiting Existing Modern Transcripts for Historical Handwritten Text Recognition" that was published in the ICFHR 2016 conference proceedings.</p> <p>Unfortunately not all of the code used in the experiments is open source, so what is provided is not enough to really reproduce the results. Nevertheless, it should be useful to understand in more detail how the experimentation was performed. The scripts make use of the library for handwritten text recognition that is available at:</p> <p>https://github.com/mauvilsa/htrsh</p> <p>The dataset is in PRImA Page XML format (http://www.primaresearch.org/tools). To visualize the what is contained in the xml files on top of the images, the following open source tool can be used:</p> <p>https://github.com/mauvilsa/nw-page-editor</p>

restrictedOct 2016View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record