Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

4

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

4 results for “Handwriting Recognition”

Learn how ShareScore rates datasets ↗
zenodo44/100

The e-NDP project : collaborative digital edition of the Chapter registers of Notre-Dame of Paris (1326-1504). Ground-truth for handwriting text recognition (HTR) on late medieval manuscripts.

<p>The <a href="https://endp.hypotheses.org/">e-NDP project</a>, funded by the ANR, is led by the <a href="https://lamop.hypotheses.org/6870">LaMOP</a> (Julie Claustre and Darwin Smith).</p> <p>The project&#39;s partners are the Archives nationales, the&nbsp;Biblioth&egrave;que nationale de France (Department of Manuscripts, Biblioth&egrave;que de l&#39;Arsenal), the &Eacute;cole nationale des chartes and the Biblioth&egrave;que Mazarine.</p> <p>The e-NDP project aims at renewing our knowledge on <strong>Notre-Dame de Paris cathedral</strong> through the creation of a collaborative digital edition of the registers of its Chapter (1326-1504, <em>AN LL 105-128</em>), the community of 51 canons meeting three times a week on set days to take all administrative, financial and practical decisions pertaining to the cathedral, its estate and the society living in its cloister. This corpus has never been the object of a comprehensive study to understand the workings and history of this urban enclave and powerful community. The collaborative digital edition is based on a process of<strong> handwriting text recognition (HTR)</strong>, tested and supervised by scholars, researchers and engineers combining expertise in Medieval history, paleography, philology and digital humanities. The edition shall allow a better insight into the Chapter&rsquo;s administration, into its economical and political power within Paris, and the relationships it maintained with other institutions in the city.</p> <p>&nbsp;</p> <p><strong>Section 1 : The e-NDP ground-truth dataset for Handwriting text recognition.</strong></p> <p>The full e-NDP corpus kept today in the French National Archives and was entirely digitized and described in its&nbsp;<a href="https://www.siv.archives-nationales.culture.gouv.fr/siv/rechercheconsultation/consultation/ir/consultationIR.action?formCaller=GENERALISTE&amp;irId=FRAN_IR_059635">catalog</a>&nbsp;in 2022.</p> <p>The first major goal of the&nbsp;e-NDP projet is to propose a first automatic transcription of the 14k pages composing the 26 chapter registers. To achieve this goal representative samples from&nbsp;each one of the volumes were selected and transcribed in order to train a specialized HTR model able to propose a high quality automatic transcription. The collected ground-truth released on this repository currently has <strong>512 pages from the 26 registers</strong> of the cathedral chapter preserved in the National Archives (LL105 - LL128, <strong>1326-1504</strong>). The transcriptions were manually completed in <strong>two rounds</strong> by a group of 12 contributors, historians and paleographers, over the course of 2021-2022 using <a href="https://escriptorium.paris.inria.fr/">eScriptorium </a>as annotation environment.&nbsp;&nbsp;</p> <p>&nbsp;</p> <p><strong>Ground-truth features :</strong></p> <p><br> <em>Number of hands </em>: according to our estimates no fewer than 18&nbsp;main hands were involved in the writing of the registers during the medieval period.&nbsp;</p> <p><em>Language</em> : More than 98% of the content of the registers was written in Latin, the rest in French. The exact percentage is hard to estimate because the vernacular language is often used in formulae, notes and comments. It is rare to find entire pages or blocks written in French.&nbsp;</p> <p><em>Script family</em> : The registers were written using a Cursive script (ca. late XIIIe - XVIe).</p> <p><em>Documental typology</em> : The volumes containing the chapter conclusions were conceived to serve&nbsp;as memorial&nbsp;records, but above all as documents for regular use and consultation in the daily practice of administration and management. In diplomatics the notion of &quot;documentary manuscripts&quot; is used to describe this kind of sources&nbsp;also by opposition to books and litterary or&nbsp;normative&nbsp;manuscripts.</p> <table align="center"> <caption><strong>Ground truth statistics</strong></caption> <tbody> <tr> <th>Text units</th> <th>Count</th> </tr> <tr> <td>Pages</td> <td>512</td> </tr> <tr> <td>Annotated regions (see section 2)</td> <td>2448</td> </tr> <tr> <td>Lines of text</td> <td>34231</td> </tr> <tr> <td>Tokens</td> <td>205083</td> </tr> <tr> <td>Characters</td> <td>3320407</td> </tr> </tbody> </table> <p>&nbsp;</p> <p><strong>Rules of transcription :</strong></p> <ul> <li>The abbreviations have been resolved, both those by suspension (<code>facimꝰ</code> ---&gt; <code>facimus</code>) and by contraction (<code>d&ntilde;i</code> --&gt; <code>domini</code>). Likewise, those using conventional signs (<code>⁊</code> --&gt; <code>et</code> ; <code>ꝓ</code> --&gt; <code>pro</code>) have been resolved.&nbsp;</li> <li>The named entities (names of persons, places and institutions) have been <code>capitalized</code>. The beginning of a block of text as well as the original capitals used by the notary are also capitalized.</li> <li>The consonantal <code>i</code> and <code>u</code> characters have been transcribed as <code>j</code> and <code>v</code> in both French and Latin.</li> <li>The punctuation marks used in the text: <code>.</code> and <code>/</code> have been transcribed, but the transcription has not been standardized with modern punctuation.</li> <li>Corrections and words that appear cancelled in the manuscript have been transcribed surrounded by the sign <code>$</code> at the beginning and at the end.</li> <li>More specific transcription rules can be found into the file <code>transcription_guidelines.pdf</code></li> </ul> <p>&nbsp;</p> <p><strong>Section 2. e-NDP Layout Segmentation.</strong></p> <p>Layout segmentation is a compulsory step before HTR recognition in order to distinguish sections and regions inside a document. This process intend to separate interdependant page zones to produce a recognition in a section-sequence order and not in a line-sequence order which mix textual and peri-textual content.</p> <p>The regions of 364&nbsp;pages (see <code>GT-layout_list</code>) of the e-NDP corpus were annotated using a 5 sections vocabulary (see <code>endp_layout_regions</code>) in order to describe&nbsp;the page distribution in all the 26 volumes :</p> <ol> <li><em>Block</em>&nbsp;: All the central text blocks, that normally corresponds to the main content called &quot;conclusions&quot; in registers.</li> <li><em>Liste</em>&nbsp;: List of names of the canons who were present during the meeting. Normally located before the <em>conclusions</em>.</li> <li><em>Entr&eacute;e</em>&nbsp;: Marginal notes or entries to inform about the content of <em>conclusions</em>.</li> <li><em>Date</em>&nbsp;: Paragraph contending the date. Normally at the head of a <em>conclusion</em>, but separate of the main body.</li> <li><em>Num&eacute;rotation</em>&nbsp;: Page numbers in roman or arabic. Usually appear in the top corners of the pages.</li> </ol> <table align="center"> <caption><strong>Layout GT statistics</strong></caption> <tbody> <tr> <th>Region</th> <th>Count</th> </tr> <tr> <td>block</td> <td>833</td> </tr> <tr> <td>liste</td> <td>431</td> </tr> <tr> <td>date</td> <td>448</td> </tr> <tr> <td>entr&eacute;e</td> <td>205</td> </tr> <tr> <td>num&eacute;rotation</td> <td>531</td> </tr> </tbody> </table> <p>&nbsp;</p> <p><strong>Section 3. The e-NDP HTR modeling.</strong></p> <p>The e-NDP project has progressively trained several HTR models adapted to work on late medieval cursive in order to accelerate the production of ground truth. Currently the best model delivers an average&nbsp;<strong>CER (Character error ratio) of 9.7%</strong> in handwriting recognition on&nbsp;the 26 registers (see <code>endp_learning_curve</code>) and can serve as generalist model&nbsp;for other manuscripts of the same period and similar script family. These models and their training implementation details can be found in the project&#39;s github <a href="https://github.com/chartes/e-NDP_HTR">repository</a>.&nbsp;</p> <p>Additionally, the automatic HTR transcriptions of the 26 registers (14k pages, 4.5M tokens) enriched with lexical and semantical information has been the subject of a first <a href="https://nosketch-engine.lamop.fr/#dashboard?corpname=endp">online publication</a> using the NoSketch engine that allows advanced data mining based on the combination of data, metadata and NLP features.&nbsp;</p> <p>&nbsp;</p> <p><strong>Section 4. Dataset content.</strong></p> <p>This zip dataset contains :</p> <p>- <code>HTR_ground_truth</code> : Two folders containing the jpg / jpeg images and their curated transcriptions in PAGE XML format.</p> <p>- <code>images_docs</code> : 4 files illustrating the different phases of the project (list of GT for layout segmentation, layout ontologie, transcription guideline and HTR evaluation curves)</p>

opencc-by-4.0Feb 2023View details →
zenodo40/100

POPP Datasets : Datasets for handwriting recognition from French population census

<p><strong>POPP datasets</strong></p> <p>This repository contains 3 datasets created within the POPP project (<a href="https://popp.hypotheses.org/#ancre2">Project for the Oceration of the Paris Population Census</a>) for the task of handwriting text recognition. These datasets have been published in <a href="https://hal.science/hal-03675614/"><em>Recognition and information extraction in historical handwritten tables: toward understanding early 20th century Paris census</em> at DAS 2022.</a></p> <p>The 3 datasets are called &ldquo;Generic dataset&rdquo;, &ldquo;Belleville&rdquo;, and &ldquo;Chauss&eacute;e d&rsquo;Antin&rdquo; and contains lines made from the extracted rows of census tables from 1926. Each table in the Paris census contains 30 rows, thus each page in these datasets corresponds to 30 lines.</p> <p>The structure of each dataset is the following:</p> <ul> <li>double-pages : images of the double pages</li> <li>pages: <ul> <li>images: images of the pages</li> <li>xml: METS and ALTO files of each page containing the coordinates of the bounding boxes of each line</li> </ul> </li> <li>lines: contains the labels in the file <code>labels.json</code> and the line images splitted into the folders <em>train</em>, <em>valid</em> and <em>test</em>. The double pages were scanned at a resolution of 200dpi and saved as PNG images with 256 gray levels. The line and page images are shared in the TIFF format, also with 256 gray levels.</li> </ul> <p>Since the lines are extracted from table rows, we defined 4 special characters to describe the structure of the text:</p> <ul> <li>&curren; : indicates an empty cell</li> <li>/ : indicates the separation into columns</li> <li>? : indicates that the content of the cell following this symbol is written above the regular baseline</li> <li>! : indicates that the content of the cell following this symbol is written below the regular baseline</li> </ul> <p>We provide a script <code>format_dataset.py</code> to define which special character you want to use in the ground-truth.</p> <p>The split for the <em>Generic Dataset</em> and <em>Belleville</em> have been made at the double-page level so that each writer only appears in one subset among train, evaluation and test. The following table summarizes the splits and the number of writers for each dataset:</p> <table> <thead> <tr> <th>Dataset</th> <th>train - # of lines</th> <th>validation - # of lines</th> <th>test - # of lines</th> <th># of writers</th> </tr> </thead> <tbody> <tr> <td>Generic</td> <td>3840 (128 pages)</td> <td>480 (16 pages)</td> <td>480 (16 pages)</td> <td>80</td> </tr> <tr> <td>Belleville</td> <td>1140 (38 pages)</td> <td>150 (5 pages)</td> <td>180 (6 pages)</td> <td>1</td> </tr> <tr> <td>Chauss&eacute;e d&rsquo;Antin</td> <td>625</td> <td>78</td> <td>77</td> <td>10</td> </tr> </tbody> </table> <p>&nbsp;</p> <p><strong>Generic dataset (or POPP dataset)</strong></p> <ul> <li>This dataset is made 4800 annotated lines extracted from 80 double pages of the 1926 Paris census.</li> <li>There is one double page for each of the 80 districts of Paris</li> <li>There is one writer per double page so the dataset contains 80 different writers.</li> </ul> <p>&nbsp;</p> <p><strong>Belleville dataset</strong></p> <p>This dataset is a mono-writer dataset made of 1470 lines (49 pages) from the <em>Belleville</em> district census of 1926.</p> <p>&nbsp;</p> <p><strong>Chauss&eacute;e d&rsquo;Antin dataset</strong></p> <p>This dataset is a multi-writer dataset made of 780 lines (26 pages) from the <em>Chauss&eacute;e d&rsquo;Antin</em> district census of 1926 and written by 10 different writers.</p> <p>&nbsp;</p> <p><strong>Error reporting</strong></p> <p>It is possible that errors persist in the ground truth, so any suggestions for correction are welcome. To do so, please make a merge request on the <a href="https://github.com/Shulk97/POPP-datasets">Github repository</a> and include the correction in both the labels.json file and in the XML file concerned.</p> <p>&nbsp;</p> <p><strong>Citation Request</strong></p> <p>If you publish material based on this database, we request you to include a reference to paper <a href="http://link.springer.com/chapter/10.1007/978-3-031-06555-2_10"><code>T. Constum, N. Kempf, T. Paquet, P. Tranouez, C. Chatelain, S. Br&eacute;e, and F. Merveille,Recognition and information extraction in historical handwritten tables: toward understanding early 20th century Paris census ,Document Analysis Systems (DAS), pp. 143- 157, La Rochelle, 2022.</code></a></p>

opencc-by-4.0Mar 2022View details →
zenodo32/100

Mathematical Subjective Questions Handwriting Recognition TestSet

Open the record for dataset details and reuse information.

opencc-by-4.0Sep 2024View details →
zenodo32/100

GoBo - A Handwriting Recognition dataset for Personalization

<p>This dataset comprises the images for the personalization described in the paper&nbsp;<em>Personalizing Handwriting Recognition Systems with Limited User-Specific Samples</em>.<br> &nbsp;</p> <p>Dataset Statistics (v.1.0)</p> <p>* Handwritten word-level images<br> * English<br> *&nbsp;40 Participants<br> * 5 sets from different sources for personalization&nbsp;<br> * 2 sets from 2 domains (same domains as 2 personalization sets) for testing<br> * 926 words/writer, 37k words in total<br> <br> More details can be found on the Github Repository:<br> <a href="https://github.com/catalpa-cl/GoBo/">Github GoBo</a><br> <br> <br> Model<br> gobo_Baselinemodel.hdf5</p>

opencc-by-4.0Sep 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record