Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
30
datasets available to search
ShareScore release 0.9.0
Dataset results
30 results for “text recognition”
ESTER-Pt: An Evaluation Suite for TExt Recognition in Portuguese
<p><em><strong>disclaimer</strong></em>: Version accepted as full paper in ICDAR 2023.</p> <p>Optical Character Recognition (OCR) is a technology that enables machines to read and interpret printed or handwritten texts from scanned images or photographs. However, the accuracy of OCR systems can vary depending on several factors, such as the quality of the input image, the font used, and the language of the document. As a general tendency, OCR algorithms perform better in resource-rich languages as they have more annotated data to train the recognition process. We propose ESTER-Pt, an Evaluation Suite for TExt Recognition in Portuguese in this work. Despite being one of the largest languages in terms of speakers, OCR in Portuguese remains largely unexplored. Our evaluation suite comprises four types of resources: synthetic text-based documents, synthetic image-based documents, real scanned documents, and a hybrid set with real image-based documents that were synthetically degraded.</p>
Dataset for Paper: Text Line Detection and Recognition of Greek Polytonic Documents
<p>Dataset for Paper: Text Line Detection and Recognition of Greek Polytonic Documents, P. Kaddas, B. Gatos, K. Palaiologos, K. Christopoulou and K. Kritsis, 4th Workshop on Machine Learning (WML), San Jose, California, USA</p> <p>We introduce a new dataset, named GTLD-small dataset, with annotated text line quadrilateral polygons of 1.642 documents, including annotations on 3 datasets (Tobacco-3482, PIOP and ShakeIT dataset)</p> <table> <caption>Overview of the datasets included in this work and the number of images used for training, validation and testing.</caption> <thead> <tr> <th scope="col">Collection</th> <th scope="col">#Total</th> <th scope="col">#train</th> <th scope="col">#val</th> <th scope="col">#test</th> </tr> </thead> <tbody> <tr> <td>PIOP-small</td> <td>950</td> <td>672</td> <td>90</td> <td>188</td> </tr> <tr> <td>ShakeIT-small</td> <td>357</td> <td>264</td> <td>27</td> <td>66</td> </tr> <tr> <td>Tobacco-3482-small</td> <td>335</td> <td>240</td> <td>30</td> <td>65</td> </tr> </tbody> </table> <p> </p>
The Belfort dataset: Handwritten Text Recognition from Crowdsourced Annotations
<p>This dataset includes minutes of Belfort municipal council drawn up between 1790 and 1946. Documents include deliberations, lists of councillors, convocations, and agendas.</p> <p>The dataset includes 24,105 text-line images that were automatically detected from pages. Up to 4 transcriptions are available for each line image: two from humans, and two from automatic models.</p> <p>We would like to thank the <em>Archives municipales de la ville de Belfort, France</em> for giving us access to these documents.</p>
Towards a general open dataset and model for late medieval Castilian text recognition (HTR/OCR). Datasets and scripts
<p>This repository contains the dataset of the article "Towards a general open dataset and models for late medieval Castilian writing (HTR/OCR)" submitted to the Journal of Data Mining and Digital Humanities (JDMDH). I refer to the paper (<a href="https://doi.org/10.5281/zenodo.7387376">https://doi.org/10.5281/zenodo.7387376</a>) for the description of the corpus and the models.</p><p><strong>The dataset is in version V2: it contains the allographetic AND graphematic transcriptions (files `*.normalized.xml`) and models.</strong></p><p><i>Caveat</i>: the allographetic transcriptions and models only are described in the data paper mentionned above. The graphematic transcriptions are produced using a Chocomuffin conversion table (see `corpus/conversion_table.csv`) to reduce each allograph to its corresponding grapheme. The abbreviations are not expanded.</p><p>Please cite the following paper if you use this dataset or the models:</p><p>@article{gille_levenson_2023_towards,<br> author = {Gille Levenson, Matthias},<br> date = {2023},<br> journaltitle = {Journal of Data Mining and Digital Humanities},<br> doi = {<a href="https://doi.org/10.46298/jdmdh.10416">10.46298/jdmdh.10416</a>},<br> editor = {Pinche, Ariane and Stokes, Peter},<br> issuetitle = {Special Issue: Historical documents and automatic text recognition},<br> title = {Towards a general open dataset and models for late medieval Castilian text recognition<br>(HTR/OCR)},</p><p>GILLE LEVENSON , Matthias, « Towards a general open dataset and models for late medieval Castilian<br>text recognition (HTR/OCR) », <i>Journal of Data Mining and Digital Humanities</i> (2023) : Special<br>Issue : Historical documents and automatic text recognition, eds. Ariane PINCHE and Peter<br>STOKES, DOI : <a href="https://doi.org/10.46298/jdmdh.10416">10.46298/jdmdh.10416</a>.</p><p>The image of the manuscript M (Esc_M) has not yet been uploaded, pending permission from the library that keeps the manuscript.</p><p>All images are kept in a directory named after the place where the manuscript is kept, and the sigla of the witness for the in-domain dataset.</p><p> </p><p>The global licence for the dataset (except for images) is CC-BY-NC-SA.</p><p>All manuscripts reproductions are published with the authorization of the libraries.</p><p><strong>©Biblioteca General Histórica de Salamanca</strong></p><p>Universidad de Salamanca (España), Biblioteca General Histórica, Ms. 2709 (L)</p><p>Universidad de Salamanca (España), Biblioteca General Histórica, Ms. 2097 (J)</p><p>Universidad de Salamanca (España), Biblioteca General Histórica, Ms. 2673</p><p>Universidad de Salamanca (España), Biblioteca General Histórica, Ms. 2011</p><p>Universidad de Salamanca (España), Biblioteca General Histórica, Ms. 2654</p><p>Universidad de Salamanca (España), Biblioteca General Histórica, Ms. 2086</p><p><strong>©Museo Lázaro Galdiano. Madrid</strong></p><p>Inv. 15304, Fundación Lázaro Galdiano (A)</p><p><strong>©Universidad de Valladolid</strong></p><p>Ms. 251, Biblioteca Santa Cruz (S)</p><p><strong>©Real Biblioteca del Escorial</strong></p><p>Ms. K.I.5, Biblioteca del Real Monasterio del Escorial (Q)</p><p>Ms. h.I.8, Biblioteca del Real Monasterio del Escorial (M): to be published</p><p>Ms. Z-I-12</p><p>Ms.Z-III-9</p><p>Ms. X-III-4</p><p>Ms. h-III-9</p><p>Ms. b-IV-15</p><p>Ms. b-II-11</p><p>Ms. a-II-17</p><p>Ms. T-III-5</p><p><strong>©Rosenbach Foundation</strong></p><p>Ms. 482/2 (U)</p><p><strong>© Gallica.bnf.fr</strong></p><p>Espagnol 12</p><p>Espagnol 36</p><p>Espagnol 218</p><p><strong>© Bodleian Library</strong></p><p>Ms. Span. d. 1</p><p>Ms. Span. d. 2/1</p><p><strong>© Biblioteca Real, Madrid</strong></p><p>Ms. II/215 (G)</p><p><strong>© Biblioteca Nacional de España</strong></p><p>Mss/4183</p><p>Inc/901 (Z)</p><p><strong>© Biblioteca Universitaria, Sevilla</strong></p><p>Ms. 332/131 (R)</p><p> </p><p>Edit: add result files</p>
NLMChem a new resource for chemical entity recognition in PubMed full-text literature
Open the record for dataset details and reuse information.
Low resolution scanned text dataset for optical character recognition
<p>A collection of scanned pages of English text designed for testing low resolution OCR systems. There are 11 different pieces of text, each of which contains 5 pages of text. Each of these 55 pages is typeset in 18 different fonts and then scanned at 300 dpi, producing a total of 990 pages of scanned text. Downsampled 60 dpi and 75 dpi versions are included.</p>
Test A for the ICDAR2017 Competition on Handwritten Text Recognition on the READ Dataset (ICDAR2017 HTR)
<p><strong>Test A. </strong>A batch of page images annotated with baselines.</p>
M-POPP datasets: Datasets for full page text recognition and information extraction from French handwritten and printed marriage records
<h1><strong>M-POPP datasets</strong></h1> <p>This repository contains 2 datasets created within the <strong>EXO-POPP project</strong> (<a href="https://exopopp.hypotheses.org/">Optical EXtraction of handwritten named entities for marriage records of the POPulation of Paris</a>) for the task of text recognition and information extraction. These datasets have been published in <a href="https://arxiv.org/abs/2404.19329"><code>End-to-end information extraction in handwritten documents: Understanding Paris marriage records from 1880 to 1940</code></a> <code>[1]</code>at ICDAR 2024.</p> <p><strong>This version contains the labels for Handwritten Text Recognition and Handwritten Text Recognition + Information Extraction as used in our new paper "<a href="https://hal.science/hal-04555188">DANIEL: A fast Document Attention Network for Information Extraction and Labelling of handwritten documents</a>" [3].</strong></p> <p><strong>This version makes corrections to the handwritten dataset. </strong>More precisely, it corrects a few errors in transcription annotations and named entities.</p> <p><strong>The printed dataset is unchanged compared to version 2.</strong></p> <p><strong>The performances of the models described in [1] and [3] are detailled in the Leaderboard section.</strong></p> <h2><strong>General information</strong></h2> <p>The <strong>EXO-POPP project</strong> aims to establish a comprehensive database comprising 300,000 marriage records from Paris and its suburbs, spanning the years 1880 to 1940, which are preserved in over 130,000 scans of double pages. Each marriage record may encompass up to 118 distinct types of information that require extraction from plain text. The M-POPP corpus (which stands for Marriage records of the POPulation of Paris) is the corpus on which the EXO-POPP project focuses. This corpus was built by gathering the marriage records of Paris and its suburb regions (Hauts- de-Seine, Seine-Saint-Denis, Val-de-Marne).</p> <p>The M-POPP corpus are a subset of the M-POPP database with annotations for full-page text recognition and named entity recognition/information extraction from both handwritten and printed documents. The first dataset comprises handwritten marriage records, while the second dataset consists of typewritten marriage records. It should be noted that even in typewritten marriage records, some handwritten information occurs, especially concerning the names of the spouses, and notes in the margin.<br>The dataset contains single-page images obtained from the original scans of double pages via page segmentation.</p> <p>The structure of the files is the following:</p> <ul> <li>handwritten: <em>the handwritten dataset</em><br> <ul> <li>images: <em>images of the dataset divided following the split used in [1]</em><br> <ul> <li>train</li> <li>valid</li> <li>test</li> </ul> </li> <li>labels: <em>labels for joint handwritten text recognition and information extraction for each encoding tested in [1]</em></li> </ul> </li> <li>printed: <em>the printed dataset</em><br> <ul> <li>images: <em>images of the dataset divided following the split used in [1]</em><br> <ul> <li>train</li> <li>valid</li> <li>test</li> </ul> </li> <li>labels: <em>labels for joint handwritten text recognition and information extraction for each encoding tested in [1]</em></li> </ul> </li> <li>encoding-2-to-encoding-5.json: <em>a JSON file giving the correspondence between the symbols of encoding 2 and encoding 5.</em></li> </ul> <p> </p> <p>Table 1: Details on the split of the handwritten dataset.</p> <table> <tbody> <tr> <td> </td> <td>Train</td> <td>Validation</td> <td>Test</td> </tr> <tr> <td>Pages</td> <td>250</td> <td>32</td> <td>32</td> </tr> <tr> <td>Acts</td> <td>344</td> <td>51</td> <td>53</td> </tr> <tr> <td>Named entities</td> <td>16727</td> <td>2223</td> <td>2517</td> </tr> </tbody> </table> <p> </p> <p>Table 2: Details on the split of the printed dataset.</p> <table> <tbody> <tr> <td> </td> <td>Train</td> <td>Validation</td> <td>Test</td> </tr> <tr> <td>Pages</td> <td>116</td> <td>14</td> <td>13</td> </tr> <tr> <td>Acts</td> <td>363</td> <td>43</td> <td>30</td> </tr> <tr> <td>Named entities</td> <td>22036</td> <td>2559</td> <td>2405</td> </tr> </tbody> </table> <p> </p> <p>Table 3: Average annotation statistics per act for the two M-POPP datasets.</p> <table> <tbody> <tr> <td>Dataset</td> <td># of characters</td> <td># of words</td> <td># of named entities</td> </tr> <tr> <td>Handwritten</td> <td>1519</td> <td>231</td> <td>48</td> </tr> <tr> <td>Printed</td> <td>1328</td> <td>200</td> <td>60</td> </tr> </tbody> </table> <p> </p> <h3><strong>Document structure Annotation</strong></h3> <p>We employ the procedure applied in [2], which involves adding opening and closing tags to the character set for each text block we want to recognize.<br>In total, we define four types of text blocks.</p> <ul> <li>Block A is located in the margin and contains the last names of the married couple, possibly with their first names and the date of the marriage.</li> <li>Block B is the body of the text. Block B is the one that contains most of the information to be extracted.</li> <li>Block C is optional and corresponds to marginal notes used in various cases, such as the mention of a divorce or a correction made to the act.</li> <li>Block D corresponds to a set containing a block A and a block B, optionally with one or more blocks C.</li> </ul> <p> </p> <h3><strong>Information Extraction annotation</strong></h3> <p>The dataset contains 118 information categories. As explained in the paper, we broke down the named entities into sub-elements pertaining to 4 hierarchical levels, which reduces the total number of categories to 23 instead of 118. Notice that level 1, 2, and 3 categories do not encode named entities but rather the relations that may occur between some lower level categories for example: (day, birth, husband) encodes the fact that the annotated piece of text is the date of birth of the husband. </p> <p>For these datasets, we chose to represent these hierarchical elements with emojis. For instance, the information <em>first name</em> is represented by the emoji 💬.<br>The meaning of each emoji can be found in Table 4. To determine the best way to encode named entities in the ground truth, we compared in [1] 5 types of encoding. To illustrate these encodings, let’s take for instance <em>Louis Alexandre MOUDEL</em> that we define as the father of the bride, where <em>Louis Alexandre</em> are his two first names, and <em>Moudel</em> is his last name. </p> <p>1) Single separate tags before each word: In this approach, each level of information is indicated by a dedicated tag, and the tags are placed before the word they encode information for. With this encoding, the ground truth for the example would be:</p> <p>💬👴👰Louis 💬👴👰Alexandre 🗨️👴👰MOUDEL</p> <p>2) Single separate tags after each word: Similar to the previous approach, except here the tags are placed after the word. With this encoding the previous example becomes:</p> <p>Louis👰👴💬 Alexandre👰👴💬 MOUDEL👰👴🗨️</p> <p>3) Open & close separate tags: Here, each word presenting information to be extracted is surrounded by one or more opening and closing tags, where each tag encodes a level of information. So the example would be as:</p> <p><👰> <👴> <💬> Louis <\💬> <\👴> <\👰><br><👰> <👴> <💬> Alexandre <\💬> <\👴> <\👰><br><👰> <👴> <🗨️> MOUDEL <\🗨️> <\👴> <\👰></p> <p>4) Nested open & close separate tags: Similar to the previous approach, but this time a tag is closed only when the encoded information is no longer the same for that level of information. We can see in the example below that the tags for wife and father are only used twice.</p> <p><👰> <👴> <💬> Louis Alexandre <\💬> <🗨️> MOUDEL <\🗨️></p> <p>5) Single combined tags after each word: In the last approach, one tag encodes all the hierarchical levels constituting information. The tags are located after the word they encode information for. </p> <p>Louis<wife_father_first_name> Alexandre<wife_father_first_name> MOUDEL<wife_father_family_name></p> <p>NB: In the labels file of encoding 5, the information are still encoded with emojis but the chosen emojis do not have a semantic meaning due to the number of information categories to be represented. The correspondence between the symbols of encoding 2 and encoding 5 can be found in the file <em>encoding-2-to-encoding-5.json</em>.<em><br></em></p> <p> </p> <p>Table 4: Details of the hierarchical breakdown of named entities. Each tag is placed in the corresponding hierarchical level and associated with the emoji representing it.</p> <table> <tbody> <tr> <td>Level</td> <td>Tags</td> <td> </td> <td> </td> <td> </td> </tr> <tr> <td>1</td> <td>Administrative 📖</td> <td> <pre>Husband<code> 👨</code></pre> </td> <td>Wife 👰</td> <td>Witness 🥸</td> </tr> <tr> <td>2</td> <td>Father 👴</td> <td>Mother 👵</td> <td>Ex-husband 💔</td> <td> </td> </tr> <tr> <td>3</td> <td>Birth 🏥</td> <td>Residence 🏠</td> <td> </td> <td> </td> </tr> <tr> <td>4</td> <td>First name 💬</td> <td>Family name 🗨️</td> <td>Age ⌛</td> <td>Occupation 🔧</td> </tr> <tr> <td>5</td> <td>Street number 🔟</td> <td>Street type 🛣</td> <td>Street name 🔠</td> <td>City 🌆</td> </tr> <tr> <td> </td> <td>Department 🗺</td> <td>Country 🗺</td> <td>Day 🌞</td> <td>Month 📅</td> </tr> <tr> <td> </td> <td>Year 🗓</td> <td>Hour ⏰</td> <td>Minute ⏱</td> <td> </td> </tr> </tbody> </table> <p> </p> <h2><strong>Leaderboard</strong></h2> <h3><strong>Results on M-POPP handwritten</strong></h3> <p><strong>HTR<br></strong></p> <p>The following table contains the current leaderboard of M-POPP v3 for HTR on the handwritten dataset.</p> <p>In this configuration, layout block C is not considered.</p> <p>These results for DAN NER are given using the named entity encoding format 5 described above.</p> <p>HTR stands for Handwritten Text Recognition and HTR+IE for combined Handwritten Text Recognition and Information Extraction.</p> <p>Metrics are expressed in percentages.</p> <table> <tbody> <tr> <td>Method</td> <td>CER</td> <td>WER</td> <td>LOER</td> <td>mAP CER</td> </tr> <tr> <td>DAN - HTR [1]</td> <td>7.21</td> <td>16.42</td> <td>5.35</td> <td>83.03</td> </tr> <tr> <td>DAN NER - HTR + IE [1]</td> <td>6.52</td> <td>14.80</td> <td>3.79</td> <td>86.29</td> </tr> <tr> <td>DANIEL - HTR [3]</td> <td>5.72</td> <td>14.08</td> <td>1.34</td> <td>89.28</td> </tr> </tbody> </table> <p> </p> <p><strong>NER</strong></p> <p>The following table contains the current leaderboard of M-POPP v3 for NER on the handwritten dataset.</p> <p>In this configuration, layout block C is not considered.</p> <p>These results are given using the named entity encoding format 5 described above.</p> <p>Metrics are expressed in percentages.</p> <table> <tbody> <tr> <td>Method</td> <td>F1</td> </tr> <tr> <td>DAN NER [1]</td> <td>76.37</td> </tr> <tr> <td>DANIEL [3]</td> <td>76.37</td> </tr> </tbody> </table> <h3> </h3> <h3><strong>Results on M-POPP printed</strong></h3> <p><strong>HTR</strong></p> <p>The following table contains the current leaderboard of this version for TR on the printed dataset.</p> <p>The labels, and therefore the results, are the same as in version 2.</p> <p>In this configuration, only layout block B is considered.</p> <p>These results for DAN NER are given using the named entity encoding format 5 described above.</p> <p>TR stands for Text Recognition and TR+IE for combined Text Recognition and Information Extraction.</p> <p>Metrics are expressed in percentages.</p> <p><strong>NB:</strong> The model DANIEL from [3] was not evaluated on the printed dataset.</p> <table> <tbody> <tr> <td>Method</td> <td>CER</td> <td>WER</td> </tr> <tr> <td>DAN - TR [1]</td> <td>0.88</td> <td>3.17</td> </tr> <tr> <td>DAN NER - TR + IE [1]</td> <td>1.54</td> <td>3.55</td> </tr> </tbody> </table> <p> </p> <p><strong>NER</strong></p> <p>The following table contains the current leaderboard of this version for NER on the printed dataset.</p> <p>The labels, and therefore the results, are the same as in version 2.</p> <p>In this configuration, only layout block B is considered.</p> <p>These results are given using the named entity encoding format 5 described above.</p> <p>Metrics are expressed in percentages.</p> <p><strong>NB:</strong> The model DANIEL from [3] was not evaluated on the printed dataset.</p> <table> <tbody> <tr> <td>Method</td> <td>F1</td> </tr> <tr> <td>DAN NER [1]</td> <td>93.04</td> </tr> </tbody> </table> <h2> </h2> <h2><strong>Citation Request</strong></h2> <p>If you publish material based on this database, we request you to include a reference to the paper <code><a href="https://arxiv.org/abs/2404.19329">T. Constum, L. Preel, T. Paquet, P. Tranouez, S. Brée, End-to-end information extraction in handwritten documents: Understanding Paris marriage records from 1880 to 1940, International Conference on Document Analysis and Recognition (ICDAR), Athens, Greece, 2024</a>.</code></p> <p> </p> <h2><strong>Bibliography</strong></h2> <p><a href="https://arxiv.org/abs/2404.19329">1: T. Constum, L. Preel, T. Paquet, P. Tranouez, S. Brée: End-to-end information extraction in handwritten documents: Understanding Paris marriage records from 1880 to 1940, International Conference on Document Analysis and Recognition (ICDAR), Athens, Greece, 2024.</a></p> <p>2: D.Coquenet, C. Chatelain, T. Paquet: DAN: a Segmentation-free Document Attention Network for Handwritten Document Recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence pp. 1–17 (2023).</p> <p><a href="https://hal.science/hal-04555188/">3: T. Constum, T. Paquet, P. Tranouez: DANIEL: A fast Document Attention Network for Information Extraction and Labelling of handwritten documents, preprint, 2024 </a></p>
NorHand / Dataset for Handwritten Text Recognition in Norwegian
<p>The dataset comprises Norwegian letter and diary line images and text from 19th and early 20th century.</p>
Dataset and scripts used in paper "Exploiting Existing Modern Transcripts for Historical Handwritten Text Recognition"
<p>This package contains the dataset and scripts used in the research paper titled "Exploiting Existing Modern Transcripts for Historical Handwritten Text Recognition" that was published in the ICFHR 2016 conference proceedings.</p> <p>Unfortunately not all of the code used in the experiments is open source, so what is provided is not enough to really reproduce the results. Nevertheless, it should be useful to understand in more detail how the experimentation was performed. The scripts make use of the library for handwritten text recognition that is available at:</p> <p>https://github.com/mauvilsa/htrsh</p> <p>The dataset is in PRImA Page XML format (http://www.primaresearch.org/tools). To visualize the what is contained in the xml files on top of the images, the following open source tool can be used:</p> <p>https://github.com/mauvilsa/nw-page-editor</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.