Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
57
datasets available to search
ShareScore release 0.9.0
Dataset results
57 results for “handwritten”
ICFHR 2020 Competition on Image Retrieval for Historical Handwritten Fragments (HisFrag20) Dataset
<p>This competition investigates the performance of large-scale retrieval of historical document fragments based on writer recognition. The analysis of historic fragments is a difficult challenge commonly solved by trained humanists.<br> We focus on the task of automatic image retrieval to simulate common scenarios of humanities research, such as fragment or writer retrieval. Therefore, we created a large dataset consisting of more than 120000 fragments.<br> The goal is then to find similar patches of the same page or manuscript. contains ~100 000 fragments using the Historical-IR19 as base dataset, they should all contain some text, however some fragments are quite small.</p> <p>Training-set: contains ~100 000 fragments using the Historical-IR19 as base dataset, they should all contain some text, however some fragments are quite small.</p> <p>Test-set: contains about 20 000 new fragments</p> <p>Naming-convention: WID_PID_FID.jpg , where WID=writer id, PID: page id, FID= fragment id</p> <p>For more information visit: <a href="https://lme.tf.fau.de/research/competitions/hisfragir20/">https://lme.tf.fau.de/research/competitions/hisfragir20/</a></p>
EPARCHOS - Historical Greek handwritten document dataset
<p>The dataset originates from a Greek handwritten codex that dates from around 1500-1530. This is the subset of the codex British Museum Addit. 6791, written by two hands, one by Antonius Eparchos and the other by Camillos Zanettus (ff. 104r-174v) and delivers texts by Hierocles (In Aureum carmen), Matthaeus Blastares (Collectio alphabetica) and, notably, texts by Michael Psellos (De omnifaria doctrina). The writing delivers the most important abbreviations, logograms and conjunctions, which are cited in virtually every Greek minuscule handwritten codex from the years of the manuscript transliteration and the prevalence of the minuscule script (9th century) to the post-Byzantine years. This dataset consists of 120 scanned handwritten text pages, containing 9285 lines of text, 18809 words (6787 unique words). For each page, a PageXML is provided containing the following groundtruth:</p> <ol> <li>Text region polygon coordinates</li> <li>Text line polygon coordinates with the corresponding transcription text</li> <li>Word polygon coordinated with the corresponding transcription text</li> </ol>
A benchmark dataset for Manipuri Meetei-Mayek handwritten character recognition
<p>A benchmark dataset is always required for any classification or recognition system. To the best of our knowledge, no benchmark dataset exists for handwritten character recognition of Manipuri Meetei-Mayek script in <strong>public domain</strong> so far. Manipuri, also referred to as Meeteilon or sometimes Meiteilon, is a Sino-Tibetan language and also one of the Eight Scheduled languages of Indian Constitution. It is the official language and lingua franca of the southeastern Himalayan state of Manipur, in northeastern India. This language is also used by a significant number of people as their communicating language over the north-east India, and some parts of Bangladesh and Myanmar. It is the most widely spoken language in Northeast India after Bengali and Assamese languages. In this work, we introduce a handwritten Manipuri Meetei-Mayek character dataset which consists of more than 5000 data samples which were collected from a diverse population group that belongs to different age groups (from 4 years to 60 years), genders, educational backgrounds, occupations, communities from three different districts of Manipur, India (Imphal East District, Thoubal District and Kangpokpi District) during March and April 2019. Each individual was asked to write down all the Manipuri characters on one A4-size paper. The recorded responses are scanned with the help of a scanner and then each character is manually segmented from the scanned images. This dataset consists of segmented scanned images of handwritten Manipuri Meetei-Mayek characters (Mapi Mayek, Lonsum Mayek, Cheitap Mayek, Cheising Mayek, Khutam Mayek) of size 128X128 pixels in .JPG format as well as in .MAT format.</p>
HWRT database of handwritten symbols
<p>The HWRT database of handwritten symbols contains on-line data of handwritten symbols such as all alphanumeric characters, arrows, greek characters and mathematical symbols like the integral symbol.</p> <p>The database can be downloaded in form of bzip2-compressed tar files. Each tar file contains:</p> <ul> <li>symbols.csv: A CSV file with the rows symbol_id, latex, training_samples, test_samples. The symbol id is an integer, the row latex contains the latex code of the symbol, the rows training_samples and test_samples contain integers with the number of labeled data.</li> <li>train-data.csv: A CSV file with the rows symbol_id, user_id, user_agent and data.</li> <li>test-data.csv: A CSV file with the rows symbol_id, user_id, user_agent and data.</li> </ul> <p>All CSV files use ";" as delimiter and "'" as quotechar. The data is given in YAML format as a list of lists of dictinaries. Each dictionary has the keys "x", "y" and "time". (x,y) are coordinates and time is the UNIX time.</p> <p> </p> <p>About 90% of the data was made available by Daniel Kirsch via github.com/kirel/detexify-data. Thank you very much, Daniel!</p>
Test A for the ICDAR2017 Competition on Handwritten Text Recognition on the READ Dataset (ICDAR2017 HTR)
<p><strong>Test A. </strong>A batch of page images annotated with baselines.</p>
Datasets for "Reading Order Independent Metrics for Information Extraction in Handwritten Documents"
<p>This repository includes the five datasets used for our paper entitled <em>Reading Order Independent Metrics for Information Extraction in Handwritten Documents</em>, in which we compare various metrics to evaluate end-to-end information extraction from scanned documents.</p> <h2>Datasets</h2> <p>Five datasets are released following the BIO format:</p> <ul> <li>IAM</li> <li>Simara</li> <li>POPP</li> <li>Esposalles</li> <li>French Military Records</li> </ul> <p>For each dataset, we provide the following data (on test sets):</p> <ul> <li>Ground truth annotations (<code>gt/</code>)</li> <li>Automatic predictions (<code>dan/</code>)</li> <li>Automatic predictions with entities appearing in random order (<code>dan_shuffled/</code>)</li> </ul> <p>The data is organized as follows:</p> <p><code>├── Dataset name/</code><br><code>│ ├── gt/</code><br><code>│ ├── dan/</code><br><code>│ └── dan_shuffled/</code></p> <h2>Metrics</h2> <p>To install the <a href="https://pypi.org/project/ie-eval/"><code>ie-eval</code></a> package, run <code>pip install ie-eval</code>.</p> <p>To compute all metrics on a specific dataset, run:<br><br><code>ie-eval all --label-dir IAM_paragraph/gt/ --prediction-dir IAM_paragraph/dan/</code><br><br></p> <p>To learn more about the various options, use the <code>--help</code> argument or read the <a href="https://ie-eval-ner-metrics-050f40e80b04480e2310d39ad338de778f6bec80e18.pages.teklia.com/">documentation</a>.</p> <p> </p>
A Benchmark Dataset for Manipuri Meetei-Mayek Handwritten Character Recognition
<p>A benchmark dataset is always required for any classification or recognition system. To the best of our knowledge, no benchmark dataset exists for handwritten character recognition of Manipuri Meetei-Mayek script in <em><strong>public domain</strong></em> so far. Manipuri, also referred to as Meeteilon or sometimes Meiteilon, is a Sino-Tibetan language and also one of the Eight Scheduled languages of Indian Constitution. It is the official language and lingua franca of the southeastern Himalayan state of Manipur, in northeastern India. This language is also used by a significant number of people as their communicating language over the north-east India, and some parts of Bangladesh and Myanmar. It is the most widely spoken language in Northeast India after Bengali and Assamese languages. In this work, we introduce a handwritten Manipuri Meetei-Mayek character dataset which consists of more than 5000 data samples which were collected from a diverse population group that belongs to different age groups (from 4 years to 60 years), genders, educational backgrounds, occupations, communities from three different districts of Manipur, India (Imphal East District, Thoubal District and Kangpokpi District) during March and April 2019. Each individual was asked to write down all the Manipuri characters on one A4-size paper. The recorded responses are scanned with the help of a scanner and then each character is manually segmented from the scanned images. This dataset consists of segmented scanned images of handwritten Manipuri Meetei-Mayek characters (Mapi Mayek, Lonsum Mayek, Cheitap Mayek, Cheising Mayek, Khutam Mayek) of size 128X128 pixels in .JPG format as well as in .MAT format.</p>
Handwritten newspapers used in Turunen, R. (2021). Shades of Red: Evolution of the Political Language of Finnish Socialism from the Nineteenth Century until the Civil War of 1918
<p>This dataset contains XLSX files of five different handwritten newspapers: <em>Kuritus </em>(1909–1911), <em>Palveliatar</em> (1907–1913, 1917), <em>Yritys</em> (1915–1917), <em>Tehtaalainen</em> (1908–1914, 1917) and <em>Nuija </em>(1899–1903, 1907–1909, 1912, 1914–1915).</p> <p>The original sources and the coding scheme for the XLSX files are described in:</p> <p>Turunen, R. (2021). <em>Shades of Red: Evolution of the Political Language of Finnish Socialism from the Nineteenth Century until the Civil War of 1918</em>. The Finnish Society for Labour History.</p>
A benchmark dataset for Manipuri Meetei-Mayek handwritten character recognition
Open the record for dataset details and reuse information.
HASY - Handwritten Symbol database
<p>HASY contains 32px x 32px images of 369 symbol classes. In total, HASY contains over 150,000 instances of handwritten symbols.</p>
HASYv2 - Handwritten Symbol database
<p>HASY contains 32px x 32px images of 369 symbol classes. In total, HASY contains over 150,000 instances of handwritten symbols.<br> <br> See "The HASYv2 dataset" paper (https://arxiv.org/abs/1701.08380) for more information.</p>
Relationship between Chinese character conformational features and the legibility of handwritten Chinese characters for CFL beginners
Open the record for dataset details and reuse information.
M-POPP datasets: Datasets for full page text recognition and information extraction from French handwritten and printed marriage records
<h1><strong>M-POPP datasets</strong></h1> <p>This repository contains 2 datasets created within the <strong>EXO-POPP project</strong> (<a href="https://exopopp.hypotheses.org/">Optical EXtraction of handwritten named entities for marriage records of the POPulation of Paris</a>) for the task of text recognition and information extraction. These datasets have been published in <a href="https://arxiv.org/abs/2404.19329"><code>End-to-end information extraction in handwritten documents: Understanding Paris marriage records from 1880 to 1940</code></a> <code>[1]</code>at ICDAR 2024.</p> <p><strong>This version contains the labels for Handwritten Text Recognition and Handwritten Text Recognition + Information Extraction as used in our new paper "<a href="https://hal.science/hal-04555188">DANIEL: A fast Document Attention Network for Information Extraction and Labelling of handwritten documents</a>" [3].</strong></p> <p><strong>This version makes corrections to the handwritten dataset. </strong>More precisely, it corrects a few errors in transcription annotations and named entities.</p> <p><strong>The printed dataset is unchanged compared to version 2.</strong></p> <p><strong>The performances of the models described in [1] and [3] are detailled in the Leaderboard section.</strong></p> <h2><strong>General information</strong></h2> <p>The <strong>EXO-POPP project</strong> aims to establish a comprehensive database comprising 300,000 marriage records from Paris and its suburbs, spanning the years 1880 to 1940, which are preserved in over 130,000 scans of double pages. Each marriage record may encompass up to 118 distinct types of information that require extraction from plain text. The M-POPP corpus (which stands for Marriage records of the POPulation of Paris) is the corpus on which the EXO-POPP project focuses. This corpus was built by gathering the marriage records of Paris and its suburb regions (Hauts- de-Seine, Seine-Saint-Denis, Val-de-Marne).</p> <p>The M-POPP corpus are a subset of the M-POPP database with annotations for full-page text recognition and named entity recognition/information extraction from both handwritten and printed documents. The first dataset comprises handwritten marriage records, while the second dataset consists of typewritten marriage records. It should be noted that even in typewritten marriage records, some handwritten information occurs, especially concerning the names of the spouses, and notes in the margin.<br>The dataset contains single-page images obtained from the original scans of double pages via page segmentation.</p> <p>The structure of the files is the following:</p> <ul> <li>handwritten: <em>the handwritten dataset</em><br> <ul> <li>images: <em>images of the dataset divided following the split used in [1]</em><br> <ul> <li>train</li> <li>valid</li> <li>test</li> </ul> </li> <li>labels: <em>labels for joint handwritten text recognition and information extraction for each encoding tested in [1]</em></li> </ul> </li> <li>printed: <em>the printed dataset</em><br> <ul> <li>images: <em>images of the dataset divided following the split used in [1]</em><br> <ul> <li>train</li> <li>valid</li> <li>test</li> </ul> </li> <li>labels: <em>labels for joint handwritten text recognition and information extraction for each encoding tested in [1]</em></li> </ul> </li> <li>encoding-2-to-encoding-5.json: <em>a JSON file giving the correspondence between the symbols of encoding 2 and encoding 5.</em></li> </ul> <p> </p> <p>Table 1: Details on the split of the handwritten dataset.</p> <table> <tbody> <tr> <td> </td> <td>Train</td> <td>Validation</td> <td>Test</td> </tr> <tr> <td>Pages</td> <td>250</td> <td>32</td> <td>32</td> </tr> <tr> <td>Acts</td> <td>344</td> <td>51</td> <td>53</td> </tr> <tr> <td>Named entities</td> <td>16727</td> <td>2223</td> <td>2517</td> </tr> </tbody> </table> <p> </p> <p>Table 2: Details on the split of the printed dataset.</p> <table> <tbody> <tr> <td> </td> <td>Train</td> <td>Validation</td> <td>Test</td> </tr> <tr> <td>Pages</td> <td>116</td> <td>14</td> <td>13</td> </tr> <tr> <td>Acts</td> <td>363</td> <td>43</td> <td>30</td> </tr> <tr> <td>Named entities</td> <td>22036</td> <td>2559</td> <td>2405</td> </tr> </tbody> </table> <p> </p> <p>Table 3: Average annotation statistics per act for the two M-POPP datasets.</p> <table> <tbody> <tr> <td>Dataset</td> <td># of characters</td> <td># of words</td> <td># of named entities</td> </tr> <tr> <td>Handwritten</td> <td>1519</td> <td>231</td> <td>48</td> </tr> <tr> <td>Printed</td> <td>1328</td> <td>200</td> <td>60</td> </tr> </tbody> </table> <p> </p> <h3><strong>Document structure Annotation</strong></h3> <p>We employ the procedure applied in [2], which involves adding opening and closing tags to the character set for each text block we want to recognize.<br>In total, we define four types of text blocks.</p> <ul> <li>Block A is located in the margin and contains the last names of the married couple, possibly with their first names and the date of the marriage.</li> <li>Block B is the body of the text. Block B is the one that contains most of the information to be extracted.</li> <li>Block C is optional and corresponds to marginal notes used in various cases, such as the mention of a divorce or a correction made to the act.</li> <li>Block D corresponds to a set containing a block A and a block B, optionally with one or more blocks C.</li> </ul> <p> </p> <h3><strong>Information Extraction annotation</strong></h3> <p>The dataset contains 118 information categories. As explained in the paper, we broke down the named entities into sub-elements pertaining to 4 hierarchical levels, which reduces the total number of categories to 23 instead of 118. Notice that level 1, 2, and 3 categories do not encode named entities but rather the relations that may occur between some lower level categories for example: (day, birth, husband) encodes the fact that the annotated piece of text is the date of birth of the husband. </p> <p>For these datasets, we chose to represent these hierarchical elements with emojis. For instance, the information <em>first name</em> is represented by the emoji 💬.<br>The meaning of each emoji can be found in Table 4. To determine the best way to encode named entities in the ground truth, we compared in [1] 5 types of encoding. To illustrate these encodings, let’s take for instance <em>Louis Alexandre MOUDEL</em> that we define as the father of the bride, where <em>Louis Alexandre</em> are his two first names, and <em>Moudel</em> is his last name. </p> <p>1) Single separate tags before each word: In this approach, each level of information is indicated by a dedicated tag, and the tags are placed before the word they encode information for. With this encoding, the ground truth for the example would be:</p> <p>💬👴👰Louis 💬👴👰Alexandre 🗨️👴👰MOUDEL</p> <p>2) Single separate tags after each word: Similar to the previous approach, except here the tags are placed after the word. With this encoding the previous example becomes:</p> <p>Louis👰👴💬 Alexandre👰👴💬 MOUDEL👰👴🗨️</p> <p>3) Open & close separate tags: Here, each word presenting information to be extracted is surrounded by one or more opening and closing tags, where each tag encodes a level of information. So the example would be as:</p> <p><👰> <👴> <💬> Louis <\💬> <\👴> <\👰><br><👰> <👴> <💬> Alexandre <\💬> <\👴> <\👰><br><👰> <👴> <🗨️> MOUDEL <\🗨️> <\👴> <\👰></p> <p>4) Nested open & close separate tags: Similar to the previous approach, but this time a tag is closed only when the encoded information is no longer the same for that level of information. We can see in the example below that the tags for wife and father are only used twice.</p> <p><👰> <👴> <💬> Louis Alexandre <\💬> <🗨️> MOUDEL <\🗨️></p> <p>5) Single combined tags after each word: In the last approach, one tag encodes all the hierarchical levels constituting information. The tags are located after the word they encode information for. </p> <p>Louis<wife_father_first_name> Alexandre<wife_father_first_name> MOUDEL<wife_father_family_name></p> <p>NB: In the labels file of encoding 5, the information are still encoded with emojis but the chosen emojis do not have a semantic meaning due to the number of information categories to be represented. The correspondence between the symbols of encoding 2 and encoding 5 can be found in the file <em>encoding-2-to-encoding-5.json</em>.<em><br></em></p> <p> </p> <p>Table 4: Details of the hierarchical breakdown of named entities. Each tag is placed in the corresponding hierarchical level and associated with the emoji representing it.</p> <table> <tbody> <tr> <td>Level</td> <td>Tags</td> <td> </td> <td> </td> <td> </td> </tr> <tr> <td>1</td> <td>Administrative 📖</td> <td> <pre>Husband<code> 👨</code></pre> </td> <td>Wife 👰</td> <td>Witness 🥸</td> </tr> <tr> <td>2</td> <td>Father 👴</td> <td>Mother 👵</td> <td>Ex-husband 💔</td> <td> </td> </tr> <tr> <td>3</td> <td>Birth 🏥</td> <td>Residence 🏠</td> <td> </td> <td> </td> </tr> <tr> <td>4</td> <td>First name 💬</td> <td>Family name 🗨️</td> <td>Age ⌛</td> <td>Occupation 🔧</td> </tr> <tr> <td>5</td> <td>Street number 🔟</td> <td>Street type 🛣</td> <td>Street name 🔠</td> <td>City 🌆</td> </tr> <tr> <td> </td> <td>Department 🗺</td> <td>Country 🗺</td> <td>Day 🌞</td> <td>Month 📅</td> </tr> <tr> <td> </td> <td>Year 🗓</td> <td>Hour ⏰</td> <td>Minute ⏱</td> <td> </td> </tr> </tbody> </table> <p> </p> <h2><strong>Leaderboard</strong></h2> <h3><strong>Results on M-POPP handwritten</strong></h3> <p><strong>HTR<br></strong></p> <p>The following table contains the current leaderboard of M-POPP v3 for HTR on the handwritten dataset.</p> <p>In this configuration, layout block C is not considered.</p> <p>These results for DAN NER are given using the named entity encoding format 5 described above.</p> <p>HTR stands for Handwritten Text Recognition and HTR+IE for combined Handwritten Text Recognition and Information Extraction.</p> <p>Metrics are expressed in percentages.</p> <table> <tbody> <tr> <td>Method</td> <td>CER</td> <td>WER</td> <td>LOER</td> <td>mAP CER</td> </tr> <tr> <td>DAN - HTR [1]</td> <td>7.21</td> <td>16.42</td> <td>5.35</td> <td>83.03</td> </tr> <tr> <td>DAN NER - HTR + IE [1]</td> <td>6.52</td> <td>14.80</td> <td>3.79</td> <td>86.29</td> </tr> <tr> <td>DANIEL - HTR [3]</td> <td>5.72</td> <td>14.08</td> <td>1.34</td> <td>89.28</td> </tr> </tbody> </table> <p> </p> <p><strong>NER</strong></p> <p>The following table contains the current leaderboard of M-POPP v3 for NER on the handwritten dataset.</p> <p>In this configuration, layout block C is not considered.</p> <p>These results are given using the named entity encoding format 5 described above.</p> <p>Metrics are expressed in percentages.</p> <table> <tbody> <tr> <td>Method</td> <td>F1</td> </tr> <tr> <td>DAN NER [1]</td> <td>76.37</td> </tr> <tr> <td>DANIEL [3]</td> <td>76.37</td> </tr> </tbody> </table> <h3> </h3> <h3><strong>Results on M-POPP printed</strong></h3> <p><strong>HTR</strong></p> <p>The following table contains the current leaderboard of this version for TR on the printed dataset.</p> <p>The labels, and therefore the results, are the same as in version 2.</p> <p>In this configuration, only layout block B is considered.</p> <p>These results for DAN NER are given using the named entity encoding format 5 described above.</p> <p>TR stands for Text Recognition and TR+IE for combined Text Recognition and Information Extraction.</p> <p>Metrics are expressed in percentages.</p> <p><strong>NB:</strong> The model DANIEL from [3] was not evaluated on the printed dataset.</p> <table> <tbody> <tr> <td>Method</td> <td>CER</td> <td>WER</td> </tr> <tr> <td>DAN - TR [1]</td> <td>0.88</td> <td>3.17</td> </tr> <tr> <td>DAN NER - TR + IE [1]</td> <td>1.54</td> <td>3.55</td> </tr> </tbody> </table> <p> </p> <p><strong>NER</strong></p> <p>The following table contains the current leaderboard of this version for NER on the printed dataset.</p> <p>The labels, and therefore the results, are the same as in version 2.</p> <p>In this configuration, only layout block B is considered.</p> <p>These results are given using the named entity encoding format 5 described above.</p> <p>Metrics are expressed in percentages.</p> <p><strong>NB:</strong> The model DANIEL from [3] was not evaluated on the printed dataset.</p> <table> <tbody> <tr> <td>Method</td> <td>F1</td> </tr> <tr> <td>DAN NER [1]</td> <td>93.04</td> </tr> </tbody> </table> <h2> </h2> <h2><strong>Citation Request</strong></h2> <p>If you publish material based on this database, we request you to include a reference to the paper <code><a href="https://arxiv.org/abs/2404.19329">T. Constum, L. Preel, T. Paquet, P. Tranouez, S. Brée, End-to-end information extraction in handwritten documents: Understanding Paris marriage records from 1880 to 1940, International Conference on Document Analysis and Recognition (ICDAR), Athens, Greece, 2024</a>.</code></p> <p> </p> <h2><strong>Bibliography</strong></h2> <p><a href="https://arxiv.org/abs/2404.19329">1: T. Constum, L. Preel, T. Paquet, P. Tranouez, S. Brée: End-to-end information extraction in handwritten documents: Understanding Paris marriage records from 1880 to 1940, International Conference on Document Analysis and Recognition (ICDAR), Athens, Greece, 2024.</a></p> <p>2: D.Coquenet, C. Chatelain, T. Paquet: DAN: a Segmentation-free Document Attention Network for Handwritten Document Recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence pp. 1–17 (2023).</p> <p><a href="https://hal.science/hal-04555188/">3: T. Constum, T. Paquet, P. Tranouez: DANIEL: A fast Document Attention Network for Information Extraction and Labelling of handwritten documents, preprint, 2024 </a></p>
NorHand / Dataset for Handwritten Text Recognition in Norwegian
<p>The dataset comprises Norwegian letter and diary line images and text from 19th and early 20th century.</p>
Handwritten Text Production in Adults With Autism
ClinicalTrials.gov study NCT06304701. IPD Sharing: NO. Countries: 1. Publications: 0.
Dataset and scripts used in paper "Exploiting Existing Modern Transcripts for Historical Handwritten Text Recognition"
<p>This package contains the dataset and scripts used in the research paper titled "Exploiting Existing Modern Transcripts for Historical Handwritten Text Recognition" that was published in the ICFHR 2016 conference proceedings.</p> <p>Unfortunately not all of the code used in the experiments is open source, so what is provided is not enough to really reproduce the results. Nevertheless, it should be useful to understand in more detail how the experimentation was performed. The scripts make use of the library for handwritten text recognition that is available at:</p> <p>https://github.com/mauvilsa/htrsh</p> <p>The dataset is in PRImA Page XML format (http://www.primaresearch.org/tools). To visualize the what is contained in the xml files on top of the images, the following open source tool can be used:</p> <p>https://github.com/mauvilsa/nw-page-editor</p>
SCANS Data Set - A Reference Data Set for Handwritten Text Detection
<p># SCANS Data Set<br> <br> A scientific paper with handwritten annotations as reference data set for handwritten text detection.<br> <br> The paper is</p> <pre><code>Sharon Fogel, Hadar Averbuch-Elor, Sarel Cohen, Shai Mazor, Roee Litman; Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogniti on (CVPR), 2020, pp. 4324-4333</code></pre> <p><br> The annotations were made by two people with different pens. The text is randomly taken from "The Fellowship Of The Ring" by JRR Tolkien.<br> <br> The documents in folder `images` were scanned with a Canon Pixima TR4550 with 300DPI. The `labels-` folders contain the labels for paragraph, line, and word level in following JSON format:<br> </p> <pre><code>{ "image_file": "paper-0001.png", # The corresponding image in folder `images` "shape": [ # The size of the image 3495, 2473 ], "properties": [ # A list of all line bounding boxes (for paragraphs, lines, or words) { "type": "HWL", # HWA: handwritten paragraph, HWL: handwritten line, HWW: handwritten word "id": "18", "points": [ # The coordinates of the bounding boxes [ 1243.87096093748, 1515.21674083916 ], [ 2149.01588303892, 1515.21674083916 ], ... ] }, ... ] }</code></pre> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.