Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
57
datasets available to search
ShareScore release 0.9.0
Dataset results
57 results for “handwritten”
Transkribus - Handwritten Text Recognition for Premodern Documents (SIMS 2020 Lightning Talk)
<p>Transkribus is a platform for text recognition and can be used via the Transkribus Expert Software (available after registration: transkribus.eu). Through Transkribus different tools for document analysis and text recognition can be directly applied. The intro demonstrates very briefly how Transkribus can help with regards to premodern documents especially since a variety of pre-trained models are already available: for Latin (prints and handwriting), for early modern vernaculars in French, Dutch, English, and German. For more information go to transkribus.eu.</p> <p>Presented as a Schoenberg Symposium 2020 Lightning Talk</p>
A Public Ground-Truth Dataset for Handwritten Circuit Diagram Images
<p><strong>CGHD</strong></p> <p>This dataset contains images of hand-drawn electrical circuit diagrams as well as accompanying annotation and segmentation ground-truth files. It is intended to train (e.g. ANN) models for extracting electrical graphs from raster graphics.</p> <p><strong>Content</strong></p> <ul> <li><strong>3.269</strong> Annotated Raw Images<br> <ul> <li>31 Main Drafters <ul> <li>12 Circuits per Drafter</li> <li>2 Drawings per Circuit</li> <li>4 Photos per Drawing</li> </ul> </li> <li>Additional Circuit Images provided by TU Dresden (from Real-World Examinations, Drafter 0)</li> <li>Additional Circuit Images provided by RPTU Kaiserslautern-Landau (Drafter -1)</li> <li><strong>248.020 </strong>Bounding Box Annotations</li> <li><strong>40.711</strong> Rotation Annotations</li> <li><strong>1.437</strong> Mirror Annotations</li> <li><strong>85.417</strong> Text String Annotations (equals <strong>93.74%</strong> completeness)<br> <ul> <li><strong>289.850</strong> Text Characters</li> <li><strong>98</strong> Character Types (Upper/Lower Case Latin, Numbers, Special Characters)</li> </ul> </li> </ul> </li> <li><strong>320</strong> Binary Segmentation Maps<br> <ul> <li>Strokes vs. Background</li> <li>Accompanying Polygon Annotation Files</li> <li><strong>22.929</strong> Polygon Annotations</li> </ul> </li> <li><strong>59 </strong>Object Classes</li> <li><strong>Scripts</strong> for Data Loading, Statistics, Consistency Check and Training Preparation</li> </ul>
Ground-Truthed Data Set of Zenon Papyri for Handwritten Text Recognition
<p>Diplomatic transcription of papyri found in the Zenon archive [see <a href="https://en.wikipedia.org/wiki/Zenon_of_Kaunos">en.wikipedia.org/wiki/Zenon_of_Kaunos</a>]</p> <p>Manually prepared as PageXML with Transkribus within <a href="http://d-scribes.philhist.unibas.ch/">D-Scribes</a> project.</p> <p> </p>
Handwritten Cree Syllabics
<p>An open dataset of labelled Cree syllabic handwriting samples.</p> <p>The dataset consists of handwriting sample images with corresponding Tesseract box files containing the label and area of each syllabic.</p>
A Brazilian Portuguese Dataset for Offline Handwritten Text Recognition (BRESSAY)
<p>The BRESSAY dataset comprises images of handwritten essays in Brazilian Portuguese, which present a series of challenges to optical recognition models. These images were sourced from multiple online platforms, limiting our ability to standardize the capture process. Due to these varied sources and the lack of a uniform collection method, the dataset provides a realistic reflection of real-world conditions. Each essay is unique, contributed by different writers, and addresses a specific content topic. Furthermore, the constraints placed on the writers often lead to various handwriting scenarios, including hard-to-read words, connected words, noise, overwriting, and struck-through texts.</p> <h3><strong>Technical Details</strong></h3> <p>The BRESSAY dataset represents a comprehensive collection of handwritten essays in Brazilian Portuguese, offering detailed insights into various handwriting scenarios. It covers a total of 1,000 pages, each contributed by a unique writer, resulting in 1,000 distinct handwriting styles. This aspect of the dataset adds a layer of diversity, which is further emphasized by the total of 4,214 paragraphs, 30,090 lines, and 416,826 words. Regarding unique tokens, we have 41,318 unique words, and 107 unique characters.</p> <h3><strong>Data Structure</strong></h3> <p>The dataset is organized as follows:</p> <ul> <li>data/: Main folder containing segmented essay images <ul> <li>lines/: Images of individual lines <ul> <li> PNG files: Line images</li> <li> TXT files: Transcriptions of lines</li> </ul> </li> <li>pages/: Full page essay images <ul> <li> PNG files: Page images</li> <li> TXT files: Transcriptions of pages</li> </ul> </li> <li>paragraphs/: Images of paragraphs <ul> <li> PNG files: Paragraph images</li> <li> TXT files: Transcriptions of paragraphs</li> </ul> </li> <li>words/: Images of individual words <ul> <li> PNG files: Word images</li> <li> TXT files: Transcriptions of words</li> </ul> </li> </ul> </li> <li>sets/: Contains partition files <ul> <li>test.txt: Names of images in the test set</li> <li>validation.txt: Names of images in the validation set</li> <li>training.txt: Names of images in the training set</li> </ul> </li> </ul> <h3><strong>Dataset Usage and Annotations</strong></h3> <p>Each name in test.txt, validation.txt and training.txt represents the name of the page and all its content (words, lines, paragraphs) must be in the respective partition.</p> <p>Annotations used in the dataset:</p> <ul> <li> <code>##@@???@@##</code>: Superscript text that has become unidentifiable and unreadable.</li> <li> <code>$$@@???@@$$</code>: Subscript text that has become unidentifiable and unreadable.</li> <li> <code>@@???@@</code>: Text that cannot be read or identified due to its illegibility.</li> <li> <code>##--xxx--##</code>: Text that has been added as a superscript and subsequently crossed out, rendering it illegible.</li> <li> <code>$$--xxx--$$</code>: Text that has been added as a subscript and subsequently crossed out, rendering it illegible.</li> <li> <code>--xxx--</code>: Text that has been crossed out in a way that makes it unreadable.</li> <li> <code>##--text--##</code>: Text that has been added as a superscript and subsequently crossed out, but remains legible.</li> <li> <code>$$--text--$$</code>: Text that has been added as a subscript and subsequently crossed out, but remains legible.</li> <li> <code>##text##</code>: Text added as a superscript in the line, typically as a correction or additional note.</li> <li> <code>$$text$$</code>: Text added as a subscript in the line, typically as a correction or additional note.</li> <li> <code>--text--</code>: Text that has been crossed out but remains readable.</li> </ul>
REPUBLIC PageXML ground truth handwritten resolutions States General
<p><strong>Using annotation software provided through the Transkribus Platform we annotated scans, concerning mostly 17th century handwritten documents from the National Archive of the Netherlands, with their textual transcriptions. The resulting ground truth was used to train a machine learning model yielding very accurate results. The ground truth was made available as an open access dataset. </strong></p>
ICDAR 2015 Competition HTRtS: Handwritten Text Recognition on the tranScriptorium Dataset
<p>This dataset comprises the dataset used for the ICDAR 2015 Competition on Handwritten Text Recognition on the tranScriptorium Dataset. The handwritten images for this contest were drawn from the English “Bentham collection” dataset used in the TRAN SCRIPTORIUM project. The selected data has been written by several hands and entails significant variabilities and difficulties regarding the quality of text images, writing styles and crossed-out text. This contest is clearly more difficult than the the first edition both for training and for testing. A portion of the training dataset and the full test dataset were provided in the form of carefully segmented line images, along with the corresponding transcripts. Another portion of the training dataset was provided as raw images and their corresponding transcripts at region level.<br> </p> <p>ICDAR 2015 competition HTRtS: handwritten text recognition on the tranScriptorium dataset<br> JA Sánchez, AH Toselli, V Romero, E Vidal. In International Conference on Document Analysis and Recognition (ICDAR), pp. 1166-1170, 2015.</p>
Relative width and height of handwritten letters. R code.
<p>Heigths and widths of handwritten letters of 21 writers. 500 letters of each writer.</p> <p>R code for data analysis. The codes are in beta phase, the codes have no help or explanations. Units mm/100.</p>
Train-B dataset for ICDAR2017 Competition on Handwritten Text Recognition on the READ Dataset (ICDAR2017 HTR). Batch 1 and Batch 2.
<p>Train-B Dataset. Dataset of pages without any layout or text line information. The corresponding transcripts are provided at page level with line breaks. It has 10k pages, though for convenience it is divided into two 5k page batches. This information is provided in PAGE format. </p> <p>This dataset is complementary to this other dataset:</p> <p>https://zenodo.org/record/439807#.WOIBZ3WLSkA</p> <p>More information at:</p> <p>https://scriptnet.iit.demokritos.gr/competitions/~icdar2017htr/</p> <p> </p>
Train-A dataset for ICDAR2017 Competition on Handwritten Text Recognition on the READ Dataset (ICDAR2017 HTR)
<p>Train-A Dataset of pages with manually revised baselines and the corresponding transcripts associated to them. This batch is small, 50 pages. Please, keep in mind that only the baselines have been manually corrected, The polygons associated to each line have not been manually reviewed. </p> <p>This dataset is complementary to this other dataset:</p> <p>https://zenodo.org/record/439811#.WOIF9HWLSkA</p> <p>More information at:</p> <p>https://scriptnet.iit.demokritos.gr/competitions/~icdar2017htr/</p>
Handwritten ASAP Short Answer Scoring
<p>This dataset is based on the Short Answer Scoring (SAS) dataset of the <a href="https://www.kaggle.com/c/asap-sas">Automated Student Assessment Prize (ASAP)</a>. <br>Although the original dataset was conducted on handwritten content, the scans are not available.<br>To analyze the full pipeline from Handwritten Answers to Automated Scoring, we let students rewrite some answers.<br>The texts used from SAS are from the test set and from the training set. </p>
NorHand v3 / Dataset for Handwritten Text Recognition in Norwegian
<p>This dataset comprises Norwegian letter and diary documents from 19th and early 20th century. It can be used to train Handwritten Text Recognition (HTR) models.</p>
Dzongkha Handwritten Digit Dataset
<p>Dzongkha, the national language of Bhutan, has limited resources available for Natural Language Processing (NLP) tasks because the language is relatively understudied. However, there is no publicly available benchmark dataset for handwritten character identification in the Dzongkha digit script. The dataset contains 1000 images of handwritten Dzongkha digits that are captured using Google Jamboard in JPG format. The image data is assembled from a total of 100 indigenous and non-indigenous people of Bhutan irrespective of age, gender, educational background, etc. In the designed dataset, there are 10 different classes of Dzongkha digits which range from 0 to 9. The labels of these classes are: 0 (༠), 1 (༡), 2 (༢), 3 (༣), 4 (༤), 5 (༥), 6 (༦), 7 (༧), 8 (༨), 9 (༩).</p>
Oficio de Hipotecas de Girona. A dataset of Spanish notarial deeds (18th Century) for Handwritten Text Recognition and Layout Analysis of historical documents.
<p>This dataset is a subset of 596 documents from the <em>Registre d'Hipoteques de Girona</em> of 1769 collection, guarded by the <a href="http://xac.gencat.cat/ca/llista_arxius_comarcals/girones/"><em>Arxiu Històric de Girona</em></a>. This collection, is composed by hundreds of thousands of notarial deeds from the XVIII-XIX century (1768-1862). Sales, redemption of censuses, inheritance and matrimonial chapters are among the most common documentary typologies in the collection.</p> <p>This dataset is composed of more than 23700 text lines written by a single hand, covering more that 50 different topics (documentary typologies) and a vocabulary of more than 2400 different words. The documents are transcribed using the so-called diplomatic criteria. Additionally, transcripts were tagged with <br> extra enriching/complementary information (e.g. expansion of the abbreviations, hyphen marks, etc.). Along with the transcripts the layout of the document is detected and recorded. Pages have been labeled using six different layout regions.</p> <p>The images along with their respective ground-truth was compiled in PAGE compliant XML format<br> by the <a href="http://www2.udg.edu/tabid/11296/Default.aspx"><em>Centre de Recerca d'Història Rural</em></a> and the HTR group of the <a href="https://www.prhlt.upv.es">Pattern Recognition and Human Language Technologies Research Center</a>.</p>
ICDAR 2019 Competition on Image Retrieval for Historical Handwritten Documents [HisIR19] Dataset
<p>This dataset contains the training and test set used in the ICDAR 2019 Competition on Image Retrieval for Historical Handwritten Documents.</p> <p>This competition investigates the performance of large-scale retrieval of historical document images based on<br> writing style. Based on large image data sets provided by cultural heritage institutions and digital libraries, providing<br> a total of 20 000 document images representing about 10 000 writers, divided in three types: writers of (i) manuscript books, (ii) letters, (iii) charters and legal documents. We focus on the task of automatic image retrieval to simulate common scenarios of humanities research, such as writer retrieval.</p> <p>The training data set encompasses images from (i) Letters A, where each writer contributed one or three images; (ii) Manuscripts, where each writer was represented by five consecutive images from a single book.<br> In total, it contains 300 writers contributing one page, 100 writers contributing three pages, and 120 writers contributing five pages resulting in 1200 images of 520 writers.</p> <p>The test data set contains 20 000 images: About 7 500 pages stem from isolated documents (partially anonymous writers, contributing one page each), and about 12 500 pages are from writers that contributed three or five pages.</p> <p> </p> <p>If you use this dataset, please cite:</p> <p>V. Christlein, A. Nicolaou, M. Seuret, D. Stutzmann, A. Maier: "ICDAR 2019 Competition on Image Retrieval for Historical Handwritten Documents", in 15th International Conference on Document Analysis and Recognition, 2019, Sydney, Australia</p> <p> </p>
H̶a̶n̶d̶w̶r̶i̶t̶i̶n̶g̶ - A collection of struck through handwritten word images using various styles
<p># H̶a̶n̶d̶w̶r̶i̶t̶i̶n̶g̶ - A collection of handwritten word images, each struck through using various styles</p> <p>This database may be used for non-commercial research purposes only. If you publish material based on this database - please cite:<br>Zesch, T., & Gold, C. (2024). H̶a̶n̶d̶w̶r̶i̶t̶i̶n̶g̶ - various struck-through handwritten words [Data set]. Zenodo.</p> <h3><br>Structure:</h3> <p>There are two types of sources: <strong>genuine</strong> and <strong>semi-genuine</strong><br>While genuine includes struck-through handwritten images that are made accidentally, semi-genuine were made on purpose and conducted as the main part of this dataset. </p> <h3>Genuine:</h3> <p>For the genuine part, only the strike-out_genuine.txt file exists. It links to struck-through words of the datasets: Handwritten ASAP Short Answer Scoring (published at: https://zenodo.org/records/8088866), IAM, and GoBo (published at: https://zenodo.org/records/8085511).<br>The struck-through types were annotated by 2 annotators and a gold version was created. </p> <p>The file has the following structure:<br><em>path status gold a1 a2 a3</em><br>e.g.:<br>genuine/Handwritten ASAP/SAS_3_6818_0.png ok wa wa? wa? wa<br>The image is part of the Handwritten ASAP dataset and refers to image "AS_3_6818_0.png". The status is ok and the gold label is "wa" which stands for wavy (see list below). The "?" indicates unsure annotation of both annotators independently.</p> <h3><br>Semi-genuine: </h3> <p><br>17 writers participated. 9 male, 8 female<br>The writers were asked to write 12 words for each struck-through type. <br>In sum over 2000 images of handwritten struck-through words (including none) were collected, with 204 images each type. <br>An instruction was presented to the writers including examples. </p> <p><strong>types for strike-out:</strong><br> no: none<br> sh: single-horizontal<br> so: single-oblique<br> mh: multiple-horizontal<br> mo: multiple-oblique<br> cr: crossed<br> ci: circled<br> wa: wavy<br> zi: zigzag<br> bl: blackened</p> <p>Afterward, the writers struck through the handwritten words according to the stated type.<br>The images are published in "boxes", "gray images" and "raw images" in color as scanned. The description file "Strike-out_semi-genuine.txt" references the "boxes" only and describes the path, status, and struck-through type of each image. </p> <h3><br>Sidenote:</h3> <p>The writers were asked to note if the struck-through type was their preferred type and if it felt natural to use this type or unnatural. More details regarding this topic can be found in the instruction pdf file.</p> <p> </p>
Stavronikita Monastery Greek handwritten document Collection no.79
<p>It comprises manuscripts made of paper, written in the 16th century and its dimensions are 220X165 mm. The manuscript is embellished with epititles and red initials. Tachygraphical symbols and abbreviations are encountered in the manuscript as well. The dataset of XΦ79 consists of 803 lines of text containing 4389 words (2069 unique words) that are distributed over 40 scanned handwritten text pages.<br> For each page, a PageXML is provided containing the following ground-truth:</p> <p>1) Text region polygon coordinates<br> 2) Text line polygon coordinates with the corresponding transcription text<br> 3) Word polygon coordinated with the corresponding transcription text</p>
Stavronikita Monastery Greek handwritten document Collection no.114
<p>It comprises manuscripts made of paper, written at the end of the 15th century and its dimensions are 218X150 mm. In various pages, we find red initials and epititles which enrich the manuscript’s decoration. <br> The dataset of ΧΦ114 consists of 1051 lines of text containing 5467 (2877 unique words) words that are distributed over 44 scanned handwritten text pages. <br> For each page, a PageXML is provided containing the following ground-truth:</p> <p>1) Text region polygon coordinates<br> 2) Text line polygon coordinates with the corresponding transcription text<br> 3) Word polygon coordinated with the corresponding transcription text </p>
Stavronikita Monastery Greek handwritten document Collection no.53
<p>The collection is one of the oldest Stavronikita Monastery on Mount Athos. It is a parchment, four-gospel manuscript which has been written between 1301 and 1350. It comprises 54 pages with dimensions that are approximately 250x185 mm. The script is elegant minuscule and the use of majuscule letters is rare. Tachygraphical symbols and abbreviations are encountered in the manuscript as well. Furthermore, the manuscript is enriched with chrysography, elegant epititles and initials. The dataset of ΧΦ53 consists of 1038 lines of text containing 5592 words (2374 unique words) that are distributed over 54 scanned handwritten text pages.</p>
Urdu Handwritten Text Dataset
<p>The dataset contains the images of handwritten text in Urdu language, one of the most widely spoken languages in South-East Asian regions. The native-speaking authors from different social domains were invited to write a pre-written text in their handwritings. The pre-written text is carefully written in a way that it includes almost all the characters, ligatures, diacritics, and dots used in writing the text Urdu script. The disabled persons are also involved to write the text to make the data collection more comprehensive. The demographic data of the authors is also recorded for supporting the research activities like author identification, text-matching etc.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.