Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
57
datasets available to search
ShareScore release 0.9.0
Dataset results
57 results for “handwritten”
HHD-Ethiopic: A Historical Handwritten Dataset for Ethiopic OCR
<p>HHD-Ethiopic is a historical handwritten dataset for Ethiopic text-image recognition.</p>
ImageCLEF 2016 Bentham Handwritten Retrieval Dataset
<p>Dataset compiled for the ImageCLEF 2016 Handwritten Scanned Document Retrieval challenge. It is derived from a subset of pages from unpublished manuscripts written by the philosopher and reformer Jeremy Bentham, that have been digitised and transcribed under the Transcribe Bentham project [Causer 2012]. More details about the dataset and the challenge are found in the overview paper at http://ceur-ws.org/Vol-1609/16090233.pdf the slides of the overview presentation at http://imageclef.org/system/files/Villegas16_CLEF_Handwritten-Overview_presentation.pdf or the evaluation web page http://imageclef.org/2016/handwritten.</p> <p>[Causer 2012] T. Causer and V. Wallace, Building a Volunteer Community: Results and Findings from Transcribe Bentham, Digital Humanities Quarterly, Vol. 6 (2012), http://www.digitalhumanities.org/dhq/vol/6/2/000125/000125.html</p>
Test-B1 ICDAR2017 Competition on Handwritten Text Recognition on the READ Dataset (ICDAR2017 HTR)
<ul> <li><strong>Test-B1</strong>: a batch of page images annotated with the geometry of regions where to detect text line and recognize.</li> </ul>
Test-B2 ICDAR2017 Competition on Handwritten Text Recognition on the READ Dataset (ICDAR2017 HTR)
<p><strong>Test-B2</strong>: a batch of page images annotated with the geometry of regions where to detect text line and recognize.</p>
Dataset for ICDAR2017 Competition on Handwritten Text Recognition on the READ Dataset (ICDAR2017 HTR)
<p><strong>Train-A:</strong> Dataset of pages with manually revised baselines and the corresponding transcripts associated to them. This batch is small, 50 pages. Please, keep in mind that only the baselines have been manually corrected, The polygons associated to each line have not been manually reviewed.</p> <p><strong>Train-B:</strong> Dataset of pages without any layout or text line information. The corresponding transcripts are provided at page level with line breaks. It has 10k pages, though for convenience it is divided into two 5k page batches. This information is provided in PAGE format.</p> <p><strong>Test A:</strong> Dataset of pages with manually revised baselines. This batch has 65 pages. The polygons associated to each line have not been manually reviewed.</p> <p><strong>Test-B1:</strong> The same dataset of pages of the Test A, but annotated only with the geometry of regions. Text line information is not provided. </p> <p><strong>Test-B2:</strong> Dataset of page images annotated with the geometry of regions where to detect text line and recognize. It has 57 pages.</p> <p><strong>Baseline.tgz:</strong> Baseline system trained using the first 40 pages of Train-A. The system is based on the deep learning toolkit to transcribe handwritten text images called Laia.</p> <p>More information at:</p> <p>https://scriptnet.iit.demokritos.gr/competitions/~icdar2017htr/</p> <p> </p>
Appendix A - The Implications of Handwritten Text Recognition for Accessing the Past at Scale
<p>List of works identified through a Grounded Theory Method (GTM) of the current and near future implications of Handwritten Text Recognition (HTR) on the historical method and wider information environment. The findings of this data collection are provided in 'The Implications of Handwritten Text Recognition for Accessing the Past at Scale'</p>
Arabic Handwritten Legal Amount (AHLA) Dataset
<p>The AHLA dataset is collected by distributing an advanced designed report with Arabic native speakers. Our dataset contains two kinds of Arabic handwritten : <br>(1) Arabic word-level images that express legal amounts of bank cheques, including the colloquial words used in writing Arabic numbers. <br>(2) Arabic legal amount sentence images.</p> <p>The primary objective of compiling this comprehensive dataset is to furnish a diverse range of Arabic language samples. These samples are intended for training and testing systems capable of autonomously recognizing and comprehending handwritten legal amounts on financial documents. Subsequently, the aim is to convert these semantic expressions into their respective numeric currency totals, facilitating digital processing and banking operations.</p>
ND250 as a prediction error signal in orthographic processing: insights from the comparison of handwritten and printed words
<p><span>This dataset contains electroencephalography (EEG) recordings and behavioral data from a study investigating the neural mechanisms of visual word recognition in native Chinese speakers. The study used a color decision task, where participants viewed printed and handwritten Chinese single-character words varying in lexical frequency (high-frequency vs. low-frequency). The primary aim was to examine the N250 ERP component, a 250-ms difference in brain activity observed between certain word types, and determine whether it reflects activation of the orthographic lexicon or a prediction error signal during orthographic processing. The findings suggest that the N250 is related to prediction error, providing support for the Interactive Account of orthographic processing.</span></p>
Appendix A - Understanding the application of Handwritten Text Recognition technology in heritage contexts
<p>This appendix lists all the categorised works mentioning the HTR software Transkribus, used in 'Understanding the application of Handwritten Text Recognition technology in heritage contexts: a systematic review of Transkribus in published research'</p>
Information Extraction in Handwritten Historical Logbooks
<p>Contains the datasets with tables used in the following paper: Information Extraction in Handwritten Historical Logbooks.</p> <p>The pages starting with "vol003" correspond to the Jeannette corpus, while the ones starting with "Albatross" correspond to the Albatross corpus.</p>
ICDAR 2015 Competition HTRtS: Handwritten Text Recognition on the tranScriptorium Dataset Rerelease
<p>A new release of the dataset used in the ICDAR 2015 HTR competition in which all Page XML files are based on the same 2013-07-15 schema. It only contains page level images, Page XML files for train and test (including the ground truth transcripts for the test and train batch 1) and plain text files for train batch 2 that have the page level ground truth transcripts. The original version of this dataset can be found at http://doi.org/10.5281/zenodo.248733<br> </p>
Handwritten Opera OMR
<p>This repository contains the dataset of more than 198.000 small images extracted from handwritten music scores of the Archivio Storico Ricordi and organized for optical music recognition (OMR) research.</p> <p>The images are categorized in 11 classes (multiclass task) or 2 classes (binary task).</p> <p>In the download there is also the code used for annotating and analyzing the dataset, as well as the model weights resulting from training.</p> <p>Additional information is available in the paper (under review).</p> <h2>Cite us</h2> <p>Simonetta F., Mondal R., Ludovico L. A., Ntalampiras S. "<em>Optical Music Recognition in Manuscripts from the Ricordi Archive</em>", AudioMostly 2024, Milan, Italy. DOI: https://doi.org/10.1145/3678299.3678324</p> <p> </p>
ICDAR2013 – Handwritten Digit and Digit String Recognition Competition
<p>The CVL Single Digit dataset consists of 7000 single digits (700 digits per class) written by approximately 60 different writers. The validation set has the same size but different writers. The validation set may be used for parameter estimation and validation but not for supervised training. The CVL Digit Strings dataset uses 10 different digit strings from a total of about 120 writers resulting in 1262 training images. The digits from the CVL Single Digit dataset were extracted from these strings.</p> <p>This database may be used for non-commercial research purpose only. If you publish material based on this database, we request you to include a reference to:</p> <p>Markus Diem, Stefan Fiel, Angelika Garz, Manuel Keglevic, Florian Kleber and Robert Sablatnig, <em>ICDAR 2013 Competition on Handwritten Digit Recognition (HDRC 2013)</em>, In Proc. of the 12th Int. Conference on Document Analysis and Recognition (ICDAR) 2013, pp. 1454-1459, 2013.</p>
Handwritten Species Names Data
<p>This dataset is part of a paper presented at the <em>41st European Conference on Information Retrieval</em> ,14th – 18th April 2019, in Cologne. DOI: https://doi.org/10.1007/978-3-030-15712-8_43</p> <p><strong>Data summary:</strong> Word images from 240 field notes from a natural history collection have been segmented and semantically annotated. This has been carried out in the context of the project ''Making Sense of Illustrated Handwritten Archives'', http://www.makingsenseproject.org/. From a field book on mammals, field notes from four different writers have been selected, to account for different handwriting styles and structures. The segmented word images were obtained from a nichesourcing effort, with the help of a group of domain expert labellers and a handwriting recognition system MONK, developed by Lambert Schomaker. The word images were subsequently manually annotated using four classes: Genus (0), Species(1), Author(2) and Other (3). </p> <p><strong>Dataframe fields:</strong> rel_xc (relative centroid x coordinate), rel_yc (relative centroid y coordinate), page (identifier of field book page), rel_x1 (relative left x coordinate bounding box), rel_x2 (relative right x coordinate bounding box), rel_y1 (relative upper y coordinate bounding box), rel_y2 (relative lower y coordinate bounding box), type (class label, 0-3), image (pixels of word image), height (height word image), size (height * width word image), width (width word image). </p> <p> </p>
Handwritten Text Recognition Ground Truth Set: StABS Ratsbücher O10, Urfehdenbuch X
<p>Ground Truth for "Urfehdenbuch X der Stadt Basel (1563-1569)" at Staatsarchiv Basel-Stadt (StABS).</p> <p>Images and text aligned, using text-to-image (provided within Transkribus, <a href="https://www.readcoop.eu">www.readcoop.eu</a>).</p> <p>ALTO and Page XML are available for the text alignment.</p> <p>TEI to txt on page basis by Peter Dängeli.</p> <p>Derived from Transcription/TEI file: Urfehdenbuch X der Stadt Basel (1563-1569), in: Die Urfehdebücher der Stadt Basel – digitale Edition, hg. v. Susanna Burghartz, Sonia Calvi und Georg Vogeler Basel/Graz 2016. (zuletzt verändert am 31.1.2017): <a href="http://hdl.handle.net/11471/1010.2.1">hdl:11471/1010.2.1</a>.</p> <p>Image source: http://dokumente.stabs.ch/view/2010/Ratsbuecher_O_10/</p> <p>CC-BY-NC-SA: The license is inherited from the project "Urfehdebücher der Stadt Basel – digitale Edition": http://gams.uni-graz.at/o:ufbas.1563.</p>
SIMARA: a database for key-value information extraction from full-page handwritten documents
<p>We propose a new database for information extraction from historical handwritten documents. The corpus includes 5,393 finding aids from six different series, dating from the 18th-20th centuries. Finding aids are handwritten documents that contain metadata describing older archives. They are stored in the National Archives of France and are used by archivists to identify and find archival documents.</p> <p><br> Each document is annotated at page-level, and contains seven fields to retrieve. The localization of each field is not available in such a way that this dataset encourages research on segmentation-free systems for information extraction.<br> </p>
The NicIcon Database of Handwritten Icons for Crisis Management
<p><strong>The NicIcon Database of Handwritten Icons for Crisis Management</strong></p> <p>The NicIcon collection comprises simultaneously acquired online pen trajectories and offline scans of 14 classes of icons that are important in the domain of crisis management and incident response systems. The icons were designed such that (i) they have a visual resemblance to the objects they represent or correspond to well known corresponding symbols (so that they are easy to learn by the users), and (ii) are distinguishable by the computer. Furthermore, each participant was requested to write the so-called "London Letter", a well-known piece of text which is used to collect handwriting for forensic document examination purposes.</p> <p>The NicIcon database contains 24,441 hand drawn icons that are available both offline (as scans of pen-on-paper drawn icons) and online (recorded as time series coordinates on a digitizing tablet).</p> <p>In total, 32 participants, all volunteers, participated inthe experiment. They were all Dutch students in the age range of 19 to 30 (μ = 21.63, σ = 2.35), of which 29 were male, and 3 were female.</p> <p>A detailed description of the data and the collection can be found in the publication "The NicIcon Database of Handwritten Icons for Crisis Management" by Ralph Niels, Don Willems and Louis Vuurpijl. This paper can for example be found at <a href="https://www.researchgate.net/publication/228664054_The_NicIcon_Database_of_Handwritten_Icons_for_Crisis_Management/citations">https://www.researchgate.net/publication/228664054_The_NicIcon_Database_of_Handwritten_Icons_for_Crisis_Management/citations</a></p> <p> </p> <p>.DATA_INFO The NicIcon collection contains data acquired from 33<br> students. The collection contains so-called "iconic<br> gestures", representing symbolic drawings of situations and<br> events to be marked on so-called "interactive maps".<br> Gesture classes are inspired by symbolic representations of,<br> e.g., cars, persons, police, floods, accidents, etcetera.<br> In total 14 classes of gestures were collected.<br> Furthermore, each participant was requested to write the so-<br> called "London Letter", a well-known piece of text which is<br> used to collect handwriting for forensic document examination<br> purposes.</p> <p> The data is segmented and labeled in pages and icons.<br> The texts remain unsegmented, but segmentation in<br> line/word/character is on our agenda.</p> <p>.SETUP A Wacom Intuos2 A4 oversize tablet was used as digital<br> input device. This tablet was set at a resolution of<br> 100 dpi, a reading height of 10 mm, and a maximum data<br> rate of 100 pps. The device distinguishes 1024 pressure<br> levels. The tablet was connected to a computer running<br> Microsoft Windows XP. Specially developed software was<br> used to record spatial coordinates and<br> pressure coordinates. An A4-sized sheet of 160 grams/$m^2$<br> paper was clamped to the tablet. A Wacom writing pen,<br> with a green Lamy M21 pen tip<br> (see www.lamy.com/eng/b2c/Refills and inks/M 21), was<br> used by the participants to draw the data.</p> <p> Each sheet of paper was scanned using a HP Scanjet 7400C<br> flatbed scanner at 300dpi resolution in 24bits color. Both<br> online and offline data are available for download from<br> http://unipen.nici.ru.nl/NicIcon</p> <p> Each paper sheet contained seven rows of five columns,<br> resulting in 35 drawing areas. For each row, 5 instances of<br> the same class were drawn at a certain size. Each of the 33<br> participants had to fill in 22 of such paper sheets,<br> resulting in 770 icon gestures per participant.</p> <p>.LEXICON_INFO The 14 gesture classes are:<br> Accident - a triangle with exclamation mark inside<br> Bomb - a circle with fuse<br> Car - side-view of a car<br> Casualty - a circle with an 'X' cross below it<br> Electricity - a "lightning" arrow with head downwards<br> Flood - two horizontal curly lines<br> FireBrigade - a triangle with character 'F' inside<br> Fire - a gestures with flames<br> Gas - three vertical curly lines<br> Injury - a person laying horizontal<br> Paramedics - a square with a '+' cross inside<br> Person - a vertical (standing) human figure<br> Police - a diamond with a 'P' inside<br> RoadBlock - a circle with a '-' sign inside</p> <p> Each label associated with an icon contains 6 fields:<br> <www>-<pp>-<r>-<c>-<l>-<size> , where:<br> www = writer identification (000-034)<br> pp = page number (00-23)<br> r = row (0-6)<br> c = column (0-4)<br> l = classlabel<br> s = size<br> For example, the icon labeled as "000-19-5-4-flood-M"<br> was written by writer "000" on page "19", at row=5 and col=4,<br> the label is "flood" and the size is "medium". Note that the<br> first two sheets were used for practising and contain icons<br> labeled with size='A' (any). The other size categories are<br> 'S' (small), 'M' (medium), and 'L' (large).</p> <p> The text from the London letter reads:</p> <p> Our Londen business is good, but Vienna an Berlin are<br> quiet. Mr. D. Lloyd has gone to Switzerland and I hope<br> for good news. He will be there for a week at 1496<br> Zermatt Street and then goes to Turin and Rome and<br> will join Colonel Parry and arrive at Athens, Greece,<br> November 27th or December 2nd. Letters there should be<br> addressed King James Blvd. 3580. We expect Charles E.<br> Fuller Tuesday. Dr. L. Mc Quaid and Robert Unger,<br> Esq., left on the 'Y. X.' Express tonight.</p> <p>.PAD Wacom Intuos2 A4 oversize</p> <p>.COMMENT Below, the settings used during data collection. So, resolution was at<br> 100 dpi and sampling rate at 100Hz. Pressure is given in tablet units.<br> Both X and Y coordinates are given in micrometer (transformed from the<br> original tablet coordinates).</p> <p>.COMMENT xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx<br> TABLET TEMPORAL RESOLUTION<br> .POINTS_PER_SECOND 100</p> <p>.COMMENT xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx<br> TABLET SPATIAL RESOLUTION = 100dpi, but coordinates are in micro-meter<br> .X_POINTS_PER_MM 1000<br> .Y_POINTS_PER_MM 1000</p> <p>.COMMENT xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx<br> actually, pressure is not in gram but in tablet units<br> TABLET PRESSURE RESOLUTION = 1024<br> .POINTS_PER_GRAM 1024</p>
The Belfort dataset: Handwritten Text Recognition from Crowdsourced Annotations
<p>This dataset includes minutes of Belfort municipal council drawn up between 1790 and 1946. Documents include deliberations, lists of councillors, convocations, and agendas.</p> <p>The dataset includes 24,105 text-line images that were automatically detected from pages. Up to 4 transcriptions are available for each line image: two from humans, and two from automatic models.</p> <p>We would like to thank the <em>Archives municipales de la ville de Belfort, France</em> for giving us access to these documents.</p>
ICDAR 2023 CROHME: Competition on Recognition of Handwritten Mathematical Expressions
<p>Here is the datasets collected for the Competitionon Recognition of Online Handwritten Mathematical Expressions in competition session of ICDAR 2023. <br> 3 tasks are proposed with different modalities, there are on-line, off-line and bi-modal. <br> For on-line task, we provide .inkml file (contain trace information, mathML and LaTeX string), and also symbol level label graph (SymLG) as ground truth. Except the new data and previous CROHME data, we also provide huge amount of artificial on-line data in the train set. <br> For off-line task, the .png images (scanned from paper or rendering from inkml) and symbol level label graph (SymLG) are provided. Except the new data and previous CROHME data, we use off-line images from OffHME to increase the size of train set. <br> For bi-modal task, both .inkml file and ,png images are provided as 2 channels input, and SymLG as ground truth. </p> <p>All the 3 tasks inherited the data collected from the previous 6 CROHME, and also the new collection 2023 in 3 sites, Nantes (France), Luleå (Sweden) and Tokyo (Japan).</p>
UHaT Dataset: Urdu Handwritten Text Dataset
<p><strong>UHaT Dataset</strong></p> <p><strong>UHaT: Urdu Handwritten Text Dataset</strong></p> <p>This dataset contains handwritten characters and digits of Urdu language. The samples are written by 900+ individuals.</p> <p><strong>Description and organization</strong></p> <p>Size of images: All the images are stored in 28 by 28 resolution.</p> <p>How many images: The training set per each character contains of 700 images on average. For example, there are 811 train set images for AYN and 697 train set images for ALIF. Similarly, the train set per each contains 700 images on average. For example, there are 678 train set images for digits one. The test set per each character contains 140 images on average. For example, there are 145 test set images for character ALIF. The test set per each digit contains 140 images on average. For example, there are 147 test set images for digit nine.</p> <p>The dataset is organized into four sub-directories. Characters Training set, Characters Test set, Digits training set and digits test set. Each sub-director contains one sub-folder per one character. For example, all the train images for character ALIF are placed in sub-folder Alif.</p> <p>The folder hierarchy is given as:</p> <p>*Data > characters train set > alif</p> <p>Data > characters train set > ayn*</p> <p>And so on….</p> <p><strong>How to load directly?</strong></p> <p>You can also load it directly from the <em>uhat_<em>dataset.npz</em> file. See the kernel </em><strong>load_dataset</strong></p> <p><strong>Acknowledgements</strong></p> <p>Thanks to all volunteers who contributed by providing handwriting samples.</p> <p><strong>Inspiration</strong></p> <p>This is an <em>MNIST</em> style dataset. The machine learning community in general will find it useful for experimentation, demonstration purposes of machine learning models.<br> The dataset will also provide an opportunity to researchers to work on Urdu text recognition.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.