Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
30
datasets available to search
ShareScore release 0.9.0
Dataset results
30 results for “text recognition”
Named Entity Recognition Dataset for Dutch Biographical Texts
<p>A dataset for Named Entity Recognition for Dutch biographies. The original data is available in the Biographical portal of the Netherlands (http://www.biografischportaal.nl/). The annotations are for 6 types of entities: PERSON, LOCATION, ORGANIZATION, DATE, ARTWORK, MISC. Additionally, the CoNLL formatted files were manually checked for tokenization and sentence splitting.</p>
Transkribus - Handwritten Text Recognition for Premodern Documents (SIMS 2020 Lightning Talk)
<p>Transkribus is a platform for text recognition and can be used via the Transkribus Expert Software (available after registration: transkribus.eu). Through Transkribus different tools for document analysis and text recognition can be directly applied. The intro demonstrates very briefly how Transkribus can help with regards to premodern documents especially since a variety of pre-trained models are already available: for Latin (prints and handwriting), for early modern vernaculars in French, Dutch, English, and German. For more information go to transkribus.eu.</p> <p>Presented as a Schoenberg Symposium 2020 Lightning Talk</p>
Ground-Truthed Data Set of Zenon Papyri for Handwritten Text Recognition
<p>Diplomatic transcription of papyri found in the Zenon archive [see <a href="https://en.wikipedia.org/wiki/Zenon_of_Kaunos">en.wikipedia.org/wiki/Zenon_of_Kaunos</a>]</p> <p>Manually prepared as PageXML with Transkribus within <a href="http://d-scribes.philhist.unibas.ch/">D-Scribes</a> project.</p> <p> </p>
A Brazilian Portuguese Dataset for Offline Handwritten Text Recognition (BRESSAY)
<p>The BRESSAY dataset comprises images of handwritten essays in Brazilian Portuguese, which present a series of challenges to optical recognition models. These images were sourced from multiple online platforms, limiting our ability to standardize the capture process. Due to these varied sources and the lack of a uniform collection method, the dataset provides a realistic reflection of real-world conditions. Each essay is unique, contributed by different writers, and addresses a specific content topic. Furthermore, the constraints placed on the writers often lead to various handwriting scenarios, including hard-to-read words, connected words, noise, overwriting, and struck-through texts.</p> <h3><strong>Technical Details</strong></h3> <p>The BRESSAY dataset represents a comprehensive collection of handwritten essays in Brazilian Portuguese, offering detailed insights into various handwriting scenarios. It covers a total of 1,000 pages, each contributed by a unique writer, resulting in 1,000 distinct handwriting styles. This aspect of the dataset adds a layer of diversity, which is further emphasized by the total of 4,214 paragraphs, 30,090 lines, and 416,826 words. Regarding unique tokens, we have 41,318 unique words, and 107 unique characters.</p> <h3><strong>Data Structure</strong></h3> <p>The dataset is organized as follows:</p> <ul> <li>data/: Main folder containing segmented essay images <ul> <li>lines/: Images of individual lines <ul> <li> PNG files: Line images</li> <li> TXT files: Transcriptions of lines</li> </ul> </li> <li>pages/: Full page essay images <ul> <li> PNG files: Page images</li> <li> TXT files: Transcriptions of pages</li> </ul> </li> <li>paragraphs/: Images of paragraphs <ul> <li> PNG files: Paragraph images</li> <li> TXT files: Transcriptions of paragraphs</li> </ul> </li> <li>words/: Images of individual words <ul> <li> PNG files: Word images</li> <li> TXT files: Transcriptions of words</li> </ul> </li> </ul> </li> <li>sets/: Contains partition files <ul> <li>test.txt: Names of images in the test set</li> <li>validation.txt: Names of images in the validation set</li> <li>training.txt: Names of images in the training set</li> </ul> </li> </ul> <h3><strong>Dataset Usage and Annotations</strong></h3> <p>Each name in test.txt, validation.txt and training.txt represents the name of the page and all its content (words, lines, paragraphs) must be in the respective partition.</p> <p>Annotations used in the dataset:</p> <ul> <li> <code>##@@???@@##</code>: Superscript text that has become unidentifiable and unreadable.</li> <li> <code>$$@@???@@$$</code>: Subscript text that has become unidentifiable and unreadable.</li> <li> <code>@@???@@</code>: Text that cannot be read or identified due to its illegibility.</li> <li> <code>##--xxx--##</code>: Text that has been added as a superscript and subsequently crossed out, rendering it illegible.</li> <li> <code>$$--xxx--$$</code>: Text that has been added as a subscript and subsequently crossed out, rendering it illegible.</li> <li> <code>--xxx--</code>: Text that has been crossed out in a way that makes it unreadable.</li> <li> <code>##--text--##</code>: Text that has been added as a superscript and subsequently crossed out, but remains legible.</li> <li> <code>$$--text--$$</code>: Text that has been added as a subscript and subsequently crossed out, but remains legible.</li> <li> <code>##text##</code>: Text added as a superscript in the line, typically as a correction or additional note.</li> <li> <code>$$text$$</code>: Text added as a subscript in the line, typically as a correction or additional note.</li> <li> <code>--text--</code>: Text that has been crossed out but remains readable.</li> </ul>
FloraNER: a Named Entity Recognition Dataset for Botanical French Text
<p>FloraNER is a Named-Entity Recognition (NER) dataset for botanical french literature. The dataset covers plant species names and plant morphological terms for both plant organs/characteristics and their descriptors. the descriptors are annotated both in a coarse-grained manner as the named entity type "DESCRIPTOR" and in a fine-grained manner, categorized into the following named-entity types: Form, Measure, Surface, Color, Position, Disposition, Structure, and Development. FloraNER is distantly annotated using a specialized botanical corpus. Consequently, it's important to note that not all named entities within the text are captured and annotated for the coarse-grained and fine-grained datasets.</p>
Raw images (photographs) of urban text scenes for camera-based Thai text recognition
<p>Raw image collection of city scenes in Thailand with text content.<br> Text is photographed from diffeent angles. Also morning and evening<br> photographs were taken in order to capture different lighting<br> conditions. The material, 309 images, was photographed in 2013 by<br> Bowornrat Sriman and volunteers.</p> <p>Example EXIF:<br> JPEG image data, Exif standard: [TIFF image data, little-endian,<br> direntries=13, height=2448, manufacturer=SAMSUNG, model=GT-I9300,<br> orientation=upper-right, xresolution=220, yresolution=228,<br> resolutionunit=2, software=I9300XXEMA2, datetime=2013:03:14 18:17:49,<br> GPS-Data, width=3264], baseline, precision 8, 3264x2448, frames 3</p> <p>The images are not labeled. The orientation (landscape/portrait) is<br> not corrected yet. This material was used in preparation of the publication:</p> <p>Sriman, B. & Schomaker, L. (2015).<br> Object Attention Patches for Text Detection and Recognition in Scene Images using SIFT,<br> Proceedings of the International Conference on Pattern Recognition Applications and<br> Methods: ICPRAM 2015. De Marsico, M., Figueiredo, M. & Fred, A.<br> (Eds.). Lisbon, Portugal: SciTePress, Vol. 1, p. 304-311 8 p.</p> <p>Please cite this publication when using these data.</p>
Dataset for ICFHR2018 Competition on Automated Text Recognition on a READ Dataset
<p>The main idea of this dataset is to analyse the impact of training data. How many training data specific to the document, you are transcribing, is necessary? </p> <p><strong>general data: </strong>This is a collection of heterogeneous documents to train an initial system. For each text line there is an image file of that line, a file with the ground truth text and an information file containing an automatically generated surrounding polygon.</p> <p><strong>specific data: </strong>The specific data contains documents related to the test data. For the specific systems only the images of the train list may be used. The file are of the same type as the general data.</p> <p><strong>test data: </strong>The test data contains only the images and the information files.</p> <p>More Information, some published results and an evaluation procedure at https://scriptnet.iit.demokritos.gr/competitions/10/</p>
The e-NDP project : collaborative digital edition of the Chapter registers of Notre-Dame of Paris (1326-1504). Ground-truth for handwriting text recognition (HTR) on late medieval manuscripts.
<p>The <a href="https://endp.hypotheses.org/">e-NDP project</a>, funded by the ANR, is led by the <a href="https://lamop.hypotheses.org/6870">LaMOP</a> (Julie Claustre and Darwin Smith).</p> <p>The project's partners are the Archives nationales, the Bibliothèque nationale de France (Department of Manuscripts, Bibliothèque de l'Arsenal), the École nationale des chartes and the Bibliothèque Mazarine.</p> <p>The e-NDP project aims at renewing our knowledge on <strong>Notre-Dame de Paris cathedral</strong> through the creation of a collaborative digital edition of the registers of its Chapter (1326-1504, <em>AN LL 105-128</em>), the community of 51 canons meeting three times a week on set days to take all administrative, financial and practical decisions pertaining to the cathedral, its estate and the society living in its cloister. This corpus has never been the object of a comprehensive study to understand the workings and history of this urban enclave and powerful community. The collaborative digital edition is based on a process of<strong> handwriting text recognition (HTR)</strong>, tested and supervised by scholars, researchers and engineers combining expertise in Medieval history, paleography, philology and digital humanities. The edition shall allow a better insight into the Chapter’s administration, into its economical and political power within Paris, and the relationships it maintained with other institutions in the city.</p> <p> </p> <p><strong>Section 1 : The e-NDP ground-truth dataset for Handwriting text recognition.</strong></p> <p>The full e-NDP corpus kept today in the French National Archives and was entirely digitized and described in its <a href="https://www.siv.archives-nationales.culture.gouv.fr/siv/rechercheconsultation/consultation/ir/consultationIR.action?formCaller=GENERALISTE&irId=FRAN_IR_059635">catalog</a> in 2022.</p> <p>The first major goal of the e-NDP projet is to propose a first automatic transcription of the 14k pages composing the 26 chapter registers. To achieve this goal representative samples from each one of the volumes were selected and transcribed in order to train a specialized HTR model able to propose a high quality automatic transcription. The collected ground-truth released on this repository currently has <strong>512 pages from the 26 registers</strong> of the cathedral chapter preserved in the National Archives (LL105 - LL128, <strong>1326-1504</strong>). The transcriptions were manually completed in <strong>two rounds</strong> by a group of 12 contributors, historians and paleographers, over the course of 2021-2022 using <a href="https://escriptorium.paris.inria.fr/">eScriptorium </a>as annotation environment. </p> <p> </p> <p><strong>Ground-truth features :</strong></p> <p><br> <em>Number of hands </em>: according to our estimates no fewer than 18 main hands were involved in the writing of the registers during the medieval period. </p> <p><em>Language</em> : More than 98% of the content of the registers was written in Latin, the rest in French. The exact percentage is hard to estimate because the vernacular language is often used in formulae, notes and comments. It is rare to find entire pages or blocks written in French. </p> <p><em>Script family</em> : The registers were written using a Cursive script (ca. late XIIIe - XVIe).</p> <p><em>Documental typology</em> : The volumes containing the chapter conclusions were conceived to serve as memorial records, but above all as documents for regular use and consultation in the daily practice of administration and management. In diplomatics the notion of "documentary manuscripts" is used to describe this kind of sources also by opposition to books and litterary or normative manuscripts.</p> <table align="center"> <caption><strong>Ground truth statistics</strong></caption> <tbody> <tr> <th>Text units</th> <th>Count</th> </tr> <tr> <td>Pages</td> <td>512</td> </tr> <tr> <td>Annotated regions (see section 2)</td> <td>2448</td> </tr> <tr> <td>Lines of text</td> <td>34231</td> </tr> <tr> <td>Tokens</td> <td>205083</td> </tr> <tr> <td>Characters</td> <td>3320407</td> </tr> </tbody> </table> <p> </p> <p><strong>Rules of transcription :</strong></p> <ul> <li>The abbreviations have been resolved, both those by suspension (<code>facimꝰ</code> ---> <code>facimus</code>) and by contraction (<code>dñi</code> --> <code>domini</code>). Likewise, those using conventional signs (<code>⁊</code> --> <code>et</code> ; <code>ꝓ</code> --> <code>pro</code>) have been resolved. </li> <li>The named entities (names of persons, places and institutions) have been <code>capitalized</code>. The beginning of a block of text as well as the original capitals used by the notary are also capitalized.</li> <li>The consonantal <code>i</code> and <code>u</code> characters have been transcribed as <code>j</code> and <code>v</code> in both French and Latin.</li> <li>The punctuation marks used in the text: <code>.</code> and <code>/</code> have been transcribed, but the transcription has not been standardized with modern punctuation.</li> <li>Corrections and words that appear cancelled in the manuscript have been transcribed surrounded by the sign <code>$</code> at the beginning and at the end.</li> <li>More specific transcription rules can be found into the file <code>transcription_guidelines.pdf</code></li> </ul> <p> </p> <p><strong>Section 2. e-NDP Layout Segmentation.</strong></p> <p>Layout segmentation is a compulsory step before HTR recognition in order to distinguish sections and regions inside a document. This process intend to separate interdependant page zones to produce a recognition in a section-sequence order and not in a line-sequence order which mix textual and peri-textual content.</p> <p>The regions of 364 pages (see <code>GT-layout_list</code>) of the e-NDP corpus were annotated using a 5 sections vocabulary (see <code>endp_layout_regions</code>) in order to describe the page distribution in all the 26 volumes :</p> <ol> <li><em>Block</em> : All the central text blocks, that normally corresponds to the main content called "conclusions" in registers.</li> <li><em>Liste</em> : List of names of the canons who were present during the meeting. Normally located before the <em>conclusions</em>.</li> <li><em>Entrée</em> : Marginal notes or entries to inform about the content of <em>conclusions</em>.</li> <li><em>Date</em> : Paragraph contending the date. Normally at the head of a <em>conclusion</em>, but separate of the main body.</li> <li><em>Numérotation</em> : Page numbers in roman or arabic. Usually appear in the top corners of the pages.</li> </ol> <table align="center"> <caption><strong>Layout GT statistics</strong></caption> <tbody> <tr> <th>Region</th> <th>Count</th> </tr> <tr> <td>block</td> <td>833</td> </tr> <tr> <td>liste</td> <td>431</td> </tr> <tr> <td>date</td> <td>448</td> </tr> <tr> <td>entrée</td> <td>205</td> </tr> <tr> <td>numérotation</td> <td>531</td> </tr> </tbody> </table> <p> </p> <p><strong>Section 3. The e-NDP HTR modeling.</strong></p> <p>The e-NDP project has progressively trained several HTR models adapted to work on late medieval cursive in order to accelerate the production of ground truth. Currently the best model delivers an average <strong>CER (Character error ratio) of 9.7%</strong> in handwriting recognition on the 26 registers (see <code>endp_learning_curve</code>) and can serve as generalist model for other manuscripts of the same period and similar script family. These models and their training implementation details can be found in the project's github <a href="https://github.com/chartes/e-NDP_HTR">repository</a>. </p> <p>Additionally, the automatic HTR transcriptions of the 26 registers (14k pages, 4.5M tokens) enriched with lexical and semantical information has been the subject of a first <a href="https://nosketch-engine.lamop.fr/#dashboard?corpname=endp">online publication</a> using the NoSketch engine that allows advanced data mining based on the combination of data, metadata and NLP features. </p> <p> </p> <p><strong>Section 4. Dataset content.</strong></p> <p>This zip dataset contains :</p> <p>- <code>HTR_ground_truth</code> : Two folders containing the jpg / jpeg images and their curated transcriptions in PAGE XML format.</p> <p>- <code>images_docs</code> : 4 files illustrating the different phases of the project (list of GT for layout segmentation, layout ontologie, transcription guideline and HTR evaluation curves)</p>
ICDAR 2015 Competition HTRtS: Handwritten Text Recognition on the tranScriptorium Dataset
<p>This dataset comprises the dataset used for the ICDAR 2015 Competition on Handwritten Text Recognition on the tranScriptorium Dataset. The handwritten images for this contest were drawn from the English “Bentham collection” dataset used in the TRAN SCRIPTORIUM project. The selected data has been written by several hands and entails significant variabilities and difficulties regarding the quality of text images, writing styles and crossed-out text. This contest is clearly more difficult than the the first edition both for training and for testing. A portion of the training dataset and the full test dataset were provided in the form of carefully segmented line images, along with the corresponding transcripts. Another portion of the training dataset was provided as raw images and their corresponding transcripts at region level.<br> </p> <p>ICDAR 2015 competition HTRtS: handwritten text recognition on the tranScriptorium dataset<br> JA Sánchez, AH Toselli, V Romero, E Vidal. In International Conference on Document Analysis and Recognition (ICDAR), pp. 1166-1170, 2015.</p>
Train-B dataset for ICDAR2017 Competition on Handwritten Text Recognition on the READ Dataset (ICDAR2017 HTR). Batch 1 and Batch 2.
<p>Train-B Dataset. Dataset of pages without any layout or text line information. The corresponding transcripts are provided at page level with line breaks. It has 10k pages, though for convenience it is divided into two 5k page batches. This information is provided in PAGE format. </p> <p>This dataset is complementary to this other dataset:</p> <p>https://zenodo.org/record/439807#.WOIBZ3WLSkA</p> <p>More information at:</p> <p>https://scriptnet.iit.demokritos.gr/competitions/~icdar2017htr/</p> <p> </p>
Train-A dataset for ICDAR2017 Competition on Handwritten Text Recognition on the READ Dataset (ICDAR2017 HTR)
<p>Train-A Dataset of pages with manually revised baselines and the corresponding transcripts associated to them. This batch is small, 50 pages. Please, keep in mind that only the baselines have been manually corrected, The polygons associated to each line have not been manually reviewed. </p> <p>This dataset is complementary to this other dataset:</p> <p>https://zenodo.org/record/439811#.WOIF9HWLSkA</p> <p>More information at:</p> <p>https://scriptnet.iit.demokritos.gr/competitions/~icdar2017htr/</p>
NorHand v3 / Dataset for Handwritten Text Recognition in Norwegian
<p>This dataset comprises Norwegian letter and diary documents from 19th and early 20th century. It can be used to train Handwritten Text Recognition (HTR) models.</p>
Oficio de Hipotecas de Girona. A dataset of Spanish notarial deeds (18th Century) for Handwritten Text Recognition and Layout Analysis of historical documents.
<p>This dataset is a subset of 596 documents from the <em>Registre d'Hipoteques de Girona</em> of 1769 collection, guarded by the <a href="http://xac.gencat.cat/ca/llista_arxius_comarcals/girones/"><em>Arxiu Històric de Girona</em></a>. This collection, is composed by hundreds of thousands of notarial deeds from the XVIII-XIX century (1768-1862). Sales, redemption of censuses, inheritance and matrimonial chapters are among the most common documentary typologies in the collection.</p> <p>This dataset is composed of more than 23700 text lines written by a single hand, covering more that 50 different topics (documentary typologies) and a vocabulary of more than 2400 different words. The documents are transcribed using the so-called diplomatic criteria. Additionally, transcripts were tagged with <br> extra enriching/complementary information (e.g. expansion of the abbreviations, hyphen marks, etc.). Along with the transcripts the layout of the document is detected and recorded. Pages have been labeled using six different layout regions.</p> <p>The images along with their respective ground-truth was compiled in PAGE compliant XML format<br> by the <a href="http://www2.udg.edu/tabid/11296/Default.aspx"><em>Centre de Recerca d'Història Rural</em></a> and the HTR group of the <a href="https://www.prhlt.upv.es">Pattern Recognition and Human Language Technologies Research Center</a>.</p>
Test-B1 ICDAR2017 Competition on Handwritten Text Recognition on the READ Dataset (ICDAR2017 HTR)
<ul> <li><strong>Test-B1</strong>: a batch of page images annotated with the geometry of regions where to detect text line and recognize.</li> </ul>
Test-B2 ICDAR2017 Competition on Handwritten Text Recognition on the READ Dataset (ICDAR2017 HTR)
<p><strong>Test-B2</strong>: a batch of page images annotated with the geometry of regions where to detect text line and recognize.</p>
Dataset for ICDAR2017 Competition on Handwritten Text Recognition on the READ Dataset (ICDAR2017 HTR)
<p><strong>Train-A:</strong> Dataset of pages with manually revised baselines and the corresponding transcripts associated to them. This batch is small, 50 pages. Please, keep in mind that only the baselines have been manually corrected, The polygons associated to each line have not been manually reviewed.</p> <p><strong>Train-B:</strong> Dataset of pages without any layout or text line information. The corresponding transcripts are provided at page level with line breaks. It has 10k pages, though for convenience it is divided into two 5k page batches. This information is provided in PAGE format.</p> <p><strong>Test A:</strong> Dataset of pages with manually revised baselines. This batch has 65 pages. The polygons associated to each line have not been manually reviewed.</p> <p><strong>Test-B1:</strong> The same dataset of pages of the Test A, but annotated only with the geometry of regions. Text line information is not provided. </p> <p><strong>Test-B2:</strong> Dataset of page images annotated with the geometry of regions where to detect text line and recognize. It has 57 pages.</p> <p><strong>Baseline.tgz:</strong> Baseline system trained using the first 40 pages of Train-A. The system is based on the deep learning toolkit to transcribe handwritten text images called Laia.</p> <p>More information at:</p> <p>https://scriptnet.iit.demokritos.gr/competitions/~icdar2017htr/</p> <p> </p>
Appendix A - The Implications of Handwritten Text Recognition for Accessing the Past at Scale
<p>List of works identified through a Grounded Theory Method (GTM) of the current and near future implications of Handwritten Text Recognition (HTR) on the historical method and wider information environment. The findings of this data collection are provided in 'The Implications of Handwritten Text Recognition for Accessing the Past at Scale'</p>
Appendix A - Understanding the application of Handwritten Text Recognition technology in heritage contexts
<p>This appendix lists all the categorised works mentioning the HTR software Transkribus, used in 'Understanding the application of Handwritten Text Recognition technology in heritage contexts: a systematic review of Transkribus in published research'</p>
ICDAR 2015 Competition HTRtS: Handwritten Text Recognition on the tranScriptorium Dataset Rerelease
<p>A new release of the dataset used in the ICDAR 2015 HTR competition in which all Page XML files are based on the same 2013-07-15 schema. It only contains page level images, Page XML files for train and test (including the ground truth transcripts for the test and train batch 1) and plain text files for train batch 2 that have the page level ground truth transcripts. The original version of this dataset can be found at http://doi.org/10.5281/zenodo.248733<br> </p>
Handwritten Text Recognition Ground Truth Set: StABS Ratsbücher O10, Urfehdenbuch X
<p>Ground Truth for "Urfehdenbuch X der Stadt Basel (1563-1569)" at Staatsarchiv Basel-Stadt (StABS).</p> <p>Images and text aligned, using text-to-image (provided within Transkribus, <a href="https://www.readcoop.eu">www.readcoop.eu</a>).</p> <p>ALTO and Page XML are available for the text alignment.</p> <p>TEI to txt on page basis by Peter Dängeli.</p> <p>Derived from Transcription/TEI file: Urfehdenbuch X der Stadt Basel (1563-1569), in: Die Urfehdebücher der Stadt Basel – digitale Edition, hg. v. Susanna Burghartz, Sonia Calvi und Georg Vogeler Basel/Graz 2016. (zuletzt verändert am 31.1.2017): <a href="http://hdl.handle.net/11471/1010.2.1">hdl:11471/1010.2.1</a>.</p> <p>Image source: http://dokumente.stabs.ch/view/2010/Ratsbuecher_O_10/</p> <p>CC-BY-NC-SA: The license is inherited from the project "Urfehdebücher der Stadt Basel – digitale Edition": http://gams.uni-graz.at/o:ufbas.1563.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.