Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
7
datasets available to search
ShareScore release 0.9.0
Dataset results
7 results for “HTR models”
Dataset and evaluation for HTR models for Latin and French Medieval Documentary Manuscripts
<p><strong>1. Dataset presentation.</strong></p> <p>This is the dataset used to produce the HTR models applied to documentary Latin and French manuscripts presented in the paper: Sergio Torres Aguilar, Vincent Jolivet. <strong>Handwritten Text Recognition for Documentary Medieval<br> Manuscripts. </strong>2022. https://hal.science/hal-03892163</p> <p>The dataset contains mostly charters and registers from the Late-medieval period (12th-15th). The training and evaluation, entailing 1855 pages, 120k lines of text and almost 1M tokens, were conducted using three freely available ground-truth corpora :</p> <p><strong>The Alcar-HOME database </strong>: https://zenodo.org/record/5600884</p> <p><strong>The e-NDP corpus </strong>: https://github.com/chartes/e-NDP_HTR</p> <p><strong>The Himanis project </strong>: https://zenodo.org/record/5535306</p> <p>The final model operates in a multilingual environment (Latin and French) and it is able to recognize several Latin script families (mostly <em>Textualis</em> and <em>Cursiva</em>) in documents produced in ca. 12th - 15th centuries. During the evaluation the models shows an accuracy of <strong>94.01%</strong> on the validation set and a CER (character error ratio) of about <strong>0.12</strong> to <strong>0.17</strong> on four external unseen datasets. A fine-tuning exercise using 10 ground-truth pages can raise these results to a CER between <strong>0.06</strong> to <strong>0.10</strong> respectively.</p> <p> </p> <p><strong>2. Dataset contents .</strong></p> <p>a) <em>GT_list : </em>List containing the GT file names which constitute the training, evaluation and test sets. The images and transcriptions can be downloaded from their original repositories.</p> <p>b) <em>Training :</em> Contains the training and testing results (evaluation and prediction files) presented in the original paper for the two training phases: Regular (Textualis and Cursiva separated training) and Quartiles (mixed training by quartiles).</p> <p>c) <em>Useful_scripts :</em> Scripts to produce the HTR metrics (CER, WER, SER) and plot the model's accuracy.</p> <p>d) <em>Best_model :</em> Contains the best multilingual and multi-script model.</p>
TRIDIS: HTR model for Multilingual Medieval and Early Modern Documentary Manuscripts (11th-16th)
<p><strong>TRIDIS (Tria Digita Scribunt)</strong> is a Handwriting Text Recognition model trained on semi-diplomatic transcriptions from medieval and Early Modern Manuscripts. It is suitable for work on documentary manuscripts, that is, manuscripts arising from legal, administrative, and memorial practices more commonly from the Late Middle Ages (13th century and onwards). It can also show good performance on documents from other domains, such as literature books, scholarly treatises and cartularies providing a versatile tool for historians and philologists in transforming and analyzing historical texts.</p> <p>A paper presenting the first version of the model is available here: Sergio Torres Aguilar, Vincent Jolivet. <strong>Handwritten Text Recognition for Documentary Medieval Manuscripts. </strong>Journal of Data Mining and Digital Humanities.<strong> </strong>2023. https://hal.science/hal-03892163</p> <p> </p> <h3>Transcriptions rules :</h3> <p>Since the majority of the training documents come from diplomatic editions, the transcriptions were <strong>normalized</strong> to contemporary reading standards, and <strong>abbreviations were expanded</strong> with the aim of facilitating a more fluid reading of the document.</p> <p>The following rules were applied:</p> <ul> <li>The abbreviations have been expanded, both those by suspension (<code>facimꝰ</code> ---> <code>facimus</code>) and by contraction (<code>dñi</code> --> <code>domini</code>). Likewise, those using conventional signs (<code>⁊</code> --> <code>et</code> ; <code>ꝓ</code> --> <code>pro</code>) have been resolved. </li> <li>The named entities (names of persons, places and institutions) have been <code>capitalized</code>. The beginning of a block of text as well as the original capitals used by the scribe are also capitalized.</li> <li>The consonantal <code>i</code> and <code>u</code> characters have been transcribed as <code>j</code> and <code>v</code> in both French and Latin.</li> <li>The punctuation marks used in the manuscript like: <code>.</code> or <code>/</code> or <code>|</code> have not been systematically transcribed as the transcription has been standardized with modern punctuation.</li> <li>Corrections and words that appear cancelled in the manuscript have been transcribed surrounded by the sign <code>$</code> at the beginning and at the end.</li> </ul> <p> </p> <h3>Versions :</h3> <p><strong>Version 1 </strong>of the model was trained on charters and registers dataset from the Late Medieval period (12th-15th centuries). The training and evaluation involved 1855 pages, 120k lines of text, and almost 1M tokens, conducted using three freely available ground-truth corpora:</p> <ul> <li>The Alcar-HOME database: <a href="../record/5600884" target="_new">https://zenodo.org/record/5600884</a></li> <li>The e-NDP corpus: <a href="../record/7575693" target="_new">https://zenodo.org/record/7575693</a></li> <li>The Himanis project: <a href="../record/5535306" target="_new">https://zenodo.org/record/5535306</a></li> </ul> <p><strong>Version 2</strong> of the model has added new datasets from feudal books and legal proceedings (14th-16th centuries), incorporating an additional 115k lines and more than 1.2M tokens to the previous version using other corpora like:</p> <ul> <li>Königsfelden Abbey corpus: <a href="../record/5179361" target="_new">https://zenodo.org/record/5179361</a></li> <li>Monumenta Luxemburgensia.</li> </ul> <p> </p> <h3>Accuracy</h3> <p>TRIDIS was trained using a CNN+RNN+CTC architecture within the Kraken suite (https://kraken.re/). This final model operates in a multilingual environment (Latin, Old French, and Old Spanish) and is capable of recognizing several Latin script families (mostly Textualis and Cursiva) in documents produced circa 11th - 16th centuries. During evaluation, the model showed an accuracy of 93.1% on the validation set and a CER (Character Error Ratio) of about 0.11 to 0.15 on four external unseen datasets. Fine-tuning the model with 10 ground-truth pages can improve these results to a CER of between 0.06 to 0.10, respectively.</p> <h3>Other formats</h3> <p>The ground truth used for version 2 was also employed to train a Transformer HTR model that combines TrOCR as the encoder with a RoBERTa medieval model as the decoder. This model exhibits a slighly better performance in terms of CER metrics to the current TRIDIS version and shows an improved WER by about 25%. The model is available on the Hugging Face Hub: <a href="https://huggingface.co/magistermilitum/tridis_HTR">magistermilitum/tridis_HTR</a></p>
HTR Model Spanish Gothic Incunabula (HSMS)
<p>The <a href="https://www.transkribus.org/model/spanish-gothic-incunabula" target="_blank" rel="noopener">Spanish Gothic Incunabula (HSMS)</a> is conceived to be uploaded inside Transkribus platform (READ Coop) to perform a training and create an PyLaia model for the automated recognition of Spanish incunabula in Gothic script printed between 1472 and 1500. It can be used for post-incunabula (up to 1520).<br><br>The transcription model follows the rules set by the Hispanic Seminary of Medieval Studies in 1977 (<a href="https://hispanicseminary.org/manual-en.htm" target="_blank" rel="noopener">newest version</a>). The rules applied are:</p> <ul> <li>Abbreviated words are expanded and the expanded text is enclosed between < >: <em>q<ue></em></li> <li>Superscripted letters are followed by a grave accent: <em>q<u>i`en</em></li> <li>All <em>ç</em> and <em>ñ</em> are transcribed as <em>ç</em> and <em>ñ</em></li> <li>No attemp has been made to normalize spacing</li> <li>All punctuations signs are kept</li> <li>All abbreviated nasals before <em>b</em> or <em>p</em> are transcribed as <n>. It is up to the editors if they should be changed into <em>m</em>.</li> <li>Abbreviated v' (tilde over v, or small slash v) that can be expanded as v<ir> or v<er> is expanded as v<er>. It is up to the editor if they should be changed to v<ir>.</li> <li>Tironian <em>et</em> is transcribed as & (ampersand)</li> <li>Pilcrows are transcribed as ¶</li> </ul> <p>The model is built on 200 openings (verso-recto) drawn from 20 books printed by five different workshops form Sevile, Zaragoza, Burgos, Toledo and Pamplona. They Train Set consist 180152 words, distribuited over 24061 lines. The CER on the Train Set is 0.20 % and on the Validation Set 0.77 %.</p> <p> </p>
HTR model NIOD_WarLet_1935-1950_NoBasemodel
<p>The HTR model ‘NIOD_WarLet_1935-1950_NoBasemodel’ was trained using 968 ‘Ground Truth’ transcriptions of high-resolution scans of various handwritten letters. These letters are all written in Dutch and originate from the period 1935-1950. The training set contains personal correspondence from a wide variety of letter writers (e.g., children, soldiers, Jewish people in hiding). These personal correspondences are all part of the archival collection known as ‘247 Correspondentie’ held by the NIOD Institute for War, Holocaust, and Genocide Studies in Amsterdam.<br> <br> This model was created as part of the project ‘First-Hand Accounts of War: War letters (1935-1950) from NIOD digitised’. All documents used for training and validation were scanned and transcribed within this project. This project ran from 2020 to 2023 and was funded by the Mondriaan Fund, the Dutch Ministry of Health, Welfare, and Sport, and the NIOD Institute for War, Holocaust, and Genocide Studies in Amsterdam.<br> <br> The ‘Ground Truth’ training set is created by project members Annelies van Nispen, Carlijn Keijzer and Milan van Lange. Additional transcription and correction of ‘Ground Truth’ transcriptions was performed under supervision of Muriël Bouman by citizen scientists Hillebrand Verkroost, Bart Cohen, Evelien Bachrach, Marjo Janssens, and Cocky Sietses. The validation set contains a sample of 17 ‘Ground Truth’ transcriptions from various writers and sub-collections. Due to legal restrictions only a limited sample of the training set is published publicly.<br> <br> The model is trained using PyLaia HTR, max. 500 epochs (321 epochs trained), learning rate 0.0003. No basemodel was used. See also: https://readcoop.eu/model/niod_warlet_1945-1950_nobasemodel/</p>
Targeting Notch signaling to restore neural development and behavior in mouse models of ASD [RNAseq_htr]
GEO Series GSE293296. Mus musculus. 12 samples. Type: Expression profiling by high throughput sequencing.
HTR model SpanishRedonda_XVI-XVII_extended DATASET (v1.2)
<pre>The SpanishRedonda_XVI-XVII_extended DATASET is conceived to be uploaded inside Transkribus platform (READ Coop) to perform a training and create an HTR+ model for the automated recognition of Spanish printed documents in Round script published between XV-XVI Century. For further information please have a look to the README file here: https://github.com/stefanobazzaco/HTR-model-SpanishRedonda_XVI-XVII_extended For system requirements and information about Transkribus platform, go to: https://readcoop.eu/transkribus/ https://readcoop.eu/transkribus/resources/ https://github.com/Transkribus</pre> <p> </p>
HTR model SpanishGothic_XV-XVI_extended DATASET (v1.2)
<pre>The SpanishGothic_XV-XVI_extended DATASET is conceived to be uploaded inside Transkribus platform (READ Coop) to perform a training and create an HTR+ model for the automated recognition of Spanish printed documents in Gothic script published between XV-XVI Century. For further information please have a look to the README file here: https://github.com/stefanobazzaco/HTR-model-SpanishGothic_XV-XVI_extended For system requirements and information about Transkribus platform, go to: https://readcoop.eu/transkribus/ https://readcoop.eu/transkribus/resources/ https://github.com/Transkribus</pre>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.