Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
50
datasets available to search
ShareScore release 0.7.1
Dataset results
50 results for “HTR”
HTR Dataset ICFHR 2016
<p>This dataset arises from the READ project (Horizon 2020).</p> <p>The dataset consists of a subset of documents from the Ratsprotokolle collection composed of minutes of the council meetings held from 1470 to 1805 (about 30.000 pages), which will be used in the READ project. This dataset is written in Early Modern German. The number of writers is unknown. Handwriting in this collection is complex enough to challenge the HTR software.</p> <p>The training dataset is composed of 400 pages; most of the pages consist of a single block with many difficulties for line detection and extraction. The ground-truth in this set is in PAGE format and it is provided annotated at line level in the PAGE files.</p> <p>The previous dataset is the same that is located at https://zenodo.org/record/218236#.WnLhaCHhBGF</p> <p>The new file includes the test set corresponding to the HTR competition held at ICFHR 2016</p>
HTRCatalogs: Dataset for historical catalogs HTR and Segmentation
<p>This release contains 465 xml files, and their corresponding images from a large corpus of 19th, 20th and 21th exhibition catalogs, manuscripts'fair catalogs and directories. The new catalogs added here were created using the HTR and segmentation models accessible in the repository. It includes a csv file describing the xml files and various tools to create a training dataset: differents bash scripts, a python programm to divide the xml files into testing, training and evaluation dataset and several fixed tests. A xsl transformation sheet is also accessible to delete the Entry and EntryEnd zones from the xml files in order to have a SegmOnto-like dataset. The xml files has been corrected since the 4.0 release thanks to the addition of a github action (SegmOntoKraken).</p>
The e-NDP project : collaborative digital edition of the Chapter registers of Notre-Dame of Paris (1326-1504). Ground-truth for handwriting text recognition (HTR) on late medieval manuscripts.
<p>The <a href="https://endp.hypotheses.org/">e-NDP project</a>, funded by the ANR, is led by the <a href="https://lamop.hypotheses.org/6870">LaMOP</a> (Julie Claustre and Darwin Smith).</p> <p>The project's partners are the Archives nationales, the Bibliothèque nationale de France (Department of Manuscripts, Bibliothèque de l'Arsenal), the École nationale des chartes and the Bibliothèque Mazarine.</p> <p>The e-NDP project aims at renewing our knowledge on <strong>Notre-Dame de Paris cathedral</strong> through the creation of a collaborative digital edition of the registers of its Chapter (1326-1504, <em>AN LL 105-128</em>), the community of 51 canons meeting three times a week on set days to take all administrative, financial and practical decisions pertaining to the cathedral, its estate and the society living in its cloister. This corpus has never been the object of a comprehensive study to understand the workings and history of this urban enclave and powerful community. The collaborative digital edition is based on a process of<strong> handwriting text recognition (HTR)</strong>, tested and supervised by scholars, researchers and engineers combining expertise in Medieval history, paleography, philology and digital humanities. The edition shall allow a better insight into the Chapter’s administration, into its economical and political power within Paris, and the relationships it maintained with other institutions in the city.</p> <p> </p> <p><strong>Section 1 : The e-NDP ground-truth dataset for Handwriting text recognition.</strong></p> <p>The full e-NDP corpus kept today in the French National Archives and was entirely digitized and described in its <a href="https://www.siv.archives-nationales.culture.gouv.fr/siv/rechercheconsultation/consultation/ir/consultationIR.action?formCaller=GENERALISTE&irId=FRAN_IR_059635">catalog</a> in 2022.</p> <p>The first major goal of the e-NDP projet is to propose a first automatic transcription of the 14k pages composing the 26 chapter registers. To achieve this goal representative samples from each one of the volumes were selected and transcribed in order to train a specialized HTR model able to propose a high quality automatic transcription. The collected ground-truth released on this repository currently has <strong>512 pages from the 26 registers</strong> of the cathedral chapter preserved in the National Archives (LL105 - LL128, <strong>1326-1504</strong>). The transcriptions were manually completed in <strong>two rounds</strong> by a group of 12 contributors, historians and paleographers, over the course of 2021-2022 using <a href="https://escriptorium.paris.inria.fr/">eScriptorium </a>as annotation environment. </p> <p> </p> <p><strong>Ground-truth features :</strong></p> <p><br> <em>Number of hands </em>: according to our estimates no fewer than 18 main hands were involved in the writing of the registers during the medieval period. </p> <p><em>Language</em> : More than 98% of the content of the registers was written in Latin, the rest in French. The exact percentage is hard to estimate because the vernacular language is often used in formulae, notes and comments. It is rare to find entire pages or blocks written in French. </p> <p><em>Script family</em> : The registers were written using a Cursive script (ca. late XIIIe - XVIe).</p> <p><em>Documental typology</em> : The volumes containing the chapter conclusions were conceived to serve as memorial records, but above all as documents for regular use and consultation in the daily practice of administration and management. In diplomatics the notion of "documentary manuscripts" is used to describe this kind of sources also by opposition to books and litterary or normative manuscripts.</p> <table align="center"> <caption><strong>Ground truth statistics</strong></caption> <tbody> <tr> <th>Text units</th> <th>Count</th> </tr> <tr> <td>Pages</td> <td>512</td> </tr> <tr> <td>Annotated regions (see section 2)</td> <td>2448</td> </tr> <tr> <td>Lines of text</td> <td>34231</td> </tr> <tr> <td>Tokens</td> <td>205083</td> </tr> <tr> <td>Characters</td> <td>3320407</td> </tr> </tbody> </table> <p> </p> <p><strong>Rules of transcription :</strong></p> <ul> <li>The abbreviations have been resolved, both those by suspension (<code>facimꝰ</code> ---> <code>facimus</code>) and by contraction (<code>dñi</code> --> <code>domini</code>). Likewise, those using conventional signs (<code>⁊</code> --> <code>et</code> ; <code>ꝓ</code> --> <code>pro</code>) have been resolved. </li> <li>The named entities (names of persons, places and institutions) have been <code>capitalized</code>. The beginning of a block of text as well as the original capitals used by the notary are also capitalized.</li> <li>The consonantal <code>i</code> and <code>u</code> characters have been transcribed as <code>j</code> and <code>v</code> in both French and Latin.</li> <li>The punctuation marks used in the text: <code>.</code> and <code>/</code> have been transcribed, but the transcription has not been standardized with modern punctuation.</li> <li>Corrections and words that appear cancelled in the manuscript have been transcribed surrounded by the sign <code>$</code> at the beginning and at the end.</li> <li>More specific transcription rules can be found into the file <code>transcription_guidelines.pdf</code></li> </ul> <p> </p> <p><strong>Section 2. e-NDP Layout Segmentation.</strong></p> <p>Layout segmentation is a compulsory step before HTR recognition in order to distinguish sections and regions inside a document. This process intend to separate interdependant page zones to produce a recognition in a section-sequence order and not in a line-sequence order which mix textual and peri-textual content.</p> <p>The regions of 364 pages (see <code>GT-layout_list</code>) of the e-NDP corpus were annotated using a 5 sections vocabulary (see <code>endp_layout_regions</code>) in order to describe the page distribution in all the 26 volumes :</p> <ol> <li><em>Block</em> : All the central text blocks, that normally corresponds to the main content called "conclusions" in registers.</li> <li><em>Liste</em> : List of names of the canons who were present during the meeting. Normally located before the <em>conclusions</em>.</li> <li><em>Entrée</em> : Marginal notes or entries to inform about the content of <em>conclusions</em>.</li> <li><em>Date</em> : Paragraph contending the date. Normally at the head of a <em>conclusion</em>, but separate of the main body.</li> <li><em>Numérotation</em> : Page numbers in roman or arabic. Usually appear in the top corners of the pages.</li> </ol> <table align="center"> <caption><strong>Layout GT statistics</strong></caption> <tbody> <tr> <th>Region</th> <th>Count</th> </tr> <tr> <td>block</td> <td>833</td> </tr> <tr> <td>liste</td> <td>431</td> </tr> <tr> <td>date</td> <td>448</td> </tr> <tr> <td>entrée</td> <td>205</td> </tr> <tr> <td>numérotation</td> <td>531</td> </tr> </tbody> </table> <p> </p> <p><strong>Section 3. The e-NDP HTR modeling.</strong></p> <p>The e-NDP project has progressively trained several HTR models adapted to work on late medieval cursive in order to accelerate the production of ground truth. Currently the best model delivers an average <strong>CER (Character error ratio) of 9.7%</strong> in handwriting recognition on the 26 registers (see <code>endp_learning_curve</code>) and can serve as generalist model for other manuscripts of the same period and similar script family. These models and their training implementation details can be found in the project's github <a href="https://github.com/chartes/e-NDP_HTR">repository</a>. </p> <p>Additionally, the automatic HTR transcriptions of the 26 registers (14k pages, 4.5M tokens) enriched with lexical and semantical information has been the subject of a first <a href="https://nosketch-engine.lamop.fr/#dashboard?corpname=endp">online publication</a> using the NoSketch engine that allows advanced data mining based on the combination of data, metadata and NLP features. </p> <p> </p> <p><strong>Section 4. Dataset content.</strong></p> <p>This zip dataset contains :</p> <p>- <code>HTR_ground_truth</code> : Two folders containing the jpg / jpeg images and their curated transcriptions in PAGE XML format.</p> <p>- <code>images_docs</code> : 4 files illustrating the different phases of the project (list of GT for layout segmentation, layout ontologie, transcription guideline and HTR evaluation curves)</p>
Lausanne Historical Censuses Dataset HTR 35k
<p>This training dataset includes a total of 34,913 manually transcribed text segments. It is dedicated to the handwritten text recognition (HTR) of historical sources, typically tabular records, such as censuses. This dataset is based on a sample of 83 pages from the 19th century (1805-1898) censuses of Lausanne, Switzerland. The primary language of the documents is French, although many germanic names and toponyms are also found.</p> <p>The training data are formatted and provided on the model of the Bentham dataset. The format thus simply consists in a list of jpeg images, one per text segments, and their corresponding transcription, stored in a txt file. The file naming convention is 'yyyy-ppp-n', where 'y' stands for the year of publication of the census, and 'p' for the page number.</p> <p>The digitized documents are provided by the <a href="http://www.lausanne.ch/vie-pratique/culture/bibliotheques-et-archives/archives.html">Archives of the City of Lausanne</a>.</p> <p>Please note that the annotation and extraction methodology, as well as the complete evaluation of performance, including HTR benchmark and post-correction performance is published in :</p> <ul> <li>Petitpierre R., Rappo L., Kramer M. (2023). <em>An end-to-end pipeline for historical censuses processing</em>. International Journal on Document Analysis and Recognition (IJDAR). doi: <a href="https://doi.org/10.1007/s10032-023-00428-9">10.1007/s10032-023-00428-9</a></li> </ul> <p>Tabular dataset resulting from automatic extraction are also available on Zenodo :</p> <ul> <li>Petitpierre R., Rappo L., Kramer M., di Lenardo I. (2023). <em>1805-1898 Census Records of Lausanne : a Long Digital Dataset for Demographic History</em>. Zenodo. doi: <a href="https://doi.org/10.5281/zenodo.7711640">10.5281/zenodo.7711640</a></li> </ul>
Train-B dataset for ICDAR2017 Competition on Handwritten Text Recognition on the READ Dataset (ICDAR2017 HTR). Batch 1 and Batch 2.
<p>Train-B Dataset. Dataset of pages without any layout or text line information. The corresponding transcripts are provided at page level with line breaks. It has 10k pages, though for convenience it is divided into two 5k page batches. This information is provided in PAGE format. </p> <p>This dataset is complementary to this other dataset:</p> <p>https://zenodo.org/record/439807#.WOIBZ3WLSkA</p> <p>More information at:</p> <p>https://scriptnet.iit.demokritos.gr/competitions/~icdar2017htr/</p> <p> </p>
Train-A dataset for ICDAR2017 Competition on Handwritten Text Recognition on the READ Dataset (ICDAR2017 HTR)
<p>Train-A Dataset of pages with manually revised baselines and the corresponding transcripts associated to them. This batch is small, 50 pages. Please, keep in mind that only the baselines have been manually corrected, The polygons associated to each line have not been manually reviewed. </p> <p>This dataset is complementary to this other dataset:</p> <p>https://zenodo.org/record/439811#.WOIF9HWLSkA</p> <p>More information at:</p> <p>https://scriptnet.iit.demokritos.gr/competitions/~icdar2017htr/</p>
6000 ground truth of VOC and notarial deeds 3.000.000 HTR of VOC, WIC and notarial deeds
<p>The National Archives of the Netherlands and Noord-Hollands Archief conducted a project using the Transkribus HTR (Handwritten Text Recognition) platform. The aim was to semi automatically transcribe 2 million pages of old Dutch texts.</p> <p>The transcribed archives are 17<sup>th</sup> and 18<sup>th</sup> century documents from the Dutch East-India Company (VOC). And 19th century notarial deeds from Noord-Hollands Archief and other archives in the provinces.</p> <p>In order to train the HTR software a team produced transcriptions of approximately 6000 scans. The scans are randomly selected from the dataset. With the transcriptions a model is trained that can recognize more than 90% of the characters correctly. Transkribus transcribed the 2 million scans automatically using the trained model.</p> <p>The following Transkribus HTR+ model has been trained for the text recognition: "IJsberg". More information about the model can be found <a href="https://readcoop.eu/transkribus/public-models/">here</a>. See the chapter "Dutch Handwriting". However, the Transkribus team retrained the model with <a href="https://readcoop.eu/transkribus/howto/how-to-train-pylaia-models-in-transkribus/">PyLaia</a> technology, which improved the HTR+ model. This PyLaia model is not publicly available.</p> <p>Later on, 1 million extra scans concerning the West India Company (WIC) were transcribed automatically without adding extra ground truth or training. These archives are from the 17<sup>th</sup> and 18<sup>th</sup> century.</p> <p>The <a href="https://github.com/knaw-huc/loghi">Loghi Handwritten Text Recognition Toolkit</a> has been added to the arsenal of the Nation Archives of the Netherlands. 1.05.11.14, Notarissen Suriname tot 1828 [digitaal duplicaat] has been processed with this tooling.</p> <p>The datasets published in Zenodo contain the ground truth (scans in JPG, transcription in PAGE XML) and the HTR results (in PAGE XML and TXT). See the overview below. Scroll to the bottom of the page to download the actual files.</p> <p>For more information on how the Dutch National Archive innovate on digital accessibility click <a href="https://www.nationaalarchief.nl/over-het-na/datalab-nationaal-archief">here</a>.</p> <p>For open data access of scans and inventories of the National Archives click <a href="https://www.nationaalarchief.nl/onderzoeken/open-data/open-data-archiefinventarissen-en-scans-van-archieven">here</a>.</p> <p><strong>Disclaimer</strong>: due to a variety of languages used and the bad state of the documents the HTR results of "1.05.21, Dutch series Guyana" can be of poor quality.</p> <p>--------------------------------------------------------------</p> <p><strong>Dataset HTR </strong><br>(Dataset, name archive, number archive, inventory numbers, link to inventory)<br><br><strong>The National Archives of the Netherlands</strong><br>HTR results VOC, VOC, 1.04.02, 7527-9540, <a href="https://www.nationaalarchief.nl/onderzoeken/archief/1.04.02/invnr/%40Deel%20I?query=1.04.02&search-type=inventory">EAD</a><br>HTR results 1.04.02, Oost-Indische Testamenten, 1.04.02, 6847-6897, <a href="https://www.nationaalarchief.nl/onderzoeken/archief/1.04.02/invnr/%40Deel%20I?query=1.04.02&search-type=inventory">EAD</a> <br>HTR results 1.05.01.01, Oude WIC, 1.05.01.01, 1-87, <a href="https://www.nationaalarchief.nl/onderzoeken/archief/1.05.01.01/invnr/%40A..?query=1.05.01.01&search-type=inventory">EAD</a><br>HTR results 1.05.01.02, Tweede WIC, 1.05.01.02, 1-1382, <a href="https://www.nationaalarchief.nl/onderzoeken/archief/1.05.01.02/invnr/%40VII~1324C2?query=1.05.01.01&search-type=inventory">EAD</a> <br>HTR results 1.05.02, Raad der Koloniën, 1.05.02, 1-192, <a href="https://www.nationaalarchief.nl/onderzoeken/archief/1.05.02/invnr/%40A?query=1.05.02&search-type=inventory">EAD</a><br>HTR results 1.05.03, Sociëteit van Suriname, 1.05.03, 1-566, <a href="https://www.nationaalarchief.nl/onderzoeken/archief/1.05.03/invnr/%40A?query=1.05.03&search-type=inventory">EAD</a><br>HTR results 1.05.05, Sociëteit van Berbice, 1.05.05, 1-445, <a href="https://www.nationaalarchief.nl/onderzoeken/archief/1.05.05/invnr/%40I?query=1.05.05&search-type=inventory">EAD</a><br>HTR results 1.05.06, Verspreide West-Indische stukken, 1.05.06, 1-1413, <a href="https://www.nationaalarchief.nl/onderzoeken/archief/1.05.06/invnr/%401?query=1.05.06&search-type=inventory">EAD</a><br>HTR results 1.05.21, Dutch series Guyana, 1.05.21, AB.1.1-BB.7.1, <a href="https://www.nationaalarchief.nl/onderzoeken/archief/1.05.21/invnr/%401.?query=1.05.21&search-type=inventory">EAD</a><br>HTR results 2.01.28.01, West-Indisch comité, 2.01.28.01, 1-254, <a href="https://www.nationaalarchief.nl/onderzoeken/archief/2.01.28.01/invnr/%40I?query=2.01.28.01&search-type=inventory">EAD</a><br>HTR results 2.01.28.02, Raad der Amerikaanse Bezittingen, 2.01.28.02, 1-264, <a href="https://www.nationaalarchief.nl/onderzoeken/archief/2.01.28.02/invnr/%40I.?query=2.01.28.02&search-type=inventory">EAD</a><br>HTR results 1.05.11.14, Notarissen Suriname tot 1828 [digitaal duplicaat], <a href="https://www.nationaalarchief.nl/onderzoeken/archief/1.05.11.14/invnr/%401?query=1.05.11.14&search-type=inventory&start=0&searchAfter=1%2C%401">EAD</a><br>HTR results 2.10.02, Koloniën, <a href="https://www.nationaalarchief.nl/onderzoeken/archief/2.10.02/invnr/%40A.?query=2.10.02&search-type=inventory&start=0&searchAfter=1%2C%40A.">EAD</a> (indices only)</p> <p><strong>Noord-Hollands archief</strong><br>HTR results NHA Notarial 1617, Oud notarieel archief Haarlem, 1617,1593-1805, <a href="https://noord-hollandsarchief.nl/bronnen/archieven?mivast=236&mizig=210&miadt=236&micode=1617&milang=nl&miview=inv2">EAD</a><br>HTR results NHA Notarial 1972, Nieuw notarieel archief Haarlem, 1972, 5-813, <a href="https://noord-hollandsarchief.nl/bronnen/archieven?mivast=236&mizig=210&miadt=236&micode=1972&milang=nl&miview=inv2">EAD</a></p> <p><strong>Brabants Historisch Informatie Centrum</strong><br>HTR results BHIC 7048 , Notarissen in Boxmeer, 1814-1935, 7048, 1-103, 162, <a href="https://proxy.archieven.nl/235/7203CD0B57624FD0B60EC6330DC7C083">EAD</a><br>HTR results BHIC 7128 , Notarissen in Grave, 1648-1935, 7128, 140-266, <a href="https://proxy.archieven.nl/235/73F9057F3EA34930829D1B55A9806B4D">EAD</a><br>HTR results BHIC 7637 , Notarissen in Sint-Oedenrode, 1642-1935, 7637, 17-78A, <a href="https://proxy.archieven.nl/235/4EBF145BB30A4CB3996DDD968914BAEE">EAD</a></p> <p><strong>Gelders Archief</strong><br>HTR results GA 0168, Notariële Archieven 1811-1925, 168, 64-69, 943-960, 1366-1395, 2472-2501, 3481-3485, 3904-3926, <a href="https://permalink.geldersarchief.nl/7A6A6A052F8A45EAAB6C7D4ECCA4041A">EAD</a></p> <p><strong>Groninger Archieven</strong><br>HTR results GRA 85, Notarissen te Appingedam (standplaats 1), 1811-1935, 85, 2-157, <a href="https://hdl.handle.net/21.12105/23FF3576C5AD4A97A847C6FB6E55323A">EAD</a><br>HTR results GRA 86, Notarissen te Appingedam (standplaats 2), 1812-1922, 86, 2-71, <a href="https://hdl.handle.net/21.12105/5C702F8661C6438EBC02025A9FC48446">EAD</a></p> <p><strong>Historisch Centrum Overijssel</strong><br>HTR results HCO 0122, Notarissen in Overijssel, 122, 5-48, 2044-2073, 3019-3047, 3733-3775, <a href="https://historischcentrumoverijssel.nl/archieven/?mivast=20&mizig=210&miadt=141&micode=0122&miview=inv2">EAD</a></p> <p><strong>The Utrecht Archives</strong><br>HTR results HUA 34-1, Notarissen in de provincie Utrecht, 1617-1895, 34-1, 928-930, 2209-2330, <a href="https://hetutrechtsarchief.nl/collectie/609C5BB45AB14642E0534701000A17FD">EAD</a></p> <p><strong>Regionaal Historisch Centrum Limburg</strong><br>HTR results RHCL 09.009, Notarissen in de Arrondissementen Maastricht en Roermond, 1896-1905, 09.009, 9147-9279, <a href="http://www.archieven.nl/mi/1540/?mivast=1540&mizig=210&miadt=38&micode=09.009&miview=inv2">EAD</a></p> <p><strong>Tresoar</strong><br>HTR results Tresoar 26, Notarieel archief, 26, 1001-9028 (met hiaten), <a href="https://www.archieven.nl/nl/zoeken?mivast=0&mizig=210&miadt=36&micode=26&miview=inv2">EAD</a></p> <p><strong>Zeeuws Archief</strong><br>HTR results ZA 13.2, Notariële Archieven Zeeland 1906-1915, (1886) 1906-1915 (1925), 13.2, 1152-1163, 1261-1320, <a href="https://hdl.handle.net/21.12113/DA16099B241C49C8843B404710612498">EAD</a></p> <p><strong>Drents Archief</strong><br>HTR results DA 114.10, Notaris jhr.mr. J.A.G.van der Wijck te Assen, 114.10, 4-7, <a href="http://www.drentsarchief.nl/onderzoeken/archiefstukken?mivast=34&mizig=210&miadt=34&micode=0114.10&miview=inv2">EAD</a><br>HTR results DA 114.11, Notaris mr. D.A.M.de Fremery te Assen, 114.11, 1, <a href="http://www.drentsarchief.nl/onderzoeken/archiefstukken?mivast=34&mizig=210&miadt=34&micode=0114.11&miview=inv2">EAD</a><br>HTR results DA 114.18, Notaris mr. Warmolt van Roijen te Borger, 114.18, 2-7, <a href="http://www.drentsarchief.nl/onderzoeken/archiefstukken?mivast=34&mizig=210&miadt=34&micode=0114.18&miview=inv2">EAD</a><br>HTR results DA 114.19, Notaris mr. Ernst Sigismund. Cornets de Groot te Borger, 114.29, 2-8, <a href="http://www.drentsarchief.nl/onderzoeken/archiefstukken?mivast=34&mizig=210&miadt=34&micode=0114.19&miview=inv2">EAD</a><br>HTR results DA 114.22, Notaris mr. Albertus Slingenberg te Coevorden, 114.22, 1-24, <a href="http://www.drentsarchief.nl/onderzoeken/archiefstukken?mivast=34&mizig=210&miadt=34&micode=0114.22&miview=inv2">EAD</a><br>HTR results DA 114.23, Notaris mr. Gozewienus Weys te Coevorden, 114.23, 4-13, <a href="http://www.drentsarchief.nl/onderzoeken/archiefstukken?mivast=34&mizig=210&miadt=34&micode=0114.23&miview=inv2">EAD</a><br>HTR results DA 114.28, Notaris mr. Johannes Beckeringh van Loenen te Dwingeloo, 114.29, 7-13, <a href="http://www.drentsarchief.nl/onderzoeken/archiefstukken?mivast=34&mizig=210&miadt=34&micode=0114.28&miview=inv2">EAD</a><br>HTR results DA 114.39, Notaris mr. Gerrit ten Raa ten Gieten, 114.39, 6-14, <a href="http://www.drentsarchief.nl/onderzoeken/archiefstukken?mivast=34&mizig=210&miadt=34&micode=0114.39&miview=inv2">EAD</a><br>HTR results DA 114.45, Notaris mr. Hendrik Jan Carsten te Hoogeveen, 114.45, 17-36, <a href="http://www.drentsarchief.nl/onderzoeken/archiefstukken?mivast=34&mizig=210&miadt=34&micode=0114.45&miview=inv2">EAD</a><br>HTR results DA 114.54, Notaris mr. Warmold Lunsingh Tonckens te Meppel, 114.54, 18-26, <a href="http://www.drentsarchief.nl/onderzoeken/archiefstukken?mivast=34&mizig=210&miadt=34&micode=0114.54&miview=inv2">EAD</a></p> <p> </p> <p><strong>Dataset Ground Truth</strong><br>(Name archive, number archive, inventory numbers, link to inventory, type of dataset)</p> <p>Dataset: Notarial deeds Ground Truths of the trainingset</p> <ul> <li>Oud notarieel archief Haarlem, 1617, 495 random scans from 1593-1805, <a href="https://noord-hollandsarchief.nl/bronnen/archieven?mivast=236&mizig=210&miadt=236&micode=1617&milang=nl&miview=inv2">EAD</a>, GT Transcriptions</li> <li>Nieuw notarieel archief Haarlem, 1972, 952 random scans from 5-813, <a href="https://noord-hollandsarchief.nl/bronnen/archieven?mivast=236&mizig=210&miadt=236&micode=1972&milang=nl&miview=inv2">EAD</a>, GT Transcriptions</li> <li>(And 168 transcripties from 7 other archives.)</li> </ul> <p>Dataset: Notarial deeds Images of the trainingset,</p> <ul> <li>Nieuw notarieel archief Haarlem, 1972, 952 random scans from 5-813, <a href="https://noord-hollandsarchief.nl/bronnen/archieven?mivast=236&mizig=210&miadt=236&micode=1972&milang=nl&miview=inv2">EAD</a>, GT Scans</li> <li>Oud notarieel archief Haarlem, 1617, 495 random scans from 1593-1805, <a href="https://noord-hollandsarchief.nl/bronnen/archieven?mivast=236&mizig=210&miadt=236&micode=1617&milang=nl&miview=inv2">EAD</a>, GT Scans</li> <li>(And 168 scans from 7 other archives.)</li> </ul> <p><br>Dataset: VOC Ground Truths of the trainingset,<br>VOC, 1.04.02, 4735 random scans from 7527-9540, <a href="https://www.nationaalarchief.nl/onderzoeken/archief/1.04.02/invnr/%40Deel%20I?query=1.04.02&search-type=inventory">EAD</a>, GT Transcriptions</p> <p><br>Dataset: VOC Images of the trainingset,<br>VOC, 1.04.02, 4735 random scans from 7527-9540, <a href="https://www.nationaalarchief.nl/onderzoeken/archief/1.04.02/invnr/%40Deel%20I?query=1.04.02&search-type=inventory">EAD</a>, GT Scans</p> <p>--------------------------------------------------------------</p> <p>Version 3.0: The first HTR results from the VOC-collection are available in .txt format, Inventory numbers 7527-9540.</p> <p>Version 3.1: The HTR results from the VOC-collection are also available in PAGE xml format. </p> <p>Version 4.0: About 30 missing inventory numbers have been added to the VOC transcriptions. The HTR results of the Notarial Deeds from the NHA archives have been added. An example on full text searchable research can be found here (Dutch): <a href="https://kia.pleio.nl/groups/view/55812425/htr-en-ocr/blog/view/55814752/reconstructie-van-een-verijdelde-slavenopstand-met-behulp-van-automatische-handschriftherkenning-en-text-mining">https://kia.pleio.nl/groups/view/55812425/htr-en-ocr/blog/view/55814752/reconstructie-van-een-verijdelde-slavenopstand-met-behulp-van-automatische-handschriftherkenning-en-text-mining</a></p> <p>Version 5.0: Around a million pages of HTR results of the following archives have been added.</p> <p>Version 6.0: The HTR results of Oost-Indische Testamenten have been added. </p> <p>Version 7.0: The HTR results of the Brabants Historisch Informatie Centrum, Gelders Archief, Groninger Archieven, Historisch Centrum Overijssel, The Utrecht Archives, Regionaal Historisch Centrum Limburg, Tresoar, Zeeuws Archief and Drents Archief have been added.</p> <p>Version 7.1: A spreadsheet "ijsberg train-val.xlsx" has been added. The division of the training- and validationset of Ground Truth of the IJsberg model can be found here</p> <p>Version 8.0: HTR results of 1.05.11.14 have been added. The scans have been inferenced with Loghi.</p> <p>Version 8.1: HTR results of 2.10.02 indices have been added. The scans have been inferenced with Loghi.</p>
Random dataset of HTR-ed police ordinances Bern - to test Annif
<p>This dataset consists of 4450 (out of the original 4932) police ordinances that were identified by <span>Claudia Schott-Volm, <em>Repertorium der Policeyordnungen der Frühen Neuzeit (#7). Orte der Schweizer Eidgenossenschaft: Bern und Zürich</em> (Frankfurt am Main, 2006) for the Swiss City-State of Bern (1528-1798). </span></p> <p><span>They have been randomly placed in a test (500), training (3550) and validation set (500) to test the tool Annif. The HTR-recognition was done in 2020 with the - now terminated HTR+-model 'German-Kurrent_XVI-XVIII++French' in the then commonly used Expert Version of Transkribus (READ-COOP SCE).</span></p>
Dataset and evaluation for HTR models for Latin and French Medieval Documentary Manuscripts
<p><strong>1. Dataset presentation.</strong></p> <p>This is the dataset used to produce the HTR models applied to documentary Latin and French manuscripts presented in the paper: Sergio Torres Aguilar, Vincent Jolivet. <strong>Handwritten Text Recognition for Documentary Medieval<br> Manuscripts. </strong>2022. https://hal.science/hal-03892163</p> <p>The dataset contains mostly charters and registers from the Late-medieval period (12th-15th). The training and evaluation, entailing 1855 pages, 120k lines of text and almost 1M tokens, were conducted using three freely available ground-truth corpora :</p> <p><strong>The Alcar-HOME database </strong>: https://zenodo.org/record/5600884</p> <p><strong>The e-NDP corpus </strong>: https://github.com/chartes/e-NDP_HTR</p> <p><strong>The Himanis project </strong>: https://zenodo.org/record/5535306</p> <p>The final model operates in a multilingual environment (Latin and French) and it is able to recognize several Latin script families (mostly <em>Textualis</em> and <em>Cursiva</em>) in documents produced in ca. 12th - 15th centuries. During the evaluation the models shows an accuracy of <strong>94.01%</strong> on the validation set and a CER (character error ratio) of about <strong>0.12</strong> to <strong>0.17</strong> on four external unseen datasets. A fine-tuning exercise using 10 ground-truth pages can raise these results to a CER between <strong>0.06</strong> to <strong>0.10</strong> respectively.</p> <p> </p> <p><strong>2. Dataset contents .</strong></p> <p>a) <em>GT_list : </em>List containing the GT file names which constitute the training, evaluation and test sets. The images and transcriptions can be downloaded from their original repositories.</p> <p>b) <em>Training :</em> Contains the training and testing results (evaluation and prediction files) presented in the original paper for the two training phases: Regular (Textualis and Cursiva separated training) and Quartiles (mixed training by quartiles).</p> <p>c) <em>Useful_scripts :</em> Scripts to produce the HTR metrics (CER, WER, SER) and plot the model's accuracy.</p> <p>d) <em>Best_model :</em> Contains the best multilingual and multi-script model.</p>
Test-B1 ICDAR2017 Competition on Handwritten Text Recognition on the READ Dataset (ICDAR2017 HTR)
<ul> <li><strong>Test-B1</strong>: a batch of page images annotated with the geometry of regions where to detect text line and recognize.</li> </ul>
Test-B2 ICDAR2017 Competition on Handwritten Text Recognition on the READ Dataset (ICDAR2017 HTR)
<p><strong>Test-B2</strong>: a batch of page images annotated with the geometry of regions where to detect text line and recognize.</p>
Dataset for ICDAR2017 Competition on Handwritten Text Recognition on the READ Dataset (ICDAR2017 HTR)
<p><strong>Train-A:</strong> Dataset of pages with manually revised baselines and the corresponding transcripts associated to them. This batch is small, 50 pages. Please, keep in mind that only the baselines have been manually corrected, The polygons associated to each line have not been manually reviewed.</p> <p><strong>Train-B:</strong> Dataset of pages without any layout or text line information. The corresponding transcripts are provided at page level with line breaks. It has 10k pages, though for convenience it is divided into two 5k page batches. This information is provided in PAGE format.</p> <p><strong>Test A:</strong> Dataset of pages with manually revised baselines. This batch has 65 pages. The polygons associated to each line have not been manually reviewed.</p> <p><strong>Test-B1:</strong> The same dataset of pages of the Test A, but annotated only with the geometry of regions. Text line information is not provided. </p> <p><strong>Test-B2:</strong> Dataset of page images annotated with the geometry of regions where to detect text line and recognize. It has 57 pages.</p> <p><strong>Baseline.tgz:</strong> Baseline system trained using the first 40 pages of Train-A. The system is based on the deep learning toolkit to transcribe handwritten text images called Laia.</p> <p>More information at:</p> <p>https://scriptnet.iit.demokritos.gr/competitions/~icdar2017htr/</p> <p> </p>
TRIDIS: HTR model for Multilingual Medieval and Early Modern Documentary Manuscripts (11th-16th)
<p><strong>TRIDIS (Tria Digita Scribunt)</strong> is a Handwriting Text Recognition model trained on semi-diplomatic transcriptions from medieval and Early Modern Manuscripts. It is suitable for work on documentary manuscripts, that is, manuscripts arising from legal, administrative, and memorial practices more commonly from the Late Middle Ages (13th century and onwards). It can also show good performance on documents from other domains, such as literature books, scholarly treatises and cartularies providing a versatile tool for historians and philologists in transforming and analyzing historical texts.</p> <p>A paper presenting the first version of the model is available here: Sergio Torres Aguilar, Vincent Jolivet. <strong>Handwritten Text Recognition for Documentary Medieval Manuscripts. </strong>Journal of Data Mining and Digital Humanities.<strong> </strong>2023. https://hal.science/hal-03892163</p> <p> </p> <h3>Transcriptions rules :</h3> <p>Since the majority of the training documents come from diplomatic editions, the transcriptions were <strong>normalized</strong> to contemporary reading standards, and <strong>abbreviations were expanded</strong> with the aim of facilitating a more fluid reading of the document.</p> <p>The following rules were applied:</p> <ul> <li>The abbreviations have been expanded, both those by suspension (<code>facimꝰ</code> ---> <code>facimus</code>) and by contraction (<code>dñi</code> --> <code>domini</code>). Likewise, those using conventional signs (<code>⁊</code> --> <code>et</code> ; <code>ꝓ</code> --> <code>pro</code>) have been resolved. </li> <li>The named entities (names of persons, places and institutions) have been <code>capitalized</code>. The beginning of a block of text as well as the original capitals used by the scribe are also capitalized.</li> <li>The consonantal <code>i</code> and <code>u</code> characters have been transcribed as <code>j</code> and <code>v</code> in both French and Latin.</li> <li>The punctuation marks used in the manuscript like: <code>.</code> or <code>/</code> or <code>|</code> have not been systematically transcribed as the transcription has been standardized with modern punctuation.</li> <li>Corrections and words that appear cancelled in the manuscript have been transcribed surrounded by the sign <code>$</code> at the beginning and at the end.</li> </ul> <p> </p> <h3>Versions :</h3> <p><strong>Version 1 </strong>of the model was trained on charters and registers dataset from the Late Medieval period (12th-15th centuries). The training and evaluation involved 1855 pages, 120k lines of text, and almost 1M tokens, conducted using three freely available ground-truth corpora:</p> <ul> <li>The Alcar-HOME database: <a href="../record/5600884" target="_new">https://zenodo.org/record/5600884</a></li> <li>The e-NDP corpus: <a href="../record/7575693" target="_new">https://zenodo.org/record/7575693</a></li> <li>The Himanis project: <a href="../record/5535306" target="_new">https://zenodo.org/record/5535306</a></li> </ul> <p><strong>Version 2</strong> of the model has added new datasets from feudal books and legal proceedings (14th-16th centuries), incorporating an additional 115k lines and more than 1.2M tokens to the previous version using other corpora like:</p> <ul> <li>Königsfelden Abbey corpus: <a href="../record/5179361" target="_new">https://zenodo.org/record/5179361</a></li> <li>Monumenta Luxemburgensia.</li> </ul> <p> </p> <h3>Accuracy</h3> <p>TRIDIS was trained using a CNN+RNN+CTC architecture within the Kraken suite (https://kraken.re/). This final model operates in a multilingual environment (Latin, Old French, and Old Spanish) and is capable of recognizing several Latin script families (mostly Textualis and Cursiva) in documents produced circa 11th - 16th centuries. During evaluation, the model showed an accuracy of 93.1% on the validation set and a CER (Character Error Ratio) of about 0.11 to 0.15 on four external unseen datasets. Fine-tuning the model with 10 ground-truth pages can improve these results to a CER of between 0.06 to 0.10, respectively.</p> <h3>Other formats</h3> <p>The ground truth used for version 2 was also employed to train a Transformer HTR model that combines TrOCR as the encoder with a RoBERTa medieval model as the decoder. This model exhibits a slighly better performance in terms of CER metrics to the current TRIDIS version and shows an improved WER by about 25%. The model is available on the Hugging Face Hub: <a href="https://huggingface.co/magistermilitum/tridis_HTR">magistermilitum/tridis_HTR</a></p>
HTR Model Spanish Gothic Incunabula (HSMS)
<p>The <a href="https://www.transkribus.org/model/spanish-gothic-incunabula" target="_blank" rel="noopener">Spanish Gothic Incunabula (HSMS)</a> is conceived to be uploaded inside Transkribus platform (READ Coop) to perform a training and create an PyLaia model for the automated recognition of Spanish incunabula in Gothic script printed between 1472 and 1500. It can be used for post-incunabula (up to 1520).<br><br>The transcription model follows the rules set by the Hispanic Seminary of Medieval Studies in 1977 (<a href="https://hispanicseminary.org/manual-en.htm" target="_blank" rel="noopener">newest version</a>). The rules applied are:</p> <ul> <li>Abbreviated words are expanded and the expanded text is enclosed between < >: <em>q<ue></em></li> <li>Superscripted letters are followed by a grave accent: <em>q<u>i`en</em></li> <li>All <em>ç</em> and <em>ñ</em> are transcribed as <em>ç</em> and <em>ñ</em></li> <li>No attemp has been made to normalize spacing</li> <li>All punctuations signs are kept</li> <li>All abbreviated nasals before <em>b</em> or <em>p</em> are transcribed as <n>. It is up to the editors if they should be changed into <em>m</em>.</li> <li>Abbreviated v' (tilde over v, or small slash v) that can be expanded as v<ir> or v<er> is expanded as v<er>. It is up to the editor if they should be changed to v<ir>.</li> <li>Tironian <em>et</em> is transcribed as & (ampersand)</li> <li>Pilcrows are transcribed as ¶</li> </ul> <p>The model is built on 200 openings (verso-recto) drawn from 20 books printed by five different workshops form Sevile, Zaragoza, Burgos, Toledo and Pamplona. They Train Set consist 180152 words, distribuited over 24061 lines. The CER on the Train Set is 0.20 % and on the Validation Set 0.77 %.</p> <p> </p>
HTR - Araucania - XIX manuscript
<p><strong>General</strong></p> <p>Ground Truth dataset for Spanish 19th typewritten OCR (XML-ALTO)</p> <p>The archives come from the events of the Occupation of Araucania (1850-1881) in Chile. They are archived in the 'Colección manuscritos' of the Archivo Central Andres Bello - Universidad de Chile.</p> <p>Thereby, it is not possible to publicly distribute the images (.jpg)</p> <p>To use them for segmentation/recognition model training, please contact me : archivo.central@uchile.cl</p> <p> </p> <p><strong>Methodology</strong></p> <p>Transcription rules :</p> <p>- xxx for blurred or unreadable characters<br> - ^+letters for superscript letters<br> - ⁋ for new paragraph</p> <p>Using the Kraken OCR engine in finetuning with the Menu_MacFrench template. A template uses the NFKD method.</p> <p>Segmonto ontology</p> <p> </p> <p><strong>Evaluation</strong></p> <table> <tbody> <tr> <td><strong>Name</strong></td> <td><strong>Quantity (GT)</strong></td> <td> <table> <thead> <tr> <th> </th> <th><strong>Val_acc</strong></th> </tr> </thead> </table> </td> <td><strong>Test_acc</strong></td> <td><strong>CER</strong></td> <td><strong>WER</strong></td> </tr> <tr> <td><strong>HTR-Araucania_XIX</strong></td> <td>180</td> <td> <table> <tbody> <tr> <td>0,90354</td> </tr> </tbody> </table> </td> <td> <table> <tbody> <tr> <td> </td> <td>0,8673</td> </tr> </tbody> </table> </td> <td>0,05598</td> <td>0.21423</td> </tr> <tr> <td><strong>HTR-Araucania_XIX_NFKD</strong></td> <td>180</td> <td>0,89872</td> <td>0,8563</td> <td> <table> <tbody> <tr> <td> </td> <td>0,06646</td> </tr> </tbody> </table> </td> <td>0.24963</td> </tr> </tbody> </table> <p><strong>Others</strong></p> <p>JSONL file for NER annotation in <code>ner/</code> (MISC, LOC, PERS, ORG, DATE)</p>
Dataset for the paper "Ground-truth Free Evaluation of HTR on Old French and Latin Medieval Literary Manuscripts"
<p>This dataset was used in the context of the article <em>Ground-truth Free Evaluation of HTR on Old French and Latin Medieval Literary Manuscripts</em>, at the Computational Humanities Research 2022 conference.</p> <p>Predictions.zip contains the XML ALTO for the HTR and Segmentation prediction of around 10 pages of 1900 manuscripts from the Bibliothèque nationale de France.</p> <p>Varying ground truth contains the original training material for training a classifier for CER classification (see the article).</p> <p>The ManuscriptsIIIF.csv contains metadata about the manuscripts.</p> <p>This work was funded by the DIM MAP under the CREMMALab funding.</p>
Test A for the ICDAR2017 Competition on Handwritten Text Recognition on the READ Dataset (ICDAR2017 HTR)
<p><strong>Test A. </strong>A batch of page images annotated with baselines.</p>
HTR model NIOD_WarLet_1935-1950_NoBasemodel
<p>The HTR model ‘NIOD_WarLet_1935-1950_NoBasemodel’ was trained using 968 ‘Ground Truth’ transcriptions of high-resolution scans of various handwritten letters. These letters are all written in Dutch and originate from the period 1935-1950. The training set contains personal correspondence from a wide variety of letter writers (e.g., children, soldiers, Jewish people in hiding). These personal correspondences are all part of the archival collection known as ‘247 Correspondentie’ held by the NIOD Institute for War, Holocaust, and Genocide Studies in Amsterdam.<br> <br> This model was created as part of the project ‘First-Hand Accounts of War: War letters (1935-1950) from NIOD digitised’. All documents used for training and validation were scanned and transcribed within this project. This project ran from 2020 to 2023 and was funded by the Mondriaan Fund, the Dutch Ministry of Health, Welfare, and Sport, and the NIOD Institute for War, Holocaust, and Genocide Studies in Amsterdam.<br> <br> The ‘Ground Truth’ training set is created by project members Annelies van Nispen, Carlijn Keijzer and Milan van Lange. Additional transcription and correction of ‘Ground Truth’ transcriptions was performed under supervision of Muriël Bouman by citizen scientists Hillebrand Verkroost, Bart Cohen, Evelien Bachrach, Marjo Janssens, and Cocky Sietses. The validation set contains a sample of 17 ‘Ground Truth’ transcriptions from various writers and sub-collections. Due to legal restrictions only a limited sample of the training set is published publicly.<br> <br> The model is trained using PyLaia HTR, max. 500 epochs (321 epochs trained), learning rate 0.0003. No basemodel was used. See also: https://readcoop.eu/model/niod_warlet_1945-1950_nobasemodel/</p>
University of Denver Collections as Data - HTR Train and Validation Set JCRS_2020_5_27
<p><a href="https://zenodo.org/api/files/333ecb88-1f48-4ffd-b39e-5e70b800c276/HTR_Train_Set_JCRS_2020_5_27.zip">HTR_Train_Set_JCRS_2020_5_27.zip</a> <br> Description</p>
Differential expression of miRNAs in EGF treated HTR-8/SVneo versus untreated control by next generation sequencing
GEO Series GSE124585. Homo sapiens. 2 samples. Type: Non-coding RNA profiling by high throughput sequencing.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.