Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
14
datasets available to search
ShareScore release 0.9.0
Dataset results
14 results for “Medieval Manuscripts”
Medieval manuscripts and their migrations: Using SPARQL to investigate the research potential of an aggregated Knowledge Graph
<p>This dataset contains the <strong>SPARQL queries</strong> presented and discussed in our article published in <em>Digital Medievalist</em> 2022 (as a PDF file), together with the <strong>results of those queries</strong> as CSV files. The query and step numbering follows that given in the article.</p> <p>The queries can be run against the SPARQL endpoint for the <strong>Mapping Manuscript Migrations</strong> project: <a href="https://ldf.fi/mmm/sparql">https://ldf.fi/mmm/sparql</a></p> <p>The full <strong>Mapping Manuscript Migrations dataset </strong>can also be downloaded from the Zenodo repository and installed in your own triple store: <a href="https://zenodo.org/record/4440464">https://zenodo.org/record/4440464</a></p> <p>When copying and pasting these SPARQL queries into a SPARQL client like <a href="https://yasgui.triply.cc/">YASGUI</a>, please check that the line numbering has been copied over correctly. Copying from a PDF file can sometimes break a single long line into multiple separate lines, which will cause a SPARQL validation error.</p> <p>The CSV files contain the results of the queries when run against the Mapping Manuscript Migrations SPARQL endpoint as of 17 December 2021. Please note that Query 2, Step 2, produces no results, so a CSV file has not been provided.</p> <p>The<strong> Mapping Manuscript Migrations portal </strong>can be found at <a href="https://mappingmanuscriptmigrations.org/en/">https://mappingmanuscriptmigrations.org/en/ </a></p> <p>SPARQL tutorials are included in the project's <strong>GitHub documentation</strong>: <a href="https://mapping-manuscript-migrations.github.io/">https://mapping-manuscript-migrations.github.io/</a></p>
MPS Data set with images of medieval charters for handwriting-style based dating of manuscripts
<pre>The MPS benchmark data set for handwritten manuscript dating ____________________________________________________________ This data set is collected for the Dutch NWO project: Medieval Paleographical Scale (MPS) by Petros Samara Project website: http://application02.target.rug.nl/monk/Projects/MPS/ Copyright (c) Huygensinstituut, Den Haag, 2016 University of Groningen, 2016. All rights reserved. Organisation of the data: Each .tar.gz file contains a number of NetPBM images. The format is chosen because of its simplicity. Also, there is no doubt about lossy compression in the processing chain. The file names are of the format 'MPS<year>_<seqnr>.ppm', for example, 'MPS1300_0056.ppm'. Note: the files are not in a separate directory, they will be extracted in place. However, due to the unique naming, there is no problem extracting them in one single current (destination) directory. The actual type of the image can be gray scale (.pgm) or color (.ppm), in '8-bit DirectClass' according to ImageMagick's 'identify' tool. The images were cropped out of larger photographs because of irrelevant elements such as a Kodak color calibrator and non-text content such as supporting surface (table) backgrounds, seals (emblems), ribbons, etc. No effort has been made to obtain a balanced set of samples over years: the given frequencies of occurrence in archives are used. There is evidently less data in years before 1375 A.D. while some periods provides us with ample data for historical reasons (e.g, 1450 A.D.). It would have been a pity if the scarce years had determined and limited the size of this data set. Selection criteria for data reduction, whether random or systematic, would have been arbitrary. In any case, these images were used in our publications, such that the performance results of future attempts on manuscript dating can be compared with earlier results. The performances that have been reached using our algorithms are in the order of an MAE (mean average error) of 10 years. If you have any questions, please contact us: Sheng He (heshengxgd@gmail.com) Petros Samara (petros.samara@huygens.knaw.nl) Jan Burgers (jan.burgers@huygens.knaw.nl) Lambert Schomaker (L.Schomaker@ai.rug.nl) Please cite our papers if you use this data set: [1] Sheng He, Petros Samara, Jan Burgers, Lambert Schomaker. Image-based historical manuscript dating using contour and stroke fragments. Pattern Recognition(PR), Vol. 59, pp. 159-171, 2016 [2] Sheng He, Petros Samara, Jan Burgers, Lambert Schomaker. Towards style-based dating of historical documents. International Conference on Frontiers in Handwriting Recognition(ICFHR), Crete, Greece, 2014 [3] Sheng He, Petros Samara, Jan Burgers, Lambert Schomaker. Multiple-Label Guided Clustering Algorithm for Historical Document Dating and Localization IEEE Trans. on Image Processing, Vol. 25(11), Nov. 2016. http://ieeexplore.ieee.org/document/7551181/</pre> <p>Data are collected thanks to Dutch NWO grant project 380-50-006</p>
The e-NDP project : collaborative digital edition of the Chapter registers of Notre-Dame of Paris (1326-1504). Ground-truth for handwriting text recognition (HTR) on late medieval manuscripts.
<p>The <a href="https://endp.hypotheses.org/">e-NDP project</a>, funded by the ANR, is led by the <a href="https://lamop.hypotheses.org/6870">LaMOP</a> (Julie Claustre and Darwin Smith).</p> <p>The project's partners are the Archives nationales, the Bibliothèque nationale de France (Department of Manuscripts, Bibliothèque de l'Arsenal), the École nationale des chartes and the Bibliothèque Mazarine.</p> <p>The e-NDP project aims at renewing our knowledge on <strong>Notre-Dame de Paris cathedral</strong> through the creation of a collaborative digital edition of the registers of its Chapter (1326-1504, <em>AN LL 105-128</em>), the community of 51 canons meeting three times a week on set days to take all administrative, financial and practical decisions pertaining to the cathedral, its estate and the society living in its cloister. This corpus has never been the object of a comprehensive study to understand the workings and history of this urban enclave and powerful community. The collaborative digital edition is based on a process of<strong> handwriting text recognition (HTR)</strong>, tested and supervised by scholars, researchers and engineers combining expertise in Medieval history, paleography, philology and digital humanities. The edition shall allow a better insight into the Chapter’s administration, into its economical and political power within Paris, and the relationships it maintained with other institutions in the city.</p> <p> </p> <p><strong>Section 1 : The e-NDP ground-truth dataset for Handwriting text recognition.</strong></p> <p>The full e-NDP corpus kept today in the French National Archives and was entirely digitized and described in its <a href="https://www.siv.archives-nationales.culture.gouv.fr/siv/rechercheconsultation/consultation/ir/consultationIR.action?formCaller=GENERALISTE&irId=FRAN_IR_059635">catalog</a> in 2022.</p> <p>The first major goal of the e-NDP projet is to propose a first automatic transcription of the 14k pages composing the 26 chapter registers. To achieve this goal representative samples from each one of the volumes were selected and transcribed in order to train a specialized HTR model able to propose a high quality automatic transcription. The collected ground-truth released on this repository currently has <strong>512 pages from the 26 registers</strong> of the cathedral chapter preserved in the National Archives (LL105 - LL128, <strong>1326-1504</strong>). The transcriptions were manually completed in <strong>two rounds</strong> by a group of 12 contributors, historians and paleographers, over the course of 2021-2022 using <a href="https://escriptorium.paris.inria.fr/">eScriptorium </a>as annotation environment. </p> <p> </p> <p><strong>Ground-truth features :</strong></p> <p><br> <em>Number of hands </em>: according to our estimates no fewer than 18 main hands were involved in the writing of the registers during the medieval period. </p> <p><em>Language</em> : More than 98% of the content of the registers was written in Latin, the rest in French. The exact percentage is hard to estimate because the vernacular language is often used in formulae, notes and comments. It is rare to find entire pages or blocks written in French. </p> <p><em>Script family</em> : The registers were written using a Cursive script (ca. late XIIIe - XVIe).</p> <p><em>Documental typology</em> : The volumes containing the chapter conclusions were conceived to serve as memorial records, but above all as documents for regular use and consultation in the daily practice of administration and management. In diplomatics the notion of "documentary manuscripts" is used to describe this kind of sources also by opposition to books and litterary or normative manuscripts.</p> <table align="center"> <caption><strong>Ground truth statistics</strong></caption> <tbody> <tr> <th>Text units</th> <th>Count</th> </tr> <tr> <td>Pages</td> <td>512</td> </tr> <tr> <td>Annotated regions (see section 2)</td> <td>2448</td> </tr> <tr> <td>Lines of text</td> <td>34231</td> </tr> <tr> <td>Tokens</td> <td>205083</td> </tr> <tr> <td>Characters</td> <td>3320407</td> </tr> </tbody> </table> <p> </p> <p><strong>Rules of transcription :</strong></p> <ul> <li>The abbreviations have been resolved, both those by suspension (<code>facimꝰ</code> ---> <code>facimus</code>) and by contraction (<code>dñi</code> --> <code>domini</code>). Likewise, those using conventional signs (<code>⁊</code> --> <code>et</code> ; <code>ꝓ</code> --> <code>pro</code>) have been resolved. </li> <li>The named entities (names of persons, places and institutions) have been <code>capitalized</code>. The beginning of a block of text as well as the original capitals used by the notary are also capitalized.</li> <li>The consonantal <code>i</code> and <code>u</code> characters have been transcribed as <code>j</code> and <code>v</code> in both French and Latin.</li> <li>The punctuation marks used in the text: <code>.</code> and <code>/</code> have been transcribed, but the transcription has not been standardized with modern punctuation.</li> <li>Corrections and words that appear cancelled in the manuscript have been transcribed surrounded by the sign <code>$</code> at the beginning and at the end.</li> <li>More specific transcription rules can be found into the file <code>transcription_guidelines.pdf</code></li> </ul> <p> </p> <p><strong>Section 2. e-NDP Layout Segmentation.</strong></p> <p>Layout segmentation is a compulsory step before HTR recognition in order to distinguish sections and regions inside a document. This process intend to separate interdependant page zones to produce a recognition in a section-sequence order and not in a line-sequence order which mix textual and peri-textual content.</p> <p>The regions of 364 pages (see <code>GT-layout_list</code>) of the e-NDP corpus were annotated using a 5 sections vocabulary (see <code>endp_layout_regions</code>) in order to describe the page distribution in all the 26 volumes :</p> <ol> <li><em>Block</em> : All the central text blocks, that normally corresponds to the main content called "conclusions" in registers.</li> <li><em>Liste</em> : List of names of the canons who were present during the meeting. Normally located before the <em>conclusions</em>.</li> <li><em>Entrée</em> : Marginal notes or entries to inform about the content of <em>conclusions</em>.</li> <li><em>Date</em> : Paragraph contending the date. Normally at the head of a <em>conclusion</em>, but separate of the main body.</li> <li><em>Numérotation</em> : Page numbers in roman or arabic. Usually appear in the top corners of the pages.</li> </ol> <table align="center"> <caption><strong>Layout GT statistics</strong></caption> <tbody> <tr> <th>Region</th> <th>Count</th> </tr> <tr> <td>block</td> <td>833</td> </tr> <tr> <td>liste</td> <td>431</td> </tr> <tr> <td>date</td> <td>448</td> </tr> <tr> <td>entrée</td> <td>205</td> </tr> <tr> <td>numérotation</td> <td>531</td> </tr> </tbody> </table> <p> </p> <p><strong>Section 3. The e-NDP HTR modeling.</strong></p> <p>The e-NDP project has progressively trained several HTR models adapted to work on late medieval cursive in order to accelerate the production of ground truth. Currently the best model delivers an average <strong>CER (Character error ratio) of 9.7%</strong> in handwriting recognition on the 26 registers (see <code>endp_learning_curve</code>) and can serve as generalist model for other manuscripts of the same period and similar script family. These models and their training implementation details can be found in the project's github <a href="https://github.com/chartes/e-NDP_HTR">repository</a>. </p> <p>Additionally, the automatic HTR transcriptions of the 26 registers (14k pages, 4.5M tokens) enriched with lexical and semantical information has been the subject of a first <a href="https://nosketch-engine.lamop.fr/#dashboard?corpname=endp">online publication</a> using the NoSketch engine that allows advanced data mining based on the combination of data, metadata and NLP features. </p> <p> </p> <p><strong>Section 4. Dataset content.</strong></p> <p>This zip dataset contains :</p> <p>- <code>HTR_ground_truth</code> : Two folders containing the jpg / jpeg images and their curated transcriptions in PAGE XML format.</p> <p>- <code>images_docs</code> : 4 files illustrating the different phases of the project (list of GT for layout segmentation, layout ontologie, transcription guideline and HTR evaluation curves)</p>
Early Medieval Latin Manuscripts Transmitting the Text of the Etymologiae of Isidore of Seville: an excel datasheet
<p>This excel file contains structured and formalized data about all surviving and identified early medieval Western manuscripts containing the text of the <em>Etymologiae</em> of Isidore of Seville, fully or partially. It records information about the place of origin, provenance, preservation, the date of origin, material properties, script, content, the state of preservation, presence of notable features, online representation, and bibliography of 507 manuscripts (v2.3.4: 496 manuscripts; v 2.3.2: 492 manuscripts; v2.1: 484 manuscripts; v2.0: 478 manuscripts) dated from the seventh to the first half of the eleventh centuries. This datasheet corresponds to the data published in the <em>Innovating Knowledge</em> database on 23 September 2024 (v2.3.4: 26 September 2023; v2.3.2: 2 August 2022; v2.1: 9 December 2021; v2.0: 12 October 2021), at: <a href="https://db.innovatingknowledge.nl/">db.innovatingknowledge.nl</a><br> </p> <p>More information about the <em>Innovating Knowledge</em> project is available at: <a href="https://innovatingknowledge.nl">innovatingknowledge.nl</a></p>
Faithful Transcriptions Data Set: TEI/XML-encoded Transcriptions of Medieval Theological Manuscripts
<p>From May to July 2021, the Berlin State Library and the Leipzig University Library jointly organized the Transcribathon <a href="https://lab.sbb.berlin/events/faithful-transcriptions/">“Faithful Transcriptions”</a>, a digital crowd souring project on medieval theological manuscripts. During the project, over 100 participants produced TEI/XML-encoded transcriptions in the IIIF workspace of the <a href="https://handschriftenportal.de/">Handschriftenportal</a>, which is currently being developed.</p> <p>The <a href="https://lab.sbb.berlin/datensets-transkribathon/">Faithful Transcriptions Data Set</a> contains 181 pages with 8.952 text lines from 12 manuscripts in German, Dutch, and Latin. The medieval scripts include Textura, Textualis, Gothic Cursiva, and Bastarda. The transcriptions are linked to the coordinates of the digitized manuscript image on text line level.</p> <p>--------------------------------------------</p> <p>Von Mai bis Juli 2021 richtete die Staatsbibliothek zu Berlin in Kooperation mit der Universitätsbibliothek Leipzig den Transkribathon <a href="https://lab.sbb.berlin/events/faithful-transcriptions/">„Faithful Transcriptions“</a> aus, ein digitales Crowd-Sourcing-Projekt zu theologischen Handschriften des Mittelalters. Über 100 Teilnehmende fertigten dabei TEI/XML-codierte Transkriptionen in der IIIF-basierten Arbeitsumgebung des aktuell in Entwicklung befindlichen <a href="https://handschriftenportal.de/">Handschriftenportals</a> an. </p> <p>Das <a href="https://lab.sbb.berlin/datensets-transkribathon/">Faithful Transcriptions-Datenset</a> enthält 181 Seiten mit 8.952 Textzeilen aus 12 Handschriften in deutscher, niederländischer und lateinischer Sprache. Die mittelalterlichen Schriften reichen von Textura über Textualis und Gotische Kursive bis hin zur Bastarda. Die Transkriptionen sind mit den Bildkoordinaten des Handschriftendigitalisats auf Textzeilenebene verknüpft. </p>
Dataset and evaluation for HTR models for Latin and French Medieval Documentary Manuscripts
<p><strong>1. Dataset presentation.</strong></p> <p>This is the dataset used to produce the HTR models applied to documentary Latin and French manuscripts presented in the paper: Sergio Torres Aguilar, Vincent Jolivet. <strong>Handwritten Text Recognition for Documentary Medieval<br> Manuscripts. </strong>2022. https://hal.science/hal-03892163</p> <p>The dataset contains mostly charters and registers from the Late-medieval period (12th-15th). The training and evaluation, entailing 1855 pages, 120k lines of text and almost 1M tokens, were conducted using three freely available ground-truth corpora :</p> <p><strong>The Alcar-HOME database </strong>: https://zenodo.org/record/5600884</p> <p><strong>The e-NDP corpus </strong>: https://github.com/chartes/e-NDP_HTR</p> <p><strong>The Himanis project </strong>: https://zenodo.org/record/5535306</p> <p>The final model operates in a multilingual environment (Latin and French) and it is able to recognize several Latin script families (mostly <em>Textualis</em> and <em>Cursiva</em>) in documents produced in ca. 12th - 15th centuries. During the evaluation the models shows an accuracy of <strong>94.01%</strong> on the validation set and a CER (character error ratio) of about <strong>0.12</strong> to <strong>0.17</strong> on four external unseen datasets. A fine-tuning exercise using 10 ground-truth pages can raise these results to a CER between <strong>0.06</strong> to <strong>0.10</strong> respectively.</p> <p> </p> <p><strong>2. Dataset contents .</strong></p> <p>a) <em>GT_list : </em>List containing the GT file names which constitute the training, evaluation and test sets. The images and transcriptions can be downloaded from their original repositories.</p> <p>b) <em>Training :</em> Contains the training and testing results (evaluation and prediction files) presented in the original paper for the two training phases: Regular (Textualis and Cursiva separated training) and Quartiles (mixed training by quartiles).</p> <p>c) <em>Useful_scripts :</em> Scripts to produce the HTR metrics (CER, WER, SER) and plot the model's accuracy.</p> <p>d) <em>Best_model :</em> Contains the best multilingual and multi-script model.</p>
Secondary ion mass spectrometry, a powerful tool for revealing ink formulations and animal skins in medieval manuscripts
<p>Book production by medieval scriptoria has gained growing interest in recent studies. In this context, identifying ink compositions and parchment animal species from illuminated manuscripts is of great importance. Here, we introduce time-of-flight secondary ion mass spectrometry (ToF-SIMS) as a non-invasive tool to identify both inks and animal skins in manuscripts, at the same time. For this purpose, both positive and negative ion spectra in inked and non-inked areas were recorded. Chemical compositions of pigments (decoration) or black inks (text) were determined by searching for characteristic ion mass peaks. Animal skins were identified by data processing of raw ToF-SIMS spectra using principal component analysis (PCA). In illuminated manuscripts from the fifteenth to sixteenth century, malachite (green), azurite (blue), cinnabar (red) inorganic pigments, as well as iron-gall black ink, were identified. Carbon black and indigo (blue) organic pigments were also identified. Animal skins were identified in modern parchments of known animal species by a two-step PCA procedure. We believe the proposed method will find extensive application in material studies of medieval manuscripts, as it is non-invasive, highly sensitive and able to identify both inks and animal skins at the same time, even from traces of pigments and tiny scanned areas.</p>
Secondary ion mass spectrometry, a powerful tool for revealing ink formulations and animal skins in medieval manuscripts
Open the record for dataset details and reuse information.
Networks of Shared Manuscript Transmission for Medieval European Vernacular Languages
<p>Codices that compile different textual units in the same physical object were common in the European Middle Ages. Many reasons guided scribes when collecting a variety of works and their analysis can offer scholars insight into scribal practices and the circulation of medieval works. In order to study this shared transmission of medieval texts, methods from network analysis offer the possibility of researching the phenomenon from a general perspective and discover fundamental trends. This paper deals with the methodological and practical foundations for such an approach. </p> <p>Modern digital databases of medieval manuscripts provide a big amount of relevant data for this type of research. In this presentation, I consider three online databases, each dealing with textual witnesses in vernacular languages: Handschriftencensus (German) Jonas (French and Occitan) and Philobiblon (Iberian languages). Each of these has different criteria on data collection and data modelling. For this reason, their analysis and comparison offers insight not only on medieval manuscripts and texts, but also on the consequences of different approaches when cataloguing and describing medieval manuscripts. </p>
etymologiae.ms: a database of the early medieval manuscripts of the Etymologiae of Isidore of Seville
<p>The Innovating Knowledge project is currently developing an online manuscript database that will bring together up-to-date information about the roughly 450 early medieval codices transmitting the most important medieval Latin encyclopaedia, the Etymologiae of Isidore of Seville. The objective of the database: not only to replace the outdated manuscript handlist of G.E. Anspach (the 1940s), but to work towards improved standards of digital publishing of manuscript data.</p> <p>Learn more about the Innovating Knowledge project at: <a href="https://www.youtube.com/redirect?redir_token=QUFFLUhqbkc1cnIySWVLTHhMNUxIMDFkVWE3UzZneUh5d3xBQ3Jtc0tseVhpZHozVkRMWG5lVTRlVExJRjc0THFnY2ItMGlYWWU0YWZPWUVkdEJkcnBpcmFrdEFRTGJ2RG1DUzJLZ0tBT0F0NUM4Z3ZuY3ZaOVRHMGRkdmI1X1JzbVUxRl94MWpkbUVnSHdzMnlVNHUyMHFKdw%3D%3D&q=https%3A%2F%2Fmittelalter.hypotheses.org%2F21234&event=video_description&v=IBIRfBHK0pU">https://mittelalter.hypotheses.org/21234</a></p> <p>Presented as a Lightning Talk for the Schoenberg Symposium 2020</p>
eCodicesNL: a virtual library for medieval manuscripts in Dutch collections
<p>This is a presentation about eCodicesNL: a new initiative to build a virtual library for medieval manuscripts in Dutch collections. The elements that go into this library are, just as in its model project e-codices Switzerland, high quality and complete images, high quality descriptions and an excellent search interface. The concept is simple, but the execution is not: what is the best model for creating such an aggregated library, and how can we deal with data that already exist? How can we facilitate collections to take care of their own material, and keep it updated within a collective system?</p>
TRIDIS: HTR model for Multilingual Medieval and Early Modern Documentary Manuscripts (11th-16th)
<p><strong>TRIDIS (Tria Digita Scribunt)</strong> is a Handwriting Text Recognition model trained on semi-diplomatic transcriptions from medieval and Early Modern Manuscripts. It is suitable for work on documentary manuscripts, that is, manuscripts arising from legal, administrative, and memorial practices more commonly from the Late Middle Ages (13th century and onwards). It can also show good performance on documents from other domains, such as literature books, scholarly treatises and cartularies providing a versatile tool for historians and philologists in transforming and analyzing historical texts.</p> <p>A paper presenting the first version of the model is available here: Sergio Torres Aguilar, Vincent Jolivet. <strong>Handwritten Text Recognition for Documentary Medieval Manuscripts. </strong>Journal of Data Mining and Digital Humanities.<strong> </strong>2023. https://hal.science/hal-03892163</p> <p> </p> <h3>Transcriptions rules :</h3> <p>Since the majority of the training documents come from diplomatic editions, the transcriptions were <strong>normalized</strong> to contemporary reading standards, and <strong>abbreviations were expanded</strong> with the aim of facilitating a more fluid reading of the document.</p> <p>The following rules were applied:</p> <ul> <li>The abbreviations have been expanded, both those by suspension (<code>facimꝰ</code> ---> <code>facimus</code>) and by contraction (<code>dñi</code> --> <code>domini</code>). Likewise, those using conventional signs (<code>⁊</code> --> <code>et</code> ; <code>ꝓ</code> --> <code>pro</code>) have been resolved. </li> <li>The named entities (names of persons, places and institutions) have been <code>capitalized</code>. The beginning of a block of text as well as the original capitals used by the scribe are also capitalized.</li> <li>The consonantal <code>i</code> and <code>u</code> characters have been transcribed as <code>j</code> and <code>v</code> in both French and Latin.</li> <li>The punctuation marks used in the manuscript like: <code>.</code> or <code>/</code> or <code>|</code> have not been systematically transcribed as the transcription has been standardized with modern punctuation.</li> <li>Corrections and words that appear cancelled in the manuscript have been transcribed surrounded by the sign <code>$</code> at the beginning and at the end.</li> </ul> <p> </p> <h3>Versions :</h3> <p><strong>Version 1 </strong>of the model was trained on charters and registers dataset from the Late Medieval period (12th-15th centuries). The training and evaluation involved 1855 pages, 120k lines of text, and almost 1M tokens, conducted using three freely available ground-truth corpora:</p> <ul> <li>The Alcar-HOME database: <a href="../record/5600884" target="_new">https://zenodo.org/record/5600884</a></li> <li>The e-NDP corpus: <a href="../record/7575693" target="_new">https://zenodo.org/record/7575693</a></li> <li>The Himanis project: <a href="../record/5535306" target="_new">https://zenodo.org/record/5535306</a></li> </ul> <p><strong>Version 2</strong> of the model has added new datasets from feudal books and legal proceedings (14th-16th centuries), incorporating an additional 115k lines and more than 1.2M tokens to the previous version using other corpora like:</p> <ul> <li>Königsfelden Abbey corpus: <a href="../record/5179361" target="_new">https://zenodo.org/record/5179361</a></li> <li>Monumenta Luxemburgensia.</li> </ul> <p> </p> <h3>Accuracy</h3> <p>TRIDIS was trained using a CNN+RNN+CTC architecture within the Kraken suite (https://kraken.re/). This final model operates in a multilingual environment (Latin, Old French, and Old Spanish) and is capable of recognizing several Latin script families (mostly Textualis and Cursiva) in documents produced circa 11th - 16th centuries. During evaluation, the model showed an accuracy of 93.1% on the validation set and a CER (Character Error Ratio) of about 0.11 to 0.15 on four external unseen datasets. Fine-tuning the model with 10 ground-truth pages can improve these results to a CER of between 0.06 to 0.10, respectively.</p> <h3>Other formats</h3> <p>The ground truth used for version 2 was also employed to train a Transformer HTR model that combines TrOCR as the encoder with a RoBERTa medieval model as the decoder. This model exhibits a slighly better performance in terms of CER metrics to the current TRIDIS version and shows an improved WER by about 25%. The model is available on the Hugging Face Hub: <a href="https://huggingface.co/magistermilitum/tridis_HTR">magistermilitum/tridis_HTR</a></p>
Dataset for the paper "Ground-truth Free Evaluation of HTR on Old French and Latin Medieval Literary Manuscripts"
<p>This dataset was used in the context of the article <em>Ground-truth Free Evaluation of HTR on Old French and Latin Medieval Literary Manuscripts</em>, at the Computational Humanities Research 2022 conference.</p> <p>Predictions.zip contains the XML ALTO for the HTR and Segmentation prediction of around 10 pages of 1900 manuscripts from the Bibliothèque nationale de France.</p> <p>Varying ground truth contains the original training material for training a classifier for CER classification (see the article).</p> <p>The ManuscriptsIIIF.csv contains metadata about the manuscripts.</p> <p>This work was funded by the DIM MAP under the CREMMALab funding.</p>
Hiding in plain sight: The biomolecular identification of pinniped use in medieval manuscripts – MALDI and mtDNA data set
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.