Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
773
datasets available to search
ShareScore release 0.9.0
Dataset results
773 results for “medieval”
Faithful Transcriptions Data Set: TEI/XML-encoded Transcriptions of Medieval Theological Manuscripts
<p>From May to July 2021, the Berlin State Library and the Leipzig University Library jointly organized the Transcribathon <a href="https://lab.sbb.berlin/events/faithful-transcriptions/">“Faithful Transcriptions”</a>, a digital crowd souring project on medieval theological manuscripts. During the project, over 100 participants produced TEI/XML-encoded transcriptions in the IIIF workspace of the <a href="https://handschriftenportal.de/">Handschriftenportal</a>, which is currently being developed.</p> <p>The <a href="https://lab.sbb.berlin/datensets-transkribathon/">Faithful Transcriptions Data Set</a> contains 181 pages with 8.952 text lines from 12 manuscripts in German, Dutch, and Latin. The medieval scripts include Textura, Textualis, Gothic Cursiva, and Bastarda. The transcriptions are linked to the coordinates of the digitized manuscript image on text line level.</p> <p>--------------------------------------------</p> <p>Von Mai bis Juli 2021 richtete die Staatsbibliothek zu Berlin in Kooperation mit der Universitätsbibliothek Leipzig den Transkribathon <a href="https://lab.sbb.berlin/events/faithful-transcriptions/">„Faithful Transcriptions“</a> aus, ein digitales Crowd-Sourcing-Projekt zu theologischen Handschriften des Mittelalters. Über 100 Teilnehmende fertigten dabei TEI/XML-codierte Transkriptionen in der IIIF-basierten Arbeitsumgebung des aktuell in Entwicklung befindlichen <a href="https://handschriftenportal.de/">Handschriftenportals</a> an. </p> <p>Das <a href="https://lab.sbb.berlin/datensets-transkribathon/">Faithful Transcriptions-Datenset</a> enthält 181 Seiten mit 8.952 Textzeilen aus 12 Handschriften in deutscher, niederländischer und lateinischer Sprache. Die mittelalterlichen Schriften reichen von Textura über Textualis und Gotische Kursive bis hin zur Bastarda. Die Transkriptionen sind mit den Bildkoordinaten des Handschriftendigitalisats auf Textzeilenebene verknüpft. </p>
FIG. 4 in Where are we now? Early Medieval archaeozoology in Slovenia: an overview
FIG. 4. — Abundance of taxonomically identified animal remains per building area for: A, Buildings I, II; B, Buildings IV, VI at Pristava. Also given is the total abundance of taxonomically identified animal remains per taxa for the whole site (NISP Σ). The lists of micro-squares included in each building area are as follows (Pleterski 2010: fig. 5.5): Building I (micro-squares Su, Dr, Ni, Mo); Building II (micro-squares He, Po, Ba, RO); Building IV (micro-squares Mi, AI) and Building VI (micro-squares Šp, Pš, Ir, Ri). Only those specimens that could be reliably assigned to individual buildings are considered, which is why the data shown here differ slightly from those in Table 3.
FIG. 1 in Where are we now? Early Medieval archaeozoology in Slovenia: an overview
FIG. 1. — Geographical location of the study area. Also shown are the archaeological sites mentioned in the text.
FIG. 3 in Where are we now? Early Medieval archaeozoology in Slovenia: an overview
FIG. 3. — Individual building areas at Pristava as defined by Pleterski (2010: 165-167, 245, figs. 5.4, 5.5).
FIG. 2 in Where are we now? Early Medieval archaeozoology in Slovenia: an overview
FIG. 2. — Final distribution of the matrix derived from multidimensional scaling of Euclidean distances among 20 early medieval mammal assemblages from Slovenia (stress = 0.0268). Each assemblage is a random subsample of the total available archaeozoological data from the sites of:, Popava (N = 5);, Pristava (N = 5);, Tonovcov grad (N = 5);, Koper (N = 5). Subsample size is constant (NISP = 30).
CREMMA Medieval - abbr and expan altos
<p>This data set is derived from the CREMMA-Medieval dataset (<a href="https://github.com/HTR-United/cremma-medieval">https://github.com/HTR-United/cremma-medieval</a>). It was modified to include both abbreviated and expanded forms in separated ALTO, for HTR training experiments.</p> <p>The original data set was created with the support of the <em>DIM MAP</em> in the context of the CREMMA project (<a href="https://www.dim-map.fr/projets-soutenus/cremma/">https://www.dim-map.fr/projets-soutenus/cremma/</a>).<br> This version was prepared in March-June 2022 as part of the research for the following paper:</p> <p><strong>Camps</strong>, Jean-Baptiste, Chahan<strong> Vidal-Gorène</strong>, Dominique <strong>Stutzmann</strong>, Marguerite <strong>Vernet</strong>, and Ariane <strong>Pinche</strong>. « Data Diversity in Handwritten Text Recognition: Challenge or Opportunity? » In <em>Digital Humanities 2022. Conference Abstracts (The University of Tokyo, Japan, 25-29 July 2022)</em>, published by DH2022 Local Organizing Committee, 160‑65. Tokyo, 2022. <a href="https://dh2022.dhii.asia/dh2022bookofabsts.pdf#page=162">https://dh2022.dhii.asia/dh2022bookofabsts.pdf#page=162</a> and <a href="https://dh2022.dhii.asia/abstracts/files/CAMPS_Jean_Baptiste_Data_Diversity_in_handwritten_text_recog.html">https://dh2022.dhii.asia/abstracts/files/CAMPS_Jean_Baptiste_Data_Diversity_in_handwritten_text_recog.html</a>.</p> <p>If you use this dataset, please quote:</p> <pre>@incollection{dh2022_local_organizing_committee_data_2022, address = {Tokyo}, title = {Data {Diversity} in handwritten text recognition: challenge or opportunity?}, url = {https://dh2022.dhii.asia/dh2022bookofabsts.pdf}, language = {en}, urldate = {2022-08-02}, booktitle = {Digital {Humanities} 2022. {Conference} {Abstracts} ({The} {University} of {Tokyo}, {Japan}, 25-29 {July} 2022)}, author = {Camps, Jean-Baptiste and Vidal-Gorène, Chahan and Stutzmann, Dominique and Vernet, Marguerite and Pinche, Ariane}, editor = {{DH2022 Local Organizing Committee}}, year = {2022}, pages = {160--165}, } </pre> <p><br> <strong>Folders</strong><br> <strong>img</strong><br> Folder with the scans of the base documents.</p> <p><strong>alto</strong><br> ALTO files, in abbreviated and expanded text form, with or without normalisations. For more details, please see the aforementioned paper.</p>
Dataset and evaluation for HTR models for Latin and French Medieval Documentary Manuscripts
<p><strong>1. Dataset presentation.</strong></p> <p>This is the dataset used to produce the HTR models applied to documentary Latin and French manuscripts presented in the paper: Sergio Torres Aguilar, Vincent Jolivet. <strong>Handwritten Text Recognition for Documentary Medieval<br> Manuscripts. </strong>2022. https://hal.science/hal-03892163</p> <p>The dataset contains mostly charters and registers from the Late-medieval period (12th-15th). The training and evaluation, entailing 1855 pages, 120k lines of text and almost 1M tokens, were conducted using three freely available ground-truth corpora :</p> <p><strong>The Alcar-HOME database </strong>: https://zenodo.org/record/5600884</p> <p><strong>The e-NDP corpus </strong>: https://github.com/chartes/e-NDP_HTR</p> <p><strong>The Himanis project </strong>: https://zenodo.org/record/5535306</p> <p>The final model operates in a multilingual environment (Latin and French) and it is able to recognize several Latin script families (mostly <em>Textualis</em> and <em>Cursiva</em>) in documents produced in ca. 12th - 15th centuries. During the evaluation the models shows an accuracy of <strong>94.01%</strong> on the validation set and a CER (character error ratio) of about <strong>0.12</strong> to <strong>0.17</strong> on four external unseen datasets. A fine-tuning exercise using 10 ground-truth pages can raise these results to a CER between <strong>0.06</strong> to <strong>0.10</strong> respectively.</p> <p> </p> <p><strong>2. Dataset contents .</strong></p> <p>a) <em>GT_list : </em>List containing the GT file names which constitute the training, evaluation and test sets. The images and transcriptions can be downloaded from their original repositories.</p> <p>b) <em>Training :</em> Contains the training and testing results (evaluation and prediction files) presented in the original paper for the two training phases: Regular (Textualis and Cursiva separated training) and Quartiles (mixed training by quartiles).</p> <p>c) <em>Useful_scripts :</em> Scripts to produce the HTR metrics (CER, WER, SER) and plot the model's accuracy.</p> <p>d) <em>Best_model :</em> Contains the best multilingual and multi-script model.</p>
Secondary ion mass spectrometry, a powerful tool for revealing ink formulations and animal skins in medieval manuscripts
<p>Book production by medieval scriptoria has gained growing interest in recent studies. In this context, identifying ink compositions and parchment animal species from illuminated manuscripts is of great importance. Here, we introduce time-of-flight secondary ion mass spectrometry (ToF-SIMS) as a non-invasive tool to identify both inks and animal skins in manuscripts, at the same time. For this purpose, both positive and negative ion spectra in inked and non-inked areas were recorded. Chemical compositions of pigments (decoration) or black inks (text) were determined by searching for characteristic ion mass peaks. Animal skins were identified by data processing of raw ToF-SIMS spectra using principal component analysis (PCA). In illuminated manuscripts from the fifteenth to sixteenth century, malachite (green), azurite (blue), cinnabar (red) inorganic pigments, as well as iron-gall black ink, were identified. Carbon black and indigo (blue) organic pigments were also identified. Animal skins were identified in modern parchments of known animal species by a two-step PCA procedure. We believe the proposed method will find extensive application in material studies of medieval manuscripts, as it is non-invasive, highly sensitive and able to identify both inks and animal skins at the same time, even from traces of pigments and tiny scanned areas.</p>
Lunar eclipses illuminate timing and climate impact of medieval volcanism
<p>This repository contains all the data and codes needed to reproduce the results and figures from the article "Lunar Eclipses Illuminate Timing and Climate Impacts of the Middle Ages" published in Nature.<br> <br> For more information, we refer the user to the readme file entitled "Guillet_et_al_Nature2023_Readme.txt".<br> <br> If you have any queries, please feel free to contact us: sebastien.guillet@unige.ch<br> <br> Thank you very much ;-)</p>
New database of fragments of medieval codices of the 11th-12th centuries
<p>The database currently has 117 codex fragments with the number of samples between 11 and 39, with an average of 17, which gives 2040 writing samples. The resolution of the width of the manuscripts varies between two thousand pixels and almost nine thousand pixels.</p>
Lunar eclipses illuminate timing and climate impact of medieval volcanism
<p>This repository contains all the data and codes needed to reproduce the results and figures from the article "Lunar Eclipses Illuminate Timing and Climate Impacts of the Middle Ages" published in Nature.<br> <br> For more information, we refer the user to the readme file entitled "Guillet_et_al_Nature2023_Readme.txt".<br> <br> If you have any queries, please feel free to contact us: sebastien.guillet@unige.ch<br> <br> Thank you very much ;-)</p>
FIG. 2 in Food taboos in medieval Iberia: the zooarchaeology of socio-cultural differences
FIG. 2. — Comparison of chicken and goose Number of Identified Specimens (NISP) % in Santa Marta-Pancorbo (left) with contemporary sites from the region. Salvatierra (Grau-Sologestoa 2015), Vitoria (Castaños et al. 2011, 2012, 2013) and Orduña (Cajigas et al. 2003) are towns; Desolado de Rada is a rural settlement (Castaños & Castaños 2003). Abbreviations:LMA, Late Middle Ages; ModE, Modern Era; Trans, Transition.
FIG. 1 in Food taboos in medieval Iberia: the zooarchaeology of socio-cultural differences
FIG. 1. — Comparison of cattle and pig Number of Identified Specimens (NISP) % in Santa Marta-Pancorbo (left) with contemporary sites from the region. Salvatierra (Grau-Sologestoa 2015), Vitoria (Castaños et al. 2011, 2012, 2013) and Orduña (Cajigas et al. 2003) are towns; Desolado de Rada is a rural settlement (Castaños & Castaños 2003). Abbreviations:LMA, Late Middle Ages; ModE, Modern Era; Trans, Transition.
Medieval sites in the hinterland of Ravenna
<p>Database with 639 site attestations coming from edited medieval written sources <a href="https://figshare.com/account/home#_ftn1">[1]</a> and published site catalogues <a href="https://figshare.com/account/home#_ftn2">[2]</a>, together with related metadata and bibliography.</p> <p><a href="https://figshare.com/account/home#_ftnref1">[1]</a> Gaddoni & Zaccherini, 1912; Mascanzoni, 1985, 2010; Benericetti, 1999, 2002a, 2002b, 2003, 2005, 2006a, 2007, 2009, 2010, 2011, 2019a, 2019b.</p> <p><a href="https://figshare.com/account/home#_ftnref2">[2]</a> Mancini & Vichi, 1959; Montevecchi, 1970, 1971, 1972; Merlini, 1982; Augenti, Ficara, & Ravaioli, 2012; Ravaioli, 2015.</p>
Zbiva, Early Medieval Data Set for the Eastern Alps. Data sub-set
<p><strong>Zbiva, Early Medieval Data Set for the Eastern Alps</strong></p><p> </p><p>Authors: Benjamin Štular, Andrej Pleterski, Mateja Belak</p><p>Institution: Znanstvenoraziskovalni center Slovenske akademije znanosti in umetnosti</p><p>Location and date: Ljubljana (Slovenia), 6 December 2021</p><p> </p><p>This dataset is a subset of Zbiva database. Zbiva is an open-access online research data base for the archaeology of the Eastern Alps in the Early Middle Ages. It consists of four parts: archaeological sites, graves, artefacts, and bibliography. Geographically, it contains data from present-day Slovenia, Austria, NW Croatia and NE Italy. For comparison purposes, it also includes selected relevant sites from neighbouring regions.</p><p>To access Zbiva directly go to https://zbiva4.zrc-sazu.si/en.</p><p>Read more about Zbiva:</p><p>Štular, B. (2019). The Zbiva Web Application: a tool for Early Medieval archaeology of the Eastern Alps . In J. D. Richards & F. Niccolucci, Eds. <i>The ARIADNE Impact;</i> (pp. 69–82). Archaeolingua, Budapest. https://doi.org/10.5281/zenodo.3476712.</p><p>Štular, B., Belak, M., Deep Data Example: Zbiva, Early Medieval Data Set for the Eastern Alps. – Research Data Journal for the Humanities and Social Sciences, September 2022; https://doi.org/10.1163/24523666-bja10024.</p><p> </p><p>This dataset is a subset of the Zbiva. It is published as a supporting material for:</p><p>Štular, B., Lozić, E., Belak, M., Rihter, J., Koch, I., Modrijan, Z., Magdič, A., Karl, S., Lehner, M., Gutjahr, Ch. 2022, Migration of Alpine Slavs and machine learning: Space-time pattern</p><p>A detailed description of the dataset including the metada is part of the downloadable dataset.</p>
Tracing 600 years of long-distance Atlantic cod trade in medieval and post-medieval Oslo using stable isotopes and ancient DNA
Open the record for dataset details and reuse information.
Secondary ion mass spectrometry, a powerful tool for revealing ink formulations and animal skins in medieval manuscripts
Open the record for dataset details and reuse information.
Medieval Italian Civic Chronicles, Part 1
<p>A recording on Civic Chronicles in Medieval Italy recorded in March 2020 for the website, <em>Middle Ages for Educators</em>.</p> <p>Beneš, Carrie and Morreale, Laura. “Medieval Italian Civic Chronicles, Part 1,” <em>Middle Ages for Educators</em>, March 27, 2020. Accessed October 25, 2020. <a href="http://middleagesforeducators.com/videos/medieval-italian-civic-chronicles-part-1/">http://middleagesforeducators.com/videos/medieval-italian-civic-chronicles-part-1/</a>.</p> <p> </p> <p> </p>
Networks of Shared Manuscript Transmission for Medieval European Vernacular Languages
<p>Codices that compile different textual units in the same physical object were common in the European Middle Ages. Many reasons guided scribes when collecting a variety of works and their analysis can offer scholars insight into scribal practices and the circulation of medieval works. In order to study this shared transmission of medieval texts, methods from network analysis offer the possibility of researching the phenomenon from a general perspective and discover fundamental trends. This paper deals with the methodological and practical foundations for such an approach. </p> <p>Modern digital databases of medieval manuscripts provide a big amount of relevant data for this type of research. In this presentation, I consider three online databases, each dealing with textual witnesses in vernacular languages: Handschriftencensus (German) Jonas (French and Occitan) and Philobiblon (Iberian languages). Each of these has different criteria on data collection and data modelling. For this reason, their analysis and comparison offers insight not only on medieval manuscripts and texts, but also on the consequences of different approaches when cataloguing and describing medieval manuscripts. </p>
An Artificial Eye for Palaeography. Applying Deep Machine Learning for the Study of Medieval Latin Scripts
<p>The project “Digital Forensics for Historical Documents” (at Huygens ING, Amsterdam) attempts to create a digital tool, based on a deep learning system, in which the unique characteristics of one medieval script sample will be matched with similar script samples by making use of digitized manuscript collections available in the world wide web.</p> <p>Project website and contact: <a href="https://www.youtube.com/redirect?q=https%3A%2F%2Fen.huygens.knaw.nl%2Fprojecten%2Fdigital-forensics-for-historical-documents%2F&v=WYtseNK-1Dc&event=video_description&redir_token=QUFFLUhqbDM0WDRJRjN0V3QzOXV6d0ZNbDB2TVYzV1hUQXxBQ3Jtc0ttLVJHMEE1RlRkZjJVV3poYnpQOHZseTFzZHVqWHRGR2k2eWpVWTJSSldtV2p4aFhKbFRIQTFjVHQ5YnM0Mkd4ajBzOTNtdE1yOVJLVktqWlhqWUgzZXc3YmNQMS1nNGFpZ3p2amNOTFNqTDVUSTdtZw%3D%3D">https://en.huygens.knaw.nl/projecten/...</a> </p> <p>Presented as a Lightning Talk for the Schoenberg Symposium 2020</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.