Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

4

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

4 results for “medieval charters”

Learn how ShareScore rates datasets ↗
zenodo44/100

MPS Data set with images of medieval charters for handwriting-style based dating of manuscripts

<pre>The MPS benchmark data set for handwritten manuscript dating ____________________________________________________________ This data set is collected for the Dutch NWO project: Medieval Paleographical Scale (MPS) by Petros Samara Project website: http://application02.target.rug.nl/monk/Projects/MPS/ Copyright (c) Huygensinstituut, Den Haag, 2016 University of Groningen, 2016. All rights reserved. Organisation of the data: Each .tar.gz file contains a number of NetPBM images. The format is chosen because of its simplicity. Also, there is no doubt about lossy compression in the processing chain. The file names are of the format &#39;MPS&lt;year&gt;_&lt;seqnr&gt;.ppm&#39;, for example, &#39;MPS1300_0056.ppm&#39;. Note: the files are not in a separate directory, they will be extracted in place. However, due to the unique naming, there is no problem extracting them in one single current (destination) directory. The actual type of the image can be gray scale (.pgm) or color (.ppm), in &#39;8-bit DirectClass&#39; according to ImageMagick&#39;s &#39;identify&#39; tool. The images were cropped out of larger photographs because of irrelevant elements such as a Kodak color calibrator and non-text content such as supporting surface (table) backgrounds, seals (emblems), ribbons, etc. No effort has been made to obtain a balanced set of samples over years: the given frequencies of occurrence in archives are used. There is evidently less data in years before 1375 A.D. while some periods provides us with ample data for historical reasons (e.g, 1450 A.D.). It would have been a pity if the scarce years had determined and limited the size of this data set. Selection criteria for data reduction, whether random or systematic, would have been arbitrary. In any case, these images were used in our publications, such that the performance results of future attempts on manuscript dating can be compared with earlier results. The performances that have been reached using our algorithms are in the order of an MAE (mean average error) of 10 years. If you have any questions, please contact us: Sheng He (heshengxgd@gmail.com) Petros Samara (petros.samara@huygens.knaw.nl) Jan Burgers (jan.burgers@huygens.knaw.nl) Lambert Schomaker (L.Schomaker@ai.rug.nl) Please cite our papers if you use this data set: [1] Sheng He, Petros Samara, Jan Burgers, Lambert Schomaker. Image-based historical manuscript dating using contour and stroke fragments. Pattern Recognition(PR), Vol. 59, pp. 159-171, 2016 [2] Sheng He, Petros Samara, Jan Burgers, Lambert Schomaker. Towards style-based dating of historical documents. International Conference on Frontiers in Handwriting Recognition(ICFHR), Crete, Greece, 2014 [3] Sheng He, Petros Samara, Jan Burgers, Lambert Schomaker. Multiple-Label Guided Clustering Algorithm for Historical Document Dating and Localization IEEE Trans. on Image Processing, Vol. 25(11), Nov. 2016. http://ieeexplore.ieee.org/document/7551181/</pre> <p>Data are collected thanks to&nbsp;Dutch NWO grant project 380-50-006</p>

opencc-by-4.0Aug 2016View details →
zenodo36/100

Multilingual named entity recognition for medieval charters. Datasets and models

<p>Annotated dataset for training named entities recognition models for medieval charters in Latin, French and Spanish.</p> <p>&nbsp;</p> <p>The original raw texts for all charters were collected from four charters collections</p> <p>- HOME-ALCAR corpus : <a href="https://zenodo.org/record/5600884">https://zenodo.org/record/5600884</a></p> <p>- CBMA : <a href="https://www.google.com/url?sa=t&amp;rct=j&amp;q=&amp;esrc=s&amp;source=web&amp;cd=&amp;ved=2ahUKEwisvNOa3qP3AhULyIUKHenpBoAQFnoECA0QAQ&amp;url=http%3A%2F%2Fwww.cbma-project.eu%2F&amp;usg=AOvVaw0blsASXzOKSNz_EkixwfJT">http://www.cbma-project.eu</a></p> <p>- Diplomata Belgica : <a href="https://www.diplomata-belgica.be">https://www.diplomata-belgica.be</a></p> <p>- CODEA corpus :<a href="http://https://corpuscodea.es/"> https://corpuscodea.es/</a></p> <p>&nbsp;</p> <p>We include (i) the annotated training datasets, (ii) the contextual and static embeddings trained on medieval multilingual texts and (iii) the named entity recognition models trained using two architectures: Bi-LSTM-CRF + stacked embeddings and fine-tuning on Bert-based models (mBert and RoBERTa)</p> <p>Codes, datasets and notebooks&nbsp;used to train models can&nbsp;be consulted in our&nbsp;gitlab repository:&nbsp;<a href="https://gitlab.com/magistermilitum/ner_medieval_multilingual">https://gitlab.com/magistermilitum/ner_medieval_multilingual</a></p> <p>Our best RoBERTa model is also available in the HuggingFace library:&nbsp;<a href="https://huggingface.co/magistermilitum/roberta-multilingual-medieval-ner">https://huggingface.co/magistermilitum/roberta-multilingual-medieval-ner</a></p>

opencc-by-4.0Apr 2022View details →
zenodo36/100

Text zones in images of medieval charters from Stiftsarchiv Seitenstetten in monasterium.net

<p>PAGE annotation (<a href="http://schema.primaresearch.org/PAGE/gts/pagecontent/2013-07-15">http://schema.primaresearch.org/PAGE/gts/pagecontent/2013-07-15</a>) of image regions containing texts in a set of scan medieval charters, scanned in the Stiftsarchiv Seitenstetten (<a href="http://monasterium.net/mom/AT-StiASei/SeitenstettenOSB/fond">http://monasterium.net/mom/AT-StiASei/SeitenstettenOSB/fond</a>). Images created by the International Center for Archival Research ICARus (<a href="http://icar-us.eu/">http://icar-us.eu/</a>) Annotations created with the support of the Austrian Science Fund, project number P 26.706 (<a href="https://illuminierte-urkunden.uni-graz.at/">https://illuminierte-urkunden.uni-graz.at/</a>). It does not contain any further segmentation for words or characters.</p>

opencc-by-nc-4.0Mar 2018View details →
zenodo16/100

Advances in Distant Diplomatics. A Stylometric Approach to Medieval Charters (data and code)

<p>The materials archived in this repository accompany and support the following paper:</p> <blockquote> <p>E. Leclerq &amp; M. Kestemont, &#39;Advances in Distant Diplomatics. A Stylometric Approach to Medieval Charters&#39;, in: <a href="https://riviste.unimi.it/interfaces/index">Interfaces. A Journal of Medieval European Literatures</a>&nbsp;[2021].</p> </blockquote> <p>The contents of this repository are the following:</p> <ul> <li><em>analysis.ipynb</em>: a Python notebook with all the code that was used for the analyses reported in the paper.</li> <li><em>CorpusCaseRogerFJeanE</em>&nbsp;(zipped folder): a collection of plain text files with the original charters. A primitive encoding scheme was applied to distinguish various subsections in the charters (see the Python notebook for the symbol legend).</li> <li><em>figures</em>&nbsp;(zipped folder): full-res versions of the automatically generated plots for the paper.</li> <li><em>embed</em>&nbsp;(zipped folder): interactive HTML scatterplots of the charters, colored depending on different kinds of metadata for a more intuitive exploration.</li> <li><em>hits</em>&nbsp;(zipped folder): HTML files containing a tabular representation of all intertexts detected between two charters (the pair of charters is identified in the filename).</li> <li><em>MetadataCaseRogerFJeanE.xlsx&nbsp;</em>(spreadsheet): a spreadsheet that encodes various kinds of metadata for each charter.</li> <li><em>README.md</em>: a README file.</li> </ul> <p>The original raw texts for all charters were collected from the <a href="https://www.diplomata-belgica.be/colophon_fr.html">Diplomata Belgica</a>&nbsp;and the <a href="http://telma.irht.cnrs.fr/outils/chartae-galliae/index/">Chartae Galliae</a>&nbsp;databases. In the metadata spreadsheet, the precise origin of the individual charters is given.</p>

restrictedOct 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record