Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

16

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

16 results for “Historical documents”

Learn how ShareScore rates datasets ↗
zenodo40/100

ScriptNet: ICDAR2017 Competition on Historical Document Writer Identification (Historical-WI)

<p>This dataset contains the test set for the ICDAR2017 Competition on Historical Document Writer Identification (Historical-WI).</p> <p>The dataset used in this competition consists of 3600 handwritten pages originating from 13th to 20th century. It contains manuscripts from 720 different writers where each writer contributed five pages.</p> <p>Competition Website: https://scriptnet.iit.demokritos.gr/competitions/6/</p> <p>Changes August 1st, 2018: uploaded trainings set in color and binarized</p> <p>if you use the dataset please cite the ICDAR2017 Competition on Historical Document Writer Identification (Historical-WI) paper: https://doi.org/10.1109/ICDAR.2017.225</p> <p>&nbsp;</p>

opencc-by-sa-4.0Aug 2017View details →
zenodo40/100

Oficio de Hipotecas de Girona. A dataset of Spanish notarial deeds (18th Century) for Handwritten Text Recognition and Layout Analysis of historical documents.

<p>This dataset is a subset of 596 documents from the&nbsp;<em>Registre d&#39;Hipoteques de Girona</em> of 1769 collection, guarded by the <a href="http://xac.gencat.cat/ca/llista_arxius_comarcals/girones/"><em>Arxiu Hist&ograve;ric de Girona</em></a>. This collection, is composed by hundreds of thousands of notarial deeds from the XVIII-XIX century (1768-1862). Sales, redemption of censuses, inheritance and matrimonial chapters are among the most common documentary&nbsp;typologies in the collection.</p> <p>This dataset is composed of more than 23700 text lines&nbsp;written by a single hand, covering more that 50 different topics (documentary typologies) and a vocabulary of more than 2400 different words. The documents are transcribed using the so-called&nbsp;diplomatic criteria. Additionally, transcripts were tagged with&nbsp;<br> extra enriching/complementary information (e.g. expansion of the&nbsp;abbreviations, hyphen marks, etc.). Along with the transcripts &nbsp;the layout of the document is detected and recorded. Pages have&nbsp;been labeled using six different layout regions.</p> <p>The images along with their respective ground-truth was compiled in PAGE compliant XML format<br> by the <a href="http://www2.udg.edu/tabid/11296/Default.aspx"><em>Centre de Recerca d&#39;Hist&ograve;ria Rural</em></a>&nbsp;and the HTR group of the <a href="https://www.prhlt.upv.es">Pattern Recognition and Human Language Technologies Research Center</a>.</p>

opencc-by-nc-4.0Jul 2018View details →
zenodo40/100

GT4HistOCR: Ground Truth for training OCR engines on historical documents in German Fraktur and Early Modern Latin

<p><strong>GT4HistOCR</strong> contains ground truth for research in Optical Character Recognition (OCR) technology applied to historical printings in German Fraktur and Early Modern Latin.</p> <p>The ground truth comes in pairs of images of single printed lines as they appear in book pages (*.png) and their corresponding diplomatic transcriptions (*.gt.txt), which are UTF-8 strings preserving the character forms (glyphs) as much as possible within the UNICODE standard. These pairs of line images and their transcriptions can be directly used to train recognition models with, e.g., the open source OCR engines <em>OCRopy</em> or <em>Tesseract</em>. A total of 313,173 ground truth lines are provided.</p> <p><strong>Please note that the subcorpora making up this collection used different transcription guidelines, so it is a bad idea to train a recognition model on the total collection! Rather train individual models for each subcorpus.</strong> Fur further information about the subcorpora, please see the README file and the accompanying publication.</p> <p>If these data are useful for you, please cite the accompanying publication:</p> <pre>@article{<a href="http://springmann.net/publications.html#springmann2018gt4hist">springmann2018gt4hist</a>, author = {Uwe Springmann and Christian Reul and Stefanie Dipper and Johannes Baiter}, title = {{Ground Truth for training {OCR} engines on historical documents in German Fraktur and Early Modern Latin}}, journal = {J. Lang. Technol. Comput. Linguistics}, volume = {33}, number = {1}, pages = {97--114}, year = {2018}, url = {https://jlcl.org/content/2-allissues/1-heft1-2018/jlcl_2018-1_5.pdf} }</pre>

opencc-by-4.0Aug 2018View details →
zenodo40/100

ICDAR 2019 Competition on Image Retrieval for Historical Handwritten Documents [HisIR19] Dataset

<p>This dataset contains the training and test set used in the ICDAR 2019 Competition on Image Retrieval for Historical Handwritten Documents.</p> <p>This competition investigates the performance of large-scale retrieval of historical document images based on<br> writing style. Based on large image data sets provided by cultural heritage institutions and digital libraries, providing<br> a total of 20 000 document images representing about 10 000 writers, divided in three types: writers of (i) manuscript books, (ii) letters, (iii) charters and legal documents. We focus on the task of automatic image retrieval to simulate common scenarios of humanities research, such as writer retrieval.</p> <p>The training data set encompasses images from (i) Letters A, where each writer contributed one or three images; (ii) Manuscripts, where each writer was represented by five consecutive images from a single book.<br> In total, it contains 300 writers contributing one page, 100 writers contributing three pages, and 120 writers contributing five pages resulting in 1200 images of 520 writers.</p> <p>The test data set contains 20 000 images: About 7 500 pages stem from isolated documents (partially anonymous writers, contributing one page each), and about 12 500 pages are from writers that contributed three or five pages.</p> <p>&nbsp;</p> <p>If you use this dataset, please cite:</p> <p>V. Christlein, A. Nicolaou, M. Seuret, D. Stutzmann, A. Maier: &quot;ICDAR 2019 Competition on Image Retrieval for Historical Handwritten Documents&quot;, in 15th International Conference on Document Analysis and Recognition, 2019, Sydney, Australia</p> <p>&nbsp;</p>

opencc-by-nc-sa-4.0Jun 2019View details →
zenodo40/100

ICDAR 2021 Historical Document Classification Test Dataset for Task 1 - Font Groups

<p>Test set for the font group classification task of the ICDAR 2021 Competition on Historical Document Classification competition. The tar.gz file contains the images and a ground truth CSV file. The csv file indicates for each image from which document and page it corresponds to, as well as whether augmentations have been applied.</p>

opencc-by-4.0May 2021View details →
zenodo36/100

Historical Photograph and Document Collection

<p>After its establishment in 1929, TREC &ldquo;began business&rdquo; in 1930, with the construction of its first lab and office buildings and planting of the first crops. The historic photographs emphasize the first 10 years of TREC, when the most visible transition occurred from pineland to farmland, though also from the following decades (1940s-1970s), when TREC started to resemble the present campus. During this time and beyond, the nearly soiless oolitic limestone with native pine rockland transitioned into ornamental, fruit, and vegetable crops, while the pine rockland habitat was becoming vastly reduced and endangered.</p> <p>The catalog for this collection is contained as `Photo&amp;DocumentScans.xlsx` with the collection as the accompanying `Photo&amp;DocumentScansCollection.zip`.</p>

opencc-by-4.0Jun 2022View details →
zenodo36/100

ICDAR 2021 Historical Document Classification Dataset for Task 2 - Dating

<p>Dataset for the localization classification task of the ICDAR 2021 Competition on Historical Document Classification competition. The tar.gz file contains the images and ground truth file (CSV format). The metadata csv file contains information, such as where the image comes from.</p>

opencc-by-4.0May 2021View details →
zenodo32/100

EPARCHOS - Historical Greek handwritten document dataset

<p>The&nbsp;dataset originates from a Greek handwritten codex that dates from around 1500-1530. This is the subset of the codex British Museum Addit. 6791, written by two hands, one by Antonius Eparchos and the other by Camillos Zanettus (ff. 104r-174v) and delivers texts by Hierocles (In Aureum carmen), Matthaeus Blastares (Collectio alphabetica) and, notably, texts by Michael Psellos (De omnifaria doctrina). The writing&nbsp;delivers the most important abbreviations, logograms and conjunctions, which are cited in virtually every Greek minuscule handwritten codex from the years of the manuscript transliteration and the prevalence of the minuscule script (9th century) to the post-Byzantine years. This dataset consists of 120 scanned handwritten text pages, containing 9285 lines of text,&nbsp;18809 words (6787 unique words). For each page, a PageXML is provided containing the following groundtruth:</p> <ol> <li>Text region polygon coordinates</li> <li>Text line polygon coordinates with the corresponding transcription text</li> <li>Word polygon coordinated with the corresponding transcription text</li> </ol>

opencc-by-4.0Oct 2020View details →
zenodo32/100

ICDAR 2021 Historical Document Classification Test Dataset for Task 3 - Location

<p>Test set for the location classification task of the ICDAR 2021 Historical Document Classification.</p>

opencc-by-4.0Apr 2023View details →
zenodo28/100

IO Islamic 549. 'Inâyatnâma, A Collection of Famous Letters and Historical Documents

<p>IO Islamic 549. &#39;In&acirc;yatn&acirc;ma, A Collection of Famous Letters and Historical Documents</p>

opencc-by-4.0Feb 2020View details →
zenodo28/100

IO Islamic 2895. Historical documents Relating to the Marattah Power in India

<p>IO Islamic 2895. Historical documents Relating to the Marattah Power in India</p>

opencc-by-4.0Mar 2020View details →
zenodo28/100

ICDAR 2021 Historical Document Classification Test Dataset for Task 1 - Scripts

<p>Test set for thescript classification task of the ICDAR 2021 Competition on Historical Document Classification competition. The tar.gz file contains the images and a ground truth CSV file. The csv file indicates for each image from which document and page it corresponds to, as well as whether augmentations have been applied.</p>

opencc-by-4.0May 2021View details →
zenodo24/100

ICDAR 2021 Historical Document Classification Dataset for Task 3 - Location

<p>Test set for the localization classification task of the ICDAR 2021 Competition on Historical Document Classification competition. The tar.gz file contains the images and a ground truth CSV file. The metadata csv file contains information, such as where the image comes from.</p>

opencc-by-4.0May 2021View details →
dryad24/100

Data from: Historical DNA documents long distance natal homing in marine fish

Open the record for dataset details and reuse information.

publicFeb 2016View details →
zenodo16/100

Bodhgayā, Bihar. Photograph of early historic slab, documented after excavation, and marked with location of inscriptions

<p>Bodhgayā, Bihar. Photograph of early historic slab, documented after excavation, and marked with location of inscriptions. British Museum 1897,0528,0.105 (a).</p>

restrictedMay 2020View details →
zenodo16/100

Bodhgayā, Bihar. Photograph of early historic slab, documented after excavation.

<p>Bodhgayā, Bihar. Photograph of early historic slab, documented after excavation. British Museum 1897,0528,0.105 (a).</p>

restrictedMay 2018View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record