Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
16
datasets available to search
ShareScore release 0.9.0
Dataset results
16 results for “Historical Documents”
ScriptNet: ICDAR2017 Competition on Historical Document Writer Identification (Historical-WI)
<p>This dataset contains the test set for the ICDAR2017 Competition on Historical Document Writer Identification (Historical-WI).</p> <p>The dataset used in this competition consists of 3600 handwritten pages originating from 13th to 20th century. It contains manuscripts from 720 different writers where each writer contributed five pages.</p> <p>Competition Website: https://scriptnet.iit.demokritos.gr/competitions/6/</p> <p>Changes August 1st, 2018: uploaded trainings set in color and binarized</p> <p>if you use the dataset please cite the ICDAR2017 Competition on Historical Document Writer Identification (Historical-WI) paper: https://doi.org/10.1109/ICDAR.2017.225</p> <p> </p>
Oficio de Hipotecas de Girona. A dataset of Spanish notarial deeds (18th Century) for Handwritten Text Recognition and Layout Analysis of historical documents.
<p>This dataset is a subset of 596 documents from the <em>Registre d'Hipoteques de Girona</em> of 1769 collection, guarded by the <a href="http://xac.gencat.cat/ca/llista_arxius_comarcals/girones/"><em>Arxiu Històric de Girona</em></a>. This collection, is composed by hundreds of thousands of notarial deeds from the XVIII-XIX century (1768-1862). Sales, redemption of censuses, inheritance and matrimonial chapters are among the most common documentary typologies in the collection.</p> <p>This dataset is composed of more than 23700 text lines written by a single hand, covering more that 50 different topics (documentary typologies) and a vocabulary of more than 2400 different words. The documents are transcribed using the so-called diplomatic criteria. Additionally, transcripts were tagged with <br> extra enriching/complementary information (e.g. expansion of the abbreviations, hyphen marks, etc.). Along with the transcripts the layout of the document is detected and recorded. Pages have been labeled using six different layout regions.</p> <p>The images along with their respective ground-truth was compiled in PAGE compliant XML format<br> by the <a href="http://www2.udg.edu/tabid/11296/Default.aspx"><em>Centre de Recerca d'Història Rural</em></a> and the HTR group of the <a href="https://www.prhlt.upv.es">Pattern Recognition and Human Language Technologies Research Center</a>.</p>
GT4HistOCR: Ground Truth for training OCR engines on historical documents in German Fraktur and Early Modern Latin
<p><strong>GT4HistOCR</strong> contains ground truth for research in Optical Character Recognition (OCR) technology applied to historical printings in German Fraktur and Early Modern Latin.</p> <p>The ground truth comes in pairs of images of single printed lines as they appear in book pages (*.png) and their corresponding diplomatic transcriptions (*.gt.txt), which are UTF-8 strings preserving the character forms (glyphs) as much as possible within the UNICODE standard. These pairs of line images and their transcriptions can be directly used to train recognition models with, e.g., the open source OCR engines <em>OCRopy</em> or <em>Tesseract</em>. A total of 313,173 ground truth lines are provided.</p> <p><strong>Please note that the subcorpora making up this collection used different transcription guidelines, so it is a bad idea to train a recognition model on the total collection! Rather train individual models for each subcorpus.</strong> Fur further information about the subcorpora, please see the README file and the accompanying publication.</p> <p>If these data are useful for you, please cite the accompanying publication:</p> <pre>@article{<a href="http://springmann.net/publications.html#springmann2018gt4hist">springmann2018gt4hist</a>, author = {Uwe Springmann and Christian Reul and Stefanie Dipper and Johannes Baiter}, title = {{Ground Truth for training {OCR} engines on historical documents in German Fraktur and Early Modern Latin}}, journal = {J. Lang. Technol. Comput. Linguistics}, volume = {33}, number = {1}, pages = {97--114}, year = {2018}, url = {https://jlcl.org/content/2-allissues/1-heft1-2018/jlcl_2018-1_5.pdf} }</pre>
ICDAR 2019 Competition on Image Retrieval for Historical Handwritten Documents [HisIR19] Dataset
<p>This dataset contains the training and test set used in the ICDAR 2019 Competition on Image Retrieval for Historical Handwritten Documents.</p> <p>This competition investigates the performance of large-scale retrieval of historical document images based on<br> writing style. Based on large image data sets provided by cultural heritage institutions and digital libraries, providing<br> a total of 20 000 document images representing about 10 000 writers, divided in three types: writers of (i) manuscript books, (ii) letters, (iii) charters and legal documents. We focus on the task of automatic image retrieval to simulate common scenarios of humanities research, such as writer retrieval.</p> <p>The training data set encompasses images from (i) Letters A, where each writer contributed one or three images; (ii) Manuscripts, where each writer was represented by five consecutive images from a single book.<br> In total, it contains 300 writers contributing one page, 100 writers contributing three pages, and 120 writers contributing five pages resulting in 1200 images of 520 writers.</p> <p>The test data set contains 20 000 images: About 7 500 pages stem from isolated documents (partially anonymous writers, contributing one page each), and about 12 500 pages are from writers that contributed three or five pages.</p> <p> </p> <p>If you use this dataset, please cite:</p> <p>V. Christlein, A. Nicolaou, M. Seuret, D. Stutzmann, A. Maier: "ICDAR 2019 Competition on Image Retrieval for Historical Handwritten Documents", in 15th International Conference on Document Analysis and Recognition, 2019, Sydney, Australia</p> <p> </p>
ICDAR 2021 Historical Document Classification Test Dataset for Task 1 - Font Groups
<p>Test set for the font group classification task of the ICDAR 2021 Competition on Historical Document Classification competition. The tar.gz file contains the images and a ground truth CSV file. The csv file indicates for each image from which document and page it corresponds to, as well as whether augmentations have been applied.</p>
Historical Photograph and Document Collection
<p>After its establishment in 1929, TREC “began business” in 1930, with the construction of its first lab and office buildings and planting of the first crops. The historic photographs emphasize the first 10 years of TREC, when the most visible transition occurred from pineland to farmland, though also from the following decades (1940s-1970s), when TREC started to resemble the present campus. During this time and beyond, the nearly soiless oolitic limestone with native pine rockland transitioned into ornamental, fruit, and vegetable crops, while the pine rockland habitat was becoming vastly reduced and endangered.</p> <p>The catalog for this collection is contained as `Photo&DocumentScans.xlsx` with the collection as the accompanying `Photo&DocumentScansCollection.zip`.</p>
ICDAR 2021 Historical Document Classification Dataset for Task 2 - Dating
<p>Dataset for the localization classification task of the ICDAR 2021 Competition on Historical Document Classification competition. The tar.gz file contains the images and ground truth file (CSV format). The metadata csv file contains information, such as where the image comes from.</p>
EPARCHOS - Historical Greek handwritten document dataset
<p>The dataset originates from a Greek handwritten codex that dates from around 1500-1530. This is the subset of the codex British Museum Addit. 6791, written by two hands, one by Antonius Eparchos and the other by Camillos Zanettus (ff. 104r-174v) and delivers texts by Hierocles (In Aureum carmen), Matthaeus Blastares (Collectio alphabetica) and, notably, texts by Michael Psellos (De omnifaria doctrina). The writing delivers the most important abbreviations, logograms and conjunctions, which are cited in virtually every Greek minuscule handwritten codex from the years of the manuscript transliteration and the prevalence of the minuscule script (9th century) to the post-Byzantine years. This dataset consists of 120 scanned handwritten text pages, containing 9285 lines of text, 18809 words (6787 unique words). For each page, a PageXML is provided containing the following groundtruth:</p> <ol> <li>Text region polygon coordinates</li> <li>Text line polygon coordinates with the corresponding transcription text</li> <li>Word polygon coordinated with the corresponding transcription text</li> </ol>
ICDAR 2021 Historical Document Classification Test Dataset for Task 3 - Location
<p>Test set for the location classification task of the ICDAR 2021 Historical Document Classification.</p>
IO Islamic 549. 'Inâyatnâma, A Collection of Famous Letters and Historical Documents
<p>IO Islamic 549. 'Inâyatnâma, A Collection of Famous Letters and Historical Documents</p>
IO Islamic 2895. Historical documents Relating to the Marattah Power in India
<p>IO Islamic 2895. Historical documents Relating to the Marattah Power in India</p>
ICDAR 2021 Historical Document Classification Test Dataset for Task 1 - Scripts
<p>Test set for thescript classification task of the ICDAR 2021 Competition on Historical Document Classification competition. The tar.gz file contains the images and a ground truth CSV file. The csv file indicates for each image from which document and page it corresponds to, as well as whether augmentations have been applied.</p>
ICDAR 2021 Historical Document Classification Dataset for Task 3 - Location
<p>Test set for the localization classification task of the ICDAR 2021 Competition on Historical Document Classification competition. The tar.gz file contains the images and a ground truth CSV file. The metadata csv file contains information, such as where the image comes from.</p>
Data from: Historical DNA documents long distance natal homing in marine fish
Open the record for dataset details and reuse information.
Bodhgayā, Bihar. Photograph of early historic slab, documented after excavation, and marked with location of inscriptions
<p>Bodhgayā, Bihar. Photograph of early historic slab, documented after excavation, and marked with location of inscriptions. British Museum 1897,0528,0.105 (a).</p>
Bodhgayā, Bihar. Photograph of early historic slab, documented after excavation.
<p>Bodhgayā, Bihar. Photograph of early historic slab, documented after excavation. British Museum 1897,0528,0.105 (a).</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.