Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

6

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

6 results for “OCR Texts”

Learn how ShareScore rates datasets ↗
zenodo40/100

OCRed text of the Allgemeine musikalische Zeitung (ALTO format) from 1798-1848 and 1863-1965

<p><span>Dataset as used in the LREC 2020 publication &#39;<em>Allgemeine Musikalische Zeitung</em></span><span> as a Searchable Online Corpus&#39;</span>.</p>

opencc-by-4.0Mar 2020View details →
zenodo40/100

Dataset of ICDAR 2019 Competition on Post-OCR Text Correction

<p><strong>Corpus for the ICDAR2019 Competition on Post-OCR Text Correction (October 2019)</strong><br> Christophe Rigaud, Antoine Doucet, Mickael Coustaty, Jean-Philippe Moreux<br> <a href="http://l3i.univ-larochelle.fr/ICDAR2019PostOCR">http://l3i.univ-larochelle.fr/ICDAR2019PostOCR</a><br> -------------------------------------------------------------------------------</p> <p>These are the supplementary materials for the ICDAR 2019 paper&nbsp;<em><a href="https://zenodo.org/record/3459116">ICDAR 2019 Competition on Post-OCR Text Correction</a></em></p> <p>Please use the following citation:</p> <pre><code>@inproceedings{rigaud2019pocr,</code> <code> </code><code>title=&quot;ICDAR 2019 Competition on Post-OCR Text Correction&quot;,</code> <code> </code><code>author={Rigaud, Christophe and Doucet, Antoine and Coustaty, Mickael and Moreux, Jean-Philippe},</code> <code> </code><code>year={2019},</code> <code> </code><code>booktitle={Proceedings of the 15th International Conference on Document Analysis and Recognition (2019)}</code> <code> </code><code>}</code></pre> <p>&nbsp;</p> <p><strong>Description</strong><br> The corpus accounts for 22M OCRed characters along with the corresponding Gold Standard (GS). The documents come from different digital collections available, among others, at the National Library of France (BnF) and the British Library (BL). The corresponding GS comes both from BnF&#39;s internal projects and external initiatives such as Europeana Newspapers, IMPACT, Project Gutenberg, Perseus and Wikisource.</p> <p><strong>Repartition of the dataset</strong><br> - <em>ICDAR2019_Post_OCR_correction_training_18M.zip</em>: 80% of the full dataset, provided to train participants&#39; methods.<br> - <em>ICDAR2019_Post_OCR_correction_evaluation_4M</em>: 20% of the full dataset used for the evaluation (with Gold Standard made publicly after the competition).<br> - <em>ICDAR2019_Post_OCR_correction_full_22M</em>: full dataset made publicly available after the competition.</p> <p><strong>Special case for Finnish language</strong><br> Material from the National Library of Finland (<em>Finnish dataset FI &gt; FI1</em>) are not allowed to be re-shared on other website. Please follow these guidelines to get and format the data from the original website.</p> <p>1. Go to <a href="https://digi.kansalliskirjasto.fi/opendata/submit?set_language=en">https://digi.kansalliskirjasto.fi/opendata/submit?set_language=en</a>;<br> 2. Download <em>OCR Ground Truth Pages (Finnish Fraktur) [v1](4.8GB)</em> from <em>Digitalia (2015-17)</em> package;<br> 3. Convert the Excel file &quot;<em>~/metadata/nlf_ocr_gt_tescomb5_2017.xlsx</em>&quot; as Comma Separated Format (.csv) by using <em>save as</em> function in a spreadsheet software (e.g. Excel, Calc) and copy it into &quot;<em>FI/FI1/HOWTO_get_data/input/</em>&quot;;<br> 4. Go to &quot;<em>FI/FI1/HOWTO_get_data/</em>&quot; and run &quot;<em>script_1.py</em>&quot; to generate the <em>full</em> &quot;<em>FI1</em>&quot; dataset in &quot;<em>output/full/</em>&quot;;<br> 4. Run &quot;<em>script_2.py</em>&quot; to split the &quot;<em>output/full/</em>&quot; dataset into &quot;<em>output/training/</em>&quot; and &quot;<em>output/evaluation/</em>&quot; sub sets.<br> At the end of the process, you should have a &quot;<em>training</em>&quot;, &quot;<em>evaluation</em>&quot; and &quot;<em>full</em>&quot; folder with 1579528, 380817 and 1960345 characters respectively.</p> <p><br> <strong>Licenses: free to use for non-commercial uses, according to sources in details</strong><br> - BG1: IMPACT - National Library of Bulgaria: CC BY NC ND<br> - CZ1: IMPACT - National Library of the Czech Republic: CC BY NC SA<br> - DE1: Front pages of Swiss newspaper NZZ: Creative Commons Attribution 4.0 International (<a href="https://zenodo.org/record/3333627">https://zenodo.org/record/3333627</a>)<br> - DE2: IMPACT - German National Library: CC BY NC ND<br> - DE3: GT4Hist-dta19 dataset: CC-BY-SA 4.0 (<a href="https://zenodo.org/record/1344132">https://zenodo.org/record/1344132</a>)<br> - DE4: GT4Hist - EarlyModernLatin: CC-BY-SA 4.0 (<a href="https://zenodo.org/record/1344132">https://zenodo.org/record/1344132</a>)<br> - DE5: GT4Hist - Kallimachos: CC-BY-SA 4.0 (<a href="https://zenodo.org/record/1344132">https://zenodo.org/record/1344132</a>)<br> - DE6: GT4Hist - RefCorpus-ENHG-Incunabula: CC-BY-SA 4.0 (<a href="https://zenodo.org/record/1344132">https://zenodo.org/record/1344132</a>)<br> - DE7: GT4Hist - RIDGES-Fraktur: CC-BY-SA 4.0 (<a href="https://zenodo.org/record/1344132">https://zenodo.org/record/1344132</a>)<br> - EN1: IMPACT - British Library: CC BY NC SA 3.0<br> - ES1: IMPACT - National Library of Spain: CC BY NC SA<br> - FI1: National Library of Finland: no re-sharing allowed, follow the above section to get the data. (<a href="https://digi.kansalliskirjasto.fi/opendata">https://digi.kansalliskirjasto.fi/opendata</a>)<br> - FR1: HIMANIS Project: CC0 (<a href="https://www.himanis.org/">https://www.himanis.org</a>)<br> - FR2: IMPACT - National Library of France: CC BY NC SA 3.0<br> - FR3: RECEIPT dataset: CC0 (<a href="http://findit.univ-lr.fr/">http://findit.univ-lr.fr</a>)<br> - NL1: IMPACT - National library of the Netherlands: CC BY<br> - PL1: IMPACT - National Library of Poland: CC BY<br> - SL1: IMPACT - Slovak National Library: CC BY NC</p> <p>Text post-processing such as cleaning and alignment have been applied on the resources mentioned above, so that the Gold Standard and the OCRs provided are not necessarily identical to the originals.</p> <p><br> <strong>Structure</strong><br> - **Content** [<em>./lang_type/sub_folder/#.txt</em>]<br> &nbsp;&nbsp; &nbsp;- &quot;<em>[OCR_toInput] </em>&quot; =&gt; Raw OCRed text to be de-noised.<br> &nbsp;&nbsp; &nbsp;- &quot;<em>[OCR_aligned] </em>&quot; =&gt; Aligned OCRed text.<br> &nbsp;&nbsp; &nbsp;- &quot;<em>[ GS_aligned] </em>&quot; =&gt; Aligned Gold Standard text.</p> <p>The aligned OCRed/GS texts are provided for training and test purposes. The alignment was made at the character level using &quot;<em>@</em>&quot; symbols. &quot;<em>#</em>&quot; symbols correspond to the absence of GS either related to alignment uncertainties or related to unreadable characters in the source document. For a better view of the alignment, make sure to disable the &quot;word wrap&quot; option in your text editor.</p> <p>The Error Rate and the quality of the alignment vary according to the nature and the state of degradation of the source documents. Periodicals (mostly historical newspapers) for example, due to their complex layout and their original fonts have been reported to be especially challenging. In addition, it should be mentioned that the quality of Gold Standard also varies as the dataset aggregates resources from different projects that have their own annotation procedure, and obviously contains some errors.</p> <p><br> <strong>ICDAR2019 competition</strong><br> Information related to the tasks, formats and the evaluation metrics are details on :<br> <a href="https://sites.google.com/view/icdar2019-postcorrectionocr/evaluation">https://sites.google.com/view/icdar2019-postcorrectionocr/evaluation</a></p> <p><br> <strong>References</strong><br> &nbsp;- IMPACT, European Commission&#39;s 7th Framework Program, grant agreement 215064<br> &nbsp;- Uwe Springmann, Christian Reul, Stefanie Dipper, Johannes Baiter (2018). Ground Truth for training OCR engines on historical documents in German Fraktur and Early Modern Latin.<br> &nbsp;- <a href="https://digi.nationallibrary.fi/">https://digi.nationallibrary.fi</a> , Wiipuri, 31.12.1904, Digital Collections of National Library of Finland<br> - EU Horizon 2020 research and innovation programme grant agreement No 770299</p> <p><br> <strong>Contact</strong><br> - christophe.rigaud(at)univ-lr.fr<br> - antoine.doucet(at)univ-lr.fr<br> - mickael.coustaty(at)univ-lr.fr<br> - jean-philippe.moreux(at)bnf.fr</p> <p>L3i - University of la Rochelle, <a href="http://l3i.univ-larochelle.fr/">http://l3i.univ-larochelle.fr</a><br> BnF - French National Library, <a href="http://www.bnf.fr/">http://www.bnf.fr</a></p>

opencc-by-4.0Oct 2019View details →
zenodo20/100

Event Registry titles dataset texts with OCR degradations

<p>This is the text of the Event Registry titles:</p> <p>Rupnik, Jan, Andrej Muhic, Gregor Leban, Primoz Skraba, Blaz Fortuna, et Marko Grobelnik. 2016. &laquo;&nbsp;News Across Languages - Cross-Lingual Document Similarity and Event Tracking&nbsp;&raquo;. <em>Journal of Artificial Intelligence Research</em> 55 (janvier): 283‑316. <a href="https://doi.org/10.1613/jair.4780">https://doi.org/10.1613/jair.4780</a>.</p> <p>Miranda, Sebasti&atilde;o, Artūrs Znotiņ&scaron;, Shay B. Cohen, et Guntis Barzdins. 2018. &laquo;&nbsp;Multilingual Clustering of Streaming News&nbsp;&raquo;. In <em>2018 Conference on Empirical Methods in Natural Language Processing</em>, 4535‑44. Brussels, Belgium: Association for Computational Linguistics. <a href="https://www.aclweb.org/anthology/D18-1483/">https://www.aclweb.org/anthology/D18-1483/</a>.</p> <p>Guillaume Bernard. (2022). Event Registry titles only dataset with multiple extracted features (both sparse and dense) (1.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.6630447</p> <p>Some degradations are applied using the DocCreator [1] tool in order to degrade the text of the tweets and to reproduce some common errors found in OCRised documents [2].</p> <p>[1]: Journet, Nicholas, Muriel Visani, Boris Mansencal, Kieu Van-Cuong, et Antoine Billy. 2017. &laquo;&nbsp;DocCreator: A New Software for Creating Synthetic Ground-Truthed Document Images&nbsp;&raquo;. <em>Journal of Imaging</em> 3 (4): 62. <a href="https://doi.org/10.3390/jimaging3040062">https://doi.org/10.3390/jimaging3040062</a>.</p> <p>[2]: Linhares Pontes, Elvys, Ahmed Hamdi, Nicolas Sidere, et Antoine Doucet. 2019. &laquo;&nbsp;Impact of OCR Quality on Named Entity Linking&nbsp;&raquo;. In <em>Digital Libraries at the Crossroads of Digital Information for the Future</em>, 11853:102‑15. Lecture Notes in Computer Science. Cham: Springer International Publishing. <a href="https://doi.org/10.1007/978-3-030-34058-2_11">https://doi.org/10.1007/978-3-030-34058-2_11</a>.</p> <p>The results of the OCR degradations are as follow:</p> <table> <caption>FibVid CER/WER</caption> <tbody> <tr> <td>&nbsp;</td> <td>&nbsp;</td> <td>Without</td> <td>Character degradation</td> <td>Phantom degradation</td> <td>Bleed</td> <td>Blur</td> <td>All</td> </tr> <tr> <td>Event Registry Titles</td> <td>CER</td> <td>2.421</td> <td>6.940</td> <td>2.414</td> <td>2.422</td> <td>2.874</td> <td>7.178</td> </tr> <tr> <td>Event Registry Titles</td> <td>WER</td> <td>1.127</td> <td>19.785</td> <td>1.124</td> <td>1.131</td> <td>2.035</td> <td>19.894</td> </tr> </tbody> </table>

restrictedJun 2022View details →
zenodo20/100

FibVid dataset texts with OCR degradations

<p>This is the text of the FibVid dataset dedicated to fake news detection that has been updated to be used in event detection.</p> <p>Kim, Jisu, Jihwan Aum, SangEun Lee, Yeonju Jang, Eunil Park, et Daejin Choi. 2021. &laquo;&nbsp;FibVID: Comprehensive Fake News Diffusion Dataset during the COVID-19 Period&nbsp;&raquo;. <em>Telematics and Informatics</em> 64 (novembre): 101688. <a href="https://doi.org/10.1016/j.tele.2021.101688">https://doi.org/10.1016/j.tele.2021.101688</a>.</p> <p>Guillaume Bernard. (2022). Fibvid dataset with multiple extracted features (both sparse and dense) (1.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.6630409</p> <p>Some degradations are applied using the DocCreator [1] tool in order to degrade the text of the tweets and to reproduce some common errors found in OCRised documents [2].</p> <p>[1]: Journet, Nicholas, Muriel Visani, Boris Mansencal, Kieu Van-Cuong, et Antoine Billy. 2017. &laquo;&nbsp;DocCreator: A New Software for Creating Synthetic Ground-Truthed Document Images&nbsp;&raquo;. <em>Journal of Imaging</em> 3 (4): 62. <a href="https://doi.org/10.3390/jimaging3040062">https://doi.org/10.3390/jimaging3040062</a>.</p> <p>[2]: Linhares Pontes, Elvys, Ahmed Hamdi, Nicolas Sidere, et Antoine Doucet. 2019. &laquo;&nbsp;Impact of OCR Quality on Named Entity Linking&nbsp;&raquo;. In <em>Digital Libraries at the Crossroads of Digital Information for the Future</em>, 11853:102‑15. Lecture Notes in Computer Science. Cham: Springer International Publishing. <a href="https://doi.org/10.1007/978-3-030-34058-2_11">https://doi.org/10.1007/978-3-030-34058-2_11</a>.</p> <p>The results of the OCR degradations are as follow:</p> <table> <caption>FibVid CER/WER</caption> <tbody> <tr> <td>&nbsp;</td> <td>&nbsp;</td> <td>Without</td> <td>Character degradation</td> <td>Phantom degradation</td> <td>Bleed</td> <td>Blur</td> <td>All</td> </tr> <tr> <td>FibVid</td> <td>CER</td> <td>1.463</td> <td>6.089</td> <td>1.461</td> <td>1.467</td> <td>1.935</td> <td>6.359</td> </tr> <tr> <td>FibVid</td> <td>WER</td> <td>2.065</td> <td>20.797</td> <td>2.041</td> <td>2.052</td> <td>2.868</td> <td>21.396</td> </tr> </tbody> </table> <p>&nbsp;</p>

restrictedJun 2022View details →
zenodo20/100

CoAID dataset texts with OCR degradations

<p>This is the text of the CoAID dataset dedicated to fake news detection that has been updated to be used in event detection.</p> <p>Cui, Limeng, et Dongwon Lee. 2020. &laquo;&nbsp;CoAID: COVID-19 Healthcare Misinformation Dataset&nbsp;&raquo;. <em>ArXiv:2006.00885 [Cs]</em>, novembre. <a href="http://arxiv.org/abs/2006.00885">http://arxiv.org/abs/2006.00885</a>.</p> <p>Guillaume Bernard. (2022). CoAID dataset with multiple extracted features (both sparse and dense) (1.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.6630405</p> <p>Some degradations are applied using the DocCreator [1] tool in order to degrade the text of the tweets and to reproduce some common errors found in OCRised documents [2].</p> <p>[1]: Journet, Nicholas, Muriel Visani, Boris Mansencal, Kieu Van-Cuong, et Antoine Billy. 2017. &laquo;&nbsp;DocCreator: A New Software for Creating Synthetic Ground-Truthed Document Images&nbsp;&raquo;. <em>Journal of Imaging</em> 3 (4): 62. <a href="https://doi.org/10.3390/jimaging3040062">https://doi.org/10.3390/jimaging3040062</a>.</p> <p>[2]: Linhares Pontes, Elvys, Ahmed Hamdi, Nicolas Sidere, et Antoine Doucet. 2019. &laquo;&nbsp;Impact of OCR Quality on Named Entity Linking&nbsp;&raquo;. In <em>Digital Libraries at the Crossroads of Digital Information for the Future</em>, 11853:102‑15. Lecture Notes in Computer Science. Cham: Springer International Publishing. <a href="https://doi.org/10.1007/978-3-030-34058-2_11">https://doi.org/10.1007/978-3-030-34058-2_11</a>.</p> <p>The results of the OCR degradations are as follow:</p> <table> <caption>CoAID CER/WER</caption> <tbody> <tr> <td>&nbsp;</td> <td>&nbsp;</td> <td>Without</td> <td>Character degradation</td> <td>Phantom degradation</td> <td>Bleed</td> <td>Blur</td> <td>All</td> </tr> <tr> <td>CoAID</td> <td>CER</td> <td>2.105</td> <td>6.358</td> <td>2.105</td> <td>2.122</td> <td>2.616</td> <td>7.898</td> </tr> <tr> <td>CoAID</td> <td>WER</td> <td>2.494</td> <td>20.230</td> <td>2.496</td> <td>2.580</td> <td>3.726</td> <td>20.230</td> </tr> </tbody> </table> <p>&nbsp;</p>

restrictedJun 2022View details →
zenodo20/100

Event Registry dataset texts with OCR degradations and synthesised segmentation

<p>This is the text of the Event Registrt dataset dedicated to fake news detection that has been updated to be used in event detection.</p> <p>Rupnik, Jan, Andrej Muhic, Gregor Leban, Primoz Skraba, Blaz Fortuna, et Marko Grobelnik. 2016. &laquo;&nbsp;News Across Languages - Cross-Lingual Document Similarity and Event Tracking&nbsp;&raquo;. <em>Journal of Artificial Intelligence Research</em> 55 (janvier): 283‑316. <a href="https://doi.org/10.1613/jair.4780">https://doi.org/10.1613/jair.4780</a>.</p> <p>Miranda, Sebasti&atilde;o, Artūrs Znotiņ&scaron;, Shay B. Cohen, et Guntis Barzdins. 2018. &laquo;&nbsp;Multilingual Clustering of Streaming News&nbsp;&raquo;. In <em>2018 Conference on Empirical Methods in Natural Language Processing</em>, 4535‑44. Brussels, Belgium: Association for Computational Linguistics. <a href="https://www.aclweb.org/anthology/D18-1483/">https://www.aclweb.org/anthology/D18-1483/</a>.</p> <p>Guillaume Bernard. (2022). Event Registry titles only dataset with multiple extracted features (both sparse and dense) (1.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.6630447</p> <p>Some degradations are applied using the DocCreator [1] tool in order to degrade the text of the tweets and to reproduce some common errors found in OCRised documents [2].</p> <p>[1]: Journet, Nicholas, Muriel Visani, Boris Mansencal, Kieu Van-Cuong, et Antoine Billy. 2017. &laquo;&nbsp;DocCreator: A New Software for Creating Synthetic Ground-Truthed Document Images&nbsp;&raquo;. <em>Journal of Imaging</em> 3 (4): 62. <a href="https://doi.org/10.3390/jimaging3040062">https://doi.org/10.3390/jimaging3040062</a>.</p> <p>[2]: Linhares Pontes, Elvys, Ahmed Hamdi, Nicolas Sidere, et Antoine Doucet. 2019. &laquo;&nbsp;Impact of OCR Quality on Named Entity Linking&nbsp;&raquo;. In <em>Digital Libraries at the Crossroads of Digital Information for the Future</em>, 11853:102‑15. Lecture Notes in Computer Science. Cham: Springer International Publishing. <a href="https://doi.org/10.1007/978-3-030-34058-2_11">https://doi.org/10.1007/978-3-030-34058-2_11</a>.</p> <p>The results of the OCR degradations are as follow:</p> <table> <caption>Event Registry CER/WER</caption> <tbody> <tr> <td>&nbsp;</td> <td>&nbsp;</td> <td>Without</td> <td>Character degradation</td> <td>Phantom degradation</td> <td>Bleed</td> <td>Blur</td> <td>All</td> </tr> <tr> <td>Event Registry</td> <td>CER</td> <td>0.282</td> <td>4.154</td> <td>0.274</td> <td>0.275</td> <td>0.582</td> <td>4.577</td> </tr> <tr> <td>Event Registry</td> <td>WER</td> <td>0.552</td> <td>16.364</td> <td>0.551</td> <td>0.548</td> <td>1.159</td> <td>16.974</td> </tr> </tbody> </table> <p>&nbsp;</p>

restrictedJun 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record