Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
237
datasets available to search
ShareScore release 0.9.0
Dataset results
237 results for “OCR”
CoAID dataset with multiple extracted features (both sparse and dense) and degraded by OCR
<p>This is the same datasets as:</p> <p>Guillaume Bernard. (2022). CoAID dataset with multiple extracted features (both sparse and dense) (1.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.6630405</p> <p>But with texts degraded by OCR as described in:</p> <p>Guillaume Bernard. (2022). CoAID dataset texts with OCR degradations (1.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.6630710</p>
Event Registry dataset with multiple extracted features (both sparse and dense) and degraded by OCR
<p>This is the same dataset as:</p> <p>Guillaume Bernard. (2022). Event Registry dataset with multiple extracted features (both sparse and dense) (1.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.6630367</p> <p>But with texts degraded by OCR as described in:</p> <p>Guillaume Bernard. (2022). Event Registry dataset texts with OCR degradations and synthesised segmentation (1.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.6631305</p>
Event Registry titles dataset with multiple extracted features (both sparse and dense) and degraded by OCR
<p>This is the same dataset as:</p> <p>Guillaume Bernard. (2022). Event Registry titles only dataset with multiple extracted features (both sparse and dense) (1.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.6630447</p> <p>But with texts degraded by OCR as described in:</p> <p>Guillaume Bernard. (2022). Event Registry titles dataset texts with OCR degradations (1.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.6630828</p>
FibVid dataset with multiple extracted features (both sparse and dense) and degraded by OCR
<p>This is the same dataset as:</p> <p>Guillaume Bernard. (2022). Fibvid dataset with multiple extracted features (both sparse and dense) (1.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.6630409</p> <p>But with texts degraded by OCR as described in:</p> <p>Guillaume Bernard. (2022). FibVid dataset texts with OCR degradations (1.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.6630758</p>
ICDAR'15 SMARTPHONE DOCUMENT CAPTURE AND OCR COMPETITION (SmartDoc) - Challenge 1 (original version)
<p><strong>CHALLENGE 1: SMARTPHONE DOCUMENT CAPTURE COMPETITION</strong></p> <p><strong>Smartphones are replacing personal scanners.</strong> They are portable, connected, powerful and affordable. They are on their way to become the new entry point in business processing applications like document archival, ID scanning, check digitization, just to name a few. In order keep our workflows streamlined, <strong>we need to make those new capture device as reliable as batch scanners</strong>.</p> <p>We believe an efficient capture process should be able to:</p> <ol> <li><em>detect and segment</em> the relevant document object during the preview phase;</li> <li><em>assess the quality</em> of the capture conditions and help the user improve them;</li> <li>optionally, <em>trigger the capture</em> at the perfect moment;</li> <li>and <em>produce a high-quality, controlled output</em> based on the high resolution captured image.</li> </ol> <p>This competition is focused on the first step of this process:<strong> </strong><strong>efficiently detect and segment document regions</strong>, as illustrated by following video showing the ideal output for the preview phase of some acquisition session: <a href="https://youtu.be/WNsI0R_rpO0">Click here to watch the video.</a> This video shows the ideal document object detection ‎(ie the ground truth, as a red frame)‎.</p> <p>For this challenge, the <strong>input</strong> consists in a set of <strong>videoclips containing a document</strong> from a predefined set, and the <strong>output</strong> should be an <strong>xml file containing the quadrilateral coordinates</strong> in which we can find the document per each frame of the video. Click <a href="https://sites.google.com/site/icdar15smartdoc/challenge-1/challenge1dataset">here</a> for detailed information about the dataset. </p> <p> </p> <p><strong>Licence</strong> for the dataset of challenge 1 (page outline detection in preview frames) :</p> <p>This work is licensed under a <strong>Creative Commons Attribution 4.0 International License</strong> <<a href="https://www.google.com/url?q=http://creativecommons.org/licenses/by/4.0/&sa=D&ust=1524734857667000&usg=AFQjCNEt4YXnUv2nXCFwkeOuBDqxDpvknQ">http://creativecommons.org/licenses/by/4.0/</a>>. Author attribution should be given by citing the following conference paper: Jean-Christophe Burie, Joseph Chazalon, Mickaël Coustaty, Sébastien Eskenazi, Muhammad Muzzamil Luqman, Maroua Mehri, Nibal Nayef, Jean-Marc OGIER, Sophea Prum and Marçal Rusinol: “ICDAR2015 Competition on Smartphone Document Capture and OCR (SmartDoc)”, In 13th International Conference on Document Analysis and Recognition (ICDAR), 2015.</p> <p><strong>If you use this dataset, please send us a short email at <icdar.smartdoc (at) gmail.com> to tell us why it was useful to you, and whether you have results or publications we can reference on our website. Thank you!</strong></p>
OCR fulltexts of the Digital Collections of the Berlin State Library (DC-SBB)
<p>The digital collections of the SBB contain 153,942 digitized works from the time period of 1470 to 1945.</p> <p>At the time of publication, 28,909 works have been OCR-processed resulting in 4,988,099 full-text pages.<br> For each page with OCR text, the language has been determined by <em>langid </em>(Lui/Baldwin 2012).</p> <p>corpus-entropy.pkl entropy rate per document page</p> <p>corpus-language.pkl language per document page</p> <p>corpus.zip fulltext corpus (extracts to .txt format)</p> <p>de_corpus.zip German sub-corpus (extracts to .txt format)</p> <p>selection_de.pkl Selection list of German documents</p> <p>xml2csv_alto.csv fulltext corpus per document page (incl.OCR word confidences)</p> <p> </p> <p><em>Sources</em></p> <p>Marco Lui and Timothy Baldwin. 2012. Langid.py:</p> <p>An off-the-shelf language identification tool. In Proceedings of the ACL 2012 System Demonstrations,</p> <p>ACL ’12, pages 25–30, Stroudsburg, PA, USA. Association for Computational Linguistics</p>
Pathway Figure OCR GMT
<p>Sample Pathway Figure OCR results from 17 December 2018, condensed by PMCID into GMT format for use in gene set analysis. Columns are PMCID, URL, Entrez Gene IDs</p>
CT-OCR-2022 (scroll001)
<p><strong>CT-OCR-2022 dataset</strong></p> <p>CT-OCR-2022 dataset contains optically scanned images for source paper document, 400 X-ray projections, 2687 CT-reconstructed cross-sections and segmentation markups for 6 model objects.</p> <p>Description of the data for each model object is presented in the table.</p> <table> <tbody><tr> <th>Files</th> <th>Data description</th> </tr> </tbody><tbody> <tr> <td>2022-ICMV-CT-OCR.[package].png</td> <td>sample projection and slice visualization</td> </tr> <tr> <td>[package].proj_src/proj_s_.log</td> <td>X-ray measurement log file</td> </tr> <tr> <td>[package].proj_src/*.tif</td> <td>preprocessed projections before rotation axe correction</td> </tr> <tr> <td>2022-ICMV-CT-OCR.[package].png</td> <td>package single projection and slice vizualization</td> </tr> <tr> <td>[package].proj_src/*.tif</td> <td>preprocessed projections before rotation axe correction</td> </tr> <tr> <td>[package].proj_norm/proj_s_.log</td> <td>X-ray measurement and geometry correction log file</td> </tr> <tr> <td>[package].proj_norm/*.tif</td> <td>preprocessed projections after rotation axe correction</td> </tr> <tr> <td>[package].rec_XXXX/metadata.json</td> <td>reconstruction metadata</td> </tr> <tr> <td>[package].rec_XXXX/*.tif</td> <td>CT-reconstructed volume, slices with size XXXX×XXXX</td> </tr> <tr> <td>[package].seg_XXXX/*.tif.seg.png</td> <td>segmentation markup</td> </tr> <tr> <td>[package].blank.png</td> <td>sample croped from pdf</td> </tr> <tr> <td>[package].scan.png</td> <td>sample croped from scanned image</td> </tr> </tbody> </table> <p>Due to the large amount of data, folders were packed into multi-volume zip-archives. Dataset published in Zenodo service in several linked repositories.</p> <p>scroll01 - <a href="https://doi.org/10.5281/zenodo.7123495">10.5281/zenodo.7123495</a><br>scroll02 - <a href="https://doi.org/10.5281/zenodo.7157600">10.5281/zenodo.7157600</a><br>scroll03 - <a href="https://doi.org/10.5281/zenodo.7157610">10.5281/zenodo.7157610</a><br>scroll04 - <a href="https://doi.org/10.5281/zenodo.7161350">10.5281/zenodo.7161350</a><br>folded01 - <a href="https://doi.org/10.5281/zenodo.7162001">10.5281/zenodo.7162001</a>, <a href="https://doi.org/10.5281/zenodo.7164152">10.5281/zenodo.7164152</a><br>folded02 - <a href="https://doi.org/10.5281/zenodo.7267141">10.5281/zenodo.7267141</a>, <a href="https://doi.org/10.5281/zenodo.7272064">10.5281/zenodo.7272064</a></p> <p>Any questions, complaints, etc. can be directed to: polevoy@smartengines.com (Dmitry Polevoy)</p> <p><strong>Share and Cite:</strong></p> <p>D. V. Polevoy, P. A. Kulagin, A. S. Ingacheva, Zh. V. Soldatova, M. V. Chukalina, D. P. Nikolaev, V. V. Arlazarov, "From tomographic reconstruction to automatic text recognition: the next frontier task for the artificial intelligence," Proc. SPIE 12701, Fifteenth International Conference on Machine Vision (ICMV 2022), 127010P (7 June 2023); https://doi.org/10.1117/12.2680132</p> <p>in BibTex format:</p> <p>@inproceedings{10.1117/12.2680132,<br>author = {D. V. Polevoy and P. A. Kulagin and A. S. Ingacheva and Zh. V. Soldatova and M. V. Chukalina and D. P. Nikolaev and V. V. Arlazarov},<br>title = {{From tomographic reconstruction to automatic text recognition: the next frontier task for the artificial intelligence}},<br>volume = {12701},<br>booktitle = {Fifteenth International Conference on Machine Vision (ICMV 2022)},<br>editor = {Wolfgang Osten and Dmitry P. Nikolaev and Jianhong (Jessica) Zhou},<br>organization = {International Society for Optics and Photonics},<br>publisher = {SPIE},<br>pages = {127010P},<br>keywords = {virtual unrolling, virtual unwrapping, digital unfolding, computational flattening, computed tomography, non-destructive analysis, open dataset},<br>year = {2023},<br>doi = {10.1117/12.2680132},<br>URL = {https://doi.org/10.1117/12.2680132}<br>}</p> <p><strong>See also</strong></p> <p>P. A. Kulagin, D. V. Polevoy, M. V. Chukalina, D. P. Nikolaev and V. V. Arlazarov, “Fully automatic virtual unwrapping method for documents imaged by X-ray tomography,” Proc. ICDAR 2024, to be published.</p> <p><a href="https://github.com/SmartEngines/virtual-unwrapping-article-code">https://github.com/SmartEngines/virtual-unwrapping-article-code</a></p>
Drinov Orthography for Post-OCR Correction dataset
<p>The Drinov Orthography for Post-OCR Correction (DOPOC) dataset was created by annotating a historical newspaper collection provided by the <a href="https://digital.libplovdiv.com/en">National Library "Ivan Vazov"</a> (NLIV) in Plovdiv, Bulgaria. We consider printed versions of these documents, which we manually annotate and align at the character level in the same format as the one from the ICDAR 2019 post-OCR correction competition.</p>
CT-OCR-2022 (folded002, part2)
<p><strong>CT-OCR-2022 dataset</strong></p> <p>CT-OCR-2022 dataset contains optically scanned images for source paper document, X-ray projections, CT-reconstructed cross-sections and segmentation markups for model objects.</p> <p>Description of the structure of the dataset, contacts and information about the reference to the dataset see at <a href="https://doi.org/10.5281/zenodo.7123495">https://doi.org/10.5281/zenodo.7123495</a>.</p>
CT-OCR-2022 (scroll002)
<p><strong>CT-OCR-2022 dataset</strong></p> <p>CT-OCR-2022 dataset contains optically scanned images for source paper document, X-ray projections, CT-reconstructed cross-sections and segmentation markups for model objects.</p> <p>Description of the structure of the dataset, contacts and information about the reference to the dataset see at <a href="https://doi.org/10.5281/zenodo.7123495">https://doi.org/10.5281/zenodo.7123495</a>.</p>
CT-OCR-2022 (folded002, part1)
<p><strong>CT-OCR-2022 dataset</strong></p> <p>CT-OCR-2022 dataset contains optically scanned images for source paper document, X-ray projections, CT-reconstructed cross-sections and segmentation markups for model objects.</p> <p>Description of the structure of the dataset, contacts and information about the reference to the dataset see at <a href="https://doi.org/10.5281/zenodo.7123495">https://doi.org/10.5281/zenodo.7123495</a>.</p>
CT-OCR-2022 (folded001, part2)
<p><strong>CT-OCR-2022 dataset</strong></p> <p>CT-OCR-2022 dataset contains optically scanned images for source paper document, X-ray projections, CT-reconstructed cross-sections and segmentation markups for model objects.</p> <p>Description of the structure of the dataset, contacts and information about the reference to the dataset see at <a href="https://doi.org/10.5281/zenodo.7123495">https://doi.org/10.5281/zenodo.7123495</a>.</p>
OCRed text of the Allgemeine musikalische Zeitung (ALTO format) from 1798-1848 and 1863-1965
<p><span>Dataset as used in the LREC 2020 publication '<em>Allgemeine Musikalische Zeitung</em></span><span> as a Searchable Online Corpus'</span>.</p>
NewsEye / READ OCR training dataset from French Newspapers (18th, 19th, early 20th C.)
<p>The dataset comprises French newspaper pages from 18th, 19th and early 20th century with carefully corrected text. The page images were provided by the <a href="https://www.bnf.fr/en">French National Library</a> and comprise 127 pages (training set) and 8 pages (validation set). The data are formed according to the PAGE format (cf. Cf. <a href="https://github.com/PRImA-Research-Lab/PAGE-XML/">https://github.com/PRImA-Research-Lab/PAGE-XML/</a>) and were produced with the <a href="http://read.transkribus.eu/">Transkribus </a>platform with support of the <a href="http://newseye.eu/">NewsEye</a> and the <a href="http://read.transkribus.eu/">READ </a>project.</p>
CIS OCR Workshop v1.0: OCR and postcorrection of early printings for digital humanities
<p>The 2-day CIS OCR Workshop on "OCR and postcorrection of early printings for digital humanities" originally held at LMU, Munich 14/15 September 2015 (see http://www.cis.lmu.de/ocrworkshop).</p> <p>Release date: 2016-02-25</p> <p><br /> CIS OCR Workshop by Uwe Springmann, Florian Fink is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.</p>
Castren 1844: Elementa grammaticae Syrjaenae, OCR Ground Truth
<p>Matthias Alexander Castrén published Elementa grammaticae Syrjaenae in 1844. This dataset contains two scans of this work, the higher quality version originating from the Internet Archive: </p> <p>https://archive.org/details/elementagrammati00cast/mode/2up</p> <p>All pages are layout detected with Transkribus, and there are 26 proofread pages. Three pages contain table layouts. </p>
Tesseract OCR of IIT-CDIP Dataset
<p>This is Tesseract generated <strong>transcriptions (no images)</strong> of (most of) the IIT-CDIP dataset. To download the images of the IIT-CDIP dataset go to <a href="https://data.nist.gov/od/id/mds2-2531">https://data.nist.gov/od/id/mds2-2531</a> </p> <p>The directory struture of this dataset is the same as the IIT-CDIP dataset (although has everything in one tar, with "a.a", "a.b", ... directories) and can thus be combine with the image IIT-CDIP dataset using rsync or similar tool. This dataset contains a "X.layout.json" for each "X.png" in the IIT-CDIP dataset (doesn't have sections 'a', 'w', 'x', 'y', and 'z').</p> <p>The jsons contain block/paragraph, line and word bounding boxes, with transcriptions for the words following the Tesseract format. The line and word annotations are directly taken from Tesseract. The block and paragraph output of Tesseract was discarded. The images were then run through both the Publaynet and PrimaNet models available on LayoutParser (<a href="https://layout-parser.github.io/">https://layout-parser.github.io/</a>). The combine output of these models became the block/paragraph annotations (we kept the Tesseract output format, but each block has 1 paragraph of exactly the same shape).</p> <p><strong>Important:</strong> There is also a "rotation" value in the json (0, 90, 180, or 270) indicating the json may be for a rotated version of the IIT-CDIP image by the given amount (attempted to rotated documents to upright position to get better OCR results).</p> <p>These are the annotations used to pre-train <strong>Dessurt</strong> (<a href="https://arxiv.org/abs/2203.16618">https://arxiv.org/abs/2203.16618</a>).</p> <p>These annotations will be worse than those that would be obtained using a commercial OCR system (like those used to pre-train LayoutLMv2/v3).</p> <p>The code used to produce these annotations is available here: <a href="https://github.com/herobd/ocr">https://github.com/herobd/ocr</a></p>
OCR model for Pracalit for Sanskrit and Newar MSS 16th to 19th C., Ground Truth
<p>Ground truth data (png and xml files) for a an OCR model. Will be continually updated.</p> <p>Originally trained on Transkribus with a PyLaia model created from ground truth data based on transcripts into Pracalit Unicode of four Nepalese manuscripts. The manuscripts used to create this model are Staatsbibliothek zu Berlin's Hitopadeśa (MIK I 4851) (mixed Newar and Sanskrit dating to 1561) and Vetālapañcaviṃśati (HS. Or. 6414) (Newar dating to 1675) as well as Cambridge Digital Library's Avalokiteśvaraguṇakāraṇḍavyūha (MS Add. 1322) (Sanskrit, 18th century) and the Royal Asiatic Society Online Collection's Madhyamasvayaṃbhūpurāṇa (RAS Hodgson MS 23) (Newar and Sanskrit dating to c. 1800).</p> <p>The training was done on 441 pages and validation on 242 pages.</p> <p>This model does not recognise spacing, except for large gaps (i.e. for pictures or string holes). Newar word divider markers may not be represented or may be transcribed as virama. In general, the model is made for MSS with scriptio continua and will transcribe into scriptio continua into Pracalit Unicode.</p> <p>Transcription was performed by Dr Alexander O'Neill (SOAS University of London). Transcription of the Vetālapañcaviṃśati (HS. Or. 6414) and Madhyamasvayaṃbhūpurāṇa (RAS Hodgson MS 23) was aided by unpublished materials provided by Dr Felix Otter (Philipps-Universität Marburg), as well as the published transcription in Shakya, Min Bahadur, and Shanta Harsha Bajracharya, eds. "Svayambhū Purāṇa." Lalitpur: Nagarjuna Institute of Exact Methods, 2001. The transcription of Avalokiteśvaraguṇakāraṇḍavyūha (MS Add. 1322) was aided by the transcription provided by the Digital Sanskrit Buddhist Canon Project based on Lokesh Chandra, "Guṇakāraṇḍavyūhasūtram," New Delhi: International Academy of Indian Culture, 1999.</p>
GT4HistOCR: Ground Truth for training OCR engines on historical documents in German Fraktur and Early Modern Latin
<p><strong>GT4HistOCR</strong> contains ground truth for research in Optical Character Recognition (OCR) technology applied to historical printings in German Fraktur and Early Modern Latin.</p> <p>The ground truth comes in pairs of images of single printed lines as they appear in book pages (*.png) and their corresponding diplomatic transcriptions (*.gt.txt), which are UTF-8 strings preserving the character forms (glyphs) as much as possible within the UNICODE standard. These pairs of line images and their transcriptions can be directly used to train recognition models with, e.g., the open source OCR engines <em>OCRopy</em> or <em>Tesseract</em>. A total of 313,173 ground truth lines are provided.</p> <p><strong>Please note that the subcorpora making up this collection used different transcription guidelines, so it is a bad idea to train a recognition model on the total collection! Rather train individual models for each subcorpus.</strong> Fur further information about the subcorpora, please see the README file and the accompanying publication.</p> <p>If these data are useful for you, please cite the accompanying publication:</p> <pre>@article{<a href="http://springmann.net/publications.html#springmann2018gt4hist">springmann2018gt4hist</a>, author = {Uwe Springmann and Christian Reul and Stefanie Dipper and Johannes Baiter}, title = {{Ground Truth for training {OCR} engines on historical documents in German Fraktur and Early Modern Latin}}, journal = {J. Lang. Technol. Comput. Linguistics}, volume = {33}, number = {1}, pages = {97--114}, year = {2018}, url = {https://jlcl.org/content/2-allissues/1-heft1-2018/jlcl_2018-1_5.pdf} }</pre>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.