Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

8

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

8 results for “Tesseract”

Learn how ShareScore rates datasets ↗
zenodo40/100

Tesseract OCR of IIT-CDIP Dataset

<p>This is&nbsp;Tesseract generated&nbsp;<strong>transcriptions (no&nbsp;images)</strong>&nbsp;of (most of) the IIT-CDIP dataset. To download the images of the IIT-CDIP dataset go to&nbsp;<a href="https://data.nist.gov/od/id/mds2-2531">https://data.nist.gov/od/id/mds2-2531</a>&nbsp;</p> <p>The directory struture of this dataset is the same as the IIT-CDIP dataset (although has everything in one tar, with &quot;a.a&quot;, &quot;a.b&quot;, ... directories)&nbsp;and can thus be combine with the image&nbsp;IIT-CDIP dataset&nbsp;using rsync or similar tool. This dataset&nbsp;contains a &quot;X.layout.json&quot; for each &quot;X.png&quot; in the IIT-CDIP dataset (doesn&#39;t have sections &#39;a&#39;, &#39;w&#39;, &#39;x&#39;, &#39;y&#39;, and &#39;z&#39;).</p> <p>The jsons contain block/paragraph, line and word bounding boxes, with&nbsp;transcriptions for the words following the Tesseract format. The line and word annotations are directly taken from Tesseract. The block and paragraph output of Tesseract was discarded. The images were then run through both the&nbsp;Publaynet&nbsp;and PrimaNet&nbsp;models available on LayoutParser (<a href="https://layout-parser.github.io/">https://layout-parser.github.io/</a>). The combine output of these models became the block/paragraph annotations (we kept the Tesseract output format, but each block has 1 paragraph of exactly the same shape).</p> <p><strong>Important:</strong> There is also a &quot;rotation&quot; value in the json (0, 90, 180, or 270) indicating the json may be&nbsp;for a rotated version of the IIT-CDIP image by the given amount (attempted to rotated documents to upright position to get better OCR results).</p> <p>These are the annotations used to pre-train <strong>Dessurt</strong> (<a href="https://arxiv.org/abs/2203.16618">https://arxiv.org/abs/2203.16618</a>).</p> <p>These annotations will be worse than&nbsp;those that would be obtained&nbsp;using&nbsp;a commercial OCR system&nbsp;(like those used to pre-train LayoutLMv2/v3).</p> <p>The code used to produce these&nbsp;annotations is available here:&nbsp;<a href="https://github.com/herobd/ocr">https://github.com/herobd/ocr</a></p>

opencc-by-4.0May 2022View details →
zenodo40/100

Fig. 3 in The Wettstein tesseract: A tool for conceptualising species-rank decisions and illustrating speciation trajectories

Fig. 3. The Wettstein tesseract as a tool for illustrating geographical (allopatric, peripatric, parapatric) or non-geographical (sympatric) pathways of speciation. Arrows indicate different paths through the tesseract from a coherent population to final stage of two evolutionary independent entities (species). Note that these paths end whenever species rank is achieved; subsequent changes of biogeographical patterns (allopatrically, peripatrically, or parapatrically formed species becoming sympatric and sympatrically formed species becoming allopatric) are not illustrated.

opencc-by-4.0Feb 2023View details →
zenodo40/100

Fig. 2 in The Wettstein tesseract: A tool for conceptualising species-rank decisions and illustrating speciation trajectories

Fig. 2. The Wettstein tesseract unfolded, with vertices representing binary-coded (statistically significant) genealogical, geographical, ecological, and morphological differences between two sister-taxa under study (left) and their taxonomical counterparts (right). Colours of vertices are just for clarity of the illustration and have no semantic implication.

opencc-by-4.0Feb 2023View details →
zenodo40/100

Fig. 1. The Wettstein tesseract – a in The Wettstein tesseract: A tool for conceptualising species-rank decisions and illustrating speciation trajectories

Fig. 1. The Wettstein tesseract – a four-dimensional hypercube for illustrating species delimitation and speciation trajectories, with the four dimensions representing geographical differentiation, ecological divergence, morphological difference, and genealogical independence, respectively, and species rank attributed to the vertices in red (while vertices in black and white indicate infraspecific diversification patterns without or with significant contribution of a genealogical factor, respectively).

opencc-by-4.0Feb 2023View details →
zenodo36/100

Tesseract OCR models for the Alsatian dialects

<p>This dataset provides trained Tesseract (<a href="https://github.com/tesseract-ocr/tesseract">https://github.com/tesseract-ocr/tesseract</a>) OCR models for the Alsatian dialects. These models were developed in the context of the RESTAURE project, funded by the French ANR.&nbsp;</p> <p>Two models are provided :</p> <p>The first model, ISKO_2015, has been presented in the following article: <a href="http://hal.archives-ouvertes.fr/hal-01252241">https://hal.archives-ouvertes.fr/hal-01252241</a>. The Tesseract model has been trained using the jTessBoxEditor tool (<a href="http://vietocr.sourceforge.net/training.html">http://vietocr.sourceforge.net/training.html</a>), Version 1.4 (2 May 2015), based on images automatically generated from the training texts (excerpts from 7 different printed works, totalling about 9,000 words). The generation of the images used a 36pt font size, and two fonts were used (Arial and Times New Roman), with their normal and italic variants.<br> The Tesseract model (gsw.traineddata) can be used with Tesseract 3.0x.</p> <p>The second model, 2018, has been trained for Tesseract 4.0x, using jTessBoxEditor version 2.0.1 (28 July 2018). Again, images were automatically generated from the training text. The training text is different from the one used for the ISKO_2015 model and is &quot;artificial&quot;, in the sense that it has been built by appending word n-grams extracted from a large variety of published texts in Alsatian, for a time period spanning 2 centuries and for different text genres. The images corresponding to this training text have been automatically generated with the Tesseract text2image tool, using the following parameters: --ptsize=36 --leading=20. The fonts used are listed in the gsw.font_properties file.</p> <p>Dictionary data has also been used for training. We conflated Alsatian words found in several lexicons and corpora:</p> <ul> <li>Lexicons produced by the OLCA (Office pour la Langue et les Cultures d&#39;Alsace et de Moselle): <a href="http://www.olcalsace.org/fr/lexiques">http://www.olcalsace.org/fr/lexiques</a></li> <li>Lexicon from a Wiktionary user page: <a href="http://fr.wiktionary.org/wiki/Utilisateur:Laurent_Bouvier/alsacien-fran%C3%A7ais">https://fr.wiktionary.org/wiki/Utilisateur:Laurent_Bouvier/alsacien-fran%C3%A7ais</a></li> <li>Lexicon from the ACPA association: <a href="http://web.archive.org/web/20160302234127/http:/culture.alsace.pagesperso-orange.fr/dictionnaire_alsacien.htm">http://web.archive.org/web/20160302234127/http:/culture.alsace.pagesperso-orange.fr/dictionnaire_alsacien.htm</a></li> <li>Chronicles published by Raymond Matzen in the local newspaper &quot;Les Derni&egrave;res Nouvelles d&#39;Alsace&quot;</li> <li>Transcriptions of television shows found in Erhart, P. (2012). <em>Les dialectes dans les m&eacute;dias: quelle image de l&rsquo;Alsace v&eacute;hiculent-ils dans les &eacute;missions de la t&eacute;l&eacute;vision r&eacute;gionale?</em>, Universit&eacute; de Strasbourg, <a href="http://www.theses.fr/167563386">http://www.theses.fr/167563386</a></li> <li>French-Alsatian parallel corpus provided by the OLCA</li> <li>Excerpts from Adolf, P. (2006). <em>Dictionnaire comparatif multilingue: fran&ccedil;ais-allemand-alsacien-anglais.</em>, Strasbourg, France, Midgard, 2006, 373 p.</li> </ul> <p>The Tesseract models can be used&nbsp; for instance using the gImageReader tool (<a href="https://github.com/manisandro/gImageReader">https://github.com/manisandro/gImageReader</a>), which provides a graphical user interface for the Tesseract tool.&nbsp;</p> <p>When evaluated against the same test corpus (prose by Marie Hart, theater and poetry by Gustave Stokopf and prose by Charles Zumstein, totalling about 4,900 words), both models achieve roughly the same performance levels. Usually, even better performance levels can be achieved by combining the Alsatian-specific model with the French and German models available for Tesseract (available from <a href="https://github.com/tesseract-ocr/tessdata">https://github.com/tesseract-ocr/tessdata</a>)</p>

opencc-by-sa-4.0Aug 2018View details →
zenodo32/100

purple tesseract

just the tesseract but purple Source: Objaverse 1.0 / Sketchfab

opencc-byJun 2022View details →
zenodo32/100

Tesseract

<p>Demonstration&nbsp;video of the Tesseract project, using a painting robot arm to visualise a live performance with solo augmented violin and another with piano.</p>

opencc-by-4.0Sep 2023View details →
zenodo28/100

Trained tesseract networks for low-resolution optical character recognition

<p>Tesseract eng.traineddata files for low-resolution optical character recognition (English).</p>

opengpl-2.0-or-laterJul 2020View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record