Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
8
datasets available to search
ShareScore release 0.9.0
Dataset results
8 results for “Tesseract”
Tesseract OCR of IIT-CDIP Dataset
<p>This is Tesseract generated <strong>transcriptions (no images)</strong> of (most of) the IIT-CDIP dataset. To download the images of the IIT-CDIP dataset go to <a href="https://data.nist.gov/od/id/mds2-2531">https://data.nist.gov/od/id/mds2-2531</a> </p> <p>The directory struture of this dataset is the same as the IIT-CDIP dataset (although has everything in one tar, with "a.a", "a.b", ... directories) and can thus be combine with the image IIT-CDIP dataset using rsync or similar tool. This dataset contains a "X.layout.json" for each "X.png" in the IIT-CDIP dataset (doesn't have sections 'a', 'w', 'x', 'y', and 'z').</p> <p>The jsons contain block/paragraph, line and word bounding boxes, with transcriptions for the words following the Tesseract format. The line and word annotations are directly taken from Tesseract. The block and paragraph output of Tesseract was discarded. The images were then run through both the Publaynet and PrimaNet models available on LayoutParser (<a href="https://layout-parser.github.io/">https://layout-parser.github.io/</a>). The combine output of these models became the block/paragraph annotations (we kept the Tesseract output format, but each block has 1 paragraph of exactly the same shape).</p> <p><strong>Important:</strong> There is also a "rotation" value in the json (0, 90, 180, or 270) indicating the json may be for a rotated version of the IIT-CDIP image by the given amount (attempted to rotated documents to upright position to get better OCR results).</p> <p>These are the annotations used to pre-train <strong>Dessurt</strong> (<a href="https://arxiv.org/abs/2203.16618">https://arxiv.org/abs/2203.16618</a>).</p> <p>These annotations will be worse than those that would be obtained using a commercial OCR system (like those used to pre-train LayoutLMv2/v3).</p> <p>The code used to produce these annotations is available here: <a href="https://github.com/herobd/ocr">https://github.com/herobd/ocr</a></p>
Fig. 3 in The Wettstein tesseract: A tool for conceptualising species-rank decisions and illustrating speciation trajectories
Fig. 3. The Wettstein tesseract as a tool for illustrating geographical (allopatric, peripatric, parapatric) or non-geographical (sympatric) pathways of speciation. Arrows indicate different paths through the tesseract from a coherent population to final stage of two evolutionary independent entities (species). Note that these paths end whenever species rank is achieved; subsequent changes of biogeographical patterns (allopatrically, peripatrically, or parapatrically formed species becoming sympatric and sympatrically formed species becoming allopatric) are not illustrated.
Fig. 2 in The Wettstein tesseract: A tool for conceptualising species-rank decisions and illustrating speciation trajectories
Fig. 2. The Wettstein tesseract unfolded, with vertices representing binary-coded (statistically significant) genealogical, geographical, ecological, and morphological differences between two sister-taxa under study (left) and their taxonomical counterparts (right). Colours of vertices are just for clarity of the illustration and have no semantic implication.
Fig. 1. The Wettstein tesseract – a in The Wettstein tesseract: A tool for conceptualising species-rank decisions and illustrating speciation trajectories
Fig. 1. The Wettstein tesseract – a four-dimensional hypercube for illustrating species delimitation and speciation trajectories, with the four dimensions representing geographical differentiation, ecological divergence, morphological difference, and genealogical independence, respectively, and species rank attributed to the vertices in red (while vertices in black and white indicate infraspecific diversification patterns without or with significant contribution of a genealogical factor, respectively).
Tesseract OCR models for the Alsatian dialects
<p>This dataset provides trained Tesseract (<a href="https://github.com/tesseract-ocr/tesseract">https://github.com/tesseract-ocr/tesseract</a>) OCR models for the Alsatian dialects. These models were developed in the context of the RESTAURE project, funded by the French ANR. </p> <p>Two models are provided :</p> <p>The first model, ISKO_2015, has been presented in the following article: <a href="http://hal.archives-ouvertes.fr/hal-01252241">https://hal.archives-ouvertes.fr/hal-01252241</a>. The Tesseract model has been trained using the jTessBoxEditor tool (<a href="http://vietocr.sourceforge.net/training.html">http://vietocr.sourceforge.net/training.html</a>), Version 1.4 (2 May 2015), based on images automatically generated from the training texts (excerpts from 7 different printed works, totalling about 9,000 words). The generation of the images used a 36pt font size, and two fonts were used (Arial and Times New Roman), with their normal and italic variants.<br> The Tesseract model (gsw.traineddata) can be used with Tesseract 3.0x.</p> <p>The second model, 2018, has been trained for Tesseract 4.0x, using jTessBoxEditor version 2.0.1 (28 July 2018). Again, images were automatically generated from the training text. The training text is different from the one used for the ISKO_2015 model and is "artificial", in the sense that it has been built by appending word n-grams extracted from a large variety of published texts in Alsatian, for a time period spanning 2 centuries and for different text genres. The images corresponding to this training text have been automatically generated with the Tesseract text2image tool, using the following parameters: --ptsize=36 --leading=20. The fonts used are listed in the gsw.font_properties file.</p> <p>Dictionary data has also been used for training. We conflated Alsatian words found in several lexicons and corpora:</p> <ul> <li>Lexicons produced by the OLCA (Office pour la Langue et les Cultures d'Alsace et de Moselle): <a href="http://www.olcalsace.org/fr/lexiques">http://www.olcalsace.org/fr/lexiques</a></li> <li>Lexicon from a Wiktionary user page: <a href="http://fr.wiktionary.org/wiki/Utilisateur:Laurent_Bouvier/alsacien-fran%C3%A7ais">https://fr.wiktionary.org/wiki/Utilisateur:Laurent_Bouvier/alsacien-fran%C3%A7ais</a></li> <li>Lexicon from the ACPA association: <a href="http://web.archive.org/web/20160302234127/http:/culture.alsace.pagesperso-orange.fr/dictionnaire_alsacien.htm">http://web.archive.org/web/20160302234127/http:/culture.alsace.pagesperso-orange.fr/dictionnaire_alsacien.htm</a></li> <li>Chronicles published by Raymond Matzen in the local newspaper "Les Dernières Nouvelles d'Alsace"</li> <li>Transcriptions of television shows found in Erhart, P. (2012). <em>Les dialectes dans les médias: quelle image de l’Alsace véhiculent-ils dans les émissions de la télévision régionale?</em>, Université de Strasbourg, <a href="http://www.theses.fr/167563386">http://www.theses.fr/167563386</a></li> <li>French-Alsatian parallel corpus provided by the OLCA</li> <li>Excerpts from Adolf, P. (2006). <em>Dictionnaire comparatif multilingue: français-allemand-alsacien-anglais.</em>, Strasbourg, France, Midgard, 2006, 373 p.</li> </ul> <p>The Tesseract models can be used for instance using the gImageReader tool (<a href="https://github.com/manisandro/gImageReader">https://github.com/manisandro/gImageReader</a>), which provides a graphical user interface for the Tesseract tool. </p> <p>When evaluated against the same test corpus (prose by Marie Hart, theater and poetry by Gustave Stokopf and prose by Charles Zumstein, totalling about 4,900 words), both models achieve roughly the same performance levels. Usually, even better performance levels can be achieved by combining the Alsatian-specific model with the French and German models available for Tesseract (available from <a href="https://github.com/tesseract-ocr/tessdata">https://github.com/tesseract-ocr/tessdata</a>)</p>
purple tesseract
just the tesseract but purple Source: Objaverse 1.0 / Sketchfab
Tesseract
<p>Demonstration video of the Tesseract project, using a painting robot arm to visualise a live performance with solo augmented violin and another with piano.</p>
Trained tesseract networks for low-resolution optical character recognition
<p>Tesseract eng.traineddata files for low-resolution optical character recognition (English).</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.