Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

30

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

30 results for “text recognition”

Learn how ShareScore rates datasets ↗
zenodo48/100

Named Entity Recognition Dataset for Dutch Biographical Texts

<p>A dataset for Named Entity Recognition for Dutch biographies. The original data is available in the Biographical portal of the Netherlands (http://www.biografischportaal.nl/). The annotations are for 6 types of entities: PERSON, LOCATION, ORGANIZATION, DATE, ARTWORK, MISC. Additionally, the CoNLL formatted files were manually checked for tokenization and sentence splitting.</p>

opencc-by-4.0Dec 2023View details →
zenodo44/100

Transkribus - Handwritten Text Recognition for Premodern Documents (SIMS 2020 Lightning Talk)

<p>Transkribus is a platform for text recognition and can be used via the Transkribus Expert Software (available after registration: transkribus.eu). Through Transkribus different tools for document analysis and text recognition can be directly applied. The intro demonstrates very briefly how Transkribus can help with regards to premodern documents especially since a variety of pre-trained models are already available: for Latin (prints and handwriting), for early modern vernaculars in French, Dutch, English, and German. For more information go to transkribus.eu.</p> <p>Presented as a Schoenberg Symposium 2020 Lightning Talk</p>

opencc-by-4.0Nov 2020View details →
zenodo44/100

Ground-Truthed Data Set of Zenon Papyri for Handwritten Text Recognition

<p>Diplomatic transcription of papyri found in the Zenon archive [see <a href="https://en.wikipedia.org/wiki/Zenon_of_Kaunos">en.wikipedia.org/wiki/Zenon_of_Kaunos</a>]</p> <p>Manually prepared as PageXML with Transkribus within <a href="http://d-scribes.philhist.unibas.ch/">D-Scribes</a> project.</p> <p>&nbsp;</p>

opencc-by-4.0May 2022View details →
zenodo44/100

A Brazilian Portuguese Dataset for Offline Handwritten Text Recognition (BRESSAY)

<p>The BRESSAY dataset comprises images of handwritten essays in Brazilian Portuguese, which present a series of challenges to optical recognition models. These images were sourced from multiple online platforms, limiting our ability to standardize the capture process. Due to these varied sources and the lack of a uniform collection method, the dataset provides a realistic reflection of real-world conditions. Each essay is unique, contributed by different writers, and addresses a specific content topic. Furthermore, the constraints placed on the writers often lead to various handwriting scenarios, including hard-to-read words, connected words, noise, overwriting, and struck-through texts.</p> <h3><strong>Technical Details</strong></h3> <p>The BRESSAY dataset represents a comprehensive collection of handwritten essays in Brazilian Portuguese, offering detailed insights into various handwriting scenarios. It covers a total of 1,000 pages, each contributed by a unique writer, resulting in 1,000 distinct handwriting styles. This aspect of the dataset adds a layer of diversity, which is further emphasized by the total of 4,214 paragraphs, 30,090 lines, and 416,826 words. Regarding unique tokens, we have 41,318 unique words, and 107 unique characters.</p> <h3><strong>Data Structure</strong></h3> <p>The dataset is organized as follows:</p> <ul> <li>data/: Main folder containing segmented essay images <ul> <li>lines/: Images of individual lines <ul> <li>&nbsp; PNG files: Line images</li> <li>&nbsp; TXT files: Transcriptions of lines</li> </ul> </li> <li>pages/: Full page essay images <ul> <li>&nbsp; PNG files: Page images</li> <li>&nbsp; TXT files: Transcriptions of pages</li> </ul> </li> <li>paragraphs/: Images of paragraphs <ul> <li>&nbsp; PNG files: Paragraph images</li> <li>&nbsp; TXT files: Transcriptions of paragraphs</li> </ul> </li> <li>words/: Images of individual words <ul> <li>&nbsp; PNG files: Word images</li> <li>&nbsp; TXT files: Transcriptions of words</li> </ul> </li> </ul> </li> <li>sets/: Contains partition files <ul> <li>test.txt: Names of images in the test set</li> <li>validation.txt: Names of images in the validation set</li> <li>training.txt: Names of images in the training set</li> </ul> </li> </ul> <h3><strong>Dataset Usage and Annotations</strong></h3> <p>Each name in test.txt, validation.txt and training.txt represents the name of the page and all its content (words, lines, paragraphs) must be in the respective partition.</p> <p>Annotations used in the dataset:</p> <ul> <li>&nbsp; <code>##@@???@@##</code>: Superscript text that has become unidentifiable and unreadable.</li> <li>&nbsp; <code>$$@@???@@$$</code>: Subscript text that has become unidentifiable and unreadable.</li> <li>&nbsp; <code>@@???@@</code>: Text that cannot be read or identified due to its illegibility.</li> <li>&nbsp; <code>##--xxx--##</code>: Text that has been added as a superscript and subsequently crossed out, rendering it illegible.</li> <li>&nbsp; <code>$$--xxx--$$</code>: Text that has been added as a subscript and subsequently crossed out, rendering it illegible.</li> <li>&nbsp; <code>--xxx--</code>: Text that has been crossed out in a way that makes it unreadable.</li> <li>&nbsp; <code>##--text--##</code>: Text that has been added as a superscript and subsequently crossed out, but remains legible.</li> <li>&nbsp; <code>$$--text--$$</code>: Text that has been added as a subscript and subsequently crossed out, but remains legible.</li> <li>&nbsp; <code>##text##</code>: Text added as a superscript in the line, typically as a correction or additional note.</li> <li>&nbsp; <code>$$text$$</code>: Text added as a subscript in the line, typically as a correction or additional note.</li> <li>&nbsp; <code>--text--</code>: Text that has been crossed out but remains readable.</li> </ul>

opencc-by-4.0May 2024View details →
zenodo44/100

FloraNER: a Named Entity Recognition Dataset for Botanical French Text

<p>FloraNER is a Named-Entity Recognition (NER) dataset for botanical french literature. The dataset covers plant species names and plant morphological terms for both plant organs/characteristics and their descriptors. the descriptors are annotated both in a coarse-grained manner as the named entity type "DESCRIPTOR" and in a fine-grained manner, categorized into the following named-entity types: Form, Measure, Surface, Color, Position, Disposition, Structure, and Development. FloraNER is distantly annotated using a specialized botanical corpus. Consequently, it's important to note that not all named entities within the text are captured and annotated for the coarse-grained and fine-grained datasets.</p>

opencc-by-4.0May 2024View details →
zenodo44/100

Raw images (photographs) of urban text scenes for camera-based Thai text recognition

<p>Raw image collection of city scenes in Thailand with text content.<br> Text is photographed from diffeent angles. Also morning and evening<br> photographs were taken in order to capture different lighting<br> conditions. The material, 309 images, was photographed in 2013 by<br> Bowornrat Sriman and volunteers.</p> <p>Example EXIF:<br> JPEG image data, Exif standard: [TIFF image data, little-endian,<br> direntries=13, height=2448, manufacturer=SAMSUNG, model=GT-I9300,<br> orientation=upper-right, xresolution=220, yresolution=228,<br> resolutionunit=2, software=I9300XXEMA2, datetime=2013:03:14 18:17:49,<br> GPS-Data, width=3264], baseline, precision 8, 3264x2448, frames 3</p> <p>The images are not labeled. The orientation (landscape/portrait) is<br> not corrected yet. This material was used in preparation of the publication:</p> <p>Sriman, B.&nbsp; &amp; Schomaker, L. (2015).<br> Object Attention Patches for Text Detection and Recognition in Scene Images using SIFT,<br> Proceedings of the International Conference on Pattern Recognition Applications and<br> Methods: ICPRAM 2015.&nbsp; De Marsico, M., Figueiredo, M.&nbsp; &amp; Fred, A.<br> (Eds.).&nbsp; Lisbon, Portugal: SciTePress, Vol.&nbsp; 1, p.&nbsp; 304-311 8 p.</p> <p>Please cite this publication when using these data.</p>

opencc-by-4.0Mar 2013View details →
zenodo44/100

Dataset for ICFHR2018 Competition on Automated Text Recognition on a READ Dataset

<p>The main idea of this dataset is to analyse the impact of training data. How many training data&nbsp;specific to the document, you are transcribing, is necessary?&nbsp;&nbsp;&nbsp;</p> <p><strong>general data: </strong>This is a&nbsp;collection of heterogeneous documents to train an initial system. For each text line there is an image file of that line, a file with the ground truth text and an information file containing an automatically generated surrounding polygon.</p> <p><strong>specific data: </strong>The&nbsp;specific data contains documents related to the test data. For the specific systems only the images of the train list may be used. The file are of the same type as the general data.</p> <p><strong>test data:&nbsp;</strong>The test data contains only the images and the information files.</p> <p>More Information, some published results and an evaluation procedure at&nbsp;https://scriptnet.iit.demokritos.gr/competitions/10/</p>

opencc-by-4.0Sep 2018View details →
zenodo44/100

The e-NDP project : collaborative digital edition of the Chapter registers of Notre-Dame of Paris (1326-1504). Ground-truth for handwriting text recognition (HTR) on late medieval manuscripts.

<p>The <a href="https://endp.hypotheses.org/">e-NDP project</a>, funded by the ANR, is led by the <a href="https://lamop.hypotheses.org/6870">LaMOP</a> (Julie Claustre and Darwin Smith).</p> <p>The project&#39;s partners are the Archives nationales, the&nbsp;Biblioth&egrave;que nationale de France (Department of Manuscripts, Biblioth&egrave;que de l&#39;Arsenal), the &Eacute;cole nationale des chartes and the Biblioth&egrave;que Mazarine.</p> <p>The e-NDP project aims at renewing our knowledge on <strong>Notre-Dame de Paris cathedral</strong> through the creation of a collaborative digital edition of the registers of its Chapter (1326-1504, <em>AN LL 105-128</em>), the community of 51 canons meeting three times a week on set days to take all administrative, financial and practical decisions pertaining to the cathedral, its estate and the society living in its cloister. This corpus has never been the object of a comprehensive study to understand the workings and history of this urban enclave and powerful community. The collaborative digital edition is based on a process of<strong> handwriting text recognition (HTR)</strong>, tested and supervised by scholars, researchers and engineers combining expertise in Medieval history, paleography, philology and digital humanities. The edition shall allow a better insight into the Chapter&rsquo;s administration, into its economical and political power within Paris, and the relationships it maintained with other institutions in the city.</p> <p>&nbsp;</p> <p><strong>Section 1 : The e-NDP ground-truth dataset for Handwriting text recognition.</strong></p> <p>The full e-NDP corpus kept today in the French National Archives and was entirely digitized and described in its&nbsp;<a href="https://www.siv.archives-nationales.culture.gouv.fr/siv/rechercheconsultation/consultation/ir/consultationIR.action?formCaller=GENERALISTE&amp;irId=FRAN_IR_059635">catalog</a>&nbsp;in 2022.</p> <p>The first major goal of the&nbsp;e-NDP projet is to propose a first automatic transcription of the 14k pages composing the 26 chapter registers. To achieve this goal representative samples from&nbsp;each one of the volumes were selected and transcribed in order to train a specialized HTR model able to propose a high quality automatic transcription. The collected ground-truth released on this repository currently has <strong>512 pages from the 26 registers</strong> of the cathedral chapter preserved in the National Archives (LL105 - LL128, <strong>1326-1504</strong>). The transcriptions were manually completed in <strong>two rounds</strong> by a group of 12 contributors, historians and paleographers, over the course of 2021-2022 using <a href="https://escriptorium.paris.inria.fr/">eScriptorium </a>as annotation environment.&nbsp;&nbsp;</p> <p>&nbsp;</p> <p><strong>Ground-truth features :</strong></p> <p><br> <em>Number of hands </em>: according to our estimates no fewer than 18&nbsp;main hands were involved in the writing of the registers during the medieval period.&nbsp;</p> <p><em>Language</em> : More than 98% of the content of the registers was written in Latin, the rest in French. The exact percentage is hard to estimate because the vernacular language is often used in formulae, notes and comments. It is rare to find entire pages or blocks written in French.&nbsp;</p> <p><em>Script family</em> : The registers were written using a Cursive script (ca. late XIIIe - XVIe).</p> <p><em>Documental typology</em> : The volumes containing the chapter conclusions were conceived to serve&nbsp;as memorial&nbsp;records, but above all as documents for regular use and consultation in the daily practice of administration and management. In diplomatics the notion of &quot;documentary manuscripts&quot; is used to describe this kind of sources&nbsp;also by opposition to books and litterary or&nbsp;normative&nbsp;manuscripts.</p> <table align="center"> <caption><strong>Ground truth statistics</strong></caption> <tbody> <tr> <th>Text units</th> <th>Count</th> </tr> <tr> <td>Pages</td> <td>512</td> </tr> <tr> <td>Annotated regions (see section 2)</td> <td>2448</td> </tr> <tr> <td>Lines of text</td> <td>34231</td> </tr> <tr> <td>Tokens</td> <td>205083</td> </tr> <tr> <td>Characters</td> <td>3320407</td> </tr> </tbody> </table> <p>&nbsp;</p> <p><strong>Rules of transcription :</strong></p> <ul> <li>The abbreviations have been resolved, both those by suspension (<code>facimꝰ</code> ---&gt; <code>facimus</code>) and by contraction (<code>d&ntilde;i</code> --&gt; <code>domini</code>). Likewise, those using conventional signs (<code>⁊</code> --&gt; <code>et</code> ; <code>ꝓ</code> --&gt; <code>pro</code>) have been resolved.&nbsp;</li> <li>The named entities (names of persons, places and institutions) have been <code>capitalized</code>. The beginning of a block of text as well as the original capitals used by the notary are also capitalized.</li> <li>The consonantal <code>i</code> and <code>u</code> characters have been transcribed as <code>j</code> and <code>v</code> in both French and Latin.</li> <li>The punctuation marks used in the text: <code>.</code> and <code>/</code> have been transcribed, but the transcription has not been standardized with modern punctuation.</li> <li>Corrections and words that appear cancelled in the manuscript have been transcribed surrounded by the sign <code>$</code> at the beginning and at the end.</li> <li>More specific transcription rules can be found into the file <code>transcription_guidelines.pdf</code></li> </ul> <p>&nbsp;</p> <p><strong>Section 2. e-NDP Layout Segmentation.</strong></p> <p>Layout segmentation is a compulsory step before HTR recognition in order to distinguish sections and regions inside a document. This process intend to separate interdependant page zones to produce a recognition in a section-sequence order and not in a line-sequence order which mix textual and peri-textual content.</p> <p>The regions of 364&nbsp;pages (see <code>GT-layout_list</code>) of the e-NDP corpus were annotated using a 5 sections vocabulary (see <code>endp_layout_regions</code>) in order to describe&nbsp;the page distribution in all the 26 volumes :</p> <ol> <li><em>Block</em>&nbsp;: All the central text blocks, that normally corresponds to the main content called &quot;conclusions&quot; in registers.</li> <li><em>Liste</em>&nbsp;: List of names of the canons who were present during the meeting. Normally located before the <em>conclusions</em>.</li> <li><em>Entr&eacute;e</em>&nbsp;: Marginal notes or entries to inform about the content of <em>conclusions</em>.</li> <li><em>Date</em>&nbsp;: Paragraph contending the date. Normally at the head of a <em>conclusion</em>, but separate of the main body.</li> <li><em>Num&eacute;rotation</em>&nbsp;: Page numbers in roman or arabic. Usually appear in the top corners of the pages.</li> </ol> <table align="center"> <caption><strong>Layout GT statistics</strong></caption> <tbody> <tr> <th>Region</th> <th>Count</th> </tr> <tr> <td>block</td> <td>833</td> </tr> <tr> <td>liste</td> <td>431</td> </tr> <tr> <td>date</td> <td>448</td> </tr> <tr> <td>entr&eacute;e</td> <td>205</td> </tr> <tr> <td>num&eacute;rotation</td> <td>531</td> </tr> </tbody> </table> <p>&nbsp;</p> <p><strong>Section 3. The e-NDP HTR modeling.</strong></p> <p>The e-NDP project has progressively trained several HTR models adapted to work on late medieval cursive in order to accelerate the production of ground truth. Currently the best model delivers an average&nbsp;<strong>CER (Character error ratio) of 9.7%</strong> in handwriting recognition on&nbsp;the 26 registers (see <code>endp_learning_curve</code>) and can serve as generalist model&nbsp;for other manuscripts of the same period and similar script family. These models and their training implementation details can be found in the project&#39;s github <a href="https://github.com/chartes/e-NDP_HTR">repository</a>.&nbsp;</p> <p>Additionally, the automatic HTR transcriptions of the 26 registers (14k pages, 4.5M tokens) enriched with lexical and semantical information has been the subject of a first <a href="https://nosketch-engine.lamop.fr/#dashboard?corpname=endp">online publication</a> using the NoSketch engine that allows advanced data mining based on the combination of data, metadata and NLP features.&nbsp;</p> <p>&nbsp;</p> <p><strong>Section 4. Dataset content.</strong></p> <p>This zip dataset contains :</p> <p>- <code>HTR_ground_truth</code> : Two folders containing the jpg / jpeg images and their curated transcriptions in PAGE XML format.</p> <p>- <code>images_docs</code> : 4 files illustrating the different phases of the project (list of GT for layout segmentation, layout ontologie, transcription guideline and HTR evaluation curves)</p>

opencc-by-4.0Feb 2023View details →
zenodo40/100

ICDAR 2015 Competition HTRtS: Handwritten Text Recognition on the tranScriptorium Dataset

<p>This dataset comprises the dataset used for the ICDAR 2015 Competition on  Handwritten Text Recognition on the tranScriptorium Dataset. The handwritten images for this contest were drawn from the English “Bentham collection” dataset used in the TRAN SCRIPTORIUM project. The selected data has been written by several hands and entails significant variabilities and difficulties regarding the quality of text images, writing styles and crossed-out text. This contest is clearly more difficult than the the first edition both for training and for testing. A portion of the training dataset and the full test dataset were provided in the form of carefully segmented line images, along with the corresponding transcripts. Another portion of the training dataset was provided as raw images and their corresponding transcripts at region level.<br>  </p> <p>ICDAR 2015 competition HTRtS: handwritten text recognition on the tranScriptorium dataset<br> JA Sánchez, AH Toselli, V Romero, E Vidal.  In International Conference on Document Analysis and Recognition (ICDAR), pp. 1166-1170, 2015.</p>

opencc-by-4.0Jan 2017View details →
zenodo40/100

Train-B dataset for ICDAR2017 Competition on Handwritten Text Recognition on the READ Dataset (ICDAR2017 HTR). Batch 1 and Batch 2.

<p>Train-B Dataset.   Dataset of pages without any layout or text line information. The corresponding transcripts are provided at page level with line breaks. It has 10k pages, though for convenience it is divided into two 5k page batches. This information is provided in PAGE format. </p> <p>This dataset is complementary to this other dataset:</p> <p>https://zenodo.org/record/439807#.WOIBZ3WLSkA</p> <p>More information at:</p> <p>https://scriptnet.iit.demokritos.gr/competitions/~icdar2017htr/</p> <p> </p>

opencc-by-4.0Apr 2017View details →
zenodo40/100

Train-A dataset for ICDAR2017 Competition on Handwritten Text Recognition on the READ Dataset (ICDAR2017 HTR)

<p>Train-A Dataset of pages with manually revised baselines and the corresponding transcripts associated to them. This batch is small, 50 pages. Please, keep in mind that only the baselines have been manually corrected, The polygons associated to each line have not been manually reviewed. </p> <p>This dataset is complementary to this other dataset:</p> <p>https://zenodo.org/record/439811#.WOIF9HWLSkA</p> <p>More information at:</p> <p>https://scriptnet.iit.demokritos.gr/competitions/~icdar2017htr/</p>

opencc-by-4.0Apr 2017View details →
zenodo40/100

NorHand v3 / Dataset for Handwritten Text Recognition in Norwegian

<p>This dataset comprises Norwegian letter and diary documents from 19th and early 20th century. It can be used to train Handwritten Text Recognition (HTR) models.</p>

opencc-by-4.0Dec 2023View details →
zenodo40/100

Oficio de Hipotecas de Girona. A dataset of Spanish notarial deeds (18th Century) for Handwritten Text Recognition and Layout Analysis of historical documents.

<p>This dataset is a subset of 596 documents from the&nbsp;<em>Registre d&#39;Hipoteques de Girona</em> of 1769 collection, guarded by the <a href="http://xac.gencat.cat/ca/llista_arxius_comarcals/girones/"><em>Arxiu Hist&ograve;ric de Girona</em></a>. This collection, is composed by hundreds of thousands of notarial deeds from the XVIII-XIX century (1768-1862). Sales, redemption of censuses, inheritance and matrimonial chapters are among the most common documentary&nbsp;typologies in the collection.</p> <p>This dataset is composed of more than 23700 text lines&nbsp;written by a single hand, covering more that 50 different topics (documentary typologies) and a vocabulary of more than 2400 different words. The documents are transcribed using the so-called&nbsp;diplomatic criteria. Additionally, transcripts were tagged with&nbsp;<br> extra enriching/complementary information (e.g. expansion of the&nbsp;abbreviations, hyphen marks, etc.). Along with the transcripts &nbsp;the layout of the document is detected and recorded. Pages have&nbsp;been labeled using six different layout regions.</p> <p>The images along with their respective ground-truth was compiled in PAGE compliant XML format<br> by the <a href="http://www2.udg.edu/tabid/11296/Default.aspx"><em>Centre de Recerca d&#39;Hist&ograve;ria Rural</em></a>&nbsp;and the HTR group of the <a href="https://www.prhlt.upv.es">Pattern Recognition and Human Language Technologies Research Center</a>.</p>

opencc-by-nc-4.0Jul 2018View details →
zenodo36/100

Test-B1 ICDAR2017 Competition on Handwritten Text Recognition on the READ Dataset (ICDAR2017 HTR)

<ul> <li><strong>Test-B1</strong>:  a batch of page images annotated with the geometry of regions where to detect text line and recognize.</li> </ul>

opencc-by-4.0Jun 2017View details →
zenodo36/100

Test-B2 ICDAR2017 Competition on Handwritten Text Recognition on the READ Dataset (ICDAR2017 HTR)

<p><strong>Test-B2</strong>:  a batch of page images annotated with the geometry of regions where to detect text line and recognize.</p>

opencc-by-4.0Jun 2017View details →
zenodo36/100

Dataset for ICDAR2017 Competition on Handwritten Text Recognition on the READ Dataset (ICDAR2017 HTR)

<p><strong>Train-A:</strong> Dataset of pages with manually revised baselines and the corresponding transcripts associated to them. This batch is small, 50 pages. Please, keep in mind that only the baselines have been manually corrected, The polygons associated to each line have not been manually reviewed.</p> <p><strong>Train-B:</strong> Dataset of pages without any layout or text line information. The corresponding transcripts are provided at page level with line breaks. It has 10k pages, though for convenience it is divided into two 5k page batches. This information is provided in PAGE format.</p> <p><strong>Test A:</strong> Dataset of pages with manually revised baselines. This batch has 65 pages. The polygons associated to each line have not been manually reviewed.</p> <p><strong>Test-B1:</strong> The same dataset of pages of the Test A, but annotated only with the geometry of regions. Text line information is not provided.                                                   </p> <p><strong>Test-B2:</strong> Dataset of page images annotated with the geometry of regions where to detect text line and recognize. It has 57 pages.</p> <p><strong>Baseline.tgz:</strong> Baseline system trained using the first 40 pages of Train-A. The system is based on the deep learning toolkit to transcribe handwritten text images called Laia.</p> <p>More information at:</p> <p>https://scriptnet.iit.demokritos.gr/competitions/~icdar2017htr/</p> <p> </p>

opencc-by-4.0Jul 2017View details →
zenodo36/100

Appendix A - The Implications of Handwritten Text Recognition for Accessing the Past at Scale

<p>List of works identified through a Grounded Theory Method (GTM) of the current and near future implications of Handwritten Text Recognition (HTR) on the historical method and wider information environment. The findings of this data collection are provided in 'The Implications of Handwritten Text Recognition for Accessing the Past at Scale'</p>

opencc-by-4.0Nov 2023View details →
zenodo36/100

Appendix A - Understanding the application of Handwritten Text Recognition technology in heritage contexts

<p>This appendix lists all the categorised works mentioning the HTR software Transkribus, used in&nbsp;&#39;Understanding the application of Handwritten Text Recognition technology in heritage contexts: a systematic review of Transkribus in published research&#39;</p>

opencc-by-4.0Jan 2022View details →
zenodo36/100

ICDAR 2015 Competition HTRtS: Handwritten Text Recognition on the tranScriptorium Dataset Rerelease

<p>A new release of the dataset used in the ICDAR 2015 HTR competition in which all Page XML files are based on the same 2013-07-15 schema. It only contains page level images, Page XML files for train and test (including the ground truth transcripts for the test and train batch 1) and plain text files for train batch 2 that have the page level ground truth transcripts. The original version of this dataset can be found at http://doi.org/10.5281/zenodo.248733<br> &nbsp;</p>

opencc-by-4.0Jan 2018View details →
zenodo36/100

Handwritten Text Recognition Ground Truth Set: StABS Ratsbücher O10, Urfehdenbuch X

<p>Ground Truth for &quot;Urfehdenbuch X der Stadt Basel (1563-1569)&quot; at Staatsarchiv Basel-Stadt (StABS).</p> <p>Images and text aligned, using text-to-image (provided within Transkribus, <a href="https://www.readcoop.eu">www.readcoop.eu</a>).</p> <p>ALTO and Page XML are available for the text alignment.</p> <p>TEI to txt on page basis by Peter D&auml;ngeli.</p> <p>Derived from Transcription/TEI file: Urfehdenbuch X der Stadt Basel (1563-1569), in: Die Urfehdeb&uuml;cher der Stadt Basel &ndash; digitale Edition, hg. v. Susanna Burghartz, Sonia Calvi und Georg Vogeler Basel/Graz 2016. (zuletzt ver&auml;ndert am 31.1.2017): <a href="http://hdl.handle.net/11471/1010.2.1">hdl:11471/1010.2.1</a>.</p> <p>Image source: http://dokumente.stabs.ch/view/2010/Ratsbuecher_O_10/</p> <p>CC-BY-NC-SA: The license is inherited from the project &quot;Urfehdeb&uuml;cher der Stadt Basel &ndash; digitale Edition&quot;: http://gams.uni-graz.at/o:ufbas.1563.</p>

opencc-by-nc-sa-4.0Aug 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record