Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
9
datasets available to search
ShareScore release 0.9.0
Dataset results
9 results for “Writer Identification”
1QIsaa data collection (binarized images, feature files, and plotting scripts) for writer identification test using artificial intelligence and image-based pattern recognition techniques
<p><strong>The Great Isaiah Scroll (1QIsa<sup>a</sup>) data set for writer identification</strong></p> <p>This data set is collected for the ERC project:<br> The Hands that Wrote the Bible: Digital Palaeography and Scribal Culture of the Dead Sea Scrolls<br> PI: Mladen Popović<br> Grant agreement ID: 640497</p> <p>Project website: <a href="https://cordis.europa.eu/project/id/640497">https://cordis.europa.eu/project/id/640497</a><br> <br> <strong>Copyright (c) </strong> University of Groningen, 2021. All rights reserved.<br> <strong>Disclaimer and copyright notice for all data contained on this .tar.gz file:</strong></p> <p><strong>1)</strong> permission is hereby granted to use the data for research purposes. It is not allowed to distribute this data for commercial purposes.</p> <p><strong>2) </strong>provider gives no express or implied warranty of any kind, and any implied warranties of merchantability and fitness for purpose are disclaimed.</p> <p><strong>3) </strong>provider shall not be liable for any direct, indirect, special, incidental, or consequential damages arising out of any use of this data.</p> <p><strong>4) </strong>the user should refer to the first public article on this data set:<br> <br> <em>Popović, M., Dhali, M. A., & Schomaker, L. (2020). Artificial intelligence-based writer identification generates new evidence for the unknown scribes of the Dead Sea Scrolls exemplified by the Great Isaiah Scroll (1QIsa<sup>a</sup>). arXiv preprint arXiv:2010.14476.</em><br> <br> BibTeX:</p> <pre>@article{popovic2020artificial, title={Artificial intelligence based writer identification generates new evidence for the unknown scribes of the Dead Sea Scrolls exemplified by the Great Isaiah Scroll (1QIsaa)}, author={Popovi{\'c}, Mladen and Dhali, Maruf A and Schomaker, Lambert}, journal={arXiv preprint arXiv:2010.14476}, year={2020} }</pre> <p><strong>5) </strong>the recipient should refrain from proliferating the data set to third parties external to his/her local research group. Please refer interested researchers to this site for obtaining their own copy.</p> <p><strong>Organisation of the data:</strong></p> <p>The .tar.gz file contains three directories: images, features, and plots. The included 'README' file contains all the instructions.</p> <p>The 'images' directory contains NetPBM images of the columns of 1QIsa<sup>a</sup>. The NetPBM format is chosen because of its simplicity. Additionally, there is no doubt about lossy compression in the processing chain. There are two images for each of the Great Isaiah Scroll columns: one is the direct binarized output from the BiNet (<em>arxiv.org/abs/1911.07930</em>) system, and the other one is the manually cleaned version of the binarized output. The file names for the direct binarized output are of the format '1QIsaa_col<columnnr>.pbm', for example, '1QIsaa_col15.pbm'. And, for the cleaned version, the format is '1QIsaa_col<columnnr>_cleaned.pbm', for example, '1QIsaa_col15_cleaned.pbm'. Note: the image files are not in a separate directory; they will be extracted in the same place. However, due to the unique naming, there is no problem extracting them in one single directory.</p> <p>The 'features' directory contains feature files computed for each of the column images. There are two types of feature files: Hinge and Adjoined. They are distinguishable by their extension, for example, '1QIsaa_col15_cleaned.hinge' and '1QIsaa_col15_cleaned.adjoined'. They are also arranged in separate directories for ease of use.</p> <p>The 'plots' directory contains a simple python script to perform PCA on the feature files and then visualize them in a 3D plot. The file takes the location of feature files as an input. The 'README_plot' file contains examples of how-to-run in the terminal.</p> <p><strong>Brief description:</strong><br> According to ImageMagick's' identify' tool, the original images are in grayscale (.jpg) from Brill collection, in '8-bit Gray 256c'. These images pass through multiple preprocessing measures to become suitable for pattern recognition-based techniques. The first step in preprocessing is the image-binarization technique. In order to prevent any classification of the text-column images based on irrelevant background patterns, a specific binarization technique (BiNet) was applied, keeping the original ink traces intact. After performing the binarization, the images were cleaned further by removing the adjacent columns that partially appear on the target columns' images. Finally, few minor affine transformations and stretching corrections were performed in a restrictive manner. These corrections are also targeted for aligning the texts where the text lines get twisted due to the leather writing surface's degradation. Hence, the clean images are there in the directory along with the direct binarized images. No effort has been made to obtain a balanced set in any way.</p> <p><strong>Tools:</strong><br> <strong>Binarization:</strong><br> The BiNet tool is available for scientific use upon request (m.a.dhal(at)rug.nl)</p> <p><strong>Image Morphing:</strong><br> In the original article, data augmentation was performed using image morphing. The tool is available on GitHub:<br> https://github.com/GrHound/imagemorph.c</p> <p><strong>Features for writer identification:</strong><br> Lambert Schomaker<br> http://www.ai.rug.nl/~lambert/allographic-fraglet-codebooks/allographic-fraglet-codebooks.html<br> http://www.ai.rug.nl/~lambert/hinge/hinge-transform.html<br> <em><strong>1. </strong>L. Schomaker & M. Bulacu (2004). Automatic writer identification using connected-component contours and edge-based features of upper-case Western script. IEEE Transactions on Pattern Analysis and Machine Intelligence, Vol 26(6), June 2004, pp. 787 - 798.<br> <strong>2. </strong>Bulacu, M. & Schomaker, L.R.B. (2007). Text-independent Writer Identification and Verification Using Textural and Allographic Features, IEEE Trans. on Pattern Analysis and Machine Intelligence (PAMI), Special Issue - Biometrics: Progress and Directions, April, 29(4), p. 701-717.</em><br> <br> The features (hinge, fraglets) have been combined in a single MS Windows application, GIWIS, which is available for scientific use upon request (l.r.b.schomaker(at)rug.nl)</p> <p><strong>If you have any question, please contact us:</strong><br> Maruf A. Dhali <m.a.dhali(at)rug.nl><br> Lambert Schomaker <l.r.b.schomaker(at)rug.nl><br> Mladen Popović <m.popovic(at)rug.nl></p> <p><strong>Please cite our papers if you use this data set:</strong><br> <em><strong>1.</strong> Popović, M., Dhali, M. A., & Schomaker, L. (2020). Artificial intelligence based writer identification generates new evidence for the unknown scribes of the Dead Sea Scrolls exemplified by the Great Isaiah Scroll (1QIsa<sup>a</sup>). arXiv preprint arXiv:2010.14476.<br> <strong>2. </strong>Dhali, M. A., de Wit, J. W., & Schomaker, L. (2019). Binet: Degraded-manuscript binarization in diverse document textures and layouts using deep encoder-decoder networks. arXiv preprint arXiv:1911.07930.</em></p>
Firemaker image collection for benchmarking forensic writer identification using image-based pattern recognition
<p>Disclaimer and terms of use:<br> ============================</p> <p>/*****************************************************************************\<br> * *<br> * *<br> * This is the Firemaker NFI-images Distribution *<br> * *<br> * This distribution contains 1000 images of scanned handwritten text, *<br> * scanned at resolution 300dpi grey scale, containing pages of *<br> * handwritten text by 250 writers, four pages per writer, from four *<br> * writing conditions, one condition per page. The conditions are: *<br> * p1: copied, natural style, p2: copied, UPPER case, p3: copied and forged, *<br> * i.e.,"try to write in a different style than your natural style", and p4, *<br> * self generated, i.e., text produced to describe a given cartoon. *<br> * *<br> * *<br> * *<br> * Copyright The International Unipen Foundation, 2000, All rights reserved *<br> *******************************************************************************<br> * *<br> * *<br> * DISCLAIMER AND COPYRIGHT NOTICE FOR ALL DATA CONTAINED ON THIS CDROM: *<br> * *<br> * *<br> * 1) PERMISSION IS HEREBY GRANTED TO USE THE DATA FOR RESEARCH *<br> * PURPOSES. IT IS NOT ALLOWED TO DISTRIBUTE THIS DATA FOR COMMERCIAL *<br> * PURPOSES. *<br> * *<br> * *<br> * 2) PROVIDER GIVES NO EXPRESS OR IMPLIED WARRANTY OF ANY KIND AND ANY *<br> * IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR PURPOSE ARE *<br> * DISCLAIMED. *<br> * *<br> * 3) PROVIDER SHALL NOT BE LIABLE FOR ANY DIRECT, INDIRECT, SPECIAL, *<br> * INCIDENTAL OR CONSEQUENTIAL DAMAGES ARISING OUT OF ANY USE OF THIS *<br> * DATA. *<br> * *<br> * 4) THE USER SHOULD REFER TO THE FIRST PUBLIC ARTICLE ON THIS DATA SET: *<br> * *<br> * M. Bulacu, L. Schomaker & L. Vuurpijl (2003). *<br> * Writer identification using edge-based directional features. *<br> * ICDAR '03: Proceedings of the 7th International Conference on Document *<br> * Analysis and Recognition, pp. 937-941. *<br> * Piscataway: IEEE Computer, ISBN 0-7695-1960-1 *<br> * *<br> * 5) THE RECIPIENT SHOULD REFRAIN FROM PROLIFERATING THE DATA SET TO THIRD *<br> * PARTIES EXTERNAL TO HIS/HER LOCAL RESEARCH GROUP. PLEASE REFER INTERESTED *<br> * RESEARCHERS TO HTTP://UNIPEN.ORG FOR OBTAINING THEIR OWN COPY. *<br> \*****************************************************************************/</p> <p>BibTeX entry: </p> <p> @inproceedings{Firemaker, <br> author = {Bulacu, M. and Schomaker, L.R.B. and Vuurpijl, L.}, <br> title = {Writer Identification Using Edge-Based Directional Features},<br> booktitle = {ICDAR '03: Proceedings of the 7th International <br> Conference on Document Analysis and Recognition},<br> year = {2003},<br> isbn = {0-7695-1960-1},<br> pages = {937-941},<br> publisher = {IEEE Computer Society},<br> address = {Washington, DC, USA},<br> }</p> <p>In the project "Vergelijk", a grant obtained from the Dutch Forensic Science<br> Institute, two existing professional writer-identification systems have been <br> compared regarding usability studies and in particular recognition <br> performance (Schomaker & Vuurpijl, 2000). The results of this comparison <br> are contained in a confidential report:</p> <p> L.R.B. Schomaker and L.G. Vuurpijl (2000). <br> Forensic writer identification: A benchmark data set <br> and a comparison of two systems. Technical report, <br> Nijmegen Institute for Cognition and Information (NICI), <br> University of Nijmegen, The Netherlands.</p> <p>Informative and non-confidential details from this report are <br> given in the accompanying file: 'firemaker-dbase.pdf'</p> <p>To compare both systems, a carefully designed experiment was conducted to<br> record handwritten samples from male and female writers in several conditions:</p> <p>Condition 1: Normal constrained handwriting<br> ==============================================</p> <p>Below, the Dutch text writers had to produce in normal handwriting is given. </p> <p>--- start text ----<br> Zij bezochten veilingen en reisden met de KLM. Voor<br> korte afstanden huurden ze een auto, meestal een VW<br> of een Ford.<br> <EMPTY LINE><br> De veilingen waren van 7-4-1993 tot 3-5-1993 in New<br> York, Tokyo, Québec, Rome, Parijs, Zürich en Oslo.<br> <EMPTY LINE><br> Omdat de veilingen steeds begonnen om 12 uur en je<br> gemiddeld 200 tot 300 kilometer moest rijden,<br> stonden zij steeds om 6.30 uur op en vertrokken om<br> 8 uur uit het hotel.<br> <EMPTY LINE><br> Elke dag hadden ze vijfhonderd (f 500,-) gulden<br> nodig. Daarvoor gebruikten ze elke keer een cheque<br> van tweehonderd (f 200,-) en een cheque van<br> driehonderd (f 300,-) gulden. Aan geschenken gaven<br> ze ongeveer honderd gulden (f 100,-) uit.<br> --- end text ----</p> <p><br> Condition 2: Production of constrained block capital handwriting<br> ================================================================</p> <p>In this condition, the writers had to produce the following text<br> in block-capital handwriting:</p> <p>--- start text ----<br> NADAT ZE IN NEW YORK, TOKYO, QUÉBEC, PARIJS, ZÜRICH<br> EN OSLO WAREN GEWEEST, VLOGEN ZE UIT DE USA TERUG<br> MET VLUCHT KL 658 OM 12 UUR.<br> <empty line><br> ZE KWAMEN AAN IN DUBLIN OM 7 UUR EN IN AMSTERDAM OM<br> 9.40 UUR 'S AVONDS. DE FIAT VAN BOB EN DE VW VAN<br> DAVID STONDEN IN R3 VAN HET PARKEERTERREIN.<br> HIERVOOR MOESTEN ZE HONDERD GULDEN (F 100,-)<br> BETALEN.<br> --- end text ----</p> <p><br> Condition 3: Production of free-forged handwriting<br> ==================================================</p> <p>Below, the text writers had to produce in the free-forged handwriting<br> condition is given. No example of handwriting is given which they have to<br> mimick (forge), the condition concerns a self-conceived distorted <br> handwriting style.</p> <p>--- start text ----<br> Nog dezelfde avond reden ze naar hun vrienden<br> Chris, Emile, Jan, Irene en Henk, nadat ze hun<br> vriendinnen Greta en Maria hadden opgehaald.<br> <EMPTY LINE><br> Samen hadden ze vijfhonderd (500) zeldzame<br> postzegels gekocht, Bob driehonderd (300) en David<br> tweehonderd (200).<br> <EMPTY LINE><br> De reis was de moeite waard geweest.<br> --- end text ----</p> <p><br> Condition 4: Production of unconstrained handwriting<br> ====================================================</p> <p>The final text writers had to produce is unconstrained handwriting.<br> The cartoon, a series of pictures concerning a 'UFO' landing had<br> to be described in their own words, in at least six lines of text.<br> See image file "space.gif".</p> <p><br> Thruth labels and writer identifications<br> ========================================</p> <p>Each writer has a unique id, specified as:</p> <p> id: {num}{set}<br> num: a three-digit number<br> set: either 01, 02, 03 or 04, identifying one of the 4 experiments</p> <p>The vast majority of the writers producing sets 01, 02 and 03 mimicked the<br> content and layout (empty lines) of the constrained texts they had to copy<br> sufficiently accurately, such that the example texts are a good indication of<br> the contents. However, as set 04 ("describe cartoon story") contains<br> unconstrained self-generated handwriting, the corresponding thruth labels had<br> to be extracted manually. The resulting label files are contained in the<br> directory ./300dpi/p4-self-natural/labels/</p> <p>Note: no letter, word, line or paragraph segmentation is provided with this<br> data set. The main text can be cropped easily. Since the orientation is<br> horizontal, projection techniques can be used to extract lines, using<br> a line-spacing parameter (~94 pixels line height) as an additional check. </p> <p><br> Overview of directories:</p> <p>300dpi/<br> p1-copy-normal/ Copying task, normal writing style <br> p2-copy-upper/ Copying task, UPPER-case <br> p3-copy-forged/ Copying task, instructed to mimic another script style<br> p4-self-natural/ Self-generated text, natural writing condition</p> <p>Note: the original raw collection contained writer #155, who has been removed<br> from this data set, as his first condition (p1) was started in upper case and<br> the page was not completed. Deleted files were 15501.tif, 15502.tif, 15503.tif<br> and 15504.tif.</p> <p>Note: the name of this data set (Firemaker) is a contraction of the names<br> Vuurpijl and Schomaker.</p> <p>Note b: Example of a cutout of essential handwritten text using NetPBM tools: <br> tifftopnm 15201.tif | pnmcut -left 50 -right 2400 -top 700 -bottom 3250 > handwriting.pgm</p> <p> For an experiment, the upper and lower halves of the resulting image were<br> usually used in the Schomaker & Bulacu studies to obtain two samples of <br> handwriting for a writer.</p> <p> http://www.ai.rug.nl/~lambert<br> http://www.ai.rug.nl/~bulacu</p> <p>Our features for writer identification:</p> <p>Lambert Schomaker<br> http://www.ai.rug.nl/~lambert/allographic-fraglet-codebooks/allographic-fraglet-codebooks.html<br> L. Schomaker & M. Bulacu (2004). <br> Automatic writer identification using connected-component contours and edge-based features of upper-case Western script. <br> IEEE Transactions on Pattern Analysis and Machine Intelligence, Vol 26(6), June 2004, pp. 787 - 798.</p> <p>Marius Bulacu<br> http://www.ai.rug.nl/~lambert/hinge/hinge-transform.html<br> Bulacu, M. & Schomaker, L.R.B. (2007). <br> Text-independent Writer Identification and Verification Using Textural and Allographic Features, <br> IEEE Trans. on Pattern Analysis and Machine Intelligence (PAMI), Special Issue - Biometrics: Progress and Directions, April, 29(4), p. 701-717.</p> <p>Axel Brink<br> http://www.ai.rug.nl/~axel/ 'Quill' feature<br> A.A. Brink, J. Smit, M.L. Bulacu, and L.R.B. Schomaker (2011). <br> Writer identification using directional ink-trace width measurements, <br> Pattern Recognition (July 2011), doi: 10.1016/j.patcog.2011.07.005<br> <br> These three feature groups (hinge, fraglets, quill) have been combined in<br> a single MS Windows application, GIWIS which is available for scientific<br> use upon request (schomaker@ai.rug.nl)</p> <p>Note c.</p> <p>The accompanying file 'Firemaker-writer-info.dat' contains some<br> writer information: <br> Column 1: writer identification code<br> Column 2: sex<br> Column 3: handedness, <br> Column 4: age in years<br> Column 5: major Western script group (print,cursive or mixed)<br> </p>
ImUnipen image data set for writer identification (N=208) - vectorial handwriting converted to usable images
<p><br> ==============<br> Terms of Usage<br> ==============</p> <p>The ImUnipen data set is intended for non-commercial, scientific use,<br> and is distributed under auspices of the Unipen Foundation.</p> <p>Please always refer to the following paper in IEEE PAMI when using<br> the ImUnipen data set:</p> <p> Bulacu, M.; Schomaker, L.<br> Text-Independent Writer Identification and Verification<br> Using Textural and Allographic Features<br> Pattern Analysis and Machine Intelligence, IEEE Transactions on<br> Volume 29, Issue 4, April 2007 Page(s):701 - 717</p> <p>The ImUnipen data set is derived from the Unipen (unipen.org)<br> data set of on-line (i.e., vectorial, xy) handwriting.<br> The xy-coordinates and a line-generator algorithm are used<br> to generate a raster image, as if the data were optically scanned.</p> <p>Contents: for 208 writers, there are two PNG images per writer of<br> an artificially constructed table of naturally written words (49MByte).<br> These words are pasted onto a white page. For systematics reasons,<br> we call such a page a Paragraph, see below.</p> <p>The file names are organized as (example):</p> <p> Writ990221.Doc01.Par00.png<br> Writ990221.Doc01.Par01.png</p> <p> meaning: writer number 990221, document 01 (there exists only Doc01)<br> and the image with artificial "paragraph" of isolated words "Par00"<br> and "Par01".</p> <p>The Par00 and Pa01 images are typically used as the query<br> and best match in a leave-one-out setting for writer identification.<br> For instance, Par00 is the query, and Par01 is added to the total set<br> of all other images as the attractor for an identification search.</p> <p>For these experiments, word labels are not given in this data set,<br> on purpose, as the goal is to test recognition-free writer identification<br> methods.</p> <p>For a description of the regular<br> Unipen data set, please visit http://unipen.org</p> <p>Lambert Schomaker constructed this set in 2005</p>
TriGraphSlant - benchmark set for writer identification - writers were asked to write in unnatural slant
<p> <br> Disclaimer and terms of use:<br> ============================<br> <br> /*****************************************************************************\<br> * *<br> * *<br> * This is the TrigraphSlant (Img version) Distribution, release 18/3/2011 * *<br> * *<br> * This distribution contains 188 images of scanned handwritten text, *<br> * scanned at resolution 300dpi Canon LiDE 25, grey scale, *<br> * by 47 Dutch writers, four pages per writer, from four *<br> * writing conditions, one condition per page. The conditions are: *<br> * 1. [AN] Copy text A in your natural handwriting. *<br> * 2. [BN] Copy text B in your natural handwriting. *<br> * 3. [BL] Copy text B and slant your handwriting to the *<br> * left as much as possible. *<br> * 4. [BR] Copy text B and slant your handwriting to the *<br> * right as much as possible. *<br> * The codes AN, BN, BL and BR refer to subsets into which the collected *<br> * pages of the writers were subdivided. AN represents a collection of *<br> * authentic documents; BN, BL and BR can be seen as collections of *<br> * questioned documents. To avoid structural effects of fatigue, the order *<br> * of item 3 and 4 was randomized at each collection: half of the subjects *<br> * wrote the BR page before the BL page. The data were collected at three *<br> * sites, in three cities: The Hague: NFI (N...), Donders Institute for *<br> * Brain, Cognition and Behaviour, Radboud University Nijmegen (D...) *<br> * and the Artificial Intelligence Dept. of University of Groningen (R...) *<br> * *<br> * Copyright The International Unipen Foundation, 2010, All rights reserved *<br> *******************************************************************************<br> * *<br> * *<br> * DISCLAIMER AND COPYRIGHT NOTICE FOR ALL DATA CONTAINED ON THIS CARRIER: *<br> * *<br> * *<br> * 1) PERMISSION IS HEREBY GRANTED TO USE THE DATA FOR RESEARCH *<br> * PURPOSES. IT IS NOT ALLOWED TO DISTRIBUTE THIS DATA FOR COMMERCIAL *<br> * PURPOSES. *<br> * *<br> * *<br> * 2) PROVIDER GIVES NO EXPRESS OR IMPLIED WARRANTY OF ANY KIND AND ANY *<br> * IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR PURPOSE ARE *<br> * DISCLAIMED. *<br> * *<br> * 3) PROVIDER SHALL NOT BE LIABLE FOR ANY DIRECT, INDIRECT, SPECIAL, *<br> * INCIDENTAL OR CONSEQUENTIAL DAMAGES ARISING OUT OF ANY USE OF THIS *<br> * DATA. *<br> * *<br> * 4) THE USER SHOULD REFER TO THE FOLLOWING ARTICLE ON THIS DATA SET: *<br> * *<br> * A.A. Brink, R.M.J. Niels, R.A. van Batenburg, C.E. van den Heuvel, *<br> * L.R.B. Schomaker, Towards robust writer verification by correcting *<br> * unnatural slant, Pattern Recognition Letters, Volume 32, Issue 3, *<br> * 1 February 2011, Pages 449-457, ISSN 0167-8655, *<br> * DOI: 10.1016/j.patrec.2010.10.010. *<br> * *<br> * 5) THE RECIPIENT SHOULD REFRAIN FROM PROLIFERATING THE DATA SET TO THIRD *<br> * PARTIES EXTERNAL TO HIS/HER LOCAL RESEARCH GROUP. PLEASE REFER INTERESTED *<br> * RESEARCHERS TO HTTP://UNIPEN.ORG FOR OBTAINING THEIR OWN COPY. *<br> \*****************************************************************************/<br> <br> Abstract<br> <br> Towards robust writer verification by correcting unnatural slant<br> <br> A.A. Brink, , R.M.J. Niels, R.A. van Batenburg, C.E. van den Heuvel, <br> and L.R.B. Schomaker, <br> <br> a Institute of Artificial Intelligence and Cognitive Engineering (ALICE), <br> University of Groningen, P.O. Box 407, 9700 AK Groningen, The Netherlands<br> <br> b Donders Institute for Brain, Cognition and Behaviour, Radboud University Nijmegen, <br> P.O. Box 9104, 6500 HE Nijmegen, The Netherlands<br> <br> c Netherlands Forensic Institute, P.O. Box 24044, 2490 AA Den Haag, The Netherlands<br> <br> Received 11 September 2009. Available online 30 October 2010.<br> <br> Slant is a salient feature of Western handwriting and it is considered to be an<br> important writer-specific feature. In disguised handwriting however, slant is<br> often modified. It was tested whether slant is indeed an important factor and it<br> was tested whether the distorting effect of deliberate slant change can be<br> countered by a simple shear transform. This was done in two off-line writer<br> verification experiments in image processing conditions of slant elimination and<br> slant correction. The experiments were performed using three features based on<br> statistical pattern recognition, including the state-of-the-art features<br> Fraglets and Hinge. A new public dataset was created and used, containing<br> natural and slanted handwriting by 47 writers. A striking result is that the<br> average natural slant value is much less important for biometric systems than is<br> usually assumed: eliminating slant yields just a 1-5% performance loss. A<br> second result is that the effects of deliberate slant change cannot be fully<br> countered by a simple shear transform: it raises performance on the distorted<br> handwriting from 53-68% to 64-90%, but this is still lower than normal<br> operation on natural handwriting: 97-100%.<br> <br> Research highlights<br> - The value of slant as a writer identification feature has been overrated. <br> - Deliberate slant change can be partly countered by the shear transform. <br> - Deliberate slant change introduces non-affine distortions to the handwriting. <br> - A new dataset of deliberately slanted handwriting was introduced.<br> <br> Keywords: Handwriting biometrics; Writer verification; Slant; Disguise; Statistical<br> pattern recognition</p> <p> </p>
ScriptNet: ICDAR2017 Competition on Historical Document Writer Identification (Historical-WI)
<p>This dataset contains the test set for the ICDAR2017 Competition on Historical Document Writer Identification (Historical-WI).</p> <p>The dataset used in this competition consists of 3600 handwritten pages originating from 13th to 20th century. It contains manuscripts from 720 different writers where each writer contributed five pages.</p> <p>Competition Website: https://scriptnet.iit.demokritos.gr/competitions/6/</p> <p>Changes August 1st, 2018: uploaded trainings set in color and binarized</p> <p>if you use the dataset please cite the ICDAR2017 Competition on Historical Document Writer Identification (Historical-WI) paper: https://doi.org/10.1109/ICDAR.2017.225</p> <p> </p>
READ ABP WI Dataset - Writer Identification over decades
<p>A hand is usually considered as a unique characteristic of a person. However, it may slightly change over their whole lifespan. This change might be due to some physical or mental issues. To the best of our knowledge, there is no dataset available, which covers this aspect of evolvement of handwriting of a single person.</p> <p>When dealing with archival documents, it is important to show that methods are invariant against these changes or investigate how much of these changes are covered. Thus, a new dataset was created with data of the Passau Diocesan Archives (ABP, <a href="https://www.bistum-passau.de/bistum/archiv">https://www.bistum-passau.de/bistum/archiv</a> ).</p> <p>The documents originate from death records of different villages or towns in the Diocese of Passau. Usually the writer of these records (mostly the priest) remains the same over several years. In total, the dataset consists of 1766 pages, which originate from 28 different writers. The number of pages per writer varies from 7 up to 311. For some writers, we only have data from 3 different years, whereas the largest time span between two documents of the same writer is 31 years.</p> <p>The dataset is organized as follows:</p> <p>[ID]_[Name]\[YEAR]\[ID]_filename.png</p> <p>The corresponding PAGE XML file is provided along with the dataset and contains the regions of the image where text is included. This file can be used to calculate features of the writer solely on the handwriting and not on the table lines.</p> <p>Currently no research tasks are defined on the dataset; we leave this up to the community. Drop us a note how you are using this dataset.</p>
Script Classification and Writer Identification: Two Tasks for a Common Understanding of Cultural Heritage - Dataset
<p>Writer identification and Script classification are usually considered as two separated and very different tasks, as well in palaeography as in computer science. Following the ICDAR competition on the CLAMM corpus about script classification and dating, this dataset provides the output created by running two infrastructures created for Script classification at a large scale (medieval scripts) on a more homogeneous dataset with a focus on Writer Identification.</p> <p>If you use the present repository, its data and figures, please consult and cite:</p> <p>Stutzmann, Dominique, Christopher Tensmeyer, and Vincent Christlein. « Writer Identification and Script Classification: Two Tasks for a Common Understanding of Cultural Heritage ». manuscript cultures, 15 (2020): 11-24. <a href="https://www.csmc.uni-hamburg.de/publications/mc/files/articles/mc15-02-stutzmann.pdf">https://www.csmc.uni-hamburg.de/publications/mc/files/articles/mc15-02-stutzmann.pdf</a></p>
CVL Database - An Off-line Database for Writer Retrieval, Writer Identification and Word Spotting
<p>The CVL Database is a public database for writer retrieval, writer identification and word spotting. The database consists of 7 different handwritten texts (1 German and 6 Englisch Texts). In total 310 writers participated in the dataset. 27 of which wrote 7 texts and 283 writers had to write 5 texts. For each text a rgb color image (300 dpi) comprising the handwritten text and the printed text sample is available as well as a cropped version (only handwritten). An unique id identifies the writer, whereas the Bounding Boxes for each single word are stored in an XML file.</p> <p>The CVL-database consists of images with cursively handwritten german and english texts which has been choosen from literary works. All pages have a unique writer id and the text number (separated by a dash) at the upper right corner, followed by the printed sample text. The text is placed between two horizontal separatores. Beneath the printed text individuals have been asked to write the text using a ruled undersheet to prevent curled text lines. The layout follows the style of the IAM database. The database was updated on 12/09/2013 since one writer ID (265/266) was wrong. The version number was changed to 1.1.</p> <p>Samples of the following texts have been used:</p> <ul> <li>Edwin A. Abbot – Flatland: A Romance of Many Dimension (92 words).</li> <li>William Shakespeare – Mac Beth (49 words).</li> <li>Wikipedia – Mailüfterl (73 words, under CC Attribution-ShareALike License).</li> <li>Charles Darwin – Origin of Species (52 words).</li> <li>Johann Wolfgang von Goethe – Faust. Eine Tragödie (50 words).</li> <li>Oscar Wilde – The Picture of Dorian Gray (66 words).</li> <li>Edgar Allan Poe – The Fall of the House of Usher (78 words).</li> </ul> <p>This database may be used for non-commercial research purpose only. If you publish material based on this database, we request you to include a reference to:</p> <p>Florian Kleber, Stefan Fiel, Markus Diem and Robert Sablatnig, <em>CVL-Database: An Off-line Database for Writer Retrieval, Writer Identification and Word Spotting</em>, In Proc. of the 12th Int. Conference on Document Analysis and Recognition (ICDAR) 2013, pp. 560-564, 2013.</p>
CERUG & Firemaker: datasets for writer identification
<p>This is the word based handwritten images for writer identification based on the following two papers:</p> <p>1, Sheng He and Lambert Schomaker. FragNet: Writer Identification using Deep Fragment Networks, IEEE Transactions on Information Forensics and Security, 2020.</p> <p>2, Sheng He and Lambert Schomaker. GR-RNN: Global-Context Residual Recurrent Neural Networks for Writer Identification. Pattern Recognition, 2021.</p> <p>Code is available:</p> <p>https://github.com/shengfly/writer-identification</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.