Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
29
datasets available to search
ShareScore release 0.9.0
Dataset results
29 results for “Document Analysis”
Nuremberg Letterbooks: A Multi-Transcriptional Dataset of Early 15th Century Manuscripts for Document Analysis
<p>This dataset contains the images and labels of the Nuremberg Letterbooks dataset.</p> <p>It consists of four books (books 2 - 5) with line-wise transcriptions. Three kinds of transcriptions are reported: basic, regularized, and diplomatic, with additional expanded abbreviations. </p> <p>Code templates for text verification and writer verification are available at:</p> <ul> <li><a href="https://github.com/M4rt1nM4yr/letterbooks_text_verification">https://github.com/M4rt1nM4yr/letterbooks_text_verification</a></li> <li><a href="https://github.com/M4rt1nM4yr/letterbooks_writer_verification">https://github.com/M4rt1nM4yr/letterbooks_writer_verification</a></li> </ul> <p>When using this dataset, please cite: <br>M. Mayr, J. Krenz, K. Neumeier, A. Bub, S. Bürcky, N. Brolich, K. Herbers, M. Habermann, P. Fleischmann, A. Maier, and V. Christlein<em>.</em> <br>Nuremberg Letterbooks: A Multi-Transcriptional Dataset of Early 15th Century Manuscripts for Document Analysis. <em>Sci Data</em> <strong>12</strong>, 811 (2025).<br><a href="https://doi.org/10.1038/s41597-025-05144-z">https://doi.org/10.1038/s41597-025-05144-z</a></p>
PSM-AP Comparative document analysis data: Priorities and challenges in the policy and digital strategies of ten PSM
<p>The document consists of a list of 61 key policy and strategy documents analysed as part of the comparative work conducted in WP1 of the project Public Service Media in the Age of Platforms (PSM-AP). It also contains a series of selected quotes supporting the three key areas prioritised by the policies and PSM digital strategies: People (reaching audiences), Personalisation (developing the video-on-demand portal), and Prominence (of PSM services and content). The data was collected and analysed in 2023, from documents concerning 10 PSM organisations in seven media markets: Belgium-Flanders (VRT), Belgium-Wallonia Brussels (RTBF), Canada (CBC/Radio-Canada), Denmark (DR, TV 2) Italy (RAI), Poland (TVP), and the UK (BBC, Channel 4, ITV). All quotes were translated to English by the authors.</p>
Oficio de Hipotecas de Girona. A dataset of Spanish notarial deeds (18th Century) for Handwritten Text Recognition and Layout Analysis of historical documents.
<p>This dataset is a subset of 596 documents from the <em>Registre d'Hipoteques de Girona</em> of 1769 collection, guarded by the <a href="http://xac.gencat.cat/ca/llista_arxius_comarcals/girones/"><em>Arxiu Històric de Girona</em></a>. This collection, is composed by hundreds of thousands of notarial deeds from the XVIII-XIX century (1768-1862). Sales, redemption of censuses, inheritance and matrimonial chapters are among the most common documentary typologies in the collection.</p> <p>This dataset is composed of more than 23700 text lines written by a single hand, covering more that 50 different topics (documentary typologies) and a vocabulary of more than 2400 different words. The documents are transcribed using the so-called diplomatic criteria. Additionally, transcripts were tagged with <br> extra enriching/complementary information (e.g. expansion of the abbreviations, hyphen marks, etc.). Along with the transcripts the layout of the document is detected and recorded. Pages have been labeled using six different layout regions.</p> <p>The images along with their respective ground-truth was compiled in PAGE compliant XML format<br> by the <a href="http://www2.udg.edu/tabid/11296/Default.aspx"><em>Centre de Recerca d'Història Rural</em></a> and the HTR group of the <a href="https://www.prhlt.upv.es">Pattern Recognition and Human Language Technologies Research Center</a>.</p>
Vorau Abbey library Cod. 253 dataset for Document Layout Analysis
<p>VORAU-253 is a music manuscript referred to as Cod. 253 of the Vorau Abbey library, which was provided by the Austrian Academy of Sciences. It is written in German Gothic notation and dated around year 1450.</p> <p>This manuscript is interesting because of the complexity of its layout, where staff, text and decorations are intertwined to<br> compose the structure of the document.</p> <p>This database is a subset of 228 pages of the archive, using 128 randomly selected pages for training/validation and 100 for test.</p> <p>The database was manually annotated into the following three layout regions:</p> <p>* staff: represents the regions that contains a set of horizontal lines and spaces where each one represent a different musical pitch. This region type does not contain text lines. Hence, no baselines.</p> <p>* lyrics: are the words that are sung appear below their corresponding staff, and other text in the document. In all cases, text to be sung and the other text are assigned to different layout regions under the lyrics label.</p> <p>* drop-capital: is a decorated letter that might appear at the beginning of a word or text line. As it is a single big letter, it contain no text lines nor baselines.</p> <p>On average each page contains 12.5 [7,23] text lines distributed over an average of 10.5[7,15] ``lyrics'' regions. Moreover, each page contains 22.3[14,28] layout regions on average.</p>
Spectra belonging to XRF instrument report. Identification of ink components through XRF analysis of Azzolino documents.
<p>Accumulated spectra. Details given in supporting information </p> <p><a href="https://journals.plos.org/plosone/article/file?type=supplementary&id=10.1371/journal.pone.0283539.s002">S2 File. </a>XRF instrument report.</p> <p>Identification of ink components through XRF analysis of Azzolino documents.</p> <p><a href="https://doi.org/10.1371/journal.pone.0283539.s002">https://doi.org/10.1371/journal.pone.0283539.s002</a></p> <p>(DOCX)</p> <p>Belonging to publication </p> <p>Lagerqvist Alidoost A, Hacke M, Winther T, Sandström T (2023) A closer look at the Azzolino collection. PLOS ONE 18(4): e0283539. <a href="https://doi.org/10.1371/journal.pone.0283539">https://doi.org/10.1371/journal.pone.0283539</a></p>
XRF maps and line scans project files belonging to XRF instrument report. Identification of ink components through XRF analysis of Azzolino documents.
<p>RTX Project files. Details given in supporting information </p> <p><a href="https://journals.plos.org/plosone/article/file?type=supplementary&id=10.1371/journal.pone.0283539.s002">S2 File. </a>XRF instrument report.</p> <p>Identification of ink components through XRF analysis of Azzolino documents.</p> <p><a href="https://doi.org/10.1371/journal.pone.0283539.s002">https://doi.org/10.1371/journal.pone.0283539.s002</a></p> <p>(DOCX)</p> <p>Belonging to publication </p> <p>Lagerqvist Alidoost A, Hacke M, Winther T, Sandström T (2023) A closer look at the Azzolino collection. PLOS ONE 18(4): e0283539. <a href="https://doi.org/10.1371/journal.pone.0283539">https://doi.org/10.1371/journal.pone.0283539</a></p>
Forschungsprojekt: Digitalisierung, Klassifikationen und Gesundheits-Apps. Dataset B - Document analysis for health app users and developers
<p><strong>Allgemeine Hinweise Data Set B</strong></p> <ol> <li>Titel des Forschungsprojekts</li> </ol> <p>Digitale Gesundheitsklassifikationen in Apps - Praktiken und Probleme ihrer Entwicklung und situativen Anwendung. Projektleitung: Prof. Dr. Rainer Diaz-Bone. Bearbeitung: Valeska Cappel Dipl. Soz., Miriam Kutt (Hilfsassistenz). Laufzeit: 2019-2023. Finanzierung: Schweizer Nationalfonds.</p> <p>2. PrimärforscherInnen:</p> <ul> <li>Rainer Diaz-Bone</li> <li>Valeska Cappel</li> <li>Miriam Kutt</li> </ul> <p>3. Publikationsjahr:</p> <ul> <li>2023</li> </ul> <p>4. Hinweise zur Verfügbarkeit</p> <p>Die Daten werden über LORY (Lucerne Open Repository) dauerhaft zugänglich gemacht.</p> <ul> <li>Alle Dateien die sich auf die App-Entwicklung beziehen beginnen mit E</li> <li>Alle Dateien die sich auf die App-Nutzung beziehen beginnen mit N</li> </ul> <p>5. Fachgebiet</p> <ul> <li>Soziologie</li> </ul> <p>6. Kategorie und Schlagwörter</p> <ul> <li>Gesundheitswesen</li> <li>Selbstvermessung</li> <li>Gesundheits-Apps</li> <li>Klassifikationen</li> <li>Pragmatismus</li> <li>Economics of convention</li> <li>Soziologie der Konventionen</li> <li>Digitalisierung</li> </ul> <p>7. Abstract, wozu die Daten erhoben wurden</p> <p>Die Daten wurden für eine Dokumentenanalyse erhoben. Ausgewählt wurden (1) Daten von forschungsrelevanten Akteuren oder Institutionen, die nicht für ein Interview gewonnen werden konnten und (2) Medienberichte, Anleitungen, Beschreibungen von präventiven Gesundheits-Apps, Wissenschaftliche Berichte (White Papers, Ausschreibungen) sowie Bewertungen von diesen Gesundheits-Apps aus dem „App-Store“ und „Google-Play-Store“ (Plattformen zum Download von Apps). Ziel war es mit diesen Dokumenten weitere Analysen durchzuführen, die diskursanlytische Aussagen über die Entstehung und Nutzung von präventiven Gesundheits-Apps sowie Entwicklungen im Feld der digitalen Gesundheit zulassen. Gleichzeitig wurden auch Dokumente, wie Nutzer-Bewertungen erhoben, um die Interviews zu ergänzen aus einer pragmatischen Perspektive Aushandlungs- und Problemlösungsprozesse im Umgang mit Gesundheits-Apps und der damit verbundenen Technologie zu untersuchen.</p> <p>8. Untersuchungsgebiet</p> <p>Das Untersuchungsgebiet liegt im Bereich der digitalen Gesundheit und beschränkt sich im Speziellen auf die Prozesse der Entwicklung von Gesundheits-Apps, sowie die Nutzung der Gesundheits-Apps. Untersuchungsgebiet waren öffentlich zugängliche Medien sowie die Gesundheits-App selbst. Dabei wurden Unternehmen, Zeitschriften, Blogs und Apps, die sich konkret mit der App-Entwicklung, der App-Nutzung oder der Berichterstattung über präventive Gesundheits-Apps beschäftigen ausgewählt zur Dokumentenerhebung ausgewählt. Bei der Auswahl wurden der Fokus darauf gelegt Inhalte auszuwählen, die sich nicht auf Gesundheits-App als Medizinprodukte konzentrieren, sondern auf präventive Gesundheits-Apps, die kein spezifisches Krankheitsbild adressieren, sondern einen allgemeinen positiven Gesundheitszustand herstellen oder erhalten sollen. </p> <p>9. Gesamtheit auf die generalisiert werden könnte („Grundgesamtheit“)</p> <p>Insgesamt wurden ca. 300 Dokumente erhoben.</p> <p>10. Auswahlverfahren und Stichproben</p> <p>Die Daten im Projekt wurden anhand qualitativer Methoden gewonnen. Die Fälle und Daten wurden über die Methode der „Theoretical-Sampling-Technik“ ausgewählt. Die genauen Begründungen zur Auswahl der Fälle wurden im Verlauf des Projektes theoretisch erarbeitet. Dabei wurde sich dem Forschungsfeld der präventiven Gesundheits-Apps mit heuristischen Vermutungen angenähert, die anhand der Konzepte der Theorie der Konventionen und einer machttheoretischen Perspektive Foucaults entwickelt wurden. Gesundheit und Gesundheitshandlungen wurden dabei aus einer pragmatischen Perspektive als ein Ergebnis von Koordinationsbemühungen zwischen Akteuren, Gegenständen, Technologien und Machtstrukturen verstanden. Die Auswahl der Dokumente stützte sich besonders auf die Memos und Inhalte der vorher geführten Interviews und den daraus entwickelten heuristische Fragestellungen während des Forschungsprozesses.</p> <p>11. Erhebungszeitraum</p> <p>Die Dokumente wurden in dem Zeitraum 2019-2023 erhoben. Das Datum in den Dateinamen bezieht sich immer auf den Erhebungszeitpunkt</p> <p>Sprache</p> <ul> <li>Deutsch und Englisch</li> </ul> <p>12. Größe des Datensatzes</p> <ul> <li>ca. 45 MB</li> </ul> <p>13. Verwendete Dateiformate und notwendige Software</p> <ul> <li>Dateiformat: PDF (Portable Document Format) und RTF (Rich Text Format)</li> </ul> <p>Software:</p> <ul> <li>RTF Standard-Textprogrammen auf unterschiedlichen Betriebssystemen (Bspw. Word, Wordpad, LibreOffice, OpenOffice)</li> <li>PDF: PDF-Programme/Reader oder auch ATLAS</li> </ul> <p> </p>
Documents used in the PLANET4B analysis of biodiversity discourse by environmental NGOs
<p>These files include press releases that have been published on the internet by European environmental NGOs, and which were used in the PLANET4B project analysis of the discourse on biodiversity.</p>
Documents used in the PLANET4B analysis of biodiversity discourse by political parties
<p>These files include press releases that have been published on the internet by European political parties, and which were used in the PLANET4B project analysis of the discourse on biodiversity.</p>
Documents used in the PLANET4B D1.1 analysis of biodiversity discourse by news outlets - 2010 and 2022 Data
<p>Data used to analyse biodiversity discourse in news outlet as part of Deliverable D1.1. of the Planet4B Project.</p>
Dataset for an analysis of selected studies on software development and reuse in the field of software documentation
<p>The data record contains bibliographical metadata from 2012 to 2023, which have been extracted from the databases ScienceDirect, SpringerLink and IEEE Xplore. It was created in the context of a master thesis which will be published later. The aim was to make a selection of studies on software documentation from the perspective of software development based on the master thesis.</p>
Circular economy standards for automotive electronics: from a state of the art analysis to a new pre-standardization document
<p>The Standardization Toolkit, developed by <a href="https://www.uni.com/en/" target="_blank" rel="noopener">UNI (Italian Standards Body)</a> within the TREASURE project, is an open reserach tool designed to simplify reserach of international and european standards (EN and ISO) relevant to circular economy strategies in the automotive sector.</p> <p>This dynamic tool, powered by Microsoft Power BI, provides an extensive list of national, European, and international standards, detailing key information such as standard numbers, URLs for document retrieval, titles, publication years, and scopes.</p> <p>The Toolkit maps out 73 standards, including 61 current standards and 12 works in progress, overseen by prominent technical committees such as ISO and CEN. It covers crucial areas like life-cycle assessment, substance determination, and decision support frameworks, offering a valuable resource for stakeholders aiming to implement effective and sustainable practices.</p> <p>The Toolkit is freely available at <a href="https://www.treasureproject.eu/standardization-toolkit/">this link</a>.</p> <p>This publication includes the dateset behind the Standardization Toolkit, while D8.4 "Standardization Toolkit" describes the previous version of the mapping and the methodology behind the state of the art analysis.</p> <p>It has also been uploaded the D8.5 "Strategic standardization roadmap" which includes a description of the Standardization Toolkit, and the new pre-standardization document (CEN Workshop Agreement - CWA) developed within Treasure project. </p> <p> </p> <p><strong>Disclaimer:</strong> The standards retrieved from the Standardization Toolkit cannot be fully downloaded. The Toolkit serves to simplify the search for these standards and to provide an overview of what has been mapped during the TREASURE project’s standardization activities.</p>
Document Layout Analysis - David Hume's History of England
<p>A fine-grained text region dataset for document layout analysis, freely available for research, featuring over 2400 annotated pages from four editions of David Hume’s History of England.</p>
IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection (part 5)
<p>This is part 5 of the IDNet dataset of our research paper "IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection". Here's a link to the paper: https://arxiv.org/pdf/2408.01690</p> <p>Citation:</p> <p>@article{guan2024idnet,<br> title={IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection},<br> author={Guan, Hong and Wang, Yancheng and Xie, Lulu and Nag, Soham and Goel, Rajeev and Swamy, Niranjan Erappa Narayana and Yang, Yingzhen and Xiao, Chaowei and Prisby, Jonathan and Maciejewski, Ross and Zou, Jia},<br> journal={arXiv preprint arXiv:2408.01690},<br> year={2024}<br>}</p>
IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection (part 1)
<p>This is part 1 of the IDNet dataset of our research paper "IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection". Here's a link to the paper: https://arxiv.org/pdf/2408.01690</p> <p>Citation:</p> <p>@article{guan2024idnet,<br> title={IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection},<br> author={Guan, Hong and Wang, Yancheng and Xie, Lulu and Nag, Soham and Goel, Rajeev and Swamy, Niranjan Erappa Narayana and Yang, Yingzhen and Xiao, Chaowei and Prisby, Jonathan and Maciejewski, Ross and Zou, Jia},<br> journal={arXiv preprint arXiv:2408.01690},<br> year={2024}<br>}</p>
IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection (part 4)
<p>This is part 4 of the IDNet dataset of our research paper "IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection". Here's a link to the paper: https://arxiv.org/pdf/2408.01690</p> <p>Citation:</p> <p>@article{guan2024idnet,<br> title={IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection},<br> author={Guan, Hong and Wang, Yancheng and Xie, Lulu and Nag, Soham and Goel, Rajeev and Swamy, Niranjan Erappa Narayana and Yang, Yingzhen and Xiao, Chaowei and Prisby, Jonathan and Maciejewski, Ross and Zou, Jia},<br> journal={arXiv preprint arXiv:2408.01690},<br> year={2024}<br>}</p>
IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection (part 3)
<p>This is part 3 of the IDNet dataset of our research paper "IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection". Here's a link to the paper: https://arxiv.org/pdf/2408.01690</p> <p>Citation:</p> <p>@article{guan2024idnet,<br> title={IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection},<br> author={Guan, Hong and Wang, Yancheng and Xie, Lulu and Nag, Soham and Goel, Rajeev and Swamy, Niranjan Erappa Narayana and Yang, Yingzhen and Xiao, Chaowei and Prisby, Jonathan and Maciejewski, Ross and Zou, Jia},<br> journal={arXiv preprint arXiv:2408.01690},<br> year={2024}<br>}</p>
IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection (part 2)
<p>This is part 2 of the IDNet dataset of our research paper "IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection". Here's a link to the paper: https://arxiv.org/pdf/2408.01690</p> <p>Citation:</p> <p>@article{guan2024idnet,<br> title={IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection},<br> author={Guan, Hong and Wang, Yancheng and Xie, Lulu and Nag, Soham and Goel, Rajeev and Swamy, Niranjan Erappa Narayana and Yang, Yingzhen and Xiao, Chaowei and Prisby, Jonathan and Maciejewski, Ross and Zou, Jia},<br> journal={arXiv preprint arXiv:2408.01690},<br> year={2024}<br>}</p>
IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection (part 7)
<p>This is part 7 of the IDNet dataset of our research paper "IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection". Here's a link to the paper: https://arxiv.org/pdf/2408.01690</p> <p>Citation:</p> <p>@article{guan2024idnet,<br> title={IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection},<br> author={Guan, Hong and Wang, Yancheng and Xie, Lulu and Nag, Soham and Goel, Rajeev and Swamy, Niranjan Erappa Narayana and Yang, Yingzhen and Xiao, Chaowei and Prisby, Jonathan and Maciejewski, Ross and Zou, Jia},<br> journal={arXiv preprint arXiv:2408.01690},<br> year={2024}<br>}</p>
IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection (part 6)
<p>This is part 6 of the IDNet dataset of our research paper "IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection". Here's a link to the paper: https://arxiv.org/pdf/2408.01690</p> <p>Citation:</p> <p>@article{guan2024idnet,<br> title={IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection},<br> author={Guan, Hong and Wang, Yancheng and Xie, Lulu and Nag, Soham and Goel, Rajeev and Swamy, Niranjan Erappa Narayana and Yang, Yingzhen and Xiao, Chaowei and Prisby, Jonathan and Maciejewski, Ross and Zou, Jia},<br> journal={arXiv preprint arXiv:2408.01690},<br> year={2024}<br>}</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.