Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,085
datasets available to search
ShareScore release 0.9.0
Dataset results
1,085 results for “documentation”
Demo of the alpha version of MEMORISE Document Manager
<p>Video demonstration showing the basic workflow of the MEMORISE Document Manager</p>
Nuremberg Letterbooks: A Multi-Transcriptional Dataset of Early 15th Century Manuscripts for Document Analysis
<p>This dataset contains the images and labels of the Nuremberg Letterbooks dataset.</p> <p>It consists of four books (books 2 - 5) with line-wise transcriptions. Three kinds of transcriptions are reported: basic, regularized, and diplomatic, with additional expanded abbreviations. </p> <p>Code templates for text verification and writer verification are available at:</p> <ul> <li><a href="https://github.com/M4rt1nM4yr/letterbooks_text_verification">https://github.com/M4rt1nM4yr/letterbooks_text_verification</a></li> <li><a href="https://github.com/M4rt1nM4yr/letterbooks_writer_verification">https://github.com/M4rt1nM4yr/letterbooks_writer_verification</a></li> </ul> <p>When using this dataset, please cite: <br>M. Mayr, J. Krenz, K. Neumeier, A. Bub, S. Bürcky, N. Brolich, K. Herbers, M. Habermann, P. Fleischmann, A. Maier, and V. Christlein<em>.</em> <br>Nuremberg Letterbooks: A Multi-Transcriptional Dataset of Early 15th Century Manuscripts for Document Analysis. <em>Sci Data</em> <strong>12</strong>, 811 (2025).<br><a href="https://doi.org/10.1038/s41597-025-05144-z">https://doi.org/10.1038/s41597-025-05144-z</a></p>
Artifacts supplementing the EuroUSEC '22 paper "Assessing Real-World Applicability of Redesigned Developer Documentation for Certificate Validation Errors"
<p>This upload supplements the conference EuroUSEC 2022 submission by providing the full questionnaire, anonymized dataset and all performed analyses presented in the paper specified below.</p> <ul> <li>Title: <strong>Assessing Real-World Applicability of Redesigned Developer Documentation for Certificate Validation Errors</strong></li> <li>Authors: Martin Ukrop, Michaela Balážová, Pavol Žáčik, Eric Vincent Valčík, Vashek Matyas</li> <li>Paper details: https://crocs.fi.muni.cz/public/papers/eurousec2022</li> <li>Paper abstract: <em>We face certificate validation errors commonly, yet the related tools and documentation had been shown to have very poor usability. Previous research suggests that just improving the error messages and corresponding documentation can have significantly positive effects. Our work aims at increasing the usability of certificate validation by 1) redesigning the API error messages and the corresponding documentation, and 2) validating the real-world applicability of the redesign by investigating the opinions of 180 IT professionals. We focus on the perceived obstacles, desired ideal form and overall satisfaction. The redesigned documentation exhibits a reliable significant decrease in perceived incompleteness, with a small amount of perceived bloat and tangle. The redesigned documentation, now published on a dedicated website, is preferred by 89% of our study participants.</em></li> </ul> <p>The artifacts accompanying this paper contain three major parts:</p> <ul> <li>The questionnaire used in the main study (described in Sections 3.1 and 3.2 of the paper and mostly present in Appendices A and B of the paper).</li> <li>The anonymized dataset (multiple formats) of all valid questionnaire answers and qualitative coding performed. Analyses of this dataset are the core of the paper and are present in subsection 3.4 and all parts of Sections 4 and 5.</li> <li>The set of analyses files (IBM SPSS scripts and outputs) producing all statistical results presented in the paper are included.</li> </ul> <p>More details about the artifacts can be found in the README file in the artifacts archive.</p>
The arrival and spread of the European firebug Pyrrhocoris apterus in Australia as documented by citizen scientists
<p>Data and R script to reproduce analyses conducted in <strong>The arrival and spread of the European firebug <em>Pyrrhocoris apterus</em> in Australia as documented by citizen scientists</strong></p> <p><strong>Abstract</strong></p> <p>We present evidence of the recent introduction and quick spread of the European firebug <em>Pyrrhocoris apterus</em> in Australia, as documented on the citizen science platform iNaturalist. The first public record of the species was reported in December 2018 in the City of Brimbank (Melbourne, Victoria). Since then, the species distribution has quickly expanded into 15 local government areas surrounding this first observation, including areas in both Metropolitan Melbourne and regional Victoria. The number of records of the European firebug in Victoria has also seen a substantial increase, with a current tally of almost 100 observations in iNaturalist as of July 31<sup>st</sup>, 2021.</p> <p>The case of the European firebug in Australia adds to the list of examples of citizen scientists playing a key role in not only early detection of newly introduced species but in documenting their expansion across their non-native range. Citizen science presents an exciting opportunity to complement biosecurity efforts carried out by government agencies, which often lack resources to sufficiently fund detection and monitoring programs given the overwhelming number of current and potential invasive species. Recognising and supporting the invaluable contribution of citizen scientists to science and society can help reduce this gap by: (1) increasing the number of introduced species that are quickly detected; (2) gathering evidence of the species’ early expansion stage; and (3) prompting adequate monitoring and rapid management plans for potentially harmful species.</p> <p>Given the range expansion patterns of the European firebug worldwide, their adaptation ability, and future climate scenarios, we suspect this species will continue expanding beyond Victoria, including other parts of Australia, New Zealand, and the South Pacific. We firmly believe that most of the knowledge about how this expansion process continues to happen will be provided by citizen scientists.</p>
Document-to-document relevant assessment for TREC Genomics Track 2005
<p>Here we present a table with document-to-document relevance assessment judgements on a subset of the TREC Genomics Track 2005 which corresponds to document-to-topic relevance assessments. This data was produced by four annotators to make it possible to analyze inter-annotator agreements as part of our future work. The data was produced with and in-house annotation tool tailored to the initial TREC data and the task at hand. The "raw data document evaluation" contains six columns, first row consecutive id, second original TREC topic, third PubMed Id used as reference document, fourth PMID used to evaluate the relevance wrt the reference document, fifth the relevance score (2 definitely relevant, 1 partially relevant, 0 non-relevant), and sixth annotator id.</p> <p> </p> <p><strong>Acknowledgements</strong></p> <p>This work is part of the STELLA project funded by DFG (project no. 407518790). This work was supported by the BMBF-funded de.NBI Cloud within the German Network for Bioinformatics Infrastructure (de.NBI) (031A532B, 031A533A, 031A533B, 031A534A, 031A535A, 031A537A, 031A537B, 031A537C, 031A537D, 031A538A).</p>
Document Liveness Challenge (DLC-2021) - part 1 (or, cg)
<p>Dataset DLC-2021 consists of 1424 video clips captured in a wide range of real-world conditions and focused on ID document forensics tasks. Each clip was shot vertically and was at least 5 seconds long. Frames extracted at 10 frames per second and for the 50 first extracted frames document position is manually annotated.<br> The novelty of the dataset is that it contains shots from video with color laminated mock ID documents, color unlaminated copies, grayscale unlaminated copies, and screen recaptures of the documents. The proposed dataset complies with the GDPR because it contains images of synthetic IDs with generated owner photos and artificial personal information.</p> <p>Part 1 contains videos, frames and markup for “original” laminated documents from MIDV-2020 collection and unlaminated gray copies. <br> Part 2 contains videos, frames and markup for documents recaptured from device screen<br> Part 3 contains videos, frames and markup for unlaminated color copies.</p> <p><strong>Share and Cite</strong></p> <p><em>MDPI and ACS Style</em></p> <p>Polevoy, D.V.; Sigareva, I.V.; Ershova, D.M.; Arlazarov, V.V.; Nikolaev, D.P.; Ming, Z.; Luqman, M.M.; Burie, J.-C. Document Liveness Challenge Dataset (DLC-2021). <em>J. Imaging</em> <strong>2022</strong>, <em>8</em>, 181. https://doi.org/10.3390/jimaging8070181</p> <p><em>AMA Style</em></p> <p>Polevoy DV, Sigareva IV, Ershova DM, Arlazarov VV, Nikolaev DP, Ming Z, Luqman MM, Burie J-C. Document Liveness Challenge Dataset (DLC-2021). <em>Journal of Imaging</em>. 2022; 8(7):181. https://doi.org/10.3390/jimaging8070181</p> <p><em>Chicago/Turabian Style</em></p> <p>Polevoy, Dmitry V., Irina V. Sigareva, Daria M. Ershova, Vladimir V. Arlazarov, Dmitry P. Nikolaev, Zuheng Ming, Muhammad M. Luqman, and Jean-Christophe Burie. 2022. "Document Liveness Challenge Dataset (DLC-2021)" <em>Journal of Imaging</em> 8, no. 7: 181. https://doi.org/10.3390/jimaging8070181</p>
The TDI data and PSD/sensitivity-related files for PyCBC LISA documentation example
<p>The TDI data and PSD/sensitivity-related files for PyCBC LISA documentation example, most of them are generated from LDC-Sangria<em> </em>dataset.</p>
Artifacts for Keyword Extraction From Specification Documents for Planning Security Mechanisms
<p>This dataset contains the data used for evaluating VDocScan - a keyword extraction based security vulnerability prediction method. The repository includes an extensive list of Products and Vulnerability reports from CVE, a custom created dataset mapping vulnerability reports to product documentations, as well as, intermediate results from the study such as decision trees rendered for each vulnerability, correlation matrix of vulnerabilities etc. </p>
DOCUMENTATION OF RISIS DATASETS Doctoral Degree and Career Dataset (DDC)
<p>Documentation is presented for The Doctoral Degree and Career Dataset (DDC). In the framework of the RISIS2 project, DDC is an experimental dissertation-centric database. It primarily consists of an enriched PhD publication dataset which brings togehter information about the dissertation (e.g., topic mapping), about the degree-granting university (e.g., geolocation), as well as basic information about the individual (e.g., gender). The DDC is also leveraging linkages in the RISIS Infrastructure to develop a caeer indicator.</p> <p>This second iteration covers the PhD production for the full two cohorts (2010,2014) for six countries ( AT, DE, IL, NL, ES, NO). The documentation details the design and contents of the dataset.</p> <p> </p>
Health-Watcher Requirements Document
<p>This repository contains the Health-Watcher Requirements Document, a requirements specification for a city hall public health system developed jointly by academic and industrial contributors for research purposes.</p>
Review of documents about Open Science
<p>In a review exercise of a sample of 46 documents about Open Science governance, from institutions and regulatory actors, we identified opportunities to strengthen Open Science practices related to integrity, traceability, and preservation. </p><p>Variables to be considered: Document name, file type, institution – country, PDF/A, have DOI, have other PID, recognized by Zotero, recognized by Mendeley, metadata in PDF properties.</p><p> </p>
Data Files and Documentation - National Coordination of Data Steward Education in Denmark
<p>This site contains the Data Files and Documentation Files from the National Coordination of Data Steward Education in Denmark Project (2019-2020). The main deliverables from the project are:</p> <p>The Main Report:</p> <p><a href="https://doi.org/10.5282/zenodo.3609516">National Coordination of Data Steward Education in Denmark</a> (<strong>Main Report</strong>)</p> <p>and</p> <p>Wildgaard, Lorna</p> <p><a href="https://doi.org/10.5281/zenodo.3628375">Reframing Data Stewardship educations in Denmark and abroad</a> (<strong>Wildgaard</strong>)</p> <p> </p> <p>List of available Data and Documentation Files below:</p> <p><strong>Wildgaard </strong>and<strong> Main Report (section 2 Review of Data Steward Education)</strong></p> <p><em>1. DS_education_1_review_thematic_analysis.nvp (NVIVO-file) </em></p> <p><em>2. DS_education_2_review_coding_scheme (Excel) </em></p> <p><em>(Created by Wildgaard, Lorna)</em></p> <p><strong>Main Report (section 3 LinkedIn analysis)</strong></p> <p>No data available due to GDPR</p> <p><strong>Main Report (section 4 Job vacancies analysis)</strong></p> <p><em>3. DS_education_3_vacancies_and_method.zip (zip-file) </em></p> <p><em>4. DS_education_4_vacancies_top10_word_frequencies_geographical_location (Excel) </em></p> <p><em>(Created by Vlachos, Evgenios)</em></p> <p>The zip-file contains a text file describing the method used in the analysis (Method) and 119 pdf files with Data Steward vacancies.</p> <p><strong>Main Report (section 5 Questionnaire)</strong></p> <p><em>5. DS_education_5_questionnaire_questions </em></p> <p><em>6. DS_education_6_questionnaire_survey_report </em></p> <p><em>(Created by Vlachos, Evgenios & Knudsen, Christian B.)</em></p> <p><strong>Main Report (section 6 Interviews)</strong></p> <p>7. <em>DS_education_7_interviews_summaries_interview_1-4</em> </p> <p><em>(Created by Hüser, Falco)</em></p>
Data Documentation WeatherAggReOpt
<p>This data documentation describes a data set of the German, France and Polish electricity system compiled within the research project “WeatherAggReOpt” (Developing Aggregation and Reduction Methods for Implementing Disaggregated Renewable Infeed Profiles in Energy System Models). The project is a collaboration between the Chair for Management Science and Energy Economics at the University of Duisburg-Essen and the Fraunhofer Institute for Solar Energy Systems (ISE). With a project period of three years WeatherAggReOpt (03ET4042A) is funded by the Federal Ministry for Economic Affairs and Energy (BMWi).</p>
Kamakhya Temple কামাখ্যা দেৱালয় (Guwahati, Assam), as documented 1959.
<p>Kamakhya Temple কামাখ্যা দেৱালয় (Guwahati, Assam), as documented 1959.</p>
Sealing. Circular sealing in terracotta with a standing bull and inscription (Yaudheya?) above; on the reverse the impression of a document
<p>Sealing. Circular sealing in terracotta with a standing bull and inscription (Yaudheya?) above; on the reverse the impression of a document. British Museum 1880.100.</p>
Benchmark for the evaluation of named entity recognition over ancient documents
<p>The dataset consists of a multilingual noisy corpora for named entity recognition (NER).<br> The noisy versions are simulated from the CoNLL-02 (Spanish and Dutch) and CoNLL-03 (English) NER corpora.<br> The original collections are re-OCRed and four types of noises at two different levels are added in order to simulate various OCR output.</p> <p>More precisely, we first extracted raw texts and converted them into images. These images have been contaminated by adding some common noises when using a scanner. We further extract OCRed data using tesseract open source<br> OCR engine v-3.04.01. Consequently to the image noise insertions, OCRed data contains degradations. Original and noisy texts are finally aligned.</p> <p>This archive contains three folders (one per language). The folders contain the degraded images, the noisy texts extracted by the OCR and their aligned version with clean data.</p> <p>These are the supplementary materials for the TPDL 2020 paper <a href="https://zenodo.org/record/4734376#.YJKAcKE6-Uk">Assessing and minimizing the impact of OCR quality on named entity recognition</a>. If you end up using whole or parts of this resource,<br> please cite this paper:</p> <pre><code>@InProceedings{10.1007/978-3-030-54956-5_7, author="Hamdi, Ahmed and Jean-Caurant, Axel and Sid{\`e}re, Nicolas and Coustaty, Micka{\"e}l and Doucet, Antoine", editor="Hall, Mark and Mer{\v{c}}un, Tanja and Risse, Thomas and Duchateau, Fabien", title="Assessing and Minimizing the Impact of OCR Quality on Named Entity Recognition", booktitle="Digital Libraries for Open Knowledge", year="2020", publisher="Springer International Publishing", address="Cham", pages="87--101", isbn="978-3-030-54956-5" }</code></pre> <p><strong>Acknowledgments</strong><br> This work has been supported by the European Union's Horizon 2020 research and innovation programme under grant 770299 [NewsEye](https://www.newseye.eu/).</p>
'Documents Laid' Metadata 1922-2019
<p>As part of the centenary celebrations of the first Dáil (Irish parliament) in 2019 the Oireachtas Library embarked on a project to prepare and publish an open dataset of metadata from its Documents Laid collection.</p> <p>‘Documents Laid’ is the collective term for documents that are formally submitted to the Irish parliament to support the democratic process. They are most often laid pursuant to a statutory obligation or standing order of the Houses. Document types might include: regulations and orders; financial statements and annual reports from Departments, Committees and public bodies; post enactment reports for primary legislation, etc.</p> <p>Responsibility for the management of all procedures relating to the laying of documents is assigned to the Oireachtas Library. The Library adds metadata (titles, department names, subjects, etc.) as catalogue records for each document.</p> <p>The collection provides a unique overview of the work of successive governments, parliaments and public bodies throughout the life of the Irish state. Documents Laid are an essential element of parliamentary scrutiny, being referenced or evaluated through debate, parliamentary questions or Committee meetings.</p> <p>The catalogue records for all documents laid between 1922 and 2019 were exported from the Library’s catalogue. The dataset was cleaned in OpenRefine by a librarian. Additional cleaning and standardisation was carried out by a data analyst from the School of Computer Science in University College Dublin. The cleaned dataset contains 86,449 records, described by 26 fields. It is stored in CSV and JSON format.</p> <p>It is intended that this open dataset will be of interest to scholars of modern Irish history and the Irish parliament, particularly in the area of parliamentary oversight and scrutiny.</p>
MEDDOCAN corpus: gold standard annotations for Medical Document Anonymization on Spanish clinical case reports
<p><strong>Intro:</strong></p> <p>Meddocan shared task dataset (divided in train, dev and test). In addition, we include here the Meddocan background set.</p> <p>It contains the training, development and test sets of the Meddocan shared task with Gold Standard annotations.</p> <p>In addition, it contains the documents of the background set, without annotations.</p> <p> </p> <p><strong>Annotation quality</strong></p> <p>Inter-annotator agreement: 98% </p> <p>For more information, see the <a href="http://ceur-ws.org/Vol-2421/MEDDOCAN_overview.pdf">paper</a>. </p> <p> </p> <p><strong>Format:</strong></p> <p>Annotations are distributed in Brat format. See <a href="https://brat.nlplab.org/standoff.html">Brat webpage</a> for more information.</p> <p>In addition, annotations are also distributed in XML format (based on i2b2 XML format).</p> <p>In the <a href="https://temu.bsc.es/meddocan/index.php/resources/">Meddocan webpage</a>, there is a script to convert between MEDDOCAN-Brat, MEDDOCAN-XML, and i2b2 formats.</p> <p> </p> <p><strong>Shared task goal:</strong></p> <p>In the three subtasks, the goal will be to predict the annotations given only the plain text files. </p> <p> </p> <p><strong>Resources:</strong></p> <ul> <li><strong><a href="https://temu.bsc.es/meddocan/">Web</a></strong></li> <li><strong>Citation: </strong>Montserrat Marimon et al. “Automatic De-identification of Medical Texts in Spanish: the MEDDOCAN Track, Corpus, Guidelines, Methods and Evaluation of Results.” In: IberLEF@ SEPLN. 2019, pp. 618–638.</li> <li><strong>Silver Standard corpus</strong></li> <li><a href="https://doi.org/10.5281/zenodo.4279337"><strong>Annotation guidelines</strong></a></li> </ul> <p> </p> <p>For further information, please visit <a href="https://temu.bsc.es/meddocan/">https://temu.bsc.es/meddocan/</a> or email us at encargo-pln-life@bsc.es</p> <p>Copyright (c) 2019 Secretaría de Estado para el Avance Digital (SEAD)</p>
PARADE: Passage Representation Aggregation for Document Reranking
<p>This submission includes all pretrained models on MSMARCO, and run files on the Robust04/GOV2 dataset for the paper "PARADE: Passage Representation Aggregation for Document Reranking". Please follow the instructions in the <a href="https://github.com/canjiali/PARADE">PARADE repo</a> to reproduce the results.</p> <p> </p>
Data set 3 photographic documentation BRAD research project
<p>Photographic documentation collected during the BRAD research project. For the personal data protection reasons, the published pictures do not represent recognizable people. The pictures present the places where part of the fieldwork was done (London, Croydon in the UK, Poznań in Poland). A separate folder contains images related to EUSS application.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.