Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

75

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

75 results for “Document Datasets”

Learn how ShareScore rates datasets ↗
zenodo48/100

LEARN-COVID: Dataset and documentation

<p>The LEARN-COVID pilot study collected data on infants and their parents during the COVID-19 pandemic. Assessments took place between April and July 2021. Predominantly Swiss parents answered a baseline questionnaire on their behaviour related to the pandemic, social support, infant nutrition, and infant regulation. Subsequently, parents answered a 10-day evening diary on daily nutrition, infant regulation, parental mood, and parental soothing behaviour.</p>

opencc-by-4.0Mar 2022View details →
zenodo44/100

zbMATHOpenRec: A Gold Standard Dataset for Recommending Scientific Documents with Mathematical Content

<p>&nbsp;</p> <p>Here we include the first gold standard dataset for recommending scientific documents with mathematical content.&nbsp;</p> <p><strong>Contents:&nbsp;</strong></p> <p>As of Feb-2023, there are 421 recommendation pairs with 80 seed documents.</p> <ol> <li>All recommendation pairs are available: recommendationPairs.csv</li> <li>Each document's contents, such as title, abstract/review/summary, authors, MSC codes, Full-text link, references, etc. are available in: documentContents.csv</li> </ol> <p><strong>Dataset construction process</strong>:</p> <p>This is the first gold standard content-based RS dataset, consisting of 421 scientific research entry recommendation pairs with mathematical content. The purpose is to enable math in scientific documents for document recommendations, meaning if two documents have similar math content, one could be recommended to the other.&nbsp;</p> <p>To create this dataset, we analyzed 4.5 million research entires from zbMATH Open (https://zbmath.org/) and performed the following steps to obtain the final dataset:</p> <ol> <li>We selected 80 seeds that capture the most word and math tokens in zbMATH Open using statistical measures.</li> <li>Three experts, one with several years of experience reviewing research entries in mathematics, curated the recommendations for 80 seeds.</li> </ol> <p>Using this dataset, researchers can accelerate the development and testing of recommendation approaches for scientific literature with mathematical content, improving recommendations for the STEM fields where mathematical content is currently being ignored</p> <p>## License&nbsp;</p> <p>Legal restrictions and copyright: The zbMATH Open data is subject to the Terms and Conditions for the zbMATH Open API Service of FIZ Karlsruhe &ndash; Leibniz-Institut f&uuml;r Informationsinfrastruktur GmbH. Content generated by zbMATH Open, such as reviews, classifications, software, or author disambiguation data, are distributed under CC-BY-SA 4.0. This defines the license for the whole dataset, which also contains non-copyrighted bibliographic metadata and reference data derived from I4OSC (CC0).</p>

opencc-by-4.0Jan 2023View details →
zenodo44/100

Nuremberg Letterbooks: A Multi-Transcriptional Dataset of Early 15th Century Manuscripts for Document Analysis

<p>This dataset contains the images and labels of the Nuremberg Letterbooks dataset.</p> <p>It consists of four books (books 2 - 5) with line-wise transcriptions. Three kinds of transcriptions are reported: basic, regularized, and diplomatic, with additional expanded abbreviations.&nbsp;</p> <p>Code templates for text verification and writer verification are available at:</p> <ul> <li><a href="https://github.com/M4rt1nM4yr/letterbooks_text_verification">https://github.com/M4rt1nM4yr/letterbooks_text_verification</a></li> <li><a href="https://github.com/M4rt1nM4yr/letterbooks_writer_verification">https://github.com/M4rt1nM4yr/letterbooks_writer_verification</a></li> </ul> <p>When using this dataset, please cite:&nbsp;<br>M. Mayr, J. Krenz, K. Neumeier, A. Bub, S. B&uuml;rcky, N. Brolich, K. Herbers, M. Habermann, P. Fleischmann, A. Maier, and V. Christlein<em>.</em> <br>Nuremberg Letterbooks: A Multi-Transcriptional Dataset of Early 15th Century Manuscripts for Document Analysis. <em>Sci Data</em> <strong>12</strong>, 811 (2025).<br><a href="https://doi.org/10.1038/s41597-025-05144-z">https://doi.org/10.1038/s41597-025-05144-z</a></p>

opencc-by-4.0Oct 2024View details →
zenodo44/100

DOCUMENTATION OF RISIS DATASETS Doctoral Degree and Career Dataset (DDC)

<p>Documentation is presented for The Doctoral Degree and Career Dataset (DDC).&nbsp; In the framework of the RISIS2 project, DDC is an experimental dissertation-centric database. It primarily consists of an enriched PhD publication dataset which brings togehter information about the dissertation (e.g., topic mapping), about the degree-granting university (e.g., geolocation), as well as basic information about the individual (e.g., gender). The DDC is also leveraging linkages in the RISIS Infrastructure to develop a caeer indicator.</p> <p>This second iteration covers the PhD production for the full two cohorts (2010,2014) for six countries ( AT, DE, IL, NL, ES, NO). The documentation details the design and contents of the dataset.</p> <p>&nbsp;</p>

opencc-by-4.0Mar 2023View details →
zenodo40/100

Long document similarity datasets, Wikipedia excerptions for movies, video games and wine collections

<p>Three&nbsp;corpora in different domains extracted from Wikipedia.</p> <p>For all datasets, the figures and tables have been filtered out, as well as the categories and &quot;see also&quot; sections.</p> <p>The article structure, and&nbsp;particularly the sub-titles and paragraphs are kept in these datasets</p> <p>&nbsp;</p> <p><strong>Wines</strong></p> <p>Wikipedia wines dataset consists of 1635 articles from the wine domain. The extracted dataset consists of a non-trivial mixture of articles, including different wine categories, brands, wineries, grape types, and more. The ground-truth recommendations were crafted by a human sommelier, which annotated 92 source articles with ~10 ground-truth recommendations for each sample. Examples for ground-truth expert-based recommendations are&nbsp;</p> <ul> <li>Dom P&eacute;rignon - Mo&euml;t &amp; Chandon</li> <li>Pinot Meunier - Chardonnay</li> </ul> <p><strong>Movies</strong></p> <p>The Wikipedia movies dataset consists of 100385 articles describing different movies. The movies&#39; articles may consist of text passages describing the plot, cast, production, reception, soundtrack, and more.<br> For this dataset, we have extracted a test set of ground truth annotations for 50 source articles using the &quot;<a href="https://bestsimilar.com/">BestSimilar</a>&quot;&nbsp;database. Each source articles is associated with a list of ${\scriptsize \sim}12$ most similar movies.<br> Examples for ground-truth expert-based recommendations are&nbsp;</p> <ul> <li>Schindler&#39;s List - The Pianist</li> <li>Lion King - The Jungle Book</li> </ul> <p><strong>Video games</strong></p> <p>The Wikipedia video games dataset consists of 21,935 articles reviewing video games from all genres and consoles. Each article may consist of a different combination of sections, including summary, gameplay, plot, production, etc. Examples for ground-truth expert-based recommendations are:</p> <ul> <li>Grand Theft Auto - Mafia</li> <li>Burnout Paradise - Forza Horizon 3</li> </ul>

opencc-by-4.0Jan 2021View details →
zenodo40/100

Test Dataset for Random Document access

Open the record for dataset details and reuse information.

opencc-by-4.0Dec 2023View details →
zenodo40/100

SDADDS-Guelma : A Multi-purpose Dataset for Synthetic Degraded Arabic Documents

<h1><strong>SDADDS-Guelma : A Multi-purpose Dataset for Synthetic Degraded Arabic Documents </strong></h1> <h2><strong>Description:</strong></h2> <p>This is a partial release of the SDADDS-Guelma dataset.</p> <p>SDADDS-Guelma (Synthetic Degraded Arabic Document DataSet of the University of Guelma) is a database of synthetic noisy or degraded Arabic document images. It was created by Dr. Abderrahmane Kefali and his team to support research on preprocessing, analysis, and recognition of degraded Arabic documents, where having a large set of images for training and testing is essential. This dataset is made publicly available to researchers in the field of document analysis and recognition, with the hope that it will be useful and contribute to their research endeavors.</p> <p>In this first release of the dataset, 84 handwritten images and 120 printed images have been used, along with 25 images of historical backgrounds, forming a total of 26316 synthetic images of degraded Arabic documents along with their corresponding ground-truth files.</p> <p>This release is separated into two parts to facilitate upload and use: one for the handwritten documents and the second for the printed documents.</p> <h2><strong>Composition of the dataset:</strong></h2> <p>Each of the parts of the SDADDS-Guelma dataset is organized into directories as follows:</p> <ul> <li>TXT_Files: Contains texts in UTF-8 format. </li> <li>IMG: Contains images of printed and handwritten Arabic text constructed from the text files. </li> <li>Bin_IMG: Contains binary images corresponding to the original images. </li> <li>BG_IMG: Contains images of empty old document backgrounds used for the generation of synthetic historical document images. </li> <li>GT_Files: Contains XML annotation files corresponding to the text images.</li> <li>Degraded_IMG: This directory contains synthetically generated degraded images, separated into sub-directories based on noise types such as Local_Noise, Show_through, Rotation, Curvature, Comb_IMG, etc.</li> </ul> <h2><strong>Ground-truth information:</strong></h2> <p>Ground truth information is essential for a document dataset, as it annotates documents and represents their essential characteristics. Our dataset is designed to be a large-scale and multipurpose dataset. As such, our methodology ensures that ground truth information is provided at three levels: text level (character codes), pixel level (binary and cleaned image), and document physical structure and other annotation information level.</p> <ul> <li>Textual Ground Truth: these are identical to the original texts.&nbsp;</li> <li>Pixel-level ground truth: presented in the form of binary images.</li> <li>Ground truth at the document structure level: the structure of each document image, alongside the textual transcription of the words and PAWs, is recorded in a corresponding XML annotation file. The XML format utilized resembles that employed in similar works with adjustments made according to the specific characteristics of Arabic texts, including the presence of PAWs.&nbsp;</li> </ul> <p>Consequently, each original text image in our dataset is associated to an XML file detailing the entire ground truth and associated metadata.&nbsp;</p> <h3><em><strong>Structure of XML file:</strong></em></h3> <p>Each XML annotation file contains metadata about the document image and text content within the image, including the language, number of lines, and font attributes. It also provides detailed information about each text line, word, and Part of Arabic Words (PAWs), including their bounding boxes and textual transcriptions.</p> <p>Thus, each ground truth file takes the following form:</p> <pre><code>&lt;DOCUMENT imageName="PR1Kufi_bin.png" height="2631" width="1860" nbTextLines="8" language="Arabic" fontName="Kufi" fontSize="34"&gt; &lt;TEXTLINE id="0" nbWords="3" boundingBox="215,481,355,1379"&gt; &lt;WORD id="0" nbPAWs="2" boundingBox="217,1065,341,1379" transcription="خصائص"&gt; &lt;PAW id="0" nbCCs="2" boundingBox="217,1206,322,1379" transcription="خصا"&gt; &lt;CC id="0" nbPixels="4110" pixels="(217,1206,1206);(218,1206,1207);(219,1206,1208);(220,1206,1208);(221,1206,1211);(222,1206,1211);..."&gt; &lt;/CC&gt; &lt;CC id="1" nbPixels="80" pixels="(263,1330,1336);(264,1329,1336);(265,1328,1337);(266,1328,1337);(267,1328,1337);(268,1328,1337);(269,1328,1337);...."&gt; &lt;/CC&gt; &lt;/PAW&gt; .... &lt;/WORD&gt; &lt;WORD id="1" nbPAWs="2" boundingBox="215,817,338,1044" transcription="التفسير"&gt; &lt;PAW id="0" nbCCs="1" boundingBox="215,1030,322,1044" transcription="ا"&gt; &lt;CC id="0" nbPixels="1037" pixels="(215,1030,1030);(216,1030,1030);(217,1030,1031);..."&gt;&lt;/CC&gt; .... &lt;/PAW&gt; .... &lt;/WORD&gt; &lt;/TEXTLINE&gt; .... &lt;/DOCUMENT&gt;</code></pre> <h1><strong>Contact:</strong></h1> <p>Name: &nbsp; &nbsp; &nbsp; &nbsp; Dr. Abderrahmane Kefali<br>Affiliation: &nbsp; &nbsp; University of 8 May 1945-Guelma, Algeria<br>Email: &nbsp; &nbsp; &nbsp; &nbsp; kefali.abderrahmane@univ-guelma.dz</p>

opencc-by-4.0Apr 2024View details →
zenodo40/100

The Clarity Software Documentation Dataset

<p>This repository holds the Clarity Dataset which is a companion to the SANER&#39;22 entitled &quot;An Empirical Investigation into the Use of Image Captioning for Automated Software Documentation&quot;. The dataset consists of 45,998 captions&nbsp;10,204 GUI screenshots and xml metadata files (akin to the &quot;html&quot; for stipulating GUIs)&nbsp;of Android applications.&nbsp;The NL captions were obtained from human labelers, underwent several quality control mechanisms, and contain both high- (screen-level) and low-(component)&nbsp;level descriptions of screen functionality. This dataset is meant as a new source of data to augment techniques for software documentation that can take advantage of the rich pixel-based information contained within screenshots.</p>

opencc-by-4.0Jan 2022View details →
zenodo40/100

Data Cleaning, Translation & Split of the Dataset for the Automatic Classification of Documents for the Classification System for the Berliner Handreichungen zur Bibliotheks- und Informationswissenschaft

<ul> <li>Cleaned_Dataset.csv &ndash; The combined CSV files of all scraped documents from DABI, e-LiS, o-bib and Springer.</li> <li>Data_Cleaning.ipynb &ndash; The Jupyter Notebook with python code for the analysis and cleaning of the original dataset.</li> <li>ger_train.csv &ndash; The German training set as CSV file.</li> <li>ger_validation.csv &ndash; The German validation set as CSV file.</li> <li>en_test.csv &ndash; The English test set as CSV file.</li> <li>en_train.csv &ndash; The English training set as CSV file.</li> <li>en_validation.csv &ndash; The English validation set as CSV file.</li> <li>splitting.py &ndash; The python code for splitting a dataset into train, test and validation set.</li> <li>DataSetTrans_de.csv &ndash; The final German dataset as a CSV file.</li> <li>DataSetTrans_en.csv &ndash; The final English dataset as a CSV file.</li> <li>translation.py &ndash; The python code for translating the cleaned dataset.</li> </ul>

opencc-by-4.0Aug 2022View details →
zenodo40/100

PURE: a Dataset of Public Requirements Documents

<p>Please cite this dataset as&nbsp;<strong>Ferrari, A., Spagnolo, G. O., &amp; Gnesi, S. (2017, September). PURE: A dataset of public requirements documents. In&nbsp;<em>2017 IEEE 25th International Requirements Engineering Conference (RE)&nbsp;</em>(pp. 502-505). IEEE.</strong></p> <p><a href="https://ieeexplore.ieee.org/abstract/document/8049173">https://ieeexplore.ieee.org/abstract/document/8049173</a></p> <p>This dataset presents PURE (PUblic REquirements dataset), a dataset of 79 publicly available natural language requirements documents collected from the Web. The dataset includes 34,268 sentences and can be used for natural language processing tasks that are typical in requirements engineering, such as model synthesis, abstraction identification and document structure assessment. It can be further annotated to work as a benchmark for other tasks, such as ambiguity detection, requirements categorisation and identification of equivalent re-quirements. In the associated paper, we present the dataset and we compare its language with generic English texts, showing the peculiarities of the requirements jargon, made of a restricted vocabulary of domain-specific acronyms and words, and long sentences. We also present the common XML format to which we have manually ported a subset of the documents, with the goal of facilitating replication of NLP experiments. The XML documents are also available for download.</p> <p>The paper associated to the dataset can be found here:&nbsp;</p> <p>https://ieeexplore.ieee.org/document/8049173/</p> <p>More info about the dataset is available here:&nbsp;</p> <p>http://nlreqdataset.isti.cnr.it</p> <p>Preprint of the paper available at ResearchGate:</p> <p>https://goo.gl/HxJD7X</p> <p>The dataset includes:</p> <p>- all the documents in PDF format</p> <p>- a subset of 19 documents in XML format</p> <p>- the .xsd schema of the XML files</p> <p>The dataset has been created by gathering data from web sources and we are not aware of license agreements or intellectual property rights on the requirements. The curator took utmost diligence in minimizing the risks of copyright infringement by using non-recent data that is less likely to be critical, by sampling a subset of the original requirements collection, and by qualitatively analyzing the requirements. In case of copyright infringement, please contact the dataset curator (Alessio Ferrari, alessio.ferrari@cnr.it, alessio.ferrari@ucd.ie) to discuss the possibility of removal of that dataset [see <a href="https://support.zenodo.org/help/en-gb/13-policies/140-what-is-your-take-down-procedure" target="_blank" rel="noopener">Zenodo's policies</a>].</p>

opencc-by-4.0Sep 2018View details →
zenodo40/100

EcoregionsTreeFinder – a global dataset documenting observations of 48,129 tree species in 828 terrestrial ecoregions

<p>Check this article for a description of the methods used to develop the EcoregionsTreeFinder. Together with the citation for this Zenodo archive, it is the suggested citation for the database.</p> <p>Kindt, R. and Pedercini, F. (2025), EcoregionsTreeFinder&mdash;A Global Dataset Documenting the Abundance of Observations of &gt;45,000 Tree Species in 828 Terrestrial Ecoregions. Global Ecol Biogeogr, 34: e70064. <a href="https://doi.org/10.1111/geb.70064">https://doi.org/10.1111/geb.70064</a></p> <p>Use this shinyapp to filter native tree species for a particular ecoregion or to see ecoregions where a species is expected to be native: <a href="https://patspo.shinyapps.io/EcoregionsTreeFinder/" target="_blank" rel="noopener">https://patspo.shinyapps.io/EcoregionsTreeFinder/</a></p> <p>&nbsp;</p> <p>The database was created from observation records filtered from: GBIF.org (16 March 2021) GBIF Occurrence Download&nbsp;<a href="https://doi.org/10.15468/dl.77gcvq" target="_blank" rel="noopener">https://doi.org/10.15468/dl.77gcvq</a></p> <p>&nbsp;</p> <p><strong>Funding </strong></p> <p>Development of the EcoregionsTreeFinder was supported by the&nbsp;<strong>Bezos Earth Fund</strong> via the Quality Tree Seed for Africa project, by <strong>Norway's International Climate and Forest Initiative</strong> via the Provision of Adequate Tree Seed Portfolio in Ethiopia (PATSPO) project, by the <strong>Darwin Initiative</strong> via project DAREX001 of Developing a Global Biodiversity Standard certification for tree-planting and restoration, by the &nbsp;<strong>Green Climate Fund</strong> via the Readiness proposal Burkina Faso and TREPA projects, and by the <strong>International Climate Initiative</strong> via the Right Tree for the Right Place and Right Purpose (RTRPRP) project.</p> <p>&nbsp;</p>

opencc-by-4.0Aug 2024View details →
zenodo40/100

Oficio de Hipotecas de Girona. A dataset of Spanish notarial deeds (18th Century) for Handwritten Text Recognition and Layout Analysis of historical documents.

<p>This dataset is a subset of 596 documents from the&nbsp;<em>Registre d&#39;Hipoteques de Girona</em> of 1769 collection, guarded by the <a href="http://xac.gencat.cat/ca/llista_arxius_comarcals/girones/"><em>Arxiu Hist&ograve;ric de Girona</em></a>. This collection, is composed by hundreds of thousands of notarial deeds from the XVIII-XIX century (1768-1862). Sales, redemption of censuses, inheritance and matrimonial chapters are among the most common documentary&nbsp;typologies in the collection.</p> <p>This dataset is composed of more than 23700 text lines&nbsp;written by a single hand, covering more that 50 different topics (documentary typologies) and a vocabulary of more than 2400 different words. The documents are transcribed using the so-called&nbsp;diplomatic criteria. Additionally, transcripts were tagged with&nbsp;<br> extra enriching/complementary information (e.g. expansion of the&nbsp;abbreviations, hyphen marks, etc.). Along with the transcripts &nbsp;the layout of the document is detected and recorded. Pages have&nbsp;been labeled using six different layout regions.</p> <p>The images along with their respective ground-truth was compiled in PAGE compliant XML format<br> by the <a href="http://www2.udg.edu/tabid/11296/Default.aspx"><em>Centre de Recerca d&#39;Hist&ograve;ria Rural</em></a>&nbsp;and the HTR group of the <a href="https://www.prhlt.upv.es">Pattern Recognition and Human Language Technologies Research Center</a>.</p>

opencc-by-nc-4.0Jul 2018View details →
zenodo40/100

Dataset of "What Should Developers Be Aware Of? An Empirical Study on the Directives of API Documentation"

<p>Dataset of <em>What Should Developers Be Aware Of? An Empirical Study on the Directives of API Documentation</em> (Martin Monperrus, Michael Eichberg, Elif Tekes, Mira Mezini), In Empirical Software Engineering, Springer, 2011.</p> <p><br> * dataset-src.tar.bz2 contains the source code of the Java libraries used as raw data.<br> * dataset.xml.bz2 contains the API documentation extracted from source code.<br> * directives.xml.bz2 contains the API directives found during the exploratory case study.<br> * directive-appendix.pdf is a human-readable PDF version of directives.xml.bz2.<br> <br> All datasets are published under the Creative Commons Attribution License: if you use them, please cite:<br> &nbsp;</p>

opencc-by-4.0Apr 2012View details →
zenodo40/100

ICDAR 2019 Competition on Image Retrieval for Historical Handwritten Documents [HisIR19] Dataset

<p>This dataset contains the training and test set used in the ICDAR 2019 Competition on Image Retrieval for Historical Handwritten Documents.</p> <p>This competition investigates the performance of large-scale retrieval of historical document images based on<br> writing style. Based on large image data sets provided by cultural heritage institutions and digital libraries, providing<br> a total of 20 000 document images representing about 10 000 writers, divided in three types: writers of (i) manuscript books, (ii) letters, (iii) charters and legal documents. We focus on the task of automatic image retrieval to simulate common scenarios of humanities research, such as writer retrieval.</p> <p>The training data set encompasses images from (i) Letters A, where each writer contributed one or three images; (ii) Manuscripts, where each writer was represented by five consecutive images from a single book.<br> In total, it contains 300 writers contributing one page, 100 writers contributing three pages, and 120 writers contributing five pages resulting in 1200 images of 520 writers.</p> <p>The test data set contains 20 000 images: About 7 500 pages stem from isolated documents (partially anonymous writers, contributing one page each), and about 12 500 pages are from writers that contributed three or five pages.</p> <p>&nbsp;</p> <p>If you use this dataset, please cite:</p> <p>V. Christlein, A. Nicolaou, M. Seuret, D. Stutzmann, A. Maier: &quot;ICDAR 2019 Competition on Image Retrieval for Historical Handwritten Documents&quot;, in 15th International Conference on Document Analysis and Recognition, 2019, Sydney, Australia</p> <p>&nbsp;</p>

opencc-by-nc-sa-4.0Jun 2019View details →
zenodo40/100

Supporting Datasets for 'The longest documented distance traveled by a West Indian manatee'

<p>This dataset contains multiple sources of geospatial and environmental data used for the analysis of Tico's journey. This is a West-Indian manatee that was rescued, rehabilitated and tracked after release. Tico traveled&nbsp;approximately 4,017 km over 62 days through deep oceanic waters. Correlating Tico's trajectory and velocity with surface currents revealed the influence of the North Brazil&nbsp;Current (NBC) and its vortices on his trajectory. The data integrates satellite remote sensing, ocean reanalysis, .</p> <p>The dataset include:</p> <ol> <li> <p><strong>Geographic locations of Tico (Delemetry Data)</strong></p> </li> <li> <p><strong>General Bathymetric Chart of the Oceans (GEBCO)<br></strong></p> </li> <li> <p><strong>NEMO - Global Ocean Physics Analysis and Forecast data:&nbsp;</strong>Surface current data from the NEMO model provided by CMEMS (GLORYS reanalysis).<strong><br></strong></p> </li> <li> <p><strong>HYCOM Global Ocean Forecasting System (GOFS) 3.1 Analysis:<br></strong></p> </li> <li> <p><strong>OSCAR - Ocean Surface Current Analyses Real-time<br></strong></p> </li> <li> <p><strong>Soil Moisture Active Passive (SMAP)<br></strong></p> </li> <li> <p><strong>Integrated Multi-satellite Retrievals for Global Precipitation Measurements (IMERG)</strong></p> </li> </ol> <p><strong>Disclaimer:</strong></p> <p>This dataset is a compilation of various data sources, each of which may be subject to its own specific licensing terms and conditions. Users of this dataset are strongly advised to review the license associated with each individual sub-dataset before using, modifying, or redistributing the data. While we have provided the dataset under a Creative Commons Attribution 4.0 International (CC BY 4.0) license, certain sub-datasets may have additional restrictions or requirements, such as attribution, non-commercial use, or limitations on derivative works. It is the responsibility of the user to ensure compliance with all applicable licenses. Please refer to the original data sources and their respective licenses for detailed information.</p>

opencc-by-4.0Sep 2024View details →
zenodo40/100

Dataset of "Knowledge-based Sense Disambiguation of Multiword Expressions in Requirements Documents"

<p>This is the dataset used in the paper &quot;Knowledge-based Sense Disambiguation of Multiword Expressions in Requirements Documents&quot; at AIRE&#39;21</p> <p>&nbsp;</p> <p>In this paper, we explore the use of a multiword expression detection in combination with a knowledge-based word sense disambiguation to disambiguate expressions in requirements documents.</p> <p>The dataset comprises a gold standard for multiword expression detection and sense disambiguation for Wikipedia and WordNet 3.1.</p> <p>It covers 18 projects: CM1, EBT and GANTT as well as the 15 projects of the NFR dataset.</p> <p>&nbsp;</p> <p><strong>File format</strong></p> <p>We use a tab-separated version of the DiMSUM file format and extended it with sense information.</p> <p>The nine original DiMSUM tab-separated columns:</p> <p>1. token offset</p> <p>2. word</p> <p>3. lowercase lemma</p> <p>4. POS</p> <p>5. MWE tag</p> <p>6. offset of parent token (i.e. previous token in the same MWE), if applicable</p> <p>7. strength level encoded in the tag, if applicable. Currently not used</p> <p>8. supersense label, Currently not used</p> <p>9. sentence ID</p> <p>&nbsp;</p> <p>and the two further columns for sense information:</p> <p>10. Wikipedia article name</p> <p>11. WordNet 3.1 synset</p> <p>&nbsp;</p> <p>The last two columns might end with .1 or .0 indicating that the sense is a fully applicable or partial sense of a multiword expression.</p> <p><strong>Attribution (of datasets used)</strong></p> <p>The NFR Dataset can be attributed to Jane Cleland-Huang.<br> Jane Cleland-Huang, Sepideh Mazrouee, Huang Liguo, &amp; Dan Port. (2007). nfr [Data set]. Zenodo. Available:&nbsp;<a href="http://doi.org/10.5281/zenodo.268542">http://doi.org/10.5281/zenodo.268542</a><br> &nbsp;</p> <p>The CM1, EBT and GANTT datasets were retrieved from the Center of Excellence for Software &amp; Systems Traceability (CoEST)&nbsp;<a href="https://doi.org/10.5281/zenodo.3309669">http://coest.org/</a></p>

opencc-by-4.0Aug 2021View details →
zenodo40/100

Vorau Abbey library Cod. 253 dataset for Document Layout Analysis

<p>VORAU-253 is a music manuscript referred to as Cod. 253 of the Vorau Abbey library, which was provided by the Austrian Academy of Sciences. It is written in German Gothic notation and dated around year 1450.</p> <p>This manuscript is interesting because of the complexity of its layout, where staff, text and decorations are intertwined to<br> compose the structure of the document.</p> <p>This database is a subset of 228 pages of the archive, using 128 randomly selected pages for training/validation and 100 for test.</p> <p>The database was manually annotated into the following three layout regions:</p> <p>* staff: represents the regions that contains a set of horizontal lines and spaces where each one represent a different musical pitch. This region type does not contain text lines. Hence, no baselines.</p> <p>* lyrics: are the words that are sung appear below their corresponding staff, and other text in the document. In all cases, text to be sung and the other text are assigned to different layout regions under the lyrics label.</p> <p>* drop-capital: is a decorated letter that might appear at the beginning of a word or text line. As it is a single big letter, it contain no text lines nor baselines.</p> <p>On average each page contains 12.5 [7,23] text lines distributed over an average&nbsp;of 10.5[7,15] ``lyrics&#39;&#39; regions. Moreover, each page contains 22.3[14,28] layout regions on average.</p>

opencc-by-4.0Sep 2021View details →
zenodo40/100

ICDAR 2021 Historical Document Classification Test Dataset for Task 1 - Font Groups

<p>Test set for the font group classification task of the ICDAR 2021 Competition on Historical Document Classification competition. The tar.gz file contains the images and a ground truth CSV file. The csv file indicates for each image from which document and page it corresponds to, as well as whether augmentations have been applied.</p>

opencc-by-4.0May 2021View details →
zenodo40/100

UDAPDR Document and Question Datasets

<p>Question and document datasets for&nbsp;<a href="https://arxiv.org/abs/2303.00807">UDAPDR: Unsupervised Domain Adaptation via LLM Prompting and Distillation of Rerankers</a></p>

opencc-by-4.0Mar 2023View details →
zenodo40/100

Dataset od: "Towards a Taxonomy of Roxygen Documentation in R Packages"

<p>Replication package for the paper titled &quot;Towards a Taxonomy of Roxygen Documentation in R Packages&quot;</p>

opencc-by-4.0Mar 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record