Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
46
datasets available to search
ShareScore release 0.9.0
Dataset results
46 results for “readability”
OpenITI: a Machine-Readable Corpus of Islamicate Texts
<p><strong>Co-PIs</strong>: Matthew Thomas Miller (University of Maryland, College Park), Maxim G. Romanov (University of Hamburg), Sarah Bowen Savant (Aga Khan University—ISMC, London).</p> <p><em>Open Islamicate Texts Initiative</em> (<strong>OpenITI</strong>, see <a href="https://openiti.org/">https://openiti.org/</a>) is a multi-institutional effort to construct the first machine-actionable scholarly corpus of premodern Islamicate texts. Led by researchers at the Aga Khan University, Institute for the Study of Muslim Civilisations (AKU-ISMC), University of Hamburg (UH), and the Roshan Institute for Persian Studies at the University of Maryland (College Park) and an interdisciplinary advisory board of leading digital humanists and Islamic, Persian, and Arabic studies scholars, <strong>OpenITI</strong> aims to provide the essential textual infrastructure in Arabic, Persian and other Islamicate languages for new forms of textual analysis and digital scholarship. In the process, OpenITI will enable new synergies between Digital Humanities and the inter-related Islamicate fields of Islamic, Persian, and Arabic Studies. In addition to support from the researchers’ home institutions, it is supported by funding from the <a href="https://erc.europa.eu/">European Research Council</a> under the European Union’s Horizon 2020 research and innovation programme, awarded to the <a href="http://kitab-project.org/">KITAB</a> project (Grant Agreement No. 772989, PI Sarah Bowen Savant) and the <a href="https://www.qnl.qa/en">Qatar National Library</a>.</p> <p>Currently, <strong>OpenITI</strong> contains almost exclusively Arabic texts, which were first assembled into a corpus within the <strong>OpenArabic</strong> project, developed first at Tufts University (at <em>The Perseus Project</em>, 2013–2015) and then at Leipzig University (at the Alexander von Humboldt Chair for Digital Humanities, 2015–2017)—in both cases with the support and under the patronage of Prof. Gregory Crane. The much more limited number of Persian texts were compiled during 2015–2016 in the Persian Digital Library (PDL) pilot (see <a href="https://persdigumd.github.io/PDL/">Persian Digital Library by PersDigUMD</a>) at Roshan Institute for Persian Studies at the University of Maryland. These texts have not been made fully compatible with OpenITI mARkdown yet and will be made fully available in next releases.</p> <p>This release contains all digital versions of the same text that are available in the OpenITI corpus . <strong>We also release a <a href="https://doi.org/10.5281/zenodo.7764025">'primary' version of the corpus</a></strong> that contains a single digital version for each text in the corpus that is marked as 'PRI' in the corpus metadata and may be more convenient for some use cases.</p> <p><strong>Note on Release Numbering</strong>: Version <strong>2019.1.1</strong>—where <strong>2019</strong> is the year of the release, the first dotted number—<strong>.1</strong>—is the ordinal release number in 2019, and the second dotted number—<strong>.1</strong>—is the overall release number; the first dotted number will reset every year, while the second one will continue on increasing.</p> <p>For more details: <a href="https://github.com/OpenITI/RELEASE">https://github.com/OpenITI/RELEASE</a></p> <p><strong>Note: </strong>In case of any issues with unzipping the files on Windows using built-in utilities, please use free softwares, such as WinRAR and 7zip.</p> <p> </p>
A Study of Improving Mathematics Assessment Readability by using GPT-3
<p><strong>A Study of Improving Mathematics Assessment Readability by using GPT-3</strong></p> <p>This dataset contains 250 math word problems from EngageNY, and their automated simplifications generated using the GPT-3 engine. There are eight ways each prompt is simplified, and three samples are taken for each prompt. Some API calls can fail or return a blank string. Those responses are omitted. Readability measures for input and output passages are generated using the Common Text Analysis Platform (<a href="http://sifnos.sfs.uni-tuebingen.de/ctap/">http://sifnos.sfs.uni-tuebingen.de/ctap/</a>). Text similarity metrics are generated by computing the % of total words between input and output and cosine similarity of the text vectors that GPT-3 embedding API generates.</p>
On the Use of Artificially Degraded Manuscripts for Quality Assessment of Readability Enhancement Methods - Dataset & Code
<p>This object contains the dataset and python code used for the paper:</p> <p>S. Brenner and R. Sablatnig. On the Use of Artificially Degraded Manuscripts for Quality Assessment of Readability Enhancement Methods. Accepted for OAGM Workshop 2019<strong>, </strong>Steyr, Austria.</p> <p>The dataset is a modified subset of the UCL Multispectral Processed Images of Parchment Damage Dataset (<a href="http://dx.doi.org/10.14324/000.ds.1469099">10.14324/000.ds.1469099</a>). The accompanying code documents how the modified version was created and how the evaluations described in the paper were performed.</p>
Machine-Readable Vocabulary Files of the "Alter Realkatalog" (ARK) of Berlin State Library (SBB)
<p>This dataset contains two versions of vocabulary files of the <a href="https://ark.staatsbibliothek-berlin.de/">ARK (Alter Realkatalog)</a> in .tsv and .ttl format used for training models for automatic subject indexing with the modular <a href="https://github.com/NatLibFi/Annif">Annif</a> tool. As the ARK is a historical classification system which has been used to describe historical works in the Staatsbibliothek zu Berlin – Berlin State Library’s collections up to 1955, this dataset has been created for generating automatic indexing suggestions for historical texts which have not yet been manually classified with the help of the ARK (for a detailed description of the ARK, see also <a href="../doi/10.5281/zenodo.12783813">Metadata of the "Alter Realkatalog" (ARK) of Berlin State Library (SBB)</a>. Together with specific corpus training data, these vocabulary files serve as input to Annif, with which the corresponding models on <a href="https://huggingface.co/SBB">Hugging Face at the Staatsbibliothek zu Berlin – Preußischer Kulturbesitz</a> community have been created. Associated corpus training data have been extracted from the <a href="../doi/10.5281/zenodo.12783813" target="_blank" rel="noopener">Metadata of the "Alter Realkatalog" (ARK)</a> (title data).</p>
OpenITI: a Machine-Readable Corpus of Islamicate Texts; Primary Version
<p>This is a partial release of the data in the <a href="https://doi.org/10.5281/zenodo.3082463">OpenITI corpus</a>. Whereas the main release of the corpus often contains multiple digital versions of the same text, this release contains a single digital version for each text in the corpus (the version marked as "primary" (PRI) in the corpus metadata).</p> <p>The version numbers corresponds to the versions of the main releases.</p> <p> </p> <p> </p>
Machine-readable Northern Karelian Proper-Livvi bilingual translation dictionary
<p>This machine readable bilingual translation dictionary of Northern Karelian Proper (ISO-639: krl) to Livvi aka Olonets-Karelian (ISO-639 olo) was created by Timo Rantakaulio during Google Summer of Code 2021 at Apertium. His work was facilitated through an online back-end dictionary editing tool `https://akusanat.com/verdd' with translation suggestions generated by Khalid Alnajjar and Mika Hämäläinen developers of the Verdd editor. Workflow, coordination, instruction and advice were given by Flammie A. Pirinen and Jack Rueter.</p>
Machine-readable Finnish-Livvi bilingual translation dictionary
<p>This machine readable bilingual translation dictionary of Finnish to Livvi aka Olonets-Karelian (ISO-396: olo) was proofread and extended by Timo Rantakaulio during Google Summer of Code 2021 at Apertium. His work was facilitated through an online back-end dictionary editing tool `https://akusanat.com/verdd' with translation suggestions generated by Khalid Alnajjar and Mika Hämäläinen developers of the Verdd editor. Workflow, coordination, instruction and advice were given by Flammie A. Pirinen and Jack Rueter.</p>
Machine-readable Finnish-Karelian bilingual translation dictionary
<p>This machine readable bilingual translation dictionary of Finnish to Northern Karelian Proper (ISO-396: krl) was proofread and extended by Timo Rantakaulio during Google Summer of Code 2021 at Apertium. His work was facilitated through an online back-end dictionary editing tool `https://akusanat.com/verdd' with translation suggestions generated by Khalid Alnajjar and Mika Hämäläinen developers of the Verdd editor. Workflow, coordination, instruction and advice were given by Flammie A. Pirinen and Jack Rueter.</p>
To What Extent Cognitive-Driven Development Improves Code Readability?
<p>Cognitive-Driven Development (CDD) is a coding design technique<br> that aims to reduce the cognitive effort that developers place in<br> understanding a given code unit (e.g., a class). By following CDD de-<br> sign practices, it is expected that the coding units to be smaller, and,<br> thus, easier to maintain and evolve. However, it is so far unknown<br> whether these smaller code units coded using CDD standards are,<br> indeed, easier to understand. In this work we aim to assess to what<br> extent CDD improves code readability. To achieve this goal, we<br> conducted a two-phase study. We start by inviting professional<br> software developers to vote (and justify their rationale) on the most<br> readable pair of code snippets (from a set of 10 pairs); one of the<br> pairs was coded using CDD practices. We received 133 answers.<br> In the second phase, we applied the state-of-the art readability<br> model on the 10-pairs of CDD-driven refactorings. We observed<br> some conflicting results. On the one hand, developers perceived<br> that seven (out of 10) CDD-driven refactorings were more readable<br> than their counterparts; for two other CDD-driven refactorings,<br> developers were undecided, while only in one of the CDD-driven<br> refactorings, developers preferred the original code snippet. On<br> the other hand, we noticed that only one CDD-driven refactorings<br> have better performance readability, assessed by state-of-the-art<br> readability models. Our results provide initial evidence that CDD<br> could be an interesting approach for software design</p>
Simple Italian sentences ranked by readability
<p>The dataset contains 500,000 sentences extracted from the Paisà corpus (https://www.corpusitaliano.it/) which have been selected for being easy to read according to four parameters: token number, average word length, depth of the parse tree and verb "arity". The sentences are ranked by readability.</p>
Replication Package for "Improving the Readability of Generated Tests Using GPT-4 and ChatGPT Code Interpreter"
<p>While automated test generation can decrease the human burden associated with testing, it does not eliminate this burden. Humans must still work with generated test cases to interpret testing results, debug the code, build and maintain a comprehensive test suite, and many other tasks. Therefore, a major challenge with automated test generation is understandability of generated test test cases. </p> <p>Large language models (LLMs), machine learning models trained on massive corpora of textual data - including both natural language and programming languages - are an emerging technology with great potential for performing language-related predictive tasks such as translation, summarization, and decision support. </p> <p>In this study, we are exploring the capabilities of LLMs with regard to improving test case understandability.</p> <p>This package contains the data produced during this exploration:</p> <ul> <li>The examples directory contains the three case studies we tested our transformation process on: <ul> <li>queue_example: Tests of a basic queue data structure</li> <li>httpie_sessions: Tests of the sessions module from the httpie project. </li> <li>string_utils_validation: Tests of the validation module from the python-string-utils project.</li> <li>Each directory contains the modules-under-test, the original test cases generated by Pynguin, and the transformed test cases. </li> <li>Two trials were performed per case example of the transformation technique to assess the impact of different results from the LLM.</li> </ul> </li> <li>The survey directory contains the survey that was sent to assess the impact of the transformation on test readability. <ul> <li>survey.pdf contains the survey questions.</li> <li>responses.xlsx contains the survey results.</li> </ul> </li> </ul>
Vikidia En/Fr bilingual dataset for Automatic Readability Assessment
<p>Vikidia.org is a children's encyclopedia, with content targeting 8-13 year old children, in several European languages. Our dataset contains 24660 texts distributed across 6165 articles in 2 reading levels, for English and French respectively i.e., each text in the corpus has four versions: en, en-simple, fr and fr-simple, and there are 6165 slugs in total. The uniqueness of the current dataset is that these are parallel, document level aligned texts in four versions - en, en-simple, fr, fr-simple. While we did not create paragraph/sentence level alignments on the corpus, we hope that this will be a useful dataset for future English and French research on ARA and Automatic Text Simplification. This is the first such dataset in ARA, and perhaps the first readily available French readability dataset.</p> <p>This dataset is used in the paper "A neural pairwise ranking model for automatic readability assessment" by Justin Lee and Sowmya Vajjala, to appear in Findings of ACL 2022. </p>
WikiReaD (Wikipedia Readability Dataset)
<p><strong>Dataset Description:</strong></p> <p>The dataset contains pairs of encyclopedic articles in 14 languages. Each pair includes the same article in two levels of readability (easy/hard). The pairs are obtained by matching Wikipedia articles (hard) with the corresponding versions from different simplified or children's encyclopedias (easy).</p> <p> </p> <p><strong>Dataset Details:</strong></p> <ul> <li><strong>Number of Languages:</strong> 14</li> <li><strong>Number of files:</strong> 19</li> <li><strong>Use Case:</strong> Training and evaluating readability scoring models for articles within and outside Wikipedia.</li> <li><strong>Processing details:</strong> Text pairs are created by matching articles from Wikipedia with the corresponding article in the simplified/children encyclopedia either via the Wikidata item ID or their page titles. The text of each article is extracted directly from their parsed HTML version.</li> <li><strong>Files:</strong> The dataset consists of independent files for each type of children/simplified encyclopedia and each language (e.g., `<wiki>-<language_code>_sentences.bz2`). Also, the dataset contains train-test split files for <div> <div>simplewiki-en (trainsplit_simplewiki-en_sentences.bz2, testsplit_simplewiki-en_sentences.bz2) needed to reproduce the results of the corresponding paper. </div> <div> </div> </div> </li> </ul> <p><strong>Attribution:</strong></p> <p>The dataset was compiled from the following sources. The text of the original articles comes from the corresponding language version of Wikipedia. The text of the simplified articles comes from one of the following encyclopedias: Simple English Wikipedia, Vikidia, Klexikon, Txikipedia, or Wikikids.</p> <p>Below we provide information about the license of the original content as well as the template to generate the link to the original source for a given page (<page_title>) and language (<language_code>). For example, <a href="https://en.wikipedia.org/wiki/Spain">https://en.wikipedia.org/wiki/Spain</a> links to the page “Spain” in English Wikipedia)</p> <ul> <li><a href="https://www.wikipedia.org/">Wikipedia</a> <ul> <li>Source: <code>https://<language_code>.wikipedia.org/wiki/<page_title></code></li> <li>License: <a href="https://creativecommons.org/licenses/by-sa/4.0/deed.en">CC BY-SA 4.0</a>, <a href="https://www.gnu.org/copyleft/fdl.html">GFDL</a></li> </ul> </li> <li><a href="https://simple.wikipedia.org">Simple English Wikipedia</a> <ul> <li>Source: <code>https://simple.wikipedia.org/wiki/<page_title></code></li> <li>License: <a href="https://creativecommons.org/licenses/by-sa/4.0/deed.en">CC BY-SA 4.0</a>, <a href="https://www.gnu.org/copyleft/fdl.html">GFDL</a></li> </ul> </li> <li><a href="https://www.vikidia.org/">Vikidia</a> <ul> <li>Source: <code>https://<language_code>.vikidia.org/wiki/<page_title></code></li> <li>License: <a href="https://creativecommons.org/licenses/by-sa/3.0/deed.en">CC BY-SA 3.0</a>, <a href="https://www.gnu.org/copyleft/fdl.html">GFDL</a></li> </ul> </li> <li><a href="https://klexikon.zum.de">Klexikon</a> <ul> <li>Source: <code>https://klexikon.zum.de/wiki/<page_title></code></li> <li>License: <a href="https://creativecommons.org/licenses/by-sa/4.0/deed.en">CC BY-SA 4.0</a></li> </ul> </li> <li><a href="https://eu.wikipedia.org/wiki/Txikipedia">Txikipedia</a> <ul> <li>Source: <code>https://eu.wikipedia.org/wiki/Txikipedia:<page_title></code></li> <li>License: <a href="https://creativecommons.org/licenses/by-sa/4.0/deed.en">CC BY-SA 4.0</a>, <a href="https://www.gnu.org/copyleft/fdl.html">GFDL</a></li> </ul> </li> <li><a href="https://wikikids.nl/">Wikikids</a> <ul> <li>Source: <code>https://wikikids.nl/<page_title></code></li> <li>License: <a href="https://creativecommons.org/licenses/by-sa/3.0/deed.en">CC BY-SA 3.0</a></li> </ul> </li> </ul> <p><strong>Related paper citation: </strong></p> <blockquote> <pre><code>@inproceedings{trokhymovych-etal-2024-open, title = "An Open Multilingual System for Scoring Readability of {W}ikipedia", author = "Trokhymovych, Mykola and Sen, Indira and Gerlach, Martin", editor = "Ku, Lun-Wei and Martins, Andre and Srikumar, Vivek", booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)", month = aug, year = "2024", address = "Bangkok, Thailand", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/2024.acl-long.342/", doi = "10.18653/v1/2024.acl-long.342", pages = "6296--6311"<br>}</code></pre> </blockquote>
A machine-readable provisional transcription of the Piers Plowman text of Takamiya MS 23
<p>A machine-readable provisional transcription of the Middle English poem <em>Piers Plowman</em> as transmitted in New Haven, Beinecke Library, MS Takamiya 23. See the <a title="Technical Introduction" href="https://github.com/icornelius/s-takamiya-23/releases/latest/download/documentation.pdf">documentation</a>.</p>
Machine readable code lists for an algorithm to identify incident non-small cell lung cancer (NSCLC) in United States healthcare claims data
<p>Machine readable code lists for an algorithm to identify incident non-small cell lung cancer (NSCLC) in United States healthcare claims data</p>
DATASET - On the Investigation of Empirical Contradictions - Aggregated Results of Local Studies on Readability and Comprehensibility of Source Code
<p>Study package containing raw and analyzed data from the work entitled "On the Investigation of Empirical Contradictions - Aggregated Results of Local Studies on Readability and Comprehensibility of Source Code".</p> <p>The package comprises: 1) a summary of the information extracted from all papers mentioned in the Background; 2) the source code snippets used in the three studies; 3) the consent and characterization forms distributed to the participants; 4) the raw data, the aggregated data and other material generated with the collected data.</p>
An Exploratory Study on the Usage and Readability of Messages Within Assertion Methods of Test Cases
<p>This is the code and dataset that accompanies the study: "<strong>An Exploratory Study on the Usage and Readability of Messages Within Assertion Methods of Test Cases</strong>." This study has been accepted for publication at the 2023 International Workshop on Natural Language-based Software Engineering.</p> <p><strong><em>Following is the abstract of the study:</em></strong></p> <p>Unit testing is a vital part of the software development process and involves developers writing code to verify or assert production code. Furthermore, to help comprehend the test case and troubleshoot issues, developers have the option to provide a message that explains the reason for the assertion failure. In this exploratory empirical study, we examine the characteristics of assertion messages contained in the test methods in 20 open-source Java systems. Our findings show that while developers rarely utilize the option of supplying a message, those who do, either compose it of only string literals, identifiers, or a combination of both types. Using standard English readability measuring techniques, we observe that a beginner's knowledge of English is required to understand messages containing only identifiers, while a 4th-grade education level is required to understand messages composed of string literals. We also discuss shortcomings with using such readability measuring techniques and common anti-patterns in assert message construction. We envision our results incorporated into code quality tools that appraise the understandability of assertion messages. </p>
The 153 readable DNA sequences of a diet research on Eurasian otter of Kinmen Island
Open the record for dataset details and reuse information.
2014 SI defining constants converted to machine-readable XML format
<p>This data set provides 2014 CODATA defining constants using the XML scheme based on the Digital-SI (D-SI) data model that enables a machine-readable transmission of metrological data in digital applications.</p>
2018 CODATA values converted to the machine-readable D-SI XML format
<p>This data set provides 2018 CODATA values for fundamental physical constants using the XML scheme based on the Digital-SI (D-SI) data model that enables a machine-readable transmission of metrological data in digital applications.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.