Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

46

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

46 results for “readability”

Learn how ShareScore rates datasets ↗
zenodo32/100

2018 SI defining constants converted to machine-readable XML format

<p>This data set provides 2018 CODATA defining constants using the XML scheme based on the Digital-SI (D-SI) data model that enables a machine-readable transmission of metrological data in digital applications.</p>

openother-atFeb 2020View details →
zenodo32/100

SmartCom XML Scheme Definition (XSD) for a Machine-readable Communication of Fundamental Constants in Metrology

<p>The XSD defines a data structure for an unambiguous, easy-to-use, safe and uniform exchange of fundamental physical constants and mathematical constants in a machine-readable data format. The data elements for constants are defined in the Digital System of Units (D-SI) metadata model from the EMPIR project 17IND02 SmartCom. The development of the data structure was based on an own analysis of minimum requirement for transfering the data of fundamental physical constants &quot;as are&quot; provided by CODATA into a machine-readable form. Furthermore, the considerations for the transformation comprise traceability to the original CODATA values.</p> <p>The availability of machine-readable data of fundamental physical constants that can easy be accessed by software and that are traceable to a international accepted definition is essential for various areas of metrology where the International System of Units (SI) is of key-importance.</p>

openother-atFeb 2020View details →
zenodo32/100

A machine readable collection of lexical data on the Burmish languages

<p>This dataset includes lexical lists, mostly semantically normalized to WordNet, that brings together with near comprehensiveness the published data on Burmish languages. It was compiled by the ERC project &#39;Asia: Beyond Boundaries&#39; in collaboration with the Center for Research in Computational Linguistics.</p> <p>The Bibtex references can be found at this Zenodo deposit--</p> <p>Hill, Nathan, List, Johann-Mattis, &amp; Gong, Xun. (2020). Asia.bib: A bibtex bibliography for Asian Historical Linguistics [Data set]. Zenodo. http://doi.org/10.5281/zenodo.3759114</p>

opencc-by-4.0Apr 2020View details →
zenodo32/100

Enhancing Code Readability through Automated Consistent Formatting

<p>The provided dataset contains the data used by "Enhancing Code Readability through Automated Consistent Formatting", in order to train a personalized code formatter based on the styling used by a team of developers in a set of code files.</p>

opencc-by-4.0Apr 2024View details →
zenodo32/100

Replication Package of "Developer-Centric Code Readability Assessment: Are We There Yet?"

<p>This repository contains the datasets and the scripts to replicate our work "Developer-Centric Code Readability Assessment: Are We There Yet?".</p>

opencc-by-4.0Nov 2024View details →
zenodo32/100

Impacts of Coding Practices on Readability - Dataset

<p>Data gathered and analyzed as described on the paper &quot;Impacts of Coding Practices on Readability&quot;, presented at the ICPC 2018</p>

opencc-by-4.0Mar 2018View details →
zenodo32/100

Domain expert readability dataset

<p>Judgments gathered from 10 experts through a web-based survey on the readability of publication abstracts. The abstracts used were a subset of the AMiner&#39;s DBLP citation nework v10 dataset (<a href="https://aminer.org/citation">https://aminer.org/citation</a>) in the discipline of data and knowledge management. In particular, abstracts containing the following keywords were used: &quot;database&quot;, &quot;machine learning&quot;, &quot;information retrieval&quot;, &quot;data management&quot;, &quot;cloud computing&quot;, &quot;data mining&quot;, &quot;algorithms&quot;, &quot;classification&quot;, &quot;query processing&quot;, &quot;networks&quot;, &quot;indexing&quot;, &quot;distributed systems&quot;.</p> <p>After reading the abstract, each expert had to answer&nbsp;the following questions on a 5 point scale.</p> <ul> <li>Q1:&nbsp;Please rate how well-written the abstract is.</li> <li>Q2: Does the abstract contain linguistic errors?</li> <li>Q3: Please rate how clear the contribution of the paper is (based on the abstract).</li> </ul> <p>For each question, the interpretation of the extreme scale values (i.e., 1 and 5) were&nbsp;provided. In particular, 1 = &ldquo;very poorly written&rdquo; / &ldquo;so many ling. errors that make abstract incomprehensible&rdquo; / &ldquo;not clear at all&rdquo; (Q1/Q2/Q3) and 5 = &ldquo;excellently written&rdquo; / &ldquo;no errors&rdquo; / &ldquo;completely clear&rdquo; (Q1/Q2/Q3).</p> <p>The pairwise correlations (Kendall&rsquo;s &tau;) of expert judgments on questions Q1-Q3 are presented in this&nbsp;<a href="http://andrea.imis.athena-innovation.gr/readability/table6.png">table</a>.</p> <p>The contained dataset is a tsv file that includes the following fields:</p> <ul> <li>user_id:&nbsp;expert identifier</li> <li>paper_id:&nbsp;AMiner&#39;s identifier from DBLP citation nework v10 dataset</li> <li>rating_1:&nbsp;answer for Q1</li> <li>rating_2:&nbsp;answer for Q2</li> <li>rating_3:&nbsp;answer fro Q3&nbsp;</li> </ul> <p>&nbsp;</p> <p><strong>Please cite:</strong><br> Thanasis Vergoulis, Ilias Kanellos, Anargiros Tzerefos, Serafeim Chatzopoulos, Theodore Dalamagas, Spiros Skiadopoulos.&nbsp;A study on the readability of scientific publications.&nbsp;<em>23<sup>rd</sup> International Conference on Theory and Practice of Digital Libraries</em>. Oslo, Norway 2019 (to appear)</p>

opencc-by-4.0Apr 2019View details →
zenodo32/100

Artifact of An LLM-based Readability Measurement for Unit Tests' Context-aware Inputs

<div> <p><strong>Artifact of An LLM-based Readability Measurement for Unit Tests' Context-aware Inputs</strong></p> <p>&nbsp;</p> <p>A README.md file can be found in the base folder after unzipping the artifact.</p> <p>&nbsp;</p> </div>

opencc-by-4.0Aug 2024View details →
zenodo32/100

Assessing Code Readability in Python Programming Courses Using Eye-Tracking - Python Code Snippets

<p>Python code snippets for assessing code readability in Python programming courses using eye-tracking.</p>

opencc-by-4.0Jun 2023View details →
zenodo32/100

The code and data for paper "Reassessing Java Code Readability Models with a Human-Centered Approach"

<p>The code and data for paper &quot;Reassessing Java Code Readability Models with a Human-Centered Approach&quot;.<br> Please refer to README.md for detailed instructions.</p>

opencc-by-4.0Sep 2023View details →
ClinicalTrials.gov32/100

Consent Forms in Cancer Research: Examining the Effect of Length on Readability

ClinicalTrials.gov study NCT04548063. IPD Sharing: NO. Countries: 1. Publications: 2.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov32/100

Impact of Readability Improvements on the Comprehension of Written Information Given to Clinical Trial Patients: a RCT

ClinicalTrials.gov study NCT00908557. IPD Sharing: Not stated. Countries: 1. Publications: 5.

restrictedIPD-UNDECIDEDFeb 2026View details →
zenodo28/100

parliamentr: speeches from european parliaments in a standardized, machine-readable format

<p>Data accompagnying the R package parliamentr. Here on zenodo is the dataset, and the R package holds the code used to scrape &amp; clean the data, as well as code to download and use this dataset.</p> <p>Currently under development&nbsp;</p>

opencc-by-sa-4.0Mar 2020View details →
zenodo28/100

FWP Life History Project in the American South: Machine Readable Text and Metadata

<p>The data is created from documents in the Federal Writers&rsquo; Project (FWP) Papers, 1936-1940 held at The Southern&nbsp;Historical Collection at the Louis Round Wilson Special Collections Library at the University of North&nbsp;Carolina &ndash; Chapel Hill. From 1936-1939, writers were sent across the American South to record life&nbsp;histories as a part of the Southern Life History Project. Previously only PDFs, each life history has been&nbsp;transformed into a .txt file. The csv files include metadata about the life histories such as &nbsp;writer name&nbsp;along with their race and gender, interviewee name along with their race and gender, reviser along with race&nbsp;and gender, location of the interview, and year. The data was designed for a historical research project,&nbsp;so scholars made decisions about race and gender guided by their expertise on the era and the analytical&nbsp;questions driving the project. It should be noted that labeling people by race and gender is a complicated&nbsp;process and practice of power, so care should be taken when using these categories. The data only includes&nbsp;the life histories held at UNC-CH, so it is not a complete collection of all life histories conducted in&nbsp;the region. The data was created for the Photogrammar project with funding from an American Council of&nbsp;Learned Societies (ACLS) Digital Extension Grant. As a condition of the funding, the data is made&nbsp;available under a GNU Public License (GPL).</p> <p>&nbsp;</p>

openMay 2020View details →
zenodo28/100

On the Importance and Shortcomings of Code Readability Metrics: A Case Study on Reactive Programming - replication package

<p>This is the replication package for the conference paper submission &quot;On the Importance and Shortcomings of Code Readability Metrics: A Case Study on Reactive Programming&quot;</p> <p><strong>Contents:</strong></p> <ul> <li>measurements.zip <ul> <li>DATASET_ORIGINAL.csv</li> <li>DATASET_REACTIVE.csv</li> </ul> </li> <li>source_code.zip <ul> <li>&nbsp;source_code_orig <ul> <li>Client.java</li> <li>Connection.java</li> <li>Server.java</li> <li>TcpConnection.java</li> <li>UdpConnection.java</li> </ul> </li> <li>&nbsp;source_code_rx <ul> <li>Client.java</li> <li>Connection.java</li> <li>Server.java</li> <li>TcpConnection.java</li> <li>UdpConnection.java</li> </ul> </li> </ul> </li> </ul> <p>&nbsp;&nbsp;&nbsp;&nbsp;</p>

opencc-by-4.0Nov 2020View details →
dryad28/100

Data from: Accuracy and readability of cardiovascular entries on Wikipedia: are they reliable learning resources for medical students?

Objective: To evaluate accuracy of content and readability level of English Wikipedia articles on cardiovascular diseases, using quality and readability tools. Methods: Wikipedia was searched on the 6 October 2013 for articles on cardiovascular diseases. Using a modified DISCERN (DISCERN is an instrument widely used in assessing online resources), articles were independently scored by three assessors. The readability was calculated using Flesch-Kincaid Grade Level. The inter-rater agreement between evaluators was calculated using the Fleiss κ scale. Results: This study was based on 47 English Wikipedia entries on cardiovascular diseases. The DISCERN scores had a median=33 (IQR=6). Four articles (8.5%) were of good quality (DISCERN score 40–50), 39 (83%) moderate (DISCERN 30–39) and 4 (8.5%) were poor (DISCERN 10–29). Although the entries covered the aetiology and the clinical picture, there were deficiencies in the pathophysiology of diseases, signs and symptoms, diagnostic approaches and treatment. The number of references varied from 1 to 127 references; 25.9±29.4 (mean±SD). Several problems were identified in the list of references and citations made in the articles. The readability of articles was 14.3±1.7 (mean±SD); consistent with the readability level for college students. In comparison, Harrison's Principles of Internal Medicine 18th edition had more tables, less references and no significant difference in number of graphs, images, illustrations or readability level. The overall agreement between the evaluators was good (Fleiss κ 0.718 (95% CI 0.57 to 0.83). Conclusions: The Wikipedia entries are not aimed at a medical audience and should not be used as a substitute to recommended medical resources. Course designers and students should be aware that Wikipedia entries on cardiovascular diseases lack accuracy, predominantly due to errors of omission. Further improvement of the Wikipedia content of cardiovascular entries would be needed before they could be considered a supplementary resource.

opencc-zeroDec 2014View details →
zenodo28/100

The code and data for paper "Reassessing Java Code Readability Models with a Human-Centered Approach"

<p>The code and data for paper &quot;Reassessing Java Code Readability Models with a Human-Centered Approach&quot;.<br> Please refer to README.md for detailed instructions.</p>

opencc-by-4.0Sep 2023View details →
dryad28/100

Data from: Accuracy and readability of cardiovascular entries on Wikipedia: are they reliable learning resources for medical students?

Open the record for dataset details and reuse information.

publicSep 2015View details →
zenodo24/100

Machine readable code lists for an algorithm to identify incident lung cancer in United States healthcare claims data

<p>Machine readable code lists for an algorithm to identify incident lung cancer in United States healthcare claims data</p>

opencc-by-4.0Mar 2020View details →
zenodo24/100

Machine readable collection of rhyme judgements from Pre-Qin and Han dynasty Chinese texts

<p>This is a machine readable collection of rhyme judgements of Pre-Qin and Han dynasty Chinese texts, prepared in the style recommended by Mattis List and Chris Foster&nbsp;<br><br></p> <div>List, Johann-Mattis, Hill, Nathan W. and Foster, Christopher J.. "Towards a standardized annotation of rhyme judgments in Chinese historical phonology (and beyond)" <em>Journal of Language Relationship</em>, vol. 17, no. 1-2, 2019, pp. 26-43. <a href="https://doi.org/10.31826/jlr-2019-171-207">https://doi.org/10.31826/jlr-2019-171-207</a><br><br>The data is taken from the website of Professor Suzuki Shingo<br><br>https://suzukish.sakura.ne.jp/search/xianqin/index.php<br><br></div>

restrictedcc-zeroAug 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record