Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
46
datasets available to search
ShareScore release 0.9.0
Dataset results
46 results for “readability”
2018 SI defining constants converted to machine-readable XML format
<p>This data set provides 2018 CODATA defining constants using the XML scheme based on the Digital-SI (D-SI) data model that enables a machine-readable transmission of metrological data in digital applications.</p>
SmartCom XML Scheme Definition (XSD) for a Machine-readable Communication of Fundamental Constants in Metrology
<p>The XSD defines a data structure for an unambiguous, easy-to-use, safe and uniform exchange of fundamental physical constants and mathematical constants in a machine-readable data format. The data elements for constants are defined in the Digital System of Units (D-SI) metadata model from the EMPIR project 17IND02 SmartCom. The development of the data structure was based on an own analysis of minimum requirement for transfering the data of fundamental physical constants "as are" provided by CODATA into a machine-readable form. Furthermore, the considerations for the transformation comprise traceability to the original CODATA values.</p> <p>The availability of machine-readable data of fundamental physical constants that can easy be accessed by software and that are traceable to a international accepted definition is essential for various areas of metrology where the International System of Units (SI) is of key-importance.</p>
A machine readable collection of lexical data on the Burmish languages
<p>This dataset includes lexical lists, mostly semantically normalized to WordNet, that brings together with near comprehensiveness the published data on Burmish languages. It was compiled by the ERC project 'Asia: Beyond Boundaries' in collaboration with the Center for Research in Computational Linguistics.</p> <p>The Bibtex references can be found at this Zenodo deposit--</p> <p>Hill, Nathan, List, Johann-Mattis, & Gong, Xun. (2020). Asia.bib: A bibtex bibliography for Asian Historical Linguistics [Data set]. Zenodo. http://doi.org/10.5281/zenodo.3759114</p>
Enhancing Code Readability through Automated Consistent Formatting
<p>The provided dataset contains the data used by "Enhancing Code Readability through Automated Consistent Formatting", in order to train a personalized code formatter based on the styling used by a team of developers in a set of code files.</p>
Replication Package of "Developer-Centric Code Readability Assessment: Are We There Yet?"
<p>This repository contains the datasets and the scripts to replicate our work "Developer-Centric Code Readability Assessment: Are We There Yet?".</p>
Impacts of Coding Practices on Readability - Dataset
<p>Data gathered and analyzed as described on the paper "Impacts of Coding Practices on Readability", presented at the ICPC 2018</p>
Domain expert readability dataset
<p>Judgments gathered from 10 experts through a web-based survey on the readability of publication abstracts. The abstracts used were a subset of the AMiner's DBLP citation nework v10 dataset (<a href="https://aminer.org/citation">https://aminer.org/citation</a>) in the discipline of data and knowledge management. In particular, abstracts containing the following keywords were used: "database", "machine learning", "information retrieval", "data management", "cloud computing", "data mining", "algorithms", "classification", "query processing", "networks", "indexing", "distributed systems".</p> <p>After reading the abstract, each expert had to answer the following questions on a 5 point scale.</p> <ul> <li>Q1: Please rate how well-written the abstract is.</li> <li>Q2: Does the abstract contain linguistic errors?</li> <li>Q3: Please rate how clear the contribution of the paper is (based on the abstract).</li> </ul> <p>For each question, the interpretation of the extreme scale values (i.e., 1 and 5) were provided. In particular, 1 = “very poorly written” / “so many ling. errors that make abstract incomprehensible” / “not clear at all” (Q1/Q2/Q3) and 5 = “excellently written” / “no errors” / “completely clear” (Q1/Q2/Q3).</p> <p>The pairwise correlations (Kendall’s τ) of expert judgments on questions Q1-Q3 are presented in this <a href="http://andrea.imis.athena-innovation.gr/readability/table6.png">table</a>.</p> <p>The contained dataset is a tsv file that includes the following fields:</p> <ul> <li>user_id: expert identifier</li> <li>paper_id: AMiner's identifier from DBLP citation nework v10 dataset</li> <li>rating_1: answer for Q1</li> <li>rating_2: answer for Q2</li> <li>rating_3: answer fro Q3 </li> </ul> <p> </p> <p><strong>Please cite:</strong><br> Thanasis Vergoulis, Ilias Kanellos, Anargiros Tzerefos, Serafeim Chatzopoulos, Theodore Dalamagas, Spiros Skiadopoulos. A study on the readability of scientific publications. <em>23<sup>rd</sup> International Conference on Theory and Practice of Digital Libraries</em>. Oslo, Norway 2019 (to appear)</p>
Artifact of An LLM-based Readability Measurement for Unit Tests' Context-aware Inputs
<div> <p><strong>Artifact of An LLM-based Readability Measurement for Unit Tests' Context-aware Inputs</strong></p> <p> </p> <p>A README.md file can be found in the base folder after unzipping the artifact.</p> <p> </p> </div>
Assessing Code Readability in Python Programming Courses Using Eye-Tracking - Python Code Snippets
<p>Python code snippets for assessing code readability in Python programming courses using eye-tracking.</p>
The code and data for paper "Reassessing Java Code Readability Models with a Human-Centered Approach"
<p>The code and data for paper "Reassessing Java Code Readability Models with a Human-Centered Approach".<br> Please refer to README.md for detailed instructions.</p>
Consent Forms in Cancer Research: Examining the Effect of Length on Readability
ClinicalTrials.gov study NCT04548063. IPD Sharing: NO. Countries: 1. Publications: 2.
Impact of Readability Improvements on the Comprehension of Written Information Given to Clinical Trial Patients: a RCT
ClinicalTrials.gov study NCT00908557. IPD Sharing: Not stated. Countries: 1. Publications: 5.
parliamentr: speeches from european parliaments in a standardized, machine-readable format
<p>Data accompagnying the R package parliamentr. Here on zenodo is the dataset, and the R package holds the code used to scrape & clean the data, as well as code to download and use this dataset.</p> <p>Currently under development </p>
FWP Life History Project in the American South: Machine Readable Text and Metadata
<p>The data is created from documents in the Federal Writers’ Project (FWP) Papers, 1936-1940 held at The Southern Historical Collection at the Louis Round Wilson Special Collections Library at the University of North Carolina – Chapel Hill. From 1936-1939, writers were sent across the American South to record life histories as a part of the Southern Life History Project. Previously only PDFs, each life history has been transformed into a .txt file. The csv files include metadata about the life histories such as writer name along with their race and gender, interviewee name along with their race and gender, reviser along with race and gender, location of the interview, and year. The data was designed for a historical research project, so scholars made decisions about race and gender guided by their expertise on the era and the analytical questions driving the project. It should be noted that labeling people by race and gender is a complicated process and practice of power, so care should be taken when using these categories. The data only includes the life histories held at UNC-CH, so it is not a complete collection of all life histories conducted in the region. The data was created for the Photogrammar project with funding from an American Council of Learned Societies (ACLS) Digital Extension Grant. As a condition of the funding, the data is made available under a GNU Public License (GPL).</p> <p> </p>
On the Importance and Shortcomings of Code Readability Metrics: A Case Study on Reactive Programming - replication package
<p>This is the replication package for the conference paper submission "On the Importance and Shortcomings of Code Readability Metrics: A Case Study on Reactive Programming"</p> <p><strong>Contents:</strong></p> <ul> <li>measurements.zip <ul> <li>DATASET_ORIGINAL.csv</li> <li>DATASET_REACTIVE.csv</li> </ul> </li> <li>source_code.zip <ul> <li> source_code_orig <ul> <li>Client.java</li> <li>Connection.java</li> <li>Server.java</li> <li>TcpConnection.java</li> <li>UdpConnection.java</li> </ul> </li> <li> source_code_rx <ul> <li>Client.java</li> <li>Connection.java</li> <li>Server.java</li> <li>TcpConnection.java</li> <li>UdpConnection.java</li> </ul> </li> </ul> </li> </ul> <p> </p>
Data from: Accuracy and readability of cardiovascular entries on Wikipedia: are they reliable learning resources for medical students?
Objective: To evaluate accuracy of content and readability level of English Wikipedia articles on cardiovascular diseases, using quality and readability tools. Methods: Wikipedia was searched on the 6 October 2013 for articles on cardiovascular diseases. Using a modified DISCERN (DISCERN is an instrument widely used in assessing online resources), articles were independently scored by three assessors. The readability was calculated using Flesch-Kincaid Grade Level. The inter-rater agreement between evaluators was calculated using the Fleiss κ scale. Results: This study was based on 47 English Wikipedia entries on cardiovascular diseases. The DISCERN scores had a median=33 (IQR=6). Four articles (8.5%) were of good quality (DISCERN score 40–50), 39 (83%) moderate (DISCERN 30–39) and 4 (8.5%) were poor (DISCERN 10–29). Although the entries covered the aetiology and the clinical picture, there were deficiencies in the pathophysiology of diseases, signs and symptoms, diagnostic approaches and treatment. The number of references varied from 1 to 127 references; 25.9±29.4 (mean±SD). Several problems were identified in the list of references and citations made in the articles. The readability of articles was 14.3±1.7 (mean±SD); consistent with the readability level for college students. In comparison, Harrison's Principles of Internal Medicine 18th edition had more tables, less references and no significant difference in number of graphs, images, illustrations or readability level. The overall agreement between the evaluators was good (Fleiss κ 0.718 (95% CI 0.57 to 0.83). Conclusions: The Wikipedia entries are not aimed at a medical audience and should not be used as a substitute to recommended medical resources. Course designers and students should be aware that Wikipedia entries on cardiovascular diseases lack accuracy, predominantly due to errors of omission. Further improvement of the Wikipedia content of cardiovascular entries would be needed before they could be considered a supplementary resource.
The code and data for paper "Reassessing Java Code Readability Models with a Human-Centered Approach"
<p>The code and data for paper "Reassessing Java Code Readability Models with a Human-Centered Approach".<br> Please refer to README.md for detailed instructions.</p>
Data from: Accuracy and readability of cardiovascular entries on Wikipedia: are they reliable learning resources for medical students?
Open the record for dataset details and reuse information.
Machine readable code lists for an algorithm to identify incident lung cancer in United States healthcare claims data
<p>Machine readable code lists for an algorithm to identify incident lung cancer in United States healthcare claims data</p>
Machine readable collection of rhyme judgements from Pre-Qin and Han dynasty Chinese texts
<p>This is a machine readable collection of rhyme judgements of Pre-Qin and Han dynasty Chinese texts, prepared in the style recommended by Mattis List and Chris Foster <br><br></p> <div>List, Johann-Mattis, Hill, Nathan W. and Foster, Christopher J.. "Towards a standardized annotation of rhyme judgments in Chinese historical phonology (and beyond)" <em>Journal of Language Relationship</em>, vol. 17, no. 1-2, 2019, pp. 26-43. <a href="https://doi.org/10.31826/jlr-2019-171-207">https://doi.org/10.31826/jlr-2019-171-207</a><br><br>The data is taken from the website of Professor Suzuki Shingo<br><br>https://suzukish.sakura.ne.jp/search/xianqin/index.php<br><br></div>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.