Domain expert readability dataset
<p>Judgments gathered from 10 experts through a web-based survey on the readability of publication abstracts. The abstracts used were a subset of the AMiner's DBLP citation nework v10 dataset (<a href="https://aminer.org/citation">https://aminer.org/citation</a>) in the discipline of data and knowledge management. In particular, abstracts containing the following keywords were used: "database", "machine learning", "information retrieval", "data management", "cloud computing", "data mining", "algorithms", "classification", "query processing", "networks", "indexing", "distributed systems".</p> <p>After reading the abstract, each expert had to answer the following questions on a 5 point scale.</p> <ul> <li>Q1: Please rate how well-written the abstract is.</li> <li>Q2: Does the abstract contain linguistic errors?</li> <li>Q3: Please rate how clear the contribution of the paper is (based on the abstract).</li> </ul> <p>For each question, the interpretation of the extreme scale values (i.e., 1 and 5) were provided. In particular, 1 = “very poorly written” / “so many ling. errors that make abstract incomprehensible” / “not clear at all” (Q1/Q2/Q3) and 5 = “excellently written” / “no errors” / “completely clear” (Q1/Q2/Q3).</p> <p>The pairwise correlations (Kendall’s τ) of expert judgments on questions Q1-Q3 are presented in this <a href="http://andrea.imis.athena-innovation.gr/readability/table6.png">table</a>.</p> <p>The contained dataset is a tsv file that includes the following fields:</p> <ul> <li>user_id: expert identifier</li> <li>paper_id: AMiner's identifier from DBLP citation nework v10 dataset</li> <li>rating_1: answer for Q1</li> <li>rating_2: answer for Q2</li> <li>rating_3: answer fro Q3 </li> </ul> <p> </p> <p><strong>Please cite:</strong><br> Thanasis Vergoulis, Ilias Kanellos, Anargiros Tzerefos, Serafeim Chatzopoulos, Theodore Dalamagas, Spiros Skiadopoulos. A study on the readability of scientific publications. <em>23<sup>rd</sup> International Conference on Theory and Practice of Digital Libraries</em>. Oslo, Norway 2019 (to appear)</p>
ShareScore
32/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 8
- Harmonization
- 4
- Access
- 16
- Reuse readiness
- 0
- Engagement
- 4