Skip to main content
zenodoopen

Domain expert readability dataset

<p>Judgments gathered from 10 experts through a web-based survey on the readability of publication abstracts. The abstracts used were a subset of the AMiner&#39;s DBLP citation nework v10 dataset (<a href="https://aminer.org/citation">https://aminer.org/citation</a>) in the discipline of data and knowledge management. In particular, abstracts containing the following keywords were used: &quot;database&quot;, &quot;machine learning&quot;, &quot;information retrieval&quot;, &quot;data management&quot;, &quot;cloud computing&quot;, &quot;data mining&quot;, &quot;algorithms&quot;, &quot;classification&quot;, &quot;query processing&quot;, &quot;networks&quot;, &quot;indexing&quot;, &quot;distributed systems&quot;.</p> <p>After reading the abstract, each expert had to answer&nbsp;the following questions on a 5 point scale.</p> <ul> <li>Q1:&nbsp;Please rate how well-written the abstract is.</li> <li>Q2: Does the abstract contain linguistic errors?</li> <li>Q3: Please rate how clear the contribution of the paper is (based on the abstract).</li> </ul> <p>For each question, the interpretation of the extreme scale values (i.e., 1 and 5) were&nbsp;provided. In particular, 1 = &ldquo;very poorly written&rdquo; / &ldquo;so many ling. errors that make abstract incomprehensible&rdquo; / &ldquo;not clear at all&rdquo; (Q1/Q2/Q3) and 5 = &ldquo;excellently written&rdquo; / &ldquo;no errors&rdquo; / &ldquo;completely clear&rdquo; (Q1/Q2/Q3).</p> <p>The pairwise correlations (Kendall&rsquo;s &tau;) of expert judgments on questions Q1-Q3 are presented in this&nbsp;<a href="http://andrea.imis.athena-innovation.gr/readability/table6.png">table</a>.</p> <p>The contained dataset is a tsv file that includes the following fields:</p> <ul> <li>user_id:&nbsp;expert identifier</li> <li>paper_id:&nbsp;AMiner&#39;s identifier from DBLP citation nework v10 dataset</li> <li>rating_1:&nbsp;answer for Q1</li> <li>rating_2:&nbsp;answer for Q2</li> <li>rating_3:&nbsp;answer fro Q3&nbsp;</li> </ul> <p>&nbsp;</p> <p><strong>Please cite:</strong><br> Thanasis Vergoulis, Ilias Kanellos, Anargiros Tzerefos, Serafeim Chatzopoulos, Theodore Dalamagas, Spiros Skiadopoulos.&nbsp;A study on the readability of scientific publications.&nbsp;<em>23<sup>rd</sup> International Conference on Theory and Practice of Digital Libraries</em>. Oslo, Norway 2019 (to appear)</p>

ShareScore

32/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
8
Harmonization
4
Access
16
Reuse readiness
0
Engagement
4

Topics