Skip to main content
zenodoopen

Luminoso Input Data for SemEval-2018 Task 10: "Capturing Discriminative Attributes"

<p>This is the data required to run Luminoso&#39;s entry to the SemEval-2018 task on Capturing Discriminative Attributes.</p> <p>This data includes:</p> <ul> <li>A recently-computed version of the <a href="https://github.com/commonsense/conceptnet-numberbatch">ConceptNet Numberbatch</a> word embeddings</li> <li>The&nbsp;output of an implementation of Semantic Matching Energy over ConceptNet</li> <li>A SQLite database containing the lead section of all articles on the <a href="http://en.wikipedia.org">English Wikipedia</a> on 2017-12-20</li> <li>The text file that that database is constructed from</li> <li>A SQLite database of words that co-occur in <a href="http://storage.googleapis.com/books/ngrams/books/datasetsv2.html">Google Books 2-grams</a></li> <li>The text file containing total counts of 2-grams in the Google Books data, which that database is constructed from</li> </ul> <p>For more information, see the paper &quot;Luminoso at SemEval-2018 Task 10: Distinguishing Attributes Using Text Corpora and Relational Knowledge&quot;, by Robyn&nbsp;Speer and Joanna Lowry-Duda, to appear in the proceedings of the SemEval workshop at&nbsp;NAACL 2018.</p>

ShareScore

36/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
8
Harmonization
4
Access
16
Reuse readiness
8
Engagement
0

Topics