Skip to main content
zenodoopen

Hypernym-LIBre: A free Web-based Corpus for Hypernym Detection

<p>The task of finding hypernyms from large text corpora is a fundamental problem in NLP. It provides a basis for the main-stream natural language problems in AI. In our paper, we introduce a free new web-based corpus for hypernym detection and we show that using this corpus we achieve similar results to the state-of-the-art pattern-based methods achieved by a well known corpus that is not freely available.&nbsp; The dataset provided here is the one we use in our paper and we provide it with an open license so others can apply different methods and techniques for hypernym detection.</p> <p>&nbsp;</p> <p>The dataset is a combination of UMBC corpus and the Wikipedia corpus. Its dependency parsed and POS-tagged versions are available at this DOI:&nbsp; 10.5281/zenodo.3689303</p> <p>Contents:</p> <p>Hypernym-LIBre.zip&nbsp; 11.3GB compresssed, 32GB uncompressed raw text</p> <p>288 files of ~110 MB each</p> <p>&nbsp;</p> <p>10.5281/zenodo.3689303</p> <p>PoS and dep annotated</p> <p>~15GB compressed, 80GB uncompressed, 442 files of ~180MB each</p> <p>&nbsp;</p> <pre>10.5281/zenodo.3695237</pre> <p>hyponym-hypernym pairs extracted from Hypernym-LIBre using Hearst patterns</p>

ShareScore

36/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
8
Harmonization
8
Access
16
Reuse readiness
0
Engagement
4

Topics