Skip to main content
zenodoopen

Early Irish Analogy Dataset for Word Embedding Evaluation

<p>An embedding evaluation dataset for Early Irish described in the paper "<a href="https://aclanthology.org/2023.insights-1.10.pdf">Do not Trust the Experts: How the Lack of Standard Complicates <span>NLP</span> for Historical <span>I</span>rish</a>".</p> <p>Traditionally, analogy datasets are based on pairwise semantic proportion, and therefore every question has a single correct answer. Given the high level of variation in historical languages, such a strict definition of a correct answer seems unjustified. Therefore, Early Irish Analogy Dataset follows the <a href="https://vecto.space/projects/BATS/">Bigger Analogy Test Set (BATS)</a> and provides several correct answers to each analogy question.&nbsp;</p> <p>Morphological and spelling variation data are extracted from the <a href="https://dil.ie/">eDIL</a>, a historical dictionary of medieval Irish. Unlike BATS, no distinction is made between inflection types due to eDIL's structure. The raw data amounted to 2,370 spelling variation and 9,690 morphological variation questions, from which 150 examples were randomly selected for each of the subsets to be comparable in size with the synonym and antonym subsets. The synonym and antonym subsets are translations of the correspondent BATS parts obtained by reverse-searching the eDIL and proofread by four expert evaluators. The dataset includes 98 entries in the synonym subset and 109 entries in the antonym subset, upon which three or more experts agreed.</p>

ShareScore

40/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
12
Harmonization
8
Access
16
Reuse readiness
0
Engagement
4

Topics