Skip to main content
zenodoopen

Wikidata Lemmatization Dataset

<p>The Wikidata Lemmatization Dataset was collected using the following SPARQL query: https://w.wiki/9TwH</p> <p>Languages included in the dataset:</p> <ul> <li>Akkadian : AKK (Q35518)</li> <li>Arabic : AR (Q13955)</li> <li>Czech : CS (Q9056)</li> <li>German : DE (Q188)</li> <li>English : EN (Q1860)</li> <li>French : FR (Q150)</li> <li>Hebrew : HE (Q9288)</li> <li>Hittite : HIT (Q35668)</li> <li>Italian : IT (Q652)</li> <li>Russian : RU (Q7737)</li> <li>Sumerian : SUX (Q36790)</li> <li>Turkish : TR (Q256)</li> </ul> <p>The choice of languages to include have to do with a collection of primary and secondary source documents which we have digitized (OCR) and are using as references for the FactGrid Cuneiform project. The resulting lexemes for each language are shared in CSV with the file names references each language, their Wikidata Q-ids, the number of lexemes at that date, and the date of access (MM_YYYY).</p> <p>The format of each CSV includes the following fields:</p> <ol> <li>lexeme : the Wikidata lexeme id (L-id)</li> <li>lexemeLabel : the label assigned to the lexeme in Wikidata</li> <li>lexical_category : the Wikidata Q-item for the part of speech</li> <li>lexical_categoryLabel : the label assigned to the lexical category (e.g. noun, verb, adjective, etc.)</li> </ol> <p>This dataset will be updated periodically using standard version control.</p>

ShareScore

36/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
12
Harmonization
4
Access
16
Reuse readiness
0
Engagement
4

Topics