Skip to main content
zenodoopen

Lists of Uzbek Stopwords

<p>The dataset presents 3 lists of stopwords in the Uzbek language. The lists were constructed using three automatic methods applied to the same corpus.&nbsp;</p> <p>The corpus was&nbsp;constructed by obtaining a source of 25 school textbooks, it was named &quot;School corpus&quot;. The total number of words in the &quot;School corpus&quot; is 731156. And them 47165 words are unique words.&nbsp;The corpus can be re-constructed using the list of urls of all files comprised in the corpus. The list is part of the dataset (list_of_urls_of_school_corpus.txt).</p> <p>Description of the methods and the lists:</p> <p>A set of grammar rules and the TDIDF algorithm was used to automatically collect a list of single-word stopwords. 2357 stopwords were collected. The name of the file: stopwords_unigrams.txt.</p> <p>A bigram method was used to extract a list of 4548 bigrams (pairs) of stopwords. The name of the file: stopwords_bigrams.txt.</p> <p>A set of two-word collocations as stopwords was also extracted. The list has 24490 pairs of stopwords. The name of the file: stopwords_bigrams_with_collocations.txt.</p>

ShareScore

44/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
8
Harmonization
4
Access
16
Reuse readiness
8
Engagement
8

Topics