Skip to main content
zenodoopen

Processed Wikipedia Dataset

<p>We extract a subset of about 1,000,000 documents of Wikipedia 2020 and extract the keywords of them. The <code>wiki_kws_dict.pkl</code> is a map which maps each keyword to its total counts in files and query trend. The&nbsp;<code>wiki_doc_0.pkl</code> contains lists of keywords of each document.&nbsp;These two datasets can be loaded by the pickle package with python.</p>

ShareScore

40/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
8
Harmonization
8
Access
16
Reuse readiness
8
Engagement
0