zenodoopen
Processed Wikipedia Dataset
<p>We extract a subset of about 1,000,000 documents of Wikipedia 2020 and extract the keywords of them. The <code>wiki_kws_dict.pkl</code> is a map which maps each keyword to its total counts in files and query trend. The <code>wiki_doc_0.pkl</code> contains lists of keywords of each document. These two datasets can be loaded by the pickle package with python.</p>
ShareScore
40/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 8
- Harmonization
- 8
- Access
- 16
- Reuse readiness
- 8
- Engagement
- 0