Skip to main content
zenodoopen

Wikipedia Citations: A comprehensive dataset of citations with identifiers extracted from English Wikipedia

<p>The dataset is composed of <strong>3 parts</strong>:</p> <p>1. &nbsp;The dataset of 29.276 million citations from &nbsp;35 different citation templates, &nbsp;out of which 3.92&nbsp;million citations already contained identifiers, and approximately 260,752 citations were equipped with identifiers from Crossref. This is under the filename: <strong>citations_from_wikipedia.zip</strong></p> <p>2. &nbsp;A minimal dataset containing a few of the columns from the citations from Wikipedia dataset. These columns are as follows:&nbsp;&#39;type_of_citation&#39;, &#39;page_title&#39;, &#39;Title&#39;, &#39;ID_list&#39;, metadata_file&#39;, &#39;updated_identifier&#39;.&nbsp;This is under the filename: <strong>minimal_dataset.zip. </strong>The &#39;metadata_file&#39; column can be used to refer to the metadata collected from CrossRef and page title, the title of the citation can be used to refer to the &#39;citations_from_wikipedia.zip&#39; dataset and get more information for a particular citation (such as author, periodical, chapter).</p> <p>3. &nbsp;Citations classified as a journal and their corresponding metadata/identifier extracted from Crossref to make the dataset more complete. This is under the filename: <strong>lookup_data.zip</strong>. This zip file contains a CSV file: <strong>lookup_table.gzip</strong> (a parquet file containing all citations classified as a journal) and a folder:<strong> metadata_extracted </strong>(a folder containing the metadata from CrossRef for all the citations mentioned in the table)</p> <p><br> The data was parsed from the Wikipedia XML content dumps published in May 2020.</p> <p>The source code to extract and getting used to the pipeline can be found here: <strong>https://github.com/Harshdeep1996/cite-classifications-wiki</strong></p> <p>The taxonomy of the dataset in (1) can be found here: <strong>https://github.com/Harshdeep1996/cite-classifications-wiki/wiki/Taxonomy-of-the-parent-dataset</strong></p>

ShareScore

32/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
8
Harmonization
4
Access
16
Reuse readiness
0
Engagement
4

Topics