Skip to main content
zenodoopen

The Collaborative Organization of Knowledge: Data Set

<p>Wikipedia is an ongoing endeavor to create a free encyclopedia through an open computer-mediated collaborative effort. How does Wikipedia grow and maintain its coverage? This page contains supporing material relevant to a publication that examines this question.</p> <ul> <li>Diomidis Spinellis and Panagiotis Louridas. The collaborative organization of knowledge. Communications of the ACM, 51(8):68&ndash;73, August 2008. (<a href="http://dx.doi.org/10.1145/1378704.1378720">doi:10.1145/1378704.1378720</a>)</li> </ul> <p>In the above paper, a longitudinal study of Wikipedia&#39;s evolution shows that although Wikipedia&#39;s scope is increasing, its coverage is not deteriorating. This can be explained by the fact that referring to an non-existing entry typically leads to the establishment of an article for it. Wikipedia&#39;s evolution also demonstrates the creation of a large real world scale-free graph through a combination of incremental growth and preferential attachment.</p> <p>Though this data set you can download the processed results. The file starts with a header giving various attributes of the processed data set.</p> <pre>% Number of bins: 72 % Total revisions: 28247658 % Maximum revisions: 28273 (George W. Bush) % Maximum reverts: 9218 (George W. Bush) % Number of moves: 81380 % Total pages: 1898139 % Revisions from IP addresses: 8518913 % Total contributors: 230130 % Maximum different contributors: 2539 (George W. Bush) % Redirected pages: 631567 % Restricted pages: 2441 % Maximum number of contained references: 17577 (List of all three letter acrony ms) % Pages with at least one revert: 211704 % Total number of reverts across all pages: 1147151 % Total time between reverts: 54524346346 % Moved pages: 80332 </pre> <p>Next comes one line of data for each one of Wikipedia&#39;s entries. Here is an example.</p> <pre>A (musical note):1128386876:Mailer diablo:1130566991:MrD9:10:7:18:0:0:0:0:0:0:0: 0:0:0:0:0:0:0:0:0:0:0:0:0:0:0:0:0:0:0:0:0:0:0:0:0:0:0:0:0:0:0:0:0:0:0:0:0:0:0:0: 0:0:0:0:0:0:0:0:0:0:0:1:1:1:2:2:2:2:2:2:2:2:2:2:2:2:E </pre> <p>Each line contains the following fields.</p> <ul> <li>Entry name</li> <li>Time of first definition (in seconds since Unix epoch)</li> <li>Name of the contributor who first defined the entry</li> <li>Time of first reference (in seconds since Unix epoch)</li> <li>Name of the contributor who first referenced the entry</li> <li>Number of references</li> <li>Number of contributors</li> <li>Number of revisions</li> <li>Number of reverts</li> <li>For each one of the time period bins (72 in this file) the number of references to the entry</li> <li>The letter &quot;E&quot;</li> </ul> <p>The fields are colon-separated. Colons in the input data are converted to an underscore.</p> <p>Finally, come lines summarizing the data set&#39;s characteristics for each time period. Here is an example.</p> <pre>2001-07-01 4851 0 27106 15129 13458 531 </pre> <p>Each line contains the following fields.</p> <ul> <li>Start date of this period</li> <li>Number of entries</li> <li>Number of entries that are stubs</li> <li>Number of references</li> <li>Number of referenced articles</li> <li>Number of undefined entries</li> <li>Number of active contributors in this period</li> </ul>

ShareScore

40/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
8
Harmonization
4
Access
16
Reuse readiness
8
Engagement
4

Topics