Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

219

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

219 results for “Wikipedia”

Learn how ShareScore rates datasets ↗
zenodo48/100

LaTeX formulae from English Wikipedia

<p>Public dump of LaTeX (texvc) input used in English Wikipedia</p> <p>Initially appeared in the public in the following form</p> <p>https://archive.softwareheritage.org/swh:1:cnt:da76ae5a988839894a2ecdab29a5ab7c8df7dc80;origin=https://github.com/wikimedia/mediawiki-services-texvcjs;visit=swh:1:snp:1ea38f5e961b25c6d0adde0145f64791eb8fb67d;anchor=swh:1:rev:101925e2ab1712c70472ed144efaea34743e6700;path=/test/en-wiki-formulae.json</p> <p>with the current 2024-11-23 results of the normalized output.</p> <p>The JSON files store the data in the following form</p> <pre><code>key=&gt;Tex</code></pre>

opencc-by-4.0Nov 2014View details →
zenodo48/100

OcWikiDisc: a Corpus of Wikipedia Talk Pages in Occitan

<p>OcWikiDisc is a freely available corpus in Occitan, extracted from the talk pages associated with the Occitan Wikipedia.</p> <p>The corpus contains messages posted by users in direct user-to-user interactions as part of the discussions about the content and the editing policies on Wikipedia. The messages are associated with metadata, such as the username, the date and time of the posting, the discussion title, etc. The corpus has also been annotated with tools for automatic language identification, allowing to filter out content in languages other than Occitan. Using different filtering strategies, four versions of the corpus are published (see documentation for more details). The version with the most restrictive filtering contains 8,000 messages for a total of 618,000 tokens, produced by 520 different users.</p>

opencc-by-sa-3.0Sep 2022View details →
zenodo48/100

Wikipedia time-series graph

<p>Wikipedia temporal graph.</p> <p>The dataset is based on two Wikipedia SQL dumps:&nbsp;(1) English language articles and (2) user visit counts per page per hour (aka pagecounts). The original datasets are publicly available on the Wikimedia website.</p> <p>Static graph structure is extracted from&nbsp;English language Wikipedia articles. Redirects are removed.&nbsp;Before building the Wikipedia graph we introduce thresholds on the minimum number of visits per hour and maximum in-degree. We remove the pages that have less than 500 visits per hour at least once during the specified period. Besides, we remove the nodes (pages) with in-degree higher than 8 000 to build a more meaningful initial graph. After cleaning, the graph contains 116 016 nodes (out of total 4 856 639 pages), 6 573 475 edges. The graph can be imported in two ways: (1) using edges.csv and vertices.csv&nbsp;or (2) using&nbsp;enwiki-20150403-graph.gt&nbsp;file&nbsp;that can be opened with open source Python library Graph-Tool.</p> <p>Time-series data contains&nbsp;users&#39; visit counts from 02:00, 23 September 2014 until 23:00, 30 April 2015. The total number of hours is&nbsp; 5278. The data is stored in two formats: CSV and H5. CSV file contains data in the following format [page_id :: count_views :: layer], where layer represents an hour. In H5 file, each layer corresponds to an hour as well.</p>

opencc-by-4.0Sep 2017View details →
zenodo48/100

Dateset: Capturing the influence of geopolitical ties from Wikipedia with reduced Google matrix

<p>This dataset provides complementary material to the scientific research presented in the paper &quot;<strong>Capturing the influence of geopolitical ties from Wikipedia with reduced Google matrix</strong>&quot;<strong>,</strong> accepted for publication in PLOS ONE under number PONE-D-18-07662R1.</p> <p>A draft version of the paper is available at <a href="https://arxiv.org/abs/1803.05336">https://arxiv.org/abs/1803.05336</a></p> <p>This paper presents two studies targeting two groups of countries:</p> <ul> <li>[40] the set 40 worldwide countries set ;</li> <li>[EU] the set of 27 European Union countries as of February 2013.</li> </ul> <p>Data is derived using Reduced Google matrix analysis on the Wikipedia English, Arabic, Russian, German and French editions collected in February 2013. Networks representing each edition are available here:</p> <p><a href="http://www.quantware.ups-tlse.fr/QWLIB/topwikipeople/index.html">http://www.quantware.ups-tlse.fr/QWLIB/topwikipeople/index.html</a></p> <p>The following files are given:</p> <ul> <li>[GRedured_40.zip] GReduced matrix and its decomposition for [40] countries set</li> <li>[GRedured_EU.zip] GReduced matrix and its decomposition for [EU] countries set</li> <li>[PageRank_vs_CheiRank.zip] PageRank versus CheiRank figures for RuWiki and ArWiki</li> <li>[Sensitivity.xlsx] and [Sensitivity_html.xlsx] Sensitivity values for both [EU] and [RU] in either .xlsx or in .html format</li> </ul>

opencc-by-4.0Jul 2018View details →
zenodo48/100

WikiLinkGraphs: A complete, longitudinal and multilanguage dataset of the Wikipedia link networks

<p>This dataset contains yearly snapshots of the Wikipedia&#39;s internal link network for the 9 largest language edition (de, en, es, fr, it, nl, pl, ru, sv). The dataset spans over 17 years, from the creation of Wikipedia in 2001 to March 2018. The snapshots are taken on March 1st of every year.</p> <p>The graphs include the links extract from the wikitext of each page (i.e in the form [[wikilink]]). Links transcluded from templates are not included. Redirects are resolved to their target page.</p> <p>More detailed information and supporting datasets are available at: http://disi.unitn.it/~consonni/datasets/.</p> <p><strong>IMPORTANT NOTICE</strong></p> <p>Gzipped files are compressed two times by Zenodo, the MD5 provided by Zenodo and the SHA512 sums provided in the `.sha512sums.txt` files, match with the files compressed once. In other words, when you download a `.gz` file save it as `.gz.gz`, uncompress it once and it should match both the MD5 provided by Zenodo and the SHA512 sum provided by us. We have opened a bug report for this behavior on Zenodo&#39;s repository at: https://github.com/zenodo/zenodo/issues/1705</p> <p>&nbsp;</p>

opencc-by-4.0Jan 2019View details →
zenodo48/100

Uncovering the Semantics of Wikipedia Categories - Axioms and Assertions

<p>Resulting axioms and assertions from applying the Cat2Ax approach to the DBpedia knowledge graph.<br> The methodology is described in the conference publication &quot;N. Heist, H. Paulheim: Uncovering the Semantics of Wikipedia Categories, International Semantic Web Conference, 2019&quot;.</p>

openmit-licenseOct 2019View details →
zenodo48/100

French Entity-Linking dataset between annotated tweets collected during major crises in France and French Wikipedia corpus

<p>Most of the available datasets are not particularly adapted to our target application: geolocate natural disasters from social networks. First, social media posts are largely underrepresented in these datasets, and the only Twitter dataset lacks Entity-Linking annotations. Second, none of the datasets focuses on a crisis or natural disaster event.</p> <p>To mitigate these issues, we extracted a collection of French tweets written during earthquakes and major floods that have occurred in France in recent years. We set up Label-Studio in order to annotate these tweets. A total of 4617 tweets were annotated, including 1678 tweets posted during earthquakes and 2939 during floods. For each annotated tweet, mentions were annotated using the set of labels described earlier in the paper as well as, when possible, the target Wikipedia title.</p> <p>Named &ldquo;R&eacute;SoCIO&rdquo; in reference to the research project in which it was carried out, the dataset resulting from this work contains a total of 12 828 annotated mentions and 1 513 distinct Wikipedia entities. 85% of mentions were associated with a Wikipedia page and 94 % if we ignore the RISKNAT and DAMAGES labels, which are often difficult to map to an existing entity.</p> <table> <tbody> <tr> <td><strong>Labels</strong></td> <td><strong>#Mentions</strong></td> <td><strong>#Linked</strong></td> <td><strong>#Entities</strong></td> </tr> <tr> <td>PERSON</td> <td>315</td> <td>263</td> <td>136</td> </tr> <tr> <td>ORG</td> <td>863</td> <td>790</td> <td>281</td> </tr> <tr> <td>GEOLOC</td> <td>4375</td> <td>4234</td> <td>701</td> </tr> <tr> <td>TRANSPORT</td> <td>250</td> <td>203</td> <td>101</td> </tr> <tr> <td>EVENT</td> <td>35</td> <td>21</td> <td>16</td> </tr> <tr> <td>FACILITY</td> <td>129</td> <td>94</td> <td>49</td> </tr> <tr> <td>RISKNAT</td> <td>5502</td> <td>4994</td> <td>128</td> </tr> <tr> <td>DAMAGES</td> <td>1136</td> <td>121</td> <td>56</td> </tr> <tr> <td>OTHER</td> <td>223</td> <td>200</td> <td>46</td> </tr> <tr> <td><strong>Total</strong></td> <td><strong>12828</strong></td> <td><strong>1322</strong></td> <td><strong>1513</strong></td> </tr> </tbody> </table> <p>Overview of the mentions annotated in the Twitter dataset. #Mentions&nbsp;shows the total number of mentions per label, #Linked the number of mentions linked&nbsp;to an entity and #Entities the number of distinct entities per label present in the&nbsp;dataset.</p> <table> <tbody> <tr> <td><strong>Labels</strong></td> <td><strong>#Mentions</strong></td> <td><strong>#Linked</strong></td> <td><strong>#Entitie</strong>s</td> </tr> <tr> <td>PERSON</td> <td>1100102</td> <td>1098406</td> <td>557697</td> </tr> <tr> <td>ORG</td> <td>750925</td> <td>749504</td> <td>130394</td> </tr> <tr> <td>GEOLOC</td> <td>2729702</td> <td>2728296</td> <td>215924</td> </tr> <tr> <td>TRANSPORT</td> <td>161539</td> <td>160487</td> <td>53405</td> </tr> <tr> <td>EVENT</td> <td>798433</td> <td>798251</td> <td>86471</td> </tr> <tr> <td>FACILITY</td> <td>258835</td> <td>258513</td> <td>109867</td> </tr> <tr> <td>RISKNAT</td> <td>5502</td> <td>4994</td> <td>127</td> </tr> <tr> <td>DAMAGES</td> <td>1136</td> <td>121</td> <td>56</td> </tr> <tr> <td>OTHER</td> <td>4340621</td> <td>4339658</td> <td>682458</td> </tr> <tr> <td><strong>Total</strong></td> <td><strong>10146795</strong></td> <td><strong>10138230</strong></td> <td><strong>1836399</strong></td> </tr> </tbody> </table> <p>Overview of the mentions annotated in the full dataset. #Mentions shows&nbsp;the total number of mentions per label, #Linked the number of mentions linked to an&nbsp;entity and #Entities the number of distinct entities per label present in the dataset.</p>

opencc-by-4.0Mar 2023View details →
zenodo48/100

OcWikiAnnot: Annotated Wikipedia Corpus of Occitan

<p>OcWikiAnnot is a corpus of Wikipedia content in Occitan that is tokenized, PoS-tagged and lemmatized. The corpus contains 100 000 sentences for a total of 2&nbsp;037&nbsp;723 tokens. It is based on the Wikipedia corpus in Occitan that is part of the <a href="https://corpora.uni-leipzig.de/en?corpusId=oci_wikipedia_2021">Leipzig Corpora Collection</a>.</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Apr 2023View details →
zenodo48/100

Wikipedia Multilingual Vandalism Detection Dataset

<p>This dataset accompanies a research paper that introduces a novel system designed to support the Wikipedia community in combating vandalism on the platform. The dataset has been prepared to enhance the accuracy and efficiency of Wikipedia patrolling in multiple languages.</p> <p>The release of this comprehensive dataset aims to encourage further research and development in vandalism detection techniques, fostering a safer and more inclusive environment for the Wikipedia community. Researchers and practitioners can utilize this dataset to train and validate their models for vandalism detection and contribute to improving online platforms' content moderation strategies.</p> <p><strong>Dataset Details:</strong></p> <ul> <li><strong>Number of Languages:</strong>&nbsp;47</li> <li><strong>Observation period:&nbsp;</strong>6 months training, one week hold-out testing</li> <li><strong>Use Case:</strong>&nbsp;The dataset is primarily intended for training and evaluating vandalism detection systems.</li> <li><strong>Features:</strong>&nbsp;Each record characterizes the corresponding revision of the Wikipedia page, including revision metadata, user details, text inserted, removed, or changed, and corresponding MLMs-based features.&nbsp;</li> <li><strong>Data Filtering and Feature Engineering:</strong>&nbsp;Advanced filtering and feature engineering techniques were applied to ensure the dataset's quality and relevance for effectively training the vandalism detection system.</li> <li><strong>Files:&nbsp;</strong>Training and hold-out testing datasets of anonymous and all users.&nbsp;</li> </ul> <p>&nbsp;</p> <p><strong>Related paper citation:</strong></p> <pre><code>@inproceedings{10.1145/3580305.3599823, author = {Trokhymovych, Mykola and Aslam, Muniza and Chou, Ai-Jou and Baeza-Yates, Ricardo and Saez-Trumper, Diego}, title = {Fair Multilingual Vandalism Detection System for Wikipedia}, year = {2023}, isbn = {9798400701030}, publisher = {Association for Computing Machinery}, address = {New York, NY, USA}, url = {https://doi.org/10.1145/3580305.3599823}, doi = {10.1145/3580305.3599823}, abstract = {This paper presents a novel design of the system aimed at supporting the Wikipedia community in addressing vandalism on the platform. To achieve this, we collected a massive dataset of 47 languages, and applied advanced filtering and feature engineering techniques, including multilingual masked language modeling to build the training dataset from human-generated data. The performance of the system was evaluated through comparison with the one used in production in Wikipedia, known as ORES. Our research results in a significant increase in the number of languages covered, making Wikipedia patrolling more efficient to a wider range of communities. Furthermore, our model outperforms ORES, ensuring that the results provided are not only more accurate but also less biased against certain groups of contributors.}, booktitle = {Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining}, pages = {4981&ndash;4990}, numpages = {10}, location = {Long Beach, CA, USA}, series = {KDD '23} }</code></pre> <p>&nbsp;</p>

opencc-by-4.0Jul 2023View details →
zenodo44/100

Yearly pageviews of English Wikipedia articles with potential links to green open access scholarly articles

<p>Number of visits in 2019 for a sample of 23462 English Wikipedia articles which contain references to academic sources which have a green open access copy available but not yet used. The consultation statistics were retrieved from the Wikimedia pageviews API using the Python client (script also included). The sample was selected among articles which in April 2020 had at least one citation of an academic paper (using the &quot;cite journal&quot; template) for which OAbot (through Unpaywall data) had found a green open access URL to add (gratis open access, not necessarily libre open access). Data shows that the top 1 % most visited articles received 30 % of the visits: over 500 million in the year, corresponding to 1 million potential citation link clicks to distribute across all references assuming a 0.2 % click-through rate per Piccardi et al. (2020).</p>

opencc-zeroMay 2020View details →
zenodo44/100

Google Trends and Wikipedia Page Views

<p><strong>Abstract</strong> (our paper)</p> <p>The frequency of a web search keyword generally reflects the degree of public interest in a particular subject matter. Search logs are therefore useful resources for trend analysis. However, access to search logs is typically restricted to search engine providers. In this paper, we investigate whether search frequency can be estimated from a different resource such as Wikipedia page views of open data. We found frequently searched keywords to have remarkably high correlations with Wikipedia page views. This suggests that Wikipedia page views can be an effective tool for determining popular global web search trends.</p> <p><strong>Data</strong></p> <p>personal-name.txt.gz:<br> The first column is the Wikipedia article id, the second column is the search keyword, the third column is the Wikipedia article title, and the fourth column is the total of page views from 2008 to 2014.</p> <p>personal-name_data_google-trends.txt.gz, personal-name_data_wikipedia.txt.gz:<br> The first column is the period to be collected, the second column is the source (Google or Wikipedia), the third column is the Wikipedia article id, the fourth column is the search keyword, the fifth column is the date, and the sixth column is the value of search trend or page view.</p> <p><strong>Publication</strong></p> <p>This data set was created for our study. If you make use of this data set, please cite:<br> Mitsuo Yoshida, Yuki Arase, Takaaki Tsunoda, Mikio Yamamoto. Wikipedia Page View Reflects Web Search Trend. <em>Proceedings of the 2015 ACM Web Science Conference (WebSci '15)</em>. no.65, pp.1-2, 2015.<br> http://dx.doi.org/10.1145/2786451.2786495<br> http://arxiv.org/abs/1509.02218 (author-created version)</p> <p><strong>Note</strong></p> <p>The raw data of Wikipedia page views is available in the following page.<br> http://dumps.wikimedia.org/other/pagecounts-raw/</p>

opencc-zeroJun 2015View details →
zenodo44/100

Structured citations in the English Wikipedia

<p>This dataset contains the metadata of citations in the English Wikipedia that editors have input as citation templates (which use the Citation Style 1). It has been obtained from an XML dump of Wikipedia (2016-05-01), and was parsed with the wikiciteparser library. This library runs the Lua code used in Wikipedia to format such citations and generate structured metadata (such as COinS) from them.</p> <p>https://github.com/dissemin/wikiciteparser</p> <p>Sample:</p> <blockquote> <p>716551092&nbsp;&nbsp; &nbsp;12&nbsp;&nbsp; &nbsp;2016-04-22T10:19:33Z&nbsp;&nbsp; &nbsp;Anarchism&nbsp;&nbsp; &nbsp;cite journal &nbsp;&nbsp; &nbsp;{&quot;PublisherName&quot;: &quot;International Group of San Francisco&quot;, &quot;Title&quot;: &quot;Towards Anarchism&quot;, &quot;URL&quot;: &quot;http://www.marxists.org/archive/malatesta/1930s/xx/toanarchy.htm&quot;, &quot;Authors&quot;: [{&quot;link&quot;: &quot;Errico Malatesta&quot;, &quot;last&quot;: &quot;Malatesta&quot;, &quot;first&quot;: &quot;Errico&quot;}], &quot;ID_list&quot;: {&quot;OCLC&quot;: &quot;3930443&quot;}, &quot;Periodical&quot;: &quot;MAN!&quot;, &quot;PublicationPlace&quot;: &quot;Los Angeles&quot;}<br /> 716551092&nbsp;&nbsp; &nbsp;12&nbsp;&nbsp; &nbsp;2016-04-22T10:19:33Z&nbsp;&nbsp; &nbsp;Anarchism&nbsp;&nbsp; &nbsp;cite journal &nbsp;&nbsp; &nbsp;{&quot;Date&quot;: &quot;2007-05-14&quot;, &quot;URL&quot;: &quot;http://www.theglobeandmail.com/servlet/story/RTGAM.20070514.wxlanarchist14/BNStory/lifeWork/home/&quot;, &quot;Title&quot;: &quot;Working for The Man&quot;, &quot;Periodical&quot;: &quot;The Globe and Mail&quot;, &quot;Authors&quot;: [{&quot;last&quot;: &quot;Agrell&quot;, &quot;first&quot;: &quot;Siri&quot;}]}<br /> 716551092&nbsp;&nbsp; &nbsp;12&nbsp;&nbsp; &nbsp;2016-04-22T10:19:33Z&nbsp;&nbsp; &nbsp;Anarchism&nbsp;&nbsp; &nbsp;cite web &nbsp;&nbsp; &nbsp;{&quot;Date&quot;: &quot;2006&quot;, &quot;URL&quot;: &quot;http://www.britannica.com/eb/article-9117285&quot;, &quot;PublisherName&quot;: &quot;Encyclop\u00e6dia Britannica Premium Service&quot;, &quot;Periodical&quot;: &quot;Encyclop\u00e6dia Britannica&quot;, &quot;Title&quot;: &quot;Anarchism&quot;}<br /> 716551092&nbsp;&nbsp; &nbsp;12&nbsp;&nbsp; &nbsp;2016-04-22T10:19:33Z&nbsp;&nbsp; &nbsp;Anarchism&nbsp;&nbsp; &nbsp;cite journal &nbsp;&nbsp; &nbsp;{&quot;Date&quot;: &quot;2005&quot;, &quot;Pages&quot;: &quot;14&quot;, &quot;Periodical&quot;: &quot;The Shorter Routledge Encyclopedia of Philosophy&quot;, &quot;Title&quot;: &quot;Anarchism&quot;}</p> </blockquote> <p>Credit: these citations have been input by Wikipedia editors (this dataset is therefore distributed under a CC-BY-SA license).</p> <p>Acknowledgments: this extraction process was carried on one of OpenJournal&#39;s servers.</p> <p>See also: citation identifiers extracted from Wikipedia (so, not looking at specific templates, but using regular expressions): http://dx.doi.org/10.6084/m9.figshare.1299540</p>

opencc-by-sa-4.0Jun 2016View details →
zenodo44/100

The Online Conversation Threads Repository (Slashdot, Barrapunto, Wikipedia talk)

<p>This repository contains datasets with online conversation threads collected and analyzed by different researchers. Currently, you can find datsets from different news aggregators (Slashdot, Barrapunto) and the English Wikipedia talk pages.</p> <p>- Slashdot conversations (Aug 2005 - Aug 2006) Online conversations generated at Slashdot during a year. Posts and comments published between August 26th, 2005 and August 31th, 2006. For each discussion thread: sub-domains, title, topics and hierarchical relations between comments. For each comment: user, date, score and textual content. This dataset is different from the Slashdot Zoo social network (it is not a signed network of users) contained in the SNAP repository and represents the full version of the dataset used in the CAW 2.0 - Content Analysis for the WEB 2.0 workshop for the WWW 2009 conference that can be found in several repositories such as Konect Barrapunto conversations (Jan 2005 - Dec 2008)</p> <p>- Online conversations generated at Barrapunto (Spanish clone of Slashdot) during three years. For each discussion thread: sub-domains, title, topics and hierarchical relations between comments. For each comment: user, date, score and textual content Wikipedia (2001 - Mar 2010)</p> <p>- Data from articles discussions (talk) pages of the English Wikipedia as of March 2010. It contains comments on about 870,000 articles (i.e. all articles which had a corresponding talk page with at least one comment), in total about 9.4 million comments. The oldest comments date back to as early as 2001.</p> <p> </p>

opencc-by-4.0Apr 2008View details →
zenodo44/100

Wikipedia Page Views of Japanese Comic

<p><strong>Abstract</strong> (our paper)</p> <p>This paper investigates the page view and interlanguage link at Wikipedia for Japanese comic analysis. This paper is based on a preliminary investigation, and obtained three results, but the analysis is insufficient to use the results for a market research immediately. I am looking for research collaborators in order to conduct a more detailed analysis.</p> <p><strong>Data</strong></p> <p><strong>Publication</strong></p> <p>This data set was created for our study. If you make use of this data set, please cite:<br> Mitsuo Yoshida. Preliminary Investigation for Japanese Comic Analysis using Wikipedia. <em>Proceedings of the Fifth Asian Conference on Information Systems (ACIS 2016)</em>. pp.229-230, 2016.</p>

opencc-zeroOct 2016View details →
zenodo44/100

HUMANE Wikipedia simulation modelling bootstrapping data

<p>This data set has been derived from the Simple English Wikipedia data publicly available and post-processed in the WikiWarMonitor project. The data set this is derived from is available from: http://wwm.phy.bme.hu/light.html</p> <p>The data set comprises a collection of 15 CSV files with summary statistics of the contributors to Wikipedia (Simple English only) in the period of 18/05/2001 to 17/10/2012. The files cover:</p> <ul> <li>Statistics of registered users, anonymous users and bots.</li> <li>History of revert activity</li> <li>History of edit wars</li> <li>Activity statistics broken down into weekly snapshots</li> </ul> <p>Each CSV file has a descriptive header that is generally self-explanatory, so the data is not further described here. However, note that in the activity_snapshots_aggregated.csv file, the edit war conditions are as follows:</p> <ul> <li>Condition 1: ongoing (started before the snapshot and continues)</li> <li>Condition 2: started and finished within the snapshot</li> <li>Condition 3: started within the snapshot, but did not finish yet</li> <li>Condition 4: started before the snapshot, but finished within the snapshot</li> </ul>

opencc-by-4.0May 2017View details →
zenodo44/100

TokTrack: A Complete Token Provenance and Change Tracking Dataset for the English Wikipedia

<p><strong>Fixes in version 1.1 (= Zenodo's "version 2")</strong></p> <p>*In 20161101-revisions-part1-12-1728.csv, missing first data line is added.</p> <p>*In Current_content and Deleted_content files, some token values ('str' column) which contain regular quotes ('"') are fixed.</p> <p>*In Current_content and Deleted_content files, some wrong revision ID values for 'origin_rev_id', 'in' and 'out' columns are fixed.</p> <p> ------</p> <p><strong>This dataset contains every instance of all tokens (≈ words) ever written in undeleted, non-redirect English Wikipedia articles until October 2016, in total 13,545,349,787 instances. Each token is annotated with (i) the article revision it was originally created in, and (ii) lists with all the revisions in which the token was ever deleted and (potentially) re-added and re-deleted from its article, enabling a complete and straightforward tracking of its history.</strong></p> <p>This data would be exceedingly hard to create by an average potential user as it is (i) very expensive to compute and as (ii) accurately tracking the history of each token in revisioned documents is a non-trivial task. <br> Adapting a state-of-the-art algorithm, we have produced a dataset that allows for a range of analyses and metrics, already popular in research and going beyond, to be generated on complete-Wikipedia scale; ensuring quality and allowing researchers to forego expensive text-comparison computation, which so far has hindered scalable usage.</p> <p>This dataset, its creation process and use cases are described in a dedicated dataset paper of the same name, published at the ICWSM 2017 conference. In this paper, we show how this data enables, on token level, computation of provenance, measuring survival of content over time, very detailed conflict metrics, and fine-grained interactions of editors like partial reverts, re-additions and other metrics.</p> <p>Tokenization used: https://gist.github.com/faflo/3f5f30b1224c38b1836d63fa05d1ac94</p> <p>Toy example for how the token metadata is generated: <br> https://gist.github.com/faflo/8bd212e81e594676f8d002b175b79de8</p> <p><strong>Be sure to read the ReadMe.txt or - even more detailed - the supporting paper which is referenced under "related identifiers".</strong></p>

opencc-by-sa-4.0Mar 2017View details →
zenodo44/100

Wikipedia: wikipedia-eu (Basque)

Wikipedia is a multilingual, web-based, free-content encyclopedia project supported by the Wikimedia Foundation and based on a model of openly editable content. EOL harvests articles from wikipedia that are indexed as species or higher taxa.<p></p><p></p>https://eu.wikipedia.org/wiki/Azala

opencc-by-sa-4.0Aug 2024View details →
zenodo44/100

Wikipedia: wikipedia-id (Indonesian)

Wikipedia is a multilingual, web-based, free-content encyclopedia project supported by the Wikimedia Foundation and based on a model of openly editable content. EOL harvests articles from wikipedia that are indexed as species or higher taxa.<p></p>Wikipedia bahasa Indonesia disediakan secara gratis oleh Wikimedia Foundation, sebuah organisasi nirlaba. Selain dalam bahasa Indonesia, Wikipedia tersedia dalam bahasa daerah berikut: Aceh, Bali, Banjar, Banyumasan, Bugis, Gorontalo, Jawa, Melayu, Minangkabau, Sunda, dan Tetun. <p></p>https://id.wikipedia.org/wiki/Halaman_Utama

opencc-by-sa-4.0Aug 2024View details →
zenodo44/100

Wikipedia: wikipedia-hr (Croatian)

Wikipedia is a multilingual, web-based, free-content encyclopedia project supported by the Wikimedia Foundation and based on a model of openly editable content. EOL harvests articles from wikipedia that are indexed as species or higher taxa.<p></p><p></p>https://hr.wikipedia.org/

opencc-by-sa-4.0Aug 2024View details →
zenodo44/100

Wikipedia: wikipedia-az (Azerbaijani)

Wikipedia is a multilingual, web-based, free-content encyclopedia project supported by the Wikimedia Foundation and based on a model of openly editable content. EOL harvests articles from wikipedia that are indexed as species or higher taxa.<p></p><p></p>https://az.wikipedia.org/

opencc-by-sa-4.0Aug 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record