Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

219

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

219 results for “wikipedia”

Learn how ShareScore rates datasets ↗
dryad40/100

Causal evidence for social group sizes from Wikipedia editing data

Open the record for dataset details and reuse information.

publicApr 2024View details →
zenodo36/100

A Wikipedia dataset of Ebola disease related articles

<p>A subset of articles extracted from the French Wikipedia (on 28/04/2019). Data published here include articles related to Ebola and some tropical related diseases. Each article is a UTF8 plain text.</p>

opencc-by-4.0Dec 2019View details →
zenodo36/100

英語版Wikipediaにおける学術文献の参照記述の初出時点データセット (Dataset of First Appears of Scholarly References on English Wikipedia)

<p>概要</p> <p>筑波大学大学院図書館情報メディア研究科に提出準備中の博士論文「Wikipediaにおける学術文献の参照記述に関する研究」の付録とする予定のデータセットです。作者は<a href="https://researchmap.jp/jir_o">吉川次郎</a>です。</p> <p>ごく簡単な説明</p> <ul> <li>博論本体の付録としてドキュメントを書きました (そのドラフトは<a href="https://www.dropbox.com/s/c8h0hkkbu2wb4jm/PhDThesis_Appendix.pdf?dl=0">こちら</a>)</li> <li>英語版Wikipedia上の学術文献の参照記述に関するデータセットです <ul> <li>dataset_doi_links はDOIリンクを対象として特定・抽出し、研究分野の紐付けを行ったものです (詳細は関連論文の1番)</li> <li>dataset_refs はdataset_doi_linksの参照記述群を対象に、誰が、いつ追加したのか? の特定を行ったものです (詳細は関連論文の2番)</li> </ul> </li> <li>いまのところ、ドキュメントは日本語のみです。英語への対応等は今後の課題です;</li> </ul> <p>関連論文</p> <ol> <li>吉川次郎; 高久雅生; 芳鐘冬樹: 「DOIリンクに基づくWikipedia上の参照記述における編集者の分析」, 情報知識学会誌, Vol. 30, No. 1, pp. 21--41, 2020. <a href="https://doi.org/10.2964/jsik_2020_004">https://doi.org/10.2964/jsik_2020_004</a></li> <li>吉川次郎; 高久雅生; 芳鐘冬樹: 「Wikipediaに学術文献の参照記述を追加する編集の特定手法」, 情報知識学会誌, Vol. 30, No. 3,2020 (全20ページ,採録決定).<a href="https://doi.org/10.2964/jsik_2020_033">https://doi.org/10.2964/jsik_2020_033</a></li> <li>吉川次郎; 高久雅生; 芳鐘冬樹: 「Wikipedia上の学術文献の参照記述の追加に関する時系列分析」(投稿中)</li> </ol>

openother-openAug 2020View details →
zenodo36/100

SparkWiki: Wikipedia graph dataset and pagecounts pre-processing tools

<p>SparkWiki toolkit can be used in various scenarios where you are interested in researching Wikipedia graph and pageview statistics. Graph and pageviews can be used and studied separately.&nbsp;The code used to process Wikipedia SQL dumps, along with deployment instructions, are&nbsp;located on <a href="https://github.com/epfl-lts2/sparkwiki">GitHub</a>.</p> <p><strong>To test an example of a pre-processed graph,</strong> you can download a dump of the English Wikipedia graph (see attached wikipedia_nrc.dump), which you can directly import into a Neo4J instance.&nbsp;The dump is intended for neo4j version 3.x and can be imported using the following command (make sure you do not have an existing wikipedia.db database as the command below will overwrite its content):</p> <p><code>sudo -u neo4j neo4j-admin load --force --from=wikipedia_nrc.dump --database=wikipedia.db</code></p> <p>If you try to import it into Neo4J&nbsp;version 4.x, you need to set the property&nbsp;</p> <p><code>dbms.allow_upgrade=true</code>&nbsp;in&nbsp;<code>/etc/neo4j/neo4j.conf</code>&nbsp;</p> <p>before importing.&nbsp;When you start the neo4j server it will upgrade the database s.t. it is compatible with version 4.x.</p> <p>&nbsp;</p>

opencc-by-4.0Jan 2019View details →
zenodo36/100

NwQM: A neural quality assessment framework for Wikipedia

<p>This contains the datasets we used for implementing &quot;NwQM: A neural quality assessment framework for Wikipedia&quot;.</p> <p>&quot;wikipages.csv&quot; contains the text page content, talk page content, and the quality for each page.</p> <p>&quot;sample_wiki_images.zip&quot; contains the screenshots for a sample of&nbsp;pages.</p> <p>&quot;finetuned_inceptionv3_model.h5&quot; is the InceptionV3 model finetuned on Wikipedia.</p> <p>&quot;finetuned_inceptionv3_embeddings.json&quot; contains the finetuned InceptionV3 embeddings&nbsp;of each image generated from&nbsp;finetuned_inceptionv3_model.h5.&nbsp;</p>

opencc-by-4.0Oct 2020View details →
zenodo36/100

Tracking Knowledge Propagation Across Wikipedia Languages

<p>We present a dataset of <em>inter-language knowledge propagation</em> in Wikipedia. Covering the entire 309 language editions and 33M articles, the dataset aims to track the full propagation history of Wikipedia concepts, and allow follow up research on building predictive models of them. For this purpose, we align all the Wikipedia articles in a language-agnostic manner according to the concept they cover, which results in 13M propagation instances. To the best of our knowledge, this dataset is the first to explore the full inter-language propagation at a large scale. Together with the dataset, a holistic overview of the propagation and key insights about the underlying structural factors are provided to aid future research. For example, we find that although long cascades are unusual, the propagation tends to continue further once it reaches more than four language editions. We also find that the size of language editions are associated with the speed of propagation. We believe the dataset not only contributes to the prior literature on Wikipedia growth but also enables new use cases such as edit recommendation for addressing knowledge gaps, detection of disinformation, and cultural relationship analysis.</p>

opencc-by-4.0Mar 2021View details →
zenodo36/100

Happy 20th birthday, Wikipedia - С днём рождения, Википедия!

<p>This repository contains a video I made on the occasion of Wikipedia&#39;s 20th birthday in January 2021.</p> <p>It is in Russian upon request of <a href="https://www1.wdr.de/radio/cosmo/programm/sendungen/radio-po-russki/index.html">Radio po-russki</a>, the Russian language service of the German radio station WDR Cosmo, who posted a version of it <a href="https://www.facebook.com/1615399212084960/videos/2900159796974569/">on their Facebook page</a> (also <a href="http://web.archive.org/web/20210209154902/https://www.facebook.com/1615399212084960/videos/2900159796974569/">available via the Wayback Machine</a>) on 28 January 2021 as a supplement to their <a href="https://www1.wdr.de/mediathek/audio/cosmo/radio-po-russki/audio-radio-po-russki---audio-on-demand--2566.html">radio program on 27 January 2021</a> (also <a href="https://web.archive.org/web/20210127201134/https://www1.wdr.de/mediathek/audio/cosmo/radio-po-russki/audio-radio-po-russki---audio-on-demand--2566.html">archived at the Wayback Machine</a>; the Wikipedia part starts at about 20:20 min), which is also available as a <a href="https://wdrmedien-a.akamaihd.net/medp/podcast/weltweit/fsk0/235/2351559/cosmoradioporusski_2021-01-27_cosmoradioporusskiganzesendung27012021_cosmo.mp3">podcast</a> (with the Wikipedia part starting at about 9:55 min) and likewise <a href="http://web.archive.org/web/20210209165657/https://wdrmedien-a.akamaihd.net/medp/podcast/weltweit/fsk0/235/2351559/cosmoradioporusski_2021-01-27_cosmoradioporusskiganzesendung27012021_cosmo.mp3">via the Wayback Machine</a>.</p> <p>The video deposited here is also available on YouTube via <a href="https://youtu.be/DJyDNC47a80">https://youtu.be/DJyDNC47a80</a> and on Wikimedia Commons via <a href="https://commons.wikimedia.org/wiki/File:Happy_20th_birthday,_Wikipedia_-_%D0%A1_%D0%B4%D0%BD%D1%91%D0%BC_%D1%80%D0%BE%D0%B6%D0%B4%D0%B5%D0%BD%D0%B8%D1%8F,_%D0%92%D0%B8%D0%BA%D0%B8%D0%BF%D0%B5%D0%B4%D0%B8%D1%8F.webm">https://w.wiki/vmn</a> .</p> <p>&nbsp;</p> <p>I used a <a href="https://www.teleprompt.online/">teleprompter</a> (speed 21, font size 8.3) to improve the flow.</p> <p>The text I spoke is below.</p> <p>&nbsp;</p> <p>Здравствуйте. Меня зовут Даниэль Митхен.&nbsp;</p> <p>&nbsp;</p> <p>Я биофизик, работаю специалистом по открытым данным, а в свободное время редактирую Википедию.</p> <p>&nbsp;</p> <p>Когда&nbsp; создали Википедию в январе 2001 года, я этого не заметил, так как моя семья была занята захоронением моего деда.&nbsp;</p> <p>&nbsp;</p> <p>Но год по-позже, работая над докторской диссертацией, я проводил систематические поиски про том, что известно или нет в моей специальности, то есть применении методой магнитно-резонансной микроскопии в биологии.</p> <p>&nbsp;</p> <p>По дороге, я узнал о Википедии, которая предлагала своим участникам совместно писать онлайн-энциклопедию на разных языках и в принципе во всех областях знаний, включая науки.&nbsp;</p> <p>&nbsp;</p> <p>Идея мне понравилась, и я попробовал.&nbsp;</p> <p>&nbsp;</p> <p>Мои первые правки были небольшими - например, добавление пропущенной запятой - и сделаны без учёта.</p> <p>&nbsp;</p> <p>Со временем, я начал узнавать все больше и больше об экосистеме Википедии и ее сообществе участников.&nbsp;</p> <p>&nbsp;</p> <p>Я интегрировал этот опыт в свои привычки учиться, преподавать и структурировать знания.</p> <p>&nbsp;</p> <p>Поскольку я ученый, большинство моих правок связано с науками, хотя некоторые из них относятся к другим вещам, включая новости или историю, языки или музыку.&nbsp;</p> <p>&nbsp;</p> <p>Я также участвую в некоторых проектах, связаные с Википедией, особенно путём сбора научных медиафайлов в Викискладе и научных данных в Викиданных.</p> <p>&nbsp;</p> <p>С днём рождения, Википедия!</p> <p><br> <br> Rough English translation:</p> <p>Hello. My name is Daniel Mietchen. I am a biophysicist working as a data scientist, and in my spare time, I edit Wikipedia.</p> <p>When Wikipedia started in January 2001, I did not notice it, since my family was occupied with the burial of my grandfather. But a year later, while working on my doctoral dissertation, I was performing systematic searches about what is known or not in my specialty, i.e. applications of Magnetic Resonance Microscopy to biology.</p> <p>On the way, I found out about Wikipedia that invited its users to write an online encyclopedia collaboratively, in several languages and across fields of knowledge, including science. I liked the idea, and so I gave it a try. My first edits were small - things like adding a missing comma - and made without a user account.</p> <p>Over time, I began to discover more and more of the Wikipedia ecosystem and its community of participants. I integrated these experiences into my habits of learning, teaching and structuring knowledge.</p> <p>Since I am a scientist, the majority of my edits is related to scientific topics, though some are linked to other things that cross my mind, e.g. news or history, languages or music. I also contribute to some of Wikipedia&#39;s neighbouring projects, especially by curating scientific media files in Wikimedia Commons and scientific data in Wikidata.</p> <p>Happy birthday, Wikipedia!</p>

opencc-zeroJan 2021View details →
zenodo36/100

Citations in the German Wikipedia

<p>Using https://github.com/halfak/Extract-scholarly-article-citations-from-Wikipedia we extracted citations from the February 2015 dump of the German Wikipedia based on DOIs, PubMed IDs, and ISBNs.</p>

opencc-zeroMar 2015View details →
zenodo36/100

Palmetto position storing Lucene index of Dutch Wikipedia

<p>Dutch language resource for calculating topic coherence with Palmetto [1, 2]. The dataset is a position storing Lucene index of the Dutch Wikipedia [3]. It was created in the context of the Netherlands eScience Center Dilipad project [4]. The pdf file contains the results of a case study that shows best topic coherence measure for topics consisting of Dutch nouns is NPMI.</p> <p>More details can be found in the README.</p> <p>[1] M. Roeder, A. Both, and A. Hinneburg. Exploring the space of topic coherence measures. In <em>Proceedings of the Eighth ACM International Conference on Web Search and Data Mining</em>, pages 399&ndash;408, 2015.</p> <p>[2] http://aksw.org/Projects/Palmetto.html</p> <p>[3] https://dumps.wikimedia.org/nlwiki/20151102/</p> <p>[4] https://www.esciencecenter.nl/project/dilipad</p>

opencc-by-sa-4.0Feb 2016View details →
zenodo36/100

DOI linked by Wikipedia and available in green Open Access

<p>List of&nbsp;399684 digital object identifiers (185790 unique) linked from the pages of Wikipedia in all languages (the 140 most visited subdomains according to stats.wikimedia.org, one file for each) and available on DOAI.io as redirect to an URL other than dx.doi.org.</p> <p>Those DOIs represent a cross section of research publications which are significant to the larger community of citizens and are available to them thanks to the green Open Access repositories.</p> <p>The script used to produce the dataset is also attached.</p>

opencc-zeroJun 2016View details →
zenodo36/100

Wikipedia: wikipedia-ml (Malayalam)

Wikipedia is a multilingual, web-based, free-content encyclopedia project supported by the Wikimedia Foundation and based on a model of openly editable content. EOL harvests articles from wikipedia that are indexed as species or higher taxa.<p></p><p></p>https://ml.wikipedia.org/

opencc-by-sa-4.0Aug 2024View details →
zenodo36/100

FAIR and Open multilingual clinical trials in Wikidata and Wikipedia

<p>&quot;FAIR and Open multilingual clinical trials in Wikidata and Wikipedia&quot; was a December 2021 presentation by Lane Rasberry and Cherrie Kwok. It was made at the conference &quot;Understanding Wikipedia&rsquo;s Dark Matter - Translation and Multilingual Practice in the World&#39;&rsquo;s Largest Online Encyclopaedia&quot; hosted by the Centre for Translation and the Department of Translation, Interpreting and Intercultural Studies, both at Hong Kong Baptist University.</p> <ul> <li>conference page <a href="https://ctn.hkbu.edu.hk/wikiconf2021/">https://ctn.hkbu.edu.hk/wikiconf2021/</a></li> <li>watch video <a href="https://www.youtube.com/watch?v=5yRhCENeezQ">https://www.youtube.com/watch?v=5yRhCENeezQ</a></li> <li>slides archive <a href="https://commons.wikimedia.org/wiki/File:FAIR_and_Open_multilingual_clinical_trials_in_Wikidata_and_Wikipedia.pdf">https://commons.wikimedia.org/wiki/File:FAIR_and_Open_multilingual_clinical_trials_in_Wikidata_and_Wikipedia.pdf</a></li> </ul> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2021View details →
zenodo36/100

Word2vec models trained on English Wikipedia

<p>This repository contains Word2Vec models trained on the full text of the English Wikipedia as downloaded in December 2021.</p> <p>Preprocessing:</p> <ul> <li>lowercasing</li> <li>n-grams up to 4-grams were computed using Bouma 2009&nbsp;(https://svn.spraakdata.gu.se/repos/gerlof/pub/www/Docs/npmi-pfd.pdf),&nbsp;min freq threshold of 10</li> </ul> <p>Two models, trained with Gensim:</p> <ul> <li>wiki_300_5_word2vec --&gt; dim 300, freq threshold 5</li> <li>wiki_300_50_word2vec &nbsp;--&gt; dim 300, freq threshold 50</li> </ul> <p>Other hyperparameters set as follows: window=5, epochs=5, seed=1830, sg=1</p> <p>Note:<br> Machine learning models trained on uncurated data inevitably learn hidden or obvious biases and as a result, the models shared with here might&nbsp;contain characteristics&nbsp;including sexism, racism, antisemitism, homophobia, and other such types of unacceptable biases. I encourage whoever is using these models to make sure such biases are actually removed before using them in production settings (see eg https://aclanthology.org/N19-1061/)</p>

opencc-by-4.0May 2022View details →
zenodo36/100

Polifonia_Corpus_Wikipedia_Annotations_ES

<p>Polifonia_Corpus_Wikipedia_Annotations_ES</p>

opencc-by-4.0Jun 2022View details →
zenodo36/100

Polifonia_Corpus_Wikipedia_Annotations_FR

<p>Polifonia_Corpus_Wikipedia_Annotations_FR</p>

opencc-by-4.0Jun 2022View details →
zenodo36/100

Polifonia_Corpus_Wikipedia_Annotations_IT

<p>Polifonia_Corpus_Wikipedia_Annotations_IT</p>

opencc-by-4.0Jun 2022View details →
zenodo36/100

Wikipedia Complete Citation Corpus

<p><strong>Wikipedia Complete Citation Corpus</strong> (<strong>WCCC</strong>) is a corpus of citations, references and sources mined from the English Wikipedia. WCCC was created as a knowledge base used in a machine-learning model for recommending reliable sources to support (or refute) a given textual claim, but can be used for many other purposes.</p> <p>The solution was described in the paper <em>&quot;Countering Disinformation by Finding Reliable Sources: a Citation-Based Approach&quot;</em>, presented at the 2022 International Joint Conference on Neural Networks (IJCNN 2022). Please refer to the article (in <a href="https://doi.org/10.1109/IJCNN55064.2022.9891941">conference proceedings</a> or <a href="https://home.ipipan.waw.pl/p.przybyla/bib/Learning_to_Cite_3_CR.pdf">authors&#39; version</a>) for more information on the process of mining the corpus, comparisons with similar resources and the role it plays in recommending sources.&nbsp;The research was done within the&nbsp;<a href="https://homados.ipipan.waw.pl/">HOMADOS</a>&nbsp;project at the&nbsp;<a href="https://ipipan.waw.pl/">Institute of Computer Science</a>, Polish Academy of Sciences.</p> <p>WCCC contains 4.8 million documents with 50.8 million citations of 24.3 million sources. The dataset is divided into 10 parts (WCCC-part0.zip to WCCC-part9.zip) with approximately the same size. Each of the parts contains batch archives (e.g. batch130.zip), each covering up to 1000 Wikipedia articles. An article is identified by its ID number and described by the following files:</p> <ul> <li>&lt;ID&gt;_text.txt: the textual content of the article,</li> <li>&lt;ID&gt;_citations.txt: the citations occurring in this article, saved as tab-separated values of (1) character offset in the textual content and (2) reference ID,</li> <li>&lt;ID&gt;_references.txt: the references cited in the article, saved as tab-separated values of (1) reference ID and (one or many) pairs of (2) source ID and (3) location in the source (e.g. page number),</li> <li>&lt;ID&gt;_sources.txt: the sources referenced in the article, saved as tab-separated values of (1) source ID and (2) source description (in wikicode).</li> <li>&lt;ID&gt;.txt: human-readable text, created by enriching textual content with the article title and reference IDs.</li> </ul> <p>Additionally, article metadata are included in the meta.tsv file. Each line describes a single article through the following tab-separated fields:</p> <ul> <li>article ID,</li> <li>title of the article,</li> <li>Wikipedia ID of the article, which can be used to access the article through URL, i.e. https://en.wikipedia.org/?curid=&lt;WIKI_ID&gt;</li> <li>length of the textual content of the article (number of characters),</li> <li>number of sources in the article,</li> <li>number of references in the article,</li> <li>number of citations in the article.</li> </ul> <p>Files metaTrain.tsv and metaTest.tsv contain the same information, but split into training and test set, as used in the work.</p> <p>Please refer to <a href="https://home.ipipan.waw.pl/p.przybyla/bib/Learning_to_Cite_3_CR.pdf">the paper</a> for an in-depth explanation of the data structure (citations, references, sources, etc.). WCCC was created using Wikipedia dump from 01.02.2021, but you can repeat the mining process using a different dump (or different procedure) by using the <a href="https://github.com/piotrmp/finding_reliable_sources">published source code</a>. If you intend to apply the corpus in a fact-checking use-case, you might also look at the <a href="https://dx.doi.org/10.5281/zenodo.6539087">evaluation datasets we publish separately</a>, one of which is based on WCCC with additional elements (e.g. source identifiers: URL/ISBN/DOI).</p>

opencc-by-4.0May 2022View details →
zenodo36/100

Wikary: A Dataset of N-ary Wikipedia Tables Matched to Qualified Wikidata Statements

<p><strong>Wikary: A Dataset of N-ary Wikipedia Tables Matched to Qualified Wikidata Statements</strong></p> <p>Created for The SemTab 2022 Datasets Track challenge.</p> <p><strong>General</strong></p> <p>Explanation of columns names used in both files</p> <p>`lang` - language and Wikipedia version used</p> <p>`pageTitle` - page title</p> <p>`tableIndex` -&nbsp; index where the given table is located on the page</p> <p><strong>Tables file</strong></p> <p>Columns names are used only in the tables file</p> <p>`pageEntity` - Wikidata entity associated with the page</p> <p>`sectionTitle` - the title of the section where the table is located</p> <p>`tableCaption` - caption of the table</p> <p>`headers` - headers of the table</p> <p>`HTML` - HTML of the table</p> <p><strong>Matches file</strong></p> <p>Columns names are used only in the matches file</p> <p>`rowIndex` - index of a row where the match is found for a given table</p> <p>`wikidata_ids` - Wikidata entities in the row including `pageEntity`</p> <p>`entities_index` -&nbsp; indexes in which cell Wikidata entities were found, -1 used for Wikidata entity associated with the page, -9 used for cells that include a date in a cell</p> <p>`entities_anchor` -&nbsp; anchor text of cells where Wikidata entities were found</p> <p>`entities_cell_text` -&nbsp; cell text of cells where Wikidata entities were found</p> <p>`subject` - subject of Wikidata statement</p> <p>`property` - property of Wikidata statement</p> <p>`object` - object of Wikidata statement</p> <p>`property_qualifier` - property qualifier of Wikidata statement</p> <p>`qualifier_value` - qualifier value of Wikidata statement</p> <p>`id_match` - 1 means that row contains *Wikidata identifier match*, 0 means no match</p> <p>`year_match` - 1 means that row contains *Year cell match*, 0 means no match</p> <p>`year_part_match` - 1 means that row contains *Within cell year match*, 0 means no match</p> <p>&nbsp;</p>

opencc-by-4.0Aug 2022View details →
zenodo36/100

Semantically tagged Finnish Wikipedia 2017

<p><strong>Description of FI Wikipedia 2017 tagging</strong></p> <p><strong>Kimmo Kettunen</strong></p> <p><strong>University of Eastern Finl</strong><strong>and</strong></p> <p>The tagged data contains the texts of the Finnish Wikipedia of 2017. It has been first tagged syntactically in the Language Bank of Finland using the available UD2 tagger version of the Mylly service (https://mylly.rahtiapp.fi/home).</p> <p>Semantic tags to the UD2 parse have been added using a lexical semantic tagger FiST (Kettunen, 2019, <a href="https://aclanthology.org/W19-0306/">https://aclanthology.org/W19-0306/</a>).</p> <p>This published version has been condensed to a format where each analysed word contains the</p> <p>1. original running word form,</p> <p>2. lemma of the word form from UD2 parse,</p> <p>3. part-of-speech of the word from FiST</p> <p>4. semantic tag(s) for the word from FiST, and</p> <p>5. syntactic function of the word from UD2 parse.</p> <p>Semantic tags used are explained in this UCREL Semantic Analysis System (USAS) document: <a href="https://ucrel.lancs.ac.uk/usas/USASSemanticTagset.pdf">https://ucrel.lancs.ac.uk/usas/USASSemanticTagset.pdf</a></p> <p>Tagging includes all the semantic tags available for the word, as FiST does not perform disambiguation. Unknown words for the tagger are marked with tag Z99. Punctuation is tagged with PUNCT and numbers with NUMB. Lines beginning with # are output of UD2 and contain document, paragraph and sentence information.</p> <p>The output contains 6&nbsp;415&nbsp;027 sentences and 98.81 million lines. Lexical coverage of the semantic tagging is 76.59 %</p> <p><strong>Examples of output</strong></p> <p># newdoc</p> <p># newpar</p> <p># sent_id = 1</p> <p># text = Amsterdam</p> <p>Amsterdam#Amsterdam#Proper#Z2 root</p> <p># newpar</p> <p># sent_id = 2</p> <p># text = Amsterdam on Alankomaiden p&auml;&auml;kaupunki.</p> <p>Amsterdam#Amsterdam#Proper#Z2 nsubj:cop</p> <p>on#olla#Verb#A3+ A1.1.1 M6 Z5 cop</p> <p>Alankomaiden#Alankomaat#Proper#Z2 nmod:poss</p> <p>p&auml;&auml;kaupunki#p&auml;&auml;kaupunki#Noun#M7 root</p> <p>. PUNCT</p>

opencc-by-4.0Sep 2022View details →
zenodo36/100

WikiReaD (Wikipedia Readability Dataset)

<p><strong>Dataset Description:</strong></p> <p>The dataset contains pairs of encyclopedic articles in 14 languages. Each pair includes the same article in two levels of readability (easy/hard). The pairs are obtained by matching Wikipedia articles (hard) with the corresponding versions from different simplified or children's encyclopedias (easy).</p> <p>&nbsp;</p> <p><strong>Dataset Details:</strong></p> <ul> <li><strong>Number of Languages:</strong> 14</li> <li><strong>Number of files:</strong> 19</li> <li><strong>Use Case:</strong> Training and evaluating readability scoring models for articles within and outside Wikipedia.</li> <li><strong>Processing details:</strong> Text pairs are created by matching articles from Wikipedia with the corresponding article in the simplified/children encyclopedia either via the Wikidata item ID or their page titles. The text of each article is extracted directly from their parsed HTML version.</li> <li><strong>Files:</strong> The dataset consists of independent files for each type of children/simplified encyclopedia and each language (e.g., `&lt;wiki&gt;-&lt;language_code&gt;_sentences.bz2`). Also, the dataset contains train-test split files for&nbsp; <div> <div>simplewiki-en (trainsplit_simplewiki-en_sentences.bz2, testsplit_simplewiki-en_sentences.bz2) needed to reproduce the results of the corresponding paper.&nbsp;</div> <div>&nbsp;</div> </div> </li> </ul> <p><strong>Attribution:</strong></p> <p>The dataset was compiled from the following sources. The text of the original articles comes from the corresponding language version of Wikipedia. The text of the simplified articles comes from one of the following encyclopedias: Simple English Wikipedia, Vikidia, Klexikon, Txikipedia, or Wikikids.</p> <p>Below we provide information about the license of the original content as well as the template to generate the link to the original source for a given page (&lt;page_title&gt;) and language (&lt;language_code&gt;). For example, <a href="https://en.wikipedia.org/wiki/Spain">https://en.wikipedia.org/wiki/Spain</a> links to the page &ldquo;Spain&rdquo; in English Wikipedia)</p> <ul> <li><a href="https://www.wikipedia.org/">Wikipedia</a> <ul> <li>Source: <code>https://&lt;language_code&gt;.wikipedia.org/wiki/&lt;page_title&gt;</code></li> <li>License:&nbsp;<a href="https://creativecommons.org/licenses/by-sa/4.0/deed.en">CC BY-SA 4.0</a>, <a href="https://www.gnu.org/copyleft/fdl.html">GFDL</a></li> </ul> </li> <li><a href="https://simple.wikipedia.org">Simple English Wikipedia</a> <ul> <li>Source: <code>https://simple.wikipedia.org/wiki/&lt;page_title&gt;</code></li> <li>License: <a href="https://creativecommons.org/licenses/by-sa/4.0/deed.en">CC BY-SA 4.0</a>, <a href="https://www.gnu.org/copyleft/fdl.html">GFDL</a></li> </ul> </li> <li><a href="https://www.vikidia.org/">Vikidia</a> <ul> <li>Source: <code>https://&lt;language_code&gt;.vikidia.org/wiki/&lt;page_title&gt;</code></li> <li>License: <a href="https://creativecommons.org/licenses/by-sa/3.0/deed.en">CC BY-SA 3.0</a>, <a href="https://www.gnu.org/copyleft/fdl.html">GFDL</a></li> </ul> </li> <li><a href="https://klexikon.zum.de">Klexikon</a> <ul> <li>Source: <code>https://klexikon.zum.de/wiki/&lt;page_title&gt;</code></li> <li>License: <a href="https://creativecommons.org/licenses/by-sa/4.0/deed.en">CC BY-SA 4.0</a></li> </ul> </li> <li><a href="https://eu.wikipedia.org/wiki/Txikipedia">Txikipedia</a> <ul> <li>Source: <code>https://eu.wikipedia.org/wiki/Txikipedia:&lt;page_title&gt;</code></li> <li>License: <a href="https://creativecommons.org/licenses/by-sa/4.0/deed.en">CC BY-SA 4.0</a>, <a href="https://www.gnu.org/copyleft/fdl.html">GFDL</a></li> </ul> </li> <li><a href="https://wikikids.nl/">Wikikids</a> <ul> <li>Source: <code>https://wikikids.nl/&lt;page_title&gt;</code></li> <li>License: <a href="https://creativecommons.org/licenses/by-sa/3.0/deed.en">CC BY-SA 3.0</a></li> </ul> </li> </ul> <p><strong>Related paper citation:&nbsp;</strong></p> <blockquote> <pre><code>@inproceedings{trokhymovych-etal-2024-open, title = "An Open Multilingual System for Scoring Readability of {W}ikipedia", author = "Trokhymovych, Mykola and Sen, Indira and Gerlach, Martin", editor = "Ku, Lun-Wei and Martins, Andre and Srikumar, Vivek", booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)", month = aug, year = "2024", address = "Bangkok, Thailand", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/2024.acl-long.342/", doi = "10.18653/v1/2024.acl-long.342", pages = "6296--6311"<br>}</code></pre> </blockquote>

opencc-by-sa-4.0May 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record