Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
219
datasets available to search
ShareScore release 0.9.0
Dataset results
219 results for “wikipedia”
Wikipedia: wikipedia-sk (Slovak)
Wikipedia is a multilingual, web-based, free-content encyclopedia project supported by the Wikimedia Foundation and based on a model of openly editable content. EOL harvests articles from wikipedia that are indexed as species or higher taxa.<p></p><p></p>https://sk.wikipedia.org/
Wikipedia: wikipedia-sv (Swedish)
Wikipedia is a multilingual, web-based, free-content encyclopedia project supported by the Wikimedia Foundation and based on a model of openly editable content. EOL harvests articles from wikipedia that are indexed as species or higher taxa.<p></p>
Wikipedia: wikipedia-min (Minangkabau)
Wikipedia is a multilingual, web-based, free-content encyclopedia project supported by the Wikimedia Foundation and based on a model of openly editable content. EOL harvests articles from wikipedia that are indexed as species or higher taxa.<p></p>
Wikipedia: wikipedia_combined_languages_batch2
<div> <div>Wikipedia is a multilingual, web-based, free-content encyclopedia project supported by the Wikimedia Foundation and based on a model of openly editable content. EOL harvests articles from wikipedia that are indexed as species or higher taxa.</div> <br> <div>(lij) Ligurian en.wikipedia.org/wiki/Ligurian_(Romance_language)</div> <div>(lez) Lezgian en.wikipedia.org/wiki/Lezgian_language</div> <div>(sa) Sanskrit en.wikipedia.org/wiki/Sanskrit</div> <div>(ace) Acehnese en.wikipedia.org/wiki/Acehnese_language</div> <div>(diq) Zazaki en.wikipedia.org/wiki/Zaza_language</div> <div>(ce) Chechen en.wikipedia.org/wiki/Chechen_language</div> <div>(yo) Yoruba en.wikipedia.org/wiki/Yoruba_language</div> <div>(rw) Kinyarwanda en.wikipedia.org/wiki/Kinyarwanda</div> <div>(vec) Venetian en.wikipedia.org/wiki/Venetian_language</div> <div>(sc) Sardinian en.wikipedia.org/wiki/Sardinian_language</div> <div>(ln) Lingala en.wikipedia.org/wiki/Lingala</div> <div>(hak) Hakka en.wikipedia.org/wiki/Hakka_Chinese</div> <div>(kw) Cornish en.wikipedia.org/wiki/Cornish_language</div> <div>(bcl) Central Bicolano en.wikipedia.org/wiki/Central_Bikol</div> <div>(za) Zhuang en.wikipedia.org/wiki/Zhuang_languages</div> <div>(ang) Anglo-Saxon en.wikipedia.org/wiki/Anglo-Frisian_languages#English_(Anglo)_languages</div> <div>(eml) Emilian-Romagnol en.wikipedia.org/wiki/Emilian-Romagnol_language</div> <div>(av) Avar en.wikipedia.org/wiki/Avar_language</div> <div>(fj) Fijian en.wikipedia.org/wiki/Fijian_language</div> <div>(chy) Cheyenne en.wikipedia.org/wiki/Cheyenne_language</div> <div>(ik) Inupiak en.wikipedia.org/wiki/Inupiaq_language</div> <div>(zea) Zeelandic en.wikipedia.org/wiki/Zeelandic</div> <div>(bxr) Buryat en.wikipedia.org/wiki/Buryat_language</div> <div>(bjn) Banjar en.wikipedia.org/wiki/Banjar_language (bjn or bvu)</div> <div>(so) Somali en.wikipedia.org/wiki/Somali_language</div> <div>(zh-classical) Classical Chinese en.wikipedia.org/wiki/Classical_Chinese *(lzh)</div> <div>(mwl) Mirandese en.wikipedia.org/wiki/Mirandese_language</div> <div>(sn) Shona en.wikipedia.org/wiki/Shona_language</div> <div>(mai) Maithili en.wikipedia.org/wiki/Maithili_language</div> <div>(chr) Cherokee en.wikipedia.org/wiki/Cherokee_language</div> <div>(tk) Turkmen en.wikipedia.org/wiki/Turkmen_language</div> <div>(szy) Sakizaya en.wikipedia.org/wiki/Sakizaya_language</div> <div>(ab) Abkhazian en.wikipedia.org/wiki/Abkhaz_language</div> <div>(tcy) Tulu en.wikipedia.org/wiki/Tulu_language</div> <div>(wo) Wolof en.wikipedia.org/wiki/Wolof_language</div> <div>(ban) Balinese en.wikipedia.org/wiki/Balinese_language</div> <div>(ay) Aymara en.wikipedia.org/wiki/Aymara_language</div> <div>(tyv) Tuvan en.wikipedia.org/wiki/Tuvan_language</div> <div>(atj) Atikamekw en.wikipedia.org/wiki/Atikamekw_language</div> <div>(new) Newar en.wikipedia.org/wiki/Newar_language</div> <div>(fiu-vro) Võro en.wikipedia.org/wiki/Võro_language *(vro)</div> <div>(mg) Malagasy en.wikipedia.org/wiki/Malagasy_language</div> <div>(rm) Romansh en.wikipedia.org/wiki/Romansh_language</div> <div>(ltg) Latgalian en.wikipedia.org/wiki/Latgalian_language</div> <div>(ext) Extremaduran en.wikipedia.org/wiki/Extremaduran_language</div> <div>(kl) Greenlandic en.wikipedia.org/wiki/Greenlandic_language</div> <div>(roa-rup) Aromanian en.wikipedia.org/wiki/Aromanian_language *(rup)</div> <div>(nrm) Norman en.wikipedia.org/wiki/Norman_language</div> <div>(rn) Kirundi en.wikipedia.org/wiki/Kirundi</div> <div>(dty) Doteli en.wikipedia.org/wiki/Doteli</div> </div>
Wikipedia: wikipedia_combined_languages
<div> <div>Wikipedia is a multilingual, web-based, free-content encyclopedia project supported by the Wikimedia Foundation and based on a model of openly editable content. EOL harvests articles from wikipedia that are indexed as species or higher taxa.</div> <br> <div>en.wikipedia.org/wiki/Tamil_language (ta) Tamil</div> <div>en.wikipedia.org/wiki/Cebuano_language (ceb) Cebuano</div> <div>en.wikipedia.org/wiki/Greek_language (el) Greek</div> <div>en.wikipedia.org/wiki/Macedonian_language (mk) Macedonian</div> <div>en.wikipedia.org/wiki/Kyrgyz_language (ky) Kirghiz</div> <div>en.wikipedia.org/wiki/Scots_language Scots (sco)</div> <div>en.wikipedia.org/wiki/Hindi (fy) Hindi (hi)</div> <div>en.wikipedia.org/wiki/West_Frisian_language West Frisian</div> <div>en.wikipedia.org/wiki/Tagalog_language (tl) Tagalog</div> <div>en.wikipedia.org/wiki/Javanese_language (jv) Javanese</div> <div>en.wikipedia.org/wiki/Interlingua (ia) Interlingua</div> <div>en.wikipedia.org/wiki/Nepali_language (ne) Nepali</div> <div>en.wikipedia.org/wiki/Occitan_language (oc) Occitan</div> <div>en.wikipedia.org/wiki/Quechuan_languages (qu) Quechua</div> <div>en.wikipedia.org/wiki/Taraškievica (be-x-old) Belarusian (Taraškievica)</div> <div>en.wikipedia.org/wiki/Komi-Permyak_language (koi) Komi-Permyak</div> <div>en.wikipedia.org/wiki/North_Frisian_language (frr) North Frisian</div> <div>en.wikipedia.org/wiki/Udmurt_language (udm) Udmurt</div> <div>en.wikipedia.org/wiki/Bashkir_language (ba) Bashkir</div> <div>en.wikipedia.org/wiki/Aragonese_language (an) Aragonese</div> <div>en.wikipedia.org/wiki/Southern_Min (zh-min-nan) Min Nan</div> <div>en.wikipedia.org/wiki/Swahili_language (sw) Swahili</div> <div>en.wikipedia.org/wiki/Telugu_language (te) Telugu</div> <div>en.wikipedia.org/wiki/Uzbek_language (uz) Uzbek</div> <div>(bs) Bosnian en.wikipedia.org/wiki/Bosnian_language</div> <div>(ku) Kurdish en.wikipedia.org/wiki/Kurdish_languages</div> <div>(io) Ido en.wikipedia.org/wiki/Ido_language</div> <div>(my) Burmese en.wikipedia.org/wiki/Burmese_language</div> <div>(mn) Mongolian en.wikipedia.org/wiki/Mongolian_language</div> <div>(kv) Komi en.wikipedia.org/wiki/Komi_language</div> <div>(lb) Luxembourgish en.wikipedia.org/wiki/Luxembourgish</div> <div>(su) Sundanese en.wikipedia.org/wiki/Sundanese_language</div> <div>(kn) Kannada en.wikipedia.org/wiki/Kannada</div> <div>(tt) Tatar en.wikipedia.org/wiki/Tatar_language</div> <div>(sq) Albanian en.wikipedia.org/wiki/Albanian_language</div> <div>(csb) Kashubian en.wikipedia.org/wiki/Kashubian_language</div> <div>(mr) Marathi en.wikipedia.org/wiki/Marathi_language</div> <div>(co) Corsican en.wikipedia.org/wiki/Corsican_language</div> <div>(fo) Faroese en.wikipedia.org/wiki/Faroese_language</div> <div>(os) Ossetian en.wikipedia.org/wiki/Ossetian_language</div> <div>(cv) Chuvash en.wikipedia.org/wiki/Chuvash_language</div> <div>(kab) Kabyle en.wikipedia.org/wiki/Kabyle_language</div> <div>(sah) Sakha en.wikipedia.org/wiki/Yakut_language</div> <div>(nds) Low Saxon en.wikipedia.org/wiki/Low_German</div> <div>(lmo) Lombard en.wikipedia.org/wiki/Lombard_language</div> <div>(pa) Punjabi en.wikipedia.org/wiki/Punjabi_language</div> <div>(wa) Walloon en.wikipedia.org/wiki/Walloon_language</div> <div>(vls) West Flemish en.wikipedia.org/wiki/West_Flemish</div> <div>(gv) Manx en.wikipedia.org/wiki/Manx_language</div> <div>(wuu) Wu en.wikipedia.org/wiki/Wu_Chinese</div> <div>(nah) Nahuatl en.wikipedia.org/wiki/Nahuatl</div> <div>(as) Assamese en.wikipedia.org/wiki/Assamese_language</div> <div>(dsb) Lower Sorbian en.wikipedia.org/wiki/Lower_Sorbian_language</div> <div>(li) Limburgish en.wikipedia.org/wiki/Limburgish</div> <div>(mi) Maori en.wikipedia.org/wiki/Māori_language</div> <div>(kbd) Kabardian Circassian en.wikipedia.org/wiki/Kabardian_language</div> <div>(mdf) Moksha en.wikipedia.org/wiki/Moksha_language</div> <div>(to) Tongan en.wikipedia.org/wiki/Tongan_language</div> <div>(bat-smg) Samogitian en.wikipedia.org/wiki/Samogitian_dialect</div> <div>(olo) Livvi-Karelian en.wikipedia.org/wiki/Livvi-Karelian_language</div> <div>(mhr) Meadow Mari en.wikipedia.org/wiki/Meadow_Mari_language</div> <div>(tg) Tajik en.wikipedia.org/wiki/Tajik_language</div> <div>(pcd) Picard en.wikipedia.org/wiki/Picard_language</div> <div>(vep) Vepsian en.wikipedia.org/wiki/Veps_language</div> <div>(se) Northern Sami en.wikipedia.org/wiki/Northern_Sami_language</div> <div>(am) Amharic en.wikipedia.org/wiki/Amharic</div> <div>(si) Sinhalese en.wikipedia.org/wiki/Sinhala_language</div> <div>(ht) Haitian en.wikipedia.org/wiki/Haitian_language</div> <div>(gn) Guarani en.wikipedia.org/wiki/Guarani_language</div> <div>(rue) Rusyn en.wikipedia.org/wiki/Rusyn_language</div> <div>(mt) Maltese en.wikipedia.org/wiki/Maltese_language</div> <div>(gu) Gujarati en.wikipedia.org/wiki/Gujarati_language</div> <div>(als) Alemannic en.wikipedia.org/wiki/Alemannic_German</div> <div>(or) Oriya en.wikipedia.org/wiki/Odia_language</div> <div>(bh) Bhojpuri en.wikipedia.org/wiki/Bhojpuri_language</div> <div>(myv) Erzya en.wikipedia.org/wiki/Erzya_language</div> <div>(scn) Sicilian en.wikipedia.org/wiki/Sicilian_language</div> <div>(gd) Scottish Gaelic en.wikipedia.org/wiki/Scottish_Gaelic</div> <div>(pam) Kapampangan en.wikipedia.org/wiki/Kapampangan_language</div> <div>(xmf) Mingrelian en.wikipedia.org/wiki/Mingrelian_language</div> <div>(cdo) Min Dong en.wikipedia.org/wiki/Eastern_Min</div> <div>(bar) Bavarian en.wikipedia.org/wiki/Bavarian_language</div> <div>(nap) Neapolitan en.wikipedia.org/wiki/Neapolitan_language</div> <div>(lfn) Lingua Franca Nova en.wikipedia.org/wiki/Lingua_Franca_Nova</div> <div>(vo) Volapük en.wikipedia.org/wiki/Volapük</div> <div>(nds-nl) Dutch Low Saxon en.wikipedia.org/wiki/Dutch_Low_Saxon</div> <div>(bo) Tibetan en.wikipedia.org/wiki/Tibetan_language</div> <div>(stq) Saterland Frisian en.wikipedia.org/wiki/Saterland_Frisian</div> <div>(inh) Ingush en.wikipedia.org/wiki/Ingush_language</div> <div>(ha) Hausa en.wikipedia.org/wiki/Hausa_language</div> <div>(lbe) Lak en.wikipedia.org/wiki/Lak_language</div> </div>
Wikipedia: wikipedia-zh (Chinese)
Wikipedia is a multilingual, web-based, free-content encyclopedia project supported by the Wikimedia Foundation and based on a model of openly editable content. EOL harvests articles from wikipedia that are indexed as species or higher taxa.<p></p>維基百科(Wikipedia,聆聽i/ˌwɪkᵻˈpiːdi.ə/或聆聽i/ˌwɪkiˈpiːdi.ə/)是一個自由內容、公開編輯且多語言的網络百科全書協作计划,透過Wiki技術使得包括您在內的所有人都可以簡單地使用網頁瀏覽器修改其中的內容。維基百科的名称取自於本網站核心技術「Wiki」以及具有百科全書之意的「encyclopedia」共同創造出來的新混成詞「Wikipedia」,當前維基百科由维基媒體基金會負責運營。 維基百科主要是由来自互联网上的志願者共同合作編寫而成,任何使用网络進入維基百科的用户都可以編寫和修改裡面的文章,但是在一些情况下為了避免擾亂或者破壞可能會限制編輯功能。我們可以自由選擇使用匿名、化名或者直接用真實身份來編輯維基百科。與傳統的百科全書相比,在互联网上運作的維基百科其文字和絕大部分圖片使用創用CC 姓名標示-相同方式分享 3.0協議和GNU自由文件授權條款來提供每個人自由且免費的資訊,任何人都可以成為條目的作者,以及在遵守协议并標示來源後直接複製、使用以及发布這些內容。<p></p>https://zh.wikipedia.org
Wikipedia: Wikipedia English - traits (inferred records)
Wikipedia is a multilingual, web-based, free-content encyclopedia project supported by the Wikimedia Foundation and based on a model of openly editable content. EOL harvests articles from wikipedia that are indexed as species or higher taxa.<p></p>
Sentence/Table Pair Data from Wikipedia for Pre-training with Distant-Supervision
<p>This is the dataset used for pre-training in "<em>ReasonBERT: Pre-trained to Reason with Distant Supervision</em>", EMNLP'21.</p> <p>There are two files:</p> <p>sentence_pairs_for_pretrain_no_tokenization.tar.gz -> Contain only sentences as evidence, Text-only</p> <p>table_pairs_for_pretrain_no_tokenization.tar.gz -> At least one piece of evidence is a table, Hybrid</p> <p>The data is chunked into multiple tar files for easy loading. We use <a href="https://github.com/webdataset/webdataset">WebDataset</a>, a PyTorch Dataset (IterableDataset) implementation providing efficient sequential/streaming data access.</p> <p>For pre-training code, or if you have any questions, please check our GitHub repo https://github.com/sunlab-osu/ReasonBERT</p> <p>Below is a sample code snippet to load the data</p> <pre><code class="language-python">import webdataset as wds # path to the uncompressed files, should be a directory with a set of tar files url = './sentence_multi_pairs_for_pretrain_no_tokenization/{000000...000763}.tar' dataset = ( wds.Dataset(url) .shuffle(1000) # cache 1000 samples and shuffle .decode() .to_tuple("json") .batched(20) # group every 20 examples into a batch ) # Please see the documentation for WebDataset for more details about how to use it as dataloader for Pytorch # You can also iterate through all examples and dump them with your preferred data format</code></pre> <p>Below we show how the data is organized with two examples.</p> <p>Text-only</p> <pre><code>{'s1_text': 'Sils is a municipality in the comarca of Selva, in Catalonia, Spain.', # query sentence 's1_all_links': { 'Sils,_Girona': [[0, 4]], 'municipality': [[10, 22]], 'Comarques_of_Catalonia': [[30, 37]], 'Selva': [[41, 46]], 'Catalonia': [[51, 60]] }, # list of entities and their mentions in the sentence (start, end location) 'pairs': [ # other sentences that share common entity pair with the query, group by shared entity pairs { 'pair': ['Comarques_of_Catalonia', 'Selva'], # the common entity pair 's1_pair_locs': [[[30, 37]], [[41, 46]]], # mention of the entity pair in the query 's2s': [ # list of other sentences that contain the common entity pair, or evidence { 'md5': '2777e32bddd6ec414f0bc7a0b7fea331', 'text': 'Selva is a coastal comarque (county) in Catalonia, Spain, located between the mountain range known as the Serralada Transversal or Puigsacalm and the Costa Brava (part of the Mediterranean coast). Unusually, it is divided between the provinces of Girona and Barcelona, with Fogars de la Selva being part of Barcelona province and all other municipalities falling inside Girona province. Also unusually, its capital, Santa Coloma de Farners, is no longer among its larger municipalities, with the coastal towns of Blanes and Lloret de Mar having far surpassed it in size.', 's_loc': [0, 27], # in addition to the sentence containing the common entity pair, we also keep its surrounding context. 's_loc' is the start/end location of the actual evidence sentence 'pair_locs': [ # mentions of the entity pair in the evidence [[19, 27]], # mentions of entity 1 [[0, 5], [288, 293]] # mentions of entity 2 ], 'all_links': { 'Selva': [[0, 5], [288, 293]], 'Comarques_of_Catalonia': [[19, 27]], 'Catalonia': [[40, 49]] } } ,...] # there are multiple evidence sentences }, ,...] # there are multiple entity pairs in the query }</code></pre> <p>Hybrid</p> <pre><code>{'s1_text': 'The 2006 Major League Baseball All-Star Game was the 77th playing of the midseason exhibition baseball game between the all-stars of the American League (AL) and National League (NL), the two leagues comprising Major League Baseball.', 's1_all_links': {...}, # same as text-only 'sentence_pairs': [{'pair': ..., 's1_pair_locs': ..., 's2s': [...]}], # same as text-only 'table_pairs': [ 'tid': 'Major_League_Baseball-1', 'text':[ ['World Series Records', 'World Series Records', ...], ['Team', 'Number of Series won', ...], ['St. Louis Cardinals (NL)', '11', ...], ...] # table content, list of rows 'index':[ [[0, 0], [0, 1], ...], [[1, 0], [1, 1], ...], ...] # index of each cell [row_id, col_id]. we keep only a table snippet, but the index here is from the original table. 'value_ranks':[ [0, 0, ...], [0, 0, ...], [0, 10, ...], ...] # if the cell contain numeric value/date, this is its rank ordered from small to large, follow TAPAS 'value_inv_ranks': [], # inverse rank 'all_links':{ 'St._Louis_Cardinals': { '2': [ [[2, 0], [0, 19]], # [[row_id, col_id], [start, end]] ] # list of mentions in the second row, the key is row_id }, 'CARDINAL:11': {'2': [[[2, 1], [0, 2]]], '8': [[[8, 3], [0, 2]]]}, } 'name': '', # table name, if exists 'pairs': { 'pair': ['American_League', 'National_League'], 's1_pair_locs': [[[137, 152]], [[162, 177]]], # mention in the query 'table_pair_locs': { '17': [ # mention of entity pair in row 17 [ [[17, 0], [3, 18]], [[17, 1], [3, 18]], [[17, 2], [3, 18]], [[17, 3], [3, 18]] ], # mention of the first entity [ [[17, 0], [21, 36]], [[17, 1], [21, 36]], ] # mention of the second entity ] } } ] }</code></pre> <p> </p> <p> </p> <p> </p>
COSI-Article matrix: linking ISCB Communities of Special Interest to Wikipedia
<p>Wikipedia is regarded as one of the most important channels for the public communication of science; English Wikipedia has around 1,500 articles relating to computational biology, which are frequently accessed as an educational resource. Joint efforts between the International Society for Computational Biology (ISCB) and the Computational Biology taskforce of WikiProject Molecular Biology (a group of expert Wikipedia editors) have considerably improved computational biology representation on Wikipedia in recent years. However, there is still an urgent need for further quality improvement, primarily while comparing to related scientific fields such as genetics and medicine. Facilitating the involvement of members from ISCB COSIs (Communities of Special Interest) would improve a vital open educational resource in computational biology, additionally allowing COSIs to provide a quality educational resource particular to their subfield.</p> <p>This first version of the COSI-Article matrix is a binary matrix identifying relevant ISCB COSIs for all Wikipedia articles relating to computational biology, defining a domain-specific open educational resource for each COSI. In addition, quality and importance ratings for each article allow identification of areas where domain experts could improve computational biology representation.</p>
TWikiL - Twitter Wikipedia Link Dataset
<p>The Twitter Wikipedia Link (TWikiL) dataset contains all Tweets posted on Twitter that contain a Wikipedia URL. The data was collected via Twitters academic research access and spans 15 years of Tweets from March 2006 to January 2021. TWikiL comes in two versions: <strong>TWikiL_raw</strong> is a list of Tweet IDs in CSV format. <strong>TWikiL_curated</strong> is an SQLite database, which is a curated version of TWikiL containing only links to Wikipedia articles. The curated version has been augmented with the language edition that the URL in the Tweet links to, the Wikidata identifier and a Wikipedia topic category. <br> <br> TWikiL raw contains 44,945,098 Tweet IDs<br> TWikiL curated contains 35,252,782 URLs/Wikidata concepts with 34,543,612 unique Tweets and 474,577 Tweets linking to multiple Wikipedia articles.</p>
Wikipedia Talk Page 'Climate Change' Sentiment and Toxicity Dataset
<p>The given dataset was prepared as part of a master's research project under the Master's program in Computational Social Systems at RWTH Aachen University.</p> <p>The talk page was parsed using the <a href="https://aclanthology.org/E17-3006/">GraWiTas tool</a> in JSON format. The file <em>Climate_change.comment_list.json </em> is raw export of discussions that needs to be cleaned before using it for calculating sentiment and toxicity scores.</p> <p>The sentiment scores were calculated using VADER and toxicity scores using the Perspective API by Google.</p> <p>Date & time of dump is: 27-06-2022 12:12 UTC+02:00</p>
DOIs linked by the English Wikipedia which could be made available in green Open Access
<p>List of citations from the English Wikipedia articles extracted from the enwiki-20170720-pages-articles XML dump via https://pypi.org/project/mwcites/ , DOIs cleaned with custom regular expressions.</p> <p>The list of scholarly publications identified by the DOIs has been filtered to exclude those which are already available in Open Access and those which may not be depositable according to SHERPA/RoMEO policy summaries, first via the Dissemin API and then by the oaDOI API, with the attached Python script (https://github.com/nemobis/bots/blob/master/doi-doai-openaccess.py ).</p> <p>This produced a list of 194913 DOIs available in open access and 430230 DOIs unavailable and depositable (as of 2017-08-22 data, which for oaDOI was partly v1 and partly v2).</p>
Wikipedia. Events and collective memory detection dataset
<p>This is the data accompanying the code (https://github.com/mizvol/WikiBrain), required to reproduce results of "Wikipedia graph mining: dynamic structure of collective memory" paper (https://arxiv.org/abs/1710.00398).</p>
Collated set of all gravitational wave observations (to current date) from Wikipedia
<p>This is a dataset I needed but couldn't find. So, with the help of Chat GPT (and several hours ot time), I compiled this easy to access html file.</p> <p>It's not perfect but here it is for all of you to enjoy. </p>
Webis Wikipedia Text Reuse Corpus 2018 (Webis-Wikipedia-Text-Reuse-18)
<p>The Wikipedia Text Reuse Corpus 2018 (Webis-Wikipedia-Text-Reuse-18) containing text reuse cases extracted from within Wikipedia and in between Wikipedia and a sample of the Common Crawl.</p> <p>The corpus has following structure:</p> <ul> <li>wikipedia.jsonl.bz2: Each line, representing a Wikipedia article, contains a json array of article_id, article_title, and article_body</li> <li>within-wikipedia-tr-01.jsonl.bz2: Each line, representing a text reuse case, contains a json array of s_id (source article id), t_id (target article id), s_text (source text), t_text (target text)</li> <li>within-wikipedia-tr-02.jsonl.bz2: Each line, representing a text reuse case, contains a json array of s_id (source article id), t_id (target article id), s_text (source text), t_text (target text)</li> <li>preprocessed-web-sample.jsonl.xz: Each line, representing a web page, contains a json object of d_id, d_url, and content</li> <li>without-wikipedia-tr.jsonl.bz2: Each line, representing a text reuse case, contains a json array of s_id (Wikipedia article id), d_id (web page id), s_text (article text), d_content (web page content)</li> </ul> <p>The datasets were extracted in the work by Alshomary et al. 2018 that aimed to study the text reuse phenomena related to Wikipedia at scale. A pipeline for large scale text reuse extraction was developed and used on Wikipedia and the CommonCrawl.</p>
Wikipedia: wikipedia-war (Waray-Waray)
Wikipedia is a multilingual, web-based, free-content encyclopedia project supported by the Wikimedia Foundation and based on a model of openly editable content. EOL harvests articles from wikipedia that are indexed as species or higher taxa.<p></p><p></p>https://war.wikipedia.org
Wikipedia: wikipedia-sl (Slovenian)
Wikipedia is a multilingual, web-based, free-content encyclopedia project supported by the Wikimedia Foundation and based on a model of openly editable content. EOL harvests articles from wikipedia that are indexed as species or higher taxa.<p></p><p></p>https://sl.wikipedia.org
Wikipedia: wikipedia-ms (Malay)
Wikipedia is a multilingual, web-based, free-content encyclopedia project supported by the Wikimedia Foundation and based on a model of openly editable content. EOL harvests articles from wikipedia that are indexed as species or higher taxa.<p></p>
Dataset for the paper "Entity Insertion in Multilingual Linked Corpora: The Case of Wikipedia"
<p>Dataset for the EMNLP'24 Main conference paper titled "Entity Insertion in Multilingual Linked Corpora: The Case of Wikipedia".</p>
Wikipedia: wikipedia-gl (Galician)
Wikipedia is a multilingual, web-based, free-content encyclopedia project supported by the Wikimedia Foundation and based on a model of openly editable content. EOL harvests articles from wikipedia that are indexed as species or higher taxa.<p></p><p></p>https://gl.wikipedia.org/
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.