Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
47
datasets available to search
ShareScore release 0.9.0
Dataset results
47 results for “wiki”
Scrubbed data on Wikipedians in Residence in Libraries based on the Mapping GLAM-Wiki collaborations
<p><strong>Source</strong>: </p> <p><a href="https://docs.google.com/spreadsheets/d/1UVN-T19g5tE7cONFCkiBkquBJiecoU-w4Rb6F-qR6II/edit#gid=791098161">GLAM-Wiki Activities Mapping - Community Review and Feedback Sheet</a></p> <p><strong>Source's context: </strong></p> <p>Gill, Satdeep. ‘Mapping GLAM-Wiki Collaborations’. <em>This Month in GLAM</em>, March 2020. <a href="https://outreach.wikimedia.org/wiki/GLAM/Newsletter/March_2020/Contents/WMF_GLAM_report">https://outreach.wikimedia.org/wiki/GLAM/Newsletter/March_2020/Contents/WMF_GLAM_report</a>.</p> <p> </p> <p>Data was scrubbed using <a href="https://openrefine.org/download.html">OpenRefine 3.4.1</a></p> <p>The original spreadsheet had only partial information in many fields and it is a work in progress (for more see the "source's context" link above).</p> <p>I have only manually double checked those rows in which the “Primary partner institution” contains the stem “libr*” or “bibli*. The following eight rows where modified and “Library” was added in the “Type of institution” column: Municipal Library, Patiala, BRAU Library of the University of Naples Federico II, Library and Archives Canada, Eötvös Loránd University Library and Archives, National Health Library and Knowledge Service, National Doctors Training and Planning, Daniel Cosío Villegas Library, Cantonal and University Library, Nationaal Archief | Koninklijke Bibliotheek; and “Library association” was added to the Online Computer Library Center (OCLC) entry. All changes can be seen in the <a href="https://zenodo.org/api/files/ae475409-a2a8-4e51-98e2-68b3090d0fd0/WiRs-in-libraries_MGW_scrubbing-changes.json">WiRs-in-libraries_MGW_scrubbing-changes.json</a> file in this release.</p>
wiki-category-consistency-eval
<p>Experiment results produced in the context of analyzing the consistency between Wikipedia and Wikidata categories using the Wikidata JSON dump of 2022-05-02 and the Wikipedia SQL dumps of 2022-05-01.</p> <p>Detailed information can be found on the <a href="https://github.com/fusion-jena/wiki-category-consistency">Github page</a>.</p>
wiki-category-consistency-dataset
<p>Candidate generation and cleaning results produced in the context of analyzing the consistency between Wikipedia and Wikidata categories using the Wikidata JSON dump of 2022-05-02 and the Wikipedia SQL dumps of 2022-05-01.</p> <p>Detailed information can be found on the <a href="https://github.com/fusion-jena/wiki-category-consistency">Github page</a>.</p>
Adverse Outcome Pathway Wiki RDF
<p>This dataset is the RDF generated from the AOP-Wiki data release (<a href="https://aopwiki.org/downloads">aopwiki.org/downloads</a>). It was generated using a Jupyter notebook that is available on GitHub (<a href="https://github.com/marvinm2/AOPWikiRDF">github.com/marvinm2/AOPWikiRDF</a>), and the process and additional description of the RDF have been published (<a href="https://doi.org/10.1089/aivt.2021.0010">doi.org/10.1089/aivt.2021.0010</a>).</p>
Wiki-talk Datasets
<p>User interaction networks of Wikipedia of 28 different languages. Nodes (orininal wikipedia user IDs) represent users of the Wikipedia, and an edge from user A to user B denotes that user A wrote a message on the talk page of user B at a certain timestamp.</p> <p>More info: http://yfiua.github.io/academic/2016/02/14/wiki-talk-datasets.html</p>
MineDojo Internet Knowledge Base (Wiki)
<p><strong>Project website:</strong> <a href="https://minedojo.org">minedojo.org</a></p> <p><strong>Paper:</strong> <a href="https://arxiv.org/abs/2206.08853">arxiv.org/abs/2206.08853</a></p> <p><strong>GitHub:</strong> <a href="https://github.com/MineDojo/MineDojo">github.com/MineDojo/MineDojo</a></p> <p>The Minecraft Wiki pages cover almost every aspect of the game mechanics, and supply a rich source of unstructured knowledge in multimodal tables, recipes, illustrations, and step-by-step tutorials. <strong>We scrape 6,735 pages that interleave text, images, tables, and diagrams.</strong> To preserve the layout information, we also save the screenshots of entire pages and extract bounding boxes of the visual elements.</p> <p>There are two files in our Wiki knowledge base.</p> <ul> <li><strong>wiki_samples.zip:</strong> A sample version of the full knowledge base (10 pages). </li> <li><strong>wiki_full.zip:</strong> The full knowledge base (6,735 pages). </li> </ul> <p>Check out our paper!</p> <pre><code class="language-markdown">@article{fan2022minedojo, title = {MineDojo: Building Open-Ended Embodied Agents with Internet-Scale Knowledge}, author = {Linxi Fan and Guanzhi Wang and Yunfan Jiang and Ajay Mandlekar and Yuncong Yang and Haoyi Zhu and Andrew Tang and De-An Huang and Yuke Zhu and Anima Anandkumar}, year = {2022}, journal = {arXiv preprint arXiv: Arxiv-2206.08853} }</code></pre>
Adverse Outcome Pathway Wiki RDF
<p>This dataset is the RDF generated from the AOP-Wiki data release (<a href="https://aopwiki.org/downloads">aopwiki.org/downloads</a>). It was generated using a Jupyter notebook that is available on GitHub (<a href="https://github.com/marvinm2/AOPWikiRDF">github.com/marvinm2/AOPWikiRDF</a>), and the process and additional description of the RDF have been published (<a href="https://doi.org/10.1089/aivt.2021.0010">doi.org/10.1089/aivt.2021.0010</a>).</p>
Wikidata 3 Topical Subsets (Gene Wiki, Music, Ships) and 4 Random Subsets
<p>This dataset contains the N-Triples files of 3 Wikidata topical subsets corresponding to 3 Wikidata WikiProject: Gene Wiki, Music, and Ships along with 4 random subsets in different sizes: two of 100K items, one 500K items, and one 1M items. Subsets are extracted from the <a href="https://academictorrents.com/details/229cfeb2331ad43d4706efd435f6d78f40a3c438">3 January 2022 dump</a>. All subsets have been extracted with <a href="https://github.com/seyedahbr/wdumper">WDumper</a> using these <a href="https://github.com/seyedahbr/RQSS_Evaluation/tree/main/WDumper%20Specification%20Files">JSON specification files</a>. The files are:</p> <ul> <li>GeneWiki.zip: contains 25 `.nt.gz` RDF files each of which corresponds to one of the main Gene Wiki WikiProject classes, e.g. protein, gene, chemical compound, etc.</li> <li>music.nt.gz: the RDF file corresponding to the Music WikiProject.</li> <li>ships.nt.gz: the RDF file corresponding to the Ships WikiProject.</li> <li>Random100K_1.zip: contains 2 `nt.gz` RDF files each of which includes (about) 50,000 random Wikidata items, 100,000 items in total.</li> <li>Random100K_2.zip: contains 2 `nt.gz` RDF files each of which includes (about) 50,000 random Wikidata items, 100,000 items in total.</li> <li>Random500K.zip: contains 10 `nt.gz` RDF files each of which includes (about) 50,000 random Wikidata items, 500,000 items in total.</li> <li>Random1M.zip: contains 20 `nt.gz` RDF files each of which includes (about) 50,000 random Wikidata items, 1,000,000 items in total.</li> </ul> <p> </p>
Output files of the RQSS extractor and framework on 3 Topical (Gene Wiki, Music, Ships) subsets and 4 Random Subsets
<p>This dataset contains the `.csv` output files of the Referencing Quality Scoring System - RQSS performed on 3 topical subsets (Gene Wiki, Music, Ships) and 4 random subsets. Subset RDF files are in <a href="https://doi.org/10.5281/zenodo.7332161">this dataset</a>.</p>
Wiki-based Communities of Interest: Demographics and Outliers
<p>These datasets contains statements about demographics and outliers of Wiki-based Communities of Interest. </p> <p><strong>Group-centric dataset (sample):</strong></p> <pre><code class="language-json">{ "title": "winners of Priestley Medal", "recorded_members": 83, "topics": ["STEM.Chemistry"], "demographics": [ "occupation-chemist", "gender-male", "citizen-U.S." ], "outliers": [ { "reason": "NOT(chemist) unlike 82 recorded members", "members": [ "Francis Garvan (lawyer, art collector)" ] }, { "reason": "NOT(male) unlike 80 recorded members", "members": [ "Mary L. Good (female)", "Darleane Hoffman (female)", "Jacqueline Barton (female)" ] } ] }</code></pre> <p><strong>Subject-centric dataset (sample):</strong></p> <pre><code class="language-json">{ "subject": "Serena Williams", "statements": [ { "statement": "NOT(sport-basketball) but (tennis) unlike 4 recorded winners of Best Female Athlete ESPY Award.", "score": 0.36 }, { "statement": "NOT(occupation-politician) but (tennis player, businessperson, autobiographer) unlike 20 recorded winners of Michigan Women's Hall of Fame.", "score": 0.17 } ] }</code></pre> <p><strong>This data can be also browsed at: <a href="https://wikiknowledge.onrender.com/demographics/">https://wikiknowledge.onrender.com/demographics/</a></strong></p>
Wikimedia wikis' wiki segmentation dataset
<p>We maintain a <strong>wiki comparison dataset</strong> (which we used to call a wiki <em>segmentation</em> dataset) to show a simple, snapshot comparison of our wikis (as opposed to, say, Wikistats 2, which is mean to show simple trends within individual wikis or wiki groups).</p>
Noscemus Wiki
<p><a href="https://wiki.uibk.ac.at/noscemus/Main_Page">This database</a> in the form of a Semantic Media Wiki (SMW) presents a <strong>selection of early modern works on science written in Latin</strong> along with their authors and some secondary literature. This selection does not constitute an anthology, as works famous for their authors, their groundbreaking character or other reasons are neither included nor excluded in a systematic way. Rather, the collection is intended to be <strong>representative of Latin literature’s engagement with contemporary science</strong> in terms of chronological spread, literary forms and scientific disciplines. In accordance with the project’s epistemological interest, the descriptions of the single works focus on their <strong>literary features</strong> such as structure, style and terminology. Their contents are outlined as far as necessary to that end.</p> <p>The <a href="https://www.uibk.ac.at/projects/noscemus/">NOSCEMUS</a> project has received funding from the European Research Council (ERC) under the European Union’s Horizon 2020 research and innovation programme (grant agreement No. [741374]).</p>
Wiki-MLM: Multiple Languages and Modalities
<p>International organizations and companies encounter web data in a range of modalities and languages. At the same time, applications developed by these users result from pipelines that perform multiple tasks. Systems that handle diverse inputs and multiple objectives hold the promise of limiting complexity in the application work-flow and improving generalization. We present a Wikidata-generated re-source designed to train and evaluate multitask systems on samples in four modalities and three languages</p>
EAGLE Media Wiki RDF data from Wikibase
<p>This is a dump of the triples entered with the Wikibase Extension into the EAGLE Media Wiki for translations.</p> <p>https://wiki.eagle-network.eu/wiki/Main_Page </p> <p>The same data is accessible via the Mediawiki API. </p> <p>It is part of the EAGLE project https://www.eagle-network.eu/.</p> <p> </p> <p>Contributors of the translations are in the data.</p>
Wikidata Dump wiki_misc
<p> RDF dump of wikidata produced with <a href="https://tools.wmflabs.org/wdumps/">wdumps</a>. </p> <p> <br> <a href="https://tools.wmflabs.org/wdumps/dump/1813">View on wdumper</a> </p> <p> <b>entity count</b>: 0, <b>statement count</b>: 0, <b>triple count</b>: 0 </p>
Gene Wiki subsets
<p>Gene Wiki Subset of Wikidata created with wdsub (https://github.com/weso/wdsub).</p> <p>Subsets are described using Shape Expressions: https://github.com/weso/genewikisub/tree/master/shex/directP31</p> <p>Input: https://dumps.wikimedia.org/wikidatawiki/entities/latest-all.json.gz downloaded between 2021/12/06 10:26:10 pm (CET) and 2021/12/07 08:40:42 am (CET)</p>
Wikidata Subsets of 4 Gene Wiki Classes
<p>Chemical compound (Q11173), disease (Q12136), gene (Q7187), and protein (Q8054) are some of the main classes containing the Gene Wiki WikiProjects. In this repository, we put the corresponding subsets of each class. The subsets contain all instances of the four main classes (no sub-classes). All subsets are extracted from the Wikidata JSON dump of 3 January 2022 (<a href="https://t.co/vdmJc8V1v2">https://t.co/vdmJc8V1v2</a>)</p>
Wikidata Subsets of 4 Gene Wiki Classes (wdsub)
<p>Chemical compound (Q11173), disease (Q12136), gene (Q7187), and protein (Q8054) are some of the main classes containing the Gene Wiki WikiProjects. In this repository, we put the corresponding subsets of each class. The subsets contain all instances of the four main classes (no sub-classes). All subsets are extracted from the Wikidata JSON dump of 3 January 2022 (<a href="https://t.co/vdmJc8V1v2">https://t.co/vdmJc8V1v2</a>) using wdsub (https://github.com/weso/wdsub)</p>
wiki-category-consistency-cache
<p>A collection of SQLite database files containing all the data retrieved from the Wikidata JSON dump of 2022-05-02 and the Wikipedia SQL dumps of 2022-05-01 in the context of analyzing the consistency between Wikipedia and Wikidata categories.<br> <br> Detailed information can be found on the <a href="https://github.com/fusion-jena/wiki-category-consistency">Github page</a>.</p>
Wikidata Dump partial-wiki
<p>RDF dump of wikidata produced with <a href="//wdumps.toolforge.org/">wdumper</a>.</p><p>basic filter<br><a href="//wdumps.toolforge.org/dump/2592">View on wdumper</a></p><p><b>entity count<b>: 0, <b>statement count</b>: 0, <b>triple count</b>: 38</b></b></p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.