Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
47
datasets available to search
ShareScore release 0.9.0
Dataset results
47 results for “wiki”
Wikidata Dump wiki_en
<p>RDF dump of wikidata produced with <a href="//wdumps.toolforge.org/">wdumper</a>.</p><p><br><a href="//wdumps.toolforge.org/dump/2702">View on wdumper</a></p><p><b>entity count<b>: 0, <b>statement count</b>: 0, <b>triple count</b>: 0</b></b></p>
Wikidata Dump wiki_en
<p>RDF dump of wikidata produced with <a href="//wdumps.toolforge.org/">wdumper</a>.</p><p><br><a href="//wdumps.toolforge.org/dump/2702">View on wdumper</a></p><p><b>entity count<b>: 0, <b>statement count</b>: 0, <b>triple count</b>: 0</b></b></p>
Wiki-Disease-Benchmark
<p>This benchmark consist in 255 randomly selected disease descriptions, as of February 2024. Each disease description was labeled by two data annotators who reviewed each other's annotations to ensure accuracy and consistency across the dataset. </p> <p>This procedure involves collecting, parsing and extracting data from Wikipedia using a software routine that interfaces with an API \footnote{https://pypi.org/project/Wikipedia-API/} to systematically retrieve and collate information related to a predefined disease. Specifically, it searches for pages with a certain disease and, within those pages, extracts the "Sings and Symptoms" section.</p> <p>This process has two steps:</p> <ul> <li>Retrieve all the labels rdfs:label of triples in DBpedia \footnote{https://dbpedia.org/} that are a disease rdf:type dbo:Disease.</li> <li>With these labels, go to each page of Wikipedia and scrape the section "Signs and Symptoms".<br><br></li> </ul> <p>After extracting the text from Wikipedia, the phenotypical entities were annotated.</p>
Supplementary tables for the paper: "Comprehensive Mapping of the AOP-Wiki Database: Identifying Biological and Disease Gaps"
<p><strong>Supplementary tables for the paper: "Comprehensive Mapping of the AOP-Wiki Database: Identifying Biological and Disease Gaps"</strong></p> <p><em>Original Research Article</em><br><strong>Frontiers in Toxicology</strong>, March 8, 2024<br>Section: Regulatory Toxicology<br><strong>Volume 6 - 2024</strong> | <a href="https://doi.org/10.3389/ftox.2024.1285768" target="_new" rel="noopener">https://doi.org/10.3389/ftox.2024.1285768</a></p> <p><strong>Authors</strong>:<br>Thomas Jaylet, Thibaut Coustillet, Nicola M. Smith, Barbara Viviani, Birgitte Lindeman, Lucia Vergauwen, Oddvar Myhre, Nurettin Yarar, Johanna M. Gostner, Pablo Monfort-Lanzas, Florence Jornod, Henrik Holbech, Xavier Coumoul, Dimosthenis A. Sarigiannis, Philipp Antczak, Anna Bal-Price, Ellen Fritsche, Eliska Kuchovska, Antonios K. Stratidakis, Robert Barouki, Min Ji Kim, Olivier Taboureau, Marcin W. Wojewodzic, Dries Knapen, Karine Audouze</p>
Wikidata Subsets of 6 Wikiproject (Gene Wiki, Taxonomy, Astronomy, Music, Law, Ships)
<p>This dataset contains the N-Triples files of 6 Wikidata subsets corresponding to 6 different Wikidata WikiProject extracted from October 2016 and June 2021 dumps. For each project, there are .nt.gz files containing RDF data. The projects are:</p> <ul> <li>Gene Wiki: There is a GeneWiki_2016.zip (extracted from the 2016 dump) and a GeneWiki_2021.zip (extracted from the 2021 dump), Each of which contains 25 .nt.gz files extracted by the WDumper tool.</li> <li>Astronomy: Astronomy_2016.nt.gz (extracted from the 2016 dump) and Astronomy_2021.nt.gz (extracted from the 2021 dump).</li> <li>Taxonomy: There is a Taxonomy_2016.zip (extracted from the 2016 dump) and a Taxonomy_2021.zip (extracted from the 2021 dump), Each of which contains two .nt.gz files.</li> <li>Law: Law_2016.nt.gz (extracted from the 2016 dump) and Law_2021.nt.gz (extracted from the 2021 dump).</li> <li>Music: Music_2016.nt.gz (extracted from the 2016 dump) and Music_2021.nt.gz (extracted from the 2021 dump).</li> <li>Ships: Ships_2016.nt.gz (extracted from the 2016 dump) and Ships_2021.nt.gz (extracted from the 2021 dump).</li> </ul> <p> </p> <ul> </ul>
Wikidata Dump wiki-songwriter
<p> RDF dump of wikidata produced with <a href="https://tools.wmflabs.org/wdumps/">wdumps</a>. </p> <p> <br> <a href="https://tools.wmflabs.org/wdumps/dump/1707">View on wdumper</a> </p> <p> <b>entity count</b>: 0, <b>statement count</b>: 0, <b>triple count</b>: 0 </p>
Wikidata Dump full_wiki_dump
<p>RDF dump of wikidata produced with <a href="//wdumps.toolforge.org/">wdumper</a>.</p><p><br><a href="//wdumps.toolforge.org/dump/2860">View on wdumper</a></p><p><b>entity count<b>: 0, <b>statement count</b>: 0, <b>triple count</b>: 0</b></b></p>
External References of English Wikipedia (ref-wiki-en)
<p><strong>External References of English Wikipedia </strong>(<strong>ref-wiki-en</strong>) is a corpus of the plain-text content of 2,475,461 external webpages linked from the reference section of articles in English Wikipedia. Specifically:</p> <ul> <li>32,329,989 external reference URLs were extracted from a <a href="https://zenodo.org/record/3483254">2018 HTML dump of English Wikipedia</a>. Removing repeated and ill-formed URLs yielded 23,036,318 unique URLs.</li> <li>These URLs were filtered to remove file extensions for unsupported formats (videos, audio, etc.), yielding 17,781,974 downloadable URLs. The URLs were loaded into <a href="http://nutch.apache.org/">Apache Nutch</a> and continuously downloaded from August 2019 to December 2019, resulting in 2,475,461 successfully downloaded URLs. Not all URLs could be accessed. The order in which URLs were accessed was determined by Nutch, which partitions URLs by <a href="https://docs.oracle.com/javase/7/docs/api/java/net/URL.html#getHost()">host</a> and then randomly chooses amongst the URLs for each host.</li> <li>The content of these webpages were indexed in <a href="https://lucene.apache.org/solr/">Apache Solr</a> by Nutch. From Solr we extracted a JSON dump of the content.</li> <li>Many URLs offer a redirect; unfortunately Nutch does not index redirect information. This means that connecting the Wikipedia article (with the pre-direct link) to the downloaded webpage (at the post-redirect link) was complicated. However, by inspecting the order of download in the Nutch log files, we managed to recover links for 2,058,896 documents (83%) from their original Wikipedia article(s).</li> <li>We further managed to associate 3,899,953 unique Wikidata items with at least one external reference webpage in the corpus.</li> </ul> <p>The ref-en-wiki corpus is incomplete, i.e., we did not attempt to download all reference URLs for English Wikipedia. We thus also collect a smaller complete corpus for the external references of 5,000 Wikipedia articles (<strong>ref-wiki-en-5k</strong>). We sampled from 5 ranges of Wikidata items: Q1-10000, Q10001-100000, Q100001-1000000, Q1000001-10000000, and Q10000001-100000000. From each range we sampled 1000 items. We then scraped the external reference URLs for the Wikipedia article corresponding to these items and downloaded them. The resulting corpus contains 37,983 webpages.</p> <p>Each line of the corpus (ref-wiki-en, ref-wiki-en-5k) encodes the webpage of an external reference in JSON format. Specifically, we provide:</p> <ul> <li><em>tstamp</em>: When the webpage was accessed</li> <li><em>host</em>: The domain (FQDN post-redirect) from which the webpage was retrieved.</li> <li><em>title</em>: The title (meta) of the document</li> <li><em>url</em>: The URL (post-redirect) of the webpage</li> <li><em>Q</em>: The Q-code identifiers of the Wikidata items whose corresponding Wikipedia article is confirmed to link to this webpage.</li> <li><em>content</em>: A plain-text encoding of the content of the webpage.</li> </ul> <p>Below we provide an abbreviated example of a line from the corpus:</p> <pre><code>{"tstamp":"2019-09-26T01:22:43.621Z","host":"geology.isu.edu","title":"Digital Geology of Idaho - Basin And Range","url":"http://geology.isu.edu/Digital_Geology_Idaho/Module9/mod9.htm","Q":[810178],"content":"Digital Geology of Idaho - Basin And Range\n1 - Idaho Basement Rock\n2 - Belt Supergroup\n3 - Rifting & Passive Margin\n4 - Accreted Terranes\n5 - Thrust Belt\n6 - Idaho Batholith\n7 - North Idaho & Mining\n8 - Challis Volcanics\n9 - Basin and Range\n10 - Columbia River Basalts\n11 - SRP & Yellowstone\n12 - Pleistocene Glaciation\n13 - Palouse & Lake Missoula\n14 - Lake Bonneville Flood\n15 - Snake River Plain Aquifer\nBasin and Range Province - Teritiary Extension\nGeneral geology of the Basin and Range Province\nMechanisms of Basin and Range faulting\nIdaho Basin and Range south of the Snake River Plain\nIdaho Basin and Range north of the Snake River Plain\nLocal areas of active and recent Basin & Range faulting: Borah Peak\nPDF Slideshows: North of SRP , South of SRP , Borah Earthquake\nFlythroughs: Teton Valley , Henry's Fork , Big Lost River , Blackfoot , Portneuf , Raft River Valley , Bear River , Salmon Falls Creek , Snake River , Big Wood River\nVocabulary Words\nthrust fault\nBasin and Range\nSnake River Plain\nhalf-graben\ntransfer zone\n \n \n \n \nFly-throughs\nGeneral geology of the Basin and Range Province\nThe Basin and Range Province generally includes most of eastern California, eastern Oregon, eastern Washington, Nevada, western Utah, southern and western Arizona, and southeastern Idaho. ..."},</code></pre> <p>A summary of the files we make available:</p> <ul> <li><strong>ref-wiki-en.json.gz</strong>: 2,475,461 external reference webpages (JSON format)</li> <li><strong>ref-wiki-en_urls.txt.gz</strong>: 23,036,318 unique raw links to external references (plain-text format)</li> <li><strong>ref-wiki-en-5k.json.gz</strong>: 37,983 external reference webpages (JSON format)</li> <li><strong>ref-wiki-en-5k_urls.json.gz</strong>: 70,375 unique raw links to external references (plain-text format)</li> <li><strong>ref-wiki-en-5k_Q.txt.gz</strong>: 5,000 Wikidata Q identifiers forming the 5k dataset (plain-text format)</li> </ul> <p>Further details can be found in the publication:</p> <ul> <li><em>Suggesting References for Wikidata Claims based on Wikipedia's External References</em>. Paolo Curotto, Aidan Hogan. Wikidata Workshop @ISWC 2020.</li> </ul> <p>Further material relating to this publication (including code for a proof-of-concept interface) is <a href="http://aidanhogan.com/wikiref/">also available</a>.</p>
mR-PODCAST 8: Mit Julia Meer über Wiki Women
<p>Wie bekommen wir mehr Informationen zu Künstlerinnen und Gestalterinnen in Wikipedia? Durch Workshops mit guter Recherche und aufmerksamer Leitung, sagt Dr. Julia Meer vom <a href="https://www.mkg-hamburg.de/" target="_blank" rel="noreferrer noopener">Hamburger Museum für Kunst und Gewerbe</a> im mR-Podcast. Und Snacks, das sei auch sehr hilfreich. Mit diesem pragmatischen Ansatz und Kolleg:innen hat sie die interaktive Ausstellung <a href="https://www.mkg-hamburg.de/ausstellungen/wiki-women-2" target="_blank" rel="noreferrer noopener">“Wiki Women #2”</a> auf die Beine gestellt.</p>
Egg 2018 Wiki
Easter Egg 2018 with slighly adapted [photo](https://en.wikipedia.org/wiki/Church_of_the_Savior_on_Blood#/media/File:St.Petersburg_Russia_Church_Park-2.jpg) from [Wikipedia article](https://en.wikipedia.org/wiki/Church_of_the_Savior_on_Blood) for Facebook [3D post](https://www.facebook.com/alexander.y.vlasov/posts/1893122770730611). Source: Objaverse 1.0 / Sketchfab
Wiki Head CT Choice Study: Adaptation of US Two Decision Aids to a Québec Local Context
ClinicalTrials.gov study NCT04140084. IPD Sharing: NO. Countries: 1. Publications: 35.
Wiki pathways (mirror https://wikipathways-data.wmcloud.org/20201010/)
Open the record for dataset details and reuse information.
RSS Wiki Example 2 Data Files
<p>Data files associated with RSS Wiki Example 2: https://stephenslab.github.io/rss/example_2.html</p>
RSS Wiki Example 1 Data Files
<p>Data files associated with RSS Wiki Example 1: https://stephenslab.github.io/rss/example_1.html</p>
RSS Wiki Example 5 Data Files
<p>Data files associated with RSS Wiki Example 5: https://stephenslab.github.io/rss/example_5.html</p>
Figure 2 from: Penev L, Hagedorn G, Mietchen D, Georgiev T, Stoev P, Sautter G, Agosti D, Plank A, Balke M, Hendrich L, Erwin T (2011) Interlinking journal and wiki publications through joint citation: Working examples from ZooKeys and Plazi on Species-ID. ZooKeys 90: 1-12. https://doi.org/10.3897/zookeys.90.1369
Figure 2 - The original description of Sinocallipus catba Stoev & Enghoff, 2011 displaying the generic URL of the wiki page (http://species-id.net/wiki/Sinocallipus_catba) right below the ZooBank LSID (see arrow).
Figure 1 from: Penev L, Hagedorn G, Mietchen D, Georgiev T, Stoev P, Sautter G, Agosti D, Plank A, Balke M, Hendrich L, Erwin T (2011) Interlinking journal and wiki publications through joint citation: Working examples from ZooKeys and Plazi on Species-ID. ZooKeys 90: 1-12. https://doi.org/10.3897/zookeys.90.1369
Figure 1 - Citation template for the simultaneous journal and wiki publication of Sinocallipus catba Stoev & Enghoff, 2011 (generic link: http://species-id.net/wiki/Sinocallipus_catba, permanent link of the version depicted in the figure: http://species-id.net/w/index.php?title=Sinocallipus_catba&oldid=XXXX). The generic link always points to the most recent version of the page, while a permanent link is specific to one particular revision.
Figure 4 from: Penev L, Hagedorn G, Mietchen D, Georgiev T, Stoev P, Sautter G, Agosti D, Plank A, Balke M, Hendrich L, Erwin T (2011) Interlinking journal and wiki publications through joint citation: Working examples from ZooKeys and Plazi on Species-ID. ZooKeys 90: 1-12. https://doi.org/10.3897/zookeys.90.1369
Figure 4 - Wiki page of Anochetus boltoni Fisher (http://species-id.net/wiki/Anochetus_boltoni) exported from the Plazi Treatment Repository to Species-ID.
Figure 3 from: Penev L, Hagedorn G, Mietchen D, Georgiev T, Stoev P, Sautter G, Agosti D, Plank A, Balke M, Hendrich L, Erwin T (2011) Interlinking journal and wiki publications through joint citation: Working examples from ZooKeys and Plazi on Species-ID. ZooKeys 90: 1-12. https://doi.org/10.3897/zookeys.90.1369
Figure 3 - Treatment of Anochetus boltoni Fisher extracted through XML markup from the original paper of Fisher and Smith (2008) and deposited at the Plazi Treatment Repository (www.plazi.org).
Figures 5-7 5 from: Hendrich L, Balke M (2011) A simultaneous journal / wiki publication and dissemination of a new species description: Neobidessodes darwiniensis sp. n. from northern Australia (Coleoptera, Dytiscidae, Bidessini). ZooKeys 79: 11-20. https://doi.org/10.3897/zookeys.79.803
Figures 5-7 5 - Figures 5–7. 5 Distribution of Neobidessodes darwiniensis sp. n. in Northern Australia. 6–7 Habitat of Neobidessodes darwiniensis sp. n., Neobidessodes grossus, Neobidessodes mjobergi and Neobidessodes thoracicus, Northern Territory Kakadu Hwy, Harriet Creek at Hwy Crossing (NT 14) (Photos: L. Hendrich).
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.