Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
411
datasets available to search
ShareScore release 0.9.0
Dataset results
411 results for “WIkidata”
Wikidata's linked data for cultural heritage digital resources: an evaluation based on the Europeana Data Model
<p>Wikidata is an open data source with many potential applications. Our study aims to evaluate the usability of Wikidata as a linked data source for acquiring richer descriptions of digital objects within the context of Europeana, a data aggregator from the cultural heritage domain. Specifically, we aim to crawl and convert Wikidata using the standard approaches and operations developed for the (Semantic) Web of Data, i.e. using technologies like linked data consumption and RDF(S)/OWL ontology expression and reasoning. We also seek to re-use existing “semantic” specifications, such as conversions to and from generic data models like Schema.org and SKOS. We have developed an experimental set-up and accompanying software to test the feasibility of this approach. We conclude that Wikidata’s linked data is able to express an interesting level of semantics for cultural heritage, but quality can still be improved and a human operator still must assist linked data applications to interpret Wikidata’s RDF.</p>
Video: Wikidata in de praktijk bij de Koninklijke Bibliotheek, masterclass Wikidata, 28-05-2021
<div><strong>Nederlands: </strong> Videoregistratie van de presentatie over 'Wikidata in de praktijk bij de Koninklijke Bibliotheek' door Olaf Janssen tijdens de masterclass <a title="nl:Wikipedia:GLAM/Wikimediatraining Suriname en het Caribisch gebied" href="https://nl.wikipedia.org/wiki/Wikipedia:GLAM/Wikimediatraining_Suriname_en_het_Caribisch_gebied#Vrijdag_28_mei_2021:_Masterclass_3_-_Meertalige_gestructureerde_data:_Wikidata">Meertalige gestructureerde data: Wikidata</a>, onderdeel van de <a title="nl:Wikipedia:GLAM/Wikimediatraining Suriname en het Caribisch gebied" href="https://nl.wikipedia.org/wiki/Wikipedia:GLAM/Wikimediatraining_Suriname_en_het_Caribisch_gebied">Wikimediatraining Suriname en het Caribisch gebied</a> op 28 mei 2021.</div> <div> </div> <div><strong>English: </strong> Wikimedia training Suriname and the Caribbean - May 28, 2021 - Wikidata at Koninklijke Bibliotheek - recording of session (recorded with participants' permission) <div> <p> </p> <p>Presentatie die in de video gebruikt wordt:</p> <ul> <li><a href="https://doi.org/10.5281/zenodo.7673599" target="_blank" rel="noopener">https://doi.org/10.5281/zenodo.7673599</a></li> <li><a title="File:Wikidata in de praktijk bij de KB door Olaf Janssen - Masterclass Wikidata 28-05-2021.pdf" href="https://commons.wikimedia.org/wiki/File:Wikidata_in_de_praktijk_bij_de_KB_door_Olaf_Janssen_-_Masterclass_Wikidata_28-05-2021.pdf">Wikidata in de praktijk bij de Koninklijke Bibliotheek</a>, door Olaf Janssen, 28 mei 2021</li> </ul> </div> </div>
Video: KB-beelden op Wikimedia Commons verbinden met Wikidata, NDE Hackalod, 9 april 2021
<p><strong>Nederlands: </strong> <a title="Commons:Structured data" href="https://commons.wikimedia.org/wiki/Commons:Structured_data">Structured Data on Commons</a> (SDoC) is een project om afbeeldingen in Wikimedia Commons te voorzien van gestructureerde gegevens uit Wikidata. Hierdoor worden die afbeeldingen beter vindbaar, zichtbaar en herbruikbaar, en worden ze verbonden met de LOD-cloud.</p> <p>Olaf Janssen (KB) legt in deze video uit wat SDoC is, <a title="Commons:Structured data/GLAM/Why" href="https://commons.wikimedia.org/wiki/Commons:Structured_data/GLAM/Why">waarom het een tof project is</a>, wat het oplevert en hoe je er als GLAM zelf mee aan de slag kunt gaan. Hij gebruikt hierbij de <a title="Category:Media contributed by Koninklijke Bibliotheek" href="https://commons.wikimedia.org/wiki/Category:Media_contributed_by_Koninklijke_Bibliotheek">Commons-beelden van de KB</a> als voorbeeld.</p> <p>Zie ook</p> <ul> <li><a href="https://web.archive.org/web/20210402070033/https://hackalod.com/index.php/2021/03/26/nieuwe-hackalod-online-9-april-thema-wikidata/" rel="nofollow">https://web.archive.org/web/20210402070033/https://hackalod.com/index.php/2021/03/26/nieuwe-hackalod-online-9-april-thema-wikidata/</a></li> <li><a href="https://commons.wikimedia.org/wiki/Commons:Structured_data">https://commons.wikimedia.org/wiki/Commons:Structured_data</a></li> <li><a href="https://commons.wikimedia.org/wiki/Commons:Structured_data/About/Why">https://commons.wikimedia.org/wiki/Commons:Structured_data/About/Why</a></li> <li><a href="https://commons.wikimedia.org/wiki/Commons:Structured_data/GLAM">https://commons.wikimedia.org/wiki/Commons:Structured_data/GLAM</a></li> <li><a href="https://commons.wikimedia.org/wiki/Commons:Depicts">https://commons.wikimedia.org/wiki/Commons:Depicts</a></li> </ul> <div> <p>Presentatie die in deze video getoond wordt:</p> <ul> <li><a title="File:KB-beelden op Wikimedia Commons verbinden met Wikidata - NDE Hackalod 9 april 2021 Olaf Janssen.pdf" href="https://commons.wikimedia.org/wiki/File:KB-beelden_op_Wikimedia_Commons_verbinden_met_Wikidata_-_NDE_Hackalod_9_april_2021_Olaf_Janssen.pdf">KB-beelden op Wikimedia Commons verbinden met Wikidata - NDE Hackalod, 9 april 2021</a></li> <li><a href="../records/7673765" target="_blank" rel="noopener">https://zenodo.org/records/7673765</a></li> </ul> </div> <p>Een langere versie van deze video is ook beschikbaar op <a href="https://www.youtube.com/watch?v=vQ6ZTYEiwTU" rel="nofollow">Youtube</a>.</p>
List of EuDML items in Wikidata (nanocontribution)
<p>CSV table of the results from the query</p> <p><code>SELECT ?item ?value</code><br><code>{</code><br><code> ?item wdt:P11166 ?value .</code><br><code>}</code></p> <p>Executed on https://query.wikidata.org/ at 2024-10-06 10:24</p>
Wikidata and DBpedia Space Travel Data Comparison with ABECTO
<p>This is an <a href="https://github.com/fusion-jena/abecto">ABECTO</a> execution plan to compare space travel data from <a href="https://www.wikidata.org">Wikidata</a> and <a href="https://www.dbpedia.org">DBpedia</a> and the according results.</p> <p>The generated result data are derived from the compared knowledge graphs, which are licensed as follows:</p> <ul> <li><a href="https://www.dbpedia.org">DBpedia</a> by DBpedia Association (<a href="http://en.wikipedia.org/wiki/Wikipedia:Text_of_Creative_Commons_Attribution-ShareAlike_3.0_Unported_License">CC BY-SA 3.0</a>)</li> <li><a href="https://wikidata.org">Wikidata</a> (<a href="https://creativecommons.org/publicdomain/zero/1.0/">CC0 1.0</a>)</li> </ul>
Datasets for Creating and Querying Personalized Versions of Wikidata in a Laptop (Wikidata workshop, 2021)
<p>These datasets are used to support the results of the paper "Datasets for Creating and Querying Personalized Versions of Wikidata in a Laptop", submitted to the Wikidata workshop 2021 (https://wikidataworkshop.github.io/2021/) at the International Semantic Web Conference.</p> <p>The datasets have been derived from Wikidata dump 20210215. To help querying purposes, the dump is organized in different files:</p> <ul> <li>claims.time.tsv.gz: time-related assertions</li> <li>claims.wikibase-item.tsv-006.gz: item-related assertions</li> <li>derived.P279.tsv.gz: statements that are subclass of another statement</li> <li>derived.P279star.tsv.gz: statement that are subclass of another statement, including their chains.</li> <li>derived.P31.tsv.gz: instance of statements.</li> <li>labels.en.tsv-004.gz: labels in English</li> <li>claims.external-id.tsv-005.gz: External identifiers for each item.</li> <li>ulan.tsv: ULAN ids (used to link external identifiers to Wikidata identifiers)</li> <li>wikidata_infobox.tsv.gz: Information about dbpedia infoboxes.</li> </ul> <p>Upload by: Daniel Garijo</p>
Cache for KGTK queries in Creating and Querying Personalized Versions of Wikidata on a Laptop
<p>Sqlite cache used to store the KGTK Kypher (https://kgtk.readthedocs.io/en/dev/transform/query/) queries for paper "Creating and Querying Personalized Versions of Wikidata on a Laptop"</p>
Uncertainty and debate in statements describing Wikidata Works of art
<p>This dataset comprises a selection of statements from <a href="https://doi.org/10.5281/zenodo.7307852">all artworks in Wikidata</a>. </p> <p>In particular, </p> <ul> <li><strong>natures.json</strong> stores statements with a “Nature of statements” qualifier. Statements, independently of rank, can be decorated with an additional triple using predicate P5102. 54 terms among 283 available may mark the statement as uncertain or debated (e.g. debated, hypothesis, possibly). For example, the painting “Abstract Speed + Sound” (Q19882431) by Giacomo Balla is deemed to be possibly part of a triptych. </li> <li><strong>non-asserted.json </strong>contains those statements which are not asserted. Competing statements are represented via a ranking mechanism (e.g., Preferred, Normal and Deprecated). Individual statements are not actually asserted, but an extra triple is added those that are deemed true. For example, the painting “Madonna with the Blue Diadem” (Q738038) has been attributed to Raphael (non asserted statement, ranked as normal) and Gianfrancesco Penni (asserted statement, ranked as preferred and additionally asserted). </li> <li><strong>null-valued.json</strong> contains all statements with a null-valued objects. A statement can be associated with a blank node. This is meant to imply that the statement is associated with an unknown value, rather than a missing statement. For example, “Missal for the use of the ecclesiastics of Clermont' (Q113302686), an illuminated manuscript from the 14th century, has been recorded with both an unknown creator and author.</li> <li><strong>sourcing-circumstances.json </strong>stores statements with a “Sourcing circumstance” qualifier. As for Natures of statements only statements with an uncertain qualifier have been selected.</li> </ul>
Wikidata 3 Topical Subsets (Gene Wiki, Music, Ships) and 4 Random Subsets
<p>This dataset contains the N-Triples files of 3 Wikidata topical subsets corresponding to 3 Wikidata WikiProject: Gene Wiki, Music, and Ships along with 4 random subsets in different sizes: two of 100K items, one 500K items, and one 1M items. Subsets are extracted from the <a href="https://academictorrents.com/details/229cfeb2331ad43d4706efd435f6d78f40a3c438">3 January 2022 dump</a>. All subsets have been extracted with <a href="https://github.com/seyedahbr/wdumper">WDumper</a> using these <a href="https://github.com/seyedahbr/RQSS_Evaluation/tree/main/WDumper%20Specification%20Files">JSON specification files</a>. The files are:</p> <ul> <li>GeneWiki.zip: contains 25 `.nt.gz` RDF files each of which corresponds to one of the main Gene Wiki WikiProject classes, e.g. protein, gene, chemical compound, etc.</li> <li>music.nt.gz: the RDF file corresponding to the Music WikiProject.</li> <li>ships.nt.gz: the RDF file corresponding to the Ships WikiProject.</li> <li>Random100K_1.zip: contains 2 `nt.gz` RDF files each of which includes (about) 50,000 random Wikidata items, 100,000 items in total.</li> <li>Random100K_2.zip: contains 2 `nt.gz` RDF files each of which includes (about) 50,000 random Wikidata items, 100,000 items in total.</li> <li>Random500K.zip: contains 10 `nt.gz` RDF files each of which includes (about) 50,000 random Wikidata items, 500,000 items in total.</li> <li>Random1M.zip: contains 20 `nt.gz` RDF files each of which includes (about) 50,000 random Wikidata items, 1,000,000 items in total.</li> </ul> <p> </p>
Wikidata subset from 2018 dumps created with WDSub at Biohackathon 2022 - Turtle format
<p>Subset of Wikidata obtained using this Shape Expression: https://github.com/kg-subsetting/datasets-biohackathon2022/blob/main/GeneWiki/GeneWiki.shex</p> <p>And the wdsub tool version 0.0.28: https://github.com/weso/wdsub</p> <p>The input dump is: wikidata-20180115-all</p> <p>And the dumpformat is TURTLE</p>
The LOTUS Initiative for Open Natural Products Research: wikidata query results
<p>Wikidata query results returned by the downloadLotus module of the <a href="https://github.com/lotusnprod/lotus-wikidata-interact">https://github.com/lotusnprod/lotus-wikidata-interact</a> program.</p> <p>See details of the module here <a href="https://github.com/lotusnprod/lotus-wikidata-interact/blob/main/downloadLotus/README.md">https://github.com/lotusnprod/lotus-wikidata-interact/blob/main/downloadLotus/README.md</a></p> <p>This dataset is constituted of 4 tables.</p> <ol> <li>compounds.tsv - chemical structures metadata (wikidataId, canonicalSmiles, isomericSmiles, inchi, inchiKey)</li> <li>references.tsv - bibliographical references metadata (wikidataId, pipe separated DOIs, titles)</li> <li>taxa.tsv - biological organisms metadata (wikidataId, pipe separated names, taxa rank)</li> <li>compound_reference_taxon.tsv - the documented structure-organism pairs</li> </ol> <p>This dataset includes not only the outputs of the LOTUS processing pipeline (available here <a href="https://doi.org/10.5281/zenodo.5665295">https://doi.org/10.5281/zenodo.5665295</a> ) but also any of wikidata chemical compounds having the found in taxon property (<a href="https://www.wikidata.org/wiki/Property:P703">https://www.wikidata.org/wiki/Property:P703</a>) and their associated organisms and documenting references.</p> <p> </p> <p><br> </p>
Automatically Extracted SHACL Shapes for WikiData, DBpedia, YAGO-4, and LUBM & Associated Coverage Statistics
<p>The uploaded datasets contain <strong>automatically extracted </strong>SHACL shapes for the following datasets:</p> <ul> <li>WikiData (the truthy dump from September 2021 filtered by removing non-English strings) [1]</li> <li>DBpedia [2]</li> <li>YAGO-4 [3] </li> <li>LUBM (scale factor 500) [4]</li> </ul> <p>The validating shapes for these datasets are generated by a program that parses the corresponding RDF files (in `.nt` format). The extracted shapes encode various SHACL constraints, e.g., sh:minCount, sh:path, sh:class, sh:datatype etc. For each shape we encode coverage in terms of number of entities satisfying such shape, this information is encoded using the <a href="http://vocab.deri.ie/void#entities">void:entities</a> predicate. </p> <p>We have provided as executable Jar file the program we developed to extract these SHACL shapes.<br> More details about the datasets used to extract these shapes and <em>how to run the Jar</em> are available on our GitHub repository <a href="https://github.com/dkw-aau/qse">https://github.com/dkw-aau/qse</a>.</p> <p>Read more about our Quality Shapes Extraction (QSE) tool on our website <a href="https://relweb.cs.aau.dk/qse/">https://relweb.cs.aau.dk/qse/</a></p> <p>[1] Vrandečić, Denny, and Markus Krötzsch. "Wikidata: a free collaborative knowledgebase." Communications of the ACM 57.10 (2014): 78-85.</p> <p>[2] Auer, Sören, et al. "Dbpedia: A nucleus for a web of open data." The semantic web. Springer, Berlin, Heidelberg, 2007. 722-735.</p> <p>[3] Pellissier Tanon, Thomas, Gerhard Weikum, and Fabian Suchanek. "Yago 4: A reason-able knowledge base." European Semantic Web Conference. Springer, Cham, 2020.</p> <p>[4] Guo, Yuanbo, Zhengxiang Pan, and Jeff Heflin. "LUBM: A benchmark for OWL knowledge base systems." Journal of Web Semantics 3.2-3 (2005): 158-182.</p>
Simple dataset obtained as a Wikidata subset from 2022 dump using entity schema about taxon
<p>The subset has been obtained using wdsub version 0.0.33 and the schema:</p> <p> </p> <pre>PREFIX p: <http://www.wikidata.org/prop/> PREFIX ps: <http://www.wikidata.org/prop/statement/> PREFIX prov: <http://www.w3.org/ns/prov#> PREFIX wd: <http://www.wikidata.org/entity/> PREFIX wdt: <http://www.wikidata.org/prop/direct/> start = @<taxon_by_wd_ontology> OR @<taxon_by_identifier> <taxon_by_wd_ontology> { wdt:P31 [wd:Q16521] ; } <taxon_by_identifier> {wdt:P685 . +;} OR # NCBI taxonomy ID {wdt:P846 . +;} OR # GBIF taxon ID {wdt:P3151 . +;} OR # iNaturalist taxon ID {wdt:P3444 . +;} # eBird taxon ID</pre>
Datalog subsetting input files (Wikidata 2015 NTriple-to-CSV dump)
<p>This is a Wikidata 2015 NTriple dump in which the delimiter is changed to ','. The file is used in subsetting experiment via <a href="https://github.com/seyedahbr/radlog">Radlog</a>.</p>
Wikidata Thematic Subgraph Selection
<p><strong>Wikidata Thematic Subgraph Selection</strong></p> <p>These datasets have been designed to train and evaluate algorithms to select thematic subgraphs of interest in a large knowledge graph from seed entities of interest. Specifically, we consider Wikidata. Given a set of seed QIDs of interest, a graph expansion is performed following P31, P279, and (-)P279 edges. Traversed classes that thematically deviates from seed QIDs of interest should be pruned. Datasets thus consist of classes reached from seed QIDs that are labeled as "to prune" or "to keep".</p> <p><strong>Available datasets</strong></p> <table> <tbody><tr> <th>Dataset</th> <th># Seed QIDs</th> <th># Labeled decisions</th> <th># Prune decisions</th> <th>Min prune depth</th> <th>Max prune depth</th> <th># Keep decisions</th> <th>Min keep depth</th> <th>Max keep depth</th> <th># Reached nodes up</th> <th># Reached nodes down</th> </tr> </tbody><tbody> <tr> <td><a href="data/dataset1">dataset1</a></td> <td>455</td> <td>5233</td> <td>3464</td> <td>1</td> <td>4</td> <td>1769</td> <td>1</td> <td>4</td> <td>1507</td> <td>2593609</td> </tr> <tr> <td><a href="data/dataset2">dataset2</a></td> <td>105</td> <td>982</td> <td>388</td> <td>1</td> <td>2</td> <td>594</td> <td>1</td> <td>3</td> <td>1159</td> <td>1247385</td> </tr> </tbody> </table> <p>Each dataset folder contains</p> <ul> <li><code>datasetX.csv</code>: a CSV file containing one seed QID per line (not the complete URL, just the QID). This CSV file has no header.</li> <li><code>datasetX_labels.csv</code>: a CSV file containing one seed QID per line and its label (not the complete URL, just the QID)</li> <li><code>datasetX_gold_decisions.csv</code>: a CSV file with seed QIDs, reached QIDs, and the labeled decision (1: keep, 0: prune)</li> <li><code>datasetX_Y_folds.pkl</code>: folds to train and test models based on the labeled decisions</li> </ul> <p><code>dataset1-2</code> consists of using <code>dataset1</code> for training and <code>dataset2</code> for testing.</p> <p><strong>License</strong></p> <p>Datasets are available under the <a href="https://creativecommons.org/licenses/by-nc/4.0/">CC BY-NC</a> license.</p>
Wikidata Dump viaf humans
<p> RDF dump of wikidata produced with <a href="https://tools.wmflabs.org/wdumps/">wdumps</a>. </p> <p> <br> <a href="https://tools.wmflabs.org/wdumps/dump/38">View on wdumper</a> </p> <p> <b>entity count</b>: 0, <b>statement count</b>: 0, <b>triple count</b>: 0 </p>
Wikidata Dump viaf humans nkcr
<p> RDF dump of wikidata produced with <a href="https://tools.wmflabs.org/wdumps/">wdumps</a>. </p> <p> <br> <a href="https://tools.wmflabs.org/wdumps/dump/39">View on wdumper</a> </p> <p> <b>entity count</b>: 0, <b>statement count</b>: 0, <b>triple count</b>: 0 </p>
Wikidata Dump Indian_Politician_Properties
<p> RDF dump of wikidata produced with <a href="https://tools.wmflabs.org/wdumps/">wdumps</a>. </p> <p> <br> <a href="https://tools.wmflabs.org/wdumps/dump/94">View on wdumper</a> </p> <p> <b>entity count</b>: 0, <b>statement count</b>: 0, <b>triple count</b>: 0 </p>
Wikidata Dump Indian_Politician_Properties
<p> RDF dump of wikidata produced with <a href="https://tools.wmflabs.org/wdumps/">wdumps</a>. </p> <p> <br> <a href="https://tools.wmflabs.org/wdumps/dump/94">View on wdumper</a> </p> <p> <b>entity count</b>: 0, <b>statement count</b>: 0, <b>triple count</b>: 0 </p>
Wikidata Dump KitchenOBJs
<p> RDF dump of wikidata produced with <a href="https://tools.wmflabs.org/wdumps/">wdumps</a>. </p> <p> Kitchenware subclasses<br> <a href="https://tools.wmflabs.org/wdumps/dump/120">View on wdumper</a> </p> <p> <b>entity count</b>: 0, <b>statement count</b>: 0, <b>triple count</b>: 0 </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.