Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
411
datasets available to search
ShareScore release 0.7.1
Dataset results
411 results for “Wikidata”
Generated Wikidata Subset for Taxons based on dump: 20170821-all
<p>Source file: GeneTaxon_wikidata-20170821-all.ttl.gz</p> <p>ShEx: https://github.com/kg-subsetting/paper-wikidata-subsetting-2023/blob/master/flexibility-experiments/genes%2Btaxons/GeneTaxon.shex</p> <p>More information: https://www.semantic-web-journal.net/content/wikidata-subsetting-approaches-tools-and-evaluation</p>
Generated Wikidata Subset for Taxons based on dump: wikidata-20150601-all
<p>Source file: GeneTaxon_wikidata-20150601-all.ttl.gz</p> <p>ShEx: https://github.com/kg-subsetting/paper-wikidata-subsetting-2023/blob/master/flexibility-experiments/genes%2Btaxons/GeneTaxon.shex</p> <p>More information: https://www.semantic-web-journal.net/content/wikidata-subsetting-approaches-tools-and-evaluation</p>
Generated Wikidata Subset for Taxons based on dump: 20160613-all
<p>Source file: GeneTaxon_wikidata-20160613-all.ttl.gz</p> <p>ShEx: https://github.com/kg-subsetting/paper-wikidata-subsetting-2023/blob/master/flexibility-experiments/genes%2Btaxons/GeneTaxon.shex</p> <p>More information: https://www.semantic-web-journal.net/content/wikidata-subsetting-approaches-tools-and-evaluation</p>
Generated Wikidata Subset for Taxons based on dump: 20220630-all
<p>Source file: GeneTaxon_wikidata-20220630-all.ttl.gz</p> <p>ShEx: https://github.com/kg-subsetting/paper-wikidata-subsetting-2023/blob/master/flexibility-experiments/genes%2Btaxons/GeneTaxon.shex</p> <p>More information: https://www.semantic-web-journal.net/content/wikidata-subsetting-approaches-tools-and-evaluation</p>
Generated Wikidata Subset for Taxons based on dump: 20210531-all
<p>Source file: GeneTaxon_wikidata-20210531-all.ttl.gz</p> <p>ShEx: https://github.com/kg-subsetting/paper-wikidata-subsetting-2023/blob/master/flexibility-experiments/genes%2Btaxons/GeneTaxon.shex</p> <p>More information: https://www.semantic-web-journal.net/content/wikidata-subsetting-approaches-tools-and-evaluation</p>
Wikidata Subsetting: Reference-based Subsetting Experiment Datasets
<p>Files in this dataset have been produced during Flexibility experiments of Wikidata subsetting practical tools: Subsetting based on references using WDSub.</p>
Wikidata Subsetting: Performance and Accuracy Experiment Datasets
<p>Files in this dataset have been produced during Performance and Accuracy experiments of Wikidata subsetting practical tools: WDumper, KGTK, WDSub, WDF.</p>
Wikidata-CS
<p>Wikidata-CS (Commonsense) datasets associated with the submission entitled 'Commonsense Knowledge in Wikidata', under review at the Wikidata workshop, ISWC 2020</p>
wikidata jsons (in pre-processed format)
<p>This is a collection of pre-processed wikidata jsons which were used in the creation of CSQA dataset (Ref: <a href="https://arxiv.org/abs/1801.10314">https://arxiv.org/abs/1801.10314</a>).</p> <p>Please refer to <a href="https://amritasaha1812.github.io/CSQA/download/">https://amritasaha1812.github.io/CSQA/download/</a> for more details.</p>
Wikidata Subsets and Specification Files Created by WDumper
<p>WDumper is a third-party tool enables users to craete custom dump of Wikidata. Here we create some topical subset of Wikidata by WDumper. There are 4 usecases:</p> <ol> <li>politicians: People with occupation of politicians in Wikidata</li> <li>militPoliticians: People with occupation of politicians which are military and also have a military rank of General in Wikidata</li> <li>ukUniversities: All United Kingdom universities in Wikidata</li> <li>geneWiki: A subset from <a href="https://elifesciences.org/articles/52614/figures#fig1">this class-diagram</a></li> </ol> <p>The subsets are in .nt.gz format. For each usecase there are two types of subsets in terms of content, one with References and Qualifiers that has a "withRQFS" in the name, and one without this feature. Also for each use case, there are two extracted subsets one from 27 April 2015 and another from 13 November 2020. The corresponding JSON file for each use case is the WDumper specification file.</p>
SSSOM-like mapping of PanglaoDB to Wikidata
<p>Mappings from PanglaoDB to Wikidata. <br>See details at https://jvfe.github.io/paper_wdt_panglao/. </p>
Scholarly Wikidata: Population and Exploration of Conference Data in Wikidata using LLMs
<p>This dataset provides the input data and intermediate results of the paper titled "Scholarly Wikidata: Population and Exploration of Conference Data in Wikidata using Large Language Models and Semantic Web Techniques". It contains the following resources.</p> <ul> <li>conference proceedings front matter links - these links can be used to download the pdf files of the conference proceeding front matters that include information about the number of submitted and accepted papers that can be used to calculate acceptance rates, names of all conference organization committee members, list of programme committee and senior programme member names for each track with other interesting facts such as the main topics of the submitted papers and emerging topics according to the editors, etc.</li> <li>web crawl of conference websites - this contains a set of crawled content from each conference website in both HTML and text formats. Each file contains web pages from a specific conference along with the page URL, page title, and page content. Information such as important dates (deadlines) and other announcements can be extracted from the content of the web sites. </li> <li>papers and paper-authors list for each conference in a given conference series - this contains the paper list along with their corresponding authors for each conference series extracted from DBLP. </li> <li>OpenRefine projects - this contains examples of open refile projects that were used to perform entity linking and reconciliation as well as the schemas that was used to map the tabular data columns to Wikidata properties, and qualifiers and cell values to Wikidata entities.</li> <li>evaluation benchmark - this contains the outputs of LLM generations for the tasks (a) extracting the number of submitted and accepted papers per each track at a given conference, (b) extraction of organizers with their roles for each conference, (c) extraction of programme committee members with track and their role (member, SPC member), and (d) extraction of important dates or deadlines for each activity (submission, notification, etc.) in each track. </li> </ul> <p>The corresponding source code is available at the <a href="https://github.com/scholarly-wikidata/scholarly-wikidata/">scholary-data repo</a>.</p>
UMLS-Wikidata
<p><em>UMLS_Wikidata</em> is a German biomedical entity linking knowledge base that provides good coverage for German entity linking datasets such as <a href="../records/8188966">WikiMed-DE-BEL</a>. The knowledge base is created by filtering out the Wikidata items that contain the Concept Unique Itentifier (CUI) of UMLS. Each entry in the knowledge base consists of Wikidata QID, label, description, UMLS CUI and aliases. The resulting KB has 731,414 Wikidata QIDs, 599,330 unique CUIs and 671,797 unique (mention, CUI) pairs where mention includes <em>label</em> and <em>aliases</em>. </p>
Topic Model for English Wikipedia's Biographies with list of all 1.8M articles linked to Wikidata
<p>A Genism LDA Topic Model of English Wikipedia biographical articles with list of all 1.8M articles, and some associated Wikidata information</p> <p>The model has 150 Topics.</p> <p>This model was developed in the process of isolating a set of visual arts biographical articles, as described in "Clowns in the Visual Artists: Topic Modeling Wikipedia and Wikidata" in the Spring 2022 issue of <em>Art Documentation - </em><a href="https://doi.org/10.1086/719999">https://doi.org/10.1086/719999</a></p> <p>Because names, nationalities, and birthdays are so prominent in biographies, the stopwords list removed 170,000 names, surnames, city names, place names, countries, days, months and other time related words (https://github.com/mandiberg/Names-Surnames-and-Countries-for-Stopwords). We also directly removed each article subject’s given and surname, which were almost always the most frequently occurring words in any given article. Otherwise, the model just produced topics based on nationality, and common names and surnames.</p> <p><strong>Files:</strong></p> <p>all_enwiki_bios_from_wikidata.csv<br> The list of all Wikidata items for humans with an enwiki page (e.g biographical article) was extracted from Wikidata JSON dump; list includes gender, occupation, and nationality. This was joined with the converted plaintext from an English Wikipedia dump. This data was downloaded in March 2021.</p> <p>Wikipedia Biographies LDA Topic Model human readable summary.csv<br> A human readable file with the 150 topics ranked by count of articles per topic from the 1.8M corpus. The most popular topics have categorical descriptions of the occupations of each cluster. Some are marked as not an occupation cluster. </p> <p>BoW_corpus.mm*<br> model_lda_full_Sep2_150Tv2*<br> These six files comprise the topic model. The code to load them is present in the python files. </p> <p>dict_full_Aug-28-2021<br> processed_docs_full_Aug-28-2021.txt<br> processed_docs_1000_Aug-18-2021.txt<br> These are the dictionary and processed corpuses required to build and implement the model using this code. The corpus with the first 1000 items is meant to be used for testing, as the full one is quite large and takes a long time to complete. </p> <p>topic-model-wikipedia-sept2021.zip<br> The code and settings used for creating and implementing this model are included in this zip and are also available here: https://github.com/mandiberg/topic-model-wikipedia</p> <p>All-Wikipedia-Biographies-with-topic1.csv<br> All-Wikipedia-Biographies-with-topic1and2.csv<br> These are the list of 1.8M biographies matched to topics. The "topic1" file just includes the first topic, this is a slightly larger list. The "topic1and2" file is slightly smaller because about 2% articles do not match to a second topic.</p> <p>Analysis-for-Clowns-Visual-Arts.zip<br> These are the raw data and final data produced for the "Clowns in the Visual Artists." Please see the article for context.</p>
Wikidata subset with revision history information [RDF]
<p>This dataset is composed of 300 instances from the 100 most important classes in Wikidata, for a total of around 30000 entities and 390000 triples. The dataset is geared towards knowledge graph refinement models that leverage edit history information from the graph. There are two versions of the dataset:</p> <ul> <li>The <strong>static</strong> version (files postfixed with '_static') contains the simple statements of each entity fetched from Wikidata.</li> <li>The <strong>dynamic</strong> version (files postfixed with '_dynamic') contains information about the operations and revisions made to these entities, and the triples that were added or removed.</li> </ul> <p>Each version is split into three subsets: train, validation (val), and test. Each split contains every entity from the dataset. The train split contains the first 70% of revisions made to each entity, the validation split contains the 70% to 85% revisions, and the test set contains the last 15% revisions.</p> <p>This is a sample from the static datasets:</p> <pre><code>wd:Q217432 a uo:entity ; wdt:P1082 1.005904e+06 ; wdt:P1296 "0052280" ; wdt:P1791 wd:Q18704103 ; wdt:P18 "Pitakwa.jpg" ; wdt:P244 "n80066826" ; wdt:P571 "+1912-00-00T00:00:00Z" ; wdt:P6766 "421180027" .</code></pre> <p>Each entity has the type <em>uo:entity</em>, and contains the statements added during that time period following Wikidata's data model.</p> <p>In the following code snippet we show an example from the dynamic dataset:</p> <pre><code>uo:rev703872813 a uo:revision ; uo:timestamp "2018-06-28T22:31:32Z" . uo:op703872813_0 a uo:operation ; uo:fromRevision uo:rev703872813 ; uo:newObject wd:Q82955 ; uo:opType uo:add ; uo:revProp wdt:P106 ; uo:revSubject wd:Q6097419 . uo:op703878666_0 a uo:operation ; uo:fromRevision uo:rev703878666 ; uo:opType uo:remove ; uo:prevObject wd:Q1108445 ; uo:revProp wdt:P460 ; uo:revSubject wd:Q1147883 .</code></pre> <p>This dataset is composed of revisions, which have a timestamp. Each revision is composed of 1 to n operations, in which there is a change to a statement from the entity. There are two types of operations: <em>uo:add</em> and <em>uo:remove</em>. In both cases, the property and the subject being modified are shown with the <em>uo:revProp</em> and <em>uo:revSubject</em> properties. In the case of additions, <em>uo:newObject</em> and <em>uo:prevObject</em> properties are added to show the previous and new objects after the addition. In the case of removals, there is a <em>uo:prevObject </em>property to record the object that was removed.</p>
Wikidata ALPINE-based WINE & SPINE embedding
<p>100-dimensional Wikidata graph embedding obtained using degree-based WINE and SPINE.</p>
Wikidata enriched Slovenian geographical names
<p>A subset of data from the official Register of Slovenian Geographical Names (<a href="https://eprostor.gov.si/imps/srv/eng/catalog.search#/metadata/0f5fd804-9073-42d9-ad3c-b273f59fc16c" target="_blank" rel="noopener">REZI</a>). Names of settlements, municipalities, statistical regions and cohesion regions are enriched with entity codes from Wikidata, Geonames, Open Street Map and DBpedia. A complete list of matching codes exists only for the Wikidata source.</p> <p>The dataset was created in 2021, at which time some entities in Wikidata were also updated and aligned with the official names.</p> <p>File formats:</p> <ul> <li>db; SQLite file</li> <li>xlsx; Excel file</li> <li>zip; zipped csv tables with TAB separator </li> </ul>
Wikiproyecto: Visibilizando el dominio público colombiano en Wikidata
<p>Listado de personas en dominio público e identificadores (Q) de Wikidata.</p>
Wikidata dump 2017-12-27
<p>Wikidata dump retrieved from <a href="https://www.google.com/url?q=https://dumps.wikimedia.org/wikidatawiki/entities/latest-all.json.bz2&sa=D&ust=1522795489326000&usg=AFQjCNFKS4q0PRJ3VfsGUVQjno2fN9otlQ">https://dumps.wikimedia.org/wikidatawiki/entities/latest-all.json.bz2</a> on 27 Dec 2017</p>
Wikidata Vandalism Corpus 2015 (WDVC-15)
<p>The Wikidata vandalism corpus 2015 (WDVC-15) is a corpus for the evaluation of automatic vandalism detectors for Wikidata. For research purposes the corpus can be used free of charge.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.