Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
219
datasets available to search
ShareScore release 0.7.1
Dataset results
219 results for “Knowledge Graph”
Wikidata5m - knowledge graph (transductive)
<p>Wikidata5m is a million-scale knowledge graph dataset with aligned corpus.This dataset integrates the <a href="https://www.wikidata.org/">Wikidata</a> knowledge graph and <a href="https://www.wikipedia.org/">Wikipedia</a> pages. Each entity in Wikidata5m is described by a corresponding Wikipedia page, which enables the evaluation of link prediction over unseen entities.</p> <p>This file contains the transductive split of Wikidata5m knowledge graph.</p>
DH ATLAS: Knowledge Graph v2.0
<p>The <a href="https://w3id.org/dh-atlas/">ATLAS Ontology</a> has been implemented to describe the metadata of selected pilot projects and their related entities. This effort has resulted in a Knowledge Graph, currently available as a <a href="https://github.com/dh-atlas/knowledge-graph/blob/main/releases/v2.0/knowledge-graph-2.0.zip">compressed archive</a> containing Turtle (.ttl) serialization.</p> <h3>Version 2.0</h3> <p>The current version of the Knowledge Graph revises the Turtle serializations according to the new properties of the <a href="https://w3id.org/dh-atlas/2.0">ATLAS Ontology v2.0</a>.</p> <p>Described resources include 54 Research Products among:</p> <ul> <li>Digital Scholarly Editions</li> <li>Text Collections</li> <li>Software</li> <li>Ontologies</li> <li>Linked Open Data</li> <li>New classes: Language Model and 3D Digital Twin</li> </ul> <p>These are linked to their contextual entities: Research Projects, Organizations, People, Websites, and Computer Programs.</p> <p>In addition, this version introduces a set of extraction graphs containing subjects corresponding to some of the described research products. These graphs were produced through semi-automatic knowledge extraction directly from the available sources, including APIs, SPARQL endpoints, static files, websites, and textual descriptions.</p>
[Re] Object Detection Meets Knowledge Graphs - Supporting Datasets
<p><strong>Supporting Datasets for [Re] Object Detection Meets Knowledge Graphs</strong></p> <p>The supporting data needed to reproduce the results of <a href="https://www.ijcai.org/Proceedings/2017/230">Object Detection Meets Knowledge Graphs</a> as a part of the submission to ReScience C journal.</p>
The Nova Scotia Disease Knowledge Graph
<p>Nova Scotia's government has an abundance of resources in terms of data and information. All this data has been collected and stored on the NSOD portal (<a href="https://data.novascotia.ca/">https://data.novascotia.ca</a>) in the form of datasets. The Nova Scotia Open Data Portal was built and managed by Socrata API (<a href="https://dev.socrata.com/">https://dev.socrata.com</a>). We transformed the disease-related datasets of Nova Scotia Open Data into RDF, enriched them by disease ontology, and it is available to be used under the MIT licence.</p> <p>- Added DBpedia linking</p>
Stratigraphic Knowledge Graph (StraKG)
<p>The background of this work is the increasing amount of geoscience literature data shared online, including those from governmental agencies, research institutions, and the crowdsourcing encyclopedia. The knowledge graph is an effective way to explore and analyze open-text data. In this work, we designed and constructed a knowledge graph for the field of stratigraphy, called StraKG, to help process records in the Baidu Encyclopedia, a big open-text data resource in Chinese. The files shared in this repository are code for building the layered structure of the StraKG and extracting instance records and relationships from the open text.</p> <p>StraKG has a two-layer structure, representing ontologies and instances, respectively. At the top is the schema layer for classes and properties of the ontologies. Community-level geological dictionaries were used as a foundation to build ontologies. At the bottom is the instance layer. Text mining techniques were used to analyze open text from Baidu Encyclopedia, extracting instances of strata, rocks, and locations, as well as the relationships between those entities. In our work, we also established mapping between the schema layer and instance layer and implemented a list of experiments to test the utility of the resulting StraKG.</p>
Dynamic Knowledge Graphs for Continual Learning of Embeddings
<p>These datasets are generated from real world usecases. They are treated as Knowledge graphs and include 20 snapshots, where between two snapshots there are 10% added links and 10% deleted links, making the first and last snapshot non-overlapping.</p>
INGRIDKG: A FAIR Knowledge Graph of Graffiti
<p>Graffiti is an urban phenomenon that is increasingly attracting the interest of the sciences. To the best of our knowledge, no<br> suitable data corpora are available for systematic research until now. The Information System Graffiti in Germany project<br> (INGRID) closes this gap by dealing with graffiti image collections that have been made available to the project for public<br> use. Within INGRID, the graffiti images are collected, digitized and annotated. With this work, we aim to support the rapid<br> access to a comprehensive data source on INGRID targeted especially by researchers. In particular, we present INGRIDKG, an<br> RDF knowledge graph of annotated graffiti, abides by the Linked Data and FAIR principles. We weekly update INGRIDKG<br> by augmenting the new annotated graffiti to our knowledge graph. Our generation pipeline applies RDF data conversion,<br> link discovery and data fusion approaches to the original data. The current version of INGRIDKG contains 460,640,154 triples<br> and is linked to 3 other knowledge graphs by over 200,000 links. In our use case studies, we demonstrate the usefulness of<br> our knowledge graph for different applications. INGRIDKG is publicly available under the Creative Commons Attribution 4.0<br> International license.</p>
Semantic Web resources and Machine Learning systems - Knowledge Graph (SWeMLS-KG)
<p>This resource is part of our submission to ESWC 2023 resource track, which includes:</p> <p>Datasets:<br> - Folder "pattern" - a set of SWeMLS patterns represented based on OPMW and P-Plan ontology,<br> - Folder "shapes" - a set of SHACL constraints to check the conformance of SWeML Systems against SWeMLS patterns as well as a set of SHACL-AF rules to generate links between system components,<br> - File "swemls-ontology.ttl" - an ontology to represent Semantic Web resources and Machine Learning systems (SWeMLS),<br> - File "swemls-instances.ttl" - a set of triples representing the extracted metadata from 476 SWeML systems and papers,<br> - File "swemls-kg.ttl" - an integrated and validated KG containing all above files, including enrichment from SHACL-AF rules using "swemls-toolkit" [2].</p> <p>These resources are produced based on the result of the Systematic Mapping Study (SMS) reported in [1]. The latest SNAPSHOT-version of the resource can be accessed through our resource landing page: <a href="https://w3id.org/semsys/sites/swemls-kg/">https://w3id.org/semsys/sites/swemls-kg/</a></p> <p>[1] Breit, A., Waltersdorfer, L., Ekaputra, J.F., Sabou, M., Ekelhart, A., Iana, A., Paulheim, H., Portisch, J., Revenko, A., Ten Teije, A., van Harmelen, F.: Combining Machine Learning and Semantic Web -A Systematic Mapping Study (under review). ACM CSUR (2022)<br> [2] Source code of swemls-toolkit is available at: https://github.com/semanticsystems/swemls-toolkit</p>
The knowledge graphs generated from CRE and The Session datasets
<p>This dataset contains the knowledge graphs generated from CRE and the Session datasets. All the knowledge graphs are in TTL format.</p>
MMpedia: A Large-scale Multi-modal Knowledge Graph
<p><strong>List of files:</strong></p> <ul> <li>entity2image.json: The entity to map images file</li> <li>MMpedia_triples.ttl: The triples file</li> <li>EntlistXX.tar: The image file</li> </ul> <p>MMpedia dataset is split into 132 subsets and each subset is compressed into a <code>.tar</code> file. After unziping files under the folder "MMpedia", the data structure are as following:</p> <pre><code>|-MMpedia |-Entitylist1 |-Entity1 |-1.jpg |-2.jpg |-3.jpg ... |-Entity2 |-Entity3 ... |-Entitylist2 |-Entitylist3</code></pre> <p>For example, the path <code>MMpedia/Entlist141/Bart Tanski/Bart Tanski+1.jpg</code> means the image corresponding to the entity "Bart Tanski"</p> <p><strong>Other image files can be found in following URLs:</strong></p> <p><a href="https://zenodo.org/record/7854781#.ZEU7Uc7iu38">MMpedia2 | Zenodo</a></p> <p><a href="https://zenodo.org/record/7855010">MMpedia3 | Zenodo</a></p> <p><a href="https://zenodo.org/record/7855226">MMpedia4 | Zenodo</a></p> <p> <strong>For more details, please refer to </strong><a href="https://github.com/Delicate2000/MMpedia">Delicate2000/MMpedia (github.com)</a></p>
Workflow for structured literature reviews using the Open Research Knowledge Graph (ORKG)
<p>Figure showing a workflow of making a structured literature review using the core features of the Open Research Knowledge Graph (ORKG). </p>
An Evaluation Framework for Mapping News Headlines to Event Classes in a Knowledge Graph
<p>News headline to event classes corpus derived from Wikidata.</p> <p>Github: https://github.com/mbouadeus/news-headline-event-linking</p>
Freebase Datasets for Robust Evaluation of Knowledge Graph Link Prediction Models
<p><strong>Freebase</strong> is amongst the largest public cross-domain knowledge graphs. It possesses three main data modeling idiosyncrasies. It has a strong <strong>type system</strong>; its properties are purposefully represented in <strong>reverse pairs</strong>; and it uses <strong>mediator objects</strong> to represent multiary relationships. These design choices are important in modeling the real-world. But they also pose nontrivial challenges in research of embedding models for knowledge graph completion, especially when models are developed and evaluated agnostically of these idiosyncrasies. We make available several variants of the Freebase dataset by inclusion and exclusion of these data modeling idiosyncrasies. This is the first-ever publicly available <strong>full-scale</strong> Freebase dataset that has gone through <strong>proper preparation</strong>. </p><p> </p><p>Dataset Details</p><p>The dataset consists of the four variants of Freebase dataset as well as related mapping/support files. For each variant, we made three kinds of files available:</p><ul><li>Subject matter triples file<ul><li><i>fb+/-CVT+/-REV</i> One folder for each variant. In each folder there are 5 files: train.txt, valid.txt, test.txt, entity2id.txt, relation2id.txt Subject matter triples are the triples belong to subject matters domains—domains describing real-world facts.<ul><li>Example of a row in train.txt, valid.txt, and test.txt: <ul><li>2, 192, 0</li></ul></li><li>Example of a row in entity2id.txt:<ul><li>/g/112yfy2xr, 2</li></ul></li><li>Example of a row in relation2id.txt:<ul><li>/music/album/release_type, 192</li></ul></li><li>Explaination<ul><li>"/g/112yfy2xr" and "/m/02lx2r" are the MID of the subject entity and object entity, respectively. "/music/album/release_type" is the realtionship between the two entities. 2, 192, and 0 are the IDs assigned by the authors to the objects.</li></ul></li></ul></li></ul></li><li>Type system file<ul><li><i>freebase_endtypes</i>: Each row maps an edge type to its required subject type and object type.<ul><li>Example<ul><li>92, 47178872, 90</li></ul></li><li>Explanation<ul><li>"92" and "90" are the type id of the subject and object which has the relationship id "47178872".</li></ul></li></ul></li></ul></li><li>Metadata files<ul><li><i>object_types</i>: Each row maps the MID of a Freebase object to a type it belongs to.<ul><li>Example<ul><li>/g/11b41c22g, /type/object/type, /people/person</li></ul></li><li>Explanation<ul><li>The entity with MID "/g/11b41c22g" has a type "/people/person"</li></ul></li></ul></li><li><i>object_names</i>: Each row maps the MID of a Freebase object to its textual label.<ul><li>Example<ul><li>/g/11b78qtr5m, /type/object/name, "Viroliano Tries Jazz"@en</li></ul></li><li>Explanation<ul><li>The entity with MID "/g/11b78qtr5m" has name "Viroliano Tries Jazz" in English.</li></ul></li></ul></li><li><i>object_ids</i>: Each row maps the MID of a Freebase object to its user-friendly identifier.<ul><li>Example<ul><li>/m/05v3y9r, /type/object/id, "/music/live_album/concert"</li></ul></li><li>Explanation<ul><li>The entity with MID "/m/05v3y9r" can be interpreted by human as a music concert live album.</li></ul></li></ul></li><li><i>domains_id_label</i>: Each row maps the MID of a Freebase domain to its label.<ul><li>Example<ul><li>/m/05v4pmy, geology, 77</li></ul></li><li>Explanation<ul><li>The object with MID "/m/05v4pmy" in Freebase is the domain "geology", and has id "77" in our dataset.</li></ul></li></ul></li><li><i>types_id_label</i>: Each row maps the MID of a Freebase type to its label.<ul><li>Example<ul><li>/m/01xljxh, /government/political_party, 147</li></ul></li><li>Explanation<ul><li>The object with MID "/m/01xljxh" in Freebase is the type "/government/political_party", and has id "147" in our dataset.</li></ul></li></ul></li><li><i>entities_id_label</i>: Each row maps the MID of a Freebase entity to its label.<ul><li>Example<ul><li>/g/11b78qtr5m, Viroliano Tries Jazz, 2234</li></ul></li><li>Explanation<ul><li>The entity with MID "/g/11b78qtr5m" in Freebase is "Viroliano Tries Jazz", and has id "2234" in our dataset.</li></ul></li><li><i>properties_id_label</i>: Each row maps the MID of a Freebase property to its label.<ul><li>Example<ul><li>/m/010h8tp2, /comedy/comedy_group/members, 47178867</li></ul></li><li>Explanation<ul><li>The object with MID "/m/010h8tp2" in Freebase is a property(relation/edge), it has label "/comedy/comedy_group/members" and has id "47178867" in our dataset.</li></ul></li></ul></li><li><i>uri_original2simplified</i> and <i>uri_simplified2original</i>: The mapping between original URI and simplified URI and the mapping between simplified URI and original URI repectively.<ul><li>Example<ul><li><i>uri_original2simplified</i><ul><li>"<a href="http://rdf.freebase.com/ns/type.property.unique">http://rdf.freebase.com/ns/type.property.unique</a>": "/type/property/unique"</li></ul></li><li><i>uri_simplified2original</i><ul><li>"/type/property/unique": "<a href="http://rdf.freebase.com/ns/type.property.unique">http://rdf.freebase.com/ns/type.property.unique</a>"</li></ul></li></ul></li><li>Explanation<ul><li>The URI "<a href="http://rdf.freebase.com/ns/type.property.unique">http://rdf.freebase.com/ns/type.property.unique</a>" in the original Freebase RDF dataset is simplified into "/type/property/unique" in our dataset.</li><li>The identifier "/type/property/unique" in our dataset has URI <a href="http://rdf.freebase.com/ns/type.property.unique">http://rdf.freebase.com/ns/type.property.unique</a> in the original Freebase RDF dataset.</li></ul></li></ul></li></ul></li></ul></li></ul>
Software Mentions Knowledge Graph
<p>Software plays an essential role in modern society. Its growth and evolution have given rise to a wide variety of pseudonyms to refer to the same software. This multiplicity of names can be confusing, complicating the accurate identification of software and its relationship to other programs and applications.<br> It is necessary to identify mentions of software in texts because they provide key information about its use and application in different contexts. Such identification makes it possible to establish a clearer link between programs, applications and their usefulness in various spheres. In addition, being able to group the different pseudonyms of a software under a single name facilitates its search and study, simplifying the acquisition of knowledge and the exchange of information between users and developers.<br> The term "Alias" refers to the name chosen to represent a group of pseudonyms that identify the same software tool. This alias acts as a common denominator, unifying the different ways in which a software tool can be mentioned in the scientific articles under analysis. From now on, we will refer to this unifying term as "alias" or "group", while "pseudonym" will designate the various forms of mentions of the same software that have appeared in the scientific literature.<br> Therefore, the main objective of this project is the construction of a knowledge graph [1] that groups the pseudonyms of scientific software tools into a common alias or group. This grouping will be done by analyzing and classifying the mentions of such software in academic publications provided by the article "CZ Software Mentions", published by the Chan Zuckerberg Initiative.</p>
MMiKG: A Knowledge Graph-based Platform for Path Mining of Microbiota-Mental Diseases Interactions
<p><strong>The original datasets released in MMiKG, containing relevant information like PMID of labels and relations (triples).</strong></p> <p><strong>Background:</strong></p> <p>Gut microbiota has been demonstrated to be crucial in gut-brain axis. In this research, knowledge graphs was leveraged to aggregate and assimilate relevant information on the microbiome-gut-brain axis and its intricate relationships with mental diseases and formed a knowledge graph platform named MMiKG</p> <p> ► <strong>Advantages of MMiKG:</strong></p> <ul> <li> Assist users in semantic search and visualization operations</li> <li> Make the scattered knowledge machine-readable and interpretable</li> <li> Boost users’ confidence in the accuracy of the information</li> <li> Support better decision-making</li> </ul> <p> ► <strong>What users can do with MMiKG:</strong></p> <ul> <li>Integrate diverse resources</li> <li>Infer potential associations between gut microbiota and mental diseases</li> </ul> <p><strong>Tools:</strong></p> <p>MMiKG contains '770' entities and '1,257' triples among them. These items cover '20' common mental illnesses, '270' types of gut microbes, and '480' distinct intermediates.</p> <p> ► <strong>MMiKG's original data includes two folders:</strong></p> <p> ♦ The labels folder:</p> <ul> <li>Present information of each entity’s 'Id', 'Name', 'Degree', 'Type'</li> <li>'Disease.csv', 'Intermediate.csv', 'Microbiota.csv</li> </ul> <p> ♦ The relationships folder: </p> <ul> <li>Present information of 'SourceId', 'TargetId', 'Number', 'Reference PMID'</li> <li>'Disease.csv', 'Intermediate.csv', 'Microbiota.csv' </li> </ul>
Results and log of LLM-KG-Bench runs described in article "Benchmarking the Abilities of Large Language Models for RDF Knowledge Graph Creation and Comprehension: How Well Do LLMs Speak Turtle?", Frey et al. 2023
<p>Results and log of LLM-KG-Bench runs described in article ""Benchmarking the Abilities of Large Language Models for RDF Knowledge Graph Creation and Comprehension: How Well Do LLMs Speak Turtle?", Frey et al. 2023, to appear in proceedings for workshop DL4KG@ISWC 2023.</p> <p>For data on task FactExtractStatic please contact authors.</p>
CLARA Knowledge Graph of licensed educational resources (using RDF-star, Standard reification, Singleton properties, or Named graphs)
<p><strong>CLARA</strong><br>This deposit is part of the <a href="https://project.inria.fr/clara/">CLARA project</a>. The CLARA project aims to empower teachers in the task of creating new educational resources. And in particular with the task of handling the licenses of reused educational resources.</p> <p>The present deposit contains the RDF files created using an RDF mapping (<a href="https://rml.io/">RML</a>) and a mapper (<a href="https://github.com/morph-kgc/morph-kgc">Morph-KGC</a>). It also contains the files JSON used as input. The corresponding pipeline can be found on <a href="https://gitlab.univ-nantes.fr/clara/pipeline">Gitlab</a>. The data used in that pipeline originate from <a href="https://www.x5gon.org/">X5GON</a>, a European project aiming to generate and gather open educational resources.</p> <p><strong>Knowledge graph content</strong><br>The present Knowledge Graph contains information about 45K Educational Resources (ERs) and 135K subjects (extracted from DBpedia).<br>That information contains </p> <ul> <li>the author,</li> <li>its title and description</li> <li>the license,</li> <li>a URL to the resource itself,</li> <li>the language of the ER,</li> <li>its mimetype,</li> <li>and finally which subject it talks about, and to what extent.</li> </ul> <p><br>That extent is given by two scores: a PageRank score and a Cosinus score.</p> <p>A particularity of the knowledge graph is its heavy use of RDF reification, across large multi-valued properties.<br>Thus four versions of the knowledge graph exist, using Standard reification, Singleton property, Named graphs, and RDF-star.</p> <p>The Knowledge Graph also contains <a href="https://databus.dbpedia.org/dbpedia/generic/categories">categories</a> originating from DBpedia. They help precise the subjects that are also extracted from DBpedia.</p> <p>The KG.zip files contain five types of files:</p> <ul> <li><strong>Authors_[</strong>X<strong>].nt</strong> - Those contain the authors' nodes, their type, and name.</li> <li><strong>ER_[</strong>X<strong>].nt/nq/ttl</strong> - Those contain the ERs and their information using the respective RDF reification model.</li> <li><strong>categories_skos_[</strong>X<strong>].ttl</strong> - Those contain the hierarchy of DBpedia categories.</li> <li><strong>categories_labels.ttl </strong>- This file contains additional information about the categories.</li> <li><strong>categories_article.ttl</strong> - This file contains the RDF triples that link the DBpedia subjects to the DBpedia categories.</li> </ul> <p> </p> <p><strong>JSON content</strong></p> <p>The original dataset was cut into multiple JSON files in order to make its processing easier. DBpedia categories were extracted as RDF and aren't present in the JSON files.<br><br>There are two types of files in the input-json.zip file:</p> <ul> <li><strong>authors_[</strong>X<strong>].json</strong> - Which lists the authors names</li> <li><strong>ER_[</strong>X<strong>].json</strong> - Which lists the ERs and their related information.<br>That information contains: <ul> <li>their <em>title.</em></li> <li>their <em>description.</em></li> <li>their <em>language</em> (and <em>language_detected</em>, only the first one is used in the pipeline here).</li> <li>their <em>license.</em></li> <li>their <em>mimetype.</em></li> <li>the <em>authors.</em></li> <li>the <em>date</em> of creation of the resource.</li> <li>a <em>url</em> linking to the resource itself.</li> <li>the subjects (named <em>concepts</em>) associated with the resource. With the corresponding scores.</li> </ul> </li> </ul> <p> </p> <p>If you do use this dataset, you can cite the corresponding paper:</p> <ul> <li>Kieffer, M., Fakih, G. & Serrano-Alvarado, P. (2023). Evaluating Reification with Multi-valued Properties in a Knowledge Graph of Licensed Educational Resources. Semantics, Leipzig, Germany.</li> </ul>
Dataset of A User-driven Hybrid Neuro-symbolic Approach for Knowledge Graph Creation from Relational Data
<p>This dataset contains the following:</p> <p>1. achieved percentage values of the generated RML rules using LXS and manually</p> <p>2. basic information about the example used and with which creation type users started</p> <p>3. all answers of users to the User Experience Questionnaire</p> <p>4. Answers to the structured part of the user interview</p>
Data from: Knowledge graphs for seismic data and metadata
Open the record for dataset details and reuse information.
Knowledge Graph about altmetrics of selected papers about COVID-19
<p>The knowledge graph (KG) contains data about altmetrics as well traditional indicators associated with 212 papers resulting from an early literature review. Publication dates encompass a time-window ranging from January 15th 2020 to February 24th 2020. The KG is represented as RDF and modelled by using the Indicators Ontology (I-Ont). I-Ont is an ontology for representing scholarly artefacts and their associated indicators, e.g. citation count or altmetrics such as the number of readers on Mendeley.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.