Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
219
datasets available to search
ShareScore release 0.7.1
Dataset results
219 results for βKnowledge Graphβ
A LARGE INTEGRATED KNOWLEDGE GRAPH OF ECNOMICS, FINANCE AND BANKING
<p>Creating the first release to obtain a DOI on Zenodo</p>
PheKnowLator Human Disease Knowledge Graph Benchmarks Archive
<h2><strong>PKT Human Disease KG Benchmark Builds</strong></h2> <p>The PheKnowLator (PKT) Human Disease KG (PKT-KG) was built to model mechanisms of human disease, which includes the Central Dogma and represents multiple biological scales of organization including molecular, cellular, tissue, and organ. The knowledge representation was designed in collaboration with a PhD-level molecular biologist (<a href="https://user-images.githubusercontent.com/8030363/195469903-86598760-40b7-4126-857c-3d6368305a86.png">Figure</a>). </p> <p>The <strong>PKT Human Disease KG</strong> was constructed using 12 OBO Foundry ontologies, 31 Linked Open Data sets, and results from two large-scale experiments (<a href="https://doi.org/10.48550/arXiv.2307.05727">Supplementary Material</a>). The 12 OBO Foundry ontologies were selected to represent chemicals and vaccines (i.e., ChEBI and Vaccine Ontology), cells and cell lines (i.e., Cell Ontology, Cell Line Ontology), gene/gene product attributes (i.e., Gene Ontology), phenotypes and diseases (i.e., Human Phenotype Ontology, Mondo Disease Ontology), proteins, including complexes and isoforms (i.e., Protein Ontology), pathways (i.e., Pathway Ontology), types and attributes of biological sequences (i.e., Sequence Ontology), and anatomical entities (Uberon ontology). The RO is used to provide relationships between the core OBO Foundry ontologies and database entities.</p> <p>The <strong>PKT Human Disease KG</strong> contained 18 node types and 33 edge types. Note that the number of nodes and edge types reflects those that are explicitly added to the core set of OBO Foundry ontologies and does not take into account the node and edge types provided by the ontologies. These nodes and edge types were used to construct 12 different PKT Human Disease benchmark KGs by altering the Knowledge Model (i.e., class- vs. instance-based), Relation Strategy (i.e., standard vs. inverse relations), and Semantic Abstraction (i.e., OWL-NETS (yes/no) with and without Knowledge Model harmonization [OWL-NETS Only vs. OWL-NETS + Harmonization]) parameters. Benchmarks within the PheKnowLator ecosystem are different versions of a KG that can be built under alternative knowledge models, relation strategies, and with or without semantic abstraction. They provide users with the ability to evaluate different modeling decisions (based on the prior mentioned parameters) and to examine the impact of these decisions on different downstream tasks.</p> <p>The Figures and Tables explaining attributes in the builds can be found <a href="https://github.com/callahantiff/PheKnowLator/wiki/Archived-Builds">here</a>.</p> <p> </p> <h3><strong>Build Data Access</strong></h3> <h4><strong>Important Build Information</strong></h4> <p>The benchmarks were originally built and stored using Google Cloud Platform (GCP) resources. For details and a complete description of this process, can be found on GitHub (<a href="https://github.com/callahantiff/PheKnowLator/tree/master/builds#readme">here</a>). Note that we have developed this Zenodo-based archive for the builds. While the original GCP resources contained all of the resources needed to generate the builds, due to the file size upload limits associated with each archive, we have limited the uploaded files to the KGs, associated metadata, and log files. The list of resources, including their URLs, and date of download, can all be found in the logs associated with each build.</p> <p>π For additional information on the KG file types please see the following <a href="https://github.com/callahantiff/PheKnowLator/wiki/KG-Construction#table-knowledge-graph-build-output">Wiki page</a>, which is also available as a download from this repository (PheKnowLator_HumanDiseaseKG_Output_FileInformation.xlsx). </p> <h4><strong>v1.0.0</strong></h4> <ul> <li>KGs: <a href="../doi/10.5281/zenodo.7030200">https://zenodo.org/doi/10.5281/zenodo.7030200</a></li> <li>Embeddings: <a href="../doi/10.5281/zenodo.7030188">https://zenodo.org/doi/10.5281/zenodo.7030188</a></li> </ul> <h4><strong>All Other Build Versions</strong></h4> <p><strong>Class-based Builds</strong></p> <p><em>Standard Relations</em></p> <ul> <li>OWL Build <ul> <li>v2.0.0: <a href="../doi/10.5281/zenodo.7029957">MAY2020</a><a href="../record/8178783">; </a><a href="../doi/10.5281/zenodo.8180239">JAN2021</a>; <a href="../doi/10.5281/zenodo.8180539">FEB2021</a></li> <li>v2.1.0: <a href="../doi/10.5281/zenodo.8180774">MAY2021</a>;<a href="../doi/10.5281/zenodo.8180825"> JUN2021</a>; <a href="../doi/10.5281/zenodo.8180972">JUL2021</a>; <a href="../doi/10.5281/zenodo.8183987">AUG2021</a>;<a href="../doi/10.5281/zenodo.8184090"> SEP2021</a></li> <li>v3.0.2: <a href="../doi/10.5281/zenodo.8184131">OCT2021</a>; <a href="../doi/10.5281/zenodo.8184205">NOV2021</a></li> </ul> </li> <li>OWL-NETS Build <ul> <li>v2.0.0: <a href="../doi/10.5281/zenodo.7029953">MAY2020</a><a href="../record/8178783">; </a><a href="../doi/10.5281/zenodo.8180255">JAN2021</a>; <a href="../doi/10.5281/zenodo.8180545">FEB2021</a></li> <li>v2.1.0: <a href="../doi/10.5281/zenodo.8180772">MAY2021</a>; <a href="../doi/10.5281/zenodo.8180827">JUN2021</a>; <a href="../doi/10.5281/zenodo.8180974">JUL2021</a>; <a href="../doi/10.5281/zenodo.8183989">AUG2021</a>; <a href="../doi/10.5281/zenodo.8184088">SEP2021</a></li> <li>v3.0.2: <a href="../doi/10.5281/zenodo.8184133">OCT2021</a>; <a href="../doi/10.5281/zenodo.8184208">NOV2021</a></li> </ul> </li> </ul> <p><em>Inverse Relations</em></p> <ul> <li>OWL Build <ul> <li>v2.0.0: <a href="../doi/10.5281/zenodo.7029893">MAY2020</a><a href="../record/8178783">; </a><a href="../doi/10.5281/zenodo.8180269">JAN2021</a>; <a href="../doi/10.5281/zenodo.8180550">FEB2021</a></li> <li>v2.1.0: <a href="../doi/10.5281/zenodo.8180766">MAY2021</a>; <a href="../doi/10.5281/zenodo.8180829">JUN2021</a>; <a href="../doi/10.5281/zenodo.8180976">JUL2021</a>;<a href="../doi/10.5281/zenodo.8183991"> AUG2021</a>; <a href="../doi/10.5281/zenodo.8184086">SEP2021</a></li> <li>v3.0.2: <a href="../doi/10.5281/zenodo.8184135">OCT2021</a>; <a href="../doi/10.5281/zenodo.8184210">NOV2021</a></li> </ul> </li> <li>OWL-NETS Build <ul> <li>v2.0.0: <a href="../doi/10.5281/zenodo.7029921">MAY2020</a><a href="../record/8178783">; </a><a href="../doi/10.5281/zenodo.8180279">JAN2021</a>; <a href="../doi/10.5281/zenodo.8180555">FEB2021</a></li> <li>v2.1.0: <a href="../doi/10.5281/zenodo.8180768">MAY2021</a>; <a href="../doi/10.5281/zenodo.8180833">JUN2021</a>; <a href="../doi/10.5281/zenodo.8180982">JUL2021</a>; <a href="../doi/10.5281/zenodo.8183993">AUG2021</a>; <a href="../doi/10.5281/zenodo.8184084">SEP2021</a></li> <li>v3.0.2: <a href="../doi/10.5281/zenodo.8184137">OCT2021</a>; <a href="../doi/10.5281/zenodo.8184212">NOV2021</a></li> </ul> </li> </ul> <p><strong>Instance-based Builds</strong></p> <p><em>Standard Relations</em></p> <ul> <li>OWL Build <ul> <li>v2.0.0: <a href="../doi/10.5281/zenodo.7029941">MAY2020</a><a href="../record/8178783">; </a><a href="../doi/10.5281/zenodo.8180333">JAN2021</a>;<a href="../doi/10.5281/zenodo.8180558"> FEB2021</a></li> <li>v2.1.0: <a href="../doi/10.5281/zenodo.8180764">MAY2021</a>; <a href="../doi/10.5281/zenodo.8180835">JUN2021</a>; <a href="../doi/10.5281/zenodo.8180984">JUL2021</a>; <a href="../doi/10.5281/zenodo.8183995">AUG2021</a>; <a href="../doi/10.5281/zenodo.8184082">SEP2021 </a></li> <li>v3.0.2: <a href="../doi/10.5281/zenodo.8184139">OCT2021</a>; <a href="../doi/10.5281/zenodo.8184216">NOV2021 </a></li> </ul> </li> <li>OWL-NETS Build <ul> <li>v2.0.0: <a href="../doi/10.5281/zenodo.7029939">MAY2020</a><a href="../record/8178783">; </a><a href="../doi/10.5281/zenodo.8180335">JAN2021</a>;<a href="../doi/10.5281/zenodo.8180564"> FEB2021</a></li> <li>v2.1.0: <a href="../doi/10.5281/zenodo.8180762">MAY2021</a>; <a href="../doi/10.5281/zenodo.8180837">JUN2021</a>; <a href="../doi/10.5281/zenodo.8180986">JUL2021</a>; <a href="../doi/10.5281/zenodo.8183997">AUG2021</a>; <a href="../doi/10.5281/zenodo.8184080">SEP2021</a></li> <li>v3.0.2: <a href="../doi/10.5281/zenodo.8184141">OCT2021</a>; <a href="../doi/10.5281/zenodo.8184218">NOV2021</a></li> </ul> </li> </ul> <p><em>Inverse Relations</em></p> <ul> <li>OWL Build <ul> <li>v2.0.0: <a href="../doi/10.5281/zenodo.7029945">MAY2020</a><a href="../record/8178783">; </a><a href="../doi/10.5281/zenodo.8180338">JAN2021</a>; <a href="../doi/10.5281/zenodo.8180588">FEB2021</a></li> <li>v2.1.0: <a href="../doi/10.5281/zenodo.8180758">MAY2021</a>; <a href="../doi/10.5281/zenodo.8180878">JUN2021</a>; <a href="../doi/10.5281/zenodo.8180992">JUL2021</a>; <a href="../doi/10.5281/zenodo.8184001">AUG2021</a>; <a href="../doi/10.5281/zenodo.8184078">SEP2021 </a></li> <li>v3.0.2: <a href="../doi/10.5281/zenodo.8184143">OCT2021</a>; <a href="../doi/10.5281/zenodo.8184220">NOV2021</a></li> </ul> </li> <li>OWL-NETS Build <ul> <li>v2.0.0: <a href="../doi/10.5281/zenodo.7029919">MAY2020</a><a href="../record/8178783">; </a><a href="../doi/10.5281/zenodo.8180340">JAN2021</a>; <a href="../doi/10.5281/zenodo.8180584">FEB2021</a></li> <li>v2.1.0: <a href="../doi/10.5281/zenodo.8180756">MAY2021</a>; <a href="../doi/10.5281/zenodo.8180823">JUN2021</a>; <a href="../doi/10.5281/zenodo.8180996">JUL2021</a>; <a href="../doi/10.5281/zenodo.8184003">AUG2021</a>; <a href="../doi/10.5281/zenodo.8184076">SEP2021 </a></li> <li>v3.0.2: <a href="../doi/10.5281/zenodo.8184145">OCT2021</a>; <a href="../doi/10.5281/zenodo.8184222">NOV2021</a></li> </ul> </li> </ul>
European Olfactory Knowledge Graph
<p>The European Olfactory Knowledge Graph (EOKG) includes information about smell from (digital) text and image collections from the European history (1600-1920), extracted in the context of the <a title="Odeuropa" href="https://odeuropa.eu/" target="_blank" rel="noopener">Odeuropa project</a> in a cultural heritage preservation perspective.</p> <p>It contains over 2,500,000 olfactory reference coming from over 43,000 images and 2,400,000 texts in six languages, organised according to the <a href="https://data.odeuropa.eu/ontology" target="_blank" rel="noopener">Odeuropa Ontology</a> and leveraging machine learning to recognise and categorise olfactory elements.</p> <h3>Additional Links</h3> <div> <div> <div> <div> <div> <div> <div> <div> <div> <div> <div> <div>EOKG Vocabularies: <a href="https://vocab.odeuropa.eu/" target="_blank" rel="noopener noreferrer">https://vocab.odeuropa.eu/</a> (vocabulary browser)<br>Odeuropa Ontology: <a href="https://data.odeuropa.eu/ontology/" target="_blank" rel="noopener noreferrer">https://data.odeuropa.eu/ontology/</a> (data model)<br>EOKG API: <a href="https://grlc.eurecom.fr/api/Odeuropa/kg-api/" target="_blank" rel="noopener noreferrer">https://grlc.eurecom.fr/api/Odeuropa/kg-api/</a> (API)<br>EOKG technical report: <a href="https://odeuropa.eu/wp-content/uploads/2024/10/D4_3_European_Olfactory_Knowledge_Graph_v2_final.pdf">https://odeuropa.eu/wp-content/uploads/2024/10/D4_3_European_Olfactory_Knowledge_Graph_v2_final.pdf</a> (documentation)<br>Odeuropa Smell Explorer: <a href="https://explorer.odeuropa.eu/" target="_blank" rel="noopener noreferrer">https://explorer.odeuropa.eu/</a> (demonstrator)</div> </div> </div> </div> </div> </div> </div> </div> <div> </div> </div> </div> </div> </div>
Datasets for Paper "MetagenomicKG: a knowledge graph for metagenomic applications"
<p>This repository contains some required data that is used for building MetagenomicKG. Please see more details in <a href="https://github.com/KoslickiLab/MetagenomicKG">https://github.com/KoslickiLab/MetagenomicKG</a>.</p>
Assessing the Overlap of Science Knowledge Graphs: A Quantitative Analysis β exact and related matches
<p>Results of the 'Assessing the Overlap of Science Knowledge Graphs: A Quantitative Analysis' papers. There are 2 datasets:</p> <ul> <li>'exact_matches.csv': contains detailed information about the concepts present both in OpenAlex and OpenAIRE.</li> <li>'related_matches.csv': contains detailed information about the concepts from OpenAlex and OpenAIRE that were not present in both KGs but got aligned following the algorithm presented in the paper.</li> </ul> <p>The detailed information refers to the following column:</p> <ul> <li>Category1: name of the first category</li> <li>Source1: source of the first category ('OpenAlex' or 'OpenAIRE')</li> <li>Category2: name of the second category</li> <li>Source2: source of the first category ('OpenAlex' or 'OpenAIRE')</li> <li>Similarity: semantic similarity value of the two categories</li> <li>PapersInC1: number of papers from the collected dataset belonging to the first category</li> <li>PapersInC2: number of papers from the collected dataset belonging to the second category</li> <li>PapersInBoth: number of papers from the collected dataset belonging to both of the categories</li> <li>Agreement: the value of the agreement of the categories in the tw KGs (Intersection over Union)</li> </ul>
SemTab 2024: Semantic Web Challenge on Tabular Data to Knowledge Graph Matching Data Sets - WikidataTables2024R1 and WikidataTables2024R2
<p>Data Sets from the ISWC 2024 Semantic Web Challenge on Tabular Data to Knowledge Graph Matching, Round 1, Wikidata Tables. Links to other datasets can be found on the challenge website: https://sem-tab-challenge.github.io/2024/ as well as the proceedings of the challenge published on CEUR.</p> <p>For details about the challenge, see: http://www.cs.ox.ac.uk/isg/challenges/sem-tab/</p> <p>For 2024 edition, see: https://sem-tab-challenge.github.io/2024/</p> <p>Note on License: This data includes data from the following sources. Refer to each source for license details:<br>- Wikidata https://www.wikidata.org/</p> <p>THIS DATA IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.</p>
Resources of IncRML: Incremental Knowledge Graph Construction from Heterogeneous Data Sources
<h2>IncRML resources</h2> <p>This Zenodo dataset contains all the resources of the paper 'IncRML: Incremental Knowledge Graph Construction from Heterogeneous Data Sources' submitted to the Semantic Web Journal's Special Issue on Knowledge Graph Construction. This resource aims to make the paper experiments fully reproducible through our <a href="https://github.com/kg-construct/exectool" target="_blank" rel="noopener">experiment tool</a> written in Python which was already used before in the <a href="https://doi.org/10.5281/zenodo.7837289" target="_blank" rel="noopener">Knowledge Graph Construction Challenge by the ESWC 2023 Workshop on Knowledge Graph Construction</a>. The exact Java JAR file of the RMLMapper (rmlmapper.jar) is also provided in this dataset which was used to execute the experiments. This JAR file was executed with Java OpenJDK 11.0.20.1 on Ubuntu 22.04.1 LTS (Linux 5.15.0-53-generic). Each experiment was executed 5 times and the median values are reported together with the standard deviation of the measurements.</p> <h2>Datasets</h2> <p>We provide both dataset dumps of the GTFS-Madrid-Benchmark and of real-life use cases from Open Data in Belgium.<br>GTFS-Madrid-Benchmark dumps are used to analyze the impact on execution time and resources, while the real-life use cases aim to verify the approach on different types of datasets since the GTFS-Madrid-Benchmark is a single type of dataset which does not advertise changes at all.</p> <h3>Benchmarks</h3> <ul> <li>GTFS-Madrid-Benchmark: change types with fixed data size and amount of changes: additions-only, modifications-only, deletions-only (11 versions)</li> <li>GTFS-Madrid-Benchmark: amount of changes with fixed data size: 0%, 25%, 50%, 75%, and 100% changes (11 versions)</li> <li>GTFS-Madrid-Benchmark: data size with fixed amount of changes: scales 1, 10, 100 (11 versions)</li> </ul> <h3>Real-world datasets</h3> <ul> <li>Traffic control center Vlaams Verkeerscentrum (Belgium): traffic board messages data (1 day, 28760 versions)</li> <li>Meteorological institute KMI (Belgium): weather sensor data (1 day, 144 versions)</li> <li>Public transport agency NMBS (Belgium): train schedule data (1 week, 7 versions)</li> <li>Public transport agency De Lijn (Belgium): busses schedule data (1 week, 7 versions)</li> <li>Bike-sharing company BlueBike (Belgium): bike-sharing availability data (1 day, 1440 versions)</li> <li>Bike-sharing company JCDecaux (EU): bike-sharing availability data (1 day, 1440 versions)</li> <li>OpenStreetMap (World): geographical map data (1 day, 1440 versions)</li> </ul> <h3>Ingestion</h3> <p>Real-world datasets LDES output was converted into SPARQL UPDATE queries and executed against Virtuoso to have an estimate for non-LDES clients how incremental generation impacted ingestion into triplestores.</p> <h2>Remarks</h2> <ol> <li>The first version of each dataset is always used as a baseline. All next versions are applied as an update on the existing version. The reported results are only focusing on the updates since these are the actual incremental generation.</li> <li>GTFS-Change-50_percent-{ALL, CHANGE}.tar.xz datasets are not uploaded as GTFS-Madrid-Benchmark scale 100 because both share the same parameters (50% changes, scale 100). Please use GTFS-Scale-100-{ALL, CHANGE}.tar.xz for GTFS-Change-50_percent-{ALL, CHANGE}.tar.xz</li> <li>All datasets are compressed with XZ and provided as a TAR archive, be aware that you need sufficient space to decompress these archives! 2 TB of free space is advised to decompress all benchmarks and use cases. The expected output is provided as a ZIP file in each TAR archive, decompressing these requires even more space (4 TB).</li> </ol> <h2>Reproducing</h2> <p>By using our <a href="https://github.com/kg-construct/exectool" target="_blank" rel="noopener">experiment tool</a>, you can easily reproduce the experiments as followed:</p> <ol> <li>Download one of the TAR.XZ archives and unpack them.</li> <li>Clone the GitHub repository of our experiment tool and install the Python dependencies with '<em>pip install -r requirements.txt'.</em></li> <li>Download the rmlmapper.jar JAR file from this Zenodo dataset and place it inside the experiment tool root folder.</li> <li>Execute the tool by running: '<em>./exectool --root=/path/to/the/root/of/the/tarxz/archive --runs=5 run</em>'. The argument '<em>--runs=5</em>' is used to perform the experiment 5 times.</li> <li>Once executed, you can generate the statistics by running: '<em>./exectool --root=/path/to/the/root/of/the/tarxz/archive stats</em>'.</li> </ol> <h2>Testcases</h2> <p>Testcases to verify the integration of RML and LDES with IncRML, see <a href="https://doi.org/10.5281/zenodo.10171394">https://doi.org/10.5281/zenodo.10171394</a></p>
PrimeKGQA, the dataset from paper: Bridging the Gap: Generating a Comprehensive Biomedical Knowledge Graph Question Answering Dataset
<p>Despite the plethora of resources such as large-scale corpora and manually curated Knowledge Graphs (KGs), the ability to perform reasoning with natural language inputs over biomedical graphs remains challenging due to insufficient training data. We propose a novel method for automatically constructing a Biomedical Knowledge Graph Question Answering (BioKGQA) dataset sourced from PrimeKG, the largest precision medicine-oriented KG. In total,<br>we create 83999 question-answer pairs along with their respective SPARQL queries. Our approach generates a diverse array of contextually relevant questions covering a wide spectrum of biomedical concepts and levels of complexity. We evaluate our method based on automatic metrics alongside manual annotations. We establish novel standards tailored for KGQA systems to highlight the linguistic correctness and semantical faithfulness of the generated questions based on extracted KG facts. The compiled dataset – PrimeKGQA – serves as a valuable benchmarking resource for advancing knowledge-driven biomedical research and evaluating KGQA system.</p>
BOCK: Biological networks and Oligogenic Combinations as a Knowledge graph
<p>BOCK is a knowledge graph integrating oligogenic disease information (originally from the Oligogenic Disease Database (Natchtegael et al. 2022)) together with multiple biological networks and ontologies.</p> <p>Compared to more generic knowledge graphs, we selected specifically networks relevant to understand the molecular mechanisms of epistasis, placing genes as the central entities, and focused on trusted resources describing a large set of human genes and their interactions.</p> <p>All entities in the KG are linked to their source database entry via an URI (Uniform Resource Identifier) to facilitate integrations within larger bioinformatics linked data repositories.</p> <p>BOCK 2.0 integrates recent versions of the used ontologies and databases, as well as additional pathway-specific (The Reactome Pathway Knowledgebase 2024, Milacic et al.) and tissue-specific information (COXPRESdb v8, Obayashi et al.). Additionally the database used for the coexpression relation between genes, has been replaced by COXPRESdb v8.</p> <p>We provide BOCK 2.0 in three formats:</p> <ol> <li><strong>GraphML (Graph Markup Language)</strong>: a network format enabling the fast import of the KG by multiple libraries (e.g networkx) and tools (e.g Cytoscape).</li> <li><strong>XML (Extensible Markup Language)</strong>: a text-encoding system that is human-readable and compatible with many systems.</li> <li><strong>Neo4J import files</strong>: tab-separated files that can be easily imported into Neo4J using the neo4j-admin utils.</li> </ol> <p> </p>
Analysis materials for "Defininig a Knowledge Graph Development Process through a Systematic Review"
<p>This depository stores the analysis materials for the article "Defining a Knowledge Graph Development Process through a Systematic Review". It includes:</p> <ul> <li><strong>Analysis of KG development process - Articles.csv </strong>- a table of summary of the articles included in the systematic review.</li> <li><strong>Analysis of KG development process - Tasks by level (count).csv</strong> - a table of counting the frequency of the tasks in the knowledge graph development process.</li> <li><strong>Analysis of KG development process - Tasks by level count (synonyms) (1).csv </strong>- a table of counting the frequency of the tasks in the knowledge graph development process after its been adjusted to synonyms.</li> <li><strong>Knowledge graph development processes - </strong>a folder of process figures from the articles that have been included in the systematic review.</li> </ul>
Zero-Shot Information Extraction to Enhance a Knowledge Graph Describing Silk Textiles - English and Spanish neighborhood sub-graphs
<p>Two language-specific sub-graphs (English and Spanish) based on the ConceptNet Knowledge Graph. These two files are required to run the code for reproducing the results reported in the paper <a href="https://aclanthology.org/2021.latechclfl-1.16/">"Zero-Shot Information Extraction to Enhancea Knowledge Graph Describing Silk Textiles"</a> at the <a href="https://sighum.wordpress.com/events/latech-clfl-2021/">LaTeCH-CLfL 2021</a> workshop co-located with <a href="https://2021.emnlp.org/">EMNLP 2021</a>.</p>
OC-782K: Knowledge Graph of "Scientometrics" modelled according to the OpenCitations Data Model
<p>This dataset is a knowledge graph extracted from a <a href="https://static.aminer.cn/misc/na-data-kdd18.zip">t</a>riplestore covering information about the journal <em>Scientometrics</em> and modelled according to the OpenCitations Data Model. The original triplestore is available <a href="https://doi.org/10.5281/zenodo.5151264">here</a>. This KG was extracted for a research project on knowledge graph embeddings (KGEs) for author disambiguation. Structural triples of the knowledge graph are split into training, testing and validation for applying representation learning methods. Textual literals and numeric literals were stored separately in order to implement multimodal approaches for KGEs (see <a href="https://arxiv.org/abs/1802.00934">arXiv:1802.00934</a>). For the same reason, textual literals and numeric literals are already stored into sentence embeddings and a numeric matrix respectively in the files <em>textual_literals.npy </em>and <em>numeric_literals.npy</em>. The file <em>and_eval</em><em>.json </em>contains the evaluation dataset used for evaluating our AND architecture. For the script used to gather this dataset see the GitHub repository: <a href="https://github.com/sntcristian/and-kge/tree/main/aminer">https://github.com/sntcristian/and-kge/tree/main/open-citations</a>.</p>
SILKNOW Knowledge Graph
<p>SILKNOW is a research project that aims at improving the understanding, conservation and dissemination of the<br> European silk heritage from the 15th to the 19th century. The SILKNOW knowledge graph (KG) lies at<br> the center of the application of Semantic Web technologies and computing research to the needs of museums and every other user of this knowledge. The underlying data model is based on CIDOC-CRM and data mappings which are realised and implementedwith conversion tools developed for SILKNOW.<br> <br> Full instructions on how to deploy this KG can be found in the README.md file. See also <a href="http://See also https://github.com/silknow/knowledge-base">https://github.com/silknow/knowledge-base</a></p>
The WASABI Dataset and RDF Knowledge Graph
<p>The WASABI Dataset and RDF Knowledge Graph is rich dataset describing more than 2 millions commercial songs, 200K albums and 77K artists (mainly from pop/rock culture). It comprises data extracted from music databases on the Web, and resulting from the processing of song lyrics and from audio analysis.</p> <p>This is version 2 of the dataset. It consists of two representation formats:</p> <ul> <li>The JSON format provides all data extracted from the MongoDB database that backs up the web application</li> <li>The RDF Knowledge Graph that represents the same data following the WASABI ontology.</li> </ul> <p>WASABI project homepage: http://wasabihome.i3s.unice.fr/</p> <p>Github: https://github.com/micbuffa/WasabiDataset</p>
Dataset for WWW2022 accepted paper "SelfKG: Self-Supervised Entity Alignment in Knowledge Graphs"
<p>Datasets for WWW2022 accepted paper "SelfKG: Self-Supervised Entity Alignment in Knowledge Graphs"</p> <p>The code repository is <a href="https://github.com/THUDM/SelfKG">here</a>, and our paper is <a href="https://arxiv.org/abs/2203.01044">here</a>.</p>
Conflict Event Knowledge Graph based on the ongoing Ukraine-Russia Conflict
<p><strong>Conflict Event Knowledge Graph</strong> is a <strong>Knowledge Graph</strong> that links the tweets and the current events of <strong>Russia-Ukraine Conflict</strong> portrayed in <strong>English Wikipedia</strong> from <strong>24th February 2022 to 4th March 2022</strong>. </p> <p><strong>Abstract</strong>: In the current situation of Russia-Ukraine Conflict, numerous contents are posted daily to engage in the discourse about different events in this conflict. The goal of this study is to provide a framework to enable analysis of these events and the Twitter data, utilizing entity linking. The relevant tweets and events are integrated into a Knowledge Graph <strong>ConflictEventKG</strong>, and the resources are made publicly available.</p> <p><strong>Homepage</strong>: <a href="https://siebeniris.github.io/ConflictEventKG/">https://siebeniris.github.io/ConflictEventKG/</a></p> <p><strong>Code</strong>: <a href="https://github.com/siebeniris/ConflictEventKG">https://github.com/siebeniris/ConflictEventKG</a></p> <p> </p>
Wikipedia Knowledge Graph dataset
<p>Wikipedia is the largest and most read online free encyclopedia currently existing. As such, Wikipedia offers a large amount of data on all its own contents and interactions around them, as well as different types of open data sources. This makes Wikipedia a unique data source that can be analyzed with quantitative data science techniques. However, the enormous amount of data makes it difficult to have an overview, and sometimes many of the analytical possibilities that Wikipedia offers remain unknown. In order to reduce the complexity of identifying and collecting data on Wikipedia and expanding its analytical potential, after collecting different data from various sources and processing them, we have generated a dedicated Wikipedia Knowledge Graph aimed at facilitating the analysis, contextualization of the activity and relations of Wikipedia pages, in this case limited to its English edition. We share this Knowledge Graph dataset in an open way, aiming to be useful for a wide range of researchers, such as informetricians, sociologists or data scientists.</p> <p>There are a total of 9 files, all of them in tsv format, and they have been built under a relational structure. The main one that acts as the core of the dataset is the <em><strong>page</strong></em> file, after it there are 4 files with different entities related to the Wikipedia pages (<em><strong>category</strong></em>, <em><strong>url</strong></em>, <em><strong>pub</strong></em> and <em><strong>page_property</strong></em> files) and 4 other files that act as "intermediate tables" making it possible to connect the pages both with the latter and between pages (<em><strong>page_category</strong></em>, <em><strong>page_url</strong></em>, <em><strong>page_pub</strong></em> and <em><strong>page_link</strong></em> files).</p> <p>The document <em><strong>Dataset_summary</strong></em> includes a detailed description of the dataset.</p> <p>Thanks to Nees Jan van Eck and the Centre for Science and Technology Studies (CWTS) for the valuable comments and suggestions.</p>
Kiez Benchmarking Knowledge Graph Embeddings
<p>This upload contains pre-calculated Knowledge Graph Embeddings produced by our study "<a href="https://dbs.uni-leipzig.de/file/KIEZ_KEOD_2021_Obraczka_Rahm.pdf">An Evaluation of Hubness Reduction Methods for Entity Alignment with Knowledge Graph Embeddings</a>"</p>
EMAKG: an enriched version of the Microsoft Academic Knowledge Graph
<p>The <strong>Enhanced MAKG</strong> (<strong>EMAKG</strong>) provides an updated and enriched version of the Microsoft Academic Knowledge Graph (MAKG).</p> <p>The EMAKG is a large dataset of <strong>scientific publications</strong> and related entities such as <strong>authors</strong>, <strong>affiliations</strong>, <strong>venues</strong>, and <strong>fields of study</strong>. Data includes <strong>authors' careers and networks of collaborations</strong>, <strong>linguistics features</strong>, together with <strong>worldwide yearly authors' stocks and flows</strong>.</p> <p>The EMAKG data is mainly based on the MAKG - <strong>Version 2020-06-19 (March 25, 2021)</strong>.</p> <p>Methods: <a href="https://github.com/LauraPollacci/EMAKG">https://github.com/LauraPollacci/EMAKG</a></p> <p><em>Version 0.0 (reduced version)</em><br> Version 0.0 provides a set of EMAKG subsets, some of which are in abridged form:</p> <p>01.AffiliationsGeo: Affiliations subset.<br> 03.ConferenceInstances: Conferences subset.<br> 04.Conference Series: ConferenceSeries subset. <br> 05.Journals: Journals subset. <br> 06.24.PaperAuthorAffiliations_Disambiguated: Relationships between papers and disambiguated authors. <br> 09.PaperResources: URLs and resources of publications.<br> 10.Papers: Papers subset.<br> 12.EntityRelatedEntities: Connections between entities.<br> 13.FieldOfStudyChildren: Field of study kinship relations.<br> 14.FieldOfStudyExtendedAttributes: Fields of study co-references between different datasets.<br> 15.FieldsOfStudy: Fields of study subset.<br> 16.PaperFieldsOfStudy: Relationships between papers and fields of study.<br> 18.RelatedFieldOfStudy: Relationships between symptoms, medical treatments, disease causes and fields of study.<br> 19.PaperCitationContexts: Contexts of citations in CiTo.<br> 20.AbstractsProcessed_Chunk0-14: Chunk of processed abstracts.<br> 22.FieldOfStudyLabeled: Tags and scores of fields of studies.<br> 23.Authors_disambiguated: Disambiguated authors subset.<br> 24.PaperAuthorAffiliation_Disambiguated: Relationships between papers, disambiguated authors and affiliations.<br> 25.AuthorORCID: Authors' ORCIDs.<br> 26.AuthorCareer: Authors' yearly publications.<br> 27.AuthorYearLocation: Authors' yearly locations.<br> 28.AuthorEgoNetworks_2000-2014: Authors' ego networks from 2000 to 2014.<br> 29.CountryAnnualFlowsAggregated: Flows aggregated by country and year.<br> 30.FlowsAnnual: Annual country to country flows.<br> 31.StocksAnnual: Annual stocks aggregated by country.<br> 32.PaperFieldsOfStudyLabeled: Publications tagged with fields of studies.<br> 33.Authors_disambiguated_Hindex: H-index of disambiguated authors.</p>
Knowledge Graph: tyrolean mining documents 15th and 16th century
<p>The dataset contains a Knowledge Graph (.nq file) of two historical mining documents: “Verleihbuch der Rattenberger Bergrichter” ( Hs. 37, 1460-1463) and “Schwazer Berglehenbuch” (Hs. 1587, approx. 1515) stored by the Tyrolean Regional Archive, Innsbruck (Austria). The user of the KG may explore the montanistic network and relations between people, claims and mines in the late medieval Tyrol. The core regions concern the districts Schwaz and Kufstein (Tyrol, Austria).</p> <p>The ontology used to represent the claims is CIDOC CRM, an ISO certified ontology for Cultural Heritage documentation. Supported by the Karma tool the KG is generated as RDF (Resource Description Framework). The generated RDF data is imported into a Triplestore, in this case GraphDB, and then displayed visually. This puts the data from the early mining texts into a semantically structured context and makes the mutual relationships between people, places and mines visible.</p> <p>Both documents and the Knowledge Graph were processed and generated by the research team of the project “Text Mining Medieval Mining Texts”. The research project (2019-2022) was carried out at the university of Innsbruck and funded by go!digital next generation programme of the Austrian Academy of Sciences.</p> <p>Citeable Transcripts of the historical documents are online available:<br> Hs. 37 DOI: 10.5281/zenodo.6274562<br> Hs. 1587 DOI: 10.5281/zenodo.6274928</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.