Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
13
datasets available to search
ShareScore release 0.7.1
Dataset results
13 results for “Knowledge Graph Embedding”
Universal Knowledge Graph Embeddings
<p>The dataset provides embeddings for entities and relations in DBpedia (English) and Wikidata. The two knowledge graphs are first merged using a novel approach that we developed by leveraging the sameAs links between them. Then, we used the state-of-the-art embedding model ConEx to compute embeddings of the merge. Our embeddings are called universal knowledge graph embeddings.</p>
Kiez Benchmarking Knowledge Graph Embeddings
<p>This upload contains pre-calculated Knowledge Graph Embeddings produced by our study "<a href="https://dbs.uni-leipzig.de/file/KIEZ_KEOD_2021_Obraczka_Rahm.pdf">An Evaluation of Hubness Reduction Methods for Entity Alignment with Knowledge Graph Embeddings</a>"</p>
Integrated knowledge graphs and embeddings vectors for drug-drug interaction prediction
<p>The associated Knowledge Graphs for predicting potential drug-drug interaction, which is used in our paper titled "Drug-Drug Interaction Prediction Based on Knowledge Graph Embeddings and Convolutional-LSTM Network". Please consider citing the following paper if you plan or used our datasets.</p> <p>Md. Rezaul Karim, Michael Cochez, Joao Bosco Jares, Mamtaz Uddin, Oya Beyan, and Stefan Decker, "Drug-Drug Interaction Prediction Based on Knowledge Graph Embeddings and Convolutional-LSTM Network", In 10th ACM Int’l Conference on Bioinformatics, Computational Biology and Health Informatics (ACM-BCB ’19), September 7–10, 2019, Niagara Falls, NY, USA.</p>
Dataset for paper: " Knowledge Graph Embeddings based Approach for Author Name Disambiguation using Literals"
<p>This dataset consists in two distinct scholarly knowledge graph created from two publicly available bibliographic datasets: 1) a triplestore covering information about the journal <em>Scientometrics</em> provided by <em>OpenCitations</em> (available <a href="https://doi.org/10.5281/zenodo.5151264">here</a>), and 2) the <em>AMiner </em>AND benchmark from 2018 available <a href="https://static.aminer.cn/misc/na-data-kdd18.zip">here</a>. This KG was extracted for a research project on knowledge graph embeddings (KGEs) for author disambiguation. Structural triples of the knowledge graphs are split into training, testing and validation for applying representation learning methods. Textual literals and numeric literals were stored separately in order to implement multimodal approaches for KGEs (see <a href="https://arxiv.org/abs/1802.00934">arXiv:1802.00934</a>). For the same reason, textual literals and numeric literals are already stored into sentence embeddings and a numeric matrix respectively in the files <em>textual_literals.npy </em>and <em>numeric_literals.npy </em>in order to simplify the representation learning task. The file <em>and_eval.json</em> of each KG contains the evaluation dataset used for evaluating our AND architecture. For the script used to gather this dataset see <a href="https://github.com/sntcristian/and-kge/tree/main/src/AMiner-534K">https://github.com/sntcristian/and-kge/tree/main/src/AMiner-534K</a> and <a href="https://github.com/sntcristian/and-kge/tree/main/src/OC-782K">https://github.com/sntcristian/and-kge/tree/main/src/OC-782K</a>.</p>
Ensembles of knowledge graph embedding models improve predictions for drug discovery
<p>This contains data described in detail in our paper, "Ensembles of knowledge graph embedding models improve predictions for drug discovery". The metadata involves the different trained models that were used for prediction analysis as well as all the predictions from the trained models.</p>
Dynamic Knowledge Graphs for Continual Learning of Embeddings
<p>These datasets are generated from real world usecases. They are treated as Knowledge graphs and include 20 snapshots, where between two snapshots there are 10% added links and 10% deleted links, making the first and last snapshot non-overlapping.</p>
Linked Papers With Code: Knowledge Graph Embeddings
<p>We provide <strong>knowledge graph embeddings</strong> for the <strong>Linked Papers With Code</strong> Knowledge Graph. More information at <a href="https://linkedpaperswithcode.com/">https://linkedpaperswithcode.com/</a></p>
PheKnowLator Human Disease Knowledge Graph Benchmarks Embeddings -- v1.0.0
<p><strong>PKT Human Disease Knowledge Graph Benchmark Builds - Embeddings (v1.0.0)</strong></p><p><strong>Build Date: September 03, 2019</strong></p><blockquote><p>Please note that all resources linked below redirect to a publicly Google Cloud Storage bucket where all data are publicly accessible. Routing users from this wiki page is perfectly safe and allows us to avoid requiring users to have a Google account and login to download data. If you have any questions or concerns, please email the project maintainer at <a href="https://github.com/callahantiff/PheKnowLator/wiki/callahantiff@gmail.com">callahantiff@gmail.com</a>.</p></blockquote><p>The KG Benchmark Builds can also be downloaded from Zenodo:<br>👉 <strong>KGs:</strong> <a href="https://doi.org/10.5281/zenodo.7030200">https://doi.org/10.5281/zenodo.7030200</a><br>👉 <strong>Embeddings:</strong> <a href="https://zenodo.org/record/7030189">https://zenodo.org/record/7030189</a></p><p> </p><p>A <a href="https://github.com/xgfs/deepwalk-c">modified version</a> of the <a href="https://github.com/phanein/deepwalk">DeepWalk algorithm</a> was implemented to generate molecular mechanism embeddings from the biomedical knowledge graph. A t-SNE plot of the dimensionality reduced mechanism embeddings is shown in <a href="https://github.com/callahantiff/PheKnowLator/wiki/v1.0.0/figure-2-t-sne-plot-of-molecular-mechanisms">Figure</a>. For this release, the hyperparameters were set to 512 dimensions, 100 walks, walk length of 20, and a window of 10. Two types of KGs were embedded: (1) the full KG; and (2) the full KG with deductive closure using the OWL 2 EL reasoner, ELK via Protégé v5.1.1. ELK is able to classify instances and supports inferences over class hierarchies and object properties. inference over disjointness, intersection, and existential quantification (ontology class hierarchies).</p>
Improving the Utility and Trustworthiness of Knowledge Graph Embeddings with Calibration
<p>This repository contains two public knowledge graph datasets used in our paper <em>Improving the Utility of Knowledge Graph Embeddings with Calibration</em>. Each dataset is described below.</p> <p>Note that for our experiments we split each dataset randomly 5 times into 80/10/10 train/validation/test splits. We recommend that users of our data do the same to avoid (potentially) overfitting models to a single dataset split.</p> <p><strong>wikidata-authors</strong></p> <p>This dataset was extracted by querying the <a href="https://www.wikidata.org/wiki/Wikidata:Main_Page">Wikidata</a> API for facts about people categorized as "authors" or "writers" on Wikidata. Note that all head entities of triples are <em>people</em> (authors or writers), and all triples describe something about that person (e.g., their place of birth, their place of death, or their spouse). The knowledge graph has 23,887 entities, 13 relations, and 86,376 triples.</p> <p>The files are as follows:</p> <p><strong><code>entities.tsv</code></strong>: A tab-separated file of all unique entities in the dataset. The fields are as follows:</p> <ul> <li><code>eid</code>: The unique Wikidata identifier of this entity. You can find the corresponding Wikidata page at <code>https://www.wikidata.org/wiki/<eid></code>.</li> <li><code>label</code>: A human-readable label of this entity (extracted from Wikidata).</li> </ul> <p><strong><code>relations.tsv</code></strong>: A tab-separated file of all unique relations in the dataset. The fields are as follows:</p> <ul> <li><code>rid</code>: The unique Wikidata identifier of this relation. You can find the corresponding Wikidata page at <code>https://www.wikidata.org/wiki/Property:<rid></code>.</li> <li><code>label</code>: A human-readable label of this relation (extracted from Wikidata).</li> </ul> <p><strong><code>triples.tsv</code></strong>: A tab-separated file of all triples in the dataset, in the form of <code><head eid></code>, <code><relation rid></code>, <code><tail eid></code>.</p> <p><strong>fb15krr-linked</strong></p> <p>This dataset is an extended version of the FB15k+ dataset provided by <a href="https://github.com/thunlp/TKRL">[Xie et al IJCAI16]</a>. It has been linked to <a href="https://www.wikidata.org/wiki/Wikidata:Main_Page">Wikidata</a> using Freebase MIDs (machine IDs) as keys; we discarded triples from the original dataset that contained entities that could not be linked to Wikidata. We also removed reverse relations following the procedure described by <a href="https://www.aclweb.org/anthology/W15-4007.pdf">[Toutanova and Chen CVSC2015]</a>. Finally, we removed existing triples labeled as <em>False</em> and added predicted triples labeled as <em>True</em> based on the crowdsourced annotations we obtained in our <em>True or False Facts</em> experiment (see our paper for details). The knowledge graph consists of 14,289 entities, 770 relations, and 272,385 triples.</p> <p>The files are as follows:</p> <p><strong><code>entities.tsv</code></strong>: A tab-separated file of all unique entities in the dataset. The fields are as follows:</p> <ul> <li><code>mid</code>: The Freebase machine ID (MID) of this entity.</li> <li><code>wiki</code>: The corresponding unique Wikidata identifier of this entity. You can find the corresponding Wikidata page at <code>https://www.wikidata.org/wiki/<eid></code>.</li> <li><code>label</code>: A human-readable label of this entity (extracted from Wikidata).</li> <li><code>types</code>: All hierarchical types of this entity, as provided by <a href="https://github.com/thunlp/TKRL">[Xie et al IJCAI16]</a>.</li> </ul> <p><strong><code>relations.tsv</code></strong>: A tab-separated file of all unique relations in the dataset. The fields are as follows:</p> <ul> <li><code>label</code>: The hierarchical Freebase label of this relation.</li> </ul> <p><strong><code>triples.tsv</code></strong>: A tab-separated file of all triples in the dataset, in the form of <code><head MID></code>, <code><relation label></code>, <code><tail MID></code>.</p>
Supplemental Material for Paper 'Taxonomy Extraction Using Knowledge Graph Embeddings and Hierarchical Clustering'
<p>Contains input data and gold standard for the non-expressive extraction task, as well as examples of extracted taxonomies for both the non-expressive and expressive cases. Extracted taxonomies can also be found at <a href="http://labowest.ca/sdb2020/">labowest.ca</a>.</p>
Data for The "Effect of Semantic Knowledge Graph Richness on Embedding Based Recommender Systems"
Open the record for dataset details and reuse information.
Embeddings of KG-COVID-19 knowledge graph (Aug 12 build), produced using node2vec, skipgram model, p=q=1, walk length = 100, num walks = 20
<p>Embeddings of KG-COVID-19 knowledge graph (Aug 12 build), produced using Embiggen, node2vec, skipgram model, p=q=1, walk length = 100, num walks = 20</p>
Embeddings of KG-COVID-19 knowledge graph (Aug 12 build), 80/20 training/test split, produced using node2vec, skipgram model, p=q=1, walk length = 100, num walks = 20
<p>KG-COVID-19 embedding data from Sep 8, 2020 experiment, for training/test split of 80/20: </p> <p>These embeddings and weights were produced from this notebook on or around Sep 8, 2020:</p> <p>https://github.com/justaddcoffee/kg_covid_19_drug_analyses/blob/master/Graph%20embedding%20using%20SkipGram%20homogeneous%20graph.ipynb</p> <p>SkipGram_80_20_training_test_epoch_500_delta_0.0001_embedding.npy<br> SkipGram_80_20_training_test_epoch_500_delta_0.0001_weights.h5</p> <p>I'm also including two runs just before this, with different epoch number and delta values:</p> <p>SkipGram_80_20_training_test_embedding_sep_6_2020_epoch_200_delta_0.001.npy</p> <p>SkipGram_80_20_training_test_weights_sep_6_2020_epoch_200_delta_0.001.h5</p> <p>SkipGram_80_20_training_test_embedding_sep_7_2020_epoch_200_delta_0.0001.npy<br> SkipGram_80_20_training_test_weights_sep_7_2020_epoch_200_delta_0.0001.h5</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.