Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

219

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

219 results for “Knowledge Graph”

Learn how ShareScore rates datasets ↗
zenodo32/100

Data underlying the article: "A patient-centric knowledge graph approach to prioritize mutants for selective anti-cancer targeting"

<p>This repository contains the data and code underlying the article: "A patient-centric knowledge graph approach to prioritize mutants for selective anti-cancer targeting" available on BioRxiv.&nbsp;</p>

opencc-by-4.0May 2024View details →
zenodo32/100

Knowledge Graph Neural Network with Spatial-Aware Capsule for Drug-Drug Interaction Prediction

Open the record for dataset details and reuse information.

opencc-by-4.0Jun 2024View details →
zenodo32/100

Cytoscape session for the potato knowledge graph extracted with IBM Watson's supervised NLP model

<p>2 cytoscape session files (.cys) representing the genotypic-phenotypic knowledge networks retrieved from scientific literature using IBM Watson. Input to these&nbsp;files were the following:&nbsp;</p> <ul> <li>cytoscapeSession_trainingSet: A training set of 34 full-test&nbsp;articles about potato flesh color</li> <li>cytoscapesession_testSet: A testing set of a&nbsp;4023 abstracts from PubMed.</li> </ul> <p>&nbsp;</p>

opencc-by-4.0Jul 2019View details →
zenodo32/100

Examining Patented Artefact Knowledge Graphs to understand Linguistic and Structural Basis

<h1>Introduction</h1> <p>This resource is uploaded in support of our research that involves examining knowledge graphs of patented artefacts to understand the linguistic and structural basis of engineering design knowledge. The research is detailed in the following paper.</p> <p><a href="https://arxiv.org/abs/2312.06355">https://arxiv.org/abs/2312.06355</a></p> <p>The resource is segregated into multiple Pandas dataframes in pickled format &ndash; &ldquo;.pkl&rdquo;. To access any dataframe, please use the following Python code.</p> <pre><code>import pandas as pd data = pd.read_pickle("PATH TO FILE.pkl") print(data.head())</code></pre> <p>The individual datasets are described as follows.</p>

opencc-by-4.0Aug 2024View details →
zenodo32/100

The C4AI Knowledge Graph on Agricultural Prices (C4AI-KGAP)

<p>This version includes a file with SPARQL queries.</p>

opencc-by-nc-4.0Sep 2024View details →
zenodo32/100

KG-IDG knowledge graph

<p>KG-IDG knowledge graph - a knowledge graph that integrates data related to the Illuminating the Druggable Genome project, for drug repurposing</p>

opencc-by-4.0Nov 2023View details →
zenodo32/100

The World Literature Knowledge Graph

<p>If you use this resource please cite:&nbsp;</p> <p>&nbsp;</p> <pre>@inproceedings{stranisci2023world, title={The World Literature Knowledge Graph}, author={Stranisci, Marco Antonio and Bernasconi, Eleonora and Patti, Viviana and Ferilli, Stefano and Ceriani, Miguel and Damiano, Rossana}, booktitle={International Semantic Web Conference}, pages={435--452}, year={2023}, organization={Springer} }</pre>

opencc-by-4.0Sep 2024View details →
zenodo32/100

Linked Papers With Code: The Latest in Machine Learning as an RDF Knowledge Graph

<p><strong>Linked Papers With Code (LPWC)</strong> is an <strong>RDF knowledge graph </strong>that comprehensively models the research field of <strong>machine learning</strong>. It contains information about almost 400,000 machine learning <strong>publications</strong>, including the <strong>tasks</strong> addressed, the <strong>datasets</strong> utilized, the <strong>methods</strong> implemented, and the <strong>evaluations</strong> conducted, along with their <strong>results</strong>.&nbsp;The data set is based on <strong>Papers With Code</strong> and licensed under the CC BY-SA 4.0 license. Furthermore, we provide <strong>knowledge graph embeddings</strong> for entities and relations represented in LPWC.</p><p>More information can be found at <a href="https://linkedpaperswithcode.com/"><strong>https://linkedpaperswithcode.com/</strong></a> and in the ISWC'23 publication <a href="https://linkedpaperswithcode.com/"><strong>"</strong></a><a href="https://aifb.kit.edu/web/Inproceedings3993"><strong>Linked Papers With Code: The Latest in Machine Learning as an RDF Knowledge Graph".</strong></a></p>

opencc-by-4.0Oct 2023View details →
zenodo32/100

Interesting Scientific Idea Generation Using Knowledge Graphs and LLMs: Evaluations with 100 Research Group Leaders

<p>Dataset for the Knowledge graph used in the paper "<a href="https://arxiv.org/abs/2405.17044">Interesting Scientific Idea Generation Using Knowledge Graphs and LLMs: Evaluations with 100 Research Group Leaders</a>" by Xueme Gu and Mario Krenn. &nbsp;</p> <p>Nodes represent scientific concepts extracted from 2.44 million paper titles and abstracts and edges are formed when two concepts co-occur in titles or abstracts of over 58 million papers from OpenAlex, augmented with citation information.&nbsp;</p>

opencc-by-4.0Oct 2024View details →
zenodo32/100

E.A.T. Knowledge Graph Project

<p>The E.A.T. Knowledge Graph Project centers on the U.S. avant-garde movement Experiments in Art and Technology (E.A.T.). Data was sourced from documents from the Robert Rauschenberg Foundation Archives, a published bibliography by E.A.T. co-founder Billy Kl&uuml;ver, and edge-notched cards from the Getty Research Institute. Processed with custom tools, allowing for semantic encoding of statements through "painting triples&rdquo;, the data is stored in and accessible through a Wikibase knowledgebase. This approach aims to reduce barriers to creating knowledge graphs and expand reuse opportunities for research, applications, and data integration in cultural heritage contexts.</p> <p>The Semantic Lab Wikibase is available at <a href="https://base.semlab.io/wiki" target="_blank" rel="noopener">https://base.semlab.io/wiki </a>and on GitHub at&nbsp;<a href="https://github.com/SemanticLab/data-export" target="_blank" rel="noopener">https://github.com/SemanticLab/data-export.</a></p> <p>Learn more about the Semantic Lab at <a href="https://semlab.io/" target="_blank" rel="noopener">https://semlab.io/.&nbsp;</a></p>

opencc-by-4.0Nov 2024View details →
zenodo32/100

Submission to Special Issue of Quantitative Science Studies "Scientific Knowledge Graphs and Research Impact Assessment"

<p>To be added</p>

opencc-by-4.0Feb 2021View details →
zenodo32/100

A unified DTI prediction framework based on knowledge graph and recommendation system

<p>## A unified DTI prediction framework based on knowledge graph and recommendation system</p> <p>&nbsp;</p> <p># Code and data description</p> <p>## Scripts</p> <p>- `kge_nfm.py`: the complement of the KGE_NFM &amp; NFM methods.</p> <p>- `kge_rf.py`: the complement of the KGE_RF &amp; RF methods.</p> <p>- `deepdit.py`: the complement of the MPNN_CNN &amp; DeepDTI methods.</p> <p>- the complement of DTINet and DTiGEMS is tested based on their source packages (more in Prerequisites)</p> <p><br> &nbsp;</p> <p>## `data/` directory</p> <p>#### `yamanishi_08/` directory</p> <p>- `data_folds/`: 10 folds training set and test set in the three scenarios</p> <p>- `warm_start_1_1/`</p> <p>- `warm_start_1_10/`</p> <p>- `drug_coldstart/`</p> <p>- `protein_coldstart/`</p> <p>- `kg_data/`: supporting knowledge graph data</p> <p>- `dt_all_08.csv`: whole DTI dataset</p> <p>- `791drug_struc.csv`: drugbank id and smiles of drugs</p> <p>- `989proseq.csv`: kegg id and sequences of proteins</p> <p>- `morganfp.txt`: list of drug morgan fingerprints</p> <p>- `pro_ctd.txt`: list of protein descriptors</p> <p>&nbsp;</p> <p>#### `BioKG/` directory</p> <p>- `data_folds/`: 10 folds training set and test set in the three scenarios</p> <p>- `warm_start_1_10/`</p> <p>- `drug_coldstart/`</p> <p>- `protein_coldstart/`</p> <p>- `kg.csv`: supporting knowledge graph data</p> <p>- `dti.csv`: whole DTI dataset</p> <p>- `comp_struc.csv`: drugbank id and smiles of drugs</p> <p>- `pro_seq.csv`: sequences of proteins</p> <p>- `fp_df.csv`: list of drug morgan fingerprints</p> <p>- `prodes_df.csv`: list of protein descriptors</p> <p>&nbsp;</p> <p>#### `hetionet/` directory</p> <p>- `data_folds/`: 10 folds training set and test set in the three scenarios</p> <p>- `warm_start_1_10/`</p> <p>- `drug_coldstart/`</p> <p>- `protein_coldstart/`</p> <p>- `kg.csv`: supporting knowledge graph data</p> <p>- `dti.csv`: whole DTI dataset</p> <p>- `map_drugs_df`: drugbank id and smiles of drugs</p> <p>- `pro_seq.csv`: sequences of proteins</p> <p>- `fp_df.csv`: list of drug morgan fingerprints</p> <p>- `prodes_df.csv`: list of protein descriptors</p> <p>&nbsp;</p> <p>#### `luo&#39;s_dataset/` directory</p> <p>- `data_folds/`: 10 folds training set and test set in the three scenarios</p> <p>- `warm_start_1_1/`</p> <p>- `warm_start_1_10/`</p> <p>- `drug_coldstart/`</p> <p>- `protein_coldstart/`</p> <p>- `mapping/`: related mappings and similarity matrix (https://github.com/luoyunan/DTINet)</p> <p>- `protein.txt`: list of protein names</p> <p>- `disease.txt`: list of disease names</p> <p>- `se.txt`: list of side effect names</p> <p>- `drug_dict_map`: a complete ID mapping between drug names and DrugBank ID</p> <p>- `protein_dict_map`: a complete ID mapping between protein names and UniProt ID</p> <p>- `Similarity_Matrix_Drugs.txt` : Drug similarity scores based on chemical structures of drugs</p> <p>- `Similarity_Matrix_Proteins.txt` : Protein similarity scores based on primary sequences of proteins</p> <p>- `feature/`: related features used in methods</p> <p>- `drug_smiles.csv`: drugbank id and smiles</p> <p>- `seq.txt`: list of protein sequences</p> <p>- `morganfp.txt`: list of drug morgan fingerprints</p> <p>- `pro_ctd.txt`: list of protein descriptors</p> <p>&nbsp;</p> <p>#### `eg_model/` directory</p> <p>We provided a pre-trained kge model for example.</p> <p>- `dismult_400_warm_1_10.pkl`</p> <p><br> &nbsp;</p> <p># Prerequisites</p> <p>#### Operating system: Linux</p> <p>#### Programing language: python</p> <p>#### KGE_NFM &amp; NFM dependencies</p> <p>```</p> <p>- python 3.6</p> <p>- pandas &#39;1.1.5&#39;</p> <p>- numpy &#39;1.18.4&#39;</p> <p>- scikit-learn &#39;0.24.1&#39;</p> <p>- tensorflow &#39;1.15.0&#39;</p> <p>- ampligraph &#39;1.3.2&#39;</p> <p>- deepctr &#39;0.8.4&#39;</p> <p>```</p> <p>#### baseline dependencies</p> <p>- RF &amp; KGE_RF (included in KGE_NFM&amp;NFM dependencies)</p> <p>- MPNN_CNN &amp; DeepDTI:</p> <p>- source: https://github.com/kexinhuang12345/DeepPurpose</p> <p>```</p> <p>- deeppurpose &#39;0.0.9&#39;</p> <p>- torch &#39;1.6.0+cu101&#39;</p> <p>```</p> <p>- DTINet:</p> <p>- source: https://github.com/luoyunan/DTINet</p> <p>- note: in this work, we run the DTINet in a python environment, which need Linux system and python2. Importantly, this method requires the [Inductive Matrix Completion](http://bigdata.ices.utexas.edu/software/inductive-matrix-completion/) (IMC) library. More detailed information about the installation of this method could be found in the source code of the DTINet.</p> <p>- DTiGEMS:</p> <p>- source: https://github.com/MahaThafar/DTiGEMSplus</p> <p>- TriModel:</p> <p>- source: http://drugtargets.insight-centre.org/</p> <p><br> <br> &nbsp;</p> <p># Example (kge_nfm.py)</p> <p>&nbsp;</p> <p>#### A brief presentation of the results:</p> <p>- return average loss when training kge model</p> <p>```</p> <p>Average Loss: 0.475181: 2%|###3 | 1/50 [01:10&lt;57:31, 70.44s/epoch]</p> <p>```</p> <p>- return performance(mrr) on training set of DTI for early stopping (kge_model in `eg_model/`)</p> <p>```</p> <p>In [35]: roc = roc_auc(test_label,test_score)</p> <p>...: pr = pr_auc(test_label,test_score)</p> <p>...: print(roc)</p> <p>...: print(pr)</p> <p>0.8731770833333332</p> <p>0.44079654835037246</p> <p>```</p> <p>&nbsp;</p> <p>- nfm training process (`patience=10`)</p> <p>&nbsp;</p> <p>```</p> <p>In [45]: roc_nfm,pr_nfm,pred_y = train_nfm(feature_columns,train_model_input,train_label,test_model_input,test_label,patience)</p> <p>Train on 44851 samples</p> <p>Epoch 1/2000</p> <p>44851/44851 - 2s - loss: 0.5332 - precision: 0.0976</p> <p>Epoch 2/2000</p> <p>44851/44851 - 1s - loss: 0.4143 - precision: 0.0000e+00</p> <p>Epoch 3/2000</p> <p>44851/44851 - 1s - loss: 0.3456 - precision: 0.0000e+00</p> <p>Epoch 4/2000</p> <p>44851/44851 - 1s - loss: 0.3443 - precision: 0.0000e+00</p> <p>Epoch 5/2000</p> <p>44851/44851 - 1s - loss: 0.3470 - precision: 0.0000e+00</p> <p>Epoch 6/2000</p> <p>44851/44851 - 1s - loss: 0.3382 - precision: 0.0000e+00</p> <p>......</p> <p>Epoch 279/2000</p> <p>44851/44851 - 1s - loss: 0.0758 - precision: 0.9248</p> <p>Epoch 280/2000</p> <p>44851/44851 - 1s - loss: 0.0753 - precision: 0.9327</p> <p>Epoch 281/2000</p> <p>44851/44851 - 1s - loss: 0.0796 - precision: 0.9155</p> <p>Epoch 282/2000</p> <p>44851/44851 - 1s - loss: 0.0764 - precision: 0.9276</p> <p>Epoch 283/2000</p> <p>44851/44851 - 1s - loss: 0.0739 - precision: 0.9127</p> <p>```</p> <p>&nbsp;</p> <p>- reutrn results as type of roc_auc &amp; pr_auc</p> <p>```</p> <p>0.9812476679104477</p> <p>0.8803416284646345</p> <p>```</p>

opencc-by-4.0Jul 2021View details →
zenodo32/100

Comparison of Knowledge Graph Representations for Consumer Scenarios - Datasets

<p>These are the datasets used for the evaluations carried out in the submission &quot;Comparison of Knowledge Graph Representations for Consumer&nbsp;Scenarios&quot; to ISWC 2023</p>

opencc-by-4.0Apr 2023View details →
zenodo32/100

Relevant Datasets and Software Used for Paper "KGML-xDTD: A Knowledge Graph-based Machine Learning Framework for Drug Treatment Prediction and Mechanism Description"

<p>This repository contains relevant datasets and software&nbsp;used in a paper&nbsp;&quot;KGML-xDTD: A Knowledge Graph-based Machine Learning Framework for Drug Treatment Prediction and Mechanism Description&quot;. They are used to run the code of <em>KGML-xDTD&nbsp;</em>stored on <a href="https://github.com/chunyuma/KGML-xDTD">Github</a>&nbsp;and support the results of this paper.</p> <p><strong>About the datasets</strong></p> <p>1. <em>bkg_rtxkg2c_v2.7.3.tar.gz</em></p> <p>This tar.gz file contains three sub-folders: tsv_files, scripts, and relevant_dbs. The &quot;tsv_files&quot; sub-folder has the input files that the neo4j software uses. The &quot;scripts&quot; sub-folder contains a shell script with a relevant python script to construct&nbsp;the&nbsp;biomedical knowledge graph. The &quot;relevant_dbs&quot; sub-folder stores two auxiliary databases that <em>KGML-xDTD</em> needs to use.&nbsp;</p> <p>2. <em>indication_paths.yaml</em></p> <p>This file contains the <a href="https://sulab.github.io/DrugMechDB">DrugMechDB</a>&nbsp;MOA paths that we used to evaluate the predicted MOA paths by <em>KGML-xDTD.&nbsp;</em>It is downloaded from the official <a href="https://github.com/SuLab/DrugMechDB">GitHub repository</a> of DrugMechDB.</p> <p>3.&nbsp;<em>training_data.tar.gz</em></p> <p>This tar.gz file contains the processed training data of four data sources (e.g., <a href="https://mychem.info">MyChem</a>, <a href="https://lhncbc.nlm.nih.gov/ii/tools/SemRep_SemMedDB_SKR/SemMedDB_download.html">SemMedDB</a>, <a href="https://bioportal.bioontology.org/ontologies/NDFRT">NDF-RT</a>, <a href="https://unmtid-shinyapps.net/shiny/repodb/">RepoDB</a>) mentioned in the paper. These processed drug-disease pairs have been matched to the identifiers of biological entities used in our biomedical knowledge graph and respectively split into true positive (tp) sets and true negative (tn) sets. We also provide the names of these drug identifiers and disease identifiers under a sub-folder &quot;translated _to_name&quot;.</p> <p><strong>About the software</strong></p> <p><em>neo4j-community-3.5.26.tar.gz</em></p> <p>This tar.gz is the Neo4j community version 3.5.26 downloaded from <a href="https://neo4j.com/download-center/#community">Neo4j Download Center</a>. Although the&nbsp;newer versions are&nbsp;available, due to their big&nbsp;changes in the Neo4j setting that are not compatible with our scripts on Github, we provide the version that we used in our research. If you would like to use the newer version, modifications to our script will be required to import the biomedical knowledge graph into your local Neo4j database with the new setting.</p>

opencc-zeroJan 2023View details →
zenodo32/100

Deep learning and knowledge graph powered drug combination discovery against infectious diseases

<p>The Datasets and source codes for paper &quot;<strong>Deep learning and knowledge graph powered drug combination discovery against infectious diseases</strong>&quot;.</p>

opencc-by-4.0Mar 2023View details →
zenodo32/100

Mobile Feature-oriented Knowledge Base Generation Using Knowledge Graphs

<p>Sample study data set and derived metadata and evaluation files for the research titled &#39;Mobile Feature-oriented Knowledge Base Generation Using Knowledge Graphs&#39; (MApp-KG resource)</p>

openapgl-v3Apr 2023View details →
zenodo32/100

PheKnowLator Human Disease Knowledge Graph Benchmarks Embeddings -- v1.0.0

<p><strong>PKT Human Disease Knowledge Graph Benchmark Builds - Embeddings (v1.0.0)</strong></p><p><strong>Build Date:&nbsp;September 03, 2019</strong></p><blockquote><p>Please note that all resources linked below redirect to a publicly Google Cloud Storage bucket where all data are publicly accessible. Routing users from this wiki page is perfectly safe and allows us to avoid requiring users to have a Google account and login to download data. If you have any questions or concerns, please email the project maintainer at&nbsp;<a href="https://github.com/callahantiff/PheKnowLator/wiki/callahantiff@gmail.com">callahantiff@gmail.com</a>.</p></blockquote><p>The KG Benchmark Builds can also be downloaded from Zenodo:<br>👉&nbsp;<strong>KGs:</strong>&nbsp;<a href="https://doi.org/10.5281/zenodo.7030200">https://doi.org/10.5281/zenodo.7030200</a><br>👉&nbsp;<strong>Embeddings:</strong>&nbsp;<a href="https://zenodo.org/record/7030189">https://zenodo.org/record/7030189</a></p><p>&nbsp;</p><p>A&nbsp;<a href="https://github.com/xgfs/deepwalk-c">modified&nbsp;version</a> of the&nbsp;<a href="https://github.com/phanein/deepwalk">DeepWalk algorithm</a>&nbsp;was implemented to generate molecular mechanism embeddings from the biomedical knowledge graph. A t-SNE plot of the dimensionality reduced mechanism embeddings is shown in&nbsp;<a href="https://github.com/callahantiff/PheKnowLator/wiki/v1.0.0/figure-2-t-sne-plot-of-molecular-mechanisms">Figure</a>. For this release, the hyperparameters were set to 512 dimensions, 100 walks, walk length of 20, and a window of 10. Two types of KGs were embedded: (1) the full KG; and (2) the full KG&nbsp;with&nbsp;deductive closure using the OWL 2 EL reasoner, ELK via Protégé v5.1.1. ELK is able to classify instances and supports inferences over class hierarchies and object properties. inference over disjointness, intersection, and existential quantification (ontology class hierarchies).</p>

opencc-by-4.0Feb 2021View details →
zenodo32/100

Ideal Solutions for the Evaluation of A User-driven Hybrid Neuro-symbolic Approach for Knowledge Graph Creation from Relational Data

<p>Those two files represent one ideal solution of RML rules for the provided ontology and database by the study conductors.</p> <p>Note: in those two Turtle files&nbsp;joins are not necessary since templates could have been defined, too. This would lead to the same result.</p>

opencc-by-4.0Oct 2023View details →
zenodo28/100

Improving the Utility and Trustworthiness of Knowledge Graph Embeddings with Calibration

<p>This repository contains two public knowledge graph datasets used in our paper <em>Improving the Utility of Knowledge Graph Embeddings with Calibration</em>. Each dataset is described below.</p> <p>Note that for our experiments we split each dataset randomly 5 times into 80/10/10 train/validation/test splits. We recommend that users of our data do the same to avoid (potentially) overfitting models to a single dataset split.</p> <p><strong>wikidata-authors</strong></p> <p>This dataset was extracted by querying the <a href="https://www.wikidata.org/wiki/Wikidata:Main_Page">Wikidata</a> API for facts about people categorized as &quot;authors&quot; or &quot;writers&quot; on Wikidata. Note that all head entities of triples are <em>people</em> (authors or writers), and all triples describe something about that person (e.g., their place of birth, their place of death, or their spouse). The knowledge graph has 23,887 entities, 13 relations, and 86,376 triples.</p> <p>The files are as follows:</p> <p><strong><code>entities.tsv</code></strong>: A tab-separated file of all unique entities in the dataset. The fields are as follows:</p> <ul> <li><code>eid</code>: The unique Wikidata identifier of this entity. You can find the corresponding Wikidata page at <code>https://www.wikidata.org/wiki/&lt;eid&gt;</code>.</li> <li><code>label</code>: A human-readable label of this entity (extracted from Wikidata).</li> </ul> <p><strong><code>relations.tsv</code></strong>: A tab-separated file of all unique relations in the dataset. The fields are as follows:</p> <ul> <li><code>rid</code>: The unique Wikidata identifier of this relation. You can find the corresponding Wikidata page at <code>https://www.wikidata.org/wiki/Property:&lt;rid&gt;</code>.</li> <li><code>label</code>: A human-readable label of this relation (extracted from Wikidata).</li> </ul> <p><strong><code>triples.tsv</code></strong>: A tab-separated file of all triples in the dataset, in the form of <code>&lt;head eid&gt;</code>, <code>&lt;relation rid&gt;</code>, <code>&lt;tail eid&gt;</code>.</p> <p><strong>fb15krr-linked</strong></p> <p>This dataset is an extended version of the FB15k+ dataset provided by <a href="https://github.com/thunlp/TKRL">[Xie et al IJCAI16]</a>. It has been linked to <a href="https://www.wikidata.org/wiki/Wikidata:Main_Page">Wikidata</a> using Freebase MIDs (machine IDs) as keys; we discarded triples from the original dataset that contained entities that could not be linked to Wikidata. We also removed reverse relations following the procedure described by <a href="https://www.aclweb.org/anthology/W15-4007.pdf">[Toutanova and Chen CVSC2015]</a>. Finally, we removed existing triples labeled as <em>False</em> and added predicted triples labeled as <em>True</em> based on the crowdsourced annotations we obtained in our <em>True or False Facts</em> experiment (see our paper for details). The knowledge graph consists of 14,289 entities, 770 relations, and 272,385 triples.</p> <p>The files are as follows:</p> <p><strong><code>entities.tsv</code></strong>: A tab-separated file of all unique entities in the dataset. The fields are as follows:</p> <ul> <li><code>mid</code>: The Freebase machine ID (MID) of this entity.</li> <li><code>wiki</code>: The corresponding unique Wikidata identifier of this entity. You can find the corresponding Wikidata page at <code>https://www.wikidata.org/wiki/&lt;eid&gt;</code>.</li> <li><code>label</code>: A human-readable label of this entity (extracted from Wikidata).</li> <li><code>types</code>: All hierarchical types of this entity, as provided by <a href="https://github.com/thunlp/TKRL">[Xie et al IJCAI16]</a>.</li> </ul> <p><strong><code>relations.tsv</code></strong>: A tab-separated file of all unique relations in the dataset. The fields are as follows:</p> <ul> <li><code>label</code>: The hierarchical Freebase label of this relation.</li> </ul> <p><strong><code>triples.tsv</code></strong>: A tab-separated file of all triples in the dataset, in the form of <code>&lt;head MID&gt;</code>, <code>&lt;relation label&gt;</code>, <code>&lt;tail MID&gt;</code>.</p>

opencc-by-4.0Apr 2020View details →
zenodo28/100

Knowledge Graph Example, calculated PageRank and HITS

<p>The file contains the graph as json.</p> <p>Where do I add a repository? ;)</p> <p><a href="https://github.com/AntonioNoack/WebPageRank/commit/0be63ef218f32676ef74b1077e375470233f2b9b">https://github.com/AntonioNoack/WebPageRank/commit/0be63ef218f32676ef74b1077e375470233f2b9b</a>&nbsp;was my last commit, when I created this file. The project there was used to calculate the PageRank and HITS values.</p> <p>&nbsp;</p> <p>Normalization: None<br> PageRank random jump probability: 15%<br> PageRank preference vector: None</p>

opencc-by-4.0Jun 2020View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record