Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
219
datasets available to search
ShareScore release 0.7.1
Dataset results
219 results for “knowledge graph”
Evaluation Set - Contributions Similarity in the Open Research Knowledge Graph
<p>This evaluation set has been created for evaluating a content-based recommender system in the context of the Open Research Knowledge Graph (ORKG). The recommender system accepts structured ORKG contribution as input and recommends existing contributions in the ORKG semantically relevant to the given one.</p> <p> </p> <p>The evaluation set is manually annotated based on the <a href="https://www.orkg.org/orkg/featured-comparisons">featured comparisons</a> in the ORKG. In the course of this, it has been distinguished between homogeneous (those who are dissimilar in 2-3 properties) and heterogeneous (otherwise) instances. Multiple annotations have been obtained for the former and exactly one for the latter.</p> <p> </p> <p>It has been also distinguished between "with_response" and "without_response" instances (50 instances for each). The former are those contributions for them the initial version of the contributions similarity service has found similarities and the latter are the opposite case.</p> <p> </p> <p>This evaluation set has been created and applied on a modified version of the contributions similarity service in the context of <a href="https://doi.org/10.15488/11834">this master's thesis</a>. The modified version of the service has simplified the document representation of contributions that are stored in an ElasticSearch index by omitting redundant terms.</p> <p>The evaluation set has the following schema:</p> <pre><code class="language-json">{ "with_response": [ { "contribution_id": "some_id", "comparison_id": "some_id", "comparison_label": "some_label", "contribution_label": "some_label", "paper": "some_id", "research_field": "some_id", "research_problems": [ "some_id" ], "annotations": [ "some_id of a similar contribution", ... ] }, ... ], "without_response": [ ... ] }</code></pre> <p> </p>
TecKnoGraph: Knowledge Graph from patents in C4ISTAR
<p><img src="https://github.com/nicolamelluso/TecKnoGraph-demo/blob/main/TecKnoGraph-Example%20Graph.png" alt="TecKnoGraph"></p> <p>This dataset contains a sample of Knowledge Graph (KG) created with TecKnoGraph.</p> <p>There are two files:</p> <p><strong>- TecKnoGraph-C4ISTAR-sample.csv</strong>: this file contains the KG in the form of triples where each element of the triple (source, relation, target) is tagged with categories.</p> <p><strong>- patents.zip:</strong> this file contains data about 10,000 patents; each patent corresponds to a txt file.</p> <p>There is available a demo for using TecKnoGraph from examples of input text:<br> https://nicolamelluso-tecknograph-demo-tecknograph-streamlit-cnq4sm.streamlitapp.com/</p> <p>In this repository it is possible to find also the appendix of the corresponding paper.</p>
PheKnowLator Human Disease Knowledge Graphs - Build Data (Processed)
<p><strong>RELEASE V2.1.0 KNOWLEDGE GRAPH: PROCESSED DATA SOURCES </strong></p> <p><strong>Release:</strong> <a href="https://github.com/callahantiff/PheKnowLator/wiki/v2.0.0">v2.1.0 </a></p> <p>The goal of this build was to create a knowledge graph that represented human disease mechanisms and included the central dogma. The data sources utilized in this release include many of the sources used in the initial release, as well as some new data made available by the <a href="https://ctdbase.org/">Comparative Toxicogenomics Database</a> and experimental data from the <a href="https://www.proteinatlas.org/">Human Protein Atlas</a>.</p> <p>Data sources are listed by type (Ontology and Data not represented in an ontology [Database Sources]). Additional details are provided for each data source below. Please see documentation on the primary release (<a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources">https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources</a>) for additional details on each data source as well as citation information.</p> <p><strong>Data Access:</strong></p> <ul> <li><a href="https://console.cloud.google.com/storage/browser/pheknowlator/archived_builds/release_v2.1.0/build_01MAY2021?project=pheknowlator">https://console.cloud.google.com/storage/browser/pheknowlator/archived_builds/release_v2.1.0/build_01MAY2021</a></li> </ul> <p> </p> <p><strong>ONTOLOGIES</strong></p> <ul> <li>Cell Ontology</li> <li>Cell Line Ontology</li> <li>Chemical Entities of Biological Interest (ChEBI) Ontology</li> <li>Gene Ontology</li> <li>Human Phenotype Ontology</li> <li>Mondo Disease Ontology</li> <li>Pathway Ontology</li> <li>Protein Ontology</li> <li>Relations Ontology</li> <li>Sequence Ontology</li> <li>Uber-Anatomy Ontology</li> <li>Vaccine Ontology</li> </ul> <p> </p> <p><strong>Cell Ontology (CL)</strong></p> <p><strong>Homepage:</strong> <strong><a href="https://github.com/obophenotype/cell-ontology"><code>GitHub</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Bard J, Rhee SY, Ashburner M. <a href="https://genomebiology.biomedcentral.com/articles/10.1186/gb-2005-6-2-r21">An ontology for cell types</a>. Genome Biology. 2005;6(2):R21</p> </blockquote> <p><strong>Usage:</strong> Utilized to connect <code>transcripts</code> and <code>proteins</code> to <code>cells</code>. Additionally, the edges between this ontology and its dependencies are utilized:</p> <ul> <li><strong><a href="https://www.ebi.ac.uk/chebi/"><code>ChEBI</code></a></strong></li> <li><strong><a href="http://geneontology.org/"><code>GO</code></a></strong></li> <li><strong><a href="https://github.com/pato-ontology/pato/"><code>PATO</code></a></strong></li> <li><strong><a href="https://proconsortium.org/"><code>PRO</code></a></strong></li> <li><strong><a href="https://github.com/oborel/obo-relations/"><code>RO</code></a></strong></li> <li><strong><a href="https://uberon.github.io/"><code>UBERON</code></a></strong></li> </ul> <p> </p> <p><strong>Cell Line Ontology (CLO)</strong></p> <p><strong>Homepage:</strong> <strong><a href="http://www.clo-ontology.org/"><code>http://www.clo-ontology.org/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Sarntivijai S, Lin Y, Xiang Z, Meehan TF, Diehl AD, Vempati UD, Schürer SC, Pang C, Malone J, Parkinson H, Liu Y. <a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4387853/">CLO: the cell line ontology</a>. Journal of Biomedical Semantics. 2014;5(1):37</p> </blockquote> <p><strong>Usage:</strong> Utilized this ontology to map <code>cell lines</code> to <code>transcripts</code> and <code>proteins</code>. Additionally, the edges between this ontology and its dependencies are utilized:</p> <ul> <li><strong><a href="http://www.obofoundry.org/ontology/cl.html"><code>CL</code></a></strong></li> <li><strong><a href="http://disease-ontology.org/"><code>DOID</code></a></strong></li> <li><strong><a href="https://github.com/obophenotype/ncbitaxon"><code>NCBITaxon</code></a></strong></li> <li><strong><a href="https://uberon.github.io/"><code>UBERON</code></a></strong></li> </ul> <p> </p> <p><strong>Chemical Entities of Biological Interest (ChEBI)</strong></p> <p><strong>Homepage:</strong> <strong><a href="https://www.ebi.ac.uk/chebi/"><code>https://www.ebi.ac.uk/chebi/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Hastings J, Owen G, Dekker A, Ennis M, Kale N, Muthukrishnan V, Turner S, Swainston N, Mendes P, Steinbeck C. <a href="https://academic.oup.com/nar/article-abstract/44/D1/D1214/2502583">ChEBI in 2016: Improved services and an expanding collection of metabolites</a>. Nucleic Acids Research. 2015;44(D1):D1214-9</p> </blockquote> <p><strong>Usage:</strong> Utilized to connect <code>chemicals</code> to <code>complexes</code>, <code>diseases</code>, <code>genes</code>, <code>GO biological processes</code>, <code>GO cellular components</code>, <code>GO molecular functions</code>, <code>pathways</code>, <code>phenotypes</code>, <code>reactions</code>, and <code>transcripts</code>.</p> <p> </p> <p><strong>Gene Ontology (GO)</strong></p> <p><strong>Homepage:</strong> <strong><a href="http://geneontology.org/"><code>http://geneontology.org/</code></a></strong><br> <strong>Citations:</strong></p> <blockquote> <p>Ashburner M, Ball CA, Blake JA, Botstein D, Butler H, Cherry JM, Davis AP, Dolinski K, Dwight SS, Eppig JT, Harris MA. <a href="https://www.nature.com/articles/ng0500_25">Gene ontology: tool for the unification of biology</a>. Nature Genetics. 2000;25(1):25</p> <p>The Gene Ontology Consortium. <a href="https://academic.oup.com/nar/article/47/D1/D330/5160994">The Gene Ontology Resource: 20 years and still GOing strong</a>. Nucleic Acids Research. 2018;47(D1):D330-8</p> </blockquote> <p><strong>Usage:</strong> Utilized to connect <code>biological processes</code>, <code>cellular components</code>, and <code>molecular functions</code> to <code>chemicals</code>, <code>pathways</code>, and <code>proteins</code>. Additionally, the edges between this ontology and its dependencies are utilized:</p> <ul> <li><strong><a href="http://www.obofoundry.org/ontology/cl.html"><code>CL</code></a></strong></li> <li><strong><a href="https://github.com/obophenotype/ncbitaxon"><code>NCBITaxon</code></a></strong></li> <li><strong><a href="https://github.com/oborel/obo-relations/"><code>RO</code></a></strong></li> <li><strong><a href="https://uberon.github.io/"><code>UBERON</code></a></strong></li> </ul> <p><strong>Other Gene Ontology Data Used:</strong> <a href="http://geneontology.org/gene-associations/goa_human.gaf.gz"><code>goa_human.gaf.gz</code></a></p> <p> </p> <p><strong>Human Phenotype Ontology (HPO)</strong></p> <p><strong>Homepage:</strong> <strong><a href="https://hpo.jax.org/"><code>https://hpo.jax.org/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Köhler S, Carmody L, Vasilevsky N, Jacobsen JO, Danis D, Gourdine JP, Gargano M, Harris NL, Matentzoglu N, McMurry JA, Osumi-Sutherland D. <a href="https://academic.oup.com/nar/article-abstract/47/D1/D1018/5198478">Expansion of the Human Phenotype Ontology (HPO) knowledge base and resources</a>. Nucleic Acids Research. 2018;47(D1):D1018-27</p> </blockquote> <p><strong>Usage:</strong> Utilized to connect <code>phenotypes</code> to <code>chemicals</code>, <code>diseases</code>, <code>genes</code>, and <code>variants</code>. Additionally, the edges between this ontology and its dependencies are utilized:</p> <ul> <li><strong><a href="http://www.obofoundry.org/ontology/cl.html"><code>CL</code></a></strong></li> <li><strong><a href="https://www.ebi.ac.uk/chebi/"><code>ChEBI</code></a></strong></li> <li><strong><a href="http://geneontology.org/"><code>GO</code></a></strong></li> <li><strong><a href="https://uberon.github.io/"><code>UBERON</code></a></strong></li> </ul> <p><strong>Files</strong></p> <ul> <li>Other Human Phenotype Ontology Data Used: <a href="http://purl.obolibrary.org/obo/hp/hpoa/phenotype.hpoa"><code>phenotype.hpoa</code></a></li> </ul> <p> </p> <p><strong>Mondo Disease Ontology (Mondo)</strong></p> <p><strong>Homepage:</strong> <strong><a href="https://mondo.monarchinitiative.org/"><code>https://mondo.monarchinitiative.org/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Mungall CJ, McMurry JA, Köhler S, Balhoff JP, Borromeo C, Brush M, Carbon S, Conlin T, Dunn N, Engelstad M, Foster E. <a href="https://academic.oup.com/nar/article-abstract/45/D1/D712/2605791">The Monarch Initiative: an integrative data and analytic platform connecting phenotypes to genotypes across species</a>. Nucleic Acids Research. 2017;45(D1):D712-22</p> </blockquote> <p><strong>Usage:</strong> Utilized to connect <code>diseases</code> to <code>chemicals</code>, <code>phenotypes</code>, <code>genes</code>, and <code>variants</code>. Additionally, the edges between this ontology and its dependencies are utilized:</p> <ul> <li><strong><a href="http://www.obofoundry.org/ontology/cl.html"><code>CL</code></a></strong></li> <li><strong><a href="https://www.ncbi.nlm.nih.gov/taxonomy"><code>NCBITaxon</code></a></strong></li> <li><strong><a href="http://geneontology.org/"><code>GO</code></a></strong></li> <li><strong><a href="https://hpo.jax.org/"><code>HPO</code></a></strong></li> <li><strong><a href="https://uberon.github.io/"><code>UBERON</code></a></strong></li> </ul> <p> </p> <p><strong>Pathway Ontology (PW)</strong></p> <p><strong>Homepage:</strong> <strong><a href="https://rgd.mcw.edu/wg/home/pathway2/"><code>rgd.mcw.edu</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Petri V, Jayaraman P, Tutaj M, Hayman GT, Smith JR, De Pons J, Laulederkind SJ, Lowry TF, Nigam R, Wang SJ, Shimoyama M. <a href="https://www.ncbi.nlm.nih.gov/pubmed/24499703">The pathway ontology–updates and applications</a>. Journal of Biomedical Semantics. 2014;5(1):7.</p> </blockquote> <p><strong>Usage:</strong> Utilized to connect <code>pathways</code> to <code>GO biological processes</code>, <code>GO cellular components</code>, <code>GO molecular functions</code>, <code>Reactome pathways</code>. Several steps are taken in order to connect <code>Pathway Ontology</code> identifiers to <code>Reactome</code> pathways and <code>GO biological processes</code>. To connect <code>Pathway Ontology</code> identifiers to <code>Reactome</code> pathways, we use <a href="https://github.com/ComPath/resources/tree/master/mappings">ComPath Pathway Database Mappings</a> developed by Daniel Domingo-Fernández (<a href="https://www.ncbi.nlm.nih.gov/pubmed/30564458">PMID:30564458</a>).</p> <p><strong>Files</strong></p> <ul> <li>Downloaded Mapping Data <ul> <li><a href="http://compath.scai.fraunhofer.de/export_mappings"><code>curated_mappings.txt</code></a></li> <li><a href="https://github.com/ComPath/resources/blob/master/mappings/kegg_reactome.csv"><code>kegg_reactome.csv</code></a></li> </ul> </li> <li>Generated Mapping Data <ul> <li><a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/REACTOME_PW_GO_MAPPINGS.txt"><code>REACTOME_PW_GO_MAPPINGS.txt</code></a></li> </ul> </li> </ul> <p> </p> <p><strong>Protein Ontology (PRO)</strong></p> <p><strong>Homepage:</strong> <strong><a href="https://proconsortium.org/"><code>https://proconsortium.org/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Natale DA, Arighi CN, Barker WC, Blake JA, Bult CJ, Caudy M, Drabkin HJ, D’Eustachio P, Evsikov AV, Huang H, Nchoutmboube J. <a href="https://academic.oup.com/nar/article-abstract/39/suppl_1/D539/2508558">The Protein Ontology: a structured representation of protein forms and complexes</a>. Nucleic Acids Research. 2010;39(suppl_1):D539-45</p> </blockquote> <p><strong>Usage:</strong> Utilized to connect <code>proteins</code> to <code>chemicals</code>, <code>genes</code>, <code>anatomy</code>, <code>catalysts</code>, <code>cell lines</code>, <code>cofactors</code>, <code>complexes</code>, <code>GO biological processes</code>, <code>GO cellular components</code>, <code>GO molecular functions</code>, <code>pathways</code>, <code>proteins</code>, <code>reactions</code>, and <code>transcripts</code>. Additionally, the edges between this ontology and its dependencies are utilized:</p> <ul> <li><strong><a href="https://www.ebi.ac.uk/chebi/"><code>ChEBI</code></a></strong></li> <li><strong><a href="http://disease-ontology.org/"><code>DOID</code></a></strong></li> <li><strong><a href="http://geneontology.org/"><code>GO</code></a></strong></li> </ul> <p><strong>Notes:</strong> A partial, human-only version of this ontology was used. Details on how this version of the ontology was generated can be found under the Protein Ontology section of the <a href="https://github.com/callahantiff/PheKnowLator/blob/master/notebooks/Data_Preparation.ipynb"><code>Data_Preparation.ipynb</code></a> Jupyter Notebook.</p> <p><strong>Files</strong></p> <ul> <li> <p>Generated Human Version Protein Ontology (PRO)</p> <ul> <li><a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/human_pro.owl"><code>human_pro.owl</code></a> (closed with <a href="http://www.hermit-reasoner.com/">hermit reasoner</a>)</li> </ul> </li> <li> <p>Other PRO Data Used: <a href="https://proconsortium.org/download/current/promapping.txt"><code>promapping.txt</code></a></p> </li> <li> <p>Generated Mapping Data</p> <ul> <li>Merged Gene, RNA, Protein Map: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/Merged_gene_rna_protein_identifiers.pkl"><code>Merged_gene_rna_protein_identifiers.pkl</code></a></li> <li>Ensembl Transcript-PRO Identifier Mapping: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/ENTREZ_GENE_ENSEMBL_TRANSCRIPT_MAP.txt"><code>ENSEMBL_TRANSCRIPT_PROTEIN_ONTOLOGY_MAP.txt</code></a></li> <li>Entrez Gene-PRO Identifier Mapping: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/ENTREZ_GENE_PRO_ONTOLOGY_MAP.txt"><code>ENTREZ_GENE_PRO_ONTOLOGY_MAP.txt</code></a></li> <li>UniProt Accession-PRO Identifier Mapping: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/UNIPROT_ACCESSION_PRO_ONTOLOGY_MAP.txt"><code>UNIPROT_ACCESSION_PRO_ONTOLOGY_MAP.txt</code></a></li> <li>STRING-PRO Identifier Mapping: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/STRING_PRO_ONTOLOGY_MAP.txt"><code>STRING_PRO_ONTOLOGY_MAP.txt</code></a></li> </ul> </li> </ul> <p> </p> <p><strong>Relations Ontology (RO)</strong></p> <p><strong>Homepage:</strong> <strong><a href="https://github.com/oborel/obo-relations/"><code>GitHub</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Smith B, Ceusters W, Klagges B, Köhler J, Kumar A, Lomax J, Mungall C, Neuhaus F, Rector AL, Rosse C. <a href="https://genomebiology.biomedcentral.com/articles/10.1186/gb-2005-6-5-r46">Relations in biomedical ontologies</a>. Genome Biology. 2005;6(5):R46.</p> </blockquote> <p><strong>Usage:</strong> Utilizing this ontology to connect all data sources in knowledge graph. Additionally, the ontology is queried prior to building the knowledge graph to identify all relations, their inverse properties, and their labels.</p> <p><strong>Files</strong></p> <ul> <li>Generated RO Data <ul> <li><a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/INVERSE_RELATIONS.txt"><code>INVERSE_RELATIONS.txt</code></a></li> <li><a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/RELATIONS_LABELS.txt"><code>RELATIONS_LABELS.txt</code></a></li> </ul> </li> </ul> <p> </p> <p><strong>Sequence Ontology (SO)</strong></p> <p><strong>Homepage:</strong> <strong><a href="https://github.com/The-Sequence-Ontology/SO-Ontologies"><code>GitHub</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Eilbeck K, Lewis SE, Mungall CJ, Yandell M, Stein L, Durbin R, Ashburner M. <a href="https://link.springer.com/article/10.1186/gb-2005-6-5-r44">The Sequence Ontology: a tool for the unification of genome annotations</a>. Genome Biology. 2005;6(5):R44</p> </blockquote> <p><strong>Usage:</strong> Utilized to connect <code>transcripts</code> and other genomic material like <code>genes</code> and <code>variants</code>.</p> <p><strong>Files</strong></p> <ul> <li>Generated Mapping Data <ul> <li><a href="https://storage.googleapis.com/pheknowlator/curated_data/genomic_sequence_ontology_mappings.xlsx"><code>genomic_sequence_ontology_mappings.xlsx</code></a></li> <li><a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/SO_GENE_TRANSCRIPT_VARIANT_TYPE_MAPPING.txt"><code>SO_GENE_TRANSCRIPT_VARIANT_TYPE_MAPPING.txt</code></a></li> </ul> </li> </ul> <p> </p> <p><strong>Uber-Anatomy Ontology (Uberon)</strong></p> <p><strong>Homepage:</strong> <strong><a href="https://uberon.github.io/"><code>GitHub</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Mungall CJ, Torniai C, Gkoutos GV, Lewis SE, Haendel MA. <a href="https://genomebiology.biomedcentral.com/articles/10.1186/gb-2012-13-1-r5">Uberon, an integrative multi-species anatomy ontology</a>. Genome Biology. 2012;13(1):R5</p> </blockquote> <p><strong>Usage:</strong> Utilized to connect <code>tissues</code>, <code>fluids</code>, and <code>cells</code> to <code>proteins</code> and <code>transcripts</code>. Additionally, the edges between this ontology and its dependencies are utilized:</p> <ul> <li><strong><a href="https://www.ebi.ac.uk/chebi/"><code>ChEBI</code></a></strong></li> <li><strong><a href="http://www.obofoundry.org/ontology/cl.html"><code>CL</code></a></strong></li> <li><strong><a href="http://geneontology.org/"><code>GO</code></a></strong></li> <li><strong><a href="https://proconsortium.org/"><code>PRO</code></a></strong></li> </ul> <p> </p> <p><strong>Vaccine Ontology (VO)</strong></p> <p><strong>Homepage:</strong> <strong><a href="http://www.violinet.org/vaccineontology/"><code>http://www.violinet.org/vaccineontology/</code></a></strong><br> <strong>Citations:</strong></p> <blockquote> <p>He Y, Racz R, Sayers S, Lin Y, Todd T, Hur J, Li X, Patel M, Zhao B, Chung M, Ostrow J. <a href="https://academic.oup.com/nar/article-abstract/42/D1/D1124/1053128">Updates on the web-based VIOLIN vaccine database and analysis system</a>. Nucleic Acids Research. 2013;42(D1):D1124-32</p> <p>Xiang Z, Todd T, Ku KP, Kovacic BL, Larson CB, Chen F, Hodges AP, Tian Y, Olenzek EA, Zhao B, Colby LA. <a href="https://academic.oup.com/nar/article-abstract/36/suppl_1/D923/2505793">VIOLIN: vaccine investigation and online information network</a>. Nucleic Acids Research. 2007;36(suppl_1):D923-8</p> </blockquote> <p><strong>Usage:</strong> Utilized the edges between this ontology and its dependencies:</p> <ul> <li><strong><a href="https://www.ebi.ac.uk/chebi/"><code>ChEBI</code></a></strong></li> <li><strong><a href="http://disease-ontology.org/"><code>DOID</code></a></strong></li> <li><strong><a href="http://geneontology.org/"><code>GO</code></a></strong></li> <li><strong><a href="https://proconsortium.org/"><code>PRO</code></a></strong></li> <li><strong><a href="https://uberon.github.io/"><code>UBERON</code></a></strong></li> </ul> <p> </p> <p><strong>DATABASE SOURCES</strong></p> <ul> <li>BioPortal</li> <li>ClinVar</li> <li>Comparative Toxicogenomics Database</li> <li>DisGeNET</li> <li>Ensembl</li> <li>GeneMANIA</li> <li>Genotype-Tissue Expression Project</li> <li>Human Genome Organisation Gene Nomenclature Committee</li> <li>Human Protein Atlas</li> <li>National Center for Biotechnology Information Gene</li> <li>Reactome Pathway Database</li> <li>Search Tool for Recurring Instances of Neighbouring Genes Database</li> <li>Universal Protein Resource Knowledgebase</li> </ul> <p> </p> <p><strong>BioPortal</strong></p> <p><strong>Homepage:</strong> <strong><a href="https://bioportal.bioontology.org/"><code>BioPortal</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>BioPortal. <a href="https://www.bioontology.org/wiki/LOOM">Lexical OWL Ontology Matcher (LOOM)</a></p> <p>Ghazvinian A, Noy NF, Musen MA. <a href="https://www.ncbi.nlm.nih.gov/pmc/articles/pmc2815474/">Creating mappings for ontologies in biomedicine: simple methods work</a>. In AMIA Annual Symposium Proceedings 2009 (Vol. 2009, p. 198). American Medical Informatics Association</p> </blockquote> <p><strong>Usage:</strong> BioPortal was utilized to obtain mappings between <code>MeSH identifiers</code> and <code>ChEBI identifiers</code> for <code>chemicals-diseases</code>, <code>chemicals-genes</code>, <code>chemical-GO biological processes</code>, <code>chemicals-GO cellular components</code>, <code>chemicals-GO molecular functions</code>, <code>chemicals-phenotypes</code>, <code>chemicals-proteins</code>, and <code>chemicals-transcripts</code>. Additional information on how this data was processed can be obtained from the <a href="https://gist.github.com/callahantiff/a28fb3160782f42f104e9ec41553af0d"><code>NCBO_rest_api.py</code></a> GitHub Gist script.</p> <p>⭐ <strong>ALTERNATIVE METHOD</strong>⭐ Since the above approach can take over two days to process, we have developed an alternative solution that downloads the <a><code>mesh2021.nt</code></a> data file directly from MeSH and the <a><code>Flat_file_tab_delimited/names.tsv.gz</code></a> file directly from ChEBI. Using these files, we have recapitulated the <a href="https://www.bioontology.org/wiki/BioPortal_Mappings"><code>LOOM</code></a> algorithm implemented by BioPortal when creating mappings between these resources. The procedure is relatively straightforward and utilizes the following information from each resource:</p> <ul> <li>For all MeSH <code>SCR Chemicals</code>, obtain the following information: <ul> <li>Identifiers: MeSH identifiers</li> <li>Labels: string labels using the <code>RDFS:label</code> object property</li> <li>Synonyms: track down all synonyms using the <code>vocab:concept</code> and <code>vocab:preferredConcept</code> object properties</li> </ul> </li> <li>For all ChEBI classes, obtain the following information: <ul> <li>Labels: string labels using the <code>RDFS:label</code> object property</li> <li>Synonyms: track down all synonyms using all <code>synonym</code> object properties</li> </ul> </li> </ul> <p><strong>Files</strong></p> <ul> <li>Generated Data: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/MESH_CHEBI_MAP.txt"><code>MESH_CHEBI_MAP.txt</code></a></li> </ul> <p> </p> <p><strong>ClinVar</strong></p> <p><strong>Homepage:</strong> <strong><a href="https://www.ncbi.nlm.nih.gov/clinvar/"><code>https://www.ncbi.nlm.nih.gov/clinvar/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Landrum MJ, Lee JM, Benson M, Brown GR, Chao C, Chitipiralla S, Gu B, Hart J, Hoffman D, Jang W, Karapetyan K. <a href="https://academic.oup.com/nar/article-abstract/46/D1/D1062/4641904">ClinVar: improving access to variant interpretations and supporting evidence</a>. Nucleic Acids Research. 2017;46(D1):D1062-7</p> </blockquote> <p><strong>Usage:</strong> ClinVar was utilized to create <code>variant-gene</code>, <code>variant-disease</code>, and <code>variant-phenotype</code> edges. The original data is filtered such that only records meeting the following criteria were included:</p> <ul> <li> <p><code>Assembly</code> = "GRCh38"</p> </li> <li> <p><code>ClinSigSimple</code> = <code>1</code></p> <ul> <li> <blockquote> <p>1 = at least one current record submitted with an interpretation of Likely pathogenic or Pathogenic (independent of whether that record includes assertion criteria and evidence)"</p> </blockquote> </li> </ul> </li> <li> <p><code>ReviewStatus</code> in ["criteria provided, multiple submitters, no conflicts", "reviewed by expert panel", "practice guideline"]</p> </li> </ul> <p><strong>Files</strong></p> <ul> <li> <p>Downloaded Data</p> <ul> <li><a href="https://ftp.ncbi.nlm.nih.gov/pub/clinvar/tab_delimited/variant_summary.txt.gz"><code>variant_summary.txt.gz</code></a></li> <li><a href="https://ftp.ncbi.nlm.nih.gov/pub/clinvar/tab_delimited/var_citations.txt"><code>var_citations.txt</code></a></li> <li><a href="https://ftp.ncbi.nlm.nih.gov/pub/clinvar/tab_delimited/allele_gene.txt.gz"><code>allele_gene.txt.gz</code></a></li> </ul> </li> <li> <p>Generated Edge Data: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/CLINVAR_VARIANT_GENE_DISEASE_PHENOTYPE_EDGES.txt"><code>CLINVAR_VARIANT_GENE_DISEASE_PHENOTYPE_EDGES.txt</code></a></p> </li> </ul> <p> </p> <p><strong>Comparative Toxicogenomics Database (CTD)</strong></p> <p><strong>Homepage:</strong> <strong><a href="http://ctdbase.org/"><code>http://ctdbase.org/</code></a></strong><br> <strong>Citations:</strong></p> <blockquote> <p>Curated [chemical–gene interactions|chemical-go interactions|chemical–disease interactions|gene–pathway interactions] data were retrieved from the Comparative Toxicogenomics Database (CTD), MDI Biological Laboratory, Salisbury Cove, Maine, and NC State University, Raleigh, North Carolina. World Wide Web (URL: <a href="http://ctdbase.org/">http://ctdbase.org/</a>)</p> <p>Davis AP, Grondin CJ, Johnson RJ, Sciaky D, McMorran R, Wiegers J, Wiegers TC, Mattingly CJ. <a href="https://academic.oup.com/nar/article-abstract/47/D1/D948/5106145">The comparative toxicogenomics database: update 2019</a>. Nucleic Acids Research. 2018;47(D1):D948-54</p> </blockquote> <p>Usage: Comparative Toxicogenomics Database (CTD) was utilized to create <code>chemical-disease</code>, <code>chemical-gene</code>, <code>chemical-GO biological process</code>, <code>chemical-GO cellular components</code>, <code>chemical-GO molecular functions</code>, <code>chemical-phenotype</code>, <code>chemical-protein</code>, <code>chemical-rna</code>, and <code>gene-pathway</code> edges. The original data is filtered such that only records meeting the following criteria were included:</p> <ul> <li><code>chemical-disease</code>: <code>DirectEvidence</code> != ""</li> <li><code>chemical-gene</code>: <code>Organism</code> == "Homo sapiens", <code>GeneForms</code> == "gene", and affects not in <code>InteractionActions</code></li> <li><code>chemical-GO biological process</code>: <code>PhenotypeName</code> == "Biological Process" and <code>Interaction</code> <= "1.04e-47" (10th percentile)</li> <li><code>chemical-GO cellular components</code>: <code>PhenotypeName</code> == "Cellular Component" and <code>Interaction</code> <= "1.04e-47" (10th percentile)</li> <li><code>chemical-GO molecular functions</code>: <code>PhenotypeName</code> == "Molecular Function" and <code>Interaction</code> <= "1.04e-47" (10th percentile)</li> <li><code>chemical-phenotype</code>: <code>DirectEvidence</code> != ""</li> <li><code>chemical-protein</code>: <code>Organism</code> == "Homo sapiens", <code>GeneForms</code> == "protein", and affects not in <code>InteractionActions</code></li> <li><code>chemical-rna</code>: <code>Organism</code> == "Homo sapiens", <code>GeneForms</code> == "mRNA", and affects and activity not in <code>InteractionActions</code></li> <li><code>gene-pathway edges</code>: <code>PathwayName</code> == R-HSA-</li> </ul> <p><strong>Files</strong></p> <ul> <li>Downloaded Data <ul> <li>Chemical-Gene Relations: <a href="http://ctdbase.org/reports/CTD_chem_gene_ixns.tsv.gz"><code>CTD_chem_gene_ixns.tsv.gz</code></a></li> <li>Chemical-Disease/Phenotype Relations: <a href="http://ctdbase.org/reports/CTD_chemicals_diseases.tsv.gz"><code>CTD_chemicals_diseases.tsv.gz</code></a></li> <li>Chemical-GO Relations: <a href="http://ctdbase.org/reports/CTD_chem_go_enriched.tsv.gz"><code>CTD_chem_go_enriched.tsv.gz</code></a></li> <li>Gene-Pathway Relations: <a href="http://ctdbase.org/reports/CTD_genes_pathways.tsv.gz"><code>CTD_genes_pathways.tsv.gz</code></a></li> </ul> </li> </ul> <p> </p> <p><strong>DisGeNET</strong></p> <p><strong>Homepage:</strong> <strong><a href="https://www.disgenet.org/"><code>https://www.disgenet.org/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Gene-disease association data retrieved from DisGeNET v6.0 (<a href="http://www.disgenet.org/">http://www.disgenet.org/</a>), Integrative Biomedical Informatics Group GRIB/IMIM/UPF. [December, 2019].</p> <p>Piñero J, Ramírez-Anguita JM, Saüch-Pitarch J, Ronzano F, Centeno E, Sanz F, Furlong LI. <a href="https://academic.oup.com/nar/advance-article-abstract/doi/10.1093/nar/gkz1021/5611674">The DisGeNET knowledge platform for disease genomics: 2019 update</a>. Nucleic Acids Research. 2019.</p> </blockquote> <p><strong>Usage:</strong> DisGeNET was utilized to create <code>gene-disease</code>, and <code>gene-phenotype</code> edges. The original data is filtered such that only records meeting the following criteria were included: <code>EI</code> >= "1.0" (90th percentile). Additionally, data from this source was used to create mappings between different types of disease and phenotype identifiers, including:</p> <ul> <li>OMIM, ORPHA, UMLS, ICD ➞ DOID</li> <li>OMIM, ORPHA, UMLS, ICD ➞ HPO</li> </ul> <p><strong>Files</strong></p> <ul> <li> <p>Downloaded Data</p> <ul> <li>Disease/Phenotype-Gene Relations: <a href="https://www.disgenet.org/static/disgenet_ap1/files/downloads/curated_gene_disease_associations.tsv.gz"><code>curated_gene_disease_associations.tsv.gz</code></a></li> <li>Disease Identifier Mapping: <a href="https://www.disgenet.org/static/disgenet_ap1/files/downloads/disease_mappings.tsv.gz"><code>disease_mappings.tsv.gz</code></a></li> </ul> </li> <li> <p>Generated Mapping Data</p> <ul> <li>Disease Identifier Mapping: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/PHENOTYPE_HPO_MAP.txt"><code>PHENOTPYE_HPO_MAP.txt</code></a></li> <li>Phenotype Identifier Mapping: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/DISEASE_MONDO_MAP.txt"><code>DISEASE_DOID_MAP.txt</code></a></li> </ul> </li> </ul> <p> </p> <p><strong>Ensembl</strong></p> <p><strong>Homepage:</strong> <strong><a href="https://uswest.ensembl.org/index.html"><code>https://uswest.ensembl.org/index.html</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Zerbino DR, Achuthan P, Akanni W, Amode MR, Barrell D, Bhai J, Billis K, Cummins C, Gall A, Girón CG, Gil L. <a href="https://academic.oup.com/nar/article/46/D1/D754/4634002">Ensembl 2018</a>. Nucleic Acids Research. 2017;46(D1):D754-61</p> </blockquote> <p><strong>Usage:</strong> Ensembl data was utilized to create mappings between Ensembl genes, transcripts, and proteins with <a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#national-center-for-biotechnology-information-gene">NCBI Gene identifiers</a>, <a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#human-genome-organisation-gene-nomenclature-committee">HUGO gene symbols</a>, <a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#universal-protein-resource-knowledgebase">UniProt Accession identifiers</a>, and <a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#protein-ontology">Protein Ontology identifiers</a> in the knowledge graph (for additional details on the processing of these data, see <a href="https://github.com/callahantiff/PheKnowLator/blob/master/notebooks/Data_Preparation.ipynb"><code>Data_Preparation.ipynb</code></a>):</p> <ul> <li>Ensembl Transcript IDs ➞ PRO IDs</li> <li>Gene Ensembl IDs ➞ Entrez Gene IDs</li> <li>Gene Ensembl IDs ➞ PRO IDs</li> <li>Gene Symbols ➞ Transcript Ensembl IDs</li> <li>Entrez Gene IDs ➞ Transcript Ensembl IDs</li> <li>Entrez Gene IDs ➞ PRO IDs</li> <li>Protein Ensembl IDs ➞ UniProt Protein Accession</li> <li>STRING IDs ➞ PRO IDs</li> <li>UniProt Protein Accession ➞ Entrez Gene IDs</li> </ul> <p><strong>Files</strong></p> <ul> <li> <p>Downloaded Data</p> <ul> <li><a><code>Homo_sapiens.GRCh38.102.gtf</code></a></li> <li><a><code>Homo_sapiens.GRCh38.102.uniprot.tsv.gz</code></a></li> <li><a><code>Homo_sapiens.GRCh38.102.entrez.tsv.gz</code></a></li> </ul> </li> <li> <p>Generated Mapping Data</p> <ul> <li>Cleaned Ensembl Gene Set: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/ensembl_identifier_data_cleaned.txt"><code>ensembl_identifier_data_cleaned.txt</code></a></li> <li>Merged Gene, RNA, Protein Map: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/Merged_gene_rna_protein_identifiers.pkl"><code>Merged_gene_rna_protein_identifiers.pkl</code></a></li> <li>Ensembl Transcript-PRO Identifier Mapping: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/ENSEMBL_TRANSCRIPT_PROTEIN_ONTOLOGY_MAP.txt"><code>ENSEMBL_TRANSCRIPT_PROTEIN_ONTOLOGY_MAP.txt</code></a></li> <li>Gene Symbol-Ensembl Transcript Identifier Mapping: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/GENE_SYMBOL_ENSEMBL_TRANSCRIPT_MAP.txt"><code>GENE_SYMBOL_ENSEMBL_TRANSCRIPT_MAP.txt</code></a></li> <li>Entrez Gene-Ensembl Transcript Identifier Mapping: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/ENTREZ_GENE_ENSEMBL_TRANSCRIPT_MAP.txt"><code>ENTREZ_GENE_ENSEMBL_TRANSCRIPT_MAP.txt</code></a></li> <li>Entrez Gene-PRO Identifier Mapping: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/ENTREZ_GENE_PRO_ONTOLOGY_MAP.txt"><code>ENTREZ_GENE_PRO_ONTOLOGY_MAP.txt</code></a></li> <li>Ensembl Gene-Entrez Gene Identifier Mapping: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/ENSEMBL_GENE_ENTREZ_GENE_MAP.txt"><code>ENSEMBL_GENE_ENTREZ_GENE_MAP.txt</code></a></li> </ul> </li> </ul> <p> </p> <p><strong>GeneMANIA</strong></p> <p><strong>Homepage:</strong> <strong><a href="https://genemania.org/"><code>https://genemania.org/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Warde-Farley D, Donaldson SL, Comes O, Zuberi K, Badrawi R, Chao P, Franz M, Grouios C, Kazi F, Lopes CT, Maitland A. <a href="https://academic.oup.com/nar/article-abstract/38/suppl_2/W214/1126704">The GeneMANIA prediction server: biological network integration for gene prioritization and predicting gene function</a>. Nucleic Acids Research. 2010;38(suppl_2):W214-20</p> </blockquote> <p><strong>Usage:</strong> GeneMANIA was utilized to create <code>gene-gene</code> edges.</p> <p><strong>Files</strong></p> <ul> <li>Downloaded Data: <a href="http://genemania.org/data/current/Homo_sapiens.COMBINED/COMBINED.DEFAULT_NETWORKS.BP_COMBINING.txt"><code>COMBINED.DEFAULT_NETWORKS.BP_COMBINING.txt</code></a></li> </ul> <p> </p> <p><strong>Genotype-Tissue Expression Project (GTEx)</strong></p> <p><strong>Homepage:</strong> <strong><a href="https://gtexportal.org/home/"><code>https://gtexportal.org/home/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Lonsdale J, Thomas J, Salvatore M, Phillips R, Lo E, Shad S, Hasz R, Walters G, Garcia F, Young N, Foster B. <a href="http://www.nature.com/ng/journal/v45/n6/full/ng.2653.html">The genotype-tissue expression (GTEx) project</a>. Nature Genetics. 2013;45(6):580</p> </blockquote> <p><strong>Usage:</strong> The Genotype-Tissue Expression (GTEx) Project was utilized to create edges between <code>protein-cell</code>, <code>protein-anatomy</code>, <code>rna-cell</code> and <code>rna-anatomy</code> entities. The original data were filtered such that only those edges where the median TPM was >=<code>1.0</code> and genes were of any type other than protein-coding were included. It should also be noted that we chose to use the RNASeQC file over the RSEM file as advised by the GTEx website.</p> <blockquote> <p>The RSEM estimates are based on combining isoform-level estimates, which adds uncertainty to the resulting gene-level values (the isoform-level estimates are highly inaccurate in some cases).</p> </blockquote> <p>The file contains <code>54</code> unique tissue and/or cell types. GTEx provides mappings from tissue types to UBERON and EFO. These provided <a href="https://gtexportal.org/home/samplingSitePage">mappings</a> were verified and extended, such that all samples which referenced a cell type were also mapped to the Cell and the Cell Line ontologies. This resulted in a total of <code>56</code> mappings (<code>1.04</code> mappings/concepts).</p> <p><strong>Files</strong></p> <ul> <li> <p>Downloaded Data: <a href="https://storage.googleapis.com/gtex_analysis_v8/rna_seq_data/GTEx_Analysis_2017-06-05_v8_RNASeQCv1.1.9_gene_median_tpm.gct.gz"><code>GTEx_Analysis_2017-06-05_v8_RNASeQCv1.1.9_gene_median_tpm.gct</code></a></p> </li> <li> <p>Mapping Results: <a href="https://storage.googleapis.com/pheknowlator/curated_data/zooma_tissue_cell_mapping_04JAN2020.xlsx"><code>zooma_tissue_cell_mapping_04JAN2020.xlsx</code></a></p> </li> <li> <p>Generated Data<br> The final mapping set was combined with terms from the <a href="https://www.proteinatlas.org/">Human Protein Atlas</a>, see <a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources/t#human-protein-atlas">here</a> for more information.</p> <ul> <li>All HPA tissue and cell type strings: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/HPA_tissues.txt"><code>HPA_tissues.txt</code></a></li> <li>Final Term Mapping: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/HPA_GTEx_TISSUE_CELL_MAP.txt"><code>HPA_GTEx_TISSUE_CELL_MAP.txt</code></a></li> <li>Final RNA, Gene, Protein-Tissues and Cell Types Relations: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/HPA_GTEX_RNA_GENE_PROTEIN_EDGES.txt"><code>HPA_GTEX_RNA_GENE_PROTEIN_EDGES.txt</code></a></li> </ul> </li> </ul> <p> </p> <p><strong>Human Genome Organisation Gene Nomenclature Committee (HUGO)</strong></p> <p><strong>Homepage:</strong> <strong><a href="https://www.genenames.org/"><code>https://www.genenames.org/</code></a></strong><br> <strong>Citations:</strong></p> <blockquote> <p>HGNC Database, HUGO Gene Nomenclature Committee (HGNC), European Molecular Biology Laboratory, European Bioinformatics Institute (EMBL-EBI), Wellcome Genome Campus, Hinxton, Cambridge CB10 1SD, United Kingdom <a href="https://www.genenames.org/">www.genenames.org</a></p> <p>Yates B, Braschi B, Gray K, Seal R, Tweedie S, Bruford E. <a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5210531/">Genenames.org: the HGNC and VGNC Resources in 2017</a>. Nucleic Acids Research. 2017;45(D1):D619-625</p> </blockquote> <p><strong>Usage:</strong> The Human Genome Organisation (HUGO) data was utilized to obtain mappings between <a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#national-center-for-biotechnology-information-gene">NCBI Gene identifiers</a>, HUGO gene symbols, <a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#universal-protein-resource-knowledgebase">UniProt Accession identifiers</a>, and <a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#protein-ontology">Protein Ontology identifiers</a>. For additional details on the processing of these data, see <a href="https://github.com/callahantiff/PheKnowLator/blob/master/notebooks/Data_Preparation.ipynb"><code>Data_Preparation.ipynb</code></a>:</p> <ul> <li>Ensembl Transcript IDs ➞ PRO IDs</li> <li>Gene Ensembl IDs ➞ Entrez Gene IDs</li> <li>Gene Ensembl IDs ➞ PRO IDs</li> <li>Gene Symbols ➞ Transcript Ensembl IDs</li> <li>Entrez Gene IDs ➞ Transcript Ensembl IDs</li> <li>Entrez Gene IDs ➞ PRO IDs</li> <li>Protein Ensembl IDs ➞ UniProt Protein Accession</li> <li>STRING IDs ➞ PRO IDs</li> <li>UniProt Protein Accession ➞ Entrez Gene IDs</li> </ul> <p><strong>Files</strong></p> <ul> <li> <p>Downloaded Data: <a href="http://ftp.ebi.ac.uk/pub/databases/genenames/hgnc/tsv/hgnc_complete_set.txt"><code>hgnc_complete_set.txt</code></a></p> </li> <li> <p>Generated Data</p> <ul> <li>Merged Gene, RNA, Protein Map: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/Merged_gene_rna_protein_identifiers.pkl"><code>Merged_gene_rna_protein_identifiers.pkl</code></a></li> <li>Gene Symbol-Ensembl Transcript Identifier Mapping: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/GENE_SYMBOL_ENSEMBL_TRANSCRIPT_MAP.txt"><code>GENE_SYMBOL_ENSEMBL_TRANSCRIPT_MAP.txt</code></a></li> </ul> </li> </ul> <p> </p> <p><strong>Human Protein Atlas (HPA)</strong></p> <p><strong>Homepage:</strong> <strong><a href="https://www.proteinatlas.org/"><code>https://www.proteinatlas.org/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Uhlén M, Fagerberg L, Hallström BM, Lindskog C, Oksvold P, Mardinoglu A, Sivertsson Å, Kampf C, Sjöstedt E, Asplund A, Olsson I. <a href="https://science.sciencemag.org/content/347/6220/1260419.short">Tissue-based map of the human proteome</a>. Science. 2015;347(6220):1260419</p> </blockquote> <p><strong>Usage:</strong> The Human Protein Atlas (HPA) was utilized to create <code>rna-cell</code>, <code>rna-anatomy</code>, <code>protein-cell</code>, and <code>protein-anatomy</code> edges. Evidence between gene and RNA expression in specific tissue types was derived by HPA, such that the <a href="https://www.proteinatlas.org/about/assays+annotation#normalization_rna">consensus normalized expression</a> was >=<code>1.0</code>. Zooma was utilized to automatically annotate the <code>153</code> unique tissues and cell types from Human Protein Atlas for all human protein-coding genes in the <a href="https://www.proteinatlas.org/humanproteome">Human Proteome</a> to the Cell Ontology, Cell Line Ontology, and the Uber-Anatomy Ontology. To best represent each concept, the automatic mappings from Zooma were extend through manual mapping efforts to ensure each concept cell type was matched to a Cell Ontology, Cell Line Ontology, and UBERON ontology term. This resulted in a total of <code>281</code> mappings (<code>1.84</code> mappings/concepts).</p> <p><strong>Files</strong></p> <ul> <li> <p>Downloaded Data: <a href="https://www.proteinatlas.org/api/search_download.php?search=&columns=g,eg,up,pe,rnatsm,rnaclsm,rnacasm,rnabrsm,rnabcsm,rnablsm,scl,t_RNA_adipose_tissue,t_RNA_adrenal_gland,t_RNA_amygdala,t_RNA_appendix,t_RNA_basal_ganglia,t_RNA_bone_marrow,t_RNA_breast,t_RNA_cerebellum,t_RNA_cerebral_cortex,t_RNA_cervix,_uterine,t_RNA_colon,t_RNA_corpus_callosum,t_RNA_ductus_deferens,t_RNA_duodenum,t_RNA_endometrium_1,t_RNA_epididymis,t_RNA_esophagus,t_RNA_fallopian_tube,t_RNA_gallbladder,t_RNA_heart_muscle,t_RNA_hippocampal_formation,t_RNA_hypothalamus,t_RNA_kidney,t_RNA_liver,t_RNA_lung,t_RNA_lymph_node,t_RNA_midbrain,t_RNA_olfactory_region,t_RNA_ovary,t_RNA_pancreas,t_RNA_parathyroid_gland,t_RNA_pituitary_gland,t_RNA_placenta,t_RNA_pons_and_medulla,t_RNA_prostate,t_RNA_rectum,t_RNA_retina,t_RNA_salivary_gland,t_RNA_seminal_vesicle,t_RNA_skeletal_muscle,t_RNA_skin_1,t_RNA_small_intestine,t_RNA_smooth_muscle,t_RNA_spinal_cord,t_RNA_spleen,t_RNA_stomach_1,t_RNA_testis,t_RNA_thalamus,t_RNA_thymus,t_RNA_thyroid_gland,t_RNA_tongue,t_RNA_tonsil,t_RNA_urinary_bladder,t_RNA_vagina,t_RNA_B-cells,t_RNA_dendritic_cells,t_RNA_granulocytes,t_RNA_monocytes,t_RNA_NK-cells,t_RNA_T-cells,t_RNA_total_PBMC,cell_RNA_A-431,cell_RNA_A549,cell_RNA_AF22,cell_RNA_AN3-CA,cell_RNA_ASC_diff,cell_RNA_ASC_TERT1,cell_RNA_BEWO,cell_RNA_BJ,cell_RNA_BJ_hTERT+,cell_RNA_BJ_hTERT+_SV40_Large_T+,cell_RNA_BJ_hTERT+_SV40_Large_T+_RasG12V,cell_RNA_CACO-2,cell_RNA_CAPAN-2,cell_RNA_Daudi,cell_RNA_EFO-21,cell_RNA_fHDF/TERT166,cell_RNA_HaCaT,cell_RNA_HAP1,cell_RNA_HBEC3-KT,cell_RNA_HBF_TERT88,cell_RNA_HDLM-2,cell_RNA_HEK_293,cell_RNA_HEL,cell_RNA_HeLa,cell_RNA_Hep_G2,cell_RNA_HHSteC,cell_RNA_HL-60,cell_RNA_HMC-1,cell_RNA_HSkMC,cell_RNA_hTCEpi,cell_RNA_hTEC/SVTERT24-B,cell_RNA_hTERT-HME1,cell_RNA_HUVEC_TERT2,cell_RNA_K-562,cell_RNA_Karpas-707,cell_RNA_LHCN-M2,cell_RNA_MCF7,cell_RNA_MOLT-4,cell_RNA_NB-4,cell_RNA_NTERA-2,cell_RNA_PC-3,cell_RNA_REH,cell_RNA_RH-30,cell_RNA_RPMI-8226,cell_RNA_RPTEC_TERT1,cell_RNA_RT4,cell_RNA_SCLC-21H,cell_RNA_SH-SY5Y,cell_RNA_SiHa,cell_RNA_SK-BR-3,cell_RNA_SK-MEL-30,cell_RNA_T-47d,cell_RNA_THP-1,cell_RNA_TIME,cell_RNA_U-138_MG,cell_RNA_U-2_OS,cell_RNA_U-2197,cell_RNA_U-251_MG,cell_RNA_U-266/70,cell_RNA_U-266/84,cell_RNA_U-698,cell_RNA_U-87_MG,cell_RNA_U-937,cell_RNA_WM-115,blood_RNA_basophil,blood_RNA_classical_monocyte,blood_RNA_eosinophil,blood_RNA_gdT-cell,blood_RNA_intermediate_monocyte,blood_RNA_MAIT_T-cell,blood_RNA_memory_B-cell,blood_RNA_memory_CD4_T-cell,blood_RNA_memory_CD8_T-cell,blood_RNA_myeloid_DC,blood_RNA_naive_B-cell,blood_RNA_naive_CD4_T-cell,blood_RNA_naive_CD8_T-cell,blood_RNA_neutrophil,blood_RNA_NK-cell,blood_RNA_non-classical_monocyte,blood_RNA_plasmacytoid_DC,blood_RNA_T-reg,blood_RNA_total_PBMC,brain_RNA_amygdala,brain_RNA_basal_ganglia,brain_RNA_cerebellum,brain_RNA_cerebral_cortex,brain_RNA_hippocampal_formation,brain_RNA_hypothalamus,brain_RNA_midbrain,brain_RNA_olfactory_region,brain_RNA_pons_and_medulla,brain_RNA_thalamus&format=tsv"><code>proteinatlas_search.tsv</code></a></p> </li> <li> <p>Mapping Results: <a href="https://storage.googleapis.com/pheknowlator/curated_data/zooma_tissue_cell_mapping_04JAN2020.xlsx"><code>zooma_tissue_cell_mapping_04JAN2020.xlsx</code></a></p> </li> <li> <p>Generated Data</p> <ul> <li>Final Term Mapping: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/HPA_GTEx_TISSUE_CELL_MAP.txt"><code>HPA_GTEx_TISSUE_CELL_MAP.txt</code></a></li> <li>Final RNA, Gene, Protein-Tissues and Cell Types Relations: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/HPA_GTEX_RNA_GENE_PROTEIN_EDGES.txt"><code>HPA_GTEX_RNA_GENE_PROTEIN_EDGES.txt</code></a></li> </ul> </li> </ul> <p> </p> <p><strong>National Center for Biotechnology Information (NCBI) Entrez Gene</strong></p> <p><strong>Homepage:</strong> <strong><a href="https://www.ncbi.nlm.nih.gov/gene/"><code>https://www.ncbi.nlm.nih.gov/gene/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Maglott D, Ostell J, Pruitt KD, Tatusova T. <a href="https://academic.oup.com/nar/article-abstract/33/suppl_1/D54/2505255">Entrez Gene: gene-centered information at NCBI</a>. Nucleic Acids Research. 2005;33(suppl_1):D54-8.</p> </blockquote> <p><strong>Usage:</strong> The National Center for Biotechnology Information (NCBI) Gene data was utilized to obtain mappings between <a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#ncbi-gene">NCBI Gene identifiers</a>, <a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#hugo-gene-nomenclature-committee">HUGO gene symbols</a>, <a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#uniprot-knowledgebase">UniProt Accession identifiers</a>, and <a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#protein-ontology">Protein Ontology identifiers</a>. For additional details on the processing of these data, see <a href="https://github.com/callahantiff/PheKnowLator/blob/master/notebooks/Data_Preparation.ipynb"><code>Data_Preparation.ipynb</code></a>:</p> <ul> <li>Ensembl Transcript IDs ➞ PRO IDs</li> <li>Gene Ensembl IDs ➞ Entrez Gene IDs</li> <li>Gene Ensembl IDs ➞ PRO IDs</li> <li>Gene Symbols ➞ Transcript Ensembl IDs</li> <li>Entrez Gene IDs ➞ Transcript Ensembl IDs</li> <li>Entrez Gene IDs ➞ PRO IDs</li> <li>Protein Ensembl IDs ➞ UniProt Protein Accession</li> <li>STRING IDs ➞ PRO IDs</li> <li>UniProt Protein Accession ➞ Entrez Gene IDs</li> </ul> <p><strong>Files</strong></p> <ul> <li> <p>Downloaded Data: <a href="https://ftp.ncbi.nih.gov/gene/DATA/GENE_INFO/Mammalia/Homo_sapiens.gene_info.gz"><code>Homo_sapiens.gene_info.gz</code></a></p> </li> <li> <p>Generated Data</p> <ul> <li>Merged Gene, RNA, Protein Map: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/Merged_gene_rna_protein_identifiers.pkl"><code>Merged_gene_rna_protein_identifiers.pkl</code></a></li> <li>Entrez Gene-Ensembl Transcript Identifier Mapping: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/ENTREZ_GENE_ENSEMBL_TRANSCRIPT_MAP.txt"><code>ENTREZ_GENE_ENSEMBL_TRANSCRIPT_MAP.txt</code></a></li> <li>Entrez Gene-PRO Identifier Mapping: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/ENTREZ_GENE_PRO_ONTOLOGY_MAP.txt"><code>ENTREZ_GENE_PRO_ONTOLOGY_MAP.txt</code></a></li> <li>Ensembl Gene-Entrez Gene Identifier Mapping: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/ENSEMBL_GENE_ENTREZ_GENE_MAP.txt"><code>ENSEMBL_GENE_ENTREZ_GENE_MAP.txt</code></a></li> <li>Uniprot Accession-Entrez Gene Identifier Mapping: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/UNIPROT_ACCESSION_ENTREZ_GENE_MAP.txt"><code>UNIPROT_ACCESSION_ENTREZ_GENE_MAP.txt</code></a></li> </ul> </li> </ul> <p> </p> <p><strong>Reactome Pathway Database</strong></p> <p><strong>Homepage:</strong> <strong><a href="https://reactome.org/"><code>https://reactome.org/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Fabregat A, Jupe S, Matthews L, Sidiropoulos K, Gillespie M, Garapati P, Haw R, Jassal B, Korninger F, May B, Milacic M. <a href="https://academic.oup.com/nar/article-abstract/46/D1/D649/4626770">The reactome pathway knowledgebase</a>. Nucleic Acids Research. 2017;46(D1):D649-55</p> </blockquote> <p><strong>Usage:</strong> The Reactome Database was utilized to create <code>chemical-pathway</code>, <code>GO Biological process-pathway</code>, <code>pathway-GO Cellular component</code>, <code>GO Molecular function-pathway</code>, and <code>protein-pathway</code> edges. The original data is filtered such that only records meeting the following criteria were included:</p> <ul> <li><code>chemical-pathway</code>: column[5] == "Homo sapiens"</li> <li><code>GO Biological process-pathway</code>: column[5] startswith "REACTOME", column[8] == "P", and column[12] == "taxon:9606"</li> <li><code>pathway-GO Cellular component</code>: column[5] startswith "REACTOME", column[8] == "C", and column[12] == "taxon:9606"</li> <li><code>GO Molecular function-pathway</code>: column[5] startswith "REACTOME", column[8] == "F", and column[12] == "taxon:9606"</li> <li><code>protein-pathway</code>: column[5] == "Homo sapiens"</li> </ul> <p><strong>Files</strong></p> <ul> <li>Downloaded Data <ul> <li>Chemical-Pathway Relations: <a href="https://reactome.org/download/current/ChEBI2Reactome_All_Levels.txt"><code>ChEBI2Reactome_All_Levels.txt</code></a></li> <li>Pathway-GO Relations: <a href="https://reactome.org/download/current/gene_association.reactome"><code>gene_association.reactome</code></a></li> <li>Protein-Pathway Relations: <a href="https://reactome.org/download/current/UniProt2Reactome_All_Levels.txt"><code>UniProt2Reactome_All_Levels.txt</code></a></li> </ul> </li> </ul> <p> </p> <p><strong>Search Tool for Recurring Instances of Neighbouring Genes (STRING) Database</strong></p> <p><strong>Homepage:</strong> <strong><a href="https://string-db.org/"><code>string-db.org</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Szklarczyk D, Gable AL, Lyon D, Junge A, Wyder S, Huerta-Cepas J, Simonovic M, Doncheva NT, Morris JH, Bork P, Jensen LJ. <a href="https://academic.oup.com/nar/article-abstract/47/D1/D607/5198476">STRING v11: protein–protein association networks with increased coverage, supporting functional discovery in genome-wide experimental datasets</a>. Nucleic Acids Research. 2018;47(D1):D607-13</p> </blockquote> <p><strong>Usage:</strong> The Search Tool for Recurring Instances of Neighbouring Genes (STRING) Database was utilized to create <code>protein-protein</code> edges. The original data is filtered such that only records meeting the following criteria were included: <code>combined_score</code> >= "700" (>90th percentile).</p> <p><strong>Files</strong></p> <ul> <li> <p>Downloaded Data: <a href="https://stringdb-static.org/download/protein.links.v11.0/9606.protein.links.v11.0.txt.gz"><code>9606.protein.links.v11.0.txt.gz</code></a></p> </li> <li> <p>Generated Data: STRING-PRO Identifier Mapping: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/STRING_PRO_ONTOLOGY_MAP.txt"><code>STRING_PRO_ONTOLOGY_MAP.txt</code></a></p> </li> </ul> <p> </p> <p><strong>Universal Protein Resource (UniProt) Knowledgebase</strong></p> <p><strong>Homepage:</strong> <strong><a href="https://www.uniprot.org/"><code>https://www.uniprot.org/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>UniProt Consortium. <a href="https://academic.oup.com/nar/article-abstract/47/D1/D506/5160987">UniProt: a worldwide hub of protein knowledge</a>. Nucleic acids research. 2018;47(D1):D506-15</p> </blockquote> <p><strong>Usage:</strong> The Universal Protein Resource (UniProt) Knowledgebase was utilized to obtain <code>cofactor</code>/<code>catalyst</code>-<code>protein</code> and <code>protein-coding gene</code>-<code>protein</code> edges as well as mappings between <a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#national-center-for-biotechnology-information-gene">NCBI Gene identifiers</a>, <a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#human-genome-organisation-gene-nomenclature-committee">HUGO gene symbols</a>, <a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#universal-protein-resource-knowledgebase">Universal Protein Resource (UniProt) Accession identifiers</a>, and <a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#protein-ontology">Protein Ontology identifiers</a>. For additional details on the processing of these data, see <a href="https://github.com/callahantiff/PheKnowLator/blob/master/notebooks/Data_Preparation.ipynb"><code>Data_Preparation.ipynb</code></a>:</p> <ul> <li>Ensembl Transcript IDs ➞ PRO IDs</li> <li>Gene Ensembl IDs ➞ Entrez Gene IDs</li> <li>Gene Ensembl IDs ➞ PRO IDs</li> <li>Gene Symbols ➞ Transcript Ensembl IDs</li> <li>Entrez Gene IDs ➞ Transcript Ensembl IDs</li> <li>Entrez Gene IDs ➞ PRO IDs</li> <li>Protein Ensembl IDs ➞ UniProt Protein Accession</li> <li>STRING IDs ➞ PRO IDs</li> <li>UniProt Protein Accession ➞ Entrez Gene IDs</li> </ul> <p><strong>Files</strong></p> <ul> <li> <p>Downloaded Data</p> <ul> <li>Cofactor and Catalyst relations: <a href="https://www.uniprot.org/uniprot/?query=&fil=organism%3A%22Homo%20sapiens%20(Human)%20%5B9606%5D%22&columns=id%2Centry%20name%2Creviewed%2Cdatabase(PRO)%2Cchebi(Cofactor)%2Cchebi(Catalytic%20activity)"><code>Cofactor/Catalyst Query Results</code></a></li> <li>UniProt Identifier Mapping: <a href="https://www.uniprot.org/uniprot/?query=&fil=organism%3A%22Homo%20sapiens%20(Human)%20%5B9606%5D%22&columns=id%2Cdatabase(GeneID)%2Cdatabase(Ensembl)%2Cdatabase(HGNC)%2Cgenes(PREFERRED)%2Cgenes(ALTERNATIVE)"><code>UniProt Identifier Query Results</code></a></li> </ul> </li> <li> <p>Generated Data</p> <ul> <li>Merged Gene, RNA, Protein Map: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/Merged_gene_rna_protein_identifiers.pkl"><code>Merged_gene_rna_protein_identifiers.pkl</code></a></li> <li>Protein-Cofactor Relations: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/UNIPROT_PROTEIN_COFACTOR.txt"><code>UNIPROT_PROTEIN_COFACTOR.txt</code></a></li> <li>Protein-Catalyst Relations: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/UNIPROT_PROTEIN_CATALYST.txt"><code>UNIPROT_PROTEIN_CATALYST.txt</code></a></li> <li>UniProt Accession-PRO Identifier Mapping: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/UNIPROT_ACCESSION_PRO_ONTOLOGY_MAP.txt"><code>UNIPROT_ACCESSION_PRO_ONTOLOGY_MAP.txt</code></a></li> <li>UniProt Accession-Entrez Gene Identifier Mapping: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/UNIPROT_ACCESSION_ENTREZ_GENE_MAP.txt"><code>UNIPROT_ACCESSION_ENTREZ_GENE_MAP.txt</code></a></li> </ul> </li> </ul> <p> </p> <p>This project is licensed under Apache License 2.0 - see the <strong><a href="https://github.com/callahantiff/PheKnowLator/blob/master/LICENSE"><code>LICENSE.md</code></a></strong> file for details. If you intend to use any of the information on this Wiki, please provide the appropriate attribution by citing this repository:</p> <pre><code>@misc{callahan_tj_2019_3401437, author = {Callahan, TJ}, title = {PheKnowLator}, month = mar, year = 2019, doi = {10.5281/zenodo.3401437}, url = {https://doi.org/10.5281/zenodo.3401437} }</code></pre>
MIRA-KG: A Knowledge Graph of Hypotheses and Findings for Social Demography Research
<p>A shift in scientific publishing from paper-based to knowledge-based practices promotes reproducibility, machine actionability and knowledge discovery. This is important for disciplines like social science, as study indicators are often social constructs such as race or education; hypothesis tests are challenging to compare in demographic research due to their limited temporal and spatial coverage; and natural language in research papers is often imprecise and ambiguous. Therefore, we present the MIRA-KG, consisting of: (1) an ontology for capturing social demography research, which links hypotheses and findings to evidence, (2) annotations of papers on health inequality in terms of the ontology, gathered by (i) prompting a Large Language Model to annotate paper abstracts using the ontology, (ii) mapping concepts to terms from NCBO BioPortal ontologies and GeoNames, and (iii) refining the final graph by a set of SHACL constraints, developed according to data quality criteria. The utility of the resource lies in its use for formally representing social demography research hypotheses, discovering research biases, discovery of knowledge, and the derivation of novel questions.<br><br>This dataset was generated using the code available on Github at <a href="https://w3id.org/mira/">https://w3id.org/mira/</a> at version v1.0. It uses the following ontology: <a href="https://w3id.org/mira/ontology/">https://w3id.org/mira/ontology/</a>. </p>
Datasets for Non-Parametric Class Completeness Estimators for Collaborative Knowledge Graphs
<p><strong>Non-Parametric Class Completeness Estimators for Collaborative Knowledge Graphs</strong></p> <p>This are intermediary datasets used for the calculation of the Class Completeness Estimators on Wikidata. For more information see: https://github.com/eXascaleInfolab/cardinal/</p> <p><strong>edits_wikidatawiki-20181001-pages.csv</strong></p> <p>This is an extract from <em>wikidatawiki-20181001-pages-meta-history</em> (All pages with complete page edit history (.bz2)) found at <a href="https://dumps.wikimedia.org/wikidatawiki/">https://dumps.wikimedia.org/wikidatawiki/</a>.</p> <p>The extract was created by the following SQL query:</p> <pre> SELECT page_title, rev_comment, rev_user_text, rev_timestamp FROM revisions WHERE rev_comment LIKE '%[[Property:%]]%[[Q%' ORDER BY rev_id INTO OUTFILE 'edits_wikidatawiki-20181001-pages.csv'; </pre> <p> </p> <p><strong>wikidata-20180813-all.json.bz2.universe.noattr.gt.bz2</strong></p> <p>This is a graph-tool representation of the WikiData graph. Output of <a href="https://github.com/eXascaleInfolab/cardinal/blob/master/1_create_inmemory_graph.py">https://github.com/eXascaleInfolab/cardinal/blob/master/1_create_inmemory_graph.py</a>.</p> <p><strong>observations_wikidatawiki-20181001-pages.pickle</strong></p> <p>Extracted observations. Output of <a href="https://github.com/eXascaleInfolab/cardinal/blob/master/2_extract_observations.py">https://github.com/eXascaleInfolab/cardinal/blob/master/2_extract_observations.py</a>.</p> <p> </p> <p><strong>estimates_wikidatawiki-20181001-pages.pickle</strong></p> <p>Extracted estimates. Output of <a href="https://github.com/eXascaleInfolab/cardinal/blob/master/3_calculate_estimates.py">https://github.com/eXascaleInfolab/cardinal/blob/master/3_calculate_estimates.py</a></p> <p> </p> <p><strong>results_wikidatawiki-20181001-pages.pickle </strong></p> <p>Results. Output of <a href="https://github.com/eXascaleInfolab/cardinal/blob/master/4_draw_graphs.py">https://github.com/eXascaleInfolab/cardinal/blob/master/4_draw_graphs.py</a></p>
Universal Knowledge Graph Embeddings
<p>The dataset provides embeddings for entities and relations in DBpedia (English) and Wikidata. The two knowledge graphs are first merged using a novel approach that we developed by leveraging the sameAs links between them. Then, we used the state-of-the-art embedding model ConEx to compute embeddings of the merge. Our embeddings are called universal knowledge graph embeddings.</p>
cultural-ai/wordsmatter: Words Matter: a knowledge graph of contentious terms
<p>The choice of words describing cultural heritage can cause debates. It is especially sensitive when artefacts relate to different cultures and peoples who have been historically marginalised. Words chosen by archivists or curators may transmit stereotypes. The cultural heritage community has produced knowledge on potentially stereotyping and offensive terminology in heritage collections. At the same time, their knowledge is difficult to incorporate into existing online collections unless this knowledge is structured and machine-readable.</p> <p>The Words Matter Knowledge Graph represents domain expert knowledge on discussions about contentious terminology in the cultural sector. In the knowledge graph, 75 English and 83 Dutch contentious terms are linked to explanations of their usage and suggested alternatives from domain experts. There are also related matches between contentious terms and sources from external datasets: Wikidata, Princeton WordNet, Open Dutch WordNet, and Getty Art & Architecture Thesaurus.</p> <p>This Zenodo publication includes the CULCO scheme used to model contentious terms in the knowledge graph. The scheme documentation is <a href="https://cultural-ai.github.io/wordsmatter/" target="_blank" rel="noopener">available on a separate page</a>.</p> <p>This knowledge graph is <a href="https://amsterdam.wereldmuseum.nl/en/about-wereldmuseum-amsterdam/research/words-matter-publication" target="_blank" rel="noopener">based</a> on the publication “Words Matter: An Unfinished Guide to Word Choices in the Cultural Sector” by the National Museum of World Cultures (NMVW). </p> <p><a href="https://doi.org/10.1007/978-3-031-33455-9_30" target="_blank" rel="noopener">Read more</a> about this work in the paper "A Knowledge Graph of Contentious Terminology for Inclusive Representation of Cultural Heritage" (2023) by Andrei Nesterov, Laura Hollink, Marieke van Erp & Jacco van Ossenbruggen.</p> <p>In this version:</p> <ul> <li>the CULCO scheme documentation is updated</li> <li>versioning is fixed</li> <li>typos are corrected</li> </ul>
LauNuts: A Knowledge Graph to identify and compare geographic regions in the European Union
<p><strong>LauNuts</strong> is a RDF Knowledge Graph consisting of:</p> <ul> <li>Local Administrative Units (LAU) and</li> <li>Nomenclature of Territorial Units for Statistics (NUTS)</li> </ul> <p><a href="https://w3id.org/launuts">https://w3id.org/launuts</a></p>
WikiCausal Corpus for Evaluation of Causal Knowledge Graph Construction
<p>Documentation on the data format and how it can be used can be found on: <a href="https://github.com/IBM/wikicausal">https://github.com/IBM/wikicausal</a> as well as our paper:</p> <pre><code>@unpublished{, author = {Oktie Hassanzadeh and Mark Feblowitz}, title = {{WikiCausal}: Corpus and Evaluation Framework for Causal Knowledge Graph Construction}, year = {2023}, doi = {10.5281/zenodo.7897996} }</code></pre> <pre>Corpus derived from Wikipedia and Wikidata. Refer to Wikipedia and Wikidata <a href="https://en.wikipedia.org/wiki/Wikipedia:Copyrights">license and terms of use</a> for more details:</pre> <ul> <li><strong>Permission is granted</strong> to copy, distribute and/or modify Wikipedia's text under the terms of the Creative Commons Attribution-ShareAlike 3.0 Unported License and, <em>unless otherwise noted</em>, the GNU Free Documentation License, unversioned, with no invariant sections, front-cover texts, or back-cover texts.</li> <li>A copy of the Creative Commons Attribution-ShareAlike 3.0 Unported License is included in the section entitled "<a href="https://en.wikipedia.org/wiki/Wikipedia:Text_of_Creative_Commons_Attribution-ShareAlike_3.0_Unported_License">Wikipedia:Text of Creative Commons Attribution-ShareAlike 3.0 Unported License</a>"</li> <li>A copy of the GNU Free Documentation License is included in the section entitled "<a href="https://en.wikipedia.org/wiki/Wikipedia:Text_of_the_GNU_Free_Documentation_License">GNU Free Documentation License</a>".</li> <li>Content on Wikipedia is covered by <a href="https://en.wikipedia.org/wiki/Wikipedia:General_disclaimer">disclaimers</a>.</li> </ul> <pre>THIS DATA IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.</pre>
Dataset: an overview of knowledge graphs in NFDI
<p>This dataset contains a list of knowledge graphs (KGs), KG software, KG publications and KG use cases in context of NFDI (German National Research Data Infrastructure). The related manuscript is submitted to a poster session at the <em>1st Conference on Research Data Infrastructure - Connecting Communities, </em>12. – 14. September 2023, Karlsruhe, Germany: 'Who is using Knowledge Graphs in NFDI? An overview by the Working Group "Knowledge Graphs"'.</p>
PheKnowLator Human Disease Knowledge Graph Benchmarks -- v1.0.0
<p><strong>PKT Human Disease Knowledge Graph Benchmark Builds (v1.0.0)</strong></p><p><strong>Build Date: September 03, 2019</strong></p><p>The KG Benchmark Builds can also be downloaded from Zenodo:<br>👉 <strong>KGs:</strong> <a href="https://doi.org/10.5281/zenodo.7030200">https://doi.org/10.5281/zenodo.7030200</a><br>👉 <strong>Embeddings:</strong> <a href="https://zenodo.org/record/7030189">https://zenodo.org/record/7030189</a></p><p> </p><p><strong>Required Input Documents</strong></p><ul><li>resource_info.txt</li><li>class_source_list.txt</li><li>instance_source_list.txt</li><li>ontology_source_list.txt</li></ul><p> </p><p><strong>Data</strong></p><p><strong>Data Download Date:</strong> November 30, 2018</p><p><i><strong>Ontologies</strong></i></p><ul><li><a href="http://purl.obolibrary.org/obo/go.owl">Gene Ontology</a></li><li><a href="http://purl.obolibrary.org/obo/hp.owl">Human Phenotype Ontology</a></li></ul><p><i><strong>Classes</strong></i></p><ul><li><a href="http://purl.obolibrary.org/obo/doid.owl">Human Disease Ontology</a></li><li><a href="http://geneontology.org/gene-associations/goa_human.gaf.gz">Gene Ontology: gene associations</a></li><li><a href="https://reactome.org/download/current/gene_association.reactome">Reactome: gene associations</a></li><li><a href="http://compbio.charite.de/jenkins/job/hpo.annotations.monthly/lastStableBuild/artifact/annotation/ALL_SOURCES_ALL_FREQUENCIES_genes_to_phenotype.txt">Human Phenotype Ontology: all source annotations - genes to phenotypes</a></li><li><a href="http://compbio.charite.de/jenkins/job/hpo.annotations.monthly/lastSt">Human Phenotype Ontology: all source annotations - diseases to genes to phenotypes</a></li></ul><p><i><strong>Instances</strong></i></p><ul><li><a href="http://ctdbase.org/reports/CTD_chem_gene_ixns.tsv.gz">CTD: chemicals-genes</a></li><li><a href="http://ctdbase.org/reports/CTD_chem_pathways_enriched.tsv.gz">CTD: chemicals-pathways</a></li><li><a href="http://ctdbase.org/reports/CTD_chemicals_diseases.tsv.gz">CTD: chemicals-diseases</a></li><li><a href="http://ctdbase.org/reports/CTD_genes_pathways.tsv.gz">CTD: genes-pathways</a></li><li><a href="http://ctdbase.org/reports/CTD_diseases_pathways.tsv.gz">CTD: diseases-pathways</a></li><li><a href="https://stringdb-static.org/download/protein.links.v10.5/9606.protein.links.v10.5.txt.gz">STRING DB: Proteins</a></li><li><a href="https://string-db.org/mapping_files/entrez_mappings/entrez_gene_id.vs.string.v10.28042015.tsv">String DB: entrez gene mappings</a></li></ul><p> </p><p><strong>Knowledge Graphs</strong></p><p><strong>Knowledge Representation</strong><br>We worked with a PhD-level biologist to develop a knowledge representation (see the figure below) that modeled mechanisms underlying human disease.</p><p> </p><p>To do this, we manually mapped all possible combinations of the following six node types:</p><ul><li>Humans Diseases</li><li>Human Phenotypes</li><li>Human Genes</li><li>Gene Ontology concepts</li><li>Reactome Pathways</li><li>Chemicals</li></ul><p>As shown in the figure above, the <a href="http://basic-formal-ontology.org/">Basic Formal Ontology</a> and <a href="https://github.com/oborel/obo-relations/">Relation Ontology</a> ontologies were then used to create edges between the node types.</p><p> </p><p>As shown in this figure, the following edge-types were created:</p><ul><li><strong>Phenotypes-Genes:</strong> The <a href="http://purl.obolibrary.org/obo/hp.owl">Human Phenotype Ontology (HP)</a> provides <a href="http://compbio.charite.de/jenkins/job/hpo.annotations.monthly/lastStableBuild/artifact/annotation/ALL_SOURCES_ALL_FREQUENCIES_genes_to_phenotype.txt">phenotype-Entrez gene annotations</a> that were used to map 6,651 HP classes to 120,288 Entrez genes.</li><li><strong>Phenotypes-Diseases:</strong> The <a href="http://purl.obolibrary.org/obo/hp.owl">HP</a> provides <a href="http://compbio.charite.de/jenkins/job/hpo.annotations.monthly/lastStableBuild/artifact/annotation/ALL_SOURCES_ALL_FREQUENCIES_diseases_to_genes_to_phenotypes.txt">HP-DOID-Gene annotations</a> that were used to map 5,438 HP concepts to 43,817 DOID concepts.</li><li><strong>Biological processes, Molecular Functions, and Cellular Locations-Genes:</strong> The <a href="http://purl.obolibrary.org/obo/go.owl">Gene Ontology (GO)</a> provides <a href="http://geneontology.org/gene-associations/goa_human.gaf.gz">GO-Gene annotations</a> that were used to map 17,505 GO concepts to 265,002 Entrez genes.</li><li><strong>Biological processes, Molecular Functions, and Cellular Locations-Pathways-Pathways:</strong> <a href="https://reactome.org/">Reactome</a> provides <a href="https://reactome.org/download/current/gene_association.reactome">GO-Gene links</a> that were used to map 17,906 pathways to 1,910 biological processes, molecular functions, and cellular locations.</li><li><strong>Chemicals-Pathways:</strong> The <a href="http://ctdbase.org/">Comparative Toxicogenomics Database (CTD)</a> provides <a href="http://ctdbase.org/reports/CTD_chem_pathways_enriched.tsv.gz">Chemical-pathway links</a> that were used to map 8,886 MESH concepts to 711,043 Reactome pathways.</li><li><strong>Chemicals-Genes:</strong> The <a href="http://ctdbase.org/">Comparative Toxicogenomics Database (CTD)</a> provides <a href="http://ctdbase.org/reports/CTD_chem_gene_ixns.tsv.gz">Chemical-Gene links</a> that were used to map 8,881 MESH concepts 410,379 Entrez genes.</li><li><strong>Chemicals-Diseases:</strong> The <a href="http://ctdbase.org/">Comparative Toxicogenomics Database (CTD)</a> provides <a href="http://ctdbase.org/reports/CTD_chemicals_diseases.tsv.gz">Chemical-Disease links</a> that were used to map 14,238 MESH concepts 1,216,900 DOID concepts.</li><li><strong>Genes-Genes:</strong> The<a href="https://string-db.org/">STRING Database</a> provides <a href="https://stringdb-static.org/download/protein.links.v10.5/9606.protein.links.v10.5.txt.gz">Gene-Gene links</a> that were used to create 594,100 gene-gene interactions. When generating these mappings, only the inferred protein-protein relationships considered to be high confidence were used (score of 700 or better).</li><li><strong>Genes-Disease:</strong> Mappings between genes and diseases were retrieved from <a href="http://www.disgenet.org/web/DisGeNET/menu">DisGeNet</a> via SPARQL endpoint and used to map 6,051 Entrez genes to 20,452 DOID concepts.</li><li><strong>Genes-Pathways:</strong> The <a href="http://ctdbase.org/">Comparative Toxicogenomics Database (CTD)</a> provides <a href="http://ctdbase.org/reports/CTD_genes_pathways.tsv.gz">Gene-Pathway links</a> that were used to map 110,370 Entrez genes to 107,029 Reactome pathways.</li><li><strong>Pathways-Disease:</strong> The <a href="http://ctdbase.org/">Comparative Toxicogenomics Database (CTD)</a> provides <a href="http://ctdbase.org/reports/CTD_diseases_pathways.tsv.gz">Pathway-Disease links</a> that were used to map 1,818 Reactome pathways to 106,727 DOID concepts.</li></ul><p> </p><p><strong>Knowledge Graph</strong><br>The knowledge graph represented above was built using the following steps: Merge Ontologies: Merge ontologies using the <a href="https://github.com/owlcollab/owltools/wiki">OWL Tools API</a><br>Express New Ontology Concept Annotations: Create new ontology annotations by asserting a relation between the instance and an instance of the ontology class. For example to assert the following relations:</p><blockquote><p><a href="https://www.ncbi.nlm.nih.gov/mesh/68009020">Morphine</a> --> <a href="https://www.ebi.ac.uk/ols/ontologies/ro/properties?iri=http%3A%2F%2Fpurl.obolibrary.org%2Fobo%2FRO_0002606">is substance that treats</a> --> <a href="https://hpo.jax.org/app/browse/term/HP:0002076">Migraine</a></p><p>We would need to create two axioms:</p><ul><li>isSubstanceThatTreats(Morphine, x1)</li><li>instanceOf(x1, Migraine)</li></ul></blockquote><p>While the instance of the HP class hemiplegic migraines can be treated as an anonymous node in the knowledge graph, we generate a new international resource identifier for each newly generated instance.</p><p><strong>Deductively Close Knowledge Graph:</strong> The knowledge graph is deductively closed by using the OWL 2 EL reasoner, ELK via Protégé v5.1.1. ELK is able to classify instances and supports inferences over class hierarchies and object properties. inference over disjointness, intersection, and existential quantification (ontology class hierarchies).</p><p><strong>Generate Edge List:</strong> The final step before exporting the edge list is to remove any nodes that are not biologically meaningful or would otherwise reduce the performance of machine learning algorithms and the algorithm used to generate embeddings.</p><p> </p><p>🚨 <strong>AVAILABLE FILES </strong>🚨Available KG benchmark files are zipped and listed below. For additional details on what each file contains, please see the associated Wiki page 👉 <a href="https://github.com/callahantiff/PheKnowLator/wiki/September-3,-2019">here</a>.</p>
WarSampo Knowledge Graph
<p>WarSampo Knowledge Graph includes harmonized data of different kinds concerning the Second World War in Finland, separated in different subgraphs representing events, actors, places, photographs, and other aspects and documentation of the war. The data covers the Winter War 1939-1940 against the Soviet attack, the Continuation War 1941-1944 where the occupied areas of the Winter War were temporarily regained, and the Lapland War 1944-1945, where the Finns pushed the German troops away from Lapland.</p> <p>To test and demonstrate its usefulness, this Knowledge Graph is in use in the semantic portal <a href="https://sotasampo.fi/en">WarSampo</a>, explained in more detail in the <a href="https://seco.cs.aalto.fi/projects/sotasampo/en/">project page</a>.</p> <p>Example SPARQL queries for the data:</p> <ul> <li><a href="http://yasgui.org/#query=PREFIX+skos%3A+%3Chttp%3A%2F%2Fwww.w3.org%2F2004%2F02%2Fskos%2Fcore%23%3E%0APREFIX+rdfs%3A+%3Chttp%3A%2F%2Fwww.w3.org%2F2000%2F01%2Frdf-schema%23%3E%0APREFIX+crm%3A+%3Chttp%3A%2F%2Fwww.cidoc-crm.org%2Fcidoc-crm%2F%3E%0APREFIX+articles%3A+%3Chttp%3A%2F%2Fldf.fi%2Fschema%2Fwarsa%2Farticles%2F%3E%0APREFIX+wet%3A+%3Chttp%3A%2F%2Fldf.fi%2Fwarsa%2Fevents%2Fevent_types%2F%3E%0APREFIX+dc%3A+%3Chttp%3A%2F%2Fpurl.org%2Fdc%2Felements%2F1.1%2F%3E%0A%0A%23+Events%2C+photographs+and+articles+that+are+situated+in+Vyborg%0ASELECT+DISTINCT+%3Ftype+%3Fresource+%3FprefLabel%0AWHERE+%7B%0A++%7B%0A++++%23+Events%0A++++BIND+(%3Chttp%3A%2F%2Fldf.fi%2Fwarsa%2Fplaces%2Fmunicipalities%2Fm_place_614%3E+as+%3Fvyborg)%0A++++%3Fresource+crm%3AP7_took_place_at+%3Fvyborg+%3B%0A++++++++++++++a+%3FtypeURI+%3B%0A++++++++++++++crm%3AP4_has_time-span+%3Ftimespan+.%0A++++FILTER(%3FtypeURI+!%3D+wet%3APhotography)%0A++++%3FtypeURI+rdfs%3AsubClassOf*+crm%3AE5_Event+.%0A++%7D%0A++UNION%0A++%7B%0A++++%23+Photographs%0A++++BIND+(%3Chttp%3A%2F%2Fldf.fi%2Fwarsa%2Fplaces%2Fmunicipalities%2Fm_place_614%3E+as+%3Fvyborg)%0A++++%3Fresource+%5Ecrm%3AP94_has_created%2Fcrm%3AP7_took_place_at+%3Fvyborg+%3B%0A++++++++++++++++++++++++++++++++++a+%3FtypeURI+.%0A++%7D%0A++UNION%0A++%7B%0A++++%23+Articles%0A++++BIND+(%3Chttp%3A%2F%2Fldf.fi%2Fwarsa%2Fplaces%2Fmunicipalities%2Fm_place_614%3E+as+%3Fvyborg)%0A++++%3Fresource+articles%3Aplace%2Fskos%3ArelatedMatch+%3Fvyborg+%3B%0A++++++++++++++++++++++++++++articles%3Aauthor+%3Fauthor+%3B%0A++++++++++++++++++++++++++++articles%3Aissue+%3Fissue+%3B%0A++++++++++++++++++++++++++++a+%3FtypeURI+.%0A++%7D%0A++OPTIONAL+%7B%0A++++%3FtypeURI+skos%3AprefLabel+%3Ftype+.%0A++++FILTER(langMatches(lang(%3Ftype)%2C+%22en%22))%0A++%7D%0A++OPTIONAL+%7B%0A++++%3FtypeURI+skos%3AprefLabel+%3Ftype+.%0A++++FILTER(langMatches(lang(%3Ftype)%2C+%22fi%22))%0A++%7D%0A++OPTIONAL+%7B%0A++++%3FtypeURI+skos%3AprefLabel+%3Ftype+.%0A++%7D%0A++OPTIONAL+%7B%0A++++%3Fresource+skos%3AprefLabel%7Cdc%3Atitle+%3FprefLabel+.%0A++++FILTER(langMatches(lang(%3FprefLabel)%2C+%22en%22))%0A++%7D%0A++OPTIONAL+%7B%0A++++%3Fresource+skos%3AprefLabel%7Cdc%3Atitle++%3FprefLabel+.%0A++++FILTER(langMatches(lang(%3FprefLabel)%2C+%22fi%22))%0A++%7D%0A++OPTIONAL+%7B%0A++++%3Fresource+skos%3AprefLabel%7Cdc%3Atitle+%3FprefLabel+.%0A++%7D%0A%7D+&contentTypeConstruct=text%2Fturtle&contentTypeSelect=application%2Fsparql-results%2Bjson&endpoint=http%3A%2F%2Fldf.fi%2Fwarsa%2Fsparql&requestMethod=POST&tabTitle=Query+1&headers=%7B%7D&outputFormat=table">Events, photographs and articles that are situated in Vyborg</a></li> <li><a href="http://yasgui.org/#query=PREFIX+%3A+%3Chttp%3A%2F%2Fldf.fi%2Fwarsa%2Factors%2F%3E+%09%0APREFIX+warsa%3A+%3Chttp%3A%2F%2Fldf.fi%2Fschema%2Fwarsa%2F%3E+%09%0APREFIX+atypes%3A+%3Chttp%3A%2F%2Fldf.fi%2Fwarsa%2Factors%2Factor_types%2F%3E+%09%0APREFIX+foaf%3A+%3Chttp%3A%2F%2Fxmlns.com%2Ffoaf%2F0.1%2F%3E+%09%0APREFIX+casualties%3A+%3Chttp%3A%2F%2Fldf.fi%2Fschema%2Fwarsa%2Fcasualties%2F%3E%09%0APREFIX+skos%3A+%3Chttp%3A%2F%2Fwww.w3.org%2F2004%2F02%2Fskos%2Fcore%23%3E+%09%0APREFIX+xsd%3A+%3Chttp%3A%2F%2Fwww.w3.org%2F2001%2FXMLSchema%23%3E+%09%0APREFIX+crm%3A+%3Chttp%3A%2F%2Fwww.cidoc-crm.org%2Fcidoc-crm%2F%3E+%09%0APREFIX+geo%3A+%3Chttp%3A%2F%2Fwww.w3.org%2F2003%2F01%2Fgeo%2Fwgs84_pos%23%3E%09%0A%0A%23+Place%2Fdate+distribution+for+casualties+of+the+1st+Division+and+its+subunits+in+time+interval+13.2.-13.3.1940%0ASELECT+%3Fplacename+%3Flat+%3Flon+(SUM(%3Fw)+AS+%3Fnum_casualties)+%3Fdate+WHERE+%7B+%09%0A++%7B%0A++++SELECT+%3Fplace+(COUNT(%3Fid)+AS+%3Fw)+%3Fdate+WHERE+%7B%0A++++++%3Aactor_1135+(%5Ecrm%3AP144_joined_with%2Fcrm%3AP143_joined)*+%3Fsubunit+.%0A%0A++++++%3Fid+a+warsa%3ADeathRecord+%3B%0A++++++++++casualties%3Aunit+%3Fsubunit+%3B%09%0A++++++++++warsa%3Adate_of_death+%3Fdate+.%0A%0A++++++FILTER(%3Fdate+%3E%3D+%221940-02-13%22%5E%5Exsd%3Adate+%26%26+%3Fdate+%3C%3D+%221940-03-13%22%5E%5Exsd%3Adate)+%09%09%0A%0A++++++%3Fid+casualties%3Amunicipality_of_death+%3Fplace+.%0A%0A++++%7D%09GROUP+BY+%3Fplace+%3Fw+%3Fdate%0A++%7D++%09%0A++FILTER+(%3Fw+%3E+0)+%0A++%3Fplace+skos%3AprefLabel+%3Fplacename+.%0A++OPTIONAL+%7B%0A++++%3Fplace+geo%3Alat+%3Flat+%3B+%0A+++++++++++geo%3Along+%3Flon+.%09%0A++%7D%0A%7D+GROUP+BY+%3Fplacename+%3Flat+%3Flon+%3Fweigth+%3Fdate+ORDER+BY+%3Fdate&contentTypeConstruct=text%2Fturtle&contentTypeSelect=application%2Fsparql-results%2Bjson&endpoint=http%3A%2F%2Fldf.fi%2Fwarsa%2Fsparql&requestMethod=POST&tabTitle=Query&headers=%7B%7D&outputFormat=table">Casualties of the 1st Division and its subunits in the time interval 13.2.-13.3.1940 by place and date</a></li> </ul> <p>WarSampo knowledge graph version history:</p> <ul> <li>1.0.0, November 2015: Initial public release</li> <li>1.1.0, November 2017: War cemeteries addition</li> <li>2.0.0, May 2018: Backwards-incompatible URI changes</li> <li>2.0.1, November 2019: Updated schema and VoiD descriptions</li> <li>2.1.0, November 2019: Prisoners of war addition</li> </ul> <p>Version 2.1.0 contains 14,322,426 triples.</p> <p>To combine the files into a single Turtle file on a Linux system:</p> <pre><code class="language-bash">find . -mindepth 2 -name "*.ttl" | xargs cat >> warsampo.ttl</code></pre> <p> </p>
Selected survey papers for creating a knowledge graph
<p>This file contains the set of selected survey papers for populating a scholarly knowledge graph. This includes the paper title, table reference, source and full paper reference. </p>
ArCo Knowledge Graph v0.1
<p>Version 0.1 of the ArCo knowledge graph contains the ontology network and the data about the cultural properties catalogued by the Italian Institute of the General Catalogue and Documentation.</p> <p>Data are represented with RDF and by using N-Triples as syntax.</p> <p>The ontologies of the network are modelled with OWL 2 and serialised with the RDF/XML syntax.</p> <p>The ontology network is released along with alignments to other ontologies/vocabularies in the Semantic Web. Those alignments are provided within separate OWL files.</p> <p>The data are contained into a single RDF dump serialised as N-TRIPLES.</p> <p>Additionally, the release provides the links between ArCO entities and other entities published in other datasets in the Linked Open Data cloud. Such links are represented by using owl:sameAs axioms and serialised as N-TRIPLES into a separate file.</p>
CoDEx: A Comprehensive Knowledge Graph Completion Benchmark
<p>This repository hosts the <strong>relational-only part</strong> of the CoDEx benchmark, which was presented at the EMNLP 2020 conference. You can access the paper <a href="https://www.aclweb.org/anthology/2020.emnlp-main.669.pdf">here</a> and the full dataset, including text and pretrained models, <a href="https://bit.ly/2EPbrJs">on GitHub</a>.</p> <p>Abstract:</p> <p><em>We present CoDEx, a set of knowledge graph completion datasets extracted from Wikidata and Wikipedia that improve upon existing knowledge graph completion benchmarks in scope and level of difficulty. In terms of scope, CoDEx comprises three knowledge graphs varying in size and structure, multilingual descriptions of entities and relations, and tens of thousands of hard negative triples that are plausible but verified to be false. To characterize CoDEx, we contribute thorough empirical analyses and benchmarking experiments. First, we analyze each CoDEx dataset in terms of logical relation patterns. Next, we report baseline link prediction and triple classification results on CoDEx for five extensively tuned embedding models. Finally, we differentiate CoDEx from the popular FB15K-237 knowledge graph completion dataset by showing that CoDEx covers more diverse and interpretable content, and is a more difficult link prediction benchmark. Data, code, and pretrained models are available <a href="https://bit.ly/2EPbrJs">here</a>.</em></p>
Data Set Knowledge Graph (DSKG)
<p>We present the <strong>Data Set Knowledge Graph (<a href="http://dskg.org">DSKG.org</a>)</strong>, an <strong>RDF</strong> <strong>dataset about datasets </strong>that are <strong>linked to publications</strong> (modeled in the Microsoft Academic Knowledge Graph, MAKG) that mention the datasets. The metadata of the datasets is based on datasets that are registered in <strong>OpenAIRE</strong> and <strong>Wikidata</strong>.</p> <p><strong>What exactly do we provide?</strong></p> <ol> <li>Periodically updated <strong><a href="http://dskg.org">RDF dump files</a></strong> of the Data Set Knowledge Graph.</li> <li><strong><a href="http://dskg.org">URI resolution</a></strong> of the Data Set Knowledge Graph within the Linked Open Data.</li> <li>A publicly accessible <strong><a href="http://dskg.org">SPARQL endpoint</a></strong> containing the latest Dataset Knowledge Graph data.</li> </ol> <p><strong>How big is the Dataset Knowledge Graph?</strong></p> <p>The <a href="http://dskg.org">Dataset Knowledge Graph</a> models, among others,</p> <ul> <li>2,208 datasets from all scientific disciplines</li> <li>813,551 links to 634,803 unique papers</li> <li>1,169 authors of datasets</li> <li>208 ORCID IDs.</li> </ul> <p><strong>Potential use cases:</strong></p> <ul> <li>Use the DSKG for the development of semantic search engines (e.g. use the metadata of the linked publications of the datasets for advanced search capabilities)</li> <li>Easier data integration by using the RDF standard vocabulary DCAT and by linking resources to other data sources (e.g., combining the DSKG with other dataset collections in RDF).</li> <li>Data analysis to measure and award the provisioning of datasets (e.g., determine the scientific influence of datasets and authors).</li> </ul>
Traditional Chinese Medicine Multidimensional Knowledge Graph
<h3>Overview of the Traditional Chinese Medicine Multi-dimensional Knowledge Graph (TCM-MKG)</h3> <p>The <strong>Traditional Chinese Medicine Multi-dimensional Knowledge Graph (TCM-MKG)</strong> is a comprehensive, open-source data platform developed by Jingqi Zeng in November 2024. This platform aims to integrate and standardize a vast array of data from multiple sources, encompassing both traditional Chinese medicine (TCM) and modern biomedical sciences. By organizing and linking this diverse information, TCM-MKG acts as a bridge that connects the ancient wisdom of TCM with contemporary medical research and applications.</p> <h3>Key Features and Objectives:</h3> <ul> <li> <p><strong>Multi-source Data Integration</strong>: TCM-MKG consolidates data from over 30 authoritative resources, covering a broad spectrum of topics, including TCM terminology, Chinese patent medicines (CPM), Chinese herbal pieces (CHP), natural products (NP), chemical components, disease targets, and more. These data sources are carefully curated and interlinked, ensuring a rich, multi-dimensional view of TCM in relation to modern biomedical research. The platform incorporates data from reputable databases such as DrugBank, BioGRID, DisGeNET, STRING, and many others, ensuring that the TCM knowledge is not only expansive but also scientifically robust and cross-referenced with global biomedical standards.</p> </li> <li> <p><strong>Standardized Design for Global Interoperability</strong>: TCM-MKG adheres to international data standards and integrates with widely-used global medical classification systems such as ICD-11, UMLS, MeSH, and DOID. This ensures that the platform’s data is globally comparable and facilitates easy integration with international research efforts, promoting collaboration and knowledge exchange across the fields of TCM and modern medicine.</p> </li> <li> <p><strong>Open Source and Collaborative</strong>: In line with its mission to enhance transparency and accessibility, TCM-MKG is open-sourced in a structured tabular format. This allows researchers worldwide to freely access, contribute to, and expand upon the data, fostering interdisciplinary collaboration and accelerating innovation in both TCM research and modern medicine.</p> </li> <li> <p><strong>Advanced Analytical Capabilities</strong>: By leveraging the power of knowledge graph technology and graph-based intelligence algorithms, TCM-MKG supports deep data mining and relational reasoning. Researchers can uncover hidden associations between TCM components, diseases, and targets, providing insights into the mechanisms of herbal interactions and offering new pathways for drug discovery and therapeutic research.</p> </li> </ul> <h3>Personal Research Application:</h3> <p>Using the TCM-MKG platform, I conducted a study titled <strong>"Graph Neural Networks for Quantifying Compatibility Mechanisms in Traditional Chinese Medicine."</strong> This research applied advanced graph intelligence algorithms to quantitatively assess the compatibility mechanisms of Chinese herbal formulas (CHF). The study provides fresh insights into the underlying principles of TCM herbal combinations.</p> <p>This research has been published:</p> <p><strong>Zeng, J., & Jia, X. (2025). Quantifying compatibility mechanisms in traditional Chinese medicine with interpretable graph neural networks. <em>Journal of Pharmaceutical Analysis</em>, 101342. <a href="https://doi.org/10.1016/j.jpha.2025.101342">https://doi.org/10.1016/j.jpha.2025.101342</a></strong></p> <p>The code and methodology for this research have been open-sourced and are available on <a href="https://github.com/ZENGJingqi/GraphAI-for-TCM" target="_new" rel="noopener">GitHub</a>.</p> <h3>Acknowledgments:</h3> <p>This work benefited from the integration of data from numerous open-access and authoritative databases. We acknowledge the valuable contributions of resources such as DrugBank, BindingDB, BioGRID, DisGeNET, and many others. These datasets provided essential insights into TCM, modern drug chemistry, genetics, diseases, and related fields, forming the foundation for the traditional Chinese medicine multi-dimensional knowledge graph (TCM-MKG) used in this study. Furthermore, we utilized the PSICHIC model (https://github. com/huankoh/PSICHIC) to analyze the binding interactions between components and targets. Full citations for these resources are included.</p> <h3>Contact Information:</h3> <p>For further inquiries or more detailed information, please feel free to contact:<br><strong>Email</strong>: <a rel="noopener">zjingqi@163.com</a></p> <p> </p>
ParliamentSampo Knowledge Graph
<p>The ParliamentSampo Knowledge Graph includes data regarding Finnish Parliamentary debates and actors. The RDF data has been converted using data from the Parliament of Finland's open data services and Wikidata.</p> <p>The Knowledge Graph contains harmonized data of</p> <ol> <li>speeches from the plenary sessions of the Finnish Parliament 1907–2024, and</li> <li>members and organizations of the Parliament.</li> </ol> <p>The data model is designed for representing speeches, interruptions, items (on agenda), documents and other aspects related to plenary session speeches and minutes as well as member of the parliament and biographical information about them focusing on their political career.</p> <p>This dataset is available on a public SPARQL endpoint (<a href="http://ldf.fi/semparl/sparql"><em>http://ldf.fi/semparl/sparql</em></a>).</p> <p>To test and demonstrate its usefulness, this Knowledge Graph is in use in the semantic portal <a href="https://parlamenttisampo.fi/">ParliamentSampo</a>, explained in more detail in the <a href="https://seco.cs.aalto.fi/projects/semparl/en/">project page</a>.</p> <p>The Knowledge Graph can be downloaded also as CSV and XML files. See the dataset page on <a href="https://www.ldf.fi/dataset/semparl">LDF.fi</a> for more details.</p> <div> <div><strong>Version history</strong></div> <ul> <li>1.0.0, February 2023: Initial public release</li> <li>1.0.1, April 2024: README.md addition</li> <li>1.1.0, December 2024: added speeches from end of the year 2022, parliamentary session 2023, plenary sessions 112/1999, 120/1999 and 54/2016; changes to speeches in the plenary session 118/1999; fixes to "group of speaker" of speeches</li> <li>1.2.0, February 2025: added speeches from the parliamentary session 2024</li> <li>1.2.1, June 2025: fixes to speeches in plenary sessions 86–132/1999 (In dataset versions 1.1.0 and 1.2.0 some speeches were erroneously mixed: the same speech id had content of two speeches. This was due to processing the source data of the plenary sessions both from PDF files and HTML files. The current data on these plenary sessions is based only on the HTML source data.)</li> </ul> </div>
The Yelp Collaborative Knowledge Graph
<p>This is the The Yelp Collaborative Knowledge Graph (YCKG) - a transformation of the Yelp Open Dataset into RDF format using Y2KG. </p> <p>The full YCKG dataset can be found in <code>yelp.ttl.gz</code> and <code>yckg.tar.xz</code></p> <p><strong>Paper Abstract</strong></p> <p>The Yelp Open Dataset (YOD) contains data about businesses, reviews, and users from the Yelp website and is available for research purposes. This dataset has been widely used to develop and test Recommender Systems (RS), especially those using Knowledge Graphs (KGs), e.g., integrating taxonomies, product categories, business locations, and social network information. Unfortunately, researchers applied naive or wrong mappings while converting YOD in KGs, consequently obtaining unrealistic results. Among the various issues, the conversion processes usually do not follow state-of-the-art methodologies, fail to properly link to other KGs and reuse existing vocabularies. In this work, we overcome these issues by introducing Y2KG, a utility to convert the Yelp dataset into a KG. Y2KG consists of two components. The first is a dataset including (1) a vocabulary that extends Schema.org with properties to describe the concepts in YOD and (2) mappings between the Yelp entities and Wikidata. The second component is a set of scripts to transform YOD in RDF and obtain the Yelp Collaborative Knowledge Graph (YCKG). The design of Y2KG was driven by 16 core competency questions. YCKG includes 150k businesses and 16.9M reviews from 1.9M distinct real users, resulting in over 244 million triples (with 144 distinct predicates) for about 72 million resources, with an average in-degree and out-degree of 3.3 and 12.2, respectively.</p> <p><strong>Links</strong></p> <p>Latest GitHub release: <a href="https://github.com/MadsCorfixen/The-Yelp-Collaborative-Knowledge-Graph">https://github.com/MadsCorfixen/The-Yelp-Collaborative-Knowledge-Graph/releases/latest</a></p> <p>PURL domain: <a href="https://purl.prod.archive.org/domain/yckg">https://purl.archive.org/domain/yckg</a></p> <p><strong>Files</strong></p> <ul> <li>Graph Data Triple Files <ul> <li><code>yelp.ttl.gz</code> full dataset</li> <li><code>yckg.tar.xz</code> full dataset</li> </ul> </li> <li>One sample file for each of the Yelp domains (Businesses, Users, Reviews, Tips and Checkins), each containing 20 entities. <ul> <li><code>yelp_schema_mappings.nt.gz</code> containing the mappings from Yelp categories to Schema things.</li> <li><code>schema_hierarchy.nt.gz</code> containing the full hierarchy of the mapped Schema things.</li> <li><code>yelp_wiki_mappings.nt.gz</code> containing the mappings from Yelp categories to Wikidata entities.</li> <li><code>wikidata_location_mappings.nt.gz</code> containing the mappings from Yelp locations to Wikidata entities.</li> </ul> </li> <li>Graph Metadata Triple Files <ul> <li><code>yelp_categories.ttl</code> contains metadata for all Yelp categories.</li> <li><code>yelp_entities.ttl</code> contains metadata regarding the dataset</li> <li><code>yelp_vocabulary.ttl</code> contains metadata on the created Yelp vocabulary and properties.</li> </ul> </li> <li>Utility Files <ul> <li><code>yelp_category_schema_mappings.csv</code>. This file contains the 310 mappings from Yelp categories to Schema types. These mappings have been manually verified to be correct.</li> <li><code>yelp_predicate_schema_mappings.csv</code>. This file contains the 14 mappings from Yelp attributes to Schema properties. These mappings are manually found.</li> <li><code>ground_truth_yelp_category_schema_mappings.csv</code>. This file contains the ground truth, based on 200 manually verified mappings from Yelp categories to Schema things. The ground truth mappings were used to calculate precision and recall for the semantic mappings.</li> <li><code>manually_split_categories.csv</code>. This file contains all Yelp categories containing either a & or /, and their manually split versions. The split versions have been used in the semantic mappings to Schema things.</li> </ul> </li> </ul>
Expanding the chemical space using a Chemical Reaction Knowledge Graph
<p>This contains:</p><ul><li>the reaction graph dataset used to train the link prediction model</li></ul><p>Homepage: https://github.com/MolecularAI/reaction-graph-link-prediction</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.