Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

219

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

219 results for “Knowledge Graph”

Learn how ShareScore rates datasets ↗
zenodo40/100

EB-KG: Knowledge Graph of the first 8 eiditions Encyclopaedia Brittanica (1768-1860)

<p>This Knowlege Graph represents the information of the first eight editions of Encyclopaedia Brittanica (years: 1768 to 1860) in RDF (ttl format).</p> <p>The raw dataset is provided by the NLS in this <a href="https://data.nls.uk/data/digitised-collections/encyclopaedia-britannica/">link</a> , and it comprises of eight editions and a total of 195 volumes with a total size of 44GB. It uses two XMLs schemas: METS&nbsp; for descriptive, structural, technical and administrative metadata (Title, Author, Publisher, etc); and ALTO&nbsp; for encoding the OCR text of a page.</p> <p>In this work, we have extracted the information from METS and ALTO XMLS using <a href="https://github.com/francesNLP/defoe">defoe</a> tool and developed <a href="https://github.com/francesNLP/defoe/tree/master/defoe/nlsArticles/queries">novel information extraction heuristics</a>. With the extracted information, we created the EB-KG Knowlege Graph, which&nbsp; uses the <a href="https://francesnlp.github.io/EB-ontology/doc/index-en.html">EB Ontolgy</a>, to represent such information. Furthermore, during the information extraction phase, we have employed several techniques to mitigate two common OCR errors: long-S and the line-break hyphenation.</p> <p>The EB-KG contains 1,638,239 RDF triples. It has information from 8 editions. Each edition can have several Volumes, references to Books, Supplements; it also has an Editor and a Publisher, which can be a Person or an Organization. A Volume has several Pages, which can contain several Terms. And a Term can be either a Topic (a term described across several pages, often combining text, pictures, and tables.) or an Article (a description of the term in one- or two-paragraph long text (similar to an entry in a dictionary)). The data model of the EB-KG can be found <a href="https://francesnlp.github.io/EB-ontology/doc/dataModel.png">here</a>.</p> <p>The original ALTO files do not indicate the start and end of each EB term, the first part of our work involved the<br> automated extraction of all terms (along with their metadata) across editions, so they can be analysed independently without the surrounding text.</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Jun 2022View details →
zenodo40/100

Biodiversity Metadata Knowledge Graph (BMKG)

<p>Biodiversity Metadata Knowledge Graph (BMKG) has the <a href="https://doi.org/10.5281/zenodo.6948519">Biodiversity Metadata Ontology (BMO)</a> as its underlying schema. BMKG has 18 datasets instances. The BMKG is automatically generated using an embedding-based technique under the scope of <a href="https://github.com/fusion-jena/Meta2KG">Meta2KG </a>project.</p>

opencc-by-4.0Jul 2022View details →
zenodo40/100

Knowledge Graph 'Bergbeschau: Schwaz / Rottenburg / Rattenberg' Tyrol, 17th Century

<p>The dataset contain a .nq file of a Knowledge Graph that is based on the extracted information / data of the historical document &ldquo;Bergbeschau Schwaz/Rottenburg/Rattenberg&rdquo; of the 17<sup>th</sup> century (approx. 1666). The document is currently stored by the Salinenarchiv Bad Ischl) using the Identifier XXD6.</p> <p>The dataset also encludes the affiliated RDF files and .csv files.</p> <p>The user of the KG may explore the mines associated mine sections as well as places of the historical mining areas in the mining district of Schwaz (Falkenstein, Ringenwechsel, Mehren, Reichental, Palleiten) as well as the mining district of Rattenberg (Gro&szlig;-, Kleinkogel, Geyer).</p> <p>The ontology used to represent the claims is CIDOC CRM, an ISO certified ontology for Cultural Heritage documentation.&nbsp;Supported by the Karma tool the data is generated as RDF (Resource Description Framework). The generated RDF data is imported into a Triplestore, in this case GraphDB, and then displayed visually. This puts the data from the early mining texts into a semantically structured context and enables the exploration of a historical mining document from another perspective.</p>

opencc-by-4.0Aug 2022View details →
zenodo40/100

PheKnowLator Human Disease Knowledge Graphs - Build Data (Original)

<p><strong>RELEASE V2.1.0 KNOWLEDGE GRAPH: ORIGINAL DATA SOURCES&nbsp;</strong></p> <p><strong>Release:</strong>&nbsp;<a href="https://github.com/callahantiff/PheKnowLator/wiki/v2.0.0">v2.1.0&nbsp;</a></p> <p>The goal of this build was to create a knowledge graph that represented human disease mechanisms and included the central dogma. The data sources utilized in this release include many of the sources used in the initial release, as well as some new data made available by the&nbsp;<a href="https://ctdbase.org/">Comparative Toxicogenomics Database</a>&nbsp;and experimental data from the&nbsp;<a href="https://www.proteinatlas.org/">Human Protein Atlas</a>.</p> <p>Data sources are listed by type (Ontology and Data not represented in an ontology [Database Sources]). Additional details are provided for each data source below.&nbsp;Please see documentation on the primary release&nbsp;(<a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources">https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources</a>) for additional details on each data source as well as citation information.</p> <p><strong>Data Access:</strong></p> <ul> <li><a href="https://console.cloud.google.com/storage/browser/pheknowlator/archived_builds/release_v2.1.0/build_01MAY2021?project=pheknowlator">https://console.cloud.google.com/storage/browser/pheknowlator/archived_builds/release_v2.1.0/build_01MAY2021</a></li> </ul> <p>&nbsp;</p> <p><strong>ONTOLOGIES</strong></p> <ul> <li>Cell Ontology</li> <li>Cell Line Ontology</li> <li>Chemical Entities of Biological Interest (ChEBI) Ontology</li> <li>Gene Ontology</li> <li>Human Phenotype Ontology</li> <li>Mondo Disease Ontology</li> <li>Pathway Ontology</li> <li>Protein Ontology</li> <li>Relations Ontology</li> <li>Sequence Ontology</li> <li>Uber-Anatomy Ontology</li> <li>Vaccine Ontology</li> </ul> <p>&nbsp;</p> <p><strong>Cell Ontology (CL)</strong></p> <p><strong>Homepage:</strong>&nbsp;<strong><a href="https://github.com/obophenotype/cell-ontology"><code>GitHub</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Bard J, Rhee SY, Ashburner M.&nbsp;<a href="https://genomebiology.biomedcentral.com/articles/10.1186/gb-2005-6-2-r21">An ontology for cell types</a>. Genome Biology. 2005;6(2):R21</p> </blockquote> <p><strong>Usage:</strong>&nbsp;Utilized to connect&nbsp;<code>transcripts</code>&nbsp;and&nbsp;<code>proteins</code>&nbsp;to&nbsp;<code>cells</code>. Additionally, the edges between this ontology and its dependencies are utilized:</p> <ul> <li><strong><a href="https://www.ebi.ac.uk/chebi/"><code>ChEBI</code></a></strong></li> <li><strong><a href="http://geneontology.org/"><code>GO</code></a></strong></li> <li><strong><a href="https://github.com/pato-ontology/pato/"><code>PATO</code></a></strong></li> <li><strong><a href="https://proconsortium.org/"><code>PRO</code></a></strong></li> <li><strong><a href="https://github.com/oborel/obo-relations/"><code>RO</code></a></strong></li> <li><strong><a href="https://uberon.github.io/"><code>UBERON</code></a></strong></li> </ul> <p>&nbsp;</p> <p><strong>Cell Line Ontology (CLO)</strong></p> <p><strong>Homepage:</strong>&nbsp;<strong><a href="http://www.clo-ontology.org/"><code>http://www.clo-ontology.org/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Sarntivijai S, Lin Y, Xiang Z, Meehan TF, Diehl AD, Vempati UD, Sch&uuml;rer SC, Pang C, Malone J, Parkinson H, Liu Y.&nbsp;<a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4387853/">CLO: the cell line ontology</a>. Journal of Biomedical Semantics. 2014;5(1):37</p> </blockquote> <p><strong>Usage:</strong>&nbsp;Utilized this ontology to map&nbsp;<code>cell lines</code>&nbsp;to&nbsp;<code>transcripts</code>&nbsp;and&nbsp;<code>proteins</code>. Additionally, the edges between this ontology and its dependencies are utilized:</p> <ul> <li><strong><a href="http://www.obofoundry.org/ontology/cl.html"><code>CL</code></a></strong></li> <li><strong><a href="http://disease-ontology.org/"><code>DOID</code></a></strong></li> <li><strong><a href="https://github.com/obophenotype/ncbitaxon"><code>NCBITaxon</code></a></strong></li> <li><strong><a href="https://uberon.github.io/"><code>UBERON</code></a></strong></li> </ul> <p>&nbsp;</p> <p><strong>Chemical Entities of Biological Interest (ChEBI)</strong></p> <p><strong>Homepage:</strong>&nbsp;<strong><a href="https://www.ebi.ac.uk/chebi/"><code>https://www.ebi.ac.uk/chebi/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Hastings J, Owen G, Dekker A, Ennis M, Kale N, Muthukrishnan V, Turner S, Swainston N, Mendes P, Steinbeck C.&nbsp;<a href="https://academic.oup.com/nar/article-abstract/44/D1/D1214/2502583">ChEBI in 2016: Improved services and an expanding collection of metabolites</a>. Nucleic Acids Research. 2015;44(D1):D1214-9</p> </blockquote> <p><strong>Usage:</strong>&nbsp;Utilized to connect&nbsp;<code>chemicals</code>&nbsp;to&nbsp;<code>complexes</code>,&nbsp;<code>diseases</code>,&nbsp;<code>genes</code>,&nbsp;<code>GO biological processes</code>,&nbsp;<code>GO cellular components</code>,&nbsp;<code>GO molecular functions</code>,&nbsp;<code>pathways</code>,&nbsp;<code>phenotypes</code>,&nbsp;<code>reactions</code>, and&nbsp;<code>transcripts</code>.</p> <p>&nbsp;</p> <p><strong>Gene Ontology (GO)</strong></p> <p><strong>Homepage:</strong>&nbsp;<strong><a href="http://geneontology.org/"><code>http://geneontology.org/</code></a></strong><br> <strong>Citations:</strong></p> <blockquote> <p>Ashburner M, Ball CA, Blake JA, Botstein D, Butler H, Cherry JM, Davis AP, Dolinski K, Dwight SS, Eppig JT, Harris MA.&nbsp;<a href="https://www.nature.com/articles/ng0500_25">Gene ontology: tool for the unification of biology</a>. Nature Genetics. 2000;25(1):25</p> <p>The Gene Ontology Consortium.&nbsp;<a href="https://academic.oup.com/nar/article/47/D1/D330/5160994">The Gene Ontology Resource: 20 years and still GOing strong</a>. Nucleic Acids Research. 2018;47(D1):D330-8</p> </blockquote> <p><strong>Usage:</strong>&nbsp;Utilized to connect&nbsp;<code>biological processes</code>,&nbsp;<code>cellular components</code>, and&nbsp;<code>molecular functions</code>&nbsp;to&nbsp;<code>chemicals</code>,&nbsp;<code>pathways</code>, and&nbsp;<code>proteins</code>. Additionally, the edges between this ontology and its dependencies are utilized:</p> <ul> <li><strong><a href="http://www.obofoundry.org/ontology/cl.html"><code>CL</code></a></strong></li> <li><strong><a href="https://github.com/obophenotype/ncbitaxon"><code>NCBITaxon</code></a></strong></li> <li><strong><a href="https://github.com/oborel/obo-relations/"><code>RO</code></a></strong></li> <li><strong><a href="https://uberon.github.io/"><code>UBERON</code></a></strong></li> </ul> <p><strong>Other Gene Ontology Data Used:</strong>&nbsp;<a href="http://geneontology.org/gene-associations/goa_human.gaf.gz"><code>goa_human.gaf.gz</code></a></p> <p>&nbsp;</p> <p><strong>Human Phenotype Ontology (HPO)</strong></p> <p><strong>Homepage:</strong>&nbsp;<strong><a href="https://hpo.jax.org/"><code>https://hpo.jax.org/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>K&ouml;hler S, Carmody L, Vasilevsky N, Jacobsen JO, Danis D, Gourdine JP, Gargano M, Harris NL, Matentzoglu N, McMurry JA, Osumi-Sutherland D.&nbsp;<a href="https://academic.oup.com/nar/article-abstract/47/D1/D1018/5198478">Expansion of the Human Phenotype Ontology (HPO) knowledge base and resources</a>. Nucleic Acids Research. 2018;47(D1):D1018-27</p> </blockquote> <p><strong>Usage:</strong>&nbsp;Utilized to connect&nbsp;<code>phenotypes</code>&nbsp;to&nbsp;<code>chemicals</code>,&nbsp;<code>diseases</code>,&nbsp;<code>genes</code>, and&nbsp;<code>variants</code>. Additionally, the edges between this ontology and its dependencies are utilized:</p> <ul> <li><strong><a href="http://www.obofoundry.org/ontology/cl.html"><code>CL</code></a></strong></li> <li><strong><a href="https://www.ebi.ac.uk/chebi/"><code>ChEBI</code></a></strong></li> <li><strong><a href="http://geneontology.org/"><code>GO</code></a></strong></li> <li><strong><a href="https://uberon.github.io/"><code>UBERON</code></a></strong></li> </ul> <p><strong>Files</strong></p> <ul> <li>Other Human Phenotype Ontology Data Used:&nbsp;<a href="http://purl.obolibrary.org/obo/hp/hpoa/phenotype.hpoa"><code>phenotype.hpoa</code></a></li> </ul> <p>&nbsp;</p> <p><strong>Mondo Disease Ontology (Mondo)</strong></p> <p><strong>Homepage:</strong>&nbsp;<strong><a href="https://mondo.monarchinitiative.org/"><code>https://mondo.monarchinitiative.org/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Mungall CJ, McMurry JA, K&ouml;hler S, Balhoff JP, Borromeo C, Brush M, Carbon S, Conlin T, Dunn N, Engelstad M, Foster E.&nbsp;<a href="https://academic.oup.com/nar/article-abstract/45/D1/D712/2605791">The Monarch Initiative: an integrative data and analytic platform connecting phenotypes to genotypes across species</a>. Nucleic Acids Research. 2017;45(D1):D712-22</p> </blockquote> <p><strong>Usage:</strong>&nbsp;Utilized to connect&nbsp;<code>diseases</code>&nbsp;to&nbsp;<code>chemicals</code>,&nbsp;<code>phenotypes</code>,&nbsp;<code>genes</code>, and&nbsp;<code>variants</code>. Additionally, the edges between this ontology and its dependencies are utilized:</p> <ul> <li><strong><a href="http://www.obofoundry.org/ontology/cl.html"><code>CL</code></a></strong></li> <li><strong><a href="https://www.ncbi.nlm.nih.gov/taxonomy"><code>NCBITaxon</code></a></strong></li> <li><strong><a href="http://geneontology.org/"><code>GO</code></a></strong></li> <li><strong><a href="https://hpo.jax.org/"><code>HPO</code></a></strong></li> <li><strong><a href="https://uberon.github.io/"><code>UBERON</code></a></strong></li> </ul> <p>&nbsp;</p> <p><strong>Pathway Ontology (PW)</strong></p> <p><strong>Homepage:</strong>&nbsp;<strong><a href="https://rgd.mcw.edu/wg/home/pathway2/"><code>rgd.mcw.edu</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Petri V, Jayaraman P, Tutaj M, Hayman GT, Smith JR, De Pons J, Laulederkind SJ, Lowry TF, Nigam R, Wang SJ, Shimoyama M.&nbsp;<a href="https://www.ncbi.nlm.nih.gov/pubmed/24499703">The pathway ontology&ndash;updates and applications</a>. Journal of Biomedical Semantics. 2014;5(1):7.</p> </blockquote> <p><strong>Usage:</strong>&nbsp;Utilized to connect&nbsp;<code>pathways</code>&nbsp;to&nbsp;<code>GO biological processes</code>,&nbsp;<code>GO cellular components</code>,&nbsp;<code>GO molecular functions</code>,&nbsp;<code>Reactome pathways</code>. Several steps are taken in order to connect&nbsp;<code>Pathway Ontology</code>&nbsp;identifiers to&nbsp;<code>Reactome</code>&nbsp;pathways and&nbsp;<code>GO biological processes</code>. To connect&nbsp;<code>Pathway Ontology</code>&nbsp;identifiers to&nbsp;<code>Reactome</code>&nbsp;pathways, we use&nbsp;<a href="https://github.com/ComPath/resources/tree/master/mappings">ComPath Pathway Database Mappings</a>&nbsp;developed by Daniel Domingo-Fern&aacute;ndez (<a href="https://www.ncbi.nlm.nih.gov/pubmed/30564458">PMID:30564458</a>).</p> <p><strong>Files</strong></p> <ul> <li>Downloaded Mapping Data <ul> <li><a href="http://compath.scai.fraunhofer.de/export_mappings"><code>curated_mappings.txt</code></a></li> <li><a href="https://github.com/ComPath/resources/blob/master/mappings/kegg_reactome.csv"><code>kegg_reactome.csv</code></a></li> </ul> </li> <li>Generated Mapping Data <ul> <li><a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/REACTOME_PW_GO_MAPPINGS.txt"><code>REACTOME_PW_GO_MAPPINGS.txt</code></a></li> </ul> </li> </ul> <p>&nbsp;</p> <p><strong>Protein Ontology (PRO)</strong></p> <p><strong>Homepage:</strong>&nbsp;<strong><a href="https://proconsortium.org/"><code>https://proconsortium.org/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Natale DA, Arighi CN, Barker WC, Blake JA, Bult CJ, Caudy M, Drabkin HJ, D&rsquo;Eustachio P, Evsikov AV, Huang H, Nchoutmboube J.&nbsp;<a href="https://academic.oup.com/nar/article-abstract/39/suppl_1/D539/2508558">The Protein Ontology: a structured representation of protein forms and complexes</a>. Nucleic Acids Research. 2010;39(suppl_1):D539-45</p> </blockquote> <p><strong>Usage:</strong>&nbsp;Utilized to connect&nbsp;<code>proteins</code>&nbsp;to&nbsp;<code>chemicals</code>,&nbsp;<code>genes</code>,&nbsp;<code>anatomy</code>,&nbsp;<code>catalysts</code>,&nbsp;<code>cell lines</code>,&nbsp;<code>cofactors</code>,&nbsp;<code>complexes</code>,&nbsp;<code>GO biological processes</code>,&nbsp;<code>GO cellular components</code>,&nbsp;<code>GO molecular functions</code>,&nbsp;<code>pathways</code>,&nbsp;<code>proteins</code>,&nbsp;<code>reactions</code>, and&nbsp;<code>transcripts</code>. Additionally, the edges between this ontology and its dependencies are utilized:</p> <ul> <li><strong><a href="https://www.ebi.ac.uk/chebi/"><code>ChEBI</code></a></strong></li> <li><strong><a href="http://disease-ontology.org/"><code>DOID</code></a></strong></li> <li><strong><a href="http://geneontology.org/"><code>GO</code></a></strong></li> </ul> <p><strong>Notes:</strong>&nbsp;A partial, human-only version of this ontology was used. Details on how this version of the ontology was generated can be found under the Protein Ontology section of the&nbsp;<a href="https://github.com/callahantiff/PheKnowLator/blob/master/notebooks/Data_Preparation.ipynb"><code>Data_Preparation.ipynb</code></a>&nbsp;Jupyter Notebook.</p> <p><strong>Files</strong></p> <ul> <li> <p>Generated Human Version Protein Ontology (PRO)</p> <ul> <li><a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/human_pro.owl"><code>human_pro.owl</code></a>&nbsp;(closed with&nbsp;<a href="http://www.hermit-reasoner.com/">hermit reasoner</a>)</li> </ul> </li> <li> <p>Other PRO Data Used:&nbsp;<a href="https://proconsortium.org/download/current/promapping.txt"><code>promapping.txt</code></a></p> </li> <li> <p>Generated Mapping Data</p> <ul> <li>Merged Gene, RNA, Protein Map:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/Merged_gene_rna_protein_identifiers.pkl"><code>Merged_gene_rna_protein_identifiers.pkl</code></a></li> <li>Ensembl Transcript-PRO Identifier Mapping:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/ENTREZ_GENE_ENSEMBL_TRANSCRIPT_MAP.txt"><code>ENSEMBL_TRANSCRIPT_PROTEIN_ONTOLOGY_MAP.txt</code></a></li> <li>Entrez Gene-PRO Identifier Mapping:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/ENTREZ_GENE_PRO_ONTOLOGY_MAP.txt"><code>ENTREZ_GENE_PRO_ONTOLOGY_MAP.txt</code></a></li> <li>UniProt Accession-PRO Identifier Mapping:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/UNIPROT_ACCESSION_PRO_ONTOLOGY_MAP.txt"><code>UNIPROT_ACCESSION_PRO_ONTOLOGY_MAP.txt</code></a></li> <li>STRING-PRO Identifier Mapping:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/STRING_PRO_ONTOLOGY_MAP.txt"><code>STRING_PRO_ONTOLOGY_MAP.txt</code></a></li> </ul> </li> </ul> <p>&nbsp;</p> <p><strong>Relations Ontology (RO)</strong></p> <p><strong>Homepage:</strong>&nbsp;<strong><a href="https://github.com/oborel/obo-relations/"><code>GitHub</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Smith B, Ceusters W, Klagges B, K&ouml;hler J, Kumar A, Lomax J, Mungall C, Neuhaus F, Rector AL, Rosse C.&nbsp;<a href="https://genomebiology.biomedcentral.com/articles/10.1186/gb-2005-6-5-r46">Relations in biomedical ontologies</a>. Genome Biology. 2005;6(5):R46.</p> </blockquote> <p><strong>Usage:</strong>&nbsp;Utilizing this ontology to connect all data sources in knowledge graph. Additionally, the ontology is queried prior to building the knowledge graph to identify all relations, their inverse properties, and their labels.</p> <p><strong>Files</strong></p> <ul> <li>Generated RO Data <ul> <li><a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/INVERSE_RELATIONS.txt"><code>INVERSE_RELATIONS.txt</code></a></li> <li><a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/RELATIONS_LABELS.txt"><code>RELATIONS_LABELS.txt</code></a></li> </ul> </li> </ul> <p>&nbsp;</p> <p><strong>Sequence Ontology (SO)</strong></p> <p><strong>Homepage:</strong>&nbsp;<strong><a href="https://github.com/The-Sequence-Ontology/SO-Ontologies"><code>GitHub</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Eilbeck K, Lewis SE, Mungall CJ, Yandell M, Stein L, Durbin R, Ashburner M.&nbsp;<a href="https://link.springer.com/article/10.1186/gb-2005-6-5-r44">The Sequence Ontology: a tool for the unification of genome annotations</a>. Genome Biology. 2005;6(5):R44</p> </blockquote> <p><strong>Usage:</strong>&nbsp;Utilized to connect&nbsp;<code>transcripts</code>&nbsp;and other genomic material like&nbsp;<code>genes</code>&nbsp;and&nbsp;<code>variants</code>.</p> <p><strong>Files</strong></p> <ul> <li>Generated Mapping Data <ul> <li><a href="https://storage.googleapis.com/pheknowlator/curated_data/genomic_sequence_ontology_mappings.xlsx"><code>genomic_sequence_ontology_mappings.xlsx</code></a></li> <li><a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/SO_GENE_TRANSCRIPT_VARIANT_TYPE_MAPPING.txt"><code>SO_GENE_TRANSCRIPT_VARIANT_TYPE_MAPPING.txt</code></a></li> </ul> </li> </ul> <p>&nbsp;</p> <p><strong>Uber-Anatomy Ontology (Uberon)</strong></p> <p><strong>Homepage:</strong>&nbsp;<strong><a href="https://uberon.github.io/"><code>GitHub</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Mungall CJ, Torniai C, Gkoutos GV, Lewis SE, Haendel MA.&nbsp;<a href="https://genomebiology.biomedcentral.com/articles/10.1186/gb-2012-13-1-r5">Uberon, an integrative multi-species anatomy ontology</a>. Genome Biology. 2012;13(1):R5</p> </blockquote> <p><strong>Usage:</strong>&nbsp;Utilized to connect&nbsp;<code>tissues</code>,&nbsp;<code>fluids</code>, and&nbsp;<code>cells</code>&nbsp;to&nbsp;<code>proteins</code>&nbsp;and&nbsp;<code>transcripts</code>. Additionally, the edges between this ontology and its dependencies are utilized:</p> <ul> <li><strong><a href="https://www.ebi.ac.uk/chebi/"><code>ChEBI</code></a></strong></li> <li><strong><a href="http://www.obofoundry.org/ontology/cl.html"><code>CL</code></a></strong></li> <li><strong><a href="http://geneontology.org/"><code>GO</code></a></strong></li> <li><strong><a href="https://proconsortium.org/"><code>PRO</code></a></strong></li> </ul> <p>&nbsp;</p> <p><strong>Vaccine Ontology (VO)</strong></p> <p><strong>Homepage:</strong>&nbsp;<strong><a href="http://www.violinet.org/vaccineontology/"><code>http://www.violinet.org/vaccineontology/</code></a></strong><br> <strong>Citations:</strong></p> <blockquote> <p>He Y, Racz R, Sayers S, Lin Y, Todd T, Hur J, Li X, Patel M, Zhao B, Chung M, Ostrow J.&nbsp;<a href="https://academic.oup.com/nar/article-abstract/42/D1/D1124/1053128">Updates on the web-based VIOLIN vaccine database and analysis system</a>. Nucleic Acids Research. 2013;42(D1):D1124-32</p> <p>Xiang Z, Todd T, Ku KP, Kovacic BL, Larson CB, Chen F, Hodges AP, Tian Y, Olenzek EA, Zhao B, Colby LA.&nbsp;<a href="https://academic.oup.com/nar/article-abstract/36/suppl_1/D923/2505793">VIOLIN: vaccine investigation and online information network</a>. Nucleic Acids Research. 2007;36(suppl_1):D923-8</p> </blockquote> <p><strong>Usage:</strong>&nbsp;Utilized the edges between this ontology and its dependencies:</p> <ul> <li><strong><a href="https://www.ebi.ac.uk/chebi/"><code>ChEBI</code></a></strong></li> <li><strong><a href="http://disease-ontology.org/"><code>DOID</code></a></strong></li> <li><strong><a href="http://geneontology.org/"><code>GO</code></a></strong></li> <li><strong><a href="https://proconsortium.org/"><code>PRO</code></a></strong></li> <li><strong><a href="https://uberon.github.io/"><code>UBERON</code></a></strong></li> </ul> <p>&nbsp;</p> <p><strong>DATABASE SOURCES</strong></p> <ul> <li>BioPortal</li> <li>ClinVar</li> <li>Comparative Toxicogenomics Database</li> <li>DisGeNET</li> <li>Ensembl</li> <li>GeneMANIA</li> <li>Genotype-Tissue Expression Project</li> <li>Human Genome Organisation Gene Nomenclature Committee</li> <li>Human Protein Atlas</li> <li>National Center for Biotechnology Information Gene</li> <li>Reactome Pathway Database</li> <li>Search Tool for Recurring Instances of Neighbouring Genes Database</li> <li>Universal Protein Resource Knowledgebase</li> </ul> <p>&nbsp;</p> <p><strong>BioPortal</strong></p> <p><strong>Homepage:</strong>&nbsp;<strong><a href="https://bioportal.bioontology.org/"><code>BioPortal</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>BioPortal.&nbsp;<a href="https://www.bioontology.org/wiki/LOOM">Lexical OWL Ontology Matcher (LOOM)</a></p> <p>Ghazvinian A, Noy NF, Musen MA.&nbsp;<a href="https://www.ncbi.nlm.nih.gov/pmc/articles/pmc2815474/">Creating mappings for ontologies in biomedicine: simple methods work</a>. In AMIA Annual Symposium Proceedings 2009 (Vol. 2009, p. 198). American Medical Informatics Association</p> </blockquote> <p><strong>Usage:</strong>&nbsp;BioPortal was utilized to obtain mappings between&nbsp;<code>MeSH identifiers</code>&nbsp;and&nbsp;<code>ChEBI identifiers</code>&nbsp;for&nbsp;<code>chemicals-diseases</code>,&nbsp;<code>chemicals-genes</code>,&nbsp;<code>chemical-GO biological processes</code>,&nbsp;<code>chemicals-GO cellular components</code>,&nbsp;<code>chemicals-GO molecular functions</code>,&nbsp;<code>chemicals-phenotypes</code>,&nbsp;<code>chemicals-proteins</code>, and&nbsp;<code>chemicals-transcripts</code>. Additional information on how this data was processed can be obtained from the&nbsp;<a href="https://gist.github.com/callahantiff/a28fb3160782f42f104e9ec41553af0d"><code>NCBO_rest_api.py</code></a>&nbsp;GitHub Gist script.</p> <p>⭐&nbsp;<strong>ALTERNATIVE METHOD</strong>⭐ Since the above approach can take over two days to process, we have developed an alternative solution that downloads the&nbsp;<a><code>mesh2021.nt</code></a>&nbsp;data file directly from MeSH and the&nbsp;<a><code>Flat_file_tab_delimited/names.tsv.gz</code></a>&nbsp;file directly from ChEBI. Using these files, we have recapitulated the&nbsp;<a href="https://www.bioontology.org/wiki/BioPortal_Mappings"><code>LOOM</code></a>&nbsp;algorithm implemented by BioPortal when creating mappings between these resources. The procedure is relatively straightforward and utilizes the following information from each resource:</p> <ul> <li>For all MeSH&nbsp;<code>SCR Chemicals</code>, obtain the following information: <ul> <li>Identifiers: MeSH identifiers</li> <li>Labels: string labels using the&nbsp;<code>RDFS:label</code>&nbsp;object property</li> <li>Synonyms: track down all synonyms using the&nbsp;<code>vocab:concept</code>&nbsp;and&nbsp;<code>vocab:preferredConcept</code>&nbsp;object properties</li> </ul> </li> <li>For all ChEBI classes, obtain the following information: <ul> <li>Labels: string labels using the&nbsp;<code>RDFS:label</code>&nbsp;object property</li> <li>Synonyms: track down all synonyms using all&nbsp;<code>synonym</code>&nbsp;object properties</li> </ul> </li> </ul> <p><strong>Files</strong></p> <ul> <li>Generated Data:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/MESH_CHEBI_MAP.txt"><code>MESH_CHEBI_MAP.txt</code></a></li> </ul> <p>&nbsp;</p> <p><strong>ClinVar</strong></p> <p><strong>Homepage:</strong>&nbsp;<strong><a href="https://www.ncbi.nlm.nih.gov/clinvar/"><code>https://www.ncbi.nlm.nih.gov/clinvar/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Landrum MJ, Lee JM, Benson M, Brown GR, Chao C, Chitipiralla S, Gu B, Hart J, Hoffman D, Jang W, Karapetyan K.&nbsp;<a href="https://academic.oup.com/nar/article-abstract/46/D1/D1062/4641904">ClinVar: improving access to variant interpretations and supporting evidence</a>. Nucleic Acids Research. 2017;46(D1):D1062-7</p> </blockquote> <p><strong>Usage:</strong>&nbsp;ClinVar was utilized to create&nbsp;<code>variant-gene</code>,&nbsp;<code>variant-disease</code>, and&nbsp;<code>variant-phenotype</code>&nbsp;edges. The original data is filtered such that only records meeting the following criteria were included:</p> <ul> <li> <p><code>Assembly</code>&nbsp;= &quot;GRCh38&quot;</p> </li> <li> <p><code>ClinSigSimple</code>&nbsp;=&nbsp;<code>1</code></p> <ul> <li> <blockquote> <p>1 = at least one current record submitted with an interpretation of Likely pathogenic or Pathogenic (independent of whether that record includes assertion criteria and evidence)&quot;</p> </blockquote> </li> </ul> </li> <li> <p><code>ReviewStatus</code>&nbsp;in [&quot;criteria provided, multiple submitters, no conflicts&quot;, &quot;reviewed by expert panel&quot;, &quot;practice guideline&quot;]</p> </li> </ul> <p><strong>Files</strong></p> <ul> <li> <p>Downloaded Data</p> <ul> <li><a href="https://ftp.ncbi.nlm.nih.gov/pub/clinvar/tab_delimited/variant_summary.txt.gz"><code>variant_summary.txt.gz</code></a></li> <li><a href="https://ftp.ncbi.nlm.nih.gov/pub/clinvar/tab_delimited/var_citations.txt"><code>var_citations.txt</code></a></li> <li><a href="https://ftp.ncbi.nlm.nih.gov/pub/clinvar/tab_delimited/allele_gene.txt.gz"><code>allele_gene.txt.gz</code></a></li> </ul> </li> <li> <p>Generated Edge Data:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/CLINVAR_VARIANT_GENE_DISEASE_PHENOTYPE_EDGES.txt"><code>CLINVAR_VARIANT_GENE_DISEASE_PHENOTYPE_EDGES.txt</code></a></p> </li> </ul> <p>&nbsp;</p> <p><strong>Comparative Toxicogenomics Database (CTD)</strong></p> <p><strong>Homepage:</strong>&nbsp;<strong><a href="http://ctdbase.org/"><code>http://ctdbase.org/</code></a></strong><br> <strong>Citations:</strong></p> <blockquote> <p>Curated [chemical&ndash;gene interactions|chemical-go interactions|chemical&ndash;disease interactions|gene&ndash;pathway interactions] data were retrieved from the Comparative Toxicogenomics Database (CTD), MDI Biological Laboratory, Salisbury Cove, Maine, and NC State University, Raleigh, North Carolina. World Wide Web (URL:&nbsp;<a href="http://ctdbase.org/">http://ctdbase.org/</a>)</p> <p>Davis AP, Grondin CJ, Johnson RJ, Sciaky D, McMorran R, Wiegers J, Wiegers TC, Mattingly CJ.&nbsp;<a href="https://academic.oup.com/nar/article-abstract/47/D1/D948/5106145">The comparative toxicogenomics database: update 2019</a>. Nucleic Acids Research. 2018;47(D1):D948-54</p> </blockquote> <p>Usage: Comparative Toxicogenomics Database (CTD) was utilized to create&nbsp;<code>chemical-disease</code>,&nbsp;<code>chemical-gene</code>,&nbsp;<code>chemical-GO biological process</code>,&nbsp;<code>chemical-GO cellular components</code>,&nbsp;<code>chemical-GO molecular functions</code>,&nbsp;<code>chemical-phenotype</code>,&nbsp;<code>chemical-protein</code>,&nbsp;<code>chemical-rna</code>, and&nbsp;<code>gene-pathway</code>&nbsp;edges. The original data is filtered such that only records meeting the following criteria were included:</p> <ul> <li><code>chemical-disease</code>:&nbsp;<code>DirectEvidence</code>&nbsp;!= &quot;&quot;</li> <li><code>chemical-gene</code>:&nbsp;<code>Organism</code>&nbsp;== &quot;Homo sapiens&quot;,&nbsp;<code>GeneForms</code>&nbsp;== &quot;gene&quot;, and affects not in&nbsp;<code>InteractionActions</code></li> <li><code>chemical-GO biological process</code>:&nbsp;<code>PhenotypeName</code>&nbsp;== &quot;Biological Process&quot; and&nbsp;<code>Interaction</code>&nbsp;&lt;= &quot;1.04e-47&quot; (10th percentile)</li> <li><code>chemical-GO cellular components</code>:&nbsp;<code>PhenotypeName</code>&nbsp;== &quot;Cellular Component&quot; and&nbsp;<code>Interaction</code>&nbsp;&lt;= &quot;1.04e-47&quot; (10th percentile)</li> <li><code>chemical-GO molecular functions</code>:&nbsp;<code>PhenotypeName</code>&nbsp;== &quot;Molecular Function&quot; and&nbsp;<code>Interaction</code>&nbsp;&lt;= &quot;1.04e-47&quot; (10th percentile)</li> <li><code>chemical-phenotype</code>:&nbsp;<code>DirectEvidence</code>&nbsp;!= &quot;&quot;</li> <li><code>chemical-protein</code>:&nbsp;<code>Organism</code>&nbsp;== &quot;Homo sapiens&quot;,&nbsp;<code>GeneForms</code>&nbsp;== &quot;protein&quot;, and affects not in&nbsp;<code>InteractionActions</code></li> <li><code>chemical-rna</code>:&nbsp;<code>Organism</code>&nbsp;== &quot;Homo sapiens&quot;,&nbsp;<code>GeneForms</code>&nbsp;== &quot;mRNA&quot;, and affects and activity not in&nbsp;<code>InteractionActions</code></li> <li><code>gene-pathway edges</code>:&nbsp;<code>PathwayName</code>&nbsp;== R-HSA-</li> </ul> <p><strong>Files</strong></p> <ul> <li>Downloaded Data <ul> <li>Chemical-Gene Relations:&nbsp;<a href="http://ctdbase.org/reports/CTD_chem_gene_ixns.tsv.gz"><code>CTD_chem_gene_ixns.tsv.gz</code></a></li> <li>Chemical-Disease/Phenotype Relations:&nbsp;<a href="http://ctdbase.org/reports/CTD_chemicals_diseases.tsv.gz"><code>CTD_chemicals_diseases.tsv.gz</code></a></li> <li>Chemical-GO Relations:&nbsp;<a href="http://ctdbase.org/reports/CTD_chem_go_enriched.tsv.gz"><code>CTD_chem_go_enriched.tsv.gz</code></a></li> <li>Gene-Pathway Relations:&nbsp;<a href="http://ctdbase.org/reports/CTD_genes_pathways.tsv.gz"><code>CTD_genes_pathways.tsv.gz</code></a></li> </ul> </li> </ul> <p>&nbsp;</p> <p><strong>DisGeNET</strong></p> <p><strong>Homepage:</strong>&nbsp;<strong><a href="https://www.disgenet.org/"><code>https://www.disgenet.org/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Gene-disease association data retrieved from DisGeNET v6.0 (<a href="http://www.disgenet.org/">http://www.disgenet.org/</a>), Integrative Biomedical Informatics Group GRIB/IMIM/UPF. [December, 2019].</p> <p>Pi&ntilde;ero J, Ram&iacute;rez-Anguita JM, Sa&uuml;ch-Pitarch J, Ronzano F, Centeno E, Sanz F, Furlong LI.&nbsp;<a href="https://academic.oup.com/nar/advance-article-abstract/doi/10.1093/nar/gkz1021/5611674">The DisGeNET knowledge platform for disease genomics: 2019 update</a>. Nucleic Acids Research. 2019.</p> </blockquote> <p><strong>Usage:</strong>&nbsp;DisGeNET was utilized to create&nbsp;<code>gene-disease</code>, and&nbsp;<code>gene-phenotype</code>&nbsp;edges. The original data is filtered such that only records meeting the following criteria were included:&nbsp;<code>EI</code>&nbsp;&gt;= &quot;1.0&quot; (90th percentile). Additionally, data from this source was used to create mappings between different types of disease and phenotype identifiers, including:</p> <ul> <li>OMIM, ORPHA, UMLS, ICD ➞ DOID</li> <li>OMIM, ORPHA, UMLS, ICD ➞ HPO</li> </ul> <p><strong>Files</strong></p> <ul> <li> <p>Downloaded Data</p> <ul> <li>Disease/Phenotype-Gene Relations:&nbsp;<a href="https://www.disgenet.org/static/disgenet_ap1/files/downloads/curated_gene_disease_associations.tsv.gz"><code>curated_gene_disease_associations.tsv.gz</code></a></li> <li>Disease Identifier Mapping:&nbsp;<a href="https://www.disgenet.org/static/disgenet_ap1/files/downloads/disease_mappings.tsv.gz"><code>disease_mappings.tsv.gz</code></a></li> </ul> </li> <li> <p>Generated Mapping Data</p> <ul> <li>Disease Identifier Mapping:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/PHENOTYPE_HPO_MAP.txt"><code>PHENOTPYE_HPO_MAP.txt</code></a></li> <li>Phenotype Identifier Mapping:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/DISEASE_MONDO_MAP.txt"><code>DISEASE_DOID_MAP.txt</code></a></li> </ul> </li> </ul> <p>&nbsp;</p> <p><strong>Ensembl</strong></p> <p><strong>Homepage:</strong>&nbsp;<strong><a href="https://uswest.ensembl.org/index.html"><code>https://uswest.ensembl.org/index.html</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Zerbino DR, Achuthan P, Akanni W, Amode MR, Barrell D, Bhai J, Billis K, Cummins C, Gall A, Gir&oacute;n CG, Gil L.&nbsp;<a href="https://academic.oup.com/nar/article/46/D1/D754/4634002">Ensembl 2018</a>. Nucleic Acids Research. 2017;46(D1):D754-61</p> </blockquote> <p><strong>Usage:</strong>&nbsp;Ensembl data was utilized to create mappings between Ensembl genes, transcripts, and proteins with&nbsp;<a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#national-center-for-biotechnology-information-gene">NCBI Gene identifiers</a>,&nbsp;<a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#human-genome-organisation-gene-nomenclature-committee">HUGO gene symbols</a>,&nbsp;<a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#universal-protein-resource-knowledgebase">UniProt Accession identifiers</a>, and&nbsp;<a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#protein-ontology">Protein Ontology identifiers</a>&nbsp;in the knowledge graph (for additional details on the processing of these data, see&nbsp;<a href="https://github.com/callahantiff/PheKnowLator/blob/master/notebooks/Data_Preparation.ipynb"><code>Data_Preparation.ipynb</code></a>):</p> <ul> <li>Ensembl Transcript IDs ➞ PRO IDs</li> <li>Gene Ensembl IDs ➞ Entrez Gene IDs</li> <li>Gene Ensembl IDs ➞ PRO IDs</li> <li>Gene Symbols ➞ Transcript Ensembl IDs</li> <li>Entrez Gene IDs ➞ Transcript Ensembl IDs</li> <li>Entrez Gene IDs ➞ PRO IDs</li> <li>Protein Ensembl IDs ➞ UniProt Protein Accession</li> <li>STRING IDs ➞ PRO IDs</li> <li>UniProt Protein Accession ➞ Entrez Gene IDs</li> </ul> <p><strong>Files</strong></p> <ul> <li> <p>Downloaded Data</p> <ul> <li><a><code>Homo_sapiens.GRCh38.102.gtf</code></a></li> <li><a><code>Homo_sapiens.GRCh38.102.uniprot.tsv.gz</code></a></li> <li><a><code>Homo_sapiens.GRCh38.102.entrez.tsv.gz</code></a></li> </ul> </li> <li> <p>Generated Mapping Data</p> <ul> <li>Cleaned Ensembl Gene Set:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/ensembl_identifier_data_cleaned.txt"><code>ensembl_identifier_data_cleaned.txt</code></a></li> <li>Merged Gene, RNA, Protein Map:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/Merged_gene_rna_protein_identifiers.pkl"><code>Merged_gene_rna_protein_identifiers.pkl</code></a></li> <li>Ensembl Transcript-PRO Identifier Mapping:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/ENSEMBL_TRANSCRIPT_PROTEIN_ONTOLOGY_MAP.txt"><code>ENSEMBL_TRANSCRIPT_PROTEIN_ONTOLOGY_MAP.txt</code></a></li> <li>Gene Symbol-Ensembl Transcript Identifier Mapping:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/GENE_SYMBOL_ENSEMBL_TRANSCRIPT_MAP.txt"><code>GENE_SYMBOL_ENSEMBL_TRANSCRIPT_MAP.txt</code></a></li> <li>Entrez Gene-Ensembl Transcript Identifier Mapping:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/ENTREZ_GENE_ENSEMBL_TRANSCRIPT_MAP.txt"><code>ENTREZ_GENE_ENSEMBL_TRANSCRIPT_MAP.txt</code></a></li> <li>Entrez Gene-PRO Identifier Mapping:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/ENTREZ_GENE_PRO_ONTOLOGY_MAP.txt"><code>ENTREZ_GENE_PRO_ONTOLOGY_MAP.txt</code></a></li> <li>Ensembl Gene-Entrez Gene Identifier Mapping:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/ENSEMBL_GENE_ENTREZ_GENE_MAP.txt"><code>ENSEMBL_GENE_ENTREZ_GENE_MAP.txt</code></a></li> </ul> </li> </ul> <p>&nbsp;</p> <p><strong>GeneMANIA</strong></p> <p><strong>Homepage:</strong>&nbsp;<strong><a href="https://genemania.org/"><code>https://genemania.org/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Warde-Farley D, Donaldson SL, Comes O, Zuberi K, Badrawi R, Chao P, Franz M, Grouios C, Kazi F, Lopes CT, Maitland A.&nbsp;<a href="https://academic.oup.com/nar/article-abstract/38/suppl_2/W214/1126704">The GeneMANIA prediction server: biological network integration for gene prioritization and predicting gene function</a>. Nucleic Acids Research. 2010;38(suppl_2):W214-20</p> </blockquote> <p><strong>Usage:</strong>&nbsp;GeneMANIA was utilized to create&nbsp;<code>gene-gene</code>&nbsp;edges.</p> <p><strong>Files</strong></p> <ul> <li>Downloaded Data:&nbsp;<a href="http://genemania.org/data/current/Homo_sapiens.COMBINED/COMBINED.DEFAULT_NETWORKS.BP_COMBINING.txt"><code>COMBINED.DEFAULT_NETWORKS.BP_COMBINING.txt</code></a></li> </ul> <p>&nbsp;</p> <p><strong>Genotype-Tissue Expression Project (GTEx)</strong></p> <p><strong>Homepage:</strong>&nbsp;<strong><a href="https://gtexportal.org/home/"><code>https://gtexportal.org/home/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Lonsdale J, Thomas J, Salvatore M, Phillips R, Lo E, Shad S, Hasz R, Walters G, Garcia F, Young N, Foster B.&nbsp;<a href="http://www.nature.com/ng/journal/v45/n6/full/ng.2653.html">The genotype-tissue expression (GTEx) project</a>. Nature Genetics. 2013;45(6):580</p> </blockquote> <p><strong>Usage:</strong>&nbsp;The Genotype-Tissue Expression (GTEx) Project was utilized to create edges between&nbsp;<code>protein-cell</code>,&nbsp;<code>protein-anatomy</code>,&nbsp;<code>rna-cell</code>&nbsp;and&nbsp;<code>rna-anatomy</code>&nbsp;entities. The original data were filtered such that only those edges where the median TPM was &gt;=<code>1.0</code>&nbsp;and genes were of any type other than protein-coding were included. It should also be noted that we chose to use the RNASeQC file over the RSEM file as advised by the GTEx website.</p> <blockquote> <p>The RSEM estimates are based on combining isoform-level estimates, which adds uncertainty to the resulting gene-level values (the isoform-level estimates are highly inaccurate in some cases).</p> </blockquote> <p>The file contains&nbsp;<code>54</code>&nbsp;unique tissue and/or cell types. GTEx provides mappings from tissue types to UBERON and EFO. These provided&nbsp;<a href="https://gtexportal.org/home/samplingSitePage">mappings</a>&nbsp;were verified and extended, such that all samples which referenced a cell type were also mapped to the Cell and the Cell Line ontologies. This resulted in a total of&nbsp;<code>56</code>&nbsp;mappings (<code>1.04</code>&nbsp;mappings/concepts).</p> <p><strong>Files</strong></p> <ul> <li> <p>Downloaded Data:&nbsp;<a href="https://storage.googleapis.com/gtex_analysis_v8/rna_seq_data/GTEx_Analysis_2017-06-05_v8_RNASeQCv1.1.9_gene_median_tpm.gct.gz"><code>GTEx_Analysis_2017-06-05_v8_RNASeQCv1.1.9_gene_median_tpm.gct</code></a></p> </li> <li> <p>Mapping Results:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/curated_data/zooma_tissue_cell_mapping_04JAN2020.xlsx"><code>zooma_tissue_cell_mapping_04JAN2020.xlsx</code></a></p> </li> <li> <p>Generated Data<br> The final mapping set was combined with terms from the&nbsp;<a href="https://www.proteinatlas.org/">Human Protein Atlas</a>, see&nbsp;<a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources/t#human-protein-atlas">here</a>&nbsp;for more information.</p> <ul> <li>All HPA tissue and cell type strings:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/HPA_tissues.txt"><code>HPA_tissues.txt</code></a></li> <li>Final Term Mapping:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/HPA_GTEx_TISSUE_CELL_MAP.txt"><code>HPA_GTEx_TISSUE_CELL_MAP.txt</code></a></li> <li>Final RNA, Gene, Protein-Tissues and Cell Types Relations:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/HPA_GTEX_RNA_GENE_PROTEIN_EDGES.txt"><code>HPA_GTEX_RNA_GENE_PROTEIN_EDGES.txt</code></a></li> </ul> </li> </ul> <p>&nbsp;</p> <p><strong>Human Genome Organisation Gene Nomenclature Committee (HUGO)</strong></p> <p><strong>Homepage:</strong>&nbsp;<strong><a href="https://www.genenames.org/"><code>https://www.genenames.org/</code></a></strong><br> <strong>Citations:</strong></p> <blockquote> <p>HGNC Database, HUGO Gene Nomenclature Committee (HGNC), European Molecular Biology Laboratory, European Bioinformatics Institute (EMBL-EBI), Wellcome Genome Campus, Hinxton, Cambridge CB10 1SD, United Kingdom&nbsp;<a href="https://www.genenames.org/">www.genenames.org</a></p> <p>Yates B, Braschi B, Gray K, Seal R, Tweedie S, Bruford E.&nbsp;<a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5210531/">Genenames.org: the HGNC and VGNC Resources in 2017</a>. Nucleic Acids Research. 2017;45(D1):D619-625</p> </blockquote> <p><strong>Usage:</strong>&nbsp;The Human Genome Organisation (HUGO) data was utilized to obtain mappings between&nbsp;<a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#national-center-for-biotechnology-information-gene">NCBI Gene identifiers</a>, HUGO gene symbols,&nbsp;<a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#universal-protein-resource-knowledgebase">UniProt Accession identifiers</a>, and&nbsp;<a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#protein-ontology">Protein Ontology identifiers</a>. For additional details on the processing of these data, see&nbsp;<a href="https://github.com/callahantiff/PheKnowLator/blob/master/notebooks/Data_Preparation.ipynb"><code>Data_Preparation.ipynb</code></a>:</p> <ul> <li>Ensembl Transcript IDs ➞ PRO IDs</li> <li>Gene Ensembl IDs ➞ Entrez Gene IDs</li> <li>Gene Ensembl IDs ➞ PRO IDs</li> <li>Gene Symbols ➞ Transcript Ensembl IDs</li> <li>Entrez Gene IDs ➞ Transcript Ensembl IDs</li> <li>Entrez Gene IDs ➞ PRO IDs</li> <li>Protein Ensembl IDs ➞ UniProt Protein Accession</li> <li>STRING IDs ➞ PRO IDs</li> <li>UniProt Protein Accession ➞ Entrez Gene IDs</li> </ul> <p><strong>Files</strong></p> <ul> <li> <p>Downloaded Data:&nbsp;<a href="http://ftp.ebi.ac.uk/pub/databases/genenames/hgnc/tsv/hgnc_complete_set.txt"><code>hgnc_complete_set.txt</code></a></p> </li> <li> <p>Generated Data</p> <ul> <li>Merged Gene, RNA, Protein Map:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/Merged_gene_rna_protein_identifiers.pkl"><code>Merged_gene_rna_protein_identifiers.pkl</code></a></li> <li>Gene Symbol-Ensembl Transcript Identifier Mapping:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/GENE_SYMBOL_ENSEMBL_TRANSCRIPT_MAP.txt"><code>GENE_SYMBOL_ENSEMBL_TRANSCRIPT_MAP.txt</code></a></li> </ul> </li> </ul> <p>&nbsp;</p> <p><strong>Human Protein Atlas (HPA)</strong></p> <p><strong>Homepage:</strong>&nbsp;<strong><a href="https://www.proteinatlas.org/"><code>https://www.proteinatlas.org/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Uhl&eacute;n M, Fagerberg L, Hallstr&ouml;m BM, Lindskog C, Oksvold P, Mardinoglu A, Sivertsson &Aring;, Kampf C, Sj&ouml;stedt E, Asplund A, Olsson I.&nbsp;<a href="https://science.sciencemag.org/content/347/6220/1260419.short">Tissue-based map of the human proteome</a>. Science. 2015;347(6220):1260419</p> </blockquote> <p><strong>Usage:</strong>&nbsp;The Human Protein Atlas (HPA) was utilized to create&nbsp;<code>rna-cell</code>,&nbsp;<code>rna-anatomy</code>,&nbsp;<code>protein-cell</code>, and&nbsp;<code>protein-anatomy</code>&nbsp;edges. Evidence between gene and RNA expression in specific tissue types was derived by HPA, such that the&nbsp;<a href="https://www.proteinatlas.org/about/assays+annotation#normalization_rna">consensus normalized expression</a>&nbsp;was &gt;=<code>1.0</code>. Zooma was utilized to automatically annotate the&nbsp;<code>153</code>&nbsp;unique tissues and cell types from Human Protein Atlas for all human protein-coding genes in the&nbsp;<a href="https://www.proteinatlas.org/humanproteome">Human Proteome</a>&nbsp;to the Cell Ontology, Cell Line Ontology, and the Uber-Anatomy Ontology. To best represent each concept, the automatic mappings from Zooma were extend through manual mapping efforts to ensure each concept cell type was matched to a Cell Ontology, Cell Line Ontology, and UBERON ontology term. This resulted in a total of&nbsp;<code>281</code>&nbsp;mappings (<code>1.84</code>&nbsp;mappings/concepts).</p> <p><strong>Files</strong></p> <ul> <li> <p>Downloaded Data:&nbsp;<a href="https://www.proteinatlas.org/api/search_download.php?search=&amp;columns=g,eg,up,pe,rnatsm,rnaclsm,rnacasm,rnabrsm,rnabcsm,rnablsm,scl,t_RNA_adipose_tissue,t_RNA_adrenal_gland,t_RNA_amygdala,t_RNA_appendix,t_RNA_basal_ganglia,t_RNA_bone_marrow,t_RNA_breast,t_RNA_cerebellum,t_RNA_cerebral_cortex,t_RNA_cervix,_uterine,t_RNA_colon,t_RNA_corpus_callosum,t_RNA_ductus_deferens,t_RNA_duodenum,t_RNA_endometrium_1,t_RNA_epididymis,t_RNA_esophagus,t_RNA_fallopian_tube,t_RNA_gallbladder,t_RNA_heart_muscle,t_RNA_hippocampal_formation,t_RNA_hypothalamus,t_RNA_kidney,t_RNA_liver,t_RNA_lung,t_RNA_lymph_node,t_RNA_midbrain,t_RNA_olfactory_region,t_RNA_ovary,t_RNA_pancreas,t_RNA_parathyroid_gland,t_RNA_pituitary_gland,t_RNA_placenta,t_RNA_pons_and_medulla,t_RNA_prostate,t_RNA_rectum,t_RNA_retina,t_RNA_salivary_gland,t_RNA_seminal_vesicle,t_RNA_skeletal_muscle,t_RNA_skin_1,t_RNA_small_intestine,t_RNA_smooth_muscle,t_RNA_spinal_cord,t_RNA_spleen,t_RNA_stomach_1,t_RNA_testis,t_RNA_thalamus,t_RNA_thymus,t_RNA_thyroid_gland,t_RNA_tongue,t_RNA_tonsil,t_RNA_urinary_bladder,t_RNA_vagina,t_RNA_B-cells,t_RNA_dendritic_cells,t_RNA_granulocytes,t_RNA_monocytes,t_RNA_NK-cells,t_RNA_T-cells,t_RNA_total_PBMC,cell_RNA_A-431,cell_RNA_A549,cell_RNA_AF22,cell_RNA_AN3-CA,cell_RNA_ASC_diff,cell_RNA_ASC_TERT1,cell_RNA_BEWO,cell_RNA_BJ,cell_RNA_BJ_hTERT+,cell_RNA_BJ_hTERT+_SV40_Large_T+,cell_RNA_BJ_hTERT+_SV40_Large_T+_RasG12V,cell_RNA_CACO-2,cell_RNA_CAPAN-2,cell_RNA_Daudi,cell_RNA_EFO-21,cell_RNA_fHDF/TERT166,cell_RNA_HaCaT,cell_RNA_HAP1,cell_RNA_HBEC3-KT,cell_RNA_HBF_TERT88,cell_RNA_HDLM-2,cell_RNA_HEK_293,cell_RNA_HEL,cell_RNA_HeLa,cell_RNA_Hep_G2,cell_RNA_HHSteC,cell_RNA_HL-60,cell_RNA_HMC-1,cell_RNA_HSkMC,cell_RNA_hTCEpi,cell_RNA_hTEC/SVTERT24-B,cell_RNA_hTERT-HME1,cell_RNA_HUVEC_TERT2,cell_RNA_K-562,cell_RNA_Karpas-707,cell_RNA_LHCN-M2,cell_RNA_MCF7,cell_RNA_MOLT-4,cell_RNA_NB-4,cell_RNA_NTERA-2,cell_RNA_PC-3,cell_RNA_REH,cell_RNA_RH-30,cell_RNA_RPMI-8226,cell_RNA_RPTEC_TERT1,cell_RNA_RT4,cell_RNA_SCLC-21H,cell_RNA_SH-SY5Y,cell_RNA_SiHa,cell_RNA_SK-BR-3,cell_RNA_SK-MEL-30,cell_RNA_T-47d,cell_RNA_THP-1,cell_RNA_TIME,cell_RNA_U-138_MG,cell_RNA_U-2_OS,cell_RNA_U-2197,cell_RNA_U-251_MG,cell_RNA_U-266/70,cell_RNA_U-266/84,cell_RNA_U-698,cell_RNA_U-87_MG,cell_RNA_U-937,cell_RNA_WM-115,blood_RNA_basophil,blood_RNA_classical_monocyte,blood_RNA_eosinophil,blood_RNA_gdT-cell,blood_RNA_intermediate_monocyte,blood_RNA_MAIT_T-cell,blood_RNA_memory_B-cell,blood_RNA_memory_CD4_T-cell,blood_RNA_memory_CD8_T-cell,blood_RNA_myeloid_DC,blood_RNA_naive_B-cell,blood_RNA_naive_CD4_T-cell,blood_RNA_naive_CD8_T-cell,blood_RNA_neutrophil,blood_RNA_NK-cell,blood_RNA_non-classical_monocyte,blood_RNA_plasmacytoid_DC,blood_RNA_T-reg,blood_RNA_total_PBMC,brain_RNA_amygdala,brain_RNA_basal_ganglia,brain_RNA_cerebellum,brain_RNA_cerebral_cortex,brain_RNA_hippocampal_formation,brain_RNA_hypothalamus,brain_RNA_midbrain,brain_RNA_olfactory_region,brain_RNA_pons_and_medulla,brain_RNA_thalamus&amp;format=tsv"><code>proteinatlas_search.tsv</code></a></p> </li> <li> <p>Mapping Results:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/curated_data/zooma_tissue_cell_mapping_04JAN2020.xlsx"><code>zooma_tissue_cell_mapping_04JAN2020.xlsx</code></a></p> </li> <li> <p>Generated Data</p> <ul> <li>Final Term Mapping:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/HPA_GTEx_TISSUE_CELL_MAP.txt"><code>HPA_GTEx_TISSUE_CELL_MAP.txt</code></a></li> <li>Final RNA, Gene, Protein-Tissues and Cell Types Relations:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/HPA_GTEX_RNA_GENE_PROTEIN_EDGES.txt"><code>HPA_GTEX_RNA_GENE_PROTEIN_EDGES.txt</code></a></li> </ul> </li> </ul> <p>&nbsp;</p> <p><strong>National Center for Biotechnology Information (NCBI) Entrez Gene</strong></p> <p><strong>Homepage:</strong>&nbsp;<strong><a href="https://www.ncbi.nlm.nih.gov/gene/"><code>https://www.ncbi.nlm.nih.gov/gene/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Maglott D, Ostell J, Pruitt KD, Tatusova T.&nbsp;<a href="https://academic.oup.com/nar/article-abstract/33/suppl_1/D54/2505255">Entrez Gene: gene-centered information at NCBI</a>. Nucleic Acids Research. 2005;33(suppl_1):D54-8.</p> </blockquote> <p><strong>Usage:</strong>&nbsp;The National Center for Biotechnology Information (NCBI) Gene data was utilized to obtain mappings between&nbsp;<a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#ncbi-gene">NCBI Gene identifiers</a>,&nbsp;<a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#hugo-gene-nomenclature-committee">HUGO gene symbols</a>,&nbsp;<a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#uniprot-knowledgebase">UniProt Accession identifiers</a>, and&nbsp;<a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#protein-ontology">Protein Ontology identifiers</a>. For additional details on the processing of these data, see&nbsp;<a href="https://github.com/callahantiff/PheKnowLator/blob/master/notebooks/Data_Preparation.ipynb"><code>Data_Preparation.ipynb</code></a>:</p> <ul> <li>Ensembl Transcript IDs ➞ PRO IDs</li> <li>Gene Ensembl IDs ➞ Entrez Gene IDs</li> <li>Gene Ensembl IDs ➞ PRO IDs</li> <li>Gene Symbols ➞ Transcript Ensembl IDs</li> <li>Entrez Gene IDs ➞ Transcript Ensembl IDs</li> <li>Entrez Gene IDs ➞ PRO IDs</li> <li>Protein Ensembl IDs ➞ UniProt Protein Accession</li> <li>STRING IDs ➞ PRO IDs</li> <li>UniProt Protein Accession ➞ Entrez Gene IDs</li> </ul> <p><strong>Files</strong></p> <ul> <li> <p>Downloaded Data:&nbsp;<a href="https://ftp.ncbi.nih.gov/gene/DATA/GENE_INFO/Mammalia/Homo_sapiens.gene_info.gz"><code>Homo_sapiens.gene_info.gz</code></a></p> </li> <li> <p>Generated Data</p> <ul> <li>Merged Gene, RNA, Protein Map:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/Merged_gene_rna_protein_identifiers.pkl"><code>Merged_gene_rna_protein_identifiers.pkl</code></a></li> <li>Entrez Gene-Ensembl Transcript Identifier Mapping:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/ENTREZ_GENE_ENSEMBL_TRANSCRIPT_MAP.txt"><code>ENTREZ_GENE_ENSEMBL_TRANSCRIPT_MAP.txt</code></a></li> <li>Entrez Gene-PRO Identifier Mapping:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/ENTREZ_GENE_PRO_ONTOLOGY_MAP.txt"><code>ENTREZ_GENE_PRO_ONTOLOGY_MAP.txt</code></a></li> <li>Ensembl Gene-Entrez Gene Identifier Mapping:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/ENSEMBL_GENE_ENTREZ_GENE_MAP.txt"><code>ENSEMBL_GENE_ENTREZ_GENE_MAP.txt</code></a></li> <li>Uniprot Accession-Entrez Gene Identifier Mapping:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/UNIPROT_ACCESSION_ENTREZ_GENE_MAP.txt"><code>UNIPROT_ACCESSION_ENTREZ_GENE_MAP.txt</code></a></li> </ul> </li> </ul> <p>&nbsp;</p> <p><strong>Reactome Pathway Database</strong></p> <p><strong>Homepage:</strong>&nbsp;<strong><a href="https://reactome.org/"><code>https://reactome.org/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Fabregat A, Jupe S, Matthews L, Sidiropoulos K, Gillespie M, Garapati P, Haw R, Jassal B, Korninger F, May B, Milacic M.&nbsp;<a href="https://academic.oup.com/nar/article-abstract/46/D1/D649/4626770">The reactome pathway knowledgebase</a>. Nucleic Acids Research. 2017;46(D1):D649-55</p> </blockquote> <p><strong>Usage:</strong>&nbsp;The Reactome Database was utilized to create&nbsp;<code>chemical-pathway</code>,&nbsp;<code>GO Biological process-pathway</code>,&nbsp;<code>pathway-GO Cellular component</code>,&nbsp;<code>GO Molecular function-pathway</code>, and&nbsp;<code>protein-pathway</code>&nbsp;edges. The original data is filtered such that only records meeting the following criteria were included:</p> <ul> <li><code>chemical-pathway</code>: column[5] == &quot;Homo sapiens&quot;</li> <li><code>GO Biological process-pathway</code>: column[5] startswith &quot;REACTOME&quot;, column[8] == &quot;P&quot;, and column[12] == &quot;taxon:9606&quot;</li> <li><code>pathway-GO Cellular component</code>: column[5] startswith &quot;REACTOME&quot;, column[8] == &quot;C&quot;, and column[12] == &quot;taxon:9606&quot;</li> <li><code>GO Molecular function-pathway</code>: column[5] startswith &quot;REACTOME&quot;, column[8] == &quot;F&quot;, and column[12] == &quot;taxon:9606&quot;</li> <li><code>protein-pathway</code>: column[5] == &quot;Homo sapiens&quot;</li> </ul> <p><strong>Files</strong></p> <ul> <li>Downloaded Data <ul> <li>Chemical-Pathway Relations:&nbsp;<a href="https://reactome.org/download/current/ChEBI2Reactome_All_Levels.txt"><code>ChEBI2Reactome_All_Levels.txt</code></a></li> <li>Pathway-GO Relations:&nbsp;<a href="https://reactome.org/download/current/gene_association.reactome"><code>gene_association.reactome</code></a></li> <li>Protein-Pathway Relations:&nbsp;<a href="https://reactome.org/download/current/UniProt2Reactome_All_Levels.txt"><code>UniProt2Reactome_All_Levels.txt</code></a></li> </ul> </li> </ul> <p>&nbsp;</p> <p><strong>Search Tool for Recurring Instances of Neighbouring Genes (STRING) Database</strong></p> <p><strong>Homepage:</strong>&nbsp;<strong><a href="https://string-db.org/"><code>string-db.org</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Szklarczyk D, Gable AL, Lyon D, Junge A, Wyder S, Huerta-Cepas J, Simonovic M, Doncheva NT, Morris JH, Bork P, Jensen LJ.&nbsp;<a href="https://academic.oup.com/nar/article-abstract/47/D1/D607/5198476">STRING v11: protein&ndash;protein association networks with increased coverage, supporting functional discovery in genome-wide experimental datasets</a>. Nucleic Acids Research. 2018;47(D1):D607-13</p> </blockquote> <p><strong>Usage:</strong>&nbsp;The Search Tool for Recurring Instances of Neighbouring Genes (STRING) Database was utilized to create&nbsp;<code>protein-protein</code>&nbsp;edges. The original data is filtered such that only records meeting the following criteria were included:&nbsp;<code>combined_score</code>&nbsp;&gt;= &quot;700&quot; (&gt;90th percentile).</p> <p><strong>Files</strong></p> <ul> <li> <p>Downloaded Data:&nbsp;<a href="https://stringdb-static.org/download/protein.links.v11.0/9606.protein.links.v11.0.txt.gz"><code>9606.protein.links.v11.0.txt.gz</code></a></p> </li> <li> <p>Generated Data: STRING-PRO Identifier Mapping:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/STRING_PRO_ONTOLOGY_MAP.txt"><code>STRING_PRO_ONTOLOGY_MAP.txt</code></a></p> </li> </ul> <p>&nbsp;</p> <p><strong>Universal Protein Resource (UniProt) Knowledgebase</strong></p> <p><strong>Homepage:</strong>&nbsp;<strong><a href="https://www.uniprot.org/"><code>https://www.uniprot.org/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>UniProt Consortium.&nbsp;<a href="https://academic.oup.com/nar/article-abstract/47/D1/D506/5160987">UniProt: a worldwide hub of protein knowledge</a>. Nucleic acids research. 2018;47(D1):D506-15</p> </blockquote> <p><strong>Usage:</strong>&nbsp;The Universal Protein Resource (UniProt) Knowledgebase was utilized to obtain&nbsp;<code>cofactor</code>/<code>catalyst</code>-<code>protein</code>&nbsp;and&nbsp;<code>protein-coding gene</code>-<code>protein</code>&nbsp;edges as well as mappings between&nbsp;<a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#national-center-for-biotechnology-information-gene">NCBI Gene identifiers</a>,&nbsp;<a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#human-genome-organisation-gene-nomenclature-committee">HUGO gene symbols</a>,&nbsp;<a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#universal-protein-resource-knowledgebase">Universal Protein Resource (UniProt) Accession identifiers</a>, and&nbsp;<a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#protein-ontology">Protein Ontology identifiers</a>. For additional details on the processing of these data, see&nbsp;<a href="https://github.com/callahantiff/PheKnowLator/blob/master/notebooks/Data_Preparation.ipynb"><code>Data_Preparation.ipynb</code></a>:</p> <ul> <li>Ensembl Transcript IDs ➞ PRO IDs</li> <li>Gene Ensembl IDs ➞ Entrez Gene IDs</li> <li>Gene Ensembl IDs ➞ PRO IDs</li> <li>Gene Symbols ➞ Transcript Ensembl IDs</li> <li>Entrez Gene IDs ➞ Transcript Ensembl IDs</li> <li>Entrez Gene IDs ➞ PRO IDs</li> <li>Protein Ensembl IDs ➞ UniProt Protein Accession</li> <li>STRING IDs ➞ PRO IDs</li> <li>UniProt Protein Accession ➞ Entrez Gene IDs</li> </ul> <p><strong>Files</strong></p> <ul> <li> <p>Downloaded Data</p> <ul> <li>Cofactor and Catalyst relations:&nbsp;<a href="https://www.uniprot.org/uniprot/?query=&amp;fil=organism%3A%22Homo%20sapiens%20(Human)%20%5B9606%5D%22&amp;columns=id%2Centry%20name%2Creviewed%2Cdatabase(PRO)%2Cchebi(Cofactor)%2Cchebi(Catalytic%20activity)"><code>Cofactor/Catalyst Query Results</code></a></li> <li>UniProt Identifier Mapping:&nbsp;<a href="https://www.uniprot.org/uniprot/?query=&amp;fil=organism%3A%22Homo%20sapiens%20(Human)%20%5B9606%5D%22&amp;columns=id%2Cdatabase(GeneID)%2Cdatabase(Ensembl)%2Cdatabase(HGNC)%2Cgenes(PREFERRED)%2Cgenes(ALTERNATIVE)"><code>UniProt Identifier Query Results</code></a></li> </ul> </li> <li> <p>Generated Data</p> <ul> <li>Merged Gene, RNA, Protein Map:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/Merged_gene_rna_protein_identifiers.pkl"><code>Merged_gene_rna_protein_identifiers.pkl</code></a></li> <li>Protein-Cofactor Relations:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/UNIPROT_PROTEIN_COFACTOR.txt"><code>UNIPROT_PROTEIN_COFACTOR.txt</code></a></li> <li>Protein-Catalyst Relations:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/UNIPROT_PROTEIN_CATALYST.txt"><code>UNIPROT_PROTEIN_CATALYST.txt</code></a></li> <li>UniProt Accession-PRO Identifier Mapping:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/UNIPROT_ACCESSION_PRO_ONTOLOGY_MAP.txt"><code>UNIPROT_ACCESSION_PRO_ONTOLOGY_MAP.txt</code></a></li> <li>UniProt Accession-Entrez Gene Identifier Mapping:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/UNIPROT_ACCESSION_ENTREZ_GENE_MAP.txt"><code>UNIPROT_ACCESSION_ENTREZ_GENE_MAP.txt</code></a></li> </ul> </li> </ul> <p>&nbsp;</p> <p>This project is licensed under Apache License 2.0 - see the&nbsp;<strong><a href="https://github.com/callahantiff/PheKnowLator/blob/master/LICENSE"><code>LICENSE.md</code></a></strong>&nbsp;file for details. If you intend to use any of the information on this Wiki, please provide the appropriate attribution by citing this repository:</p> <pre><code>@misc{callahan_tj_2019_3401437, author = {Callahan, TJ}, title = {PheKnowLator}, month = mar, year = 2019, doi = {10.5281/zenodo.3401437}, url = {https://doi.org/10.5281/zenodo.3401437} }</code></pre>

opencc-by-4.0Apr 2021View details →
zenodo40/100

DrugProt Silver Standard Knowledge Graph

<p><strong>DrugProt Silver Standard Knowledge Graph</strong></p><p>&nbsp;</p><p>&nbsp;</p><p><strong>Please cite if you use any DrugProt resource:</strong></p><p>Antonio Miranda-Escalada, Farrokh Mehryary, Jouni Luoma, Darryl Estrada-Zavala, Luis Gasco, Sampo Pyysalo, Alfonso Valencia, Martin Krallinger, Overview of&nbsp;DrugProt task at BioCreative VII: data and&nbsp;methods for&nbsp;large-scale text mining and&nbsp;knowledge graph generation of&nbsp;heterogenous chemical–protein relations, <i>Database</i>, Volume 2023, 2023, baad080</p><blockquote><p><i>@article{miranda2023overview, &nbsp;title={Overview of DrugProt task at BioCreative VII: data and methods for large-scale text mining and knowledge graph generation of heterogenous chemical--protein relations}, &nbsp;author={Miranda-Escalada, Antonio and Mehryary, Farrokh and Luoma, Jouni and Estrada-Zavala, Darryl and Gasco, Luis and Pyysalo, Sampo and Valencia, Alfonso and Krallinger, Martin}, &nbsp;journal={Database}, &nbsp;volume={2023}, &nbsp;pages={baad080}, &nbsp;year={2023}, &nbsp;publisher={Oxford University Press UK} }</i></p></blockquote><p>Miranda, Antonio, et al. "Overview of DrugProt BioCreative VII track: quality evaluation and large scale text mining of drug-gene/protein relations."&nbsp;<i>Proceedings of the seventh BioCreative challenge evaluation workshop</i>. 2021.</p><blockquote><p><i>@inproceedings{miranda2021overview, &nbsp;title={Overview of DrugProt BioCreative VII track: quality evaluation and large scale text mining of drug-gene/protein relations}, &nbsp;author={Miranda, Antonio and Mehryary, Farrokh and Luoma, Jouni and Pyysalo, Sampo and Valencia, Alfonso and Krallinger, Martin}, &nbsp;booktitle={Proceedings of the seventh BioCreative challenge evaluation workshop}, &nbsp;year={2021} }</i></p></blockquote><p><strong>Description</strong></p><p>&nbsp;</p><p>&nbsp;</p><p><strong>Files:</strong></p><ul><li>drugprot-silver-standard-kg.zip : JSON files with the relations predicted by the DrugProt systems and their precision</li><li>large_scale_network_abstracts.tsv : PubMed abstracts</li><li>large_scale_network_entities.tsv : CHEMICAL/drug and GENE/protein&nbsp;entities predicted by DrugProt NER Taggers</li><li>large_scale_network_pmids.txt : list of PMIDs</li></ul><p>&nbsp;</p><p><strong>Related resources:</strong></p><ul><li><a href="https://biocreative.bioinformatics.udel.edu/tasks/biocreative-vii/track-1/">Web</a></li><li><a href="https://doi.org/10.5281/zenodo.4955410">DrugProt corpus</a></li><li><a href="https://github.com/tonifuc3m/drugprot-evaluation-library">Evaluation library</a></li><li><a href="https://codalab.lisn.upsaclay.fr/competitions/8293">Online evaluation (CodaLab)</a></li><li><a href="https://doi.org/10.5281/zenodo.4957137">Relation annotation guidelines</a></li><li><a href="https://doi.org/10.5281/zenodo.4957576">Gene and protein annotation guidelines</a></li><li><a href="https://doi.org/10.5281/zenodo.4957518">Chemicals and drugs annotation guidelines</a></li><li><a href="https://doi.org/10.5281/zenodo.5042178">FAQ</a></li><li><a href="https://doi.org/10.5281/zenodo.5119878">DrugProt Large Scale Additional SubTrack</a></li><li><a href="https://doi.org/10.5281/zenodo.5656991">DrugProt Large Scale document collection protocol</a></li></ul>

opencc-by-4.0Oct 2022View details →
zenodo40/100

Helmholtz Knowledge Graph: RDF data dump

<p>Under this base DOI we regularly publish full RDF data dumps of the Helmholtz-Knowledge Graph.<br>Data dumps are typical associated with major releases, or major data updates.</p> <p>Dumps are serialized in .ttl and compressed with gzip.<br><br>For more information on deployment and documentation as well as data access see:<br>Search UI: https://search.unhide.helmholtz-metadaten.de/<br>SPARQL endpoint: &nbsp;https://sparql.unhide.helmholtz-metadaten.de/<br>Documentation: https://docs.unhide.helmholtz-metadaten.de/<br>Software: https://codebase.helmholtz.cloud/hmc/hmc-public/unhide</p>

opencc-by-4.0Apr 2024View details →
zenodo40/100

NASA GES-DISC Knowledge Graph for Link Prediction

<p>This dataset includes a knowledge graph of NASA GES-DISC collections, featuring interconnected nodes for datasets, data center, projects, platforms, instruments, science keywords, and publications. Designed for link prediction tasks, it aids machine learning research in satellite observation, remote sensing, and climate change. The dataset is in CSV format, ready for graph databases and ML frameworks.</p> <p>&nbsp;</p>

opencc-by-4.0Jun 2024View details →
zenodo40/100

TAXREF-LD: Knowledge Graph of the French taxonomic registry

<p>TAXREF-LD is a Linked Data knowledge graph representing <a href="https://inpn.mnhn.fr/programme/referentiel-taxonomique-taxref?lg=en">TAXREF</a>, the French national taxonomical register for fauna, flora and fungus, that covers mainland France and overseas territories.</p> <p>TAXREF-LD is a joint initiative of the <a href="http://www.patrinat.fr/">UMS Patrinat</a> of the <a href="http://www.mnhn.fr/">National Museum of Natural History</a>, and the <a href="http://www.i3s.unice.fr/">I3S laboratory</a>, <a href="https://univ-cotedazur.fr">University C&ocirc;te d&#39;Azur</a>, <a href="https://www.inria.fr">Inria</a>, <a href="https://www.cnrs.fr">CNRS</a>.</p> <p>Homepage: https://github.com/frmichel/taxref-ld/</p>

opencc-by-4.0Jan 2022View details →
zenodo40/100

BRAIN Journal-Right-Linear Languages Generated in Systems of Knowledge Representation based on LSG-Right-Figure 3. Representation of the grammar G1 in the labelled graph G0 1

<p>If we take the labeled graph G0 1 given in Figure 3 and construct the stratified graph structure over (99) such that (100) we obtain (101), (102).&nbsp;&nbsp;</p> <p>In this paper, we proposed a new system for formal language generation by means of stratified graphs structures. This mechanism can generate languages of the first type and of the second type. More precisely, we propose a new system for formal language generation by means of a system of knowledge based on stratified graphs. We exemplified that, using an interpretation&nbsp; system specially defined for stratified graphs representations, a particular formal language can be obtained by means of the resulted accepted structured paths.</p>

opencc-by-4.0Apr 2017View details →
zenodo40/100

Ozymandias: A biodiversity knowledge graph

<p>Triples for the knowledge graph described in &quot;Ozymandias: A biodiversity knowledge graph&quot;&nbsp;<a href="https://doi.org/10.7717/peerj.6739">doi:10.7717/peerj.6739</a>&nbsp;Each set of triples corresponds to a named graph in the triple store:</p> <blockquote> <p>oz-ala.nt &nbsp; &nbsp; &nbsp; &nbsp; &lt;https://bie.ala.org.au&gt;<br> oz-publication.nt &lt;https://biodiversity.org.au/afd/publication&gt;<br> oz-zenodo.nt &nbsp; &nbsp; &nbsp;&lt;https://zenodo.org&gt;<br> oz-crossref.nt &nbsp; &nbsp;&lt;https://crossref.org&gt;<br> oz-orcid.nt &nbsp; &nbsp; &nbsp; &lt;https://orcid.org&gt;<br> oz-species.nt &nbsp; &nbsp; &lt;https://species.wikimedia.org&gt;<br> oz-gbif.nt &nbsp; &nbsp; &nbsp; &nbsp;&lt;https://gbif.org/species&gt;<br> oz-bold.nt &nbsp; &nbsp; &nbsp; &nbsp;&lt;http://boldsystems.org&gt;</p> </blockquote> <pre>&nbsp;</pre> <pre>&nbsp;</pre>

opencc-by-4.0Apr 2019View details →
zenodo40/100

ArCo Knowledge Graph

<p>The ArCo knowledge graph contains the ontology network and the data about the cultural properties catalogued&nbsp;by the Italian Institute of the General Catalogue and Documentation.</p> <p>Data are represented with RDF and&nbsp;by using&nbsp;N-Triples as syntax.</p> <p>The ontologies of the network are modelled with OWL 2 and serialised with the RDF/XML syntax.</p> <p>The ontology network is released along with alignments to other ontologies/vocabularies in the Semantic Web. Those alignments are provided within separate OWL files.</p> <p>The data are contained into a single RDF dump serialised as N-TRIPLES.</p> <p>Additionally, the release provides the links between ArCO entities and other entities published in other datasets in the Linked Open Data cloud. Such links are represented by using owl:sameAs axioms and serialised as N-TRIPLES into a separate file.</p>

opencc-by-sa-4.0Apr 2019View details →
zenodo40/100

Integrated knowledge graphs and embeddings vectors for drug-drug interaction prediction

<p>The associated Knowledge Graphs for predicting potential drug-drug interaction, which is used in our paper titled &quot;Drug-Drug Interaction Prediction Based on Knowledge Graph Embeddings and Convolutional-LSTM Network&quot;.&nbsp; Please consider citing the following paper if you plan or used our datasets.</p> <p>Md. Rezaul Karim, Michael Cochez, Joao Bosco Jares, Mamtaz Uddin, Oya Beyan, and Stefan Decker, &quot;Drug-Drug Interaction Prediction Based on Knowledge Graph Embeddings and Convolutional-LSTM Network&quot;, In 10th ACM Int&rsquo;l Conference on Bioinformatics, Computational Biology and Health Informatics (ACM-BCB &rsquo;19), September 7&ndash;10, 2019, Niagara Falls, NY, USA.</p>

opencc-by-4.0Sep 2019View details →
zenodo40/100

FAIR Jupyter Knowledge Graph

<p><a href="https://w3id.org/fairjupyter/">FAIR Jupyter</a> is a knowledge graph representation of <a href="https://doi.org/10.5281/zenodo.8226725">Dataset of a Study of Computational reproducibility of Jupyter notebooks from biomedical publications</a>.&nbsp;<br>This repository contains the data, mapping, and the RDF files for each entity represented in the FAIR Jupyter KG.</p> <p>Folder structure:</p> <p>data: contains the CSV files that are exported for each entity type from the <a href="https://doi.org/10.5281/zenodo.8226725">original sqlite database</a>.</p> <p>mapping: contains the RML and YARRML mapping files for each entity type of FAIR Jupyter KG.</p> <p>kg: contains the RDF files created in N-Triples format for each entity type of FAIR Jupyter KG.</p> <p>The code for creating the FAIR Jupyter Knowledge graph is available here: &nbsp;https://github.com/fusion-jena/fairjupyter</p>

opencc-by-4.0Sep 2024View details →
zenodo40/100

EB_HQ: Knowledge Graph of the First Eight Editions of the Encyclopaedia Britannica (1768-1860) Following the Heritage Textual Ontology

<p>This EB_HQ Knowledge Graph, represents information from the first eight editions of the <em>Encyclopaedia Britannica</em>&nbsp;(1768-1860), structured according to the&nbsp;<a href="https://eur02.safelinks.protection.outlook.com/?url=http%3A%2F%2Fw3id.org%2Fhto&amp;data=05%7C02%7C%7C9596c1d96d154755723e08dce92a686d%7C2e9f06b016694589878910a06934dc61%7C0%7C0%7C638641615571225885%7CUnknown%7CTWFpbGZsb3d8eyJWIjoiMC4wLjAwMDAiLCJQIjoiV2luMzIiLCJBTiI6Ik1haWwiLCJXVCI6Mn0%3D%7C0%7C%7C%7C&amp;sdata=vr0NRZPBmtKGFCs2gliaBKJ6K%2FCCrwEEv%2FlrWfNLRpU%3D&amp;reserved=0">Heritage Textual Ontology&nbsp;</a>(HTO). This version enhances the information from the previously developed <a href="https://eur02.safelinks.protection.outlook.com/?url=https%3A%2F%2Fzenodo.org%2Frecords%2F6673897&amp;data=05%7C02%7C%7C9596c1d96d154755723e08dce92a686d%7C2e9f06b016694589878910a06934dc61%7C0%7C0%7C638641615571238085%7CUnknown%7CTWFpbGZsb3d8eyJWIjoiMC4wLjAwMDAiLCJQIjoiV2luMzIiLCJBTiI6Ik1haWwiLCJXVCI6Mn0%3D%7C0%7C%7C%7C&amp;sdata=HHP1VKXayraELBzHQEhpya01JM1vGV6YY7Zrgz8rhkg%3D&amp;reserved=0">EB-KG</a>&nbsp;by unifying data across various sources (such as&nbsp;<a href="https://eur02.safelinks.protection.outlook.com/?url=https%3A%2F%2Fdata.nls.uk%2Fdata%2Fdigitised-collections%2Fencyclopaedia-britannica%2F&amp;data=05%7C02%7C%7C9596c1d96d154755723e08dce92a686d%7C2e9f06b016694589878910a06934dc61%7C0%7C0%7C638641615571251754%7CUnknown%7CTWFpbGZsb3d8eyJWIjoiMC4wLjAwMDAiLCJQIjoiV2luMzIiLCJBTiI6Ik1haWwiLCJXVCI6Mn0%3D%7C0%7C%7C%7C&amp;sdata=4pRrz54ZNf4gPOH20g4apwPxbl3LHknXpbUyGssV7R0%3D&amp;reserved=0">National Library of Scotland</a>, or&nbsp;<a href="https://eur02.safelinks.protection.outlook.com/?url=https%3A%2F%2Ftu-plogan.github.io%2Fsource%2Fr_7th_edition.html&amp;data=05%7C02%7C%7C9596c1d96d154755723e08dce92a686d%7C2e9f06b016694589878910a06934dc61%7C0%7C0%7C638641615571265383%7CUnknown%7CTWFpbGZsb3d8eyJWIjoiMC4wLjAwMDAiLCJQIjoiV2luMzIiLCJBTiI6Ik1haWwiLCJXVCI6Mn0%3D%7C0%7C%7C%7C&amp;sdata=aoDV%2BSnXJ%2B3yzI7rsP%2BCdhAXERsu8WVUSA2xtbw6vHE%3D&amp;reserved=0">&nbsp;Nineteenth-Century Knowledge Project</a>&nbsp;), editions, allowing for more comprehensive tracking of the evolution of concepts over time. For each edition, the highest-quality text source is selected to ensure optimal content. It integrates advanced information extraction methods, deep-learning-based knowledge enrichment, facilitating richer analyses.&nbsp;</p> <p>The EB_HQ captures 3316459 RDF triples, providing structured metadata and descriptions for each edition, volume, and term. It unifies information across various editions, allowing seamless tracking of the evolution of concepts over time. Data enrichment includes semantic linkages to external knowledge bases like DBpedia and Wikidata, facilitating broader analysis and connectivity to contemporary information.</p> <p>This dataset supports historical research, offering rich semantic data for researchers exploring the evolution of knowledge and concepts in the&nbsp;<em>Encyclopaedia Britannica</em>. It features terms categorized as&nbsp;<em>Articles</em>&nbsp;or&nbsp;<em>Topics</em>, each with detailed metadata extracted from METS and ALTO XML files.&nbsp;</p>

opencc-by-4.0Oct 2024View details →
zenodo40/100

EB_Composite: Knowledge Graph of the First Eight Editions of the Encyclopaedia Britannica (1768-1860) Following the Heritage Textual Ontology

<p>The EB_Composite Knowledge Graph represents information from the first eight editions of the <em>Encyclopaedia Britannica</em> (1768&ndash;1860), structured using the <a href="https://w3id.org/hto"><em>Heritage Textual Ontology</em></a> (HTO). It extends our previously developed <a href="https://zenodo.org/records/13919115">EB_HQ</a> by incorporating multiple text sources, such as the <a href="https://data.nls.uk/data/digitised-collections/encyclopaedia-britannica/">National Library of Scotland</a> and the <a href="https://tu-plogan.github.io/source/r_7th_edition.html">Nineteenth-Century Knowledge Project</a>, for each edition. A particular source is added, comprising of post-corrected textual content generated using deep-learning-based OCR error correction methods. These sources presents different levels of text quality, and EB_Composite enables researchers to compare these texts, and track how they are digitised or extracted from these sources.</p> <p>The EB_Composite captures 4150776 RDF triples. Same as EB_HQ, it provides structured metadata and descriptions for each edition, volume, and term. By integrating information across different editions, it enables smooth tracking of concept evolution over time. Additionally, the dataset features semantic connections to external knowledge bases such as DBpedia and Wikidata, enhancing links to modern information and supporting more comprehensive analyses.</p> <p>Designed to support historical research, this dataset offers rich semantic data for exploring the development of knowledge and concepts in the <em>Encyclopaedia Britannica</em>. It categorizes terms as either <em>Articles</em> or <em>Topics</em>, each with detailed metadata extracted from METS and ALTO XML files. OCR errors common in historical texts have been mitigated using deep-learning-based corrections.</p>

opencc-by-4.0Oct 2024View details →
zenodo40/100

Knowledge graph for air quality simulations in the Region of Murcia from 2024-12-26 to 2024-12-29

<p>This resource provides the knowledge graphs with the information of the simulation done from 2023-12-26 to 2023-12-26 in the Region of Murcia with the CHIMERE-WRF model (https://www.lmd.polytechnique.fr/chimere/) and transformed into knowledge graph thanks to the SWIT framework (http://sele.inf.um.es/swit/about.html). It also links this data to the github repository including the code and the ontology generated in the work "Representation of&nbsp; chemistry transport models simulations using knowledge graphs".</p>

opencc-by-4.0Oct 2024View details →
zenodo40/100

Detecting Cross-Language Plagiarism using Open Knowledge Graphs

<p>Corresponding authors: <a href="mailto:meuschke@uni-wuppertal.de?subject=Inquiry%20about%20CL-OSA%20dataset">Norman Meuschke</a>, <a href="mailto:ruas@uni-wuppertal.de?subject=Inquiry%20about%20CL-OSA%20dataset">Terry Ruas</a><br> Venue: 2nd Workshop on Extraction and Evaluation of Knowledge Entities from Scientific Documents (EEKE2021)<br> at the ACM/IEEE Joint Conference on Digital Libraries 2021 (JCDL2021)</p> <p>==========================================================================</p> <p><strong>Source code:&nbsp;<a href="https://github.com/ag-gipp/cl-osa">https://github.com/ag-gipp/cl-osa</a>&nbsp;</strong></p> <p>==========================================================================</p> <p><strong>Dataset Details</strong></p> <p><em><a href="https://jipsti.jst.go.jp/aspec/">ASPEC</a></em>. The Asian Scientific Paper Excerpt Corpus comprises&nbsp;excepts of scientific papers in Japanese that have been manually translated to English and Chinese. We use both subsets of&nbsp;the ASPEC corpus.&nbsp;</p> <ul> <li><em>ASPEC-JC</em><strong>&nbsp;</strong>contains abstracts and paragraphs from the main text of research papers that were translated manually from Japanese&nbsp;to Chinese.</li> <li><em>ASPEC-JE</em> contains abstracts of approx. two million research papers that were translated manually from Japanese to English. &nbsp;</li> </ul> <p><em><a href="https://ec.europa.eu/jrc/en/language-technologies/jrc-acquis">JRC-Acquis</a></em>. The corpus consists of legislative texts in 22 languages, which the European Union&#39;s Joint Research Centre (JRC) selected from the cumulative body of EU laws (the so called Acquis communautaire). We sampled our test cases from the 10,000 document pairs in the English-French&nbsp;subset of the corpus.</p> <p><em><a href="https://www.statmt.org/europarl/">Europarl</a></em>. The corpus contains transcripts of European Parliament proceedings in 21 European languages. We exclusively sampled test cases from the 9,443 document pairs in the English-French&nbsp;subset of the corpus.</p> <p><em><a href="https://zenodo.org/record/3250095#.YOr25jqxVH4">PAN-PC-11</a></em>.&nbsp;The corpus contains instances of simulated monolingual and cross-language plagiarism that were used for evaluating plagiarism detection methods as part of the workshop series <a href="https://pan.webis.de/">Plagiarism Analysis, Authorship Identification, and Near-Duplicate Detection (PAN)</a>. Most of the 26,939 documents in the corpus were created by extracting text from openly available books. The documents are partially interspersed with instances of simulated plagiarism that were created and obfuscated automatically or by crowdsourced workers. We exclusively sampled test cases from the 2,921 Spanish-English aligned document pairs in the corpus,&nbsp;for which simulated plagiarism instances were either machine-generated or created manually by crowdsourced workers.</p> <ul> </ul> <p>==========================================================================</p> <p><strong>File Structure</strong></p> <p><strong>[corpus_documents] folder</strong>: Corpora of translation-aligned documents used in our experiments composed of:&nbsp;</p> <ul> <li>aspec: Japanese and English</li> <li>aspecx: Japanese and Chinese&nbsp;</li> <li>jrc: English and French&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp;</li> <li>europarl: English and French</li> <li>pan: English and Spanish</li> </ul> <p>Each sub-corpus consists of&nbsp;4,000 translation-aligned files (2,000 per language); the entire corpus has thus 20,000 files.<br> Each set of translation-aligned documents was randomly selected from the original datasets (details in the paper).<br> The Japanese files in aspec and aspecx do not necessarily overlap even though they are from the same dataset.</p> <p><br> <strong>[vectors_documents] folder</strong>: Average vector representation of the documents in the datasets from two pre-trained models:</p> <ul> <li>Universal Sentence Encoder - Multilingual (USE-ML)</li> <li>ConceptNet Numberbatch</li> </ul> <p>&nbsp;</p> <p><strong>Naming convention</strong>: &lt;model&gt;_&lt;dataset&gt;_&lt;language&gt;;</p> <ul> <li>Example: cn_jrc_es: <ul> <li>model: ConceptNet Numberbatch</li> <li>corpus: JRC-Acquis</li> <li>language: Spanish</li> </ul> </li> </ul> <ul> <li>Labels: <ul> <li>&lt;model&gt;:<br> cn - ConceptNet Numberbatch<br> um - USE-ML<br> &nbsp;</li> <li>&lt;dataset&gt;<br> aspec - ASPEC (Asian Scientific Paper Excerpt Corpus) - English and Japanese<br> aspecx - ASPEC (Asian Scientific Paper Excerpt Corpus) - Japanese and Chinese<br> jrc - JRC-Acquis<br> europarl - Europarl<br> pan - PAN-PC-11<br> &nbsp;</li> <li>&lt;language&gt;<br> en - English<br> es - Spanish<br> fr - French<br> ja - Japanese<br> zh - Chinese</li> </ul> </li> </ul>

opencc-by-4.0Jul 2020View details →
zenodo40/100

Description of the Features of the Open Research Knowledge Graph as a Crowdsourcing Platform based on the 4 Pillars of Crowdsourcing

<p>This dataset provides the description of the features of the <a href="https://www.orkg.org/orkg/">Open Research Knowledge Graph</a> (ORKG) based on the 4 pillars of crowdsourcing according to the reference model for crowdsourcing by Hosseini et al. [1]. This overview represents the features of the current implementation status of&nbsp; ORKG as a crowdsourcing platform.</p> <p>[1] M. Hosseini, K. Phalp, J. Taylor, and R. Ali, &quot;<a href="https://ieeexplore.ieee.org/abstract/document/6861072?casa_token=zWTHNBHeH6kAAAAA:1n3EJjajqgSkSk154g4DlNFAmJs_7o3KY7LnobvP_W7AUIr-FvM5OyGP1FaRn68zUX-2oYHh7A">The Four Pillars of Crowdsourcing: A Reference Model</a>&quot;, in 2014 IEEE 8th International Conference on Research Challenges in Information Science (RCIS). IEEE, 2014, pp. 1&ndash;12.</p>

opencc-by-4.0Jun 2021View details →
zenodo40/100

Not Scared of Chemistry Knowledge Graph

<p>A combination of ExCAPE-DB, BioGRID, HomoloGene, and chemical similarities in a knowledge graph. Original data uncer CC-0 1.0 and derived data redistributed with their original licenses:</p> <pre>- EXCAPE - CC-BY-SA-4.0 License - BioGRID - MIT License - HomoloGene - ??</pre> <p>Automatically generated by https://github.com/cthoyt/nsockg.</p>

opencc-zeroMar 2021View details →
zenodo40/100

Wikibase knowledge graphs

<p>A collection of open source tools,&nbsp;resources and conferences related to Wikibase knowledge graphs (https://github.com/shigapov/wikibase-knowledge-graphs)</p>

opencc-zeroDec 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record