Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

6,157

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

6,157 results for “knowledge”

Learn how ShareScore rates datasets ↗
zenodo44/100

PheKnowLator Human Disease Knowledge Graphs - Build Data (Processed)

<p><strong>RELEASE V2.1.0 KNOWLEDGE GRAPH: PROCESSED DATA SOURCES&nbsp;</strong></p> <p><strong>Release:</strong>&nbsp;<a href="https://github.com/callahantiff/PheKnowLator/wiki/v2.0.0">v2.1.0&nbsp;</a></p> <p>The goal of this build was to create a knowledge graph that represented human disease mechanisms and included the central dogma. The data sources utilized in this release include many of the sources used in the initial release, as well as some new data made available by the&nbsp;<a href="https://ctdbase.org/">Comparative Toxicogenomics Database</a>&nbsp;and experimental data from the&nbsp;<a href="https://www.proteinatlas.org/">Human Protein Atlas</a>.</p> <p>Data sources are listed by type (Ontology and Data not represented in an ontology [Database Sources]). Additional details are provided for each data source below.&nbsp;Please see documentation on the primary release&nbsp;(<a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources">https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources</a>) for additional details on each data source as well as citation information.</p> <p><strong>Data Access:</strong></p> <ul> <li><a href="https://console.cloud.google.com/storage/browser/pheknowlator/archived_builds/release_v2.1.0/build_01MAY2021?project=pheknowlator">https://console.cloud.google.com/storage/browser/pheknowlator/archived_builds/release_v2.1.0/build_01MAY2021</a></li> </ul> <p>&nbsp;</p> <p><strong>ONTOLOGIES</strong></p> <ul> <li>Cell Ontology</li> <li>Cell Line Ontology</li> <li>Chemical Entities of Biological Interest (ChEBI) Ontology</li> <li>Gene Ontology</li> <li>Human Phenotype Ontology</li> <li>Mondo Disease Ontology</li> <li>Pathway Ontology</li> <li>Protein Ontology</li> <li>Relations Ontology</li> <li>Sequence Ontology</li> <li>Uber-Anatomy Ontology</li> <li>Vaccine Ontology</li> </ul> <p>&nbsp;</p> <p><strong>Cell Ontology (CL)</strong></p> <p><strong>Homepage:</strong>&nbsp;<strong><a href="https://github.com/obophenotype/cell-ontology"><code>GitHub</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Bard J, Rhee SY, Ashburner M.&nbsp;<a href="https://genomebiology.biomedcentral.com/articles/10.1186/gb-2005-6-2-r21">An ontology for cell types</a>. Genome Biology. 2005;6(2):R21</p> </blockquote> <p><strong>Usage:</strong>&nbsp;Utilized to connect&nbsp;<code>transcripts</code>&nbsp;and&nbsp;<code>proteins</code>&nbsp;to&nbsp;<code>cells</code>. Additionally, the edges between this ontology and its dependencies are utilized:</p> <ul> <li><strong><a href="https://www.ebi.ac.uk/chebi/"><code>ChEBI</code></a></strong></li> <li><strong><a href="http://geneontology.org/"><code>GO</code></a></strong></li> <li><strong><a href="https://github.com/pato-ontology/pato/"><code>PATO</code></a></strong></li> <li><strong><a href="https://proconsortium.org/"><code>PRO</code></a></strong></li> <li><strong><a href="https://github.com/oborel/obo-relations/"><code>RO</code></a></strong></li> <li><strong><a href="https://uberon.github.io/"><code>UBERON</code></a></strong></li> </ul> <p>&nbsp;</p> <p><strong>Cell Line Ontology (CLO)</strong></p> <p><strong>Homepage:</strong>&nbsp;<strong><a href="http://www.clo-ontology.org/"><code>http://www.clo-ontology.org/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Sarntivijai S, Lin Y, Xiang Z, Meehan TF, Diehl AD, Vempati UD, Sch&uuml;rer SC, Pang C, Malone J, Parkinson H, Liu Y.&nbsp;<a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4387853/">CLO: the cell line ontology</a>. Journal of Biomedical Semantics. 2014;5(1):37</p> </blockquote> <p><strong>Usage:</strong>&nbsp;Utilized this ontology to map&nbsp;<code>cell lines</code>&nbsp;to&nbsp;<code>transcripts</code>&nbsp;and&nbsp;<code>proteins</code>. Additionally, the edges between this ontology and its dependencies are utilized:</p> <ul> <li><strong><a href="http://www.obofoundry.org/ontology/cl.html"><code>CL</code></a></strong></li> <li><strong><a href="http://disease-ontology.org/"><code>DOID</code></a></strong></li> <li><strong><a href="https://github.com/obophenotype/ncbitaxon"><code>NCBITaxon</code></a></strong></li> <li><strong><a href="https://uberon.github.io/"><code>UBERON</code></a></strong></li> </ul> <p>&nbsp;</p> <p><strong>Chemical Entities of Biological Interest (ChEBI)</strong></p> <p><strong>Homepage:</strong>&nbsp;<strong><a href="https://www.ebi.ac.uk/chebi/"><code>https://www.ebi.ac.uk/chebi/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Hastings J, Owen G, Dekker A, Ennis M, Kale N, Muthukrishnan V, Turner S, Swainston N, Mendes P, Steinbeck C.&nbsp;<a href="https://academic.oup.com/nar/article-abstract/44/D1/D1214/2502583">ChEBI in 2016: Improved services and an expanding collection of metabolites</a>. Nucleic Acids Research. 2015;44(D1):D1214-9</p> </blockquote> <p><strong>Usage:</strong>&nbsp;Utilized to connect&nbsp;<code>chemicals</code>&nbsp;to&nbsp;<code>complexes</code>,&nbsp;<code>diseases</code>,&nbsp;<code>genes</code>,&nbsp;<code>GO biological processes</code>,&nbsp;<code>GO cellular components</code>,&nbsp;<code>GO molecular functions</code>,&nbsp;<code>pathways</code>,&nbsp;<code>phenotypes</code>,&nbsp;<code>reactions</code>, and&nbsp;<code>transcripts</code>.</p> <p>&nbsp;</p> <p><strong>Gene Ontology (GO)</strong></p> <p><strong>Homepage:</strong>&nbsp;<strong><a href="http://geneontology.org/"><code>http://geneontology.org/</code></a></strong><br> <strong>Citations:</strong></p> <blockquote> <p>Ashburner M, Ball CA, Blake JA, Botstein D, Butler H, Cherry JM, Davis AP, Dolinski K, Dwight SS, Eppig JT, Harris MA.&nbsp;<a href="https://www.nature.com/articles/ng0500_25">Gene ontology: tool for the unification of biology</a>. Nature Genetics. 2000;25(1):25</p> <p>The Gene Ontology Consortium.&nbsp;<a href="https://academic.oup.com/nar/article/47/D1/D330/5160994">The Gene Ontology Resource: 20 years and still GOing strong</a>. Nucleic Acids Research. 2018;47(D1):D330-8</p> </blockquote> <p><strong>Usage:</strong>&nbsp;Utilized to connect&nbsp;<code>biological processes</code>,&nbsp;<code>cellular components</code>, and&nbsp;<code>molecular functions</code>&nbsp;to&nbsp;<code>chemicals</code>,&nbsp;<code>pathways</code>, and&nbsp;<code>proteins</code>. Additionally, the edges between this ontology and its dependencies are utilized:</p> <ul> <li><strong><a href="http://www.obofoundry.org/ontology/cl.html"><code>CL</code></a></strong></li> <li><strong><a href="https://github.com/obophenotype/ncbitaxon"><code>NCBITaxon</code></a></strong></li> <li><strong><a href="https://github.com/oborel/obo-relations/"><code>RO</code></a></strong></li> <li><strong><a href="https://uberon.github.io/"><code>UBERON</code></a></strong></li> </ul> <p><strong>Other Gene Ontology Data Used:</strong>&nbsp;<a href="http://geneontology.org/gene-associations/goa_human.gaf.gz"><code>goa_human.gaf.gz</code></a></p> <p>&nbsp;</p> <p><strong>Human Phenotype Ontology (HPO)</strong></p> <p><strong>Homepage:</strong>&nbsp;<strong><a href="https://hpo.jax.org/"><code>https://hpo.jax.org/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>K&ouml;hler S, Carmody L, Vasilevsky N, Jacobsen JO, Danis D, Gourdine JP, Gargano M, Harris NL, Matentzoglu N, McMurry JA, Osumi-Sutherland D.&nbsp;<a href="https://academic.oup.com/nar/article-abstract/47/D1/D1018/5198478">Expansion of the Human Phenotype Ontology (HPO) knowledge base and resources</a>. Nucleic Acids Research. 2018;47(D1):D1018-27</p> </blockquote> <p><strong>Usage:</strong>&nbsp;Utilized to connect&nbsp;<code>phenotypes</code>&nbsp;to&nbsp;<code>chemicals</code>,&nbsp;<code>diseases</code>,&nbsp;<code>genes</code>, and&nbsp;<code>variants</code>. Additionally, the edges between this ontology and its dependencies are utilized:</p> <ul> <li><strong><a href="http://www.obofoundry.org/ontology/cl.html"><code>CL</code></a></strong></li> <li><strong><a href="https://www.ebi.ac.uk/chebi/"><code>ChEBI</code></a></strong></li> <li><strong><a href="http://geneontology.org/"><code>GO</code></a></strong></li> <li><strong><a href="https://uberon.github.io/"><code>UBERON</code></a></strong></li> </ul> <p><strong>Files</strong></p> <ul> <li>Other Human Phenotype Ontology Data Used:&nbsp;<a href="http://purl.obolibrary.org/obo/hp/hpoa/phenotype.hpoa"><code>phenotype.hpoa</code></a></li> </ul> <p>&nbsp;</p> <p><strong>Mondo Disease Ontology (Mondo)</strong></p> <p><strong>Homepage:</strong>&nbsp;<strong><a href="https://mondo.monarchinitiative.org/"><code>https://mondo.monarchinitiative.org/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Mungall CJ, McMurry JA, K&ouml;hler S, Balhoff JP, Borromeo C, Brush M, Carbon S, Conlin T, Dunn N, Engelstad M, Foster E.&nbsp;<a href="https://academic.oup.com/nar/article-abstract/45/D1/D712/2605791">The Monarch Initiative: an integrative data and analytic platform connecting phenotypes to genotypes across species</a>. Nucleic Acids Research. 2017;45(D1):D712-22</p> </blockquote> <p><strong>Usage:</strong>&nbsp;Utilized to connect&nbsp;<code>diseases</code>&nbsp;to&nbsp;<code>chemicals</code>,&nbsp;<code>phenotypes</code>,&nbsp;<code>genes</code>, and&nbsp;<code>variants</code>. Additionally, the edges between this ontology and its dependencies are utilized:</p> <ul> <li><strong><a href="http://www.obofoundry.org/ontology/cl.html"><code>CL</code></a></strong></li> <li><strong><a href="https://www.ncbi.nlm.nih.gov/taxonomy"><code>NCBITaxon</code></a></strong></li> <li><strong><a href="http://geneontology.org/"><code>GO</code></a></strong></li> <li><strong><a href="https://hpo.jax.org/"><code>HPO</code></a></strong></li> <li><strong><a href="https://uberon.github.io/"><code>UBERON</code></a></strong></li> </ul> <p>&nbsp;</p> <p><strong>Pathway Ontology (PW)</strong></p> <p><strong>Homepage:</strong>&nbsp;<strong><a href="https://rgd.mcw.edu/wg/home/pathway2/"><code>rgd.mcw.edu</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Petri V, Jayaraman P, Tutaj M, Hayman GT, Smith JR, De Pons J, Laulederkind SJ, Lowry TF, Nigam R, Wang SJ, Shimoyama M.&nbsp;<a href="https://www.ncbi.nlm.nih.gov/pubmed/24499703">The pathway ontology&ndash;updates and applications</a>. Journal of Biomedical Semantics. 2014;5(1):7.</p> </blockquote> <p><strong>Usage:</strong>&nbsp;Utilized to connect&nbsp;<code>pathways</code>&nbsp;to&nbsp;<code>GO biological processes</code>,&nbsp;<code>GO cellular components</code>,&nbsp;<code>GO molecular functions</code>,&nbsp;<code>Reactome pathways</code>. Several steps are taken in order to connect&nbsp;<code>Pathway Ontology</code>&nbsp;identifiers to&nbsp;<code>Reactome</code>&nbsp;pathways and&nbsp;<code>GO biological processes</code>. To connect&nbsp;<code>Pathway Ontology</code>&nbsp;identifiers to&nbsp;<code>Reactome</code>&nbsp;pathways, we use&nbsp;<a href="https://github.com/ComPath/resources/tree/master/mappings">ComPath Pathway Database Mappings</a>&nbsp;developed by Daniel Domingo-Fern&aacute;ndez (<a href="https://www.ncbi.nlm.nih.gov/pubmed/30564458">PMID:30564458</a>).</p> <p><strong>Files</strong></p> <ul> <li>Downloaded Mapping Data <ul> <li><a href="http://compath.scai.fraunhofer.de/export_mappings"><code>curated_mappings.txt</code></a></li> <li><a href="https://github.com/ComPath/resources/blob/master/mappings/kegg_reactome.csv"><code>kegg_reactome.csv</code></a></li> </ul> </li> <li>Generated Mapping Data <ul> <li><a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/REACTOME_PW_GO_MAPPINGS.txt"><code>REACTOME_PW_GO_MAPPINGS.txt</code></a></li> </ul> </li> </ul> <p>&nbsp;</p> <p><strong>Protein Ontology (PRO)</strong></p> <p><strong>Homepage:</strong>&nbsp;<strong><a href="https://proconsortium.org/"><code>https://proconsortium.org/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Natale DA, Arighi CN, Barker WC, Blake JA, Bult CJ, Caudy M, Drabkin HJ, D&rsquo;Eustachio P, Evsikov AV, Huang H, Nchoutmboube J.&nbsp;<a href="https://academic.oup.com/nar/article-abstract/39/suppl_1/D539/2508558">The Protein Ontology: a structured representation of protein forms and complexes</a>. Nucleic Acids Research. 2010;39(suppl_1):D539-45</p> </blockquote> <p><strong>Usage:</strong>&nbsp;Utilized to connect&nbsp;<code>proteins</code>&nbsp;to&nbsp;<code>chemicals</code>,&nbsp;<code>genes</code>,&nbsp;<code>anatomy</code>,&nbsp;<code>catalysts</code>,&nbsp;<code>cell lines</code>,&nbsp;<code>cofactors</code>,&nbsp;<code>complexes</code>,&nbsp;<code>GO biological processes</code>,&nbsp;<code>GO cellular components</code>,&nbsp;<code>GO molecular functions</code>,&nbsp;<code>pathways</code>,&nbsp;<code>proteins</code>,&nbsp;<code>reactions</code>, and&nbsp;<code>transcripts</code>. Additionally, the edges between this ontology and its dependencies are utilized:</p> <ul> <li><strong><a href="https://www.ebi.ac.uk/chebi/"><code>ChEBI</code></a></strong></li> <li><strong><a href="http://disease-ontology.org/"><code>DOID</code></a></strong></li> <li><strong><a href="http://geneontology.org/"><code>GO</code></a></strong></li> </ul> <p><strong>Notes:</strong>&nbsp;A partial, human-only version of this ontology was used. Details on how this version of the ontology was generated can be found under the Protein Ontology section of the&nbsp;<a href="https://github.com/callahantiff/PheKnowLator/blob/master/notebooks/Data_Preparation.ipynb"><code>Data_Preparation.ipynb</code></a>&nbsp;Jupyter Notebook.</p> <p><strong>Files</strong></p> <ul> <li> <p>Generated Human Version Protein Ontology (PRO)</p> <ul> <li><a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/human_pro.owl"><code>human_pro.owl</code></a>&nbsp;(closed with&nbsp;<a href="http://www.hermit-reasoner.com/">hermit reasoner</a>)</li> </ul> </li> <li> <p>Other PRO Data Used:&nbsp;<a href="https://proconsortium.org/download/current/promapping.txt"><code>promapping.txt</code></a></p> </li> <li> <p>Generated Mapping Data</p> <ul> <li>Merged Gene, RNA, Protein Map:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/Merged_gene_rna_protein_identifiers.pkl"><code>Merged_gene_rna_protein_identifiers.pkl</code></a></li> <li>Ensembl Transcript-PRO Identifier Mapping:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/ENTREZ_GENE_ENSEMBL_TRANSCRIPT_MAP.txt"><code>ENSEMBL_TRANSCRIPT_PROTEIN_ONTOLOGY_MAP.txt</code></a></li> <li>Entrez Gene-PRO Identifier Mapping:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/ENTREZ_GENE_PRO_ONTOLOGY_MAP.txt"><code>ENTREZ_GENE_PRO_ONTOLOGY_MAP.txt</code></a></li> <li>UniProt Accession-PRO Identifier Mapping:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/UNIPROT_ACCESSION_PRO_ONTOLOGY_MAP.txt"><code>UNIPROT_ACCESSION_PRO_ONTOLOGY_MAP.txt</code></a></li> <li>STRING-PRO Identifier Mapping:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/STRING_PRO_ONTOLOGY_MAP.txt"><code>STRING_PRO_ONTOLOGY_MAP.txt</code></a></li> </ul> </li> </ul> <p>&nbsp;</p> <p><strong>Relations Ontology (RO)</strong></p> <p><strong>Homepage:</strong>&nbsp;<strong><a href="https://github.com/oborel/obo-relations/"><code>GitHub</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Smith B, Ceusters W, Klagges B, K&ouml;hler J, Kumar A, Lomax J, Mungall C, Neuhaus F, Rector AL, Rosse C.&nbsp;<a href="https://genomebiology.biomedcentral.com/articles/10.1186/gb-2005-6-5-r46">Relations in biomedical ontologies</a>. Genome Biology. 2005;6(5):R46.</p> </blockquote> <p><strong>Usage:</strong>&nbsp;Utilizing this ontology to connect all data sources in knowledge graph. Additionally, the ontology is queried prior to building the knowledge graph to identify all relations, their inverse properties, and their labels.</p> <p><strong>Files</strong></p> <ul> <li>Generated RO Data <ul> <li><a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/INVERSE_RELATIONS.txt"><code>INVERSE_RELATIONS.txt</code></a></li> <li><a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/RELATIONS_LABELS.txt"><code>RELATIONS_LABELS.txt</code></a></li> </ul> </li> </ul> <p>&nbsp;</p> <p><strong>Sequence Ontology (SO)</strong></p> <p><strong>Homepage:</strong>&nbsp;<strong><a href="https://github.com/The-Sequence-Ontology/SO-Ontologies"><code>GitHub</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Eilbeck K, Lewis SE, Mungall CJ, Yandell M, Stein L, Durbin R, Ashburner M.&nbsp;<a href="https://link.springer.com/article/10.1186/gb-2005-6-5-r44">The Sequence Ontology: a tool for the unification of genome annotations</a>. Genome Biology. 2005;6(5):R44</p> </blockquote> <p><strong>Usage:</strong>&nbsp;Utilized to connect&nbsp;<code>transcripts</code>&nbsp;and other genomic material like&nbsp;<code>genes</code>&nbsp;and&nbsp;<code>variants</code>.</p> <p><strong>Files</strong></p> <ul> <li>Generated Mapping Data <ul> <li><a href="https://storage.googleapis.com/pheknowlator/curated_data/genomic_sequence_ontology_mappings.xlsx"><code>genomic_sequence_ontology_mappings.xlsx</code></a></li> <li><a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/SO_GENE_TRANSCRIPT_VARIANT_TYPE_MAPPING.txt"><code>SO_GENE_TRANSCRIPT_VARIANT_TYPE_MAPPING.txt</code></a></li> </ul> </li> </ul> <p>&nbsp;</p> <p><strong>Uber-Anatomy Ontology (Uberon)</strong></p> <p><strong>Homepage:</strong>&nbsp;<strong><a href="https://uberon.github.io/"><code>GitHub</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Mungall CJ, Torniai C, Gkoutos GV, Lewis SE, Haendel MA.&nbsp;<a href="https://genomebiology.biomedcentral.com/articles/10.1186/gb-2012-13-1-r5">Uberon, an integrative multi-species anatomy ontology</a>. Genome Biology. 2012;13(1):R5</p> </blockquote> <p><strong>Usage:</strong>&nbsp;Utilized to connect&nbsp;<code>tissues</code>,&nbsp;<code>fluids</code>, and&nbsp;<code>cells</code>&nbsp;to&nbsp;<code>proteins</code>&nbsp;and&nbsp;<code>transcripts</code>. Additionally, the edges between this ontology and its dependencies are utilized:</p> <ul> <li><strong><a href="https://www.ebi.ac.uk/chebi/"><code>ChEBI</code></a></strong></li> <li><strong><a href="http://www.obofoundry.org/ontology/cl.html"><code>CL</code></a></strong></li> <li><strong><a href="http://geneontology.org/"><code>GO</code></a></strong></li> <li><strong><a href="https://proconsortium.org/"><code>PRO</code></a></strong></li> </ul> <p>&nbsp;</p> <p><strong>Vaccine Ontology (VO)</strong></p> <p><strong>Homepage:</strong>&nbsp;<strong><a href="http://www.violinet.org/vaccineontology/"><code>http://www.violinet.org/vaccineontology/</code></a></strong><br> <strong>Citations:</strong></p> <blockquote> <p>He Y, Racz R, Sayers S, Lin Y, Todd T, Hur J, Li X, Patel M, Zhao B, Chung M, Ostrow J.&nbsp;<a href="https://academic.oup.com/nar/article-abstract/42/D1/D1124/1053128">Updates on the web-based VIOLIN vaccine database and analysis system</a>. Nucleic Acids Research. 2013;42(D1):D1124-32</p> <p>Xiang Z, Todd T, Ku KP, Kovacic BL, Larson CB, Chen F, Hodges AP, Tian Y, Olenzek EA, Zhao B, Colby LA.&nbsp;<a href="https://academic.oup.com/nar/article-abstract/36/suppl_1/D923/2505793">VIOLIN: vaccine investigation and online information network</a>. Nucleic Acids Research. 2007;36(suppl_1):D923-8</p> </blockquote> <p><strong>Usage:</strong>&nbsp;Utilized the edges between this ontology and its dependencies:</p> <ul> <li><strong><a href="https://www.ebi.ac.uk/chebi/"><code>ChEBI</code></a></strong></li> <li><strong><a href="http://disease-ontology.org/"><code>DOID</code></a></strong></li> <li><strong><a href="http://geneontology.org/"><code>GO</code></a></strong></li> <li><strong><a href="https://proconsortium.org/"><code>PRO</code></a></strong></li> <li><strong><a href="https://uberon.github.io/"><code>UBERON</code></a></strong></li> </ul> <p>&nbsp;</p> <p><strong>DATABASE SOURCES</strong></p> <ul> <li>BioPortal</li> <li>ClinVar</li> <li>Comparative Toxicogenomics Database</li> <li>DisGeNET</li> <li>Ensembl</li> <li>GeneMANIA</li> <li>Genotype-Tissue Expression Project</li> <li>Human Genome Organisation Gene Nomenclature Committee</li> <li>Human Protein Atlas</li> <li>National Center for Biotechnology Information Gene</li> <li>Reactome Pathway Database</li> <li>Search Tool for Recurring Instances of Neighbouring Genes Database</li> <li>Universal Protein Resource Knowledgebase</li> </ul> <p>&nbsp;</p> <p><strong>BioPortal</strong></p> <p><strong>Homepage:</strong>&nbsp;<strong><a href="https://bioportal.bioontology.org/"><code>BioPortal</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>BioPortal.&nbsp;<a href="https://www.bioontology.org/wiki/LOOM">Lexical OWL Ontology Matcher (LOOM)</a></p> <p>Ghazvinian A, Noy NF, Musen MA.&nbsp;<a href="https://www.ncbi.nlm.nih.gov/pmc/articles/pmc2815474/">Creating mappings for ontologies in biomedicine: simple methods work</a>. In AMIA Annual Symposium Proceedings 2009 (Vol. 2009, p. 198). American Medical Informatics Association</p> </blockquote> <p><strong>Usage:</strong>&nbsp;BioPortal was utilized to obtain mappings between&nbsp;<code>MeSH identifiers</code>&nbsp;and&nbsp;<code>ChEBI identifiers</code>&nbsp;for&nbsp;<code>chemicals-diseases</code>,&nbsp;<code>chemicals-genes</code>,&nbsp;<code>chemical-GO biological processes</code>,&nbsp;<code>chemicals-GO cellular components</code>,&nbsp;<code>chemicals-GO molecular functions</code>,&nbsp;<code>chemicals-phenotypes</code>,&nbsp;<code>chemicals-proteins</code>, and&nbsp;<code>chemicals-transcripts</code>. Additional information on how this data was processed can be obtained from the&nbsp;<a href="https://gist.github.com/callahantiff/a28fb3160782f42f104e9ec41553af0d"><code>NCBO_rest_api.py</code></a>&nbsp;GitHub Gist script.</p> <p>⭐&nbsp;<strong>ALTERNATIVE METHOD</strong>⭐ Since the above approach can take over two days to process, we have developed an alternative solution that downloads the&nbsp;<a><code>mesh2021.nt</code></a>&nbsp;data file directly from MeSH and the&nbsp;<a><code>Flat_file_tab_delimited/names.tsv.gz</code></a>&nbsp;file directly from ChEBI. Using these files, we have recapitulated the&nbsp;<a href="https://www.bioontology.org/wiki/BioPortal_Mappings"><code>LOOM</code></a>&nbsp;algorithm implemented by BioPortal when creating mappings between these resources. The procedure is relatively straightforward and utilizes the following information from each resource:</p> <ul> <li>For all MeSH&nbsp;<code>SCR Chemicals</code>, obtain the following information: <ul> <li>Identifiers: MeSH identifiers</li> <li>Labels: string labels using the&nbsp;<code>RDFS:label</code>&nbsp;object property</li> <li>Synonyms: track down all synonyms using the&nbsp;<code>vocab:concept</code>&nbsp;and&nbsp;<code>vocab:preferredConcept</code>&nbsp;object properties</li> </ul> </li> <li>For all ChEBI classes, obtain the following information: <ul> <li>Labels: string labels using the&nbsp;<code>RDFS:label</code>&nbsp;object property</li> <li>Synonyms: track down all synonyms using all&nbsp;<code>synonym</code>&nbsp;object properties</li> </ul> </li> </ul> <p><strong>Files</strong></p> <ul> <li>Generated Data:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/MESH_CHEBI_MAP.txt"><code>MESH_CHEBI_MAP.txt</code></a></li> </ul> <p>&nbsp;</p> <p><strong>ClinVar</strong></p> <p><strong>Homepage:</strong>&nbsp;<strong><a href="https://www.ncbi.nlm.nih.gov/clinvar/"><code>https://www.ncbi.nlm.nih.gov/clinvar/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Landrum MJ, Lee JM, Benson M, Brown GR, Chao C, Chitipiralla S, Gu B, Hart J, Hoffman D, Jang W, Karapetyan K.&nbsp;<a href="https://academic.oup.com/nar/article-abstract/46/D1/D1062/4641904">ClinVar: improving access to variant interpretations and supporting evidence</a>. Nucleic Acids Research. 2017;46(D1):D1062-7</p> </blockquote> <p><strong>Usage:</strong>&nbsp;ClinVar was utilized to create&nbsp;<code>variant-gene</code>,&nbsp;<code>variant-disease</code>, and&nbsp;<code>variant-phenotype</code>&nbsp;edges. The original data is filtered such that only records meeting the following criteria were included:</p> <ul> <li> <p><code>Assembly</code>&nbsp;= &quot;GRCh38&quot;</p> </li> <li> <p><code>ClinSigSimple</code>&nbsp;=&nbsp;<code>1</code></p> <ul> <li> <blockquote> <p>1 = at least one current record submitted with an interpretation of Likely pathogenic or Pathogenic (independent of whether that record includes assertion criteria and evidence)&quot;</p> </blockquote> </li> </ul> </li> <li> <p><code>ReviewStatus</code>&nbsp;in [&quot;criteria provided, multiple submitters, no conflicts&quot;, &quot;reviewed by expert panel&quot;, &quot;practice guideline&quot;]</p> </li> </ul> <p><strong>Files</strong></p> <ul> <li> <p>Downloaded Data</p> <ul> <li><a href="https://ftp.ncbi.nlm.nih.gov/pub/clinvar/tab_delimited/variant_summary.txt.gz"><code>variant_summary.txt.gz</code></a></li> <li><a href="https://ftp.ncbi.nlm.nih.gov/pub/clinvar/tab_delimited/var_citations.txt"><code>var_citations.txt</code></a></li> <li><a href="https://ftp.ncbi.nlm.nih.gov/pub/clinvar/tab_delimited/allele_gene.txt.gz"><code>allele_gene.txt.gz</code></a></li> </ul> </li> <li> <p>Generated Edge Data:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/CLINVAR_VARIANT_GENE_DISEASE_PHENOTYPE_EDGES.txt"><code>CLINVAR_VARIANT_GENE_DISEASE_PHENOTYPE_EDGES.txt</code></a></p> </li> </ul> <p>&nbsp;</p> <p><strong>Comparative Toxicogenomics Database (CTD)</strong></p> <p><strong>Homepage:</strong>&nbsp;<strong><a href="http://ctdbase.org/"><code>http://ctdbase.org/</code></a></strong><br> <strong>Citations:</strong></p> <blockquote> <p>Curated [chemical&ndash;gene interactions|chemical-go interactions|chemical&ndash;disease interactions|gene&ndash;pathway interactions] data were retrieved from the Comparative Toxicogenomics Database (CTD), MDI Biological Laboratory, Salisbury Cove, Maine, and NC State University, Raleigh, North Carolina. World Wide Web (URL:&nbsp;<a href="http://ctdbase.org/">http://ctdbase.org/</a>)</p> <p>Davis AP, Grondin CJ, Johnson RJ, Sciaky D, McMorran R, Wiegers J, Wiegers TC, Mattingly CJ.&nbsp;<a href="https://academic.oup.com/nar/article-abstract/47/D1/D948/5106145">The comparative toxicogenomics database: update 2019</a>. Nucleic Acids Research. 2018;47(D1):D948-54</p> </blockquote> <p>Usage: Comparative Toxicogenomics Database (CTD) was utilized to create&nbsp;<code>chemical-disease</code>,&nbsp;<code>chemical-gene</code>,&nbsp;<code>chemical-GO biological process</code>,&nbsp;<code>chemical-GO cellular components</code>,&nbsp;<code>chemical-GO molecular functions</code>,&nbsp;<code>chemical-phenotype</code>,&nbsp;<code>chemical-protein</code>,&nbsp;<code>chemical-rna</code>, and&nbsp;<code>gene-pathway</code>&nbsp;edges. The original data is filtered such that only records meeting the following criteria were included:</p> <ul> <li><code>chemical-disease</code>:&nbsp;<code>DirectEvidence</code>&nbsp;!= &quot;&quot;</li> <li><code>chemical-gene</code>:&nbsp;<code>Organism</code>&nbsp;== &quot;Homo sapiens&quot;,&nbsp;<code>GeneForms</code>&nbsp;== &quot;gene&quot;, and affects not in&nbsp;<code>InteractionActions</code></li> <li><code>chemical-GO biological process</code>:&nbsp;<code>PhenotypeName</code>&nbsp;== &quot;Biological Process&quot; and&nbsp;<code>Interaction</code>&nbsp;&lt;= &quot;1.04e-47&quot; (10th percentile)</li> <li><code>chemical-GO cellular components</code>:&nbsp;<code>PhenotypeName</code>&nbsp;== &quot;Cellular Component&quot; and&nbsp;<code>Interaction</code>&nbsp;&lt;= &quot;1.04e-47&quot; (10th percentile)</li> <li><code>chemical-GO molecular functions</code>:&nbsp;<code>PhenotypeName</code>&nbsp;== &quot;Molecular Function&quot; and&nbsp;<code>Interaction</code>&nbsp;&lt;= &quot;1.04e-47&quot; (10th percentile)</li> <li><code>chemical-phenotype</code>:&nbsp;<code>DirectEvidence</code>&nbsp;!= &quot;&quot;</li> <li><code>chemical-protein</code>:&nbsp;<code>Organism</code>&nbsp;== &quot;Homo sapiens&quot;,&nbsp;<code>GeneForms</code>&nbsp;== &quot;protein&quot;, and affects not in&nbsp;<code>InteractionActions</code></li> <li><code>chemical-rna</code>:&nbsp;<code>Organism</code>&nbsp;== &quot;Homo sapiens&quot;,&nbsp;<code>GeneForms</code>&nbsp;== &quot;mRNA&quot;, and affects and activity not in&nbsp;<code>InteractionActions</code></li> <li><code>gene-pathway edges</code>:&nbsp;<code>PathwayName</code>&nbsp;== R-HSA-</li> </ul> <p><strong>Files</strong></p> <ul> <li>Downloaded Data <ul> <li>Chemical-Gene Relations:&nbsp;<a href="http://ctdbase.org/reports/CTD_chem_gene_ixns.tsv.gz"><code>CTD_chem_gene_ixns.tsv.gz</code></a></li> <li>Chemical-Disease/Phenotype Relations:&nbsp;<a href="http://ctdbase.org/reports/CTD_chemicals_diseases.tsv.gz"><code>CTD_chemicals_diseases.tsv.gz</code></a></li> <li>Chemical-GO Relations:&nbsp;<a href="http://ctdbase.org/reports/CTD_chem_go_enriched.tsv.gz"><code>CTD_chem_go_enriched.tsv.gz</code></a></li> <li>Gene-Pathway Relations:&nbsp;<a href="http://ctdbase.org/reports/CTD_genes_pathways.tsv.gz"><code>CTD_genes_pathways.tsv.gz</code></a></li> </ul> </li> </ul> <p>&nbsp;</p> <p><strong>DisGeNET</strong></p> <p><strong>Homepage:</strong>&nbsp;<strong><a href="https://www.disgenet.org/"><code>https://www.disgenet.org/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Gene-disease association data retrieved from DisGeNET v6.0 (<a href="http://www.disgenet.org/">http://www.disgenet.org/</a>), Integrative Biomedical Informatics Group GRIB/IMIM/UPF. [December, 2019].</p> <p>Pi&ntilde;ero J, Ram&iacute;rez-Anguita JM, Sa&uuml;ch-Pitarch J, Ronzano F, Centeno E, Sanz F, Furlong LI.&nbsp;<a href="https://academic.oup.com/nar/advance-article-abstract/doi/10.1093/nar/gkz1021/5611674">The DisGeNET knowledge platform for disease genomics: 2019 update</a>. Nucleic Acids Research. 2019.</p> </blockquote> <p><strong>Usage:</strong>&nbsp;DisGeNET was utilized to create&nbsp;<code>gene-disease</code>, and&nbsp;<code>gene-phenotype</code>&nbsp;edges. The original data is filtered such that only records meeting the following criteria were included:&nbsp;<code>EI</code>&nbsp;&gt;= &quot;1.0&quot; (90th percentile). Additionally, data from this source was used to create mappings between different types of disease and phenotype identifiers, including:</p> <ul> <li>OMIM, ORPHA, UMLS, ICD ➞ DOID</li> <li>OMIM, ORPHA, UMLS, ICD ➞ HPO</li> </ul> <p><strong>Files</strong></p> <ul> <li> <p>Downloaded Data</p> <ul> <li>Disease/Phenotype-Gene Relations:&nbsp;<a href="https://www.disgenet.org/static/disgenet_ap1/files/downloads/curated_gene_disease_associations.tsv.gz"><code>curated_gene_disease_associations.tsv.gz</code></a></li> <li>Disease Identifier Mapping:&nbsp;<a href="https://www.disgenet.org/static/disgenet_ap1/files/downloads/disease_mappings.tsv.gz"><code>disease_mappings.tsv.gz</code></a></li> </ul> </li> <li> <p>Generated Mapping Data</p> <ul> <li>Disease Identifier Mapping:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/PHENOTYPE_HPO_MAP.txt"><code>PHENOTPYE_HPO_MAP.txt</code></a></li> <li>Phenotype Identifier Mapping:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/DISEASE_MONDO_MAP.txt"><code>DISEASE_DOID_MAP.txt</code></a></li> </ul> </li> </ul> <p>&nbsp;</p> <p><strong>Ensembl</strong></p> <p><strong>Homepage:</strong>&nbsp;<strong><a href="https://uswest.ensembl.org/index.html"><code>https://uswest.ensembl.org/index.html</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Zerbino DR, Achuthan P, Akanni W, Amode MR, Barrell D, Bhai J, Billis K, Cummins C, Gall A, Gir&oacute;n CG, Gil L.&nbsp;<a href="https://academic.oup.com/nar/article/46/D1/D754/4634002">Ensembl 2018</a>. Nucleic Acids Research. 2017;46(D1):D754-61</p> </blockquote> <p><strong>Usage:</strong>&nbsp;Ensembl data was utilized to create mappings between Ensembl genes, transcripts, and proteins with&nbsp;<a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#national-center-for-biotechnology-information-gene">NCBI Gene identifiers</a>,&nbsp;<a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#human-genome-organisation-gene-nomenclature-committee">HUGO gene symbols</a>,&nbsp;<a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#universal-protein-resource-knowledgebase">UniProt Accession identifiers</a>, and&nbsp;<a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#protein-ontology">Protein Ontology identifiers</a>&nbsp;in the knowledge graph (for additional details on the processing of these data, see&nbsp;<a href="https://github.com/callahantiff/PheKnowLator/blob/master/notebooks/Data_Preparation.ipynb"><code>Data_Preparation.ipynb</code></a>):</p> <ul> <li>Ensembl Transcript IDs ➞ PRO IDs</li> <li>Gene Ensembl IDs ➞ Entrez Gene IDs</li> <li>Gene Ensembl IDs ➞ PRO IDs</li> <li>Gene Symbols ➞ Transcript Ensembl IDs</li> <li>Entrez Gene IDs ➞ Transcript Ensembl IDs</li> <li>Entrez Gene IDs ➞ PRO IDs</li> <li>Protein Ensembl IDs ➞ UniProt Protein Accession</li> <li>STRING IDs ➞ PRO IDs</li> <li>UniProt Protein Accession ➞ Entrez Gene IDs</li> </ul> <p><strong>Files</strong></p> <ul> <li> <p>Downloaded Data</p> <ul> <li><a><code>Homo_sapiens.GRCh38.102.gtf</code></a></li> <li><a><code>Homo_sapiens.GRCh38.102.uniprot.tsv.gz</code></a></li> <li><a><code>Homo_sapiens.GRCh38.102.entrez.tsv.gz</code></a></li> </ul> </li> <li> <p>Generated Mapping Data</p> <ul> <li>Cleaned Ensembl Gene Set:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/ensembl_identifier_data_cleaned.txt"><code>ensembl_identifier_data_cleaned.txt</code></a></li> <li>Merged Gene, RNA, Protein Map:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/Merged_gene_rna_protein_identifiers.pkl"><code>Merged_gene_rna_protein_identifiers.pkl</code></a></li> <li>Ensembl Transcript-PRO Identifier Mapping:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/ENSEMBL_TRANSCRIPT_PROTEIN_ONTOLOGY_MAP.txt"><code>ENSEMBL_TRANSCRIPT_PROTEIN_ONTOLOGY_MAP.txt</code></a></li> <li>Gene Symbol-Ensembl Transcript Identifier Mapping:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/GENE_SYMBOL_ENSEMBL_TRANSCRIPT_MAP.txt"><code>GENE_SYMBOL_ENSEMBL_TRANSCRIPT_MAP.txt</code></a></li> <li>Entrez Gene-Ensembl Transcript Identifier Mapping:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/ENTREZ_GENE_ENSEMBL_TRANSCRIPT_MAP.txt"><code>ENTREZ_GENE_ENSEMBL_TRANSCRIPT_MAP.txt</code></a></li> <li>Entrez Gene-PRO Identifier Mapping:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/ENTREZ_GENE_PRO_ONTOLOGY_MAP.txt"><code>ENTREZ_GENE_PRO_ONTOLOGY_MAP.txt</code></a></li> <li>Ensembl Gene-Entrez Gene Identifier Mapping:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/ENSEMBL_GENE_ENTREZ_GENE_MAP.txt"><code>ENSEMBL_GENE_ENTREZ_GENE_MAP.txt</code></a></li> </ul> </li> </ul> <p>&nbsp;</p> <p><strong>GeneMANIA</strong></p> <p><strong>Homepage:</strong>&nbsp;<strong><a href="https://genemania.org/"><code>https://genemania.org/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Warde-Farley D, Donaldson SL, Comes O, Zuberi K, Badrawi R, Chao P, Franz M, Grouios C, Kazi F, Lopes CT, Maitland A.&nbsp;<a href="https://academic.oup.com/nar/article-abstract/38/suppl_2/W214/1126704">The GeneMANIA prediction server: biological network integration for gene prioritization and predicting gene function</a>. Nucleic Acids Research. 2010;38(suppl_2):W214-20</p> </blockquote> <p><strong>Usage:</strong>&nbsp;GeneMANIA was utilized to create&nbsp;<code>gene-gene</code>&nbsp;edges.</p> <p><strong>Files</strong></p> <ul> <li>Downloaded Data:&nbsp;<a href="http://genemania.org/data/current/Homo_sapiens.COMBINED/COMBINED.DEFAULT_NETWORKS.BP_COMBINING.txt"><code>COMBINED.DEFAULT_NETWORKS.BP_COMBINING.txt</code></a></li> </ul> <p>&nbsp;</p> <p><strong>Genotype-Tissue Expression Project (GTEx)</strong></p> <p><strong>Homepage:</strong>&nbsp;<strong><a href="https://gtexportal.org/home/"><code>https://gtexportal.org/home/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Lonsdale J, Thomas J, Salvatore M, Phillips R, Lo E, Shad S, Hasz R, Walters G, Garcia F, Young N, Foster B.&nbsp;<a href="http://www.nature.com/ng/journal/v45/n6/full/ng.2653.html">The genotype-tissue expression (GTEx) project</a>. Nature Genetics. 2013;45(6):580</p> </blockquote> <p><strong>Usage:</strong>&nbsp;The Genotype-Tissue Expression (GTEx) Project was utilized to create edges between&nbsp;<code>protein-cell</code>,&nbsp;<code>protein-anatomy</code>,&nbsp;<code>rna-cell</code>&nbsp;and&nbsp;<code>rna-anatomy</code>&nbsp;entities. The original data were filtered such that only those edges where the median TPM was &gt;=<code>1.0</code>&nbsp;and genes were of any type other than protein-coding were included. It should also be noted that we chose to use the RNASeQC file over the RSEM file as advised by the GTEx website.</p> <blockquote> <p>The RSEM estimates are based on combining isoform-level estimates, which adds uncertainty to the resulting gene-level values (the isoform-level estimates are highly inaccurate in some cases).</p> </blockquote> <p>The file contains&nbsp;<code>54</code>&nbsp;unique tissue and/or cell types. GTEx provides mappings from tissue types to UBERON and EFO. These provided&nbsp;<a href="https://gtexportal.org/home/samplingSitePage">mappings</a>&nbsp;were verified and extended, such that all samples which referenced a cell type were also mapped to the Cell and the Cell Line ontologies. This resulted in a total of&nbsp;<code>56</code>&nbsp;mappings (<code>1.04</code>&nbsp;mappings/concepts).</p> <p><strong>Files</strong></p> <ul> <li> <p>Downloaded Data:&nbsp;<a href="https://storage.googleapis.com/gtex_analysis_v8/rna_seq_data/GTEx_Analysis_2017-06-05_v8_RNASeQCv1.1.9_gene_median_tpm.gct.gz"><code>GTEx_Analysis_2017-06-05_v8_RNASeQCv1.1.9_gene_median_tpm.gct</code></a></p> </li> <li> <p>Mapping Results:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/curated_data/zooma_tissue_cell_mapping_04JAN2020.xlsx"><code>zooma_tissue_cell_mapping_04JAN2020.xlsx</code></a></p> </li> <li> <p>Generated Data<br> The final mapping set was combined with terms from the&nbsp;<a href="https://www.proteinatlas.org/">Human Protein Atlas</a>, see&nbsp;<a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources/t#human-protein-atlas">here</a>&nbsp;for more information.</p> <ul> <li>All HPA tissue and cell type strings:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/HPA_tissues.txt"><code>HPA_tissues.txt</code></a></li> <li>Final Term Mapping:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/HPA_GTEx_TISSUE_CELL_MAP.txt"><code>HPA_GTEx_TISSUE_CELL_MAP.txt</code></a></li> <li>Final RNA, Gene, Protein-Tissues and Cell Types Relations:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/HPA_GTEX_RNA_GENE_PROTEIN_EDGES.txt"><code>HPA_GTEX_RNA_GENE_PROTEIN_EDGES.txt</code></a></li> </ul> </li> </ul> <p>&nbsp;</p> <p><strong>Human Genome Organisation Gene Nomenclature Committee (HUGO)</strong></p> <p><strong>Homepage:</strong>&nbsp;<strong><a href="https://www.genenames.org/"><code>https://www.genenames.org/</code></a></strong><br> <strong>Citations:</strong></p> <blockquote> <p>HGNC Database, HUGO Gene Nomenclature Committee (HGNC), European Molecular Biology Laboratory, European Bioinformatics Institute (EMBL-EBI), Wellcome Genome Campus, Hinxton, Cambridge CB10 1SD, United Kingdom&nbsp;<a href="https://www.genenames.org/">www.genenames.org</a></p> <p>Yates B, Braschi B, Gray K, Seal R, Tweedie S, Bruford E.&nbsp;<a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5210531/">Genenames.org: the HGNC and VGNC Resources in 2017</a>. Nucleic Acids Research. 2017;45(D1):D619-625</p> </blockquote> <p><strong>Usage:</strong>&nbsp;The Human Genome Organisation (HUGO) data was utilized to obtain mappings between&nbsp;<a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#national-center-for-biotechnology-information-gene">NCBI Gene identifiers</a>, HUGO gene symbols,&nbsp;<a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#universal-protein-resource-knowledgebase">UniProt Accession identifiers</a>, and&nbsp;<a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#protein-ontology">Protein Ontology identifiers</a>. For additional details on the processing of these data, see&nbsp;<a href="https://github.com/callahantiff/PheKnowLator/blob/master/notebooks/Data_Preparation.ipynb"><code>Data_Preparation.ipynb</code></a>:</p> <ul> <li>Ensembl Transcript IDs ➞ PRO IDs</li> <li>Gene Ensembl IDs ➞ Entrez Gene IDs</li> <li>Gene Ensembl IDs ➞ PRO IDs</li> <li>Gene Symbols ➞ Transcript Ensembl IDs</li> <li>Entrez Gene IDs ➞ Transcript Ensembl IDs</li> <li>Entrez Gene IDs ➞ PRO IDs</li> <li>Protein Ensembl IDs ➞ UniProt Protein Accession</li> <li>STRING IDs ➞ PRO IDs</li> <li>UniProt Protein Accession ➞ Entrez Gene IDs</li> </ul> <p><strong>Files</strong></p> <ul> <li> <p>Downloaded Data:&nbsp;<a href="http://ftp.ebi.ac.uk/pub/databases/genenames/hgnc/tsv/hgnc_complete_set.txt"><code>hgnc_complete_set.txt</code></a></p> </li> <li> <p>Generated Data</p> <ul> <li>Merged Gene, RNA, Protein Map:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/Merged_gene_rna_protein_identifiers.pkl"><code>Merged_gene_rna_protein_identifiers.pkl</code></a></li> <li>Gene Symbol-Ensembl Transcript Identifier Mapping:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/GENE_SYMBOL_ENSEMBL_TRANSCRIPT_MAP.txt"><code>GENE_SYMBOL_ENSEMBL_TRANSCRIPT_MAP.txt</code></a></li> </ul> </li> </ul> <p>&nbsp;</p> <p><strong>Human Protein Atlas (HPA)</strong></p> <p><strong>Homepage:</strong>&nbsp;<strong><a href="https://www.proteinatlas.org/"><code>https://www.proteinatlas.org/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Uhl&eacute;n M, Fagerberg L, Hallstr&ouml;m BM, Lindskog C, Oksvold P, Mardinoglu A, Sivertsson &Aring;, Kampf C, Sj&ouml;stedt E, Asplund A, Olsson I.&nbsp;<a href="https://science.sciencemag.org/content/347/6220/1260419.short">Tissue-based map of the human proteome</a>. Science. 2015;347(6220):1260419</p> </blockquote> <p><strong>Usage:</strong>&nbsp;The Human Protein Atlas (HPA) was utilized to create&nbsp;<code>rna-cell</code>,&nbsp;<code>rna-anatomy</code>,&nbsp;<code>protein-cell</code>, and&nbsp;<code>protein-anatomy</code>&nbsp;edges. Evidence between gene and RNA expression in specific tissue types was derived by HPA, such that the&nbsp;<a href="https://www.proteinatlas.org/about/assays+annotation#normalization_rna">consensus normalized expression</a>&nbsp;was &gt;=<code>1.0</code>. Zooma was utilized to automatically annotate the&nbsp;<code>153</code>&nbsp;unique tissues and cell types from Human Protein Atlas for all human protein-coding genes in the&nbsp;<a href="https://www.proteinatlas.org/humanproteome">Human Proteome</a>&nbsp;to the Cell Ontology, Cell Line Ontology, and the Uber-Anatomy Ontology. To best represent each concept, the automatic mappings from Zooma were extend through manual mapping efforts to ensure each concept cell type was matched to a Cell Ontology, Cell Line Ontology, and UBERON ontology term. This resulted in a total of&nbsp;<code>281</code>&nbsp;mappings (<code>1.84</code>&nbsp;mappings/concepts).</p> <p><strong>Files</strong></p> <ul> <li> <p>Downloaded Data:&nbsp;<a href="https://www.proteinatlas.org/api/search_download.php?search=&amp;columns=g,eg,up,pe,rnatsm,rnaclsm,rnacasm,rnabrsm,rnabcsm,rnablsm,scl,t_RNA_adipose_tissue,t_RNA_adrenal_gland,t_RNA_amygdala,t_RNA_appendix,t_RNA_basal_ganglia,t_RNA_bone_marrow,t_RNA_breast,t_RNA_cerebellum,t_RNA_cerebral_cortex,t_RNA_cervix,_uterine,t_RNA_colon,t_RNA_corpus_callosum,t_RNA_ductus_deferens,t_RNA_duodenum,t_RNA_endometrium_1,t_RNA_epididymis,t_RNA_esophagus,t_RNA_fallopian_tube,t_RNA_gallbladder,t_RNA_heart_muscle,t_RNA_hippocampal_formation,t_RNA_hypothalamus,t_RNA_kidney,t_RNA_liver,t_RNA_lung,t_RNA_lymph_node,t_RNA_midbrain,t_RNA_olfactory_region,t_RNA_ovary,t_RNA_pancreas,t_RNA_parathyroid_gland,t_RNA_pituitary_gland,t_RNA_placenta,t_RNA_pons_and_medulla,t_RNA_prostate,t_RNA_rectum,t_RNA_retina,t_RNA_salivary_gland,t_RNA_seminal_vesicle,t_RNA_skeletal_muscle,t_RNA_skin_1,t_RNA_small_intestine,t_RNA_smooth_muscle,t_RNA_spinal_cord,t_RNA_spleen,t_RNA_stomach_1,t_RNA_testis,t_RNA_thalamus,t_RNA_thymus,t_RNA_thyroid_gland,t_RNA_tongue,t_RNA_tonsil,t_RNA_urinary_bladder,t_RNA_vagina,t_RNA_B-cells,t_RNA_dendritic_cells,t_RNA_granulocytes,t_RNA_monocytes,t_RNA_NK-cells,t_RNA_T-cells,t_RNA_total_PBMC,cell_RNA_A-431,cell_RNA_A549,cell_RNA_AF22,cell_RNA_AN3-CA,cell_RNA_ASC_diff,cell_RNA_ASC_TERT1,cell_RNA_BEWO,cell_RNA_BJ,cell_RNA_BJ_hTERT+,cell_RNA_BJ_hTERT+_SV40_Large_T+,cell_RNA_BJ_hTERT+_SV40_Large_T+_RasG12V,cell_RNA_CACO-2,cell_RNA_CAPAN-2,cell_RNA_Daudi,cell_RNA_EFO-21,cell_RNA_fHDF/TERT166,cell_RNA_HaCaT,cell_RNA_HAP1,cell_RNA_HBEC3-KT,cell_RNA_HBF_TERT88,cell_RNA_HDLM-2,cell_RNA_HEK_293,cell_RNA_HEL,cell_RNA_HeLa,cell_RNA_Hep_G2,cell_RNA_HHSteC,cell_RNA_HL-60,cell_RNA_HMC-1,cell_RNA_HSkMC,cell_RNA_hTCEpi,cell_RNA_hTEC/SVTERT24-B,cell_RNA_hTERT-HME1,cell_RNA_HUVEC_TERT2,cell_RNA_K-562,cell_RNA_Karpas-707,cell_RNA_LHCN-M2,cell_RNA_MCF7,cell_RNA_MOLT-4,cell_RNA_NB-4,cell_RNA_NTERA-2,cell_RNA_PC-3,cell_RNA_REH,cell_RNA_RH-30,cell_RNA_RPMI-8226,cell_RNA_RPTEC_TERT1,cell_RNA_RT4,cell_RNA_SCLC-21H,cell_RNA_SH-SY5Y,cell_RNA_SiHa,cell_RNA_SK-BR-3,cell_RNA_SK-MEL-30,cell_RNA_T-47d,cell_RNA_THP-1,cell_RNA_TIME,cell_RNA_U-138_MG,cell_RNA_U-2_OS,cell_RNA_U-2197,cell_RNA_U-251_MG,cell_RNA_U-266/70,cell_RNA_U-266/84,cell_RNA_U-698,cell_RNA_U-87_MG,cell_RNA_U-937,cell_RNA_WM-115,blood_RNA_basophil,blood_RNA_classical_monocyte,blood_RNA_eosinophil,blood_RNA_gdT-cell,blood_RNA_intermediate_monocyte,blood_RNA_MAIT_T-cell,blood_RNA_memory_B-cell,blood_RNA_memory_CD4_T-cell,blood_RNA_memory_CD8_T-cell,blood_RNA_myeloid_DC,blood_RNA_naive_B-cell,blood_RNA_naive_CD4_T-cell,blood_RNA_naive_CD8_T-cell,blood_RNA_neutrophil,blood_RNA_NK-cell,blood_RNA_non-classical_monocyte,blood_RNA_plasmacytoid_DC,blood_RNA_T-reg,blood_RNA_total_PBMC,brain_RNA_amygdala,brain_RNA_basal_ganglia,brain_RNA_cerebellum,brain_RNA_cerebral_cortex,brain_RNA_hippocampal_formation,brain_RNA_hypothalamus,brain_RNA_midbrain,brain_RNA_olfactory_region,brain_RNA_pons_and_medulla,brain_RNA_thalamus&amp;format=tsv"><code>proteinatlas_search.tsv</code></a></p> </li> <li> <p>Mapping Results:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/curated_data/zooma_tissue_cell_mapping_04JAN2020.xlsx"><code>zooma_tissue_cell_mapping_04JAN2020.xlsx</code></a></p> </li> <li> <p>Generated Data</p> <ul> <li>Final Term Mapping:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/HPA_GTEx_TISSUE_CELL_MAP.txt"><code>HPA_GTEx_TISSUE_CELL_MAP.txt</code></a></li> <li>Final RNA, Gene, Protein-Tissues and Cell Types Relations:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/HPA_GTEX_RNA_GENE_PROTEIN_EDGES.txt"><code>HPA_GTEX_RNA_GENE_PROTEIN_EDGES.txt</code></a></li> </ul> </li> </ul> <p>&nbsp;</p> <p><strong>National Center for Biotechnology Information (NCBI) Entrez Gene</strong></p> <p><strong>Homepage:</strong>&nbsp;<strong><a href="https://www.ncbi.nlm.nih.gov/gene/"><code>https://www.ncbi.nlm.nih.gov/gene/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Maglott D, Ostell J, Pruitt KD, Tatusova T.&nbsp;<a href="https://academic.oup.com/nar/article-abstract/33/suppl_1/D54/2505255">Entrez Gene: gene-centered information at NCBI</a>. Nucleic Acids Research. 2005;33(suppl_1):D54-8.</p> </blockquote> <p><strong>Usage:</strong>&nbsp;The National Center for Biotechnology Information (NCBI) Gene data was utilized to obtain mappings between&nbsp;<a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#ncbi-gene">NCBI Gene identifiers</a>,&nbsp;<a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#hugo-gene-nomenclature-committee">HUGO gene symbols</a>,&nbsp;<a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#uniprot-knowledgebase">UniProt Accession identifiers</a>, and&nbsp;<a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#protein-ontology">Protein Ontology identifiers</a>. For additional details on the processing of these data, see&nbsp;<a href="https://github.com/callahantiff/PheKnowLator/blob/master/notebooks/Data_Preparation.ipynb"><code>Data_Preparation.ipynb</code></a>:</p> <ul> <li>Ensembl Transcript IDs ➞ PRO IDs</li> <li>Gene Ensembl IDs ➞ Entrez Gene IDs</li> <li>Gene Ensembl IDs ➞ PRO IDs</li> <li>Gene Symbols ➞ Transcript Ensembl IDs</li> <li>Entrez Gene IDs ➞ Transcript Ensembl IDs</li> <li>Entrez Gene IDs ➞ PRO IDs</li> <li>Protein Ensembl IDs ➞ UniProt Protein Accession</li> <li>STRING IDs ➞ PRO IDs</li> <li>UniProt Protein Accession ➞ Entrez Gene IDs</li> </ul> <p><strong>Files</strong></p> <ul> <li> <p>Downloaded Data:&nbsp;<a href="https://ftp.ncbi.nih.gov/gene/DATA/GENE_INFO/Mammalia/Homo_sapiens.gene_info.gz"><code>Homo_sapiens.gene_info.gz</code></a></p> </li> <li> <p>Generated Data</p> <ul> <li>Merged Gene, RNA, Protein Map:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/Merged_gene_rna_protein_identifiers.pkl"><code>Merged_gene_rna_protein_identifiers.pkl</code></a></li> <li>Entrez Gene-Ensembl Transcript Identifier Mapping:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/ENTREZ_GENE_ENSEMBL_TRANSCRIPT_MAP.txt"><code>ENTREZ_GENE_ENSEMBL_TRANSCRIPT_MAP.txt</code></a></li> <li>Entrez Gene-PRO Identifier Mapping:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/ENTREZ_GENE_PRO_ONTOLOGY_MAP.txt"><code>ENTREZ_GENE_PRO_ONTOLOGY_MAP.txt</code></a></li> <li>Ensembl Gene-Entrez Gene Identifier Mapping:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/ENSEMBL_GENE_ENTREZ_GENE_MAP.txt"><code>ENSEMBL_GENE_ENTREZ_GENE_MAP.txt</code></a></li> <li>Uniprot Accession-Entrez Gene Identifier Mapping:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/UNIPROT_ACCESSION_ENTREZ_GENE_MAP.txt"><code>UNIPROT_ACCESSION_ENTREZ_GENE_MAP.txt</code></a></li> </ul> </li> </ul> <p>&nbsp;</p> <p><strong>Reactome Pathway Database</strong></p> <p><strong>Homepage:</strong>&nbsp;<strong><a href="https://reactome.org/"><code>https://reactome.org/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Fabregat A, Jupe S, Matthews L, Sidiropoulos K, Gillespie M, Garapati P, Haw R, Jassal B, Korninger F, May B, Milacic M.&nbsp;<a href="https://academic.oup.com/nar/article-abstract/46/D1/D649/4626770">The reactome pathway knowledgebase</a>. Nucleic Acids Research. 2017;46(D1):D649-55</p> </blockquote> <p><strong>Usage:</strong>&nbsp;The Reactome Database was utilized to create&nbsp;<code>chemical-pathway</code>,&nbsp;<code>GO Biological process-pathway</code>,&nbsp;<code>pathway-GO Cellular component</code>,&nbsp;<code>GO Molecular function-pathway</code>, and&nbsp;<code>protein-pathway</code>&nbsp;edges. The original data is filtered such that only records meeting the following criteria were included:</p> <ul> <li><code>chemical-pathway</code>: column[5] == &quot;Homo sapiens&quot;</li> <li><code>GO Biological process-pathway</code>: column[5] startswith &quot;REACTOME&quot;, column[8] == &quot;P&quot;, and column[12] == &quot;taxon:9606&quot;</li> <li><code>pathway-GO Cellular component</code>: column[5] startswith &quot;REACTOME&quot;, column[8] == &quot;C&quot;, and column[12] == &quot;taxon:9606&quot;</li> <li><code>GO Molecular function-pathway</code>: column[5] startswith &quot;REACTOME&quot;, column[8] == &quot;F&quot;, and column[12] == &quot;taxon:9606&quot;</li> <li><code>protein-pathway</code>: column[5] == &quot;Homo sapiens&quot;</li> </ul> <p><strong>Files</strong></p> <ul> <li>Downloaded Data <ul> <li>Chemical-Pathway Relations:&nbsp;<a href="https://reactome.org/download/current/ChEBI2Reactome_All_Levels.txt"><code>ChEBI2Reactome_All_Levels.txt</code></a></li> <li>Pathway-GO Relations:&nbsp;<a href="https://reactome.org/download/current/gene_association.reactome"><code>gene_association.reactome</code></a></li> <li>Protein-Pathway Relations:&nbsp;<a href="https://reactome.org/download/current/UniProt2Reactome_All_Levels.txt"><code>UniProt2Reactome_All_Levels.txt</code></a></li> </ul> </li> </ul> <p>&nbsp;</p> <p><strong>Search Tool for Recurring Instances of Neighbouring Genes (STRING) Database</strong></p> <p><strong>Homepage:</strong>&nbsp;<strong><a href="https://string-db.org/"><code>string-db.org</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Szklarczyk D, Gable AL, Lyon D, Junge A, Wyder S, Huerta-Cepas J, Simonovic M, Doncheva NT, Morris JH, Bork P, Jensen LJ.&nbsp;<a href="https://academic.oup.com/nar/article-abstract/47/D1/D607/5198476">STRING v11: protein&ndash;protein association networks with increased coverage, supporting functional discovery in genome-wide experimental datasets</a>. Nucleic Acids Research. 2018;47(D1):D607-13</p> </blockquote> <p><strong>Usage:</strong>&nbsp;The Search Tool for Recurring Instances of Neighbouring Genes (STRING) Database was utilized to create&nbsp;<code>protein-protein</code>&nbsp;edges. The original data is filtered such that only records meeting the following criteria were included:&nbsp;<code>combined_score</code>&nbsp;&gt;= &quot;700&quot; (&gt;90th percentile).</p> <p><strong>Files</strong></p> <ul> <li> <p>Downloaded Data:&nbsp;<a href="https://stringdb-static.org/download/protein.links.v11.0/9606.protein.links.v11.0.txt.gz"><code>9606.protein.links.v11.0.txt.gz</code></a></p> </li> <li> <p>Generated Data: STRING-PRO Identifier Mapping:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/STRING_PRO_ONTOLOGY_MAP.txt"><code>STRING_PRO_ONTOLOGY_MAP.txt</code></a></p> </li> </ul> <p>&nbsp;</p> <p><strong>Universal Protein Resource (UniProt) Knowledgebase</strong></p> <p><strong>Homepage:</strong>&nbsp;<strong><a href="https://www.uniprot.org/"><code>https://www.uniprot.org/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>UniProt Consortium.&nbsp;<a href="https://academic.oup.com/nar/article-abstract/47/D1/D506/5160987">UniProt: a worldwide hub of protein knowledge</a>. Nucleic acids research. 2018;47(D1):D506-15</p> </blockquote> <p><strong>Usage:</strong>&nbsp;The Universal Protein Resource (UniProt) Knowledgebase was utilized to obtain&nbsp;<code>cofactor</code>/<code>catalyst</code>-<code>protein</code>&nbsp;and&nbsp;<code>protein-coding gene</code>-<code>protein</code>&nbsp;edges as well as mappings between&nbsp;<a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#national-center-for-biotechnology-information-gene">NCBI Gene identifiers</a>,&nbsp;<a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#human-genome-organisation-gene-nomenclature-committee">HUGO gene symbols</a>,&nbsp;<a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#universal-protein-resource-knowledgebase">Universal Protein Resource (UniProt) Accession identifiers</a>, and&nbsp;<a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#protein-ontology">Protein Ontology identifiers</a>. For additional details on the processing of these data, see&nbsp;<a href="https://github.com/callahantiff/PheKnowLator/blob/master/notebooks/Data_Preparation.ipynb"><code>Data_Preparation.ipynb</code></a>:</p> <ul> <li>Ensembl Transcript IDs ➞ PRO IDs</li> <li>Gene Ensembl IDs ➞ Entrez Gene IDs</li> <li>Gene Ensembl IDs ➞ PRO IDs</li> <li>Gene Symbols ➞ Transcript Ensembl IDs</li> <li>Entrez Gene IDs ➞ Transcript Ensembl IDs</li> <li>Entrez Gene IDs ➞ PRO IDs</li> <li>Protein Ensembl IDs ➞ UniProt Protein Accession</li> <li>STRING IDs ➞ PRO IDs</li> <li>UniProt Protein Accession ➞ Entrez Gene IDs</li> </ul> <p><strong>Files</strong></p> <ul> <li> <p>Downloaded Data</p> <ul> <li>Cofactor and Catalyst relations:&nbsp;<a href="https://www.uniprot.org/uniprot/?query=&amp;fil=organism%3A%22Homo%20sapiens%20(Human)%20%5B9606%5D%22&amp;columns=id%2Centry%20name%2Creviewed%2Cdatabase(PRO)%2Cchebi(Cofactor)%2Cchebi(Catalytic%20activity)"><code>Cofactor/Catalyst Query Results</code></a></li> <li>UniProt Identifier Mapping:&nbsp;<a href="https://www.uniprot.org/uniprot/?query=&amp;fil=organism%3A%22Homo%20sapiens%20(Human)%20%5B9606%5D%22&amp;columns=id%2Cdatabase(GeneID)%2Cdatabase(Ensembl)%2Cdatabase(HGNC)%2Cgenes(PREFERRED)%2Cgenes(ALTERNATIVE)"><code>UniProt Identifier Query Results</code></a></li> </ul> </li> <li> <p>Generated Data</p> <ul> <li>Merged Gene, RNA, Protein Map:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/Merged_gene_rna_protein_identifiers.pkl"><code>Merged_gene_rna_protein_identifiers.pkl</code></a></li> <li>Protein-Cofactor Relations:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/UNIPROT_PROTEIN_COFACTOR.txt"><code>UNIPROT_PROTEIN_COFACTOR.txt</code></a></li> <li>Protein-Catalyst Relations:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/UNIPROT_PROTEIN_CATALYST.txt"><code>UNIPROT_PROTEIN_CATALYST.txt</code></a></li> <li>UniProt Accession-PRO Identifier Mapping:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/UNIPROT_ACCESSION_PRO_ONTOLOGY_MAP.txt"><code>UNIPROT_ACCESSION_PRO_ONTOLOGY_MAP.txt</code></a></li> <li>UniProt Accession-Entrez Gene Identifier Mapping:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/UNIPROT_ACCESSION_ENTREZ_GENE_MAP.txt"><code>UNIPROT_ACCESSION_ENTREZ_GENE_MAP.txt</code></a></li> </ul> </li> </ul> <p>&nbsp;</p> <p>This project is licensed under Apache License 2.0 - see the&nbsp;<strong><a href="https://github.com/callahantiff/PheKnowLator/blob/master/LICENSE"><code>LICENSE.md</code></a></strong>&nbsp;file for details. If you intend to use any of the information on this Wiki, please provide the appropriate attribution by citing this repository:</p> <pre><code>@misc{callahan_tj_2019_3401437, author = {Callahan, TJ}, title = {PheKnowLator}, month = mar, year = 2019, doi = {10.5281/zenodo.3401437}, url = {https://doi.org/10.5281/zenodo.3401437} }</code></pre>

opencc-by-4.0Apr 2021View details →
zenodo44/100

A comprehensive floristic knowledge of the largest Atlantic Forest fragment of the Fluminense Paraíba do Sul River Valley, Rio de Janeiro, Brazil

<p>The &ldquo;<em>Serra da Conc&oacute;rdia</em>&rdquo; is part of the Atlantic Forest phytogeographical domain in the Brazilian state of Rio de Janeiro and it has a predominant phytophysiognomy of Semideciduous Seasonal Forest. This region underwent intense habitat loss and fragmentation during the 19<sup>th</sup> century, due to coffee plantations and later pastures. With the decline of these activities, the areas were abandoned, triggering secondary succession. In 2002, the &ldquo;<em>Parque Estadual da Serra da Conc&oacute;rdia</em>&rdquo; was established in this region to preserve the remaining forest fragments. The updated list of vascular plants recorded in this protected area, published in the &ldquo;<em>Cat&aacute;logo de Plantas das Unidades de Conserva&ccedil;&atilde;o do Brasil</em>&rdquo;, is presented here, along with information on richness, endemism, and conservation status.</p>

opencc-zeroMar 2024View details →
zenodo44/100

Knowledge of Social Networks for Health is Associated with COVID-19 Health Protective Behaviors

<p>This is the dataset and stata code for the paper "Knowledge of Social Networks for Health is Associated with COVID-19 Health Protective Behaviors&rdquo; submitted to Plos One May 1st, 2024.</p>

opencc-by-4.0Apr 2024View details →
zenodo44/100

Context-based Entity Recommendation on Real-Life Knowledge Work in Context (RLKWiC dataset)

<h2><a href="../records/11059573">RLKWiC</a> Add-on: Benchmarking Dataset for Entity Recommendation</h2> <p>This benchmark, built on top of the Real-Life Knowledge Work in Context (<a href="../records/11059573">RLKWiC</a>) dataset, is designed to evaluate context-based entity recommendation by simulating a scenario where participants receive entities extracted from their activities across their defined contexts.&nbsp;</p> <p>In total, 1850 entity recommendations were generated across 56 contexts. After deduplication, these entities were presented to participants for explicit relevance assessment on a 3-point scale:</p> <ul> <li>0 [Irrelevant]: Signifying a lack of relevance between the recommended entity and the context.</li> <li>1 [Relevant] Denoting a connection between the entity and the context, although it may not fully represent it.</li> <li>2 [Representative]: The entity closely aligns with the context, indicating a high level of relevance where the context can be inferred to be about this entity.</li> </ul> <p>Participants could also suggest additional relevant entities. The resulting dataset comprises 1067 entities with explicit relevance scores, offering a resource for benchmarking entity recommendation in real-life knowledge work.</p> <h3><strong>Paper: </strong><a href="https://dl.acm.org/doi/10.1145/3640457.3688068" target="_blank" rel="noopener">Context-based Entity Recommendation for Knowledge Workers: Establishing a Benchmark on Real-life Data</a></h3>

opencc-by-4.0May 2024View details →
zenodo44/100

MIRA-KG: A Knowledge Graph of Hypotheses and Findings for Social Demography Research

<p>A shift in scientific publishing from paper-based to knowledge-based practices promotes reproducibility, machine actionability and knowledge discovery. This is important for disciplines like social science, as study indicators are often social constructs such as race or education; hypothesis tests are challenging to compare in demographic research due to their limited temporal and spatial coverage; and natural language in research papers is often imprecise and ambiguous. Therefore, we present the MIRA-KG, consisting of: (1) an ontology for capturing social demography research, which links hypotheses and findings to evidence, (2) annotations of papers on health inequality in terms of the ontology, gathered by (i) prompting a Large Language Model to annotate paper abstracts using the ontology, (ii) mapping concepts to terms from NCBO BioPortal ontologies and GeoNames, and (iii) refining the final graph by a set of SHACL constraints, developed according to data quality criteria. The utility of the resource lies in its use for formally representing social demography research hypotheses, discovering research biases, discovery of knowledge, and the derivation of novel questions.<br><br>This dataset was generated using the code available on Github at <a href="https://w3id.org/mira/">https://w3id.org/mira/</a> at version v1.0. It uses the following ontology: <a href="https://w3id.org/mira/ontology/">https://w3id.org/mira/ontology/</a>.&nbsp;</p>

opencc-by-4.0Dec 2023View details →
zenodo44/100

FaaS Characteristics and Constraints Knowledge Base

<p>YAML-formatted and timestamped description of Function-as-a-Service (FaaS) service characteristics and constraints such as maximum execution time and pricing. This dataset allows for adaptive software and workflow generation in dynamically evolving Serverless Computing environments. We envision the inclusion of the dataset into code generators, code transformers, workflow schedulers and compatibility modes of open source FaaS runtimes.</p> <p>Furthermore, due to evidences of evolving values being given by hyperlinks, the dataset will serve as single source of truth about the technological development in the FaaS space.</p>

opencc-by-sa-4.0Apr 2018View details →
zenodo44/100

Sensitivity Datasets - Leveraging Implicit Knowledge in Neural Networks for Functional Dissection and Engineering of Proteins

<p><strong>Leveraging Implicit Knowledge in Neural Networks for Functional Dissection and Engineering of Proteins</strong></p> <p>The Sensitivity datasets cover more than 800 proteins and are structured as follows. The sensitivity values are the mean of four DeeProtein replicates.</p> <p>It is uploaded as tar.gz. and contains one directory.</p> <p>File names contain the PDB<sup>1</sup> identifier and the respective chain identifier.&nbsp;</p> <p>The sequences and secondary structure information were downloaded from the RCSB Protein Databank and are available here: <a href="https://cdn.rcsb.org/etl/kabschSander/ss_dis.txt.gz">https://cdn.rcsb.org/etl/kabschSander/ss_dis.txt.gz</a> This URL can be found with some explanation at <a href="http://www.rcsb.org/pdb/static.do?p=download/http/index.html">http://www.rcsb.org/pdb/static.do?p=download/http/index.html</a></p> <p>The secondary structure annotation relies on the DSSP Algorithm by Kabsch and Sander<sup>2</sup>.</p> <p>&nbsp;</p> <p><strong>The files are tab-separated and contain the following columns:</strong></p> <ul> <li><strong>Pos</strong>&nbsp;Position in the sequence, starting from zero</li> <li><strong>AA</strong>&nbsp;Amino acid in that position</li> <li><strong>sec</strong> Secondary structure as annotated in the RCSB Protein Databank</li> <li><strong>dis</strong>&nbsp;if a region has not been experimentally observed (sometimes explains mismatches with crystal structures)</li> <li><strong>GO:_______</strong>&nbsp;Sensitivity for the GO term</li> </ul> <p><strong>References</strong></p> <ol> <li>The Protein Data Bank H.M. Berman, J. Westbrook, Z. Feng, G. Gilliland, T.N. Bhat, H. Weissig, I.N. Shindyalov, P.E. Bourne (2000) Nucleic Acids Research, 28: 235-242. doi:10.1093/nar/28.1.235</li> <li>Kabsch, W. &amp; Sander, C. Dictionary of protein secondary structure: pattern recognition of hydrogen-bonded and geometrical features. Biopolymers 22, 2577-2637, doi:10.1002/bip.360221211 (1983).</li> </ol>

opencc-by-4.0Aug 2018View details →
zenodo44/100

Big Data to Knowledge (BD2K) Training Coordinating Center (TCC) Educational Resource Discovery Index (ERuDIte) as Linked Data

<p>This is a release of the Big Data to Knowledge (BD2K) Training Coordinating Center (TCC) Educational Resource Discovery Index (ERuDIte)&nbsp;as Linked Data.<br> <br> ERuDIte contains over 11,000 training resources on data science including courses (MOOCs), video tutorials, conference talks, and other materials. The metadata of these resources is described uniformly using schema.org. In addition, we use machine learning techniques to tag each resource with concepts from the Data Science Education Ontology (DSEO), which we developed to further describe the contents of the training resources. Resource relevance and tags are curated by experts to ensure high quality. Finally, we map the references to people and organizations in the learning resource metadata to entities in DBpedia, DBLP, and ORCID, thus embedding our collection in the web of linked data. Our collection is continually growing. We hope that ERuDIte will provide a framework to foster open linked educational resources on the web.<br> <br> &nbsp;Distributed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License (https://creativecommons.org/licenses/by-nc-sa/4.0/)</p>

openother-openMay 2018View details →
zenodo44/100

Datasets for Non-Parametric Class Completeness Estimators for Collaborative Knowledge Graphs

<p><strong>Non-Parametric Class Completeness Estimators for Collaborative Knowledge Graphs</strong></p> <p>This are intermediary datasets used for the calculation of the Class Completeness Estimators on Wikidata. For more information see:&nbsp;https://github.com/eXascaleInfolab/cardinal/</p> <p><strong>edits_wikidatawiki-20181001-pages.csv</strong></p> <p>This is an extract from&nbsp;<em>wikidatawiki-20181001-pages-meta-history</em> (All pages with complete page edit history (.bz2)) found at&nbsp;<a href="https://dumps.wikimedia.org/wikidatawiki/">https://dumps.wikimedia.org/wikidatawiki/</a>.</p> <p>The extract&nbsp;was created by the following SQL query:</p> <pre> SELECT page_title, rev_comment, rev_user_text, rev_timestamp FROM revisions WHERE rev_comment LIKE &#39;%[[Property:%]]%[[Q%&#39; ORDER BY rev_id INTO OUTFILE &#39;edits_wikidatawiki-20181001-pages.csv&#39;; </pre> <p>&nbsp;</p> <p><strong>wikidata-20180813-all.json.bz2.universe.noattr.gt.bz2</strong></p> <p>This is a graph-tool representation of the WikiData graph. Output of&nbsp;<a href="https://github.com/eXascaleInfolab/cardinal/blob/master/1_create_inmemory_graph.py">https://github.com/eXascaleInfolab/cardinal/blob/master/1_create_inmemory_graph.py</a>.</p> <p><strong>observations_wikidatawiki-20181001-pages.pickle</strong></p> <p>Extracted observations. Output of&nbsp;<a href="https://github.com/eXascaleInfolab/cardinal/blob/master/2_extract_observations.py">https://github.com/eXascaleInfolab/cardinal/blob/master/2_extract_observations.py</a>.</p> <p>&nbsp;</p> <p><strong>estimates_wikidatawiki-20181001-pages.pickle</strong></p> <p>Extracted estimates. Output of&nbsp;<a href="https://github.com/eXascaleInfolab/cardinal/blob/master/3_calculate_estimates.py">https://github.com/eXascaleInfolab/cardinal/blob/master/3_calculate_estimates.py</a></p> <p>&nbsp;</p> <p><strong>results_wikidatawiki-20181001-pages.pickle&nbsp;</strong></p> <p>Results. Output of&nbsp;<a href="https://github.com/eXascaleInfolab/cardinal/blob/master/4_draw_graphs.py">https://github.com/eXascaleInfolab/cardinal/blob/master/4_draw_graphs.py</a></p>

opencc-zeroJul 2019View details →
zenodo44/100

Time-resolved compound repositioning predictions on a text-mined knowledge network

<p><strong>gs_positives.csv</strong>: The re-processed version of DrugCentral indications, utilized as training and testing positives in the analysis.</p> <p><strong>top_5000_predictions.csv</strong>: The top 5000 drug-disease pairs, by probability, produced by this analysis pipeline.</p> <p><strong>file_info.txt</strong>: Information about the column headings in of the two files.</p> <p>&nbsp;</p>

opencc-by-sa-4.0Sep 2019View details →
zenodo44/100

Monitoring knowledge, risk perceptions, preventive behaviours and trust to inform pandemic outbreak response.

<p>The study is part of the large project promoted by WHO Regional Office for Europe called &ldquo;<em>Monitoring knowledge, risk perceptions, preventive behaviours and trust to inform pandemic outbreak response</em>&rdquo; and carried out in over 30 countries of the WHO European Region (Registered ISRCTN on 11/05/2021, ID: ISRCTN26200758). In Italy, the survey was conducted administering an online questionnaire developed <em>ad hoc</em> by the WHO in four waves (January-May 2021) to a sample of 10.000 individuals aged 18-70 years. A detailed sampling plan was developed to obtain a representative sample of the Italian adult population. The following variables were taken into account for stratification of the participants: gender by age (four age groups: 18-34 years, 35-44 years, 45-54 years, 55-70 years); geographical area (four areas: North West, North East, Centre, South and Islands); size of living centers (two classes: above and below 100,000 inhabitants); level of education (up to lower middle school, beyond lower middle school); and employment situation (employed, not employed). At the end of each survey&rsquo;s wave, a weighting procedure has been applied to accurately restore the proportionality of the total sample examined with the reference population, according to the most recent data of the Italian Statistics Institute (ISTAT, 12/31/2019). In particular, data have been weighted for the main socio-demographic and geographic variables (e.g., sex by age by geographical area, occupation, educational qualification, geographical area by size of living centers). The sample size made it possible to maintain a sampling error of less than 2% (at the significance level of 95%) and to control the error of estimates within groups or subgroups of interest. The interviews were conducted by Doxa S.p.a. and carried out with the CAWI technique (Computer Assisted Web Interviewing) on an online panel and on the Confirmit software platform used by Doxa S.p.a. The average administration time was about 18-20 minutes. This study was approved by the Ethics Committee of the IRCCS San John of God Fatebenefratelli of Brescia (n&deg; 72-2020), and all participants provided written informed consent.</p> <p>The primary objectives are to:</p> <p>● Monitor variables that are critical for population behaviour to control transmission of the novel coronavirus, including risk perceptions, knowledge, self-efficacy, confidence in institutions, behaviours, rumours, affect, worry, resilience, trust in/use of information sources and more.<br> ● Document changes over time in these factors to understand the effect of the pandemic process, new developments, events or measures taken.<br> ● Monitor possible issues, e.g. related to misinformation or distrust, as they emerge, to allow early response.<br> ● Identify relationships between variables to identify levers for effective and appropriate responses.<br> ● Explore the relationship of psychological variables (e.g. worry, resilience, trust, affect) with the epidemiological situation and the events and measures taken.<br> ● Identify gaps between perceived and actual knowledge.<br> ● Evaluate the effectiveness of pandemic response measures, and the acceptance and effectiveness of policies and restrictions implemented, including the easing of such restrictions.<br> The secondary objectives are to:<br> ● Contribute to post-outbreak evaluation, thereby contributing to the continued regional/global efforts to better understand mechanisms of crisis response.<br> ● If additional research capacity is available, the data can be triangulated with data on media reporting, COVID-19 cases and other.● If additional research capacity is available, the data can be triangulated with data on media reporting, imported or confirmed cases, etc.: The relationship between psychological variables and characteristics of the outbreak situation can be explored (i.e. how closely the perceived risk mirrors reported cases, relative import risk, media reports).<br> This approach allows a citizen-centred approach where insights into population perceptions and behaviours inform COVID-19 actions, alongside epidemiological data and considerations of economic, cultural, ethical, structural political nature and other.</p> <p>The WHO questionnaire includes 21 different thematic areas noteworthy for the investigation of COVID-19 experience. The questionnaire was translated into specific country language by each recruiting site, following the WHO&rsquo;s guidelines for translations of tools into other languages. The process included the following steps: forward translation, panel experts, back-translation, pre-test and cognitive interviews and, finally, development of the final version. Variables being surveyed include the following:<br> &bull; Socio-demography;<br> &bull; COVID-19 personal experience;<br> &bull; Health literacy;<br> &bull; COVID-19 risk perception;<br> &bull; Probability and Severity;<br> &bull; Preparedness and Perceived self-efficacy;<br> &bull; Prevention &ndash; own behaviours;<br> &bull; Affect;<br> &bull; Trust in sources of information;<br> &bull; Use of sources of information;<br> &bull; Frequency of Information;<br> &bull; Trust in institutions (perceptions);<br> &bull; Policies, interventions (perceptions);<br> &bull; Conspiracies (perceptions);<br> &bull; Resilience (perceptions);<br> &bull; Testing and tracing;<br> &bull; Fairness (perceptions);<br> &bull; Lifting restrictions (pandemic transition phase);<br> &bull; Unwanted behaviour;<br> &bull; Wellbeing;<br> &bull; COVID-19 vaccine.</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2020View details →
zenodo44/100

Knowledge repository for Non-Wood Forest Products - Dataset of factsheets from INCREDIBLE

<p>The<strong><em> Knowledge Repository for Non-Wood Forest Products</em></strong> is a collection of information on innovation about non-wood forest products gathered from experts and practitioners during the <a href="https://incredibleforest.net/">INCREDIBLE project</a>. It brings together, in a single platform, knowledge about cork, pine oleoresin, wild mushrooms &amp; truffles, aromatic &amp; medicinal plants and wild nuts &amp; berries, around various themes, from Portugal, Spain, France, Italy, Croatia, Greece and Tunisia.</p> <p>Each piece of knowledge is summarised in a factsheet, that can concern one or several non-wood forest products, as some issues or solutions are transversal. The factsheets can either contain practical knowledge (success stories, good practices, technical reports) or more theoretical or academic results (research results, databases, policies). The factsheets can also be identified by the position in the value chain to which the knowledge applies, from forestry to end-consumers.</p>

opencc-by-4.0Jun 2021View details →
zenodo44/100

Software Engineering Education Knowledge versus Industrial Needs

<p>Dataset of the research paper:&nbsp;<strong>Software Engineering Education Knowledge versus Industrial&nbsp;Needs</strong></p> <p><em>Contribution</em>: Determine and analyze the gap between software practitioners&rsquo; education outlined in the 2014 IEEE/ACM Software Engineering Education Knowledge (SEEK) and industrial needs pointed by Wikipedia articles referenced in Stack Overflow (SO) posts.<br> <em>Background</em>: Previous work has uncovered deficiencies in the coverage of computer fundamentals, people skills, software processes, and human-computer interaction, suggesting rebalancing.<br> <em>Research Questions</em>: 1) To what extent are developers&rsquo; needs, in terms of Wikipedia articles referenced in SO posts, covered by the SEEK knowledge units? 2) How does the popularity of Wikipedia articles relate to their SEEK coverage? 3) What areas of computing knowledge can be better covered by the SEEK knowledge units? 4) Why are Wikipedia articles covered by the SEEK knowledge units cited on SO?<br> <em>Methodology</em>: Wikipedia articles were systematically collected from SO posts. The most cited were manually mapped to the SEEK knowledge units, assessed according to their degree of coverage. Articles insufficiently covered by the SEEK were classified by hand using the 2012 ACM Computing Classification System. A sample of posts referencing sufficiently covered articles was manually analyzed. A survey was conducted on software practitioners to validate the study findings.<br> <em>Findings</em>: SEEK appears to cover sufficiently computer science fundamentals, software design and mathematical concepts, but less so areas like the World Wide Web, software engineering components, and computer graphics. Developers seek advice, best practices and explanations about software topics, and code review assistance. Future SEEK models and the computing education could dive deeper in information systems, design, testing, security, and soft skills.</p> <p>The following data files are included.</p> <ul> <li><strong>wikipedia_articles.csv</strong>: Wikipedia articles mapped to the knowledge units of the 2014 IEEE/ACM Software Engineering Education Knowledge (SEEK) and the first and second level categories of the 2012 ACM Computing Classification System (CCS).</li> <li> <p><strong>posts_analysis.csv</strong>: Stack Overflow post data and metadata.</p> </li> <li> <p><strong>posts_aggregated_codes.csv</strong>: The aggregated codes that resulted from the manual analysis of the Stack Overflow posts by grouping individual keywords assigned to the posts.</p> </li> <li> <p><strong>survey_questionnaire.csv</strong>:&nbsp;The final survey questionnaire.</p> </li> <li> <p><strong>survey_responses.csv</strong>:&nbsp;Anonymized responses of the final survey questionnaire. (E-mail addresses have been excluded for privacy reasons.)</p> </li> </ul>

opencc-by-4.0Jul 2021View details →
zenodo44/100

Code and data for "KGML-ag: A Modeling Framework of Knowledge-Guided Machine Learning to Simulate Agroecosystems: A Case Study of Estimating N2O Emission using Data from Mesocosm Experiments "

<p>This is code and data for manuscript:&nbsp;<br> &quot;KGML-ag: A Modeling Framework of Knowledge-Guided Machine Learning to Simulate Agroecosystems:&nbsp;<br> A Case Study of Estimating N<sub>2</sub>O Emission using Data from Mesocosm Experiments&quot;<br> Licheng Liu, Shaoming Xu, Zhenong Jin*, Jinyun Tang, Kaiyu Guan, Timothy J. Griffis,&nbsp;<br> Matt D. Erickson, Alexander L. Frie, Xiaowei Jia, Taegon Kim, Lee T. Miller, Bin Peng, Shaowei Wu, Yufeng Yang, Wang Zhou, Vipin Kumar</p> <p>All the files belong to Prof. Zhenong Jin, University of Minnesota, UA. jinzn@umn.edu<br> &quot;code&quot; foler includes code for data processing, model training, and results plotting.<br> &quot;trained_model_saved&quot; includes all trained model so you can use to reproduce the results showed in the study;<br> &quot;data&quot; includes all data presented in the study. Finetuning data is refering to&nbsp;Miller, L.T. , Griffis, T. J., Erickson, M. D.,&nbsp; Turner, P. A., Deventer, M. J., Chen, Z., Yu,&nbsp; Z., Venterea, R.T., Baker, J. M., and Frie, A. L. (2021). Response of nitrous oxide emissions to future changes in precipitation and individual rain events. Journal of Environmental Quality, In review</p>

opencc-by-4.0Sep 2021View details →
zenodo44/100

Iterative Bleaching Extends Multiplexity (IBEX) Knowledge-Base

<p>The Iterative Bleaching Extends Multiplexity (IBEX) imaging method is an iterative immunolabeling and chemical bleaching method that enables highly multiplexed imaging of diverse tissues. Development of the <a href="https://doi.org/10.1038/s41596-021-00644-9">IBEX method</a> and <a href="https://github.com/niaid/imaris_extensions">related software</a> was led by Dr. Andrea Radtke and Dr. Ziv Yaniv. <a href="https://doi.org/10.1073/pnas.2018488117">IBEX</a> and related methods, <a href="https://doi.org/10.1073/pnas.1708981114">Ce3D</a>, <a href="https://doi.org/10.1111/imr.13052">Ce3D-IBEX</a>, <a href="https://doi.org/10.1073/pnas.2018488117">Opal-plex</a>, were originally developed in the laboratory of <a href="https://www.niaid.nih.gov/research/ronald-n-germain-md-phd">Dr. Ronald N. Germain</a>, US National Institutes of Health.</p><p>The IBEX Imaging Community is an international group of scientists committed to sharing knowledge related to multiplexed imaging in a transparent and collaborative manner. This open, global repository is a central resource for reagents, protocols, panels, publications, software, and datasets. In addition to IBEX, we support standard, single cycle multiplexed imaging (Multiplexed 2D imaging), volume imaging of cleared tissues with clearing enhanced 3D (Ce3D), highly multiplexed 3D imaging (Ce3D-IBEX), and extension of the IBEX dye inactivation protocol to the Leica Cell DIVE (Cell DIVE-IBEX). This dataset contains the current state of knowledge with respect to the IBEX microscopy imaging protocol.</p><p>How to use the Knowledge-Base:</p><ol><li>Save a copy to your computer.</li><li>To find a reagent: Open the reagent_resources.csv file found in the data directory. Use a spreadsheet application to filter the columns based on target name, target species, vendor, etc.</li><li>To view a complete list of fluorescent probes tested by the IBEX imaging community: Open the fluorescent_probes.csv file. This file reports the spectral properties and inactivation conditions of each fluorescent probe.</li><li>To import publications cited in the Knowledge-Base, import the publications.bib file found in the data directory to your reference manager.</li><li>To view a local copy of the website: Open the index.md file found in the docs directory using a markdown editor such as the free <a href="https://code.visualstudio.com/">Visual Studio Code</a>.</li><li>To view supporting information for a reagent (images, publications, notes): Open a specific target-conjugate-orcid combination under the docs-supporting_material directory structure using a markdown editor. This can also be visualized from the <a href="https://ibeximagingcommunity.github.io/ibex_imaging_knowledge_base/reagent_resources.html">Reagent Resources page</a> and filtered using a catalog number or other unique identifier in your web browser.</li></ol><p></p><p>Join the <a href="https://ibeximagingcommunity.github.io/ibex_imaging_knowledge_base/">online IBEX Imaging community</a> and contribute your knowledge. For more details on how to contribute, see <a href="https://ibeximagingcommunity.github.io/ibex_imaging_knowledge_base/contrib.html">these instructions</a>.</p><p>This research was supported by:</p><ul><li>The Intramural Research Program of the NIH, National Institute of Allergy and Infectious Diseases and National Cancer Institute, under grants 1ZIAAI001290-02, 1ZIAAI000545-33, 1ZIAAI000758-24, 1ZIAAI000974-16, 1ZIAAI001034-14.</li><li> The Wellcome Trust, under grant 224586/Z/21/Z.</li><li> The National Institute of Allergy and Infectious Diseases, NIH, under grant 1ZIAAI001343-01.</li></ul><p></p>

opencc-by-4.0Jul 2024View details →
zenodo44/100

Global knowledge and use of soil biodiversity: Results of an expert survey

<p>A global survey on soil biodiversity (see file Global Biodiversity Survey.pdf provided as an attachment) was conducted over a three-week period in March 2022 by the Global Soil Partnership (GSP) of the Food and Agriculture Organization (FAO) of the United Nations, as part of the activities of the International Network on Soil Biodiversity (NETSOB). The survey intended to obtain information on the current status of knowledge and use of soil organisms worldwide, i.e., to identify who is doing what, where, and how, as well as the main gaps, pitfalls, and opportunities across existing national initiatives and research.</p> <p>The survey included 122 questions that characterized the work undertaken by experts regarding microbes, fauna, and their activity in soils, community &amp; functional assessments, inventories, mapping and monitoring activities, ecosystem services, applications, and threats to soil biodiversity, education, and communication activities, as well as public policies related to soil biodiversity. The online survey was created using the software Survey Monkey v. 11 and was sent out to over 70 thousand e-mail addresses with a link to complete the survey.&nbsp;</p> <p>Over 2,600 responses were received, representing &gt;1,350 institutions from 135 countries, mainly from experts active in research and academia. The number of respondents was not equal for all questions, as the survey guided the respondents to different parts, depending on their replies.</p> <p>The 122 questions and the replies of the respondents are presented as separate tabs in the attached Excel file (Results survey for Zenodo.xlsx). The respondents and their identities, as well as their e-mails and any personal websites were removed in the current file to maintain anonymity. Institutional websites were maintained as long as they did not identify the respondent(s) directly.&nbsp;</p> <p>A detailed written description of the survey results was prepared as a manuscript for a special issue of the journal Soil Organisms, volume 97 (Brown et al., 2025). The survey was prepared by a team of scientists from the Brazilian Corporation for Agricultural Research (Embrapa) and collaborating institutions, with assistance from the board of the International Network on Soil Biodiversity (NETSOB), and with funding provided by the FAO. The work was further supported by the Funda&ccedil;&atilde;o de Apoio a Pesquisa e Desenvolvimento Agropecu&aacute;rio Edmundo Gastal (FAPEG), Brazil, a grant of CNPq (Processo No. 312824/2022-0) to GGB, and of the Natural Sciences and Engineering Research Council of Canada (NSERC) Discovery Grant program (# 05901&ndash;2019) to ZL, who was also supported by Western University.</p>

opencc-by-4.0Oct 2024View details →
zenodo44/100

Universal Knowledge Graph Embeddings

<p>The dataset provides embeddings for entities and relations in DBpedia (English) and Wikidata. The two knowledge graphs are first merged using a&nbsp;novel approach that we developed&nbsp;by leveraging the sameAs links between them. Then, we used the state-of-the-art embedding model ConEx to compute embeddings of the merge. Our embeddings are called universal knowledge graph embeddings.</p>

opencc-by-4.0Dec 2022View details →
zenodo44/100

Social knowledge and irony understanding

<p>Dataset accompanying the publication titled &quot;Impact of social knowledge about the speaker on irony understanding: Evidence from neural oscillations.&quot; It contains EEG recording (.set and .fdt files) of 22 participants reading stories during a task of irony undertsanding. Stimuli were manipulated according to a context condition (ironic, literal) and a speaker occupation stereotype (no occupation, non-sarcastic, sarcastic).</p> <p>Matlab scripts used for preprocessing and time-frequency analysis are available on github:&nbsp;&nbsp;<a href="https://github.com/deebeebolger/project_Ironie2022.git">https://github.com/deebeebolger/project_Ironie2022.git</a></p>

opencc-by-4.0Dec 2022View details →
zenodo44/100

cultural-ai/wordsmatter: Words Matter: a knowledge graph of contentious terms

<p>The choice of words describing cultural heritage can cause debates. It is especially sensitive when artefacts relate to different cultures and peoples who have been historically marginalised. Words chosen by archivists or curators may transmit stereotypes. The cultural heritage community has produced knowledge on potentially stereotyping and offensive terminology in heritage collections. At the same time, their knowledge is difficult to incorporate into existing online collections unless this knowledge is structured and machine-readable.</p> <p>The Words Matter Knowledge Graph represents domain expert knowledge on discussions about contentious terminology in the cultural sector. In the knowledge graph, 75 English and 83 Dutch contentious terms are linked to explanations of their usage and suggested alternatives from domain experts. There are also related matches between contentious terms and sources from external datasets: Wikidata, Princeton WordNet, Open Dutch WordNet, and Getty Art &amp; Architecture Thesaurus.</p> <p>This Zenodo publication includes the CULCO scheme used to model contentious terms in the knowledge graph. The scheme documentation is <a href="https://cultural-ai.github.io/wordsmatter/" target="_blank" rel="noopener">available on a separate page</a>.</p> <p>This knowledge graph is <a href="https://amsterdam.wereldmuseum.nl/en/about-wereldmuseum-amsterdam/research/words-matter-publication" target="_blank" rel="noopener">based</a> on the publication &ldquo;Words Matter: An Unfinished Guide to Word Choices in the Cultural Sector&rdquo; by the National Museum of World Cultures (NMVW).&nbsp;</p> <p><a href="https://doi.org/10.1007/978-3-031-33455-9_30" target="_blank" rel="noopener">Read more</a> about this work in the paper "A Knowledge Graph of Contentious Terminology for Inclusive Representation of Cultural Heritage" (2023) by Andrei Nesterov,&nbsp;Laura Hollink,&nbsp;Marieke van Erp &amp;&nbsp;Jacco van Ossenbruggen.</p> <p>In this version:</p> <ul> <li>the CULCO scheme documentation is updated</li> <li>versioning is fixed</li> <li>typos are corrected</li> </ul>

opencc-by-sa-4.0Mar 2023View details →
zenodo44/100

LauNuts: A Knowledge Graph to identify and compare geographic regions in the European Union

<p><strong>LauNuts</strong> is a RDF Knowledge Graph consisting of:</p> <ul> <li>Local Administrative Units (LAU) and</li> <li>Nomenclature of Territorial Units for Statistics (NUTS)</li> </ul> <p><a href="https://w3id.org/launuts">https://w3id.org/launuts</a></p>

opencc-by-4.0Dec 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record