Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

915

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

915 results for “Graph”

Learn how ShareScore rates datasets ↗
zenodo44/100

scGraph2Vec: a deep generative model for gene embedding augmented by Graph Neural Network and single-cell omics data

<p>This repository contains the training data and source code to reproduce the results of our paper:<br>scGraph2Vec: a deep generative model for gene embedding augmented by Graph Neural Network and single-cell omics data</p> <p>More description can be also found in GitHub (https://github.com/LPH-BIG/scGraph2Vec).</p>

opencc-zeroJun 2024View details →
zenodo44/100

Tree Annotation Vocabulary (TAV) - Knowledge Graph and Annotated Dataset

<p>This dataset contains all the files used in developing the Tree-KG, the knowledge graph to capture the tree annotations in the works of Vladimir Nabokov.&nbsp;</p> <p>In the Annotated Dataset folder, 6 spreadsheets in excel (.xlsx) format are provided. They are numbered. Note that annotated data are all in English as the consulted works are the English translations of the literary works of Nabokov.</p> <p>(1) contains the tree annotations from the novels originally written in Russian by Vladimir Nabokov.</p> <p>(2) contains the tree annotations from the novels originally written in English by Vladimir Nabokov.</p> <p>(3) contains the tree annotations from the short stories originally written in Russian and English by Vladimir Nabokov.</p> <p>(4) is the knowledge base (KB) developed to link the annotated trees to Wikidata and DBPedia.</p> <p>(5) is the benchmarking results of some entity recognition tools. It includes the relevant passages from Nabokov's novels that were used in the experiments as well as the prompts used in getting the results.</p> <p>(6) represents the complete bibliographic details of the works of Vladimir Nabokov (https://thenabokovian.org/abbreviations).</p> <p>In the Ontology Versions folder, four ontology (TAV) files in turtle (.ttl) format are provided. They are all numbered and dated to represent their different versions. Some sample SPARQL queries are provided in a .txt file. The KG was developed on Prot&eacute;g&eacute;.&nbsp;</p> <p>(1) contains the essential schema for the TAV vocabulary.</p> <p>(2) contains the schema for TAV vocabulary with links to external vocabularies (Schema.Org; Open Annotation, etc.).&nbsp;</p> <p>(3) contains the Tree-KG in so far it reflects data from three novels (Mary; King, Queen, Knave; Glory).</p> <p>(4) contains the entire Tree-KG based on all the works mentioned in the excel sheets (20 books).</p> <p>(5) contains some sample SPARQL queries (.txt) file.</p>

opencc-by-4.0Aug 2024View details →
zenodo44/100

Global Biotic Interactions: Taxon Graph hash://sha256/0b58753e4ff5519442689d866c0f1d19ffa7d97f917144df1d1cd56ea756921d hash://md5/b23bd0210c88ca10c3e3253091f4fdfa

<p>Global Biotic Interactions: Taxon Cache and Taxon Map</p> <p>Global Biotic Interactions (GloBI) provides access to existing species interaction datasets (Poelen et al. 2014, http://globalbioticinteractions.org). As part of the dataset integration and aggregation, a best effort is made to resolve, match and link taxonomic names and associated vernacular/common names, hierarchies and thumbnails.&nbsp;</p> <p>The data archives included in this publication contain established taxonomic links (taxonMap.tsv.gz) and taxonomic information (taxonCache.tsv.gz) that GloBI retrieved and integrated from taxonomic name sources and web services associated with http://itis.gov, http://globalnames.org, http://eol.org and others open data services.&nbsp;</p> <p>While GloBI is not a naming authority and the primary goal of the name matching process is to detect incorrect or outdates names, the archives may serve as an example of how to publish denormalized taxonomic records and their interrelatioships in a pragmatic way.</p> <p>For related discussion threads, see https://github.com/globalbioticinteractions/globalbioticinteractions/issues/145 , https://github.com/globalbioticinteractions/globalbioticinteractions/issues/274 , https://github.com/globalbioticinteractions/globalbioticinteractions/issues/70 , https://github.com/EOL/tramea/issues/10 and https://github.com/globalbioticinteractions/globalbioticinteractions/issues/274 .</p> <p>Files<br>&nbsp;&nbsp;<br>&nbsp; README&nbsp;<br>&nbsp; &nbsp; &nbsp; this file</p> <p>&nbsp; taxonCache.tsv.gz&nbsp;<br>&nbsp; &nbsp; &nbsp;Taxonomic name, ids, hierarchies, common names and thumbnail associated to taxa known to GloBI.&nbsp;<br>&nbsp;<br>&nbsp; taxonCache.tsv.sha256<br>&nbsp; &nbsp; &nbsp;sha256 hash of taxonCache.tsv</p> <p>&nbsp; taxonCacheFirst10.tsv<br>&nbsp; &nbsp; &nbsp; Header and 10 following lines from taxonCache.tsv</p> <p>&nbsp; taxonCacheFirst10.tsv.sha256<br>&nbsp; &nbsp; &nbsp; sha256 hash of taxonCacheFirst10.tsv<br>&nbsp; &nbsp; &nbsp; &nbsp;<br>&nbsp; taxonMap.tsv.gz&nbsp;<br>&nbsp; &nbsp; &nbsp; Links between taxon name and ids across various taxon providers.&nbsp;</p> <p>&nbsp; taxonMap.tsv.sha256&nbsp;<br>&nbsp; &nbsp; &nbsp; sha256 hash of taxonMap.tsv</p> <p>&nbsp; taxonMapFirst10.tsv<br>&nbsp; &nbsp; &nbsp; Header and 10 following lines from taxonMap.tsv<br>&nbsp;<br>&nbsp; taxonMapFirst10.tsv.sha256<br>&nbsp; &nbsp; &nbsp; sha256 hash of taxonMapFirst10.tsv</p> <p>&nbsp; prefixes.tsv<br>&nbsp; &nbsp; &nbsp; Term prefixes and their associated uri schemes.&nbsp;</p> <p>&nbsp; names.tsv.gz<br>&nbsp; &nbsp; &nbsp; Corpus of names used to resolve and link. Generated using https://github.com/globalbioticinteractions/elton .</p> <p>&nbsp; names.tsv.sha256<br>&nbsp; &nbsp; &nbsp; sha256 hash of names.tsv</p> <p>&nbsp; namesUnresolved.tsv.gz<br>&nbsp; &nbsp; &nbsp; Names that are not (yet) linked to name sources using https://github.com/globalbioticinteractions/nomer .</p> <p>&nbsp; namesUnresolved.tsv.sha256<br>&nbsp; &nbsp; &nbsp; sha256 hash of namesUnresolved.tsv&nbsp;</p> <p>Column Descriptions</p> <p>&nbsp; taxonCache.tsv.gz&nbsp;</p> <p>&nbsp; &nbsp; 1 | id<br>&nbsp; &nbsp; 2 | name<br>&nbsp; &nbsp; 3 | rank<br>&nbsp; &nbsp; 4 | commonNames<br>&nbsp; &nbsp; 5 | path<br>&nbsp; &nbsp; 6 | pathIds&nbsp;<br>&nbsp; &nbsp; 7 | pathNames<br>&nbsp; &nbsp; 8 | externalUrl<br>&nbsp; &nbsp; 9 | thumbnailUrl<br>&nbsp;<br>&nbsp; taxonMap.tsv.gz</p> <p>&nbsp; &nbsp; 1 | providedTaxonId<br>&nbsp; &nbsp; 2 | providedTaxonName<br>&nbsp; &nbsp; 3 | resolvedTaxonId<br>&nbsp; &nbsp; 4 | resolvedTaxonName</p> <p>&nbsp; names.tsv.gz</p> <p>&nbsp; &nbsp; 1 | providedTaxonId<br>&nbsp; &nbsp; 2 | providedTaxonName</p> <p>&nbsp; &nbsp;namesUnresolved.tsv.gz</p> <p>&nbsp; &nbsp; 1 | providedTaxonId<br>&nbsp; &nbsp; 2 | providedTaxonName</p> <p>References</p> <p>Jorrit H. Poelen, James D. Simons and Chris J. Mungall. (2014). Global Biotic Interactions: An open infrastructure to share and analyze species-interaction datasets. Ecological Informatics. https://doi.org/10.1016/j.ecoinf.2014.08.005.</p> <p>Updates</p> <p>org.globalbioticinteractions.taxon v0.3, 2018-03-02</p> <p>This taxon archive version was created by taking GloBI taxon v0.2 (Jan 2018) and appending a semi-automatically created WikiData taxon mapping and taxon cache.</p> <p>org.globalbioticinteractions.taxon v0.3.1, 2018-04-05</p> <p>This taxon archive version was created by taking GloBI taxon v0.2 (Jan 2018) and appending an automatically created WikiData taxon mapping and taxon cache using Apache Spark scripts at https://github.com/bio-guoda/guoda-datasets/tree/master/wikidata .</p> <p>org.globalbioticinteractions.taxon v0.3.2, 2018-05-21</p> <p>This taxon archive version includes the following:</p> <p>1. all lines in taxonMap.tsv.gz v0.3.1 that passed all validate-term-link tests defined in nomer v0.0.7 (see https://doi.org/10.5281/zenodo.1249964 or https://github.com/globalbioticinteractions/nomer/releases/tag/0.0.7).</p> <p>2. all lines in taxonCache.tsv.gz. v0.3.1 that passed all validate-term tests defined in nomer v0.0.7&nbsp;</p> <p>3. all lines in 1. that did *not* pass the validate-term test, were re-resolved using nomer v0.0.7 commands "append globi-enrich" and "append globi-globalnames". Only SAME_AS and SYNONYM_OF matches were used to generate new entries for taxonCache and taxonMap.</p> <p>4. in addition, elton v0.4.5 (see https://doi.org/10.5281/zenodo.1212599 or https://github.com/globalbioticinteractions/elton/releases/tag/0.4.5) was used to generate an up-to-date names list by running the "update" and "names" commands on 18-19 May 2018. Of the resulting names, only id/names pairs that were unknown to the taxon graph were resolved using the "append globi-enrich" and "append globi-globalnames" commands of nomer v0.0.7. Only matches classified as SAME_AS and SYNONYM_OF were used to generate new entries for taxonCache and taxonMap.</p> <p>5. the updated versions of taxonMap.tsv.gz and taxonCache.tsv.gz were produced by appending result of 1., 2., 3. and 4. , removing duplicate lines and sorting the result.&nbsp;</p> <p>6. finally, the resulting taxonMap.tsv.gz. and taxonCache.tsv.gz files were validated using the nomer v0.0.7 validate-term-link and validate-term commands, respectively. The result indicated that all lines (other than the header) passed the validation tests.</p> <p>org.globalbioticinteractions.taxon v0.3.3, 2018-06-12</p> <p>This taxon archive version includes the following:</p> <p>1. normalizing taxonomic ranks using nomer's taxon rank matcher</p> <p>2. include more manual taxonomic name mappings provided by Brian Hayden and collaborators.</p> <p>3. update taxonCache and taxonMap using automated scripts available at https://doi.org/10.5281/zenodo.1286023 .&nbsp;</p> <p>4. remove mapping to NCBI taxa with name "Small" (and associated OTT).</p> <p><br>org.globalbioticinteractions.taxon v0.3.4, 2018-06-27</p> <p>1. update taxonCache and taxonMap using automated scripts available at https://doi.org/10.5281/zenodo.1286023</p> <p>Please note that nomer and elton rely on web accessible apis like taxonomy resolution services and data portals. This dependence on external web-only accessible services might make reproduction of the results tricky due to network outages, server failures, upgrades, downgrades, data loss and/or abandonment of informatics projects/ datasets.&nbsp;</p> <p>org.globalbioticinteractions.taxon v0.3.5, 2018-06-28</p> <p>1. remove dubious provided name from taxon map. Names include "no name", "unidentified".<br>2. remove dubious mappings to Pavlova (e.g., Unidentified Amoebozoa -&gt; Pavlova). Related to 1.<br>3. remove dubious mappings to resolve taxa that include names like "unidentified" or "organic species"<br>4. removed dubious mappings to "Boiga dendrophila"<br>5. removed dubious mappings from "Chaetognatha" (arrowworm) to a suspected homonym Lepidoptera GBIF:3257692 and IRMNG:1252651<br>6. removed dubious mappings from "small sharks" to multiple NCBI/OTT terms with name "Small"</p> <p>Please note that nomer and elton rely on web accessible apis like taxonomy resolution services and data portals. This dependence on external web-only accessible services might make reproduction of the results tricky due to network outages, server failures, upgrades, downgrades, data loss and/or abandonment of informatics projects/ datasets.</p> <p>org.globalbioticinteractions.taxon v0.3.6, 2018-09-10</p> <p>1. update taxonCache and taxonMap using automated scripts available at https://doi.org/10.5281/zenodo.1286023</p> <p>org.globalbioticinteractions.taxon v0.3.7, 2018-10-18</p> <p>1. update taxonCache and taxonMap using automated scripts available at https://doi.org/10.5281/zenodo.1286023<br>2. remove dubious mapping to Vertebrata (WORMS:370321 , http://www.marinespecies.org/aphia.php?p=taxdetails&amp;id=370321). Also see https://github.com/globalbioticinteractions/globalbioticinteractions/issues/361 .<br>3. remove dubious mapping to NCBITaxon:1585532 (Beta vulgaris/Cercospora beticola mixed EST library). Also see https://github.com/globalbioticinteractions/globalbioticinteractions/issues/346 and https://github.com/Planteome/samara/issues/50&nbsp;</p> <p>org.globalbioticinteractions.taxon v0.3.8, 2018-11-15</p> <p>1. update taxonCache and taxonMap using automated scripts available at https://doi.org/10.5281/zenodo.1286023</p> <p>org.globalbioticinteractions.taxon v0.3.9, 2018-11-23</p> <p>1. label deprecated EOL ids by applying patches in http://doi.org/10.5281/zenodo.1495266 to taxonMap.tsv.gz and taxonCache.tsv.gz . Related to https://github.com/globalbioticinteractions/globalbioticinteractions/issues/383 .<br>2. remove all Encyclopedia of Life thumbnail urls from taxonCache. Related to https://github.com/globalbioticinteractions/globalbioticinteractions/issues/381 .<br>3. remove Encyclopedia of Life external urls associated with deprecated ids from taxonCache.&nbsp;</p> <p><br>org.globalbioticinteractions.taxon v0.3.10, 2018-11-26</p> <p>1. Remove suspicious name mappings related to Humpback scorpionfish (Scorpaenopsis gibbosa) by applying patch published in Poelen, Jorrit H. (2018). Global Biotic Interactions: Taxon Graph Patches (Version 0.2. [Data set]. Zenodo. http://doi.org/10.5281/zenodo.1560662&nbsp;</p> <p>org.globalbioticinteractions.taxon v0.3.11, 2018-12-21</p> <p>1. update taxonCache and taxonMap using automated scripts available at https://doi.org/10.5281/zenodo.1286023<br>2. remove suspicious name mappings using: ```zcat taxonMap.tsv.gz | grep -v -i -P "\tnone\t" | grep -v -P "(GBIF|IRMNG):.*\tBrachyura$" | grep -v -P "Gamarus" | &nbsp;grep -v -P "^EOL:1047365\ttrachurus trachurus" | grep -v -P "Loros\t.*Psittacidae" | grep -v -P "(GBIF|IRMNG).*Lucifer$" | grep -v -P "GBIF.*Diadema$" | gzip &gt; taxonMapUpdated.tsv.gz```</p> <p>org.globalbioticinteractions.taxon v0.3.12, 2019-06-05</p> <p>1. update taxonCache and taxonMap using automated scripts available at https://doi.org/10.5281/zenodo.3240558</p> <p>org.globalbioticinteractions.taxon v0.3.13, 2019-06-12</p> <p>1. update taxonCache and taxonMap using automated scripts available at https://doi.org/10.5281/zenodo.3240558</p> <p>org.globalbioticinteractions.taxon v0.3.14, 2019-08-19</p> <p>1. revisit deprecated EOL ids by applying patches in http://doi.org/10.5281/zenodo.3371634 to taxonMap.tsv.gz and taxonCache.tsv.gz . Related to https://github.com/jhpoelen/eol-globi-data/issues/403 .</p> <p>org.globalbioticinteractions.taxon v0.3.15, 2019-08-26</p> <p>1. update taxonCache and taxonMap using automated scripts available at https://doi.org/10.5281/zenodo.3240558</p> <p>org.globalbioticinteractions.taxon v0.3.16, 2019-09-22</p> <p>1. revisit deprecated EOL ids by applying patches in http://doi.org/10.5281/zenodo.3457626 to taxonMap.tsv.gz and taxonCache.tsv.gz of http://doi.org/10.5281/zenodo.3378125. Related to https://github.com/globalbioticinteractions/globalbioticinteractions/issues/408 .</p> <p>org.globalbioticinteractions.taxon v0.3.17, 2019-09-27</p> <p>1. update taxonCache and taxonMap using automated scripts available at https://doi.org/10.5281/zenodo.3240558</p> <p>org.globalbioticinteractions.taxon v0.3.18, 2019-10-30</p> <p>1. update taxonCache and taxonMap using automated scripts available at https://doi.org/10.5281/zenodo.3240558</p> <p>org.globalbioticinteractions.taxon v0.3.19, 2019-11-07</p> <p>1. update taxonCache and taxonMap using automated scripts available at https://doi.org/10.5281/zenodo.3240558</p> <p>org.globalbioticinteractions.taxon v0.3.20, 2020-01-17</p> <p>1. update taxonCache and taxonMap using automated scripts available at https://doi.org/10.5281/zenodo.3240558</p> <p>org.globalbioticinteractions.taxon v0.3.21, 2020-03-11</p> <p>1. update taxonCache and taxonMap using automated scripts available at https://doi.org/10.5281/zenodo.3240558</p> <p>org.globalbioticinteractions.taxon v0.3.22, 2020-04-14</p> <p>1. update taxonCache and taxonMap using automated scripts available at https://doi.org/10.5281/zenodo.3240558</p> <p>org.globalbioticinteractions.taxon v0.3.23, 2020-05-22</p> <p>1. update taxonCache and taxonMap using automated scripts available at https://doi.org/10.5281/zenodo.3240558</p> <p>org.globalbioticinteractions.taxon v0.3.24, 2020-06-23</p> <p>1. update taxonCache and taxonMap using automated scripts available at https://doi.org/10.5281/zenodo.3240558</p> <p>org.globalbioticinteractions.taxon v0.3.25, 2020-08-19</p> <p>1. update taxonCache and taxonMap using automated scripts available at https://doi.org/10.5281/zenodo.3240558</p> <p>org.globalbioticinteraction.taxon v0.3.26, 2020-10-01</p> <p>1. adding links to Plazi treatment via nomer append plazi (see https://github.com/globalbioticinteractions/nomer/issues/23)<br>by applying patches available via https://doi.org/10.5281/zenodo.4062711 .</p> <p>org.globalbioticinteraction.taxon v0.3.27, 2020-10-22</p> <p>1. update taxonCache and taxonMap using automated scripts available at https://doi.org/10.5281/zenodo.3240558</p> <p>org.globalbioticinteraction.taxon v0.3.28, 2021-01-19</p> <p>1. update taxonCache and taxonMap using patch 20210114-01 available via Poelen, Jorrit H. (2021). Global Biotic Interactions: Taxon Graph Patches (Version 0.6) [Data set]. Zenodo. http://doi.org/10.5281/zenodo.4451462 .</p> <p>org.globalbioticinteractions.taxon v0.3.29, 2021-01-26</p> <p>1. update taxonCache and taxonMap using automated scripts available at https://doi.org/10.5281/zenodo.3240558</p> <p>org.globalbioticinteractions.taxon v0.3.30, 2021-03-10</p> <p>1. update taxonCache and taxonMap using automated scripts available at https://doi.org/10.5281/zenodo.3240558</p> <p>org.globalbioticinteractions.taxon v0.3.31, 2021-03-31</p> <p>1. update taxonCache and taxonMap using patch 20210331-01 available via Poelen, Jorrit H. (2021). Global Biotic Interactions: Taxon Graph Patches (Version 0.7) [Data set]. Zenodo. http://doi.org/10.5281/zenodo.4655153 .</p> <p>org.globalbioticinteractions.taxon v0.3.32, 2021-05-12</p> <p>1. update taxonCache and taxonMap using automated scripts available at https://doi.org/10.5281/zenodo.3240558<br>2. remove suspicious mappings from Fungal to some virus name described in https://www.gbif.org/species/4904189 Fungal see https://github.com/globalbioticinteractions/mangal/issues/1#issuecomment-833956239 .</p> <p>org.globalbioticinteractions.taxon v0.3.33, 2021-06-23</p> <p>1. remove suspicious viral name mappings as reported in https://github.com/globalbioticinteractions/globalbioticinteractions/issues/672 by updating taxonMap.tsv.gz using patch 20210623-01 available via Poelen, Jorrit H. (2021). Global Biotic Interactions: Taxon Graph Patches (Version 0.8) [Data set]. Zenodo. http://doi.org/10.5281/zenodo.5021824 .</p> <p>org.globalbioticinteractions.taxon v0.3.34, 2021-09-24</p> <p>1. update taxonCache and taxonMap using automated scripts available at https://doi.org/10.5281/zenodo.3240558</p> <p>org.globalbioticinteractions.taxon v0.3.35, 2021-11-19</p> <p>1. update taxonCache and taxonMap using automated scripts available at https://doi.org/10.5281/zenodo.3240558</p> <p>org.globalbioticinteractions.taxon v0.3.36, 2022-03-29</p> <p>1. update taxonCache and taxonMap using automated scripts available at https://doi.org/10.5281/zenodo.6394931</p> <p>org.globalbioticinteractions.taxon v0.4.0, 2023-03-21</p> <p>1. update elton, nomer, and globi taxon graph versions<br>2. attempt to align all names, including those aligned previously. Replaced incremental name alignment. Incremental name alignment was a optimization needed because of web api performance. Now, no web apis are used, so the optimization is no longer needed.<br>take names from https://globalbioticinteractions.org/data verbatim-interactions.tsv.gz instead of parsing verbatim names from their sources</p> <p>org.globalbioticinteractions.taxon v0.4.1, 2023-03-23</p> <p>update taxon graph build script to fit into existing taxonMap/taxonCache schema<br>fix various bugs<br>remove internal validation until a more up-to-date validation method is available</p> <p>org.globalbioticinteractions.taxon v0.4.2, 2022-10-14</p> <p>update taxonCache and taxonMap using automated scripts available at globalbioticinteractions. (2023). globalbioticinteractions/taxon-graph-builder: 0.0.7 (0.0.7). Zenodo. https://doi.org/10.5281/zenodo.10037579</p> <p>org.globalbioticinteractions.taxon v0.4.3, 2022-10-26</p> <p>apply patch 20231026-01 to address https://github.com/globalbioticinteractions/globalwebdb/issues/1 and https://discuss.eol.org/t/questionable-link-in-trophic-web-for-white-tailed-jackrabbit/2296</p> <p>org.globalbioticinteractions.taxon v0.4.4, 2022-10-26</p> <p>apply patch 20231026-02 to continue to work towards addressing https://github.com/globalbioticinteractions/globalwebdb/issues/1 and https://discuss.eol.org/t/questionable-link-in-trophic-web-for-white-tailed-jackrabbit/2296</p> <p>org.globalbioticinteractions.taxon v0.4.5, 2022-10-26</p> <p>apply patch 20231026-03 to continue to work towards addressing https://github.com/globalbioticinteractions/globalwebdb/issues/1 and https://discuss.eol.org/t/questionable-link-in-trophic-web-for-white-tailed-jackrabbit/2296</p> <p>org.globalbioticinteractions.taxon v0.4.6, 2024-06-17</p> <p>apply patch 20240617 to work towards addressing suspicious Candidatus name mapping reported in https://github.com/globalbioticinteractions/globalbioticinteractions/issues/968</p> <p>org.globalbioticinteractions.taxon v0.5.0, 2024-07-05</p> <p>1. update taxonCache and taxonMap using automated scripts available via Taxon Graph Builder v0.1.0 https://github.com/globalbioticinteractions/taxon-graph-builder/releases/tag/0.1.0 and/or https://doi.org/10.5281/zenodo.1286023 .&nbsp;</p> <p>org.globalbioticinteractions.taxon v0.5.1, 2024-07-08</p> <p>1. update taxonCache and taxonMap using automated scripts available via Taxon Graph Builder v0.1.1 https://github.com/globalbioticinteractions/taxon-graph-builder/releases/tag/0.1.1 and/or https://doi.org/10.5281/zenodo.12687693 .&nbsp;</p> <p>org.globalbioticinteractions.taxon v0.5.2, 2024-07-11</p> <p>1. update taxonCache and taxonMap using automated scripts available via Taxon Graph Builder v0.1.2 https://github.com/globalbioticinteractions/taxon-graph-builder/releases/tag/0.1.2 and/or https://doi.org/10.5281/zenodo.12687693 .&nbsp;</p> <p>org.globalbioticinteractions.taxon v0.5.3, 2024-07-24</p> <p>1. update taxonCache and taxonMap using automated scripts available via Taxon Graph Builder v0.1.2 https://github.com/globalbioticinteractions/taxon-graph-builder/releases/tag/0.1.2 and/or https://doi.org/10.5281/zenodo.12687693 .&nbsp;</p> <p><br>org.globalbioticinteractions.taxon v0.5.4, 2025-02-12</p> <p>1. update taxonCache and taxonMap using automated scripts available via Taxon Graph Builder v0.1.2 https://github.com/globalbioticinteractions/taxon-graph-builder/releases/tag/0.1.2 and/or https://doi.org/10.5281/zenodo.12687693 .&nbsp;</p>

opencc-zeroJul 2024View details →
zenodo44/100

Deep learning models predicting gene functions and pathways using public DRKG knowledge graph and graph neural network

<p>The attached dataset contains pretrained link prediction models, as described in our paper 'Morphological Map of Under- and Over-Expression of Genes in Human Cells'.</p>

opencc-by-4.0Nov 2024View details →
zenodo44/100

Dataset variants used in "Task-Driven Knowledge Graph Filtering Improves Prioritizing Drugs for Repurposing"

<p>This file contains all datasets and variants thereof used in the linked paper. We do not take credit for constructing the datasets, which has been done by the respective original authors (<a href="https://github.com/hetio/hetionet">https://github.com/hetio/hetionet</a>,&nbsp;<a href="https://github.com/gnn4dr/DRKG">https://github.com/gnn4dr/DRKG</a>). For our work we produced modified versions (called &quot;subset&quot; in the file) by applying our metapath based filtering approach. For validation purposed we also constructed ablation versions where one specific type of entities is missing (i.e. &quot;nogene&quot;, &quot;noside&quot;, etc).</p>

opencc-by-4.0Nov 2021View details →
zenodo44/100

UC Santa Barbara Invertebrate Zoology Collection (UCSB-IZC) Data Archive and Biodiversity Dataset Graph hash://md5/10663911550bb52a0f5741993f82db9d hash://sha256/80c0f5fc598be1446d23c95141e87880c9e53773cb2e0b5b54cb57a8ea00b20c

<p>A biodiversity dataset graph: UCSB-IZC</p> <p>The intended use of this archive is to facilitate (meta-)analysis of the UC Santa Barbara Invertebrate Zoology Collection (UCSB-IZC). UCSB-IZC is a natural history collection of invertebrate zoology at Cheadle Center of Biodiversity and Ecological Restoration, University of California Santa Barbara.</p> <p>This dataset provides versioned snapshots of the UCSB-IZC network as tracked by Preston [2,3] between 2021-10-08 and 2021-11-04 using [preston track &quot;https://api.gbif.org/v1/occurrence/search/?datasetKey=d6097f75-f99e-4c2a-b8a5-b0fc213ecbd0&quot;].</p> <p>This archive contains 14349 images related to 32533 occurrence/specimen records. See included sample-image.jpg and their associated meta-data sample-image.json [4].</p> <p>The images were counted using:</p> <p>$ preston cat hash://sha256/80c0f5fc598be1446d23c95141e87880c9e53773cb2e0b5b54cb57a8ea00b20c\<br> &nbsp;| grep -o -P &quot;.*depict&quot;\<br> &nbsp;| sort\<br> &nbsp;| uniq\<br> &nbsp;| wc -l</p> <p>And the occurrences were counted using:</p> <p>$ preston cat hash://sha256/80c0f5fc598be1446d23c95141e87880c9e53773cb2e0b5b54cb57a8ea00b20c\<br> &nbsp;| grep -o -P &quot;occurrence/([0-9])+&quot;\<br> &nbsp;| sort\<br> &nbsp;| uniq\<br> &nbsp;| wc -l</p> <p>The archive consists of 256 individual parts (e.g., preston-00.tar.gz, preston-01.tar.gz, ...) to allow for parallel file downloads. The archive contains three types of files: index files, provenance files and data files. Only two index and provenance files are included and have been individually included in this dataset publication. Index files provide a way to links provenance files in time to establish a versioning mechanism.</p> <p>To retrieve and verify the downloaded UCSB-IZC biodiversity dataset graph, first download preston-*.tar.gz. Then, extract the archives into a &quot;data&quot; folder. Alternatively, you can use the Preston [2,3] command-line tool to &quot;clone&quot; this dataset using:</p> <p>$ java -jar preston.jar clone --remote https://archive.org/download/preston-ucsb-izc/data.zip/,https://zenodo.org/record/5557670/files,https://zenodo.org/record/5660088/files/</p> <p>After that, verify the index of the archive by reproducing the following provenance log history:</p> <p>$ java -jar preston.jar history<br> &lt;urn:uuid:0659a54f-b713-4f86-a917-5be166a14110&gt; &lt;http://purl.org/pav/hasVersion&gt; &lt;hash://sha256/d5eb492d3e0304afadcc85f968de1e23042479ad670a5819cee00f2c2c277f36&gt; .<br> &lt;hash://sha256/80c0f5fc598be1446d23c95141e87880c9e53773cb2e0b5b54cb57a8ea00b20c&gt; &lt;http://purl.org/pav/previousVersion&gt; &lt;hash://sha256/d5eb492d3e0304afadcc85f968de1e23042479ad670a5819cee00f2c2c277f36&gt; .</p> <p>To check the integrity of the extracted archive, confirm that each line produce by the command &quot;preston verify&quot; produces lines as shown below, with each line including &quot;CONTENT_PRESENT_VALID_HASH&quot;. Depending on hardware capacity, this may take a while.</p> <p>$ java -jar preston.jar verify<br> hash://sha256/ce1dc2468dfb1706a6f972f11b5489dc635bdcf9c9fd62a942af14898c488b2c&nbsp;&nbsp;&nbsp; file:/home/jhpoelen/ucsb-izc/data/ce/1d/ce1dc2468dfb1706a6f972f11b5489dc635bdcf9c9fd62a942af14898c488b2c&nbsp;&nbsp;&nbsp; OK&nbsp;&nbsp;&nbsp; CONTENT_PRESENT_VALID_HASH&nbsp;&nbsp;&nbsp; 66438&nbsp;&nbsp;&nbsp; hash://sha256/ce1dc2468dfb1706a6f972f11b5489dc635bdcf9c9fd62a942af14898c488b2c<br> hash://sha256/f68d489a9275cb9d1249767244b594c09ab23fd00b82374cb5877cabaa4d0844&nbsp;&nbsp;&nbsp; file:/home/jhpoelen/ucsb-izc/data/f6/8d/f68d489a9275cb9d1249767244b594c09ab23fd00b82374cb5877cabaa4d0844&nbsp;&nbsp;&nbsp; OK&nbsp;&nbsp;&nbsp; CONTENT_PRESENT_VALID_HASH&nbsp;&nbsp;&nbsp; 4093&nbsp;&nbsp;&nbsp; hash://sha256/f68d489a9275cb9d1249767244b594c09ab23fd00b82374cb5877cabaa4d0844<br> hash://sha256/3e70b7adc1a342e5551b598d732c20b96a0102bb1e7f42cfc2ae8a2c4227edef&nbsp;&nbsp;&nbsp; file:/home/jhpoelen/ucsb-izc/data/3e/70/3e70b7adc1a342e5551b598d732c20b96a0102bb1e7f42cfc2ae8a2c4227edef&nbsp;&nbsp;&nbsp; OK&nbsp;&nbsp;&nbsp; CONTENT_PRESENT_VALID_HASH&nbsp;&nbsp;&nbsp; 5746&nbsp;&nbsp;&nbsp; hash://sha256/3e70b7adc1a342e5551b598d732c20b96a0102bb1e7f42cfc2ae8a2c4227edef<br> hash://sha256/995806159ae2fdffdc35eef2a7eccf362cb663522c308aa6aa52e2faca8bb25b&nbsp;&nbsp;&nbsp; file:/home/jhpoelen/ucsb-izc/data/99/58/995806159ae2fdffdc35eef2a7eccf362cb663522c308aa6aa52e2faca8bb25b&nbsp;&nbsp;&nbsp; OK&nbsp;&nbsp;&nbsp; CONTENT_PRESENT_VALID_HASH&nbsp;&nbsp;&nbsp; 6147&nbsp;&nbsp;&nbsp; hash://sha256/995806159ae2fdffdc35eef2a7eccf362cb663522c308aa6aa52e2faca8bb25b</p> <p>Note that a copy of the java program &quot;preston&quot;, preston.jar, is included in this publication. The program runs on java 8+ virtual machine using &quot;java -jar preston.jar&quot;, or in short &quot;preston&quot;.</p> <p>Files in this data publication:</p> <p>--- start of file descriptions ---</p> <p>-- description of archive and its contents (this file) --<br> README</p> <p>-- executable java jar containing preston [2,3] v0.3.1. --<br> preston.jar</p> <p>-- preston archive containing UCSB-IZC (meta-)data/image files, associated provenance logs and a provenance index --<br> preston-[00-ff].tar.gz</p> <p>-- individual provenance index files --<br> 2a5de79372318317a382ea9a2cef069780b852b01210ef59e06b640a3539cb5a</p> <p>-- example image and meta-data --<br> sample-image.jpg (with hash://sha256/916ba5dc6ad37a3c16634e1a0e3d2a09969f2527bb207220e3dbdbcf4d6b810c)<br> sample-image.json (with hash://sha256/f68d489a9275cb9d1249767244b594c09ab23fd00b82374cb5877cabaa4d0844)</p> <p>--- end of file descriptions ---</p> <p><br> References</p> <p>[1] Cheadle Center for Biodiversity and Ecological Restoration (2021). University of California Santa Barbara Invertebrate Zoology Collection. Occurrence dataset https://doi.org/10.15468/w6hvhv accessed via GBIF.org on 2021-11-04 as indexed by the Global Biodiversity Informatics Facility (GBIF) with provenance hash://sha256/d5eb492d3e0304afadcc85f968de1e23042479ad670a5819cee00f2c2c277f36 hash://sha256/80c0f5fc598be1446d23c95141e87880c9e53773cb2e0b5b54cb57a8ea00b20c.<br> [2] https://preston.guoda.bio, https://doi.org/10.5281/zenodo.1410543 .<br> [3] MJ Elliott, JH Poelen, JAB Fortes (2020). Toward Reliable Biodiversity Dataset References. Ecological Informatics. https://doi.org/10.1016/j.ecoinf.2020.101132<br> [4] Cheadle Center for Biodiversity and Ecological Restoration (2021). University of California Santa Barbara Invertebrate Zoology Collection. Occurrence dataset https://doi.org/10.15468/w6hvhv accessed via GBIF.org on 2021-10-08. https://www.gbif.org/occurrence/3323647301 . hash://sha256/f68d489a9275cb9d1249767244b594c09ab23fd00b82374cb5877cabaa4d0844 hash://sha256/916ba5dc6ad37a3c16634e1a0e3d2a09969f2527bb207220e3dbdbcf4d6b810c</p>

opencc-zeroNov 2021View details →
zenodo44/100

Medieval manuscripts and their migrations: Using SPARQL to investigate the research potential of an aggregated Knowledge Graph

<p>This dataset contains the <strong>SPARQL queries</strong> presented and discussed in our article published in <em>Digital Medievalist</em> 2022 (as a PDF file), together with the <strong>results of those queries</strong> as CSV files. The query and step numbering follows that given in the article.</p> <p>The queries can be run against the SPARQL endpoint for the <strong>Mapping Manuscript Migrations</strong> project:&nbsp;<a href="https://ldf.fi/mmm/sparql">https://ldf.fi/mmm/sparql</a></p> <p>The full <strong>Mapping Manuscript Migrations dataset </strong>can also be downloaded from the Zenodo repository and installed in your own triple store: <a href="https://zenodo.org/record/4440464">https://zenodo.org/record/4440464</a></p> <p>When copying and pasting these SPARQL queries into a SPARQL client like <a href="https://yasgui.triply.cc/">YASGUI</a>, please check that the line numbering has been copied over correctly. Copying from a PDF file can sometimes break a single long line into multiple separate lines, which will cause a SPARQL validation error.</p> <p>The CSV files contain the results of the queries when run against the Mapping Manuscript Migrations SPARQL endpoint as of 17 December 2021. Please note that Query 2, Step 2, produces no results, so a CSV file has not been provided.</p> <p>The<strong> Mapping Manuscript Migrations portal </strong>can be found at&nbsp;<a href="https://mappingmanuscriptmigrations.org/en/">https://mappingmanuscriptmigrations.org/en/&nbsp;</a></p> <p>SPARQL tutorials are included in the project&#39;s <strong>GitHub documentation</strong>:&nbsp;<a href="https://mapping-manuscript-migrations.github.io/">https://mapping-manuscript-migrations.github.io/</a></p>

opencc-by-4.0Dec 2021View details →
zenodo44/100

Anonymized Graph Data with Friend Connections of 189505 VKontakte Users

<p>The dataset contains anonymized graph data with friend connections of 189505 VKontakte users. The dataset was used in <a href="http://github.com/filipp134/vk_bot_detection">this</a> Github project on the detection of social bots on VKontakte. The script which was used for the collection of the dataset is <a href="https://github.com/filipp134/vk_bot_detection/blob/main/Collecting%20datasets%20and%20merging%20them%20into%20one/collect_graph_data.py">here</a>.</p> <p>The dataset was collected in the following 2 steps by using the official <a href="https://dev.vk.com/api/getting-started">VKontakte API</a>:</p> <p>1. Friend connections&nbsp;of 11766 VKontakte users, who had 177739 unique friends, were collected.</p> <p>2. Friend connections of these 177739 users were collected.&nbsp;</p> <p>The dataset is in JSON format and is quite heavy: 424.5 MB.</p> <p>&nbsp;</p>

opencc-by-4.0Feb 2022View details →
zenodo44/100

FabWave Product Design Knowlsdge Graph (FPD-KG)

<p>FabWave Product Design Knowledge Graph: A Knowledge Graph for product design &amp; manufacturing, constructed using the openly available and academia-sourced&nbsp;3D CAD data. (Starly, Binil; Bharadwaj, Akshay; Angrish, Atin. (2019). FabWave CAD Repository Categorized Part Classes. 10.13140/RG.2.2.31167.87201.)</p>

opencc-by-4.0Feb 2022View details →
zenodo44/100

Bread wheat genomes graph pangenome

<p>A Giraffe and a GFA-formatted minigraph assembly of sixteen bread wheat cultivar genome assemblies.</p> <p>15-wheat10+.bed.gz is the relinearised graph from gfatools gfa2bed</p> <p>15-wheat10+.gfa.gz is the graph in GFA format as built by minigraph</p> <p>index.min, index.dist, index.giraffe.gbz are the same graph formatted for Giraffe alignments with vg giraffe v1.34.0 or later.</p> <pre><code class="language-bash">vg autoindex -w giraffe -g 15-wheat10+.gfa -t 16 -T ./ -V 1 --target-mem 850G</code></pre> <p>It is possible to convert the gfa graph to vg format:</p> <pre><code class="language-bash">vg convert -v -g 15-wheat10+.gfa &gt; 15-wheat10+.vg</code></pre> <p>To use the index in alignments using vg v1.34.0 or later:</p> <pre><code class="language-bash">vg giraffe -Z index.giraffe.gbz -m index.min -d index.dist -f your_reads.fq &gt; mapped.gam</code></pre> <p>&nbsp;</p>

opencc-by-4.0Feb 2022View details →
zenodo44/100

Supporting Online Toxicity Detection with Knowledge Graphs: Data

<p>This data repository contains the output files from the analysis of the paper &quot;Supporting Online Toxicity Detection with Knowledge Graphs&quot; presented at the International Conference on Web and Social Media 2022 (ICWSM-2022).</p> <p>&nbsp;</p> <p>The data contains annotations of gender and sexual orientation entities provided by the Gender and Sexual Orientation Ontology (https://bioportal.bioontology.org/ontologies/GSSO).</p> <p>We analyse demographic group samples from the Civil Comments Identities dataset (https://www.tensorflow.org/datasets/catalog/civil_comments).</p>

opencc-by-4.0Mar 2022View details →
zenodo44/100

DBpedia RDF2Vec Graph Embeddings

<p>DBpedia graph embeddings using RDF2Vec.&nbsp;RDF2Vec embedding generation code can be found&nbsp;<a href="https://github.com/dwslab/jRDF2Vec">here</a>&nbsp;and is based on a publication by Portisch et al.&nbsp;[1].</p> <p>The embeddings dataset consists of 200-dimensional vectors of DBpedia entities (from 1/9/2021).</p> <p>Figure of cosine similarities between a selected set of DBpedia entities are provided in the dataset <a href="https://zenodo.org/record/6384728/files/heatmap.pdf?download=1">here</a>.</p> <p>&nbsp;</p> <p><strong>Generating Embeddings</strong></p> <p>The code for generating these embeddings can be found&nbsp;<a href="https://github.com/EDAO-Project/DBpediaEmbedding">here</a>.</p> <p>Run the run.sh script&nbsp;that wraps all the necessary commmands to generate embeddings</p> <pre><code class="language-bash">bash run.sh</code></pre> <p>The script downloads a set of DBpedia files, which are listed in&nbsp;<code>dbpedia_files.txt</code>. It then builds a Docker image and runs a container of that image that generates the embeddings for the DBpedia graph defined by the DBpedia files.</p> <p>A folder&nbsp;<code>files</code>&nbsp;is created containing all the downloaded DBpedia files, and a folder&nbsp;<code>embeddings/dbpedia</code>&nbsp;is created containing the embeddings in&nbsp;<code>vectors.txt</code>&nbsp;along a set of random walk files.</p> <p>&nbsp;</p> <p><strong>Run Time of Embeddings Generation</strong></p> <p>Generating embeddings can take more than a day, but it depends on the number of DBpedia files chosen to be downloaded. Following are some basic run time statistics when embeddings are generated on a 64 GB RAM, 8 cores (AMD EPYC), 1 TB SSD, 1996.221 MHz machine.</p> <ul> <li><strong>Total</strong>: 1 day, 8 hours, 52 minutes, 41 seconds</li> <li><strong>Walk generation</strong>: 0 days, 7 minutes, 24 minutes, 36 seconds</li> <li><strong>Training</strong>: 1 day, 1 hour, 28 minutes, 5 seconds</li> </ul> <p>&nbsp;</p> <p><strong>Parameters Used</strong></p> <p>Here is listed the parameters used to generate the embeddings provided here:</p> <ul> <li><strong>Number of walks per entity</strong>: 100</li> <li><strong>Depth (hops) per walk</strong>: 4</li> <li><strong>Walk generation mode</strong>: RANDOM_WALKS_DUPLICATE_FREE</li> <li><strong>Threads</strong>: # of processors / 2</li> <li><strong>Training mode</strong>: sg</li> <li><strong>Embeddings vector dimension</strong>: 200</li> <li><strong>Minimum word2vec word count</strong>: 1</li> <li><strong>Sample rate</strong>: 0.0</li> <li><strong>Training window size</strong>: 5</li> <li><strong>Training epochs</strong>: 5</li> </ul>

opencc-by-4.0Dec 2021View details →
zenodo44/100

Graph Theoretical Measures of Fast Ripple Networks Support the Epileptic Network Hypothesis

<p>MongoDB JSON files of the (high-frequency oscillation) HFO and electrode databases used for this study and others.</p>

opencc-by-4.0Apr 2022View details →
zenodo44/100

SeaLiT Knowledge Graphs - Maritime History Data in RDF using a CIDOC-CRM extension (SeaLiT Ontology)

<p><strong>SeaLiT Knowledge Graphs</strong> is an RDF dataset of maritime history data that has been transcribed (and then transformed) from original archival sources&nbsp;in the context of the <a href="http://www.sealitproject.eu/">SeaLiT Project</a>&nbsp;(Seafaring Lives in Transition, Mediterranean Maritime Labour and Shipping, 1850s-1920s).&nbsp;The underlying data model is the <a href="https://zenodo.org/record/5964240">SeaLiT Ontology</a>, an extension of the ISO standard&nbsp;<strong>CIDOC-CRM</strong>&nbsp;(ISO 21127:2014) for the modelling and integration of maritime history information.&nbsp;</p> <p>The knowledge graphs integrate data of totally 16 different types of archival sources:</p> <ul> <li>Crew Lists <ul> <li>Crew and displacement list (Roll)</li> <li>Crew List (Ruoli di Equipaggio)</li> <li>General Spanish Crew List</li> </ul> </li> <li>Registers / Lists <ul> <li>Students Register</li> <li>Civil Register</li> <li>Register of Maritime Personnel</li> <li>Register of Maritime Workers (Matricole della gente di mare)</li> <li>Sailors Register (Libro de registro de marineros)</li> <li>Naval Ship Register List</li> <li>Seagoing Personnel</li> <li>Lists of ships</li> </ul> </li> <li>Censuses <ul> <li>Census La Ciotat</li> <li>First National all-Russian Census of the Russian Empire</li> </ul> </li> <li>Payrolls <ul> <li>Payrolls&nbsp;of private archives and libraries in Greece</li> <li>Payrolls of Russian Steam Navigation and Trading Company</li> </ul> </li> <li>Employment records <ul> <li>Shipyards of Messageries Maritimes, La Ciotat</li> </ul> </li> </ul> <p>More information about the archival sources are available through the <a href="https://sealitproject.eu/dictionary-of-source-types-list">SeaLiT website</a>. Data exploration applications over these sources are also publicly available (<a href="https://catalogues.sealitproject.eu/">SeaLiT Catalogues</a>,&nbsp;<a href="http://rs.sealitproject.eu/">SeaLiT ResearchSpace</a>).&nbsp;</p> <p>Data from these archival sources has been transcribed in tabular form&nbsp;and then curated&nbsp;by historians of SeaLiT using the <a href="https://www.ics.forth.gr/isl/fast-cat">FAST CAT</a> system. The transcripts (records), together with the curated vocabulary terms and entity instances (ships, persons, locations, organizations), are then transformed to RDF using the SeaLiT Ontology as the target (domain) model.&nbsp;To this end, the corresponding schema mappings between the original schemata and the&nbsp;ontology were defined using the <a href="https://github.com/isl/x3ml">X3ML</a> mapping definition language, that were subsequently used for delivering the RDF datasets.&nbsp;</p> <p>More information about the FAST CAT system and the data transcription, curation and&nbsp;transformation processes can be found in the following paper:</p> <blockquote> <p>P. Fafalios, K. Petrakis, G. Samaritakis, K. Doerr, A. Kritsotaki, Y. Tzitzikas, M. Doerr, &quot;FAST CAT: Collaborative Data Entry and Curation for Semantic Interoperability in Digital Humanities&quot;, ACM Journal on Computing and Cultural Heritage, 2021. <a href="https://doi.org/10.1145/3461460">https://doi.org/10.1145/3461460</a>&nbsp;[<a href="http://users.ics.forth.gr/~fafalios/files/pubs/fafaliosJOCCH2021.pdf">pdf</a>, <a href="http://users.ics.forth.gr/~fafalios/files/bibs/fafaliosJOCCH2021.bib">bib</a>]</p> </blockquote> <p>The RDF dataset is provided as a set of TriG files per record per archival source. For each record, the dataset provides: i) one trig file for the record&#39;s data (<em>records.trig</em>), ii) one trig file for the record&#39;s (curated) vocabulary terms (<em>vocabularies.trig</em>), and iii) four trig files for the record&#39;s (curated) entity instances (<em>ships.trig, persons.trig, persons.trig, organizations.trig</em>).</p> <p>We also provide the RDFS files of the used ontologies&nbsp;(SeaLiT Ontology verson 1.0, CIDOC-CRM version 7.1.1).&nbsp;</p>

opencc-by-4.0Apr 2022View details →
zenodo44/100

Publication and Maintenance of Relational Data in Enterprise Knowledge Graphs Created (Files used in the experiments)

<p>This dataset contains two files created for the experiments presented in the article: Publication and Maintenance of RDB2RDF Views Externally Materialized in Enterprise Knowledge Graphs.</p> <p><strong>mapR2RML_MusicBrainz_completo.txt</strong><strong>:</strong>&nbsp;We created the&nbsp;R2RML mapping&nbsp;for translating MBD data into the&nbsp;Music Ontology vocabulary, which is used for publishing the LMB view. The LMB view was materialized using the&nbsp;D2RQ tool. It took 67 minutes to materialize the view with approximately 41.1 GB of NTriples. We also provided&nbsp;SPARQL endpoint&nbsp;for querying LMB View.</p> <p><strong>TriggersAndProcedures.txt</strong>: We created the&nbsp;triggers, procedures, and&nbsp;class in java&nbsp;to implement the rules required to compute and publish the changesets.</p> <p><strong>relationalViewDefinition.pdf</strong>: This document&nbsp;gives details about the process of creating the relational views used in the experiments.</p>

opencc-by-4.0Feb 2022View details →
zenodo44/100

Cubic vertex-transitive graphs on up to 1280 vertices.

<p>The <strong>Census of cubic vertex-transitive graphs</strong>&nbsp;contains the list of all cubic vertex-transitive graphs on at most 1280 vertices, together with a number of graph-theoretic properties (listed below).&nbsp;</p> <p>The&nbsp;list of graphs was&nbsp;originally compiled by Pablo Spiga, Gabriel Verret, and Primož Potočnik. The authors described&nbsp;the theoretical results and computations that were needed to compile the list in the paper&nbsp;<a href="https://doi.org/10.1016/j.jsc.2012.09.002">Cubic vertex-transitive graphs on up to 1280 vertices</a>.&nbsp;The original dataset files are available on <a href="https://www.fmf.uni-lj.si/~potocnik/work.htm">Potočnik&#39;s website</a>.</p> <p>The graph-theoretic properties were either computed or verified with the SageMath system, with a few exceptions. Some of the properties were not supported by SageMath at the time of computation, and were contributed: is_cayley,&nbsp;odd_girth,&nbsp;is_partial_cube (based on some previous code). The methods&nbsp;<a href="https://github.com/DiscreteZOO/DiscreteZOO-sage/blob/df8c8368a4912bd5396001139a13a9f3b9b2863f/discretezoo/entities/cvt/cvtgraph.py#L177-L249">is_moebius_ladder, is_prism, and is_spx</a>&nbsp;(together with <a href="https://github.com/DiscreteZOO/DiscreteZOO-sage/blob/df8c8368a4912bd5396001139a13a9f3b9b2863f/discretezoo/entities/spx/spxgraph.py#L182-L300">SPX constructions</a>) were computed in DiscreteZOO.</p> <p>A searchable version of this dataset is available on an instance of the <a href="https://data.mathhub.info/">MathDataHub</a> platform hosted at <a href="http://mdh.graphsym.net/collection/CVT">mdh.graphsym.net/collection/CVT</a>.</p> <p>You can download this collection either as an&nbsp;SQLite database or as CSV files. Contents:</p> <ul> <li>master_db_zenodo.sqlite3: the&nbsp;SQLite database,</li> <li>create_database.sql: the&nbsp;SQL script that creates empty database schema used in the SQLite database (not necessary for opening the database),</li> <li>main_Graph.txt: the main table containing graphs as a CSV file,</li> <li>main_Graph.sample.txt: a sample of the&nbsp;main table,</li> <li>main_CLTime.txt: the table containing canonical labelling computation times as a CSV file.</li> <li>main_CLTime.sample.txt:&nbsp;a sample of the canonical labeling&nbsp;table</li> </ul> <p><strong>Structure and contents</strong></p> <p>Each graph is given a census-specific identifier&nbsp;<em>CVT[n,i]</em>&nbsp;(the <em>i</em>-th graph of order <em>n</em> in the census).&nbsp;Some graphs have a list of more easily recognized names&nbsp;(such as the Petersen graph).&nbsp;The dataset contains graphs in two formats: the&nbsp;<a href="https://users.cecs.anu.edu.au/~bdm/data/formats.html">sparse6 format</a>&nbsp;and&nbsp;a format readable by the&nbsp;computer algebra system Magma.&nbsp;For compatibility with SQLite, boolean values are represented with ones and zeroes&nbsp;(true and false, respectively).</p> <p>Data columns with types:</p> <ul> <li>canonical_label (string): the graph in sparse6 format, canonically labelled with nauty (call with no additional arguments), version: nauty-27r3,</li> <li>cvt_index (string): the census-specific identifier&nbsp;<em>CVT[n,i]</em>&nbsp;(the <em>i</em>-th graph of order <em>n</em> in the census),</li> <li>data (string): the graph in sparse6 format,</li> <li>raw_magma_code (string): the list of vertex neighbourhoods, readable by the computer algebra system&nbsp;Magma,</li> <li>name (string): a&nbsp;list of names of the graph,</li> <li>number_of_vertices (integer): number of vertices&nbsp;in the graph,</li> <li>clique_number (integer): the number of vertices in the largest clique subgraph,</li> <li>diameter (integer): the greatest distance between any pair of points,</li> <li>girth (integer): the length of the shortest cycle in the graph,</li> <li>is_arc_transitive (boolean): for every&nbsp;two ordered pairs of adjacent vertices, does there exist an automorphism, mapping one to the other,</li> <li>is_bipartite (boolean): can the vertices of the graph be partitioned into two sets, such that every edge connects a vertex in one set to a vertex in the other set,</li> <li>is_cayley (boolean): can the graph be constructed as a Cayley graph of some group for some generating set,</li> <li>is_distance_regular (boolean): for any two vertices <em>v</em>&nbsp;and <em>w</em>, does the number of vertices at distance <em>j</em>&nbsp;from <em>v</em>&nbsp;and at distance <em>k</em>&nbsp;from <em>w</em>&nbsp;depend&nbsp;only upon <em>j</em>, <em>k</em>, and <em>i = d(v, w)</em>,</li> <li>is_distance_transitive (boolean): for any two vertices <em>v</em>&nbsp;and <em>w</em> at any distance <em>i</em>, and any other two vertices <em>x</em>&nbsp;and <em>y</em>&nbsp;at the same distance, is there&nbsp;an automorphism of the graph that carries <em>v</em>&nbsp;to <em>x</em>&nbsp;and <em>w</em>&nbsp;to <em>y</em>,</li> <li>is_edge_transitive (boolean):&nbsp;for every two edges, does there exist an automorphism, mapping one to the other,</li> <li>is_hamiltonian (boolean): does the graph have a cycle that visits each vertex exactly once,</li> <li>is_partial_cube (boolean): is&nbsp;the graph isometric to a subgraph of a hypercube,</li> <li>is_split (boolean): can the vertices of the graph be partitioned into a clique and an independent set,</li> <li>is_strongly_regular (boolean): do there exists <span class="math-tex">\(\lambda\)</span>&nbsp;and <span class="math-tex">\(\mu\)</span>&nbsp;such that every two adjacent vertices have <span class="math-tex">\(\lambda\)</span> common neighbours and every&nbsp;two non-adjacent vertices have <span class="math-tex">\(\mu\)</span> common neighbours,</li> <li>odd_girth (integer): length of the shortest odd cycle,</li> <li>triangles_count (integer): the number of cycles of length <em>3</em>&nbsp;in the graph,</li> <li>is_moebius_ladder (boolean): is&nbsp;the graph a M&ouml;bius ladder,</li> <li>is_prism (boolean): is the graph a skeleton of a prism,</li> <li>is_spx (boolean): does the graph belong to the Split Praeger-Xu family of graphs,</li> <li>vertex_stabilizer (list of integer pairs): the order of the vertex stabilizer of the graph&#39;s automorphism group, given as a prime factorization; a list of pairs <span class="math-tex">\((p_i, e_i)\)</span>&nbsp;of primes and exponents such that <span class="math-tex">\(\prod_i p_i^{e_i}\)</span>&nbsp;is the order of the vertex stabilizer.</li> </ul> <p>This work was partially supported by ARRS research project no. J1-1691.</p>

opencc-by-4.0May 2022View details →
zenodo44/100

Evaluation Set - Contributions Similarity in the Open Research Knowledge Graph

<p>This evaluation set has been created for evaluating a content-based recommender system in the context of the Open Research Knowledge Graph (ORKG). The recommender system accepts structured ORKG contribution as input and recommends existing contributions in the ORKG semantically relevant to the given one.</p> <p>&nbsp;</p> <p>The evaluation set is manually annotated based on the <a href="https://www.orkg.org/orkg/featured-comparisons">featured comparisons</a> in the ORKG. In the course of this, it has been distinguished between homogeneous (those who are dissimilar in 2-3 properties) and heterogeneous (otherwise) instances. Multiple annotations have been obtained for the former and exactly one for the latter.</p> <p>&nbsp;</p> <p>It has been also distinguished between &quot;with_response&quot; and &quot;without_response&quot; instances (50 instances for each). The former are those contributions for them the initial version of the contributions similarity service has found similarities and the latter are the opposite case.</p> <p>&nbsp;</p> <p>This evaluation set has been created and applied on a modified version of the contributions similarity service in the context of <a href="https://doi.org/10.15488/11834">this master&#39;s thesis</a>. The modified version of the service has simplified the document representation of contributions that are stored in an ElasticSearch index by omitting redundant terms.</p> <p>The evaluation set has the following schema:</p> <pre><code class="language-json">{ "with_response": [ { "contribution_id": "some_id", "comparison_id": "some_id", "comparison_label": "some_label", "contribution_label": "some_label", "paper": "some_id", "research_field": "some_id", "research_problems": [ "some_id" ], "annotations": [ "some_id of a similar contribution", ... ] }, ... ], "without_response": [ ... ] }</code></pre> <p>&nbsp;</p>

opencc-by-4.0Jan 2022View details →
zenodo44/100

TecKnoGraph: Knowledge Graph from patents in C4ISTAR

<p>&lt;img src=&quot;https://github.com/nicolamelluso/TecKnoGraph-demo/blob/main/TecKnoGraph-Example%20Graph.png&quot; alt=&quot;TecKnoGraph&quot;&gt;</p> <p>This dataset contains a sample of Knowledge Graph (KG) created with TecKnoGraph.</p> <p>There are two files:</p> <p><strong>- TecKnoGraph-C4ISTAR-sample.csv</strong>: this file&nbsp;contains the KG in the form of triples where each element of the triple (source, relation, target) is tagged with categories.</p> <p><strong>- patents.zip:</strong>&nbsp;this file contains data about 10,000 patents; each patent corresponds to a txt file.</p> <p>There is available a demo for using TecKnoGraph from examples of input text:<br> https://nicolamelluso-tecknograph-demo-tecknograph-streamlit-cnq4sm.streamlitapp.com/</p> <p>In this repository it is possible to find also the appendix of the corresponding paper.</p>

opencc-by-4.0Jul 2022View details →
zenodo44/100

PheKnowLator Human Disease Knowledge Graphs - Build Data (Processed)

<p><strong>RELEASE V2.1.0 KNOWLEDGE GRAPH: PROCESSED DATA SOURCES&nbsp;</strong></p> <p><strong>Release:</strong>&nbsp;<a href="https://github.com/callahantiff/PheKnowLator/wiki/v2.0.0">v2.1.0&nbsp;</a></p> <p>The goal of this build was to create a knowledge graph that represented human disease mechanisms and included the central dogma. The data sources utilized in this release include many of the sources used in the initial release, as well as some new data made available by the&nbsp;<a href="https://ctdbase.org/">Comparative Toxicogenomics Database</a>&nbsp;and experimental data from the&nbsp;<a href="https://www.proteinatlas.org/">Human Protein Atlas</a>.</p> <p>Data sources are listed by type (Ontology and Data not represented in an ontology [Database Sources]). Additional details are provided for each data source below.&nbsp;Please see documentation on the primary release&nbsp;(<a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources">https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources</a>) for additional details on each data source as well as citation information.</p> <p><strong>Data Access:</strong></p> <ul> <li><a href="https://console.cloud.google.com/storage/browser/pheknowlator/archived_builds/release_v2.1.0/build_01MAY2021?project=pheknowlator">https://console.cloud.google.com/storage/browser/pheknowlator/archived_builds/release_v2.1.0/build_01MAY2021</a></li> </ul> <p>&nbsp;</p> <p><strong>ONTOLOGIES</strong></p> <ul> <li>Cell Ontology</li> <li>Cell Line Ontology</li> <li>Chemical Entities of Biological Interest (ChEBI) Ontology</li> <li>Gene Ontology</li> <li>Human Phenotype Ontology</li> <li>Mondo Disease Ontology</li> <li>Pathway Ontology</li> <li>Protein Ontology</li> <li>Relations Ontology</li> <li>Sequence Ontology</li> <li>Uber-Anatomy Ontology</li> <li>Vaccine Ontology</li> </ul> <p>&nbsp;</p> <p><strong>Cell Ontology (CL)</strong></p> <p><strong>Homepage:</strong>&nbsp;<strong><a href="https://github.com/obophenotype/cell-ontology"><code>GitHub</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Bard J, Rhee SY, Ashburner M.&nbsp;<a href="https://genomebiology.biomedcentral.com/articles/10.1186/gb-2005-6-2-r21">An ontology for cell types</a>. Genome Biology. 2005;6(2):R21</p> </blockquote> <p><strong>Usage:</strong>&nbsp;Utilized to connect&nbsp;<code>transcripts</code>&nbsp;and&nbsp;<code>proteins</code>&nbsp;to&nbsp;<code>cells</code>. Additionally, the edges between this ontology and its dependencies are utilized:</p> <ul> <li><strong><a href="https://www.ebi.ac.uk/chebi/"><code>ChEBI</code></a></strong></li> <li><strong><a href="http://geneontology.org/"><code>GO</code></a></strong></li> <li><strong><a href="https://github.com/pato-ontology/pato/"><code>PATO</code></a></strong></li> <li><strong><a href="https://proconsortium.org/"><code>PRO</code></a></strong></li> <li><strong><a href="https://github.com/oborel/obo-relations/"><code>RO</code></a></strong></li> <li><strong><a href="https://uberon.github.io/"><code>UBERON</code></a></strong></li> </ul> <p>&nbsp;</p> <p><strong>Cell Line Ontology (CLO)</strong></p> <p><strong>Homepage:</strong>&nbsp;<strong><a href="http://www.clo-ontology.org/"><code>http://www.clo-ontology.org/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Sarntivijai S, Lin Y, Xiang Z, Meehan TF, Diehl AD, Vempati UD, Sch&uuml;rer SC, Pang C, Malone J, Parkinson H, Liu Y.&nbsp;<a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4387853/">CLO: the cell line ontology</a>. Journal of Biomedical Semantics. 2014;5(1):37</p> </blockquote> <p><strong>Usage:</strong>&nbsp;Utilized this ontology to map&nbsp;<code>cell lines</code>&nbsp;to&nbsp;<code>transcripts</code>&nbsp;and&nbsp;<code>proteins</code>. Additionally, the edges between this ontology and its dependencies are utilized:</p> <ul> <li><strong><a href="http://www.obofoundry.org/ontology/cl.html"><code>CL</code></a></strong></li> <li><strong><a href="http://disease-ontology.org/"><code>DOID</code></a></strong></li> <li><strong><a href="https://github.com/obophenotype/ncbitaxon"><code>NCBITaxon</code></a></strong></li> <li><strong><a href="https://uberon.github.io/"><code>UBERON</code></a></strong></li> </ul> <p>&nbsp;</p> <p><strong>Chemical Entities of Biological Interest (ChEBI)</strong></p> <p><strong>Homepage:</strong>&nbsp;<strong><a href="https://www.ebi.ac.uk/chebi/"><code>https://www.ebi.ac.uk/chebi/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Hastings J, Owen G, Dekker A, Ennis M, Kale N, Muthukrishnan V, Turner S, Swainston N, Mendes P, Steinbeck C.&nbsp;<a href="https://academic.oup.com/nar/article-abstract/44/D1/D1214/2502583">ChEBI in 2016: Improved services and an expanding collection of metabolites</a>. Nucleic Acids Research. 2015;44(D1):D1214-9</p> </blockquote> <p><strong>Usage:</strong>&nbsp;Utilized to connect&nbsp;<code>chemicals</code>&nbsp;to&nbsp;<code>complexes</code>,&nbsp;<code>diseases</code>,&nbsp;<code>genes</code>,&nbsp;<code>GO biological processes</code>,&nbsp;<code>GO cellular components</code>,&nbsp;<code>GO molecular functions</code>,&nbsp;<code>pathways</code>,&nbsp;<code>phenotypes</code>,&nbsp;<code>reactions</code>, and&nbsp;<code>transcripts</code>.</p> <p>&nbsp;</p> <p><strong>Gene Ontology (GO)</strong></p> <p><strong>Homepage:</strong>&nbsp;<strong><a href="http://geneontology.org/"><code>http://geneontology.org/</code></a></strong><br> <strong>Citations:</strong></p> <blockquote> <p>Ashburner M, Ball CA, Blake JA, Botstein D, Butler H, Cherry JM, Davis AP, Dolinski K, Dwight SS, Eppig JT, Harris MA.&nbsp;<a href="https://www.nature.com/articles/ng0500_25">Gene ontology: tool for the unification of biology</a>. Nature Genetics. 2000;25(1):25</p> <p>The Gene Ontology Consortium.&nbsp;<a href="https://academic.oup.com/nar/article/47/D1/D330/5160994">The Gene Ontology Resource: 20 years and still GOing strong</a>. Nucleic Acids Research. 2018;47(D1):D330-8</p> </blockquote> <p><strong>Usage:</strong>&nbsp;Utilized to connect&nbsp;<code>biological processes</code>,&nbsp;<code>cellular components</code>, and&nbsp;<code>molecular functions</code>&nbsp;to&nbsp;<code>chemicals</code>,&nbsp;<code>pathways</code>, and&nbsp;<code>proteins</code>. Additionally, the edges between this ontology and its dependencies are utilized:</p> <ul> <li><strong><a href="http://www.obofoundry.org/ontology/cl.html"><code>CL</code></a></strong></li> <li><strong><a href="https://github.com/obophenotype/ncbitaxon"><code>NCBITaxon</code></a></strong></li> <li><strong><a href="https://github.com/oborel/obo-relations/"><code>RO</code></a></strong></li> <li><strong><a href="https://uberon.github.io/"><code>UBERON</code></a></strong></li> </ul> <p><strong>Other Gene Ontology Data Used:</strong>&nbsp;<a href="http://geneontology.org/gene-associations/goa_human.gaf.gz"><code>goa_human.gaf.gz</code></a></p> <p>&nbsp;</p> <p><strong>Human Phenotype Ontology (HPO)</strong></p> <p><strong>Homepage:</strong>&nbsp;<strong><a href="https://hpo.jax.org/"><code>https://hpo.jax.org/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>K&ouml;hler S, Carmody L, Vasilevsky N, Jacobsen JO, Danis D, Gourdine JP, Gargano M, Harris NL, Matentzoglu N, McMurry JA, Osumi-Sutherland D.&nbsp;<a href="https://academic.oup.com/nar/article-abstract/47/D1/D1018/5198478">Expansion of the Human Phenotype Ontology (HPO) knowledge base and resources</a>. Nucleic Acids Research. 2018;47(D1):D1018-27</p> </blockquote> <p><strong>Usage:</strong>&nbsp;Utilized to connect&nbsp;<code>phenotypes</code>&nbsp;to&nbsp;<code>chemicals</code>,&nbsp;<code>diseases</code>,&nbsp;<code>genes</code>, and&nbsp;<code>variants</code>. Additionally, the edges between this ontology and its dependencies are utilized:</p> <ul> <li><strong><a href="http://www.obofoundry.org/ontology/cl.html"><code>CL</code></a></strong></li> <li><strong><a href="https://www.ebi.ac.uk/chebi/"><code>ChEBI</code></a></strong></li> <li><strong><a href="http://geneontology.org/"><code>GO</code></a></strong></li> <li><strong><a href="https://uberon.github.io/"><code>UBERON</code></a></strong></li> </ul> <p><strong>Files</strong></p> <ul> <li>Other Human Phenotype Ontology Data Used:&nbsp;<a href="http://purl.obolibrary.org/obo/hp/hpoa/phenotype.hpoa"><code>phenotype.hpoa</code></a></li> </ul> <p>&nbsp;</p> <p><strong>Mondo Disease Ontology (Mondo)</strong></p> <p><strong>Homepage:</strong>&nbsp;<strong><a href="https://mondo.monarchinitiative.org/"><code>https://mondo.monarchinitiative.org/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Mungall CJ, McMurry JA, K&ouml;hler S, Balhoff JP, Borromeo C, Brush M, Carbon S, Conlin T, Dunn N, Engelstad M, Foster E.&nbsp;<a href="https://academic.oup.com/nar/article-abstract/45/D1/D712/2605791">The Monarch Initiative: an integrative data and analytic platform connecting phenotypes to genotypes across species</a>. Nucleic Acids Research. 2017;45(D1):D712-22</p> </blockquote> <p><strong>Usage:</strong>&nbsp;Utilized to connect&nbsp;<code>diseases</code>&nbsp;to&nbsp;<code>chemicals</code>,&nbsp;<code>phenotypes</code>,&nbsp;<code>genes</code>, and&nbsp;<code>variants</code>. Additionally, the edges between this ontology and its dependencies are utilized:</p> <ul> <li><strong><a href="http://www.obofoundry.org/ontology/cl.html"><code>CL</code></a></strong></li> <li><strong><a href="https://www.ncbi.nlm.nih.gov/taxonomy"><code>NCBITaxon</code></a></strong></li> <li><strong><a href="http://geneontology.org/"><code>GO</code></a></strong></li> <li><strong><a href="https://hpo.jax.org/"><code>HPO</code></a></strong></li> <li><strong><a href="https://uberon.github.io/"><code>UBERON</code></a></strong></li> </ul> <p>&nbsp;</p> <p><strong>Pathway Ontology (PW)</strong></p> <p><strong>Homepage:</strong>&nbsp;<strong><a href="https://rgd.mcw.edu/wg/home/pathway2/"><code>rgd.mcw.edu</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Petri V, Jayaraman P, Tutaj M, Hayman GT, Smith JR, De Pons J, Laulederkind SJ, Lowry TF, Nigam R, Wang SJ, Shimoyama M.&nbsp;<a href="https://www.ncbi.nlm.nih.gov/pubmed/24499703">The pathway ontology&ndash;updates and applications</a>. Journal of Biomedical Semantics. 2014;5(1):7.</p> </blockquote> <p><strong>Usage:</strong>&nbsp;Utilized to connect&nbsp;<code>pathways</code>&nbsp;to&nbsp;<code>GO biological processes</code>,&nbsp;<code>GO cellular components</code>,&nbsp;<code>GO molecular functions</code>,&nbsp;<code>Reactome pathways</code>. Several steps are taken in order to connect&nbsp;<code>Pathway Ontology</code>&nbsp;identifiers to&nbsp;<code>Reactome</code>&nbsp;pathways and&nbsp;<code>GO biological processes</code>. To connect&nbsp;<code>Pathway Ontology</code>&nbsp;identifiers to&nbsp;<code>Reactome</code>&nbsp;pathways, we use&nbsp;<a href="https://github.com/ComPath/resources/tree/master/mappings">ComPath Pathway Database Mappings</a>&nbsp;developed by Daniel Domingo-Fern&aacute;ndez (<a href="https://www.ncbi.nlm.nih.gov/pubmed/30564458">PMID:30564458</a>).</p> <p><strong>Files</strong></p> <ul> <li>Downloaded Mapping Data <ul> <li><a href="http://compath.scai.fraunhofer.de/export_mappings"><code>curated_mappings.txt</code></a></li> <li><a href="https://github.com/ComPath/resources/blob/master/mappings/kegg_reactome.csv"><code>kegg_reactome.csv</code></a></li> </ul> </li> <li>Generated Mapping Data <ul> <li><a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/REACTOME_PW_GO_MAPPINGS.txt"><code>REACTOME_PW_GO_MAPPINGS.txt</code></a></li> </ul> </li> </ul> <p>&nbsp;</p> <p><strong>Protein Ontology (PRO)</strong></p> <p><strong>Homepage:</strong>&nbsp;<strong><a href="https://proconsortium.org/"><code>https://proconsortium.org/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Natale DA, Arighi CN, Barker WC, Blake JA, Bult CJ, Caudy M, Drabkin HJ, D&rsquo;Eustachio P, Evsikov AV, Huang H, Nchoutmboube J.&nbsp;<a href="https://academic.oup.com/nar/article-abstract/39/suppl_1/D539/2508558">The Protein Ontology: a structured representation of protein forms and complexes</a>. Nucleic Acids Research. 2010;39(suppl_1):D539-45</p> </blockquote> <p><strong>Usage:</strong>&nbsp;Utilized to connect&nbsp;<code>proteins</code>&nbsp;to&nbsp;<code>chemicals</code>,&nbsp;<code>genes</code>,&nbsp;<code>anatomy</code>,&nbsp;<code>catalysts</code>,&nbsp;<code>cell lines</code>,&nbsp;<code>cofactors</code>,&nbsp;<code>complexes</code>,&nbsp;<code>GO biological processes</code>,&nbsp;<code>GO cellular components</code>,&nbsp;<code>GO molecular functions</code>,&nbsp;<code>pathways</code>,&nbsp;<code>proteins</code>,&nbsp;<code>reactions</code>, and&nbsp;<code>transcripts</code>. Additionally, the edges between this ontology and its dependencies are utilized:</p> <ul> <li><strong><a href="https://www.ebi.ac.uk/chebi/"><code>ChEBI</code></a></strong></li> <li><strong><a href="http://disease-ontology.org/"><code>DOID</code></a></strong></li> <li><strong><a href="http://geneontology.org/"><code>GO</code></a></strong></li> </ul> <p><strong>Notes:</strong>&nbsp;A partial, human-only version of this ontology was used. Details on how this version of the ontology was generated can be found under the Protein Ontology section of the&nbsp;<a href="https://github.com/callahantiff/PheKnowLator/blob/master/notebooks/Data_Preparation.ipynb"><code>Data_Preparation.ipynb</code></a>&nbsp;Jupyter Notebook.</p> <p><strong>Files</strong></p> <ul> <li> <p>Generated Human Version Protein Ontology (PRO)</p> <ul> <li><a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/human_pro.owl"><code>human_pro.owl</code></a>&nbsp;(closed with&nbsp;<a href="http://www.hermit-reasoner.com/">hermit reasoner</a>)</li> </ul> </li> <li> <p>Other PRO Data Used:&nbsp;<a href="https://proconsortium.org/download/current/promapping.txt"><code>promapping.txt</code></a></p> </li> <li> <p>Generated Mapping Data</p> <ul> <li>Merged Gene, RNA, Protein Map:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/Merged_gene_rna_protein_identifiers.pkl"><code>Merged_gene_rna_protein_identifiers.pkl</code></a></li> <li>Ensembl Transcript-PRO Identifier Mapping:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/ENTREZ_GENE_ENSEMBL_TRANSCRIPT_MAP.txt"><code>ENSEMBL_TRANSCRIPT_PROTEIN_ONTOLOGY_MAP.txt</code></a></li> <li>Entrez Gene-PRO Identifier Mapping:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/ENTREZ_GENE_PRO_ONTOLOGY_MAP.txt"><code>ENTREZ_GENE_PRO_ONTOLOGY_MAP.txt</code></a></li> <li>UniProt Accession-PRO Identifier Mapping:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/UNIPROT_ACCESSION_PRO_ONTOLOGY_MAP.txt"><code>UNIPROT_ACCESSION_PRO_ONTOLOGY_MAP.txt</code></a></li> <li>STRING-PRO Identifier Mapping:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/STRING_PRO_ONTOLOGY_MAP.txt"><code>STRING_PRO_ONTOLOGY_MAP.txt</code></a></li> </ul> </li> </ul> <p>&nbsp;</p> <p><strong>Relations Ontology (RO)</strong></p> <p><strong>Homepage:</strong>&nbsp;<strong><a href="https://github.com/oborel/obo-relations/"><code>GitHub</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Smith B, Ceusters W, Klagges B, K&ouml;hler J, Kumar A, Lomax J, Mungall C, Neuhaus F, Rector AL, Rosse C.&nbsp;<a href="https://genomebiology.biomedcentral.com/articles/10.1186/gb-2005-6-5-r46">Relations in biomedical ontologies</a>. Genome Biology. 2005;6(5):R46.</p> </blockquote> <p><strong>Usage:</strong>&nbsp;Utilizing this ontology to connect all data sources in knowledge graph. Additionally, the ontology is queried prior to building the knowledge graph to identify all relations, their inverse properties, and their labels.</p> <p><strong>Files</strong></p> <ul> <li>Generated RO Data <ul> <li><a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/INVERSE_RELATIONS.txt"><code>INVERSE_RELATIONS.txt</code></a></li> <li><a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/RELATIONS_LABELS.txt"><code>RELATIONS_LABELS.txt</code></a></li> </ul> </li> </ul> <p>&nbsp;</p> <p><strong>Sequence Ontology (SO)</strong></p> <p><strong>Homepage:</strong>&nbsp;<strong><a href="https://github.com/The-Sequence-Ontology/SO-Ontologies"><code>GitHub</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Eilbeck K, Lewis SE, Mungall CJ, Yandell M, Stein L, Durbin R, Ashburner M.&nbsp;<a href="https://link.springer.com/article/10.1186/gb-2005-6-5-r44">The Sequence Ontology: a tool for the unification of genome annotations</a>. Genome Biology. 2005;6(5):R44</p> </blockquote> <p><strong>Usage:</strong>&nbsp;Utilized to connect&nbsp;<code>transcripts</code>&nbsp;and other genomic material like&nbsp;<code>genes</code>&nbsp;and&nbsp;<code>variants</code>.</p> <p><strong>Files</strong></p> <ul> <li>Generated Mapping Data <ul> <li><a href="https://storage.googleapis.com/pheknowlator/curated_data/genomic_sequence_ontology_mappings.xlsx"><code>genomic_sequence_ontology_mappings.xlsx</code></a></li> <li><a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/SO_GENE_TRANSCRIPT_VARIANT_TYPE_MAPPING.txt"><code>SO_GENE_TRANSCRIPT_VARIANT_TYPE_MAPPING.txt</code></a></li> </ul> </li> </ul> <p>&nbsp;</p> <p><strong>Uber-Anatomy Ontology (Uberon)</strong></p> <p><strong>Homepage:</strong>&nbsp;<strong><a href="https://uberon.github.io/"><code>GitHub</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Mungall CJ, Torniai C, Gkoutos GV, Lewis SE, Haendel MA.&nbsp;<a href="https://genomebiology.biomedcentral.com/articles/10.1186/gb-2012-13-1-r5">Uberon, an integrative multi-species anatomy ontology</a>. Genome Biology. 2012;13(1):R5</p> </blockquote> <p><strong>Usage:</strong>&nbsp;Utilized to connect&nbsp;<code>tissues</code>,&nbsp;<code>fluids</code>, and&nbsp;<code>cells</code>&nbsp;to&nbsp;<code>proteins</code>&nbsp;and&nbsp;<code>transcripts</code>. Additionally, the edges between this ontology and its dependencies are utilized:</p> <ul> <li><strong><a href="https://www.ebi.ac.uk/chebi/"><code>ChEBI</code></a></strong></li> <li><strong><a href="http://www.obofoundry.org/ontology/cl.html"><code>CL</code></a></strong></li> <li><strong><a href="http://geneontology.org/"><code>GO</code></a></strong></li> <li><strong><a href="https://proconsortium.org/"><code>PRO</code></a></strong></li> </ul> <p>&nbsp;</p> <p><strong>Vaccine Ontology (VO)</strong></p> <p><strong>Homepage:</strong>&nbsp;<strong><a href="http://www.violinet.org/vaccineontology/"><code>http://www.violinet.org/vaccineontology/</code></a></strong><br> <strong>Citations:</strong></p> <blockquote> <p>He Y, Racz R, Sayers S, Lin Y, Todd T, Hur J, Li X, Patel M, Zhao B, Chung M, Ostrow J.&nbsp;<a href="https://academic.oup.com/nar/article-abstract/42/D1/D1124/1053128">Updates on the web-based VIOLIN vaccine database and analysis system</a>. Nucleic Acids Research. 2013;42(D1):D1124-32</p> <p>Xiang Z, Todd T, Ku KP, Kovacic BL, Larson CB, Chen F, Hodges AP, Tian Y, Olenzek EA, Zhao B, Colby LA.&nbsp;<a href="https://academic.oup.com/nar/article-abstract/36/suppl_1/D923/2505793">VIOLIN: vaccine investigation and online information network</a>. Nucleic Acids Research. 2007;36(suppl_1):D923-8</p> </blockquote> <p><strong>Usage:</strong>&nbsp;Utilized the edges between this ontology and its dependencies:</p> <ul> <li><strong><a href="https://www.ebi.ac.uk/chebi/"><code>ChEBI</code></a></strong></li> <li><strong><a href="http://disease-ontology.org/"><code>DOID</code></a></strong></li> <li><strong><a href="http://geneontology.org/"><code>GO</code></a></strong></li> <li><strong><a href="https://proconsortium.org/"><code>PRO</code></a></strong></li> <li><strong><a href="https://uberon.github.io/"><code>UBERON</code></a></strong></li> </ul> <p>&nbsp;</p> <p><strong>DATABASE SOURCES</strong></p> <ul> <li>BioPortal</li> <li>ClinVar</li> <li>Comparative Toxicogenomics Database</li> <li>DisGeNET</li> <li>Ensembl</li> <li>GeneMANIA</li> <li>Genotype-Tissue Expression Project</li> <li>Human Genome Organisation Gene Nomenclature Committee</li> <li>Human Protein Atlas</li> <li>National Center for Biotechnology Information Gene</li> <li>Reactome Pathway Database</li> <li>Search Tool for Recurring Instances of Neighbouring Genes Database</li> <li>Universal Protein Resource Knowledgebase</li> </ul> <p>&nbsp;</p> <p><strong>BioPortal</strong></p> <p><strong>Homepage:</strong>&nbsp;<strong><a href="https://bioportal.bioontology.org/"><code>BioPortal</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>BioPortal.&nbsp;<a href="https://www.bioontology.org/wiki/LOOM">Lexical OWL Ontology Matcher (LOOM)</a></p> <p>Ghazvinian A, Noy NF, Musen MA.&nbsp;<a href="https://www.ncbi.nlm.nih.gov/pmc/articles/pmc2815474/">Creating mappings for ontologies in biomedicine: simple methods work</a>. In AMIA Annual Symposium Proceedings 2009 (Vol. 2009, p. 198). American Medical Informatics Association</p> </blockquote> <p><strong>Usage:</strong>&nbsp;BioPortal was utilized to obtain mappings between&nbsp;<code>MeSH identifiers</code>&nbsp;and&nbsp;<code>ChEBI identifiers</code>&nbsp;for&nbsp;<code>chemicals-diseases</code>,&nbsp;<code>chemicals-genes</code>,&nbsp;<code>chemical-GO biological processes</code>,&nbsp;<code>chemicals-GO cellular components</code>,&nbsp;<code>chemicals-GO molecular functions</code>,&nbsp;<code>chemicals-phenotypes</code>,&nbsp;<code>chemicals-proteins</code>, and&nbsp;<code>chemicals-transcripts</code>. Additional information on how this data was processed can be obtained from the&nbsp;<a href="https://gist.github.com/callahantiff/a28fb3160782f42f104e9ec41553af0d"><code>NCBO_rest_api.py</code></a>&nbsp;GitHub Gist script.</p> <p>⭐&nbsp;<strong>ALTERNATIVE METHOD</strong>⭐ Since the above approach can take over two days to process, we have developed an alternative solution that downloads the&nbsp;<a><code>mesh2021.nt</code></a>&nbsp;data file directly from MeSH and the&nbsp;<a><code>Flat_file_tab_delimited/names.tsv.gz</code></a>&nbsp;file directly from ChEBI. Using these files, we have recapitulated the&nbsp;<a href="https://www.bioontology.org/wiki/BioPortal_Mappings"><code>LOOM</code></a>&nbsp;algorithm implemented by BioPortal when creating mappings between these resources. The procedure is relatively straightforward and utilizes the following information from each resource:</p> <ul> <li>For all MeSH&nbsp;<code>SCR Chemicals</code>, obtain the following information: <ul> <li>Identifiers: MeSH identifiers</li> <li>Labels: string labels using the&nbsp;<code>RDFS:label</code>&nbsp;object property</li> <li>Synonyms: track down all synonyms using the&nbsp;<code>vocab:concept</code>&nbsp;and&nbsp;<code>vocab:preferredConcept</code>&nbsp;object properties</li> </ul> </li> <li>For all ChEBI classes, obtain the following information: <ul> <li>Labels: string labels using the&nbsp;<code>RDFS:label</code>&nbsp;object property</li> <li>Synonyms: track down all synonyms using all&nbsp;<code>synonym</code>&nbsp;object properties</li> </ul> </li> </ul> <p><strong>Files</strong></p> <ul> <li>Generated Data:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/MESH_CHEBI_MAP.txt"><code>MESH_CHEBI_MAP.txt</code></a></li> </ul> <p>&nbsp;</p> <p><strong>ClinVar</strong></p> <p><strong>Homepage:</strong>&nbsp;<strong><a href="https://www.ncbi.nlm.nih.gov/clinvar/"><code>https://www.ncbi.nlm.nih.gov/clinvar/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Landrum MJ, Lee JM, Benson M, Brown GR, Chao C, Chitipiralla S, Gu B, Hart J, Hoffman D, Jang W, Karapetyan K.&nbsp;<a href="https://academic.oup.com/nar/article-abstract/46/D1/D1062/4641904">ClinVar: improving access to variant interpretations and supporting evidence</a>. Nucleic Acids Research. 2017;46(D1):D1062-7</p> </blockquote> <p><strong>Usage:</strong>&nbsp;ClinVar was utilized to create&nbsp;<code>variant-gene</code>,&nbsp;<code>variant-disease</code>, and&nbsp;<code>variant-phenotype</code>&nbsp;edges. The original data is filtered such that only records meeting the following criteria were included:</p> <ul> <li> <p><code>Assembly</code>&nbsp;= &quot;GRCh38&quot;</p> </li> <li> <p><code>ClinSigSimple</code>&nbsp;=&nbsp;<code>1</code></p> <ul> <li> <blockquote> <p>1 = at least one current record submitted with an interpretation of Likely pathogenic or Pathogenic (independent of whether that record includes assertion criteria and evidence)&quot;</p> </blockquote> </li> </ul> </li> <li> <p><code>ReviewStatus</code>&nbsp;in [&quot;criteria provided, multiple submitters, no conflicts&quot;, &quot;reviewed by expert panel&quot;, &quot;practice guideline&quot;]</p> </li> </ul> <p><strong>Files</strong></p> <ul> <li> <p>Downloaded Data</p> <ul> <li><a href="https://ftp.ncbi.nlm.nih.gov/pub/clinvar/tab_delimited/variant_summary.txt.gz"><code>variant_summary.txt.gz</code></a></li> <li><a href="https://ftp.ncbi.nlm.nih.gov/pub/clinvar/tab_delimited/var_citations.txt"><code>var_citations.txt</code></a></li> <li><a href="https://ftp.ncbi.nlm.nih.gov/pub/clinvar/tab_delimited/allele_gene.txt.gz"><code>allele_gene.txt.gz</code></a></li> </ul> </li> <li> <p>Generated Edge Data:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/CLINVAR_VARIANT_GENE_DISEASE_PHENOTYPE_EDGES.txt"><code>CLINVAR_VARIANT_GENE_DISEASE_PHENOTYPE_EDGES.txt</code></a></p> </li> </ul> <p>&nbsp;</p> <p><strong>Comparative Toxicogenomics Database (CTD)</strong></p> <p><strong>Homepage:</strong>&nbsp;<strong><a href="http://ctdbase.org/"><code>http://ctdbase.org/</code></a></strong><br> <strong>Citations:</strong></p> <blockquote> <p>Curated [chemical&ndash;gene interactions|chemical-go interactions|chemical&ndash;disease interactions|gene&ndash;pathway interactions] data were retrieved from the Comparative Toxicogenomics Database (CTD), MDI Biological Laboratory, Salisbury Cove, Maine, and NC State University, Raleigh, North Carolina. World Wide Web (URL:&nbsp;<a href="http://ctdbase.org/">http://ctdbase.org/</a>)</p> <p>Davis AP, Grondin CJ, Johnson RJ, Sciaky D, McMorran R, Wiegers J, Wiegers TC, Mattingly CJ.&nbsp;<a href="https://academic.oup.com/nar/article-abstract/47/D1/D948/5106145">The comparative toxicogenomics database: update 2019</a>. Nucleic Acids Research. 2018;47(D1):D948-54</p> </blockquote> <p>Usage: Comparative Toxicogenomics Database (CTD) was utilized to create&nbsp;<code>chemical-disease</code>,&nbsp;<code>chemical-gene</code>,&nbsp;<code>chemical-GO biological process</code>,&nbsp;<code>chemical-GO cellular components</code>,&nbsp;<code>chemical-GO molecular functions</code>,&nbsp;<code>chemical-phenotype</code>,&nbsp;<code>chemical-protein</code>,&nbsp;<code>chemical-rna</code>, and&nbsp;<code>gene-pathway</code>&nbsp;edges. The original data is filtered such that only records meeting the following criteria were included:</p> <ul> <li><code>chemical-disease</code>:&nbsp;<code>DirectEvidence</code>&nbsp;!= &quot;&quot;</li> <li><code>chemical-gene</code>:&nbsp;<code>Organism</code>&nbsp;== &quot;Homo sapiens&quot;,&nbsp;<code>GeneForms</code>&nbsp;== &quot;gene&quot;, and affects not in&nbsp;<code>InteractionActions</code></li> <li><code>chemical-GO biological process</code>:&nbsp;<code>PhenotypeName</code>&nbsp;== &quot;Biological Process&quot; and&nbsp;<code>Interaction</code>&nbsp;&lt;= &quot;1.04e-47&quot; (10th percentile)</li> <li><code>chemical-GO cellular components</code>:&nbsp;<code>PhenotypeName</code>&nbsp;== &quot;Cellular Component&quot; and&nbsp;<code>Interaction</code>&nbsp;&lt;= &quot;1.04e-47&quot; (10th percentile)</li> <li><code>chemical-GO molecular functions</code>:&nbsp;<code>PhenotypeName</code>&nbsp;== &quot;Molecular Function&quot; and&nbsp;<code>Interaction</code>&nbsp;&lt;= &quot;1.04e-47&quot; (10th percentile)</li> <li><code>chemical-phenotype</code>:&nbsp;<code>DirectEvidence</code>&nbsp;!= &quot;&quot;</li> <li><code>chemical-protein</code>:&nbsp;<code>Organism</code>&nbsp;== &quot;Homo sapiens&quot;,&nbsp;<code>GeneForms</code>&nbsp;== &quot;protein&quot;, and affects not in&nbsp;<code>InteractionActions</code></li> <li><code>chemical-rna</code>:&nbsp;<code>Organism</code>&nbsp;== &quot;Homo sapiens&quot;,&nbsp;<code>GeneForms</code>&nbsp;== &quot;mRNA&quot;, and affects and activity not in&nbsp;<code>InteractionActions</code></li> <li><code>gene-pathway edges</code>:&nbsp;<code>PathwayName</code>&nbsp;== R-HSA-</li> </ul> <p><strong>Files</strong></p> <ul> <li>Downloaded Data <ul> <li>Chemical-Gene Relations:&nbsp;<a href="http://ctdbase.org/reports/CTD_chem_gene_ixns.tsv.gz"><code>CTD_chem_gene_ixns.tsv.gz</code></a></li> <li>Chemical-Disease/Phenotype Relations:&nbsp;<a href="http://ctdbase.org/reports/CTD_chemicals_diseases.tsv.gz"><code>CTD_chemicals_diseases.tsv.gz</code></a></li> <li>Chemical-GO Relations:&nbsp;<a href="http://ctdbase.org/reports/CTD_chem_go_enriched.tsv.gz"><code>CTD_chem_go_enriched.tsv.gz</code></a></li> <li>Gene-Pathway Relations:&nbsp;<a href="http://ctdbase.org/reports/CTD_genes_pathways.tsv.gz"><code>CTD_genes_pathways.tsv.gz</code></a></li> </ul> </li> </ul> <p>&nbsp;</p> <p><strong>DisGeNET</strong></p> <p><strong>Homepage:</strong>&nbsp;<strong><a href="https://www.disgenet.org/"><code>https://www.disgenet.org/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Gene-disease association data retrieved from DisGeNET v6.0 (<a href="http://www.disgenet.org/">http://www.disgenet.org/</a>), Integrative Biomedical Informatics Group GRIB/IMIM/UPF. [December, 2019].</p> <p>Pi&ntilde;ero J, Ram&iacute;rez-Anguita JM, Sa&uuml;ch-Pitarch J, Ronzano F, Centeno E, Sanz F, Furlong LI.&nbsp;<a href="https://academic.oup.com/nar/advance-article-abstract/doi/10.1093/nar/gkz1021/5611674">The DisGeNET knowledge platform for disease genomics: 2019 update</a>. Nucleic Acids Research. 2019.</p> </blockquote> <p><strong>Usage:</strong>&nbsp;DisGeNET was utilized to create&nbsp;<code>gene-disease</code>, and&nbsp;<code>gene-phenotype</code>&nbsp;edges. The original data is filtered such that only records meeting the following criteria were included:&nbsp;<code>EI</code>&nbsp;&gt;= &quot;1.0&quot; (90th percentile). Additionally, data from this source was used to create mappings between different types of disease and phenotype identifiers, including:</p> <ul> <li>OMIM, ORPHA, UMLS, ICD ➞ DOID</li> <li>OMIM, ORPHA, UMLS, ICD ➞ HPO</li> </ul> <p><strong>Files</strong></p> <ul> <li> <p>Downloaded Data</p> <ul> <li>Disease/Phenotype-Gene Relations:&nbsp;<a href="https://www.disgenet.org/static/disgenet_ap1/files/downloads/curated_gene_disease_associations.tsv.gz"><code>curated_gene_disease_associations.tsv.gz</code></a></li> <li>Disease Identifier Mapping:&nbsp;<a href="https://www.disgenet.org/static/disgenet_ap1/files/downloads/disease_mappings.tsv.gz"><code>disease_mappings.tsv.gz</code></a></li> </ul> </li> <li> <p>Generated Mapping Data</p> <ul> <li>Disease Identifier Mapping:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/PHENOTYPE_HPO_MAP.txt"><code>PHENOTPYE_HPO_MAP.txt</code></a></li> <li>Phenotype Identifier Mapping:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/DISEASE_MONDO_MAP.txt"><code>DISEASE_DOID_MAP.txt</code></a></li> </ul> </li> </ul> <p>&nbsp;</p> <p><strong>Ensembl</strong></p> <p><strong>Homepage:</strong>&nbsp;<strong><a href="https://uswest.ensembl.org/index.html"><code>https://uswest.ensembl.org/index.html</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Zerbino DR, Achuthan P, Akanni W, Amode MR, Barrell D, Bhai J, Billis K, Cummins C, Gall A, Gir&oacute;n CG, Gil L.&nbsp;<a href="https://academic.oup.com/nar/article/46/D1/D754/4634002">Ensembl 2018</a>. Nucleic Acids Research. 2017;46(D1):D754-61</p> </blockquote> <p><strong>Usage:</strong>&nbsp;Ensembl data was utilized to create mappings between Ensembl genes, transcripts, and proteins with&nbsp;<a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#national-center-for-biotechnology-information-gene">NCBI Gene identifiers</a>,&nbsp;<a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#human-genome-organisation-gene-nomenclature-committee">HUGO gene symbols</a>,&nbsp;<a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#universal-protein-resource-knowledgebase">UniProt Accession identifiers</a>, and&nbsp;<a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#protein-ontology">Protein Ontology identifiers</a>&nbsp;in the knowledge graph (for additional details on the processing of these data, see&nbsp;<a href="https://github.com/callahantiff/PheKnowLator/blob/master/notebooks/Data_Preparation.ipynb"><code>Data_Preparation.ipynb</code></a>):</p> <ul> <li>Ensembl Transcript IDs ➞ PRO IDs</li> <li>Gene Ensembl IDs ➞ Entrez Gene IDs</li> <li>Gene Ensembl IDs ➞ PRO IDs</li> <li>Gene Symbols ➞ Transcript Ensembl IDs</li> <li>Entrez Gene IDs ➞ Transcript Ensembl IDs</li> <li>Entrez Gene IDs ➞ PRO IDs</li> <li>Protein Ensembl IDs ➞ UniProt Protein Accession</li> <li>STRING IDs ➞ PRO IDs</li> <li>UniProt Protein Accession ➞ Entrez Gene IDs</li> </ul> <p><strong>Files</strong></p> <ul> <li> <p>Downloaded Data</p> <ul> <li><a><code>Homo_sapiens.GRCh38.102.gtf</code></a></li> <li><a><code>Homo_sapiens.GRCh38.102.uniprot.tsv.gz</code></a></li> <li><a><code>Homo_sapiens.GRCh38.102.entrez.tsv.gz</code></a></li> </ul> </li> <li> <p>Generated Mapping Data</p> <ul> <li>Cleaned Ensembl Gene Set:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/ensembl_identifier_data_cleaned.txt"><code>ensembl_identifier_data_cleaned.txt</code></a></li> <li>Merged Gene, RNA, Protein Map:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/Merged_gene_rna_protein_identifiers.pkl"><code>Merged_gene_rna_protein_identifiers.pkl</code></a></li> <li>Ensembl Transcript-PRO Identifier Mapping:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/ENSEMBL_TRANSCRIPT_PROTEIN_ONTOLOGY_MAP.txt"><code>ENSEMBL_TRANSCRIPT_PROTEIN_ONTOLOGY_MAP.txt</code></a></li> <li>Gene Symbol-Ensembl Transcript Identifier Mapping:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/GENE_SYMBOL_ENSEMBL_TRANSCRIPT_MAP.txt"><code>GENE_SYMBOL_ENSEMBL_TRANSCRIPT_MAP.txt</code></a></li> <li>Entrez Gene-Ensembl Transcript Identifier Mapping:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/ENTREZ_GENE_ENSEMBL_TRANSCRIPT_MAP.txt"><code>ENTREZ_GENE_ENSEMBL_TRANSCRIPT_MAP.txt</code></a></li> <li>Entrez Gene-PRO Identifier Mapping:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/ENTREZ_GENE_PRO_ONTOLOGY_MAP.txt"><code>ENTREZ_GENE_PRO_ONTOLOGY_MAP.txt</code></a></li> <li>Ensembl Gene-Entrez Gene Identifier Mapping:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/ENSEMBL_GENE_ENTREZ_GENE_MAP.txt"><code>ENSEMBL_GENE_ENTREZ_GENE_MAP.txt</code></a></li> </ul> </li> </ul> <p>&nbsp;</p> <p><strong>GeneMANIA</strong></p> <p><strong>Homepage:</strong>&nbsp;<strong><a href="https://genemania.org/"><code>https://genemania.org/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Warde-Farley D, Donaldson SL, Comes O, Zuberi K, Badrawi R, Chao P, Franz M, Grouios C, Kazi F, Lopes CT, Maitland A.&nbsp;<a href="https://academic.oup.com/nar/article-abstract/38/suppl_2/W214/1126704">The GeneMANIA prediction server: biological network integration for gene prioritization and predicting gene function</a>. Nucleic Acids Research. 2010;38(suppl_2):W214-20</p> </blockquote> <p><strong>Usage:</strong>&nbsp;GeneMANIA was utilized to create&nbsp;<code>gene-gene</code>&nbsp;edges.</p> <p><strong>Files</strong></p> <ul> <li>Downloaded Data:&nbsp;<a href="http://genemania.org/data/current/Homo_sapiens.COMBINED/COMBINED.DEFAULT_NETWORKS.BP_COMBINING.txt"><code>COMBINED.DEFAULT_NETWORKS.BP_COMBINING.txt</code></a></li> </ul> <p>&nbsp;</p> <p><strong>Genotype-Tissue Expression Project (GTEx)</strong></p> <p><strong>Homepage:</strong>&nbsp;<strong><a href="https://gtexportal.org/home/"><code>https://gtexportal.org/home/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Lonsdale J, Thomas J, Salvatore M, Phillips R, Lo E, Shad S, Hasz R, Walters G, Garcia F, Young N, Foster B.&nbsp;<a href="http://www.nature.com/ng/journal/v45/n6/full/ng.2653.html">The genotype-tissue expression (GTEx) project</a>. Nature Genetics. 2013;45(6):580</p> </blockquote> <p><strong>Usage:</strong>&nbsp;The Genotype-Tissue Expression (GTEx) Project was utilized to create edges between&nbsp;<code>protein-cell</code>,&nbsp;<code>protein-anatomy</code>,&nbsp;<code>rna-cell</code>&nbsp;and&nbsp;<code>rna-anatomy</code>&nbsp;entities. The original data were filtered such that only those edges where the median TPM was &gt;=<code>1.0</code>&nbsp;and genes were of any type other than protein-coding were included. It should also be noted that we chose to use the RNASeQC file over the RSEM file as advised by the GTEx website.</p> <blockquote> <p>The RSEM estimates are based on combining isoform-level estimates, which adds uncertainty to the resulting gene-level values (the isoform-level estimates are highly inaccurate in some cases).</p> </blockquote> <p>The file contains&nbsp;<code>54</code>&nbsp;unique tissue and/or cell types. GTEx provides mappings from tissue types to UBERON and EFO. These provided&nbsp;<a href="https://gtexportal.org/home/samplingSitePage">mappings</a>&nbsp;were verified and extended, such that all samples which referenced a cell type were also mapped to the Cell and the Cell Line ontologies. This resulted in a total of&nbsp;<code>56</code>&nbsp;mappings (<code>1.04</code>&nbsp;mappings/concepts).</p> <p><strong>Files</strong></p> <ul> <li> <p>Downloaded Data:&nbsp;<a href="https://storage.googleapis.com/gtex_analysis_v8/rna_seq_data/GTEx_Analysis_2017-06-05_v8_RNASeQCv1.1.9_gene_median_tpm.gct.gz"><code>GTEx_Analysis_2017-06-05_v8_RNASeQCv1.1.9_gene_median_tpm.gct</code></a></p> </li> <li> <p>Mapping Results:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/curated_data/zooma_tissue_cell_mapping_04JAN2020.xlsx"><code>zooma_tissue_cell_mapping_04JAN2020.xlsx</code></a></p> </li> <li> <p>Generated Data<br> The final mapping set was combined with terms from the&nbsp;<a href="https://www.proteinatlas.org/">Human Protein Atlas</a>, see&nbsp;<a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources/t#human-protein-atlas">here</a>&nbsp;for more information.</p> <ul> <li>All HPA tissue and cell type strings:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/HPA_tissues.txt"><code>HPA_tissues.txt</code></a></li> <li>Final Term Mapping:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/HPA_GTEx_TISSUE_CELL_MAP.txt"><code>HPA_GTEx_TISSUE_CELL_MAP.txt</code></a></li> <li>Final RNA, Gene, Protein-Tissues and Cell Types Relations:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/HPA_GTEX_RNA_GENE_PROTEIN_EDGES.txt"><code>HPA_GTEX_RNA_GENE_PROTEIN_EDGES.txt</code></a></li> </ul> </li> </ul> <p>&nbsp;</p> <p><strong>Human Genome Organisation Gene Nomenclature Committee (HUGO)</strong></p> <p><strong>Homepage:</strong>&nbsp;<strong><a href="https://www.genenames.org/"><code>https://www.genenames.org/</code></a></strong><br> <strong>Citations:</strong></p> <blockquote> <p>HGNC Database, HUGO Gene Nomenclature Committee (HGNC), European Molecular Biology Laboratory, European Bioinformatics Institute (EMBL-EBI), Wellcome Genome Campus, Hinxton, Cambridge CB10 1SD, United Kingdom&nbsp;<a href="https://www.genenames.org/">www.genenames.org</a></p> <p>Yates B, Braschi B, Gray K, Seal R, Tweedie S, Bruford E.&nbsp;<a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5210531/">Genenames.org: the HGNC and VGNC Resources in 2017</a>. Nucleic Acids Research. 2017;45(D1):D619-625</p> </blockquote> <p><strong>Usage:</strong>&nbsp;The Human Genome Organisation (HUGO) data was utilized to obtain mappings between&nbsp;<a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#national-center-for-biotechnology-information-gene">NCBI Gene identifiers</a>, HUGO gene symbols,&nbsp;<a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#universal-protein-resource-knowledgebase">UniProt Accession identifiers</a>, and&nbsp;<a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#protein-ontology">Protein Ontology identifiers</a>. For additional details on the processing of these data, see&nbsp;<a href="https://github.com/callahantiff/PheKnowLator/blob/master/notebooks/Data_Preparation.ipynb"><code>Data_Preparation.ipynb</code></a>:</p> <ul> <li>Ensembl Transcript IDs ➞ PRO IDs</li> <li>Gene Ensembl IDs ➞ Entrez Gene IDs</li> <li>Gene Ensembl IDs ➞ PRO IDs</li> <li>Gene Symbols ➞ Transcript Ensembl IDs</li> <li>Entrez Gene IDs ➞ Transcript Ensembl IDs</li> <li>Entrez Gene IDs ➞ PRO IDs</li> <li>Protein Ensembl IDs ➞ UniProt Protein Accession</li> <li>STRING IDs ➞ PRO IDs</li> <li>UniProt Protein Accession ➞ Entrez Gene IDs</li> </ul> <p><strong>Files</strong></p> <ul> <li> <p>Downloaded Data:&nbsp;<a href="http://ftp.ebi.ac.uk/pub/databases/genenames/hgnc/tsv/hgnc_complete_set.txt"><code>hgnc_complete_set.txt</code></a></p> </li> <li> <p>Generated Data</p> <ul> <li>Merged Gene, RNA, Protein Map:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/Merged_gene_rna_protein_identifiers.pkl"><code>Merged_gene_rna_protein_identifiers.pkl</code></a></li> <li>Gene Symbol-Ensembl Transcript Identifier Mapping:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/GENE_SYMBOL_ENSEMBL_TRANSCRIPT_MAP.txt"><code>GENE_SYMBOL_ENSEMBL_TRANSCRIPT_MAP.txt</code></a></li> </ul> </li> </ul> <p>&nbsp;</p> <p><strong>Human Protein Atlas (HPA)</strong></p> <p><strong>Homepage:</strong>&nbsp;<strong><a href="https://www.proteinatlas.org/"><code>https://www.proteinatlas.org/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Uhl&eacute;n M, Fagerberg L, Hallstr&ouml;m BM, Lindskog C, Oksvold P, Mardinoglu A, Sivertsson &Aring;, Kampf C, Sj&ouml;stedt E, Asplund A, Olsson I.&nbsp;<a href="https://science.sciencemag.org/content/347/6220/1260419.short">Tissue-based map of the human proteome</a>. Science. 2015;347(6220):1260419</p> </blockquote> <p><strong>Usage:</strong>&nbsp;The Human Protein Atlas (HPA) was utilized to create&nbsp;<code>rna-cell</code>,&nbsp;<code>rna-anatomy</code>,&nbsp;<code>protein-cell</code>, and&nbsp;<code>protein-anatomy</code>&nbsp;edges. Evidence between gene and RNA expression in specific tissue types was derived by HPA, such that the&nbsp;<a href="https://www.proteinatlas.org/about/assays+annotation#normalization_rna">consensus normalized expression</a>&nbsp;was &gt;=<code>1.0</code>. Zooma was utilized to automatically annotate the&nbsp;<code>153</code>&nbsp;unique tissues and cell types from Human Protein Atlas for all human protein-coding genes in the&nbsp;<a href="https://www.proteinatlas.org/humanproteome">Human Proteome</a>&nbsp;to the Cell Ontology, Cell Line Ontology, and the Uber-Anatomy Ontology. To best represent each concept, the automatic mappings from Zooma were extend through manual mapping efforts to ensure each concept cell type was matched to a Cell Ontology, Cell Line Ontology, and UBERON ontology term. This resulted in a total of&nbsp;<code>281</code>&nbsp;mappings (<code>1.84</code>&nbsp;mappings/concepts).</p> <p><strong>Files</strong></p> <ul> <li> <p>Downloaded Data:&nbsp;<a href="https://www.proteinatlas.org/api/search_download.php?search=&amp;columns=g,eg,up,pe,rnatsm,rnaclsm,rnacasm,rnabrsm,rnabcsm,rnablsm,scl,t_RNA_adipose_tissue,t_RNA_adrenal_gland,t_RNA_amygdala,t_RNA_appendix,t_RNA_basal_ganglia,t_RNA_bone_marrow,t_RNA_breast,t_RNA_cerebellum,t_RNA_cerebral_cortex,t_RNA_cervix,_uterine,t_RNA_colon,t_RNA_corpus_callosum,t_RNA_ductus_deferens,t_RNA_duodenum,t_RNA_endometrium_1,t_RNA_epididymis,t_RNA_esophagus,t_RNA_fallopian_tube,t_RNA_gallbladder,t_RNA_heart_muscle,t_RNA_hippocampal_formation,t_RNA_hypothalamus,t_RNA_kidney,t_RNA_liver,t_RNA_lung,t_RNA_lymph_node,t_RNA_midbrain,t_RNA_olfactory_region,t_RNA_ovary,t_RNA_pancreas,t_RNA_parathyroid_gland,t_RNA_pituitary_gland,t_RNA_placenta,t_RNA_pons_and_medulla,t_RNA_prostate,t_RNA_rectum,t_RNA_retina,t_RNA_salivary_gland,t_RNA_seminal_vesicle,t_RNA_skeletal_muscle,t_RNA_skin_1,t_RNA_small_intestine,t_RNA_smooth_muscle,t_RNA_spinal_cord,t_RNA_spleen,t_RNA_stomach_1,t_RNA_testis,t_RNA_thalamus,t_RNA_thymus,t_RNA_thyroid_gland,t_RNA_tongue,t_RNA_tonsil,t_RNA_urinary_bladder,t_RNA_vagina,t_RNA_B-cells,t_RNA_dendritic_cells,t_RNA_granulocytes,t_RNA_monocytes,t_RNA_NK-cells,t_RNA_T-cells,t_RNA_total_PBMC,cell_RNA_A-431,cell_RNA_A549,cell_RNA_AF22,cell_RNA_AN3-CA,cell_RNA_ASC_diff,cell_RNA_ASC_TERT1,cell_RNA_BEWO,cell_RNA_BJ,cell_RNA_BJ_hTERT+,cell_RNA_BJ_hTERT+_SV40_Large_T+,cell_RNA_BJ_hTERT+_SV40_Large_T+_RasG12V,cell_RNA_CACO-2,cell_RNA_CAPAN-2,cell_RNA_Daudi,cell_RNA_EFO-21,cell_RNA_fHDF/TERT166,cell_RNA_HaCaT,cell_RNA_HAP1,cell_RNA_HBEC3-KT,cell_RNA_HBF_TERT88,cell_RNA_HDLM-2,cell_RNA_HEK_293,cell_RNA_HEL,cell_RNA_HeLa,cell_RNA_Hep_G2,cell_RNA_HHSteC,cell_RNA_HL-60,cell_RNA_HMC-1,cell_RNA_HSkMC,cell_RNA_hTCEpi,cell_RNA_hTEC/SVTERT24-B,cell_RNA_hTERT-HME1,cell_RNA_HUVEC_TERT2,cell_RNA_K-562,cell_RNA_Karpas-707,cell_RNA_LHCN-M2,cell_RNA_MCF7,cell_RNA_MOLT-4,cell_RNA_NB-4,cell_RNA_NTERA-2,cell_RNA_PC-3,cell_RNA_REH,cell_RNA_RH-30,cell_RNA_RPMI-8226,cell_RNA_RPTEC_TERT1,cell_RNA_RT4,cell_RNA_SCLC-21H,cell_RNA_SH-SY5Y,cell_RNA_SiHa,cell_RNA_SK-BR-3,cell_RNA_SK-MEL-30,cell_RNA_T-47d,cell_RNA_THP-1,cell_RNA_TIME,cell_RNA_U-138_MG,cell_RNA_U-2_OS,cell_RNA_U-2197,cell_RNA_U-251_MG,cell_RNA_U-266/70,cell_RNA_U-266/84,cell_RNA_U-698,cell_RNA_U-87_MG,cell_RNA_U-937,cell_RNA_WM-115,blood_RNA_basophil,blood_RNA_classical_monocyte,blood_RNA_eosinophil,blood_RNA_gdT-cell,blood_RNA_intermediate_monocyte,blood_RNA_MAIT_T-cell,blood_RNA_memory_B-cell,blood_RNA_memory_CD4_T-cell,blood_RNA_memory_CD8_T-cell,blood_RNA_myeloid_DC,blood_RNA_naive_B-cell,blood_RNA_naive_CD4_T-cell,blood_RNA_naive_CD8_T-cell,blood_RNA_neutrophil,blood_RNA_NK-cell,blood_RNA_non-classical_monocyte,blood_RNA_plasmacytoid_DC,blood_RNA_T-reg,blood_RNA_total_PBMC,brain_RNA_amygdala,brain_RNA_basal_ganglia,brain_RNA_cerebellum,brain_RNA_cerebral_cortex,brain_RNA_hippocampal_formation,brain_RNA_hypothalamus,brain_RNA_midbrain,brain_RNA_olfactory_region,brain_RNA_pons_and_medulla,brain_RNA_thalamus&amp;format=tsv"><code>proteinatlas_search.tsv</code></a></p> </li> <li> <p>Mapping Results:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/curated_data/zooma_tissue_cell_mapping_04JAN2020.xlsx"><code>zooma_tissue_cell_mapping_04JAN2020.xlsx</code></a></p> </li> <li> <p>Generated Data</p> <ul> <li>Final Term Mapping:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/HPA_GTEx_TISSUE_CELL_MAP.txt"><code>HPA_GTEx_TISSUE_CELL_MAP.txt</code></a></li> <li>Final RNA, Gene, Protein-Tissues and Cell Types Relations:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/HPA_GTEX_RNA_GENE_PROTEIN_EDGES.txt"><code>HPA_GTEX_RNA_GENE_PROTEIN_EDGES.txt</code></a></li> </ul> </li> </ul> <p>&nbsp;</p> <p><strong>National Center for Biotechnology Information (NCBI) Entrez Gene</strong></p> <p><strong>Homepage:</strong>&nbsp;<strong><a href="https://www.ncbi.nlm.nih.gov/gene/"><code>https://www.ncbi.nlm.nih.gov/gene/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Maglott D, Ostell J, Pruitt KD, Tatusova T.&nbsp;<a href="https://academic.oup.com/nar/article-abstract/33/suppl_1/D54/2505255">Entrez Gene: gene-centered information at NCBI</a>. Nucleic Acids Research. 2005;33(suppl_1):D54-8.</p> </blockquote> <p><strong>Usage:</strong>&nbsp;The National Center for Biotechnology Information (NCBI) Gene data was utilized to obtain mappings between&nbsp;<a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#ncbi-gene">NCBI Gene identifiers</a>,&nbsp;<a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#hugo-gene-nomenclature-committee">HUGO gene symbols</a>,&nbsp;<a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#uniprot-knowledgebase">UniProt Accession identifiers</a>, and&nbsp;<a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#protein-ontology">Protein Ontology identifiers</a>. For additional details on the processing of these data, see&nbsp;<a href="https://github.com/callahantiff/PheKnowLator/blob/master/notebooks/Data_Preparation.ipynb"><code>Data_Preparation.ipynb</code></a>:</p> <ul> <li>Ensembl Transcript IDs ➞ PRO IDs</li> <li>Gene Ensembl IDs ➞ Entrez Gene IDs</li> <li>Gene Ensembl IDs ➞ PRO IDs</li> <li>Gene Symbols ➞ Transcript Ensembl IDs</li> <li>Entrez Gene IDs ➞ Transcript Ensembl IDs</li> <li>Entrez Gene IDs ➞ PRO IDs</li> <li>Protein Ensembl IDs ➞ UniProt Protein Accession</li> <li>STRING IDs ➞ PRO IDs</li> <li>UniProt Protein Accession ➞ Entrez Gene IDs</li> </ul> <p><strong>Files</strong></p> <ul> <li> <p>Downloaded Data:&nbsp;<a href="https://ftp.ncbi.nih.gov/gene/DATA/GENE_INFO/Mammalia/Homo_sapiens.gene_info.gz"><code>Homo_sapiens.gene_info.gz</code></a></p> </li> <li> <p>Generated Data</p> <ul> <li>Merged Gene, RNA, Protein Map:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/Merged_gene_rna_protein_identifiers.pkl"><code>Merged_gene_rna_protein_identifiers.pkl</code></a></li> <li>Entrez Gene-Ensembl Transcript Identifier Mapping:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/ENTREZ_GENE_ENSEMBL_TRANSCRIPT_MAP.txt"><code>ENTREZ_GENE_ENSEMBL_TRANSCRIPT_MAP.txt</code></a></li> <li>Entrez Gene-PRO Identifier Mapping:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/ENTREZ_GENE_PRO_ONTOLOGY_MAP.txt"><code>ENTREZ_GENE_PRO_ONTOLOGY_MAP.txt</code></a></li> <li>Ensembl Gene-Entrez Gene Identifier Mapping:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/ENSEMBL_GENE_ENTREZ_GENE_MAP.txt"><code>ENSEMBL_GENE_ENTREZ_GENE_MAP.txt</code></a></li> <li>Uniprot Accession-Entrez Gene Identifier Mapping:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/UNIPROT_ACCESSION_ENTREZ_GENE_MAP.txt"><code>UNIPROT_ACCESSION_ENTREZ_GENE_MAP.txt</code></a></li> </ul> </li> </ul> <p>&nbsp;</p> <p><strong>Reactome Pathway Database</strong></p> <p><strong>Homepage:</strong>&nbsp;<strong><a href="https://reactome.org/"><code>https://reactome.org/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Fabregat A, Jupe S, Matthews L, Sidiropoulos K, Gillespie M, Garapati P, Haw R, Jassal B, Korninger F, May B, Milacic M.&nbsp;<a href="https://academic.oup.com/nar/article-abstract/46/D1/D649/4626770">The reactome pathway knowledgebase</a>. Nucleic Acids Research. 2017;46(D1):D649-55</p> </blockquote> <p><strong>Usage:</strong>&nbsp;The Reactome Database was utilized to create&nbsp;<code>chemical-pathway</code>,&nbsp;<code>GO Biological process-pathway</code>,&nbsp;<code>pathway-GO Cellular component</code>,&nbsp;<code>GO Molecular function-pathway</code>, and&nbsp;<code>protein-pathway</code>&nbsp;edges. The original data is filtered such that only records meeting the following criteria were included:</p> <ul> <li><code>chemical-pathway</code>: column[5] == &quot;Homo sapiens&quot;</li> <li><code>GO Biological process-pathway</code>: column[5] startswith &quot;REACTOME&quot;, column[8] == &quot;P&quot;, and column[12] == &quot;taxon:9606&quot;</li> <li><code>pathway-GO Cellular component</code>: column[5] startswith &quot;REACTOME&quot;, column[8] == &quot;C&quot;, and column[12] == &quot;taxon:9606&quot;</li> <li><code>GO Molecular function-pathway</code>: column[5] startswith &quot;REACTOME&quot;, column[8] == &quot;F&quot;, and column[12] == &quot;taxon:9606&quot;</li> <li><code>protein-pathway</code>: column[5] == &quot;Homo sapiens&quot;</li> </ul> <p><strong>Files</strong></p> <ul> <li>Downloaded Data <ul> <li>Chemical-Pathway Relations:&nbsp;<a href="https://reactome.org/download/current/ChEBI2Reactome_All_Levels.txt"><code>ChEBI2Reactome_All_Levels.txt</code></a></li> <li>Pathway-GO Relations:&nbsp;<a href="https://reactome.org/download/current/gene_association.reactome"><code>gene_association.reactome</code></a></li> <li>Protein-Pathway Relations:&nbsp;<a href="https://reactome.org/download/current/UniProt2Reactome_All_Levels.txt"><code>UniProt2Reactome_All_Levels.txt</code></a></li> </ul> </li> </ul> <p>&nbsp;</p> <p><strong>Search Tool for Recurring Instances of Neighbouring Genes (STRING) Database</strong></p> <p><strong>Homepage:</strong>&nbsp;<strong><a href="https://string-db.org/"><code>string-db.org</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Szklarczyk D, Gable AL, Lyon D, Junge A, Wyder S, Huerta-Cepas J, Simonovic M, Doncheva NT, Morris JH, Bork P, Jensen LJ.&nbsp;<a href="https://academic.oup.com/nar/article-abstract/47/D1/D607/5198476">STRING v11: protein&ndash;protein association networks with increased coverage, supporting functional discovery in genome-wide experimental datasets</a>. Nucleic Acids Research. 2018;47(D1):D607-13</p> </blockquote> <p><strong>Usage:</strong>&nbsp;The Search Tool for Recurring Instances of Neighbouring Genes (STRING) Database was utilized to create&nbsp;<code>protein-protein</code>&nbsp;edges. The original data is filtered such that only records meeting the following criteria were included:&nbsp;<code>combined_score</code>&nbsp;&gt;= &quot;700&quot; (&gt;90th percentile).</p> <p><strong>Files</strong></p> <ul> <li> <p>Downloaded Data:&nbsp;<a href="https://stringdb-static.org/download/protein.links.v11.0/9606.protein.links.v11.0.txt.gz"><code>9606.protein.links.v11.0.txt.gz</code></a></p> </li> <li> <p>Generated Data: STRING-PRO Identifier Mapping:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/STRING_PRO_ONTOLOGY_MAP.txt"><code>STRING_PRO_ONTOLOGY_MAP.txt</code></a></p> </li> </ul> <p>&nbsp;</p> <p><strong>Universal Protein Resource (UniProt) Knowledgebase</strong></p> <p><strong>Homepage:</strong>&nbsp;<strong><a href="https://www.uniprot.org/"><code>https://www.uniprot.org/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>UniProt Consortium.&nbsp;<a href="https://academic.oup.com/nar/article-abstract/47/D1/D506/5160987">UniProt: a worldwide hub of protein knowledge</a>. Nucleic acids research. 2018;47(D1):D506-15</p> </blockquote> <p><strong>Usage:</strong>&nbsp;The Universal Protein Resource (UniProt) Knowledgebase was utilized to obtain&nbsp;<code>cofactor</code>/<code>catalyst</code>-<code>protein</code>&nbsp;and&nbsp;<code>protein-coding gene</code>-<code>protein</code>&nbsp;edges as well as mappings between&nbsp;<a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#national-center-for-biotechnology-information-gene">NCBI Gene identifiers</a>,&nbsp;<a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#human-genome-organisation-gene-nomenclature-committee">HUGO gene symbols</a>,&nbsp;<a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#universal-protein-resource-knowledgebase">Universal Protein Resource (UniProt) Accession identifiers</a>, and&nbsp;<a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#protein-ontology">Protein Ontology identifiers</a>. For additional details on the processing of these data, see&nbsp;<a href="https://github.com/callahantiff/PheKnowLator/blob/master/notebooks/Data_Preparation.ipynb"><code>Data_Preparation.ipynb</code></a>:</p> <ul> <li>Ensembl Transcript IDs ➞ PRO IDs</li> <li>Gene Ensembl IDs ➞ Entrez Gene IDs</li> <li>Gene Ensembl IDs ➞ PRO IDs</li> <li>Gene Symbols ➞ Transcript Ensembl IDs</li> <li>Entrez Gene IDs ➞ Transcript Ensembl IDs</li> <li>Entrez Gene IDs ➞ PRO IDs</li> <li>Protein Ensembl IDs ➞ UniProt Protein Accession</li> <li>STRING IDs ➞ PRO IDs</li> <li>UniProt Protein Accession ➞ Entrez Gene IDs</li> </ul> <p><strong>Files</strong></p> <ul> <li> <p>Downloaded Data</p> <ul> <li>Cofactor and Catalyst relations:&nbsp;<a href="https://www.uniprot.org/uniprot/?query=&amp;fil=organism%3A%22Homo%20sapiens%20(Human)%20%5B9606%5D%22&amp;columns=id%2Centry%20name%2Creviewed%2Cdatabase(PRO)%2Cchebi(Cofactor)%2Cchebi(Catalytic%20activity)"><code>Cofactor/Catalyst Query Results</code></a></li> <li>UniProt Identifier Mapping:&nbsp;<a href="https://www.uniprot.org/uniprot/?query=&amp;fil=organism%3A%22Homo%20sapiens%20(Human)%20%5B9606%5D%22&amp;columns=id%2Cdatabase(GeneID)%2Cdatabase(Ensembl)%2Cdatabase(HGNC)%2Cgenes(PREFERRED)%2Cgenes(ALTERNATIVE)"><code>UniProt Identifier Query Results</code></a></li> </ul> </li> <li> <p>Generated Data</p> <ul> <li>Merged Gene, RNA, Protein Map:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/Merged_gene_rna_protein_identifiers.pkl"><code>Merged_gene_rna_protein_identifiers.pkl</code></a></li> <li>Protein-Cofactor Relations:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/UNIPROT_PROTEIN_COFACTOR.txt"><code>UNIPROT_PROTEIN_COFACTOR.txt</code></a></li> <li>Protein-Catalyst Relations:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/UNIPROT_PROTEIN_CATALYST.txt"><code>UNIPROT_PROTEIN_CATALYST.txt</code></a></li> <li>UniProt Accession-PRO Identifier Mapping:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/UNIPROT_ACCESSION_PRO_ONTOLOGY_MAP.txt"><code>UNIPROT_ACCESSION_PRO_ONTOLOGY_MAP.txt</code></a></li> <li>UniProt Accession-Entrez Gene Identifier Mapping:&nbsp;<a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/UNIPROT_ACCESSION_ENTREZ_GENE_MAP.txt"><code>UNIPROT_ACCESSION_ENTREZ_GENE_MAP.txt</code></a></li> </ul> </li> </ul> <p>&nbsp;</p> <p>This project is licensed under Apache License 2.0 - see the&nbsp;<strong><a href="https://github.com/callahantiff/PheKnowLator/blob/master/LICENSE"><code>LICENSE.md</code></a></strong>&nbsp;file for details. If you intend to use any of the information on this Wiki, please provide the appropriate attribution by citing this repository:</p> <pre><code>@misc{callahan_tj_2019_3401437, author = {Callahan, TJ}, title = {PheKnowLator}, month = mar, year = 2019, doi = {10.5281/zenodo.3401437}, url = {https://doi.org/10.5281/zenodo.3401437} }</code></pre>

opencc-by-4.0Apr 2021View details →
zenodo44/100

Transaction Graph Dataset for the Bitcoin Blockchain - Part 2 of 4

<p>This dataset contains bitcoin transfer transactions extracted from the&nbsp;Bitcoin Mainnet blockchain.</p> <p>Part1 is available at <a href="https://zenodo.org/deposit/7157356">https://zenodo.org/deposit/7157356</a><br> Part3 is available at <a href="https://zenodo.org/deposit/7158133">https://zenodo.org/deposit/7158133</a><br> Part4 is available at <a href="https://zenodo.org/deposit/7158328">https://zenodo.org/deposit/7158328</a></p> <p>Details of the datasets are given below:</p> <p><strong>FILENAME FORMAT:</strong></p> <p>The filenames have the following format:</p> <p>btc-tx-&lt;start blockno&gt;-&lt;end blockno&gt;-&lt;part&gt;.bz2&nbsp;</p> <p>where &lt;start blockno&gt; is the starting block number, &lt;end blockno&gt; final block number, and &lt;part&gt; is the split part of the file.&nbsp;</p> <p>For example file btc-tx-100000-149999-aa.bz2 &nbsp;and the rest of the parts if any contain transactions from&nbsp;</p> <p>block 100000 to block 149999&nbsp;&nbsp;inclusive.&nbsp;</p> <p>The files are compressed with bzip2. They can be uncompressed using command bunzip2.</p> <p>&nbsp;</p> <p><strong>TRANSACTION FORMAT:</strong></p> <p>Each line in a file corresponds to a transaction. The transaction has the following&nbsp;format:</p> <p>&lt;SYMBOL&gt; &lt;blockno&gt; &lt;txno&gt; &lt;from&gt; &lt;to&gt; &lt;value&gt;&nbsp;</p> <p><br> &lt;SYMBOL&gt;&nbsp;&nbsp;Type of transaction (i.e. BTC-IN or BTC-OUT).</p> <p>&lt;blockno&gt; &nbsp;Number of the block which contains the transaction.&nbsp;</p> <p>&lt;txno&gt;&nbsp;&nbsp;Position of the transaction in the block (i.e. transaction number in the block).</p> <p>&lt;from&gt; &nbsp;Source bitcoin address/transaction of the transfer.</p> <p>&lt;to&gt;&nbsp;&nbsp;Destination bitcoin address/transaction of the transfer.</p> <p>&lt;value&gt; &nbsp;Amount of transfer.</p> <p>&nbsp;</p> <p><strong>BLOCK TIME FORMAT:</strong></p> <p>The block time file has the following&nbsp;format:</p> <p>&lt;block no&gt; &lt;timestamp&gt;</p> <p><br> &lt;block no&gt; &nbsp;Number of the block.&nbsp;</p> <p>&lt;timestamp&gt; &nbsp;Unix timestamp at which the block is mined as a hexadecimal number.</p> <p>&nbsp;</p> <p><strong>IMPORTANT NOTE:</strong></p> <p>Public Bitcoin Mainnet blockchain data is open and can be obtained by connecting as a node on the blockchain or by using the block explorer web sites such as <a href="https://btcscan.org">https://btcscan.org</a>&nbsp;. The downloaders and users of this dataset accept&nbsp;the full responsibility of using the data in GDPR compliant manner or any other regulations. We provide the data as is and we cannot be held responsible for anything.</p> <p>&nbsp;</p> <p><strong>NOTE:</strong></p> <p>If you use this dataset, please do not forget to add the DOI number to the citation.</p> <p>If you use our dataset in your research, please also cite our paper: <a href="https://link.springer.com/chapter/10.1007/978-3-030-94590-9_14">https://link.springer.com/chapter/10.1007/978-3-030-94590-9_14</a></p> <pre><code>@incollection{kilicc2022analyzing, title={Analyzing Large-Scale Blockchain Transaction Graphs for Fraudulent Activities}, author={K{\i}l{\i}{\c{c}}, Baran and {\"O}zturan, Can and {\c{S}}en, Alper}, booktitle={Big Data and Artificial Intelligence in Digital Finance}, pages={253--267}, year={2022}, publisher={Springer, Cham} }</code></pre> <p>&nbsp;</p>

opencc-by-4.0Oct 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record