Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,481
datasets available to search
ShareScore release 0.9.0
Dataset results
1,481 results for “processed data”
PheKnowLator Human Disease Knowledge Graphs - Build Data (Processed)
<p><strong>RELEASE V2.1.0 KNOWLEDGE GRAPH: PROCESSED DATA SOURCES </strong></p> <p><strong>Release:</strong> <a href="https://github.com/callahantiff/PheKnowLator/wiki/v2.0.0">v2.1.0 </a></p> <p>The goal of this build was to create a knowledge graph that represented human disease mechanisms and included the central dogma. The data sources utilized in this release include many of the sources used in the initial release, as well as some new data made available by the <a href="https://ctdbase.org/">Comparative Toxicogenomics Database</a> and experimental data from the <a href="https://www.proteinatlas.org/">Human Protein Atlas</a>.</p> <p>Data sources are listed by type (Ontology and Data not represented in an ontology [Database Sources]). Additional details are provided for each data source below. Please see documentation on the primary release (<a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources">https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources</a>) for additional details on each data source as well as citation information.</p> <p><strong>Data Access:</strong></p> <ul> <li><a href="https://console.cloud.google.com/storage/browser/pheknowlator/archived_builds/release_v2.1.0/build_01MAY2021?project=pheknowlator">https://console.cloud.google.com/storage/browser/pheknowlator/archived_builds/release_v2.1.0/build_01MAY2021</a></li> </ul> <p> </p> <p><strong>ONTOLOGIES</strong></p> <ul> <li>Cell Ontology</li> <li>Cell Line Ontology</li> <li>Chemical Entities of Biological Interest (ChEBI) Ontology</li> <li>Gene Ontology</li> <li>Human Phenotype Ontology</li> <li>Mondo Disease Ontology</li> <li>Pathway Ontology</li> <li>Protein Ontology</li> <li>Relations Ontology</li> <li>Sequence Ontology</li> <li>Uber-Anatomy Ontology</li> <li>Vaccine Ontology</li> </ul> <p> </p> <p><strong>Cell Ontology (CL)</strong></p> <p><strong>Homepage:</strong> <strong><a href="https://github.com/obophenotype/cell-ontology"><code>GitHub</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Bard J, Rhee SY, Ashburner M. <a href="https://genomebiology.biomedcentral.com/articles/10.1186/gb-2005-6-2-r21">An ontology for cell types</a>. Genome Biology. 2005;6(2):R21</p> </blockquote> <p><strong>Usage:</strong> Utilized to connect <code>transcripts</code> and <code>proteins</code> to <code>cells</code>. Additionally, the edges between this ontology and its dependencies are utilized:</p> <ul> <li><strong><a href="https://www.ebi.ac.uk/chebi/"><code>ChEBI</code></a></strong></li> <li><strong><a href="http://geneontology.org/"><code>GO</code></a></strong></li> <li><strong><a href="https://github.com/pato-ontology/pato/"><code>PATO</code></a></strong></li> <li><strong><a href="https://proconsortium.org/"><code>PRO</code></a></strong></li> <li><strong><a href="https://github.com/oborel/obo-relations/"><code>RO</code></a></strong></li> <li><strong><a href="https://uberon.github.io/"><code>UBERON</code></a></strong></li> </ul> <p> </p> <p><strong>Cell Line Ontology (CLO)</strong></p> <p><strong>Homepage:</strong> <strong><a href="http://www.clo-ontology.org/"><code>http://www.clo-ontology.org/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Sarntivijai S, Lin Y, Xiang Z, Meehan TF, Diehl AD, Vempati UD, Schürer SC, Pang C, Malone J, Parkinson H, Liu Y. <a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4387853/">CLO: the cell line ontology</a>. Journal of Biomedical Semantics. 2014;5(1):37</p> </blockquote> <p><strong>Usage:</strong> Utilized this ontology to map <code>cell lines</code> to <code>transcripts</code> and <code>proteins</code>. Additionally, the edges between this ontology and its dependencies are utilized:</p> <ul> <li><strong><a href="http://www.obofoundry.org/ontology/cl.html"><code>CL</code></a></strong></li> <li><strong><a href="http://disease-ontology.org/"><code>DOID</code></a></strong></li> <li><strong><a href="https://github.com/obophenotype/ncbitaxon"><code>NCBITaxon</code></a></strong></li> <li><strong><a href="https://uberon.github.io/"><code>UBERON</code></a></strong></li> </ul> <p> </p> <p><strong>Chemical Entities of Biological Interest (ChEBI)</strong></p> <p><strong>Homepage:</strong> <strong><a href="https://www.ebi.ac.uk/chebi/"><code>https://www.ebi.ac.uk/chebi/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Hastings J, Owen G, Dekker A, Ennis M, Kale N, Muthukrishnan V, Turner S, Swainston N, Mendes P, Steinbeck C. <a href="https://academic.oup.com/nar/article-abstract/44/D1/D1214/2502583">ChEBI in 2016: Improved services and an expanding collection of metabolites</a>. Nucleic Acids Research. 2015;44(D1):D1214-9</p> </blockquote> <p><strong>Usage:</strong> Utilized to connect <code>chemicals</code> to <code>complexes</code>, <code>diseases</code>, <code>genes</code>, <code>GO biological processes</code>, <code>GO cellular components</code>, <code>GO molecular functions</code>, <code>pathways</code>, <code>phenotypes</code>, <code>reactions</code>, and <code>transcripts</code>.</p> <p> </p> <p><strong>Gene Ontology (GO)</strong></p> <p><strong>Homepage:</strong> <strong><a href="http://geneontology.org/"><code>http://geneontology.org/</code></a></strong><br> <strong>Citations:</strong></p> <blockquote> <p>Ashburner M, Ball CA, Blake JA, Botstein D, Butler H, Cherry JM, Davis AP, Dolinski K, Dwight SS, Eppig JT, Harris MA. <a href="https://www.nature.com/articles/ng0500_25">Gene ontology: tool for the unification of biology</a>. Nature Genetics. 2000;25(1):25</p> <p>The Gene Ontology Consortium. <a href="https://academic.oup.com/nar/article/47/D1/D330/5160994">The Gene Ontology Resource: 20 years and still GOing strong</a>. Nucleic Acids Research. 2018;47(D1):D330-8</p> </blockquote> <p><strong>Usage:</strong> Utilized to connect <code>biological processes</code>, <code>cellular components</code>, and <code>molecular functions</code> to <code>chemicals</code>, <code>pathways</code>, and <code>proteins</code>. Additionally, the edges between this ontology and its dependencies are utilized:</p> <ul> <li><strong><a href="http://www.obofoundry.org/ontology/cl.html"><code>CL</code></a></strong></li> <li><strong><a href="https://github.com/obophenotype/ncbitaxon"><code>NCBITaxon</code></a></strong></li> <li><strong><a href="https://github.com/oborel/obo-relations/"><code>RO</code></a></strong></li> <li><strong><a href="https://uberon.github.io/"><code>UBERON</code></a></strong></li> </ul> <p><strong>Other Gene Ontology Data Used:</strong> <a href="http://geneontology.org/gene-associations/goa_human.gaf.gz"><code>goa_human.gaf.gz</code></a></p> <p> </p> <p><strong>Human Phenotype Ontology (HPO)</strong></p> <p><strong>Homepage:</strong> <strong><a href="https://hpo.jax.org/"><code>https://hpo.jax.org/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Köhler S, Carmody L, Vasilevsky N, Jacobsen JO, Danis D, Gourdine JP, Gargano M, Harris NL, Matentzoglu N, McMurry JA, Osumi-Sutherland D. <a href="https://academic.oup.com/nar/article-abstract/47/D1/D1018/5198478">Expansion of the Human Phenotype Ontology (HPO) knowledge base and resources</a>. Nucleic Acids Research. 2018;47(D1):D1018-27</p> </blockquote> <p><strong>Usage:</strong> Utilized to connect <code>phenotypes</code> to <code>chemicals</code>, <code>diseases</code>, <code>genes</code>, and <code>variants</code>. Additionally, the edges between this ontology and its dependencies are utilized:</p> <ul> <li><strong><a href="http://www.obofoundry.org/ontology/cl.html"><code>CL</code></a></strong></li> <li><strong><a href="https://www.ebi.ac.uk/chebi/"><code>ChEBI</code></a></strong></li> <li><strong><a href="http://geneontology.org/"><code>GO</code></a></strong></li> <li><strong><a href="https://uberon.github.io/"><code>UBERON</code></a></strong></li> </ul> <p><strong>Files</strong></p> <ul> <li>Other Human Phenotype Ontology Data Used: <a href="http://purl.obolibrary.org/obo/hp/hpoa/phenotype.hpoa"><code>phenotype.hpoa</code></a></li> </ul> <p> </p> <p><strong>Mondo Disease Ontology (Mondo)</strong></p> <p><strong>Homepage:</strong> <strong><a href="https://mondo.monarchinitiative.org/"><code>https://mondo.monarchinitiative.org/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Mungall CJ, McMurry JA, Köhler S, Balhoff JP, Borromeo C, Brush M, Carbon S, Conlin T, Dunn N, Engelstad M, Foster E. <a href="https://academic.oup.com/nar/article-abstract/45/D1/D712/2605791">The Monarch Initiative: an integrative data and analytic platform connecting phenotypes to genotypes across species</a>. Nucleic Acids Research. 2017;45(D1):D712-22</p> </blockquote> <p><strong>Usage:</strong> Utilized to connect <code>diseases</code> to <code>chemicals</code>, <code>phenotypes</code>, <code>genes</code>, and <code>variants</code>. Additionally, the edges between this ontology and its dependencies are utilized:</p> <ul> <li><strong><a href="http://www.obofoundry.org/ontology/cl.html"><code>CL</code></a></strong></li> <li><strong><a href="https://www.ncbi.nlm.nih.gov/taxonomy"><code>NCBITaxon</code></a></strong></li> <li><strong><a href="http://geneontology.org/"><code>GO</code></a></strong></li> <li><strong><a href="https://hpo.jax.org/"><code>HPO</code></a></strong></li> <li><strong><a href="https://uberon.github.io/"><code>UBERON</code></a></strong></li> </ul> <p> </p> <p><strong>Pathway Ontology (PW)</strong></p> <p><strong>Homepage:</strong> <strong><a href="https://rgd.mcw.edu/wg/home/pathway2/"><code>rgd.mcw.edu</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Petri V, Jayaraman P, Tutaj M, Hayman GT, Smith JR, De Pons J, Laulederkind SJ, Lowry TF, Nigam R, Wang SJ, Shimoyama M. <a href="https://www.ncbi.nlm.nih.gov/pubmed/24499703">The pathway ontology–updates and applications</a>. Journal of Biomedical Semantics. 2014;5(1):7.</p> </blockquote> <p><strong>Usage:</strong> Utilized to connect <code>pathways</code> to <code>GO biological processes</code>, <code>GO cellular components</code>, <code>GO molecular functions</code>, <code>Reactome pathways</code>. Several steps are taken in order to connect <code>Pathway Ontology</code> identifiers to <code>Reactome</code> pathways and <code>GO biological processes</code>. To connect <code>Pathway Ontology</code> identifiers to <code>Reactome</code> pathways, we use <a href="https://github.com/ComPath/resources/tree/master/mappings">ComPath Pathway Database Mappings</a> developed by Daniel Domingo-Fernández (<a href="https://www.ncbi.nlm.nih.gov/pubmed/30564458">PMID:30564458</a>).</p> <p><strong>Files</strong></p> <ul> <li>Downloaded Mapping Data <ul> <li><a href="http://compath.scai.fraunhofer.de/export_mappings"><code>curated_mappings.txt</code></a></li> <li><a href="https://github.com/ComPath/resources/blob/master/mappings/kegg_reactome.csv"><code>kegg_reactome.csv</code></a></li> </ul> </li> <li>Generated Mapping Data <ul> <li><a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/REACTOME_PW_GO_MAPPINGS.txt"><code>REACTOME_PW_GO_MAPPINGS.txt</code></a></li> </ul> </li> </ul> <p> </p> <p><strong>Protein Ontology (PRO)</strong></p> <p><strong>Homepage:</strong> <strong><a href="https://proconsortium.org/"><code>https://proconsortium.org/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Natale DA, Arighi CN, Barker WC, Blake JA, Bult CJ, Caudy M, Drabkin HJ, D’Eustachio P, Evsikov AV, Huang H, Nchoutmboube J. <a href="https://academic.oup.com/nar/article-abstract/39/suppl_1/D539/2508558">The Protein Ontology: a structured representation of protein forms and complexes</a>. Nucleic Acids Research. 2010;39(suppl_1):D539-45</p> </blockquote> <p><strong>Usage:</strong> Utilized to connect <code>proteins</code> to <code>chemicals</code>, <code>genes</code>, <code>anatomy</code>, <code>catalysts</code>, <code>cell lines</code>, <code>cofactors</code>, <code>complexes</code>, <code>GO biological processes</code>, <code>GO cellular components</code>, <code>GO molecular functions</code>, <code>pathways</code>, <code>proteins</code>, <code>reactions</code>, and <code>transcripts</code>. Additionally, the edges between this ontology and its dependencies are utilized:</p> <ul> <li><strong><a href="https://www.ebi.ac.uk/chebi/"><code>ChEBI</code></a></strong></li> <li><strong><a href="http://disease-ontology.org/"><code>DOID</code></a></strong></li> <li><strong><a href="http://geneontology.org/"><code>GO</code></a></strong></li> </ul> <p><strong>Notes:</strong> A partial, human-only version of this ontology was used. Details on how this version of the ontology was generated can be found under the Protein Ontology section of the <a href="https://github.com/callahantiff/PheKnowLator/blob/master/notebooks/Data_Preparation.ipynb"><code>Data_Preparation.ipynb</code></a> Jupyter Notebook.</p> <p><strong>Files</strong></p> <ul> <li> <p>Generated Human Version Protein Ontology (PRO)</p> <ul> <li><a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/human_pro.owl"><code>human_pro.owl</code></a> (closed with <a href="http://www.hermit-reasoner.com/">hermit reasoner</a>)</li> </ul> </li> <li> <p>Other PRO Data Used: <a href="https://proconsortium.org/download/current/promapping.txt"><code>promapping.txt</code></a></p> </li> <li> <p>Generated Mapping Data</p> <ul> <li>Merged Gene, RNA, Protein Map: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/Merged_gene_rna_protein_identifiers.pkl"><code>Merged_gene_rna_protein_identifiers.pkl</code></a></li> <li>Ensembl Transcript-PRO Identifier Mapping: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/ENTREZ_GENE_ENSEMBL_TRANSCRIPT_MAP.txt"><code>ENSEMBL_TRANSCRIPT_PROTEIN_ONTOLOGY_MAP.txt</code></a></li> <li>Entrez Gene-PRO Identifier Mapping: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/ENTREZ_GENE_PRO_ONTOLOGY_MAP.txt"><code>ENTREZ_GENE_PRO_ONTOLOGY_MAP.txt</code></a></li> <li>UniProt Accession-PRO Identifier Mapping: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/UNIPROT_ACCESSION_PRO_ONTOLOGY_MAP.txt"><code>UNIPROT_ACCESSION_PRO_ONTOLOGY_MAP.txt</code></a></li> <li>STRING-PRO Identifier Mapping: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/STRING_PRO_ONTOLOGY_MAP.txt"><code>STRING_PRO_ONTOLOGY_MAP.txt</code></a></li> </ul> </li> </ul> <p> </p> <p><strong>Relations Ontology (RO)</strong></p> <p><strong>Homepage:</strong> <strong><a href="https://github.com/oborel/obo-relations/"><code>GitHub</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Smith B, Ceusters W, Klagges B, Köhler J, Kumar A, Lomax J, Mungall C, Neuhaus F, Rector AL, Rosse C. <a href="https://genomebiology.biomedcentral.com/articles/10.1186/gb-2005-6-5-r46">Relations in biomedical ontologies</a>. Genome Biology. 2005;6(5):R46.</p> </blockquote> <p><strong>Usage:</strong> Utilizing this ontology to connect all data sources in knowledge graph. Additionally, the ontology is queried prior to building the knowledge graph to identify all relations, their inverse properties, and their labels.</p> <p><strong>Files</strong></p> <ul> <li>Generated RO Data <ul> <li><a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/INVERSE_RELATIONS.txt"><code>INVERSE_RELATIONS.txt</code></a></li> <li><a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/RELATIONS_LABELS.txt"><code>RELATIONS_LABELS.txt</code></a></li> </ul> </li> </ul> <p> </p> <p><strong>Sequence Ontology (SO)</strong></p> <p><strong>Homepage:</strong> <strong><a href="https://github.com/The-Sequence-Ontology/SO-Ontologies"><code>GitHub</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Eilbeck K, Lewis SE, Mungall CJ, Yandell M, Stein L, Durbin R, Ashburner M. <a href="https://link.springer.com/article/10.1186/gb-2005-6-5-r44">The Sequence Ontology: a tool for the unification of genome annotations</a>. Genome Biology. 2005;6(5):R44</p> </blockquote> <p><strong>Usage:</strong> Utilized to connect <code>transcripts</code> and other genomic material like <code>genes</code> and <code>variants</code>.</p> <p><strong>Files</strong></p> <ul> <li>Generated Mapping Data <ul> <li><a href="https://storage.googleapis.com/pheknowlator/curated_data/genomic_sequence_ontology_mappings.xlsx"><code>genomic_sequence_ontology_mappings.xlsx</code></a></li> <li><a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/SO_GENE_TRANSCRIPT_VARIANT_TYPE_MAPPING.txt"><code>SO_GENE_TRANSCRIPT_VARIANT_TYPE_MAPPING.txt</code></a></li> </ul> </li> </ul> <p> </p> <p><strong>Uber-Anatomy Ontology (Uberon)</strong></p> <p><strong>Homepage:</strong> <strong><a href="https://uberon.github.io/"><code>GitHub</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Mungall CJ, Torniai C, Gkoutos GV, Lewis SE, Haendel MA. <a href="https://genomebiology.biomedcentral.com/articles/10.1186/gb-2012-13-1-r5">Uberon, an integrative multi-species anatomy ontology</a>. Genome Biology. 2012;13(1):R5</p> </blockquote> <p><strong>Usage:</strong> Utilized to connect <code>tissues</code>, <code>fluids</code>, and <code>cells</code> to <code>proteins</code> and <code>transcripts</code>. Additionally, the edges between this ontology and its dependencies are utilized:</p> <ul> <li><strong><a href="https://www.ebi.ac.uk/chebi/"><code>ChEBI</code></a></strong></li> <li><strong><a href="http://www.obofoundry.org/ontology/cl.html"><code>CL</code></a></strong></li> <li><strong><a href="http://geneontology.org/"><code>GO</code></a></strong></li> <li><strong><a href="https://proconsortium.org/"><code>PRO</code></a></strong></li> </ul> <p> </p> <p><strong>Vaccine Ontology (VO)</strong></p> <p><strong>Homepage:</strong> <strong><a href="http://www.violinet.org/vaccineontology/"><code>http://www.violinet.org/vaccineontology/</code></a></strong><br> <strong>Citations:</strong></p> <blockquote> <p>He Y, Racz R, Sayers S, Lin Y, Todd T, Hur J, Li X, Patel M, Zhao B, Chung M, Ostrow J. <a href="https://academic.oup.com/nar/article-abstract/42/D1/D1124/1053128">Updates on the web-based VIOLIN vaccine database and analysis system</a>. Nucleic Acids Research. 2013;42(D1):D1124-32</p> <p>Xiang Z, Todd T, Ku KP, Kovacic BL, Larson CB, Chen F, Hodges AP, Tian Y, Olenzek EA, Zhao B, Colby LA. <a href="https://academic.oup.com/nar/article-abstract/36/suppl_1/D923/2505793">VIOLIN: vaccine investigation and online information network</a>. Nucleic Acids Research. 2007;36(suppl_1):D923-8</p> </blockquote> <p><strong>Usage:</strong> Utilized the edges between this ontology and its dependencies:</p> <ul> <li><strong><a href="https://www.ebi.ac.uk/chebi/"><code>ChEBI</code></a></strong></li> <li><strong><a href="http://disease-ontology.org/"><code>DOID</code></a></strong></li> <li><strong><a href="http://geneontology.org/"><code>GO</code></a></strong></li> <li><strong><a href="https://proconsortium.org/"><code>PRO</code></a></strong></li> <li><strong><a href="https://uberon.github.io/"><code>UBERON</code></a></strong></li> </ul> <p> </p> <p><strong>DATABASE SOURCES</strong></p> <ul> <li>BioPortal</li> <li>ClinVar</li> <li>Comparative Toxicogenomics Database</li> <li>DisGeNET</li> <li>Ensembl</li> <li>GeneMANIA</li> <li>Genotype-Tissue Expression Project</li> <li>Human Genome Organisation Gene Nomenclature Committee</li> <li>Human Protein Atlas</li> <li>National Center for Biotechnology Information Gene</li> <li>Reactome Pathway Database</li> <li>Search Tool for Recurring Instances of Neighbouring Genes Database</li> <li>Universal Protein Resource Knowledgebase</li> </ul> <p> </p> <p><strong>BioPortal</strong></p> <p><strong>Homepage:</strong> <strong><a href="https://bioportal.bioontology.org/"><code>BioPortal</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>BioPortal. <a href="https://www.bioontology.org/wiki/LOOM">Lexical OWL Ontology Matcher (LOOM)</a></p> <p>Ghazvinian A, Noy NF, Musen MA. <a href="https://www.ncbi.nlm.nih.gov/pmc/articles/pmc2815474/">Creating mappings for ontologies in biomedicine: simple methods work</a>. In AMIA Annual Symposium Proceedings 2009 (Vol. 2009, p. 198). American Medical Informatics Association</p> </blockquote> <p><strong>Usage:</strong> BioPortal was utilized to obtain mappings between <code>MeSH identifiers</code> and <code>ChEBI identifiers</code> for <code>chemicals-diseases</code>, <code>chemicals-genes</code>, <code>chemical-GO biological processes</code>, <code>chemicals-GO cellular components</code>, <code>chemicals-GO molecular functions</code>, <code>chemicals-phenotypes</code>, <code>chemicals-proteins</code>, and <code>chemicals-transcripts</code>. Additional information on how this data was processed can be obtained from the <a href="https://gist.github.com/callahantiff/a28fb3160782f42f104e9ec41553af0d"><code>NCBO_rest_api.py</code></a> GitHub Gist script.</p> <p>⭐ <strong>ALTERNATIVE METHOD</strong>⭐ Since the above approach can take over two days to process, we have developed an alternative solution that downloads the <a><code>mesh2021.nt</code></a> data file directly from MeSH and the <a><code>Flat_file_tab_delimited/names.tsv.gz</code></a> file directly from ChEBI. Using these files, we have recapitulated the <a href="https://www.bioontology.org/wiki/BioPortal_Mappings"><code>LOOM</code></a> algorithm implemented by BioPortal when creating mappings between these resources. The procedure is relatively straightforward and utilizes the following information from each resource:</p> <ul> <li>For all MeSH <code>SCR Chemicals</code>, obtain the following information: <ul> <li>Identifiers: MeSH identifiers</li> <li>Labels: string labels using the <code>RDFS:label</code> object property</li> <li>Synonyms: track down all synonyms using the <code>vocab:concept</code> and <code>vocab:preferredConcept</code> object properties</li> </ul> </li> <li>For all ChEBI classes, obtain the following information: <ul> <li>Labels: string labels using the <code>RDFS:label</code> object property</li> <li>Synonyms: track down all synonyms using all <code>synonym</code> object properties</li> </ul> </li> </ul> <p><strong>Files</strong></p> <ul> <li>Generated Data: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/MESH_CHEBI_MAP.txt"><code>MESH_CHEBI_MAP.txt</code></a></li> </ul> <p> </p> <p><strong>ClinVar</strong></p> <p><strong>Homepage:</strong> <strong><a href="https://www.ncbi.nlm.nih.gov/clinvar/"><code>https://www.ncbi.nlm.nih.gov/clinvar/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Landrum MJ, Lee JM, Benson M, Brown GR, Chao C, Chitipiralla S, Gu B, Hart J, Hoffman D, Jang W, Karapetyan K. <a href="https://academic.oup.com/nar/article-abstract/46/D1/D1062/4641904">ClinVar: improving access to variant interpretations and supporting evidence</a>. Nucleic Acids Research. 2017;46(D1):D1062-7</p> </blockquote> <p><strong>Usage:</strong> ClinVar was utilized to create <code>variant-gene</code>, <code>variant-disease</code>, and <code>variant-phenotype</code> edges. The original data is filtered such that only records meeting the following criteria were included:</p> <ul> <li> <p><code>Assembly</code> = "GRCh38"</p> </li> <li> <p><code>ClinSigSimple</code> = <code>1</code></p> <ul> <li> <blockquote> <p>1 = at least one current record submitted with an interpretation of Likely pathogenic or Pathogenic (independent of whether that record includes assertion criteria and evidence)"</p> </blockquote> </li> </ul> </li> <li> <p><code>ReviewStatus</code> in ["criteria provided, multiple submitters, no conflicts", "reviewed by expert panel", "practice guideline"]</p> </li> </ul> <p><strong>Files</strong></p> <ul> <li> <p>Downloaded Data</p> <ul> <li><a href="https://ftp.ncbi.nlm.nih.gov/pub/clinvar/tab_delimited/variant_summary.txt.gz"><code>variant_summary.txt.gz</code></a></li> <li><a href="https://ftp.ncbi.nlm.nih.gov/pub/clinvar/tab_delimited/var_citations.txt"><code>var_citations.txt</code></a></li> <li><a href="https://ftp.ncbi.nlm.nih.gov/pub/clinvar/tab_delimited/allele_gene.txt.gz"><code>allele_gene.txt.gz</code></a></li> </ul> </li> <li> <p>Generated Edge Data: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/CLINVAR_VARIANT_GENE_DISEASE_PHENOTYPE_EDGES.txt"><code>CLINVAR_VARIANT_GENE_DISEASE_PHENOTYPE_EDGES.txt</code></a></p> </li> </ul> <p> </p> <p><strong>Comparative Toxicogenomics Database (CTD)</strong></p> <p><strong>Homepage:</strong> <strong><a href="http://ctdbase.org/"><code>http://ctdbase.org/</code></a></strong><br> <strong>Citations:</strong></p> <blockquote> <p>Curated [chemical–gene interactions|chemical-go interactions|chemical–disease interactions|gene–pathway interactions] data were retrieved from the Comparative Toxicogenomics Database (CTD), MDI Biological Laboratory, Salisbury Cove, Maine, and NC State University, Raleigh, North Carolina. World Wide Web (URL: <a href="http://ctdbase.org/">http://ctdbase.org/</a>)</p> <p>Davis AP, Grondin CJ, Johnson RJ, Sciaky D, McMorran R, Wiegers J, Wiegers TC, Mattingly CJ. <a href="https://academic.oup.com/nar/article-abstract/47/D1/D948/5106145">The comparative toxicogenomics database: update 2019</a>. Nucleic Acids Research. 2018;47(D1):D948-54</p> </blockquote> <p>Usage: Comparative Toxicogenomics Database (CTD) was utilized to create <code>chemical-disease</code>, <code>chemical-gene</code>, <code>chemical-GO biological process</code>, <code>chemical-GO cellular components</code>, <code>chemical-GO molecular functions</code>, <code>chemical-phenotype</code>, <code>chemical-protein</code>, <code>chemical-rna</code>, and <code>gene-pathway</code> edges. The original data is filtered such that only records meeting the following criteria were included:</p> <ul> <li><code>chemical-disease</code>: <code>DirectEvidence</code> != ""</li> <li><code>chemical-gene</code>: <code>Organism</code> == "Homo sapiens", <code>GeneForms</code> == "gene", and affects not in <code>InteractionActions</code></li> <li><code>chemical-GO biological process</code>: <code>PhenotypeName</code> == "Biological Process" and <code>Interaction</code> <= "1.04e-47" (10th percentile)</li> <li><code>chemical-GO cellular components</code>: <code>PhenotypeName</code> == "Cellular Component" and <code>Interaction</code> <= "1.04e-47" (10th percentile)</li> <li><code>chemical-GO molecular functions</code>: <code>PhenotypeName</code> == "Molecular Function" and <code>Interaction</code> <= "1.04e-47" (10th percentile)</li> <li><code>chemical-phenotype</code>: <code>DirectEvidence</code> != ""</li> <li><code>chemical-protein</code>: <code>Organism</code> == "Homo sapiens", <code>GeneForms</code> == "protein", and affects not in <code>InteractionActions</code></li> <li><code>chemical-rna</code>: <code>Organism</code> == "Homo sapiens", <code>GeneForms</code> == "mRNA", and affects and activity not in <code>InteractionActions</code></li> <li><code>gene-pathway edges</code>: <code>PathwayName</code> == R-HSA-</li> </ul> <p><strong>Files</strong></p> <ul> <li>Downloaded Data <ul> <li>Chemical-Gene Relations: <a href="http://ctdbase.org/reports/CTD_chem_gene_ixns.tsv.gz"><code>CTD_chem_gene_ixns.tsv.gz</code></a></li> <li>Chemical-Disease/Phenotype Relations: <a href="http://ctdbase.org/reports/CTD_chemicals_diseases.tsv.gz"><code>CTD_chemicals_diseases.tsv.gz</code></a></li> <li>Chemical-GO Relations: <a href="http://ctdbase.org/reports/CTD_chem_go_enriched.tsv.gz"><code>CTD_chem_go_enriched.tsv.gz</code></a></li> <li>Gene-Pathway Relations: <a href="http://ctdbase.org/reports/CTD_genes_pathways.tsv.gz"><code>CTD_genes_pathways.tsv.gz</code></a></li> </ul> </li> </ul> <p> </p> <p><strong>DisGeNET</strong></p> <p><strong>Homepage:</strong> <strong><a href="https://www.disgenet.org/"><code>https://www.disgenet.org/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Gene-disease association data retrieved from DisGeNET v6.0 (<a href="http://www.disgenet.org/">http://www.disgenet.org/</a>), Integrative Biomedical Informatics Group GRIB/IMIM/UPF. [December, 2019].</p> <p>Piñero J, Ramírez-Anguita JM, Saüch-Pitarch J, Ronzano F, Centeno E, Sanz F, Furlong LI. <a href="https://academic.oup.com/nar/advance-article-abstract/doi/10.1093/nar/gkz1021/5611674">The DisGeNET knowledge platform for disease genomics: 2019 update</a>. Nucleic Acids Research. 2019.</p> </blockquote> <p><strong>Usage:</strong> DisGeNET was utilized to create <code>gene-disease</code>, and <code>gene-phenotype</code> edges. The original data is filtered such that only records meeting the following criteria were included: <code>EI</code> >= "1.0" (90th percentile). Additionally, data from this source was used to create mappings between different types of disease and phenotype identifiers, including:</p> <ul> <li>OMIM, ORPHA, UMLS, ICD ➞ DOID</li> <li>OMIM, ORPHA, UMLS, ICD ➞ HPO</li> </ul> <p><strong>Files</strong></p> <ul> <li> <p>Downloaded Data</p> <ul> <li>Disease/Phenotype-Gene Relations: <a href="https://www.disgenet.org/static/disgenet_ap1/files/downloads/curated_gene_disease_associations.tsv.gz"><code>curated_gene_disease_associations.tsv.gz</code></a></li> <li>Disease Identifier Mapping: <a href="https://www.disgenet.org/static/disgenet_ap1/files/downloads/disease_mappings.tsv.gz"><code>disease_mappings.tsv.gz</code></a></li> </ul> </li> <li> <p>Generated Mapping Data</p> <ul> <li>Disease Identifier Mapping: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/PHENOTYPE_HPO_MAP.txt"><code>PHENOTPYE_HPO_MAP.txt</code></a></li> <li>Phenotype Identifier Mapping: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/DISEASE_MONDO_MAP.txt"><code>DISEASE_DOID_MAP.txt</code></a></li> </ul> </li> </ul> <p> </p> <p><strong>Ensembl</strong></p> <p><strong>Homepage:</strong> <strong><a href="https://uswest.ensembl.org/index.html"><code>https://uswest.ensembl.org/index.html</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Zerbino DR, Achuthan P, Akanni W, Amode MR, Barrell D, Bhai J, Billis K, Cummins C, Gall A, Girón CG, Gil L. <a href="https://academic.oup.com/nar/article/46/D1/D754/4634002">Ensembl 2018</a>. Nucleic Acids Research. 2017;46(D1):D754-61</p> </blockquote> <p><strong>Usage:</strong> Ensembl data was utilized to create mappings between Ensembl genes, transcripts, and proteins with <a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#national-center-for-biotechnology-information-gene">NCBI Gene identifiers</a>, <a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#human-genome-organisation-gene-nomenclature-committee">HUGO gene symbols</a>, <a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#universal-protein-resource-knowledgebase">UniProt Accession identifiers</a>, and <a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#protein-ontology">Protein Ontology identifiers</a> in the knowledge graph (for additional details on the processing of these data, see <a href="https://github.com/callahantiff/PheKnowLator/blob/master/notebooks/Data_Preparation.ipynb"><code>Data_Preparation.ipynb</code></a>):</p> <ul> <li>Ensembl Transcript IDs ➞ PRO IDs</li> <li>Gene Ensembl IDs ➞ Entrez Gene IDs</li> <li>Gene Ensembl IDs ➞ PRO IDs</li> <li>Gene Symbols ➞ Transcript Ensembl IDs</li> <li>Entrez Gene IDs ➞ Transcript Ensembl IDs</li> <li>Entrez Gene IDs ➞ PRO IDs</li> <li>Protein Ensembl IDs ➞ UniProt Protein Accession</li> <li>STRING IDs ➞ PRO IDs</li> <li>UniProt Protein Accession ➞ Entrez Gene IDs</li> </ul> <p><strong>Files</strong></p> <ul> <li> <p>Downloaded Data</p> <ul> <li><a><code>Homo_sapiens.GRCh38.102.gtf</code></a></li> <li><a><code>Homo_sapiens.GRCh38.102.uniprot.tsv.gz</code></a></li> <li><a><code>Homo_sapiens.GRCh38.102.entrez.tsv.gz</code></a></li> </ul> </li> <li> <p>Generated Mapping Data</p> <ul> <li>Cleaned Ensembl Gene Set: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/ensembl_identifier_data_cleaned.txt"><code>ensembl_identifier_data_cleaned.txt</code></a></li> <li>Merged Gene, RNA, Protein Map: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/Merged_gene_rna_protein_identifiers.pkl"><code>Merged_gene_rna_protein_identifiers.pkl</code></a></li> <li>Ensembl Transcript-PRO Identifier Mapping: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/ENSEMBL_TRANSCRIPT_PROTEIN_ONTOLOGY_MAP.txt"><code>ENSEMBL_TRANSCRIPT_PROTEIN_ONTOLOGY_MAP.txt</code></a></li> <li>Gene Symbol-Ensembl Transcript Identifier Mapping: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/GENE_SYMBOL_ENSEMBL_TRANSCRIPT_MAP.txt"><code>GENE_SYMBOL_ENSEMBL_TRANSCRIPT_MAP.txt</code></a></li> <li>Entrez Gene-Ensembl Transcript Identifier Mapping: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/ENTREZ_GENE_ENSEMBL_TRANSCRIPT_MAP.txt"><code>ENTREZ_GENE_ENSEMBL_TRANSCRIPT_MAP.txt</code></a></li> <li>Entrez Gene-PRO Identifier Mapping: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/ENTREZ_GENE_PRO_ONTOLOGY_MAP.txt"><code>ENTREZ_GENE_PRO_ONTOLOGY_MAP.txt</code></a></li> <li>Ensembl Gene-Entrez Gene Identifier Mapping: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/ENSEMBL_GENE_ENTREZ_GENE_MAP.txt"><code>ENSEMBL_GENE_ENTREZ_GENE_MAP.txt</code></a></li> </ul> </li> </ul> <p> </p> <p><strong>GeneMANIA</strong></p> <p><strong>Homepage:</strong> <strong><a href="https://genemania.org/"><code>https://genemania.org/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Warde-Farley D, Donaldson SL, Comes O, Zuberi K, Badrawi R, Chao P, Franz M, Grouios C, Kazi F, Lopes CT, Maitland A. <a href="https://academic.oup.com/nar/article-abstract/38/suppl_2/W214/1126704">The GeneMANIA prediction server: biological network integration for gene prioritization and predicting gene function</a>. Nucleic Acids Research. 2010;38(suppl_2):W214-20</p> </blockquote> <p><strong>Usage:</strong> GeneMANIA was utilized to create <code>gene-gene</code> edges.</p> <p><strong>Files</strong></p> <ul> <li>Downloaded Data: <a href="http://genemania.org/data/current/Homo_sapiens.COMBINED/COMBINED.DEFAULT_NETWORKS.BP_COMBINING.txt"><code>COMBINED.DEFAULT_NETWORKS.BP_COMBINING.txt</code></a></li> </ul> <p> </p> <p><strong>Genotype-Tissue Expression Project (GTEx)</strong></p> <p><strong>Homepage:</strong> <strong><a href="https://gtexportal.org/home/"><code>https://gtexportal.org/home/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Lonsdale J, Thomas J, Salvatore M, Phillips R, Lo E, Shad S, Hasz R, Walters G, Garcia F, Young N, Foster B. <a href="http://www.nature.com/ng/journal/v45/n6/full/ng.2653.html">The genotype-tissue expression (GTEx) project</a>. Nature Genetics. 2013;45(6):580</p> </blockquote> <p><strong>Usage:</strong> The Genotype-Tissue Expression (GTEx) Project was utilized to create edges between <code>protein-cell</code>, <code>protein-anatomy</code>, <code>rna-cell</code> and <code>rna-anatomy</code> entities. The original data were filtered such that only those edges where the median TPM was >=<code>1.0</code> and genes were of any type other than protein-coding were included. It should also be noted that we chose to use the RNASeQC file over the RSEM file as advised by the GTEx website.</p> <blockquote> <p>The RSEM estimates are based on combining isoform-level estimates, which adds uncertainty to the resulting gene-level values (the isoform-level estimates are highly inaccurate in some cases).</p> </blockquote> <p>The file contains <code>54</code> unique tissue and/or cell types. GTEx provides mappings from tissue types to UBERON and EFO. These provided <a href="https://gtexportal.org/home/samplingSitePage">mappings</a> were verified and extended, such that all samples which referenced a cell type were also mapped to the Cell and the Cell Line ontologies. This resulted in a total of <code>56</code> mappings (<code>1.04</code> mappings/concepts).</p> <p><strong>Files</strong></p> <ul> <li> <p>Downloaded Data: <a href="https://storage.googleapis.com/gtex_analysis_v8/rna_seq_data/GTEx_Analysis_2017-06-05_v8_RNASeQCv1.1.9_gene_median_tpm.gct.gz"><code>GTEx_Analysis_2017-06-05_v8_RNASeQCv1.1.9_gene_median_tpm.gct</code></a></p> </li> <li> <p>Mapping Results: <a href="https://storage.googleapis.com/pheknowlator/curated_data/zooma_tissue_cell_mapping_04JAN2020.xlsx"><code>zooma_tissue_cell_mapping_04JAN2020.xlsx</code></a></p> </li> <li> <p>Generated Data<br> The final mapping set was combined with terms from the <a href="https://www.proteinatlas.org/">Human Protein Atlas</a>, see <a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources/t#human-protein-atlas">here</a> for more information.</p> <ul> <li>All HPA tissue and cell type strings: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/HPA_tissues.txt"><code>HPA_tissues.txt</code></a></li> <li>Final Term Mapping: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/HPA_GTEx_TISSUE_CELL_MAP.txt"><code>HPA_GTEx_TISSUE_CELL_MAP.txt</code></a></li> <li>Final RNA, Gene, Protein-Tissues and Cell Types Relations: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/HPA_GTEX_RNA_GENE_PROTEIN_EDGES.txt"><code>HPA_GTEX_RNA_GENE_PROTEIN_EDGES.txt</code></a></li> </ul> </li> </ul> <p> </p> <p><strong>Human Genome Organisation Gene Nomenclature Committee (HUGO)</strong></p> <p><strong>Homepage:</strong> <strong><a href="https://www.genenames.org/"><code>https://www.genenames.org/</code></a></strong><br> <strong>Citations:</strong></p> <blockquote> <p>HGNC Database, HUGO Gene Nomenclature Committee (HGNC), European Molecular Biology Laboratory, European Bioinformatics Institute (EMBL-EBI), Wellcome Genome Campus, Hinxton, Cambridge CB10 1SD, United Kingdom <a href="https://www.genenames.org/">www.genenames.org</a></p> <p>Yates B, Braschi B, Gray K, Seal R, Tweedie S, Bruford E. <a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5210531/">Genenames.org: the HGNC and VGNC Resources in 2017</a>. Nucleic Acids Research. 2017;45(D1):D619-625</p> </blockquote> <p><strong>Usage:</strong> The Human Genome Organisation (HUGO) data was utilized to obtain mappings between <a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#national-center-for-biotechnology-information-gene">NCBI Gene identifiers</a>, HUGO gene symbols, <a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#universal-protein-resource-knowledgebase">UniProt Accession identifiers</a>, and <a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#protein-ontology">Protein Ontology identifiers</a>. For additional details on the processing of these data, see <a href="https://github.com/callahantiff/PheKnowLator/blob/master/notebooks/Data_Preparation.ipynb"><code>Data_Preparation.ipynb</code></a>:</p> <ul> <li>Ensembl Transcript IDs ➞ PRO IDs</li> <li>Gene Ensembl IDs ➞ Entrez Gene IDs</li> <li>Gene Ensembl IDs ➞ PRO IDs</li> <li>Gene Symbols ➞ Transcript Ensembl IDs</li> <li>Entrez Gene IDs ➞ Transcript Ensembl IDs</li> <li>Entrez Gene IDs ➞ PRO IDs</li> <li>Protein Ensembl IDs ➞ UniProt Protein Accession</li> <li>STRING IDs ➞ PRO IDs</li> <li>UniProt Protein Accession ➞ Entrez Gene IDs</li> </ul> <p><strong>Files</strong></p> <ul> <li> <p>Downloaded Data: <a href="http://ftp.ebi.ac.uk/pub/databases/genenames/hgnc/tsv/hgnc_complete_set.txt"><code>hgnc_complete_set.txt</code></a></p> </li> <li> <p>Generated Data</p> <ul> <li>Merged Gene, RNA, Protein Map: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/Merged_gene_rna_protein_identifiers.pkl"><code>Merged_gene_rna_protein_identifiers.pkl</code></a></li> <li>Gene Symbol-Ensembl Transcript Identifier Mapping: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/GENE_SYMBOL_ENSEMBL_TRANSCRIPT_MAP.txt"><code>GENE_SYMBOL_ENSEMBL_TRANSCRIPT_MAP.txt</code></a></li> </ul> </li> </ul> <p> </p> <p><strong>Human Protein Atlas (HPA)</strong></p> <p><strong>Homepage:</strong> <strong><a href="https://www.proteinatlas.org/"><code>https://www.proteinatlas.org/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Uhlén M, Fagerberg L, Hallström BM, Lindskog C, Oksvold P, Mardinoglu A, Sivertsson Å, Kampf C, Sjöstedt E, Asplund A, Olsson I. <a href="https://science.sciencemag.org/content/347/6220/1260419.short">Tissue-based map of the human proteome</a>. Science. 2015;347(6220):1260419</p> </blockquote> <p><strong>Usage:</strong> The Human Protein Atlas (HPA) was utilized to create <code>rna-cell</code>, <code>rna-anatomy</code>, <code>protein-cell</code>, and <code>protein-anatomy</code> edges. Evidence between gene and RNA expression in specific tissue types was derived by HPA, such that the <a href="https://www.proteinatlas.org/about/assays+annotation#normalization_rna">consensus normalized expression</a> was >=<code>1.0</code>. Zooma was utilized to automatically annotate the <code>153</code> unique tissues and cell types from Human Protein Atlas for all human protein-coding genes in the <a href="https://www.proteinatlas.org/humanproteome">Human Proteome</a> to the Cell Ontology, Cell Line Ontology, and the Uber-Anatomy Ontology. To best represent each concept, the automatic mappings from Zooma were extend through manual mapping efforts to ensure each concept cell type was matched to a Cell Ontology, Cell Line Ontology, and UBERON ontology term. This resulted in a total of <code>281</code> mappings (<code>1.84</code> mappings/concepts).</p> <p><strong>Files</strong></p> <ul> <li> <p>Downloaded Data: <a href="https://www.proteinatlas.org/api/search_download.php?search=&columns=g,eg,up,pe,rnatsm,rnaclsm,rnacasm,rnabrsm,rnabcsm,rnablsm,scl,t_RNA_adipose_tissue,t_RNA_adrenal_gland,t_RNA_amygdala,t_RNA_appendix,t_RNA_basal_ganglia,t_RNA_bone_marrow,t_RNA_breast,t_RNA_cerebellum,t_RNA_cerebral_cortex,t_RNA_cervix,_uterine,t_RNA_colon,t_RNA_corpus_callosum,t_RNA_ductus_deferens,t_RNA_duodenum,t_RNA_endometrium_1,t_RNA_epididymis,t_RNA_esophagus,t_RNA_fallopian_tube,t_RNA_gallbladder,t_RNA_heart_muscle,t_RNA_hippocampal_formation,t_RNA_hypothalamus,t_RNA_kidney,t_RNA_liver,t_RNA_lung,t_RNA_lymph_node,t_RNA_midbrain,t_RNA_olfactory_region,t_RNA_ovary,t_RNA_pancreas,t_RNA_parathyroid_gland,t_RNA_pituitary_gland,t_RNA_placenta,t_RNA_pons_and_medulla,t_RNA_prostate,t_RNA_rectum,t_RNA_retina,t_RNA_salivary_gland,t_RNA_seminal_vesicle,t_RNA_skeletal_muscle,t_RNA_skin_1,t_RNA_small_intestine,t_RNA_smooth_muscle,t_RNA_spinal_cord,t_RNA_spleen,t_RNA_stomach_1,t_RNA_testis,t_RNA_thalamus,t_RNA_thymus,t_RNA_thyroid_gland,t_RNA_tongue,t_RNA_tonsil,t_RNA_urinary_bladder,t_RNA_vagina,t_RNA_B-cells,t_RNA_dendritic_cells,t_RNA_granulocytes,t_RNA_monocytes,t_RNA_NK-cells,t_RNA_T-cells,t_RNA_total_PBMC,cell_RNA_A-431,cell_RNA_A549,cell_RNA_AF22,cell_RNA_AN3-CA,cell_RNA_ASC_diff,cell_RNA_ASC_TERT1,cell_RNA_BEWO,cell_RNA_BJ,cell_RNA_BJ_hTERT+,cell_RNA_BJ_hTERT+_SV40_Large_T+,cell_RNA_BJ_hTERT+_SV40_Large_T+_RasG12V,cell_RNA_CACO-2,cell_RNA_CAPAN-2,cell_RNA_Daudi,cell_RNA_EFO-21,cell_RNA_fHDF/TERT166,cell_RNA_HaCaT,cell_RNA_HAP1,cell_RNA_HBEC3-KT,cell_RNA_HBF_TERT88,cell_RNA_HDLM-2,cell_RNA_HEK_293,cell_RNA_HEL,cell_RNA_HeLa,cell_RNA_Hep_G2,cell_RNA_HHSteC,cell_RNA_HL-60,cell_RNA_HMC-1,cell_RNA_HSkMC,cell_RNA_hTCEpi,cell_RNA_hTEC/SVTERT24-B,cell_RNA_hTERT-HME1,cell_RNA_HUVEC_TERT2,cell_RNA_K-562,cell_RNA_Karpas-707,cell_RNA_LHCN-M2,cell_RNA_MCF7,cell_RNA_MOLT-4,cell_RNA_NB-4,cell_RNA_NTERA-2,cell_RNA_PC-3,cell_RNA_REH,cell_RNA_RH-30,cell_RNA_RPMI-8226,cell_RNA_RPTEC_TERT1,cell_RNA_RT4,cell_RNA_SCLC-21H,cell_RNA_SH-SY5Y,cell_RNA_SiHa,cell_RNA_SK-BR-3,cell_RNA_SK-MEL-30,cell_RNA_T-47d,cell_RNA_THP-1,cell_RNA_TIME,cell_RNA_U-138_MG,cell_RNA_U-2_OS,cell_RNA_U-2197,cell_RNA_U-251_MG,cell_RNA_U-266/70,cell_RNA_U-266/84,cell_RNA_U-698,cell_RNA_U-87_MG,cell_RNA_U-937,cell_RNA_WM-115,blood_RNA_basophil,blood_RNA_classical_monocyte,blood_RNA_eosinophil,blood_RNA_gdT-cell,blood_RNA_intermediate_monocyte,blood_RNA_MAIT_T-cell,blood_RNA_memory_B-cell,blood_RNA_memory_CD4_T-cell,blood_RNA_memory_CD8_T-cell,blood_RNA_myeloid_DC,blood_RNA_naive_B-cell,blood_RNA_naive_CD4_T-cell,blood_RNA_naive_CD8_T-cell,blood_RNA_neutrophil,blood_RNA_NK-cell,blood_RNA_non-classical_monocyte,blood_RNA_plasmacytoid_DC,blood_RNA_T-reg,blood_RNA_total_PBMC,brain_RNA_amygdala,brain_RNA_basal_ganglia,brain_RNA_cerebellum,brain_RNA_cerebral_cortex,brain_RNA_hippocampal_formation,brain_RNA_hypothalamus,brain_RNA_midbrain,brain_RNA_olfactory_region,brain_RNA_pons_and_medulla,brain_RNA_thalamus&format=tsv"><code>proteinatlas_search.tsv</code></a></p> </li> <li> <p>Mapping Results: <a href="https://storage.googleapis.com/pheknowlator/curated_data/zooma_tissue_cell_mapping_04JAN2020.xlsx"><code>zooma_tissue_cell_mapping_04JAN2020.xlsx</code></a></p> </li> <li> <p>Generated Data</p> <ul> <li>Final Term Mapping: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/HPA_GTEx_TISSUE_CELL_MAP.txt"><code>HPA_GTEx_TISSUE_CELL_MAP.txt</code></a></li> <li>Final RNA, Gene, Protein-Tissues and Cell Types Relations: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/HPA_GTEX_RNA_GENE_PROTEIN_EDGES.txt"><code>HPA_GTEX_RNA_GENE_PROTEIN_EDGES.txt</code></a></li> </ul> </li> </ul> <p> </p> <p><strong>National Center for Biotechnology Information (NCBI) Entrez Gene</strong></p> <p><strong>Homepage:</strong> <strong><a href="https://www.ncbi.nlm.nih.gov/gene/"><code>https://www.ncbi.nlm.nih.gov/gene/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Maglott D, Ostell J, Pruitt KD, Tatusova T. <a href="https://academic.oup.com/nar/article-abstract/33/suppl_1/D54/2505255">Entrez Gene: gene-centered information at NCBI</a>. Nucleic Acids Research. 2005;33(suppl_1):D54-8.</p> </blockquote> <p><strong>Usage:</strong> The National Center for Biotechnology Information (NCBI) Gene data was utilized to obtain mappings between <a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#ncbi-gene">NCBI Gene identifiers</a>, <a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#hugo-gene-nomenclature-committee">HUGO gene symbols</a>, <a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#uniprot-knowledgebase">UniProt Accession identifiers</a>, and <a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#protein-ontology">Protein Ontology identifiers</a>. For additional details on the processing of these data, see <a href="https://github.com/callahantiff/PheKnowLator/blob/master/notebooks/Data_Preparation.ipynb"><code>Data_Preparation.ipynb</code></a>:</p> <ul> <li>Ensembl Transcript IDs ➞ PRO IDs</li> <li>Gene Ensembl IDs ➞ Entrez Gene IDs</li> <li>Gene Ensembl IDs ➞ PRO IDs</li> <li>Gene Symbols ➞ Transcript Ensembl IDs</li> <li>Entrez Gene IDs ➞ Transcript Ensembl IDs</li> <li>Entrez Gene IDs ➞ PRO IDs</li> <li>Protein Ensembl IDs ➞ UniProt Protein Accession</li> <li>STRING IDs ➞ PRO IDs</li> <li>UniProt Protein Accession ➞ Entrez Gene IDs</li> </ul> <p><strong>Files</strong></p> <ul> <li> <p>Downloaded Data: <a href="https://ftp.ncbi.nih.gov/gene/DATA/GENE_INFO/Mammalia/Homo_sapiens.gene_info.gz"><code>Homo_sapiens.gene_info.gz</code></a></p> </li> <li> <p>Generated Data</p> <ul> <li>Merged Gene, RNA, Protein Map: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/Merged_gene_rna_protein_identifiers.pkl"><code>Merged_gene_rna_protein_identifiers.pkl</code></a></li> <li>Entrez Gene-Ensembl Transcript Identifier Mapping: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/ENTREZ_GENE_ENSEMBL_TRANSCRIPT_MAP.txt"><code>ENTREZ_GENE_ENSEMBL_TRANSCRIPT_MAP.txt</code></a></li> <li>Entrez Gene-PRO Identifier Mapping: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/ENTREZ_GENE_PRO_ONTOLOGY_MAP.txt"><code>ENTREZ_GENE_PRO_ONTOLOGY_MAP.txt</code></a></li> <li>Ensembl Gene-Entrez Gene Identifier Mapping: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/ENSEMBL_GENE_ENTREZ_GENE_MAP.txt"><code>ENSEMBL_GENE_ENTREZ_GENE_MAP.txt</code></a></li> <li>Uniprot Accession-Entrez Gene Identifier Mapping: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/UNIPROT_ACCESSION_ENTREZ_GENE_MAP.txt"><code>UNIPROT_ACCESSION_ENTREZ_GENE_MAP.txt</code></a></li> </ul> </li> </ul> <p> </p> <p><strong>Reactome Pathway Database</strong></p> <p><strong>Homepage:</strong> <strong><a href="https://reactome.org/"><code>https://reactome.org/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Fabregat A, Jupe S, Matthews L, Sidiropoulos K, Gillespie M, Garapati P, Haw R, Jassal B, Korninger F, May B, Milacic M. <a href="https://academic.oup.com/nar/article-abstract/46/D1/D649/4626770">The reactome pathway knowledgebase</a>. Nucleic Acids Research. 2017;46(D1):D649-55</p> </blockquote> <p><strong>Usage:</strong> The Reactome Database was utilized to create <code>chemical-pathway</code>, <code>GO Biological process-pathway</code>, <code>pathway-GO Cellular component</code>, <code>GO Molecular function-pathway</code>, and <code>protein-pathway</code> edges. The original data is filtered such that only records meeting the following criteria were included:</p> <ul> <li><code>chemical-pathway</code>: column[5] == "Homo sapiens"</li> <li><code>GO Biological process-pathway</code>: column[5] startswith "REACTOME", column[8] == "P", and column[12] == "taxon:9606"</li> <li><code>pathway-GO Cellular component</code>: column[5] startswith "REACTOME", column[8] == "C", and column[12] == "taxon:9606"</li> <li><code>GO Molecular function-pathway</code>: column[5] startswith "REACTOME", column[8] == "F", and column[12] == "taxon:9606"</li> <li><code>protein-pathway</code>: column[5] == "Homo sapiens"</li> </ul> <p><strong>Files</strong></p> <ul> <li>Downloaded Data <ul> <li>Chemical-Pathway Relations: <a href="https://reactome.org/download/current/ChEBI2Reactome_All_Levels.txt"><code>ChEBI2Reactome_All_Levels.txt</code></a></li> <li>Pathway-GO Relations: <a href="https://reactome.org/download/current/gene_association.reactome"><code>gene_association.reactome</code></a></li> <li>Protein-Pathway Relations: <a href="https://reactome.org/download/current/UniProt2Reactome_All_Levels.txt"><code>UniProt2Reactome_All_Levels.txt</code></a></li> </ul> </li> </ul> <p> </p> <p><strong>Search Tool for Recurring Instances of Neighbouring Genes (STRING) Database</strong></p> <p><strong>Homepage:</strong> <strong><a href="https://string-db.org/"><code>string-db.org</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Szklarczyk D, Gable AL, Lyon D, Junge A, Wyder S, Huerta-Cepas J, Simonovic M, Doncheva NT, Morris JH, Bork P, Jensen LJ. <a href="https://academic.oup.com/nar/article-abstract/47/D1/D607/5198476">STRING v11: protein–protein association networks with increased coverage, supporting functional discovery in genome-wide experimental datasets</a>. Nucleic Acids Research. 2018;47(D1):D607-13</p> </blockquote> <p><strong>Usage:</strong> The Search Tool for Recurring Instances of Neighbouring Genes (STRING) Database was utilized to create <code>protein-protein</code> edges. The original data is filtered such that only records meeting the following criteria were included: <code>combined_score</code> >= "700" (>90th percentile).</p> <p><strong>Files</strong></p> <ul> <li> <p>Downloaded Data: <a href="https://stringdb-static.org/download/protein.links.v11.0/9606.protein.links.v11.0.txt.gz"><code>9606.protein.links.v11.0.txt.gz</code></a></p> </li> <li> <p>Generated Data: STRING-PRO Identifier Mapping: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/STRING_PRO_ONTOLOGY_MAP.txt"><code>STRING_PRO_ONTOLOGY_MAP.txt</code></a></p> </li> </ul> <p> </p> <p><strong>Universal Protein Resource (UniProt) Knowledgebase</strong></p> <p><strong>Homepage:</strong> <strong><a href="https://www.uniprot.org/"><code>https://www.uniprot.org/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>UniProt Consortium. <a href="https://academic.oup.com/nar/article-abstract/47/D1/D506/5160987">UniProt: a worldwide hub of protein knowledge</a>. Nucleic acids research. 2018;47(D1):D506-15</p> </blockquote> <p><strong>Usage:</strong> The Universal Protein Resource (UniProt) Knowledgebase was utilized to obtain <code>cofactor</code>/<code>catalyst</code>-<code>protein</code> and <code>protein-coding gene</code>-<code>protein</code> edges as well as mappings between <a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#national-center-for-biotechnology-information-gene">NCBI Gene identifiers</a>, <a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#human-genome-organisation-gene-nomenclature-committee">HUGO gene symbols</a>, <a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#universal-protein-resource-knowledgebase">Universal Protein Resource (UniProt) Accession identifiers</a>, and <a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#protein-ontology">Protein Ontology identifiers</a>. For additional details on the processing of these data, see <a href="https://github.com/callahantiff/PheKnowLator/blob/master/notebooks/Data_Preparation.ipynb"><code>Data_Preparation.ipynb</code></a>:</p> <ul> <li>Ensembl Transcript IDs ➞ PRO IDs</li> <li>Gene Ensembl IDs ➞ Entrez Gene IDs</li> <li>Gene Ensembl IDs ➞ PRO IDs</li> <li>Gene Symbols ➞ Transcript Ensembl IDs</li> <li>Entrez Gene IDs ➞ Transcript Ensembl IDs</li> <li>Entrez Gene IDs ➞ PRO IDs</li> <li>Protein Ensembl IDs ➞ UniProt Protein Accession</li> <li>STRING IDs ➞ PRO IDs</li> <li>UniProt Protein Accession ➞ Entrez Gene IDs</li> </ul> <p><strong>Files</strong></p> <ul> <li> <p>Downloaded Data</p> <ul> <li>Cofactor and Catalyst relations: <a href="https://www.uniprot.org/uniprot/?query=&fil=organism%3A%22Homo%20sapiens%20(Human)%20%5B9606%5D%22&columns=id%2Centry%20name%2Creviewed%2Cdatabase(PRO)%2Cchebi(Cofactor)%2Cchebi(Catalytic%20activity)"><code>Cofactor/Catalyst Query Results</code></a></li> <li>UniProt Identifier Mapping: <a href="https://www.uniprot.org/uniprot/?query=&fil=organism%3A%22Homo%20sapiens%20(Human)%20%5B9606%5D%22&columns=id%2Cdatabase(GeneID)%2Cdatabase(Ensembl)%2Cdatabase(HGNC)%2Cgenes(PREFERRED)%2Cgenes(ALTERNATIVE)"><code>UniProt Identifier Query Results</code></a></li> </ul> </li> <li> <p>Generated Data</p> <ul> <li>Merged Gene, RNA, Protein Map: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/Merged_gene_rna_protein_identifiers.pkl"><code>Merged_gene_rna_protein_identifiers.pkl</code></a></li> <li>Protein-Cofactor Relations: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/UNIPROT_PROTEIN_COFACTOR.txt"><code>UNIPROT_PROTEIN_COFACTOR.txt</code></a></li> <li>Protein-Catalyst Relations: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/UNIPROT_PROTEIN_CATALYST.txt"><code>UNIPROT_PROTEIN_CATALYST.txt</code></a></li> <li>UniProt Accession-PRO Identifier Mapping: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/UNIPROT_ACCESSION_PRO_ONTOLOGY_MAP.txt"><code>UNIPROT_ACCESSION_PRO_ONTOLOGY_MAP.txt</code></a></li> <li>UniProt Accession-Entrez Gene Identifier Mapping: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/UNIPROT_ACCESSION_ENTREZ_GENE_MAP.txt"><code>UNIPROT_ACCESSION_ENTREZ_GENE_MAP.txt</code></a></li> </ul> </li> </ul> <p> </p> <p>This project is licensed under Apache License 2.0 - see the <strong><a href="https://github.com/callahantiff/PheKnowLator/blob/master/LICENSE"><code>LICENSE.md</code></a></strong> file for details. If you intend to use any of the information on this Wiki, please provide the appropriate attribution by citing this repository:</p> <pre><code>@misc{callahan_tj_2019_3401437, author = {Callahan, TJ}, title = {PheKnowLator}, month = mar, year = 2019, doi = {10.5281/zenodo.3401437}, url = {https://doi.org/10.5281/zenodo.3401437} }</code></pre>
Raw data for the article entitled "Facile Solution Synthesis, Processing and Characterization of n- and p-Type Binary and Ternary Bi–Sb Tellurides"
<p>Raw data for the plots in the open access article "Facile Solution Synthesis, Processing and Characterization of n- and p-Type Binary and Ternary Bi–Sb Tellurides"</p>
Raw data for the plots in the article entitled "Thermoelectric Inks and Power Factor Tunability in Hybrid Films through All Solution Process"
<p>Raw data for the plots in the article entitled "Thermoelectric Inks and Power Factor Tunability in Hybrid Films through All Solution Process"</p> <p>https://doi.org/10.1021/acsami.1c24392 </p> <p>ACS Appl. Mater. Interfaces 2022, 14, 19295−19303</p>
Pre-processed AMR data on S. aureus isolates from PATRIC database
<p>Pre-processed AMR data on S. aureus isolates from PATRIC database. The majority class size was decreased to reach the class ratio of 1:1 when the susceptible/resistant or resistant/susceptible class ratio exceeded 3.5.</p>
Data for "Three-Dimensional Broadband Interferometric Mapping and Polarization (BIMAP-3D) Observations of Lightning Discharge Processes" by Shao et al.
<p>Data set for manuscript of “Three-Dimensional Broadband Interferometric Mapping and Polarization (BIMAP-3D) Observations of Lightning Discharge Processes” by Shao et al. submitted to Journal of Geophysical Research-atmosphere</p>
The dataset of Hong Kong Housing transaction records (1997-2018, after an incoming data quality processing)
<p>These detailed housing transaction records of over 2 million property rights entries in Hong Kong's property markets over the past 23 years (1997 - 2018).</p>
Ion Implantation Sensor and Process Target Data for Predicting Ion Beam Tuning in Semiconductor Manufacturing
<h2><strong>Dataset Description:</strong></h2> <p>This dataset is designed to predict ion beam tuning setup processes in semiconductor manufacturing, in terms of tuning success or failure, and tuning duration. It is split into <strong><code>X</code></strong> and <code><strong>y</strong></code> to allow for supervised learning approaches.</p> <ul> <li><code><strong>X</strong></code> represents the current equipment condition and the process targets of the currently processed and the upcoming lot, as defined within recipes.</li> <li><code><strong>y</strong></code> represents the ion beam tuning setup report, which informs about the tuning success ratio and tuning duration. These setups are necessary, when switching between recipes to prepare the equipment for processing the next lot. <strong><code>y</code></strong> contains three labels, enabling classification of (1) tuning success or fail, and (2) prolonged tuning, as well as (3) estimation of tuning duration as a regression task.</li> </ul> <p>About <strong><code>X</code></strong>:</p> <p>Each lot is processed with a specific recipe to achieve the process target. The tuning takes place before the first wafer of the to-be-tuned recipe is processed. Each row in <strong><code>X</code></strong> includes logistical information such as the equipment used for processing and parsed recipe / process target information for the current and upcoming lot. The majority of data consists out of aggregated metrics of equipment-internally tracked sensor traces, recording physical parameters such as gas flows, temperatures, voltages and currents. When analyzed in conjunction with the processed recipe, these sensors provide insights into the current equipment condition. </p> <p>About <code><strong>y</strong></code>:</p> <p>The <code>setup_result</code> column indicates the success or failure of tuning - with <code>setup_result=0</code> indicating tuning success, while <code>setup_result=1</code> signals tuning failure. If the first tuning attempt fails, there may be follow-up attempts, but these are not included in this dataset. The <code>duration</code> column represents the tuning duration in seconds, as used for regression analysis. The <code>duration_interval</code> column is a binary label for prolonged tunings, i.e. <code>duration_interval=1</code> for instances, which take more than 6 minutes to tune.</p> <p>For reproducibility of the corresponding paper's results:</p> <ol> <li>The dataset contains the same carefully curated subset of features.</li> <li>The train_test_split() has already been performed, thus we provide <code>x_train</code> and <code>x_valid</code> separately.</li> <li>To reduce the effect of outliers in the data, the sensor data has already been scaled, as derived from <code>x_train</code>.</li> </ol> <p>In summary, these datasets (<code><strong>X</strong></code>, <code><strong>y</strong></code>) provide comprehensive information for predicting ion beam tuning in semiconductor manufacturing, making it a valuable resource for researchers and practitioners in the field.</p> <h2><strong>Python Code for Reproducibility:</strong></h2> <p>Furthermore, we share a jupyter notebook <code>ionbeamtuning.ipynb</code> with Python code to train the best performing model on the provided data, as described in the paper. To execute the code, you may need to install any missing packages specified in the <code>requirements.txt</code>, as indicated within the notebook.</p>
Data for common data models to streamline metabolomics processing and annotation, and implementation in a Python pipeline
<p>This upload contains the HZV029 Plasma and HZV029 Two-Phase dataset for reviewers of the "Data for common data models to streamline metabolomics processing and annotation, and implementation in a Python pipeline" submission. </p> <p>Both datasets will be uploaded to metabolomics workbench and the upload completed before final publication of the manuscript. For the he HZV029 Plasma datasets only the final run is included for any sample (i.e., failed injections or other samples with data quality issues that were reran during acquisition were omitted).</p> <p>Also included in the upload is the source code for the MetDataModel and the pcpfm at the time of manuscript re-submission and the pcpfm itself. If you find this upload in the future, please check out the github repos for more updated versions:</p> <p>https://github.com/shuzhao-li-lab/PythonCentricPipelineForMetabolomics</p> <p>https://github.com/shuzhao-li-lab/metDataModel</p> <p>The github repo does not store the input the data for space reasons, they only have the notebooks. However, the .zip here has both the notebooks by themselves in the notebook subdirectory and a separate directory with the notebooks and the data used to generate all the figures and results in the manuscript.</p> <p><strong>Some information that is needed to rerun this analysis:</strong></p> <p>Sequence files are critical to the functioning of the pipeline. The sequence files for all analyses are provided under sequence_files.zip. These can be used to recapitulate the analysis by eitehr changing the filepath to each acquisition to where you put it on your sytem or by placing the sequence file in the same directory as the mzml or raw. In the latter case, the pipeline will search for filenames matching the sample names. The sequence files also store some sample metadata such as the type of sample a given acquisition is (unknown, pooled, qc, etc...)</p> <p>.raw to .mzML conversion works well on MacOS but may not work well on other systems. You will need to use the ability to specify your own conversion command or convert files outside of the pipeline. </p> <p>To replicate the results, you do need to have the annotation sources downloaded which can be done using the pipeline. MS2 annotation requires the files in the AcquireX directory which is MS2 acquisitions on pooled HZV029 plasma samples.</p> <p>For the comparison between MetaboAnalystR and the pcpfm, subsets of the datasets were used. These subsets and the sequence files are in Subsets_for_performance_testing.zip. The sequences are also in the sequence_files directory as well</p> <p>The notebooks reference data in the analysis folders. Copies of these files are located with the notebooks to ease reproduction of the exact results in the paper; however, to do so, you will need to change paths to this data in the notebook. This lets the notebooks be ran during a rerun without copying intermediates back and forth and it keeps the github repo clean.</p> <p><strong>Version History:</strong></p> <p>This version is after reviewer comments and is for resubmission.</p> <p> </p> <p><strong>Contributions:</strong></p> <p>Joshua M Mitchell implemented the pipeline and was first author on the manuscript. Shuzhao Li is the corresponding author on the manuscript. </p> <p>Maheshwor Thapa performed the experiments to collect the HZV029 data. Yuanye Chi helped with testing and documenting the pipeline. </p> <p>Jiangou (Jeff) Xia and Zhiqiang Pang provided the R portion of the analysis. </p>
Supplementary Data from, "Causal health impacts of power plant emission controls under modeled and uncertain physical process interference."
<p>These data are used to conduct the analysis in, "<a href="https://arxiv.org/abs/2306.05665">Causal health impacts of power plant emission controls under modeled and uncertain physical process interference</a>," by Wikle and Zigler (2024), to appear in <em>Annals of Applied</em> Statistics. This is purely for archival purposes to facilitate access to and replication of the aforementioned analysis. Data were obtained from the following sources:</p> <ol> <li> U.S. Emissions Data [<a href="https://ampd.epa.gov/ampd">U.S. EPA, Air markets program data (AMPD)</a>] <ul> <li>AMPD_Unit_with_Sulfur_Content_and_Regulations_with_Facility_Attributes.csv</li> </ul> </li> <li> US Census 2016 American Community Survey [<a href="https://www.census.gov/programs-surveys/acs">US Census Bureau ACS</a>] <ul> <li>Census_2016_TxZCTA.RDS</li> <li><em>Note: data were obtained using the r package ‘<a href="https://walker-data.com/tidycensus/">tidycensus</a>’.</em></li> </ul> </li> <li> Daymet Annual Climate Summaries [<a href="https://daac.ornl.gov/DAYMET/guides/Daymet_V4_Annual_Climatology.html">Daymet Version 4</a>] <ul> <li>daymet_v4_prcp_annttl_na_2016.nc</li> <li>daymet_v4_tmax_annavg_na_2016.nc</li> <li>daymet_v4_tmin_annavg_na_2016.nc</li> <li>daymet_v4_vp_annavg_na_2016.nc</li> </ul> </li> <li> SO<sub>4</sub> and Black Carbon Concentrations [<a href="https://sites.wustl.edu/acag/datasets/surface-pm2-5/#V4.NA.03">Randall Martin Atmospheric Composition Analysis Group, North American Regional Estimates, version V4.NA.02</a>] <ul> <li>GWRwSPEC_BC_NA_201601_201612.nc</li> <li>GWRwSPEC_SO4_NA_201601_201612.nc</li> </ul> </li> <li> HyADS Coal-Attributed PM2.5 Concentrations [<a href="https://doi.org/10.1097/EDE.0000000000001024">Henneman et al. (2019)</a>] <ul> <li>HyADS_grids_pm25_byunit_2016.fst</li> <li>HyADS_grids_pm25_total_2016.fst</li> </ul> </li> <li> Mexico Emissions Data [<a href="https://www.epa.gov/air-emissions-modeling/2014-2016-version-7-air-emissions-modeling-platforms">National Emissions Inventory Collaborative, 2016v1 emissions modeling platform</a>] <ul> <li>Mexico_2016_point_interpolated_02mar2018_v0.csv</li> </ul> </li> <li> North American Regional Reanalysis Meteorological Data [<a href="https://psl.noaa.gov/data/gridded/data.narr.monolevel.html">NOAA</a>] <ul> <li>rhum.2m.mon.mean.nc</li> <li>uwnd.10m.mon.mean.nc</li> <li>vwnd.10m.mon.mean.nc</li> </ul> </li> <li> Cigarette smoking data [<a href="https://doi.org/10.1186/1478-7954-12-5">Dwyer-Lindgren et al. (2014)</a>] <ul> <li>smokedatwithfips_1996-2012.csv</li> </ul> </li> <li> Synthetic pediatric asthma data [<em>Note:<strong> synthetic data!</strong> Simulated to match the format, but not the observations, from the <a href="https://www.dshs.texas.gov/texas-health-care-information-collection">Texas Health Care Information Collection (THCIC), Texas DSHS</a></em>] <ul> <li>synth-ped-asthma-data.csv</li> </ul> </li> <li> Texas state shape file [<a href="https://www.census.gov/geographies/mapping-files/time-series/geo/carto-boundary-file.html">US Census</a>] <ul> <li>texas-state-sf.RDS</li> </ul> </li> <li> US ZIPcode-to-county data crosswalk [<a href="https://mcdc.missouri.edu/applications/geocorr2014.html">Missouri Census Data Center</a>] <ul> <li>tx-zip-to-county.csv</li> </ul> </li> </ol> <p>Code and supplementary material from this analysis, as well as more detailed data descriptions, are available at: <a href="https://github.com/nbwikle/estimating-interference">https://github.com/nbwikle/estimating-interference</a></p>
[Data] Qualify-As-You-Go: Sensor Fusion of Optical and Acoustic Signatures with Contrastive Deep Learning for Multi-Material Composition Monitoring in Laser Powder Bed Fusion Process
<p><br>Growing demand for multi-material Laser Powder Bed Fusion (LPBF) faces process control and quality monitoring challenges, particularly in ensuring precise material composition. This study explores optical and acoustic emission signals during LPBF processes with multiple materials, addressing challenges in process control and ensuring accurate material composition. Experimental data from processing five powder compositions were collected using a custombuilt monitoring system in a commercial LPBF machine. The research categorised signals from LPBF processing various compositions, enhancing prediction accuracy by combining optical with acoustic data and training convolutional neural networks using contrastive learning. Latent spaces of trained models using two contrastive loss functions, clustered acoustic and optical<br>emissions based on similarities, aligning with five compositions. Contrastive learning and sensor fusion were found to be essential for monitoring LPBF processes involving multiple materials. This research advances the understanding of multi-material LPBF, highlighting sensor fusion strategies’ potential for improving quality control in additive manufacturing. Data set for this work is hosted here</p>
Dense vegetation hinders sediment transport towards saltmarsh interiors - Supporting data and source code (Part IV: Post-processing)
<p>This is Part IV of the supporting data and source code for the paper entitled "Dense vegetation hinders sediment transport towards saltmarsh interiors", submitted to <em>Limnology and Oceanography Letters.</em> It contains all input and output files for the post-processing of all model results.</p> <p>To be able to run the scripts as is, the folder structure should be as follows:</p> <p>Runs (includes all model run folders from Part II and Part III)<br>Post/Basic/Channels<br>Post/Basic/Cross_sections<br>Post/Basic/Integrals<br>Post/Basic/Median_neighborhood_analysis (includes all unzipped MNA_TIGER_XX.zip folders)<br>Post/Basic/Skeleton_clean<br>Post/Basic/Skeleton_final<br>Post/Basic/Skeleton_raw<br>Post/Basic/Unchanneled_path_length<br>Post/Basic/Watersheds<br>Post/Basic/Scenarios.txt<br>Post/Basic/TIGER_2km_5m.slf<br>Post/Paper_1/Erosion-deposition<br>Post/Paper_1/Fluxes<br>Post/Paper_1/Profiles<br>Post/Paper_1/Std</p>
Raw and processed autofluorescence neuroimaging data
<p>Raw and processed <em>in vivo</em> mouse whole-cortex autofluorescence imaging data.</p> <p>This release accompanies the pub <a href="https://doi.org/10.57844/arcadia-b963-15ac">"Label-free neuroimaging in mice captures sensory activity in response to tactile stimuli and acute pain</a>.</p> <p>The raw data is split into five zipfiles because it is quite large (70GB total). Each of these five zipfiles corresponds to one of the five subdirectories in the single zipfile of processed data.</p>
CATCH-EyoU: Processes in Youth's Construction of Active EU Citizenship: Longitudinal Survey Data: Wave 1 & Wave 2: Estonia
<p>The data set was generated within the research project Constructing AcTive CitizensHip with European Youth: Policies, Practices, Challenges and Solutions (CATCH-EyoU) funded by European Union, Horizon 2020 Programme - Grant Agreement No 649538. The data set is a truncated version of the adolescents’ and young adults’ longitudinal survey that was carried out in Estonia from October 2016 to February 2018. It merges results of two polls (15-19 and 20-30 year olds). Survey was conducted by Univversity of Tartu (UT) within the WP7 research activity which aims at testing processes influencing societal and political engagement of young people.</p>
CATCH-EyoU: Processes in Youth's Construction of Active EU Citizenship: Survey Data: Estonia: Wave 2
<p>The data set was generated within the research project Constructing AcTive CitizensHip with European Youth: Policies, Practices, Challenges and Solutions (CATCH-EyoU) funded by European Union, Horizon 2020 Programme - Grant Agreement No 649538. The data set is a truncated version of the adolescents’ and young adults’ survey that was carried out in Estonia from November 2017 to February 2018. It merges results of two polls (15-19 and 20-30 year olds). Survey was conducted by University of Taartu (UT) within the WP7 research activity which aims at testing processes influencing societal and political engagement of young people.</p>
Data processing scripts and images from e-MERLIN project CY6213 used in Ghirlanda et al. 2019, Science
<p>Data processing scripts and images from e-MERLIN project CY6213 used in Ghirlanda et al. 2019, Science</p> <p> </p> <ul> <li>info.txt contains a general description of how the data was processed and the main results.</li> <li>pipeline.tar contains the data pipeline used to process the e-MERLIN observations</li> <li>imaging.py is the script used to produce the final images</li> <li>CY6213_images.tar contain the final images of the target source (not corrected by calibration factor, described in the imaging script).</li> </ul> <p> </p>
Processed data for the study on "Chromatin 3D interactions mediate genetic effects on gene expression"
<p>This repository contains the processed data that was generated as part of the following study:</p> <p>Delaneau et al. (2019) <strong>Chromatin 3D interactions mediate genetic effects on gene expression.</strong></p> <p><em>Abstract:</em> Studying the genetic basis of gene expression and chromatin organization is key to characterize the effect of genetic variability on the function and structure of the human genome. Here, we unravel how genetic variation perturbs gene regulation using a dataset combining activity of regulatory elements, gene expression and genetic variants across 317 individuals and two cell types. We show that variability in regulatory activity is structured at the intra- and inter-chromosomal levels within 12,583 Cis Regulatory Domains and 30 Trans Regulatory Hubs that highly reflect the local (i.e. Topologically Associating Domains) and global (i.e. open/close chromatin compartments) nuclear chromatin organization. These structures delimit cell type specific regulatory networks that control gene expression/co-expression and mediate the genetic effects of <em>cis</em>- and <em>trans</em>-acting regulatory variants on genes.</p> <p> </p> <p>This repository contains:</p> <ol> <li>Chromatin QTLs for H3K27ac, H3K4me1 and H3K4me3 discovered in 317 Lymphoblastoids Cell Lines (LCLs) and 78 Fibroblasts.</li> <li>Molecular QTLs affecting the activity and structure of Cis Regulatory Domains (CRDs) in LCLs.</li> <li>Basic information about the full set of genetic variants being analyzed in the study.</li> <li>The peak coordinates, their hierarchy based on inter-individual correlation and the CRD calls for both LCLs and Fibroblasts.</li> <li>The functional links discovered in LCLs between CRDs and genes.</li> <li>eQTLs for LCLs.</li> <li>A README file containing the description of the file format for each file.</li> </ol>
CellSIUS provides sensitive and specific detection of rare cell populations from complex single cell RNA-seq data: Codes and processed data
<p>Codes and processed data to reproduce the analysis discussed in: </p> <p>Wegmann <em>et Al.</em>,<strong> CellSIUS provides sensitive and specific detection of rare cell<br> populations from complex single cell RNA-seq data</strong>, Genome Biology 2019 (Accepted)<br> </p>
Measurements of Ice Nucleating Particles in Beijing, China - Data and processing code
<p>Dataset needed to replicate findings published in the journal article "Measurements of Ice Nucleating Particles in Beijing, China, published in the Journal of Geophysical Research. The dataset contains the following:</p> <p>1. Data files containing raw data from a Continuous Flow Diffusion Chamber - Ice Activation Spectrometer (CFDC-IAS), in comma-delimited format).</p> <p>2. Data and processing files for analysis of backward air trajectories as an Igor Pro 8 packed experiment package file. Igor Pro is available from www.wavemetrics.com and a free 30-day trial version can be used to export data to other formats.</p> <p>3. Data and processing files for analysis of CFDC, APS and meteorological data, including data in the form of waves, as part of an Igor Pro 8 packed experiment package.</p>
Data for the publication "Incorporation of inline warm rain diagnostics into the COSP2 satellite simulator for process-oriented model evaluation"
<p>Michibata et al. (2019), currently under peer-review for publication in <em>Geoscientific Model Development</em>, incorporated a diagnostic tool for warm rain microphysics into the CFMIP Observation Simulator Package (COSP; Bodas-Salcedo et al. 2011; Swales et al., 2018), designed to evaluate model representations of aerosol–cloud–precipitation interactions at a fundamental process-level. The tool automatically generates two diagnostics related to warm rain microphysics during COSP execution in a host model. One is the contoured frequency by optical depth diagram (CFODD), which visualizes a cloud-to-rain microphysical vertical structure (Suzuki et al., 2015). The other diagnostic is a global map of warm rain fraction classified as non-precipitating clouds (< –15 dBZ<sub>e</sub>), drizzling clouds (–15 < dBZ<sub>e</sub>< 0), and precipitating clouds (0 < dBZ<sub>e</sub>).</p> <p>This repository contains the MIROC6/COSP2 input data and A-Train satellite statistics used in Michibata et al. (2019). A sample of the post-processing scripts for visualization using the GrADS software is also included in this repository.</p>
Processed data from SnoHATS and METCRAX II: anisotropic turbulence and geometry of the Reynolds stress tensor in a streamline coordinate system
<p>Datasets used for the paper 'Interpreting turbulence anisotropy in a streamline coordinate system'. Data from SnoHATS and METCRAX II field campaigns. Datasets include turbulent quantities calculated on 30- and 1-min averaging windows for unstable and stable conditions, with prior linear detrending. Planar fit was used in METCRAX II and double rotation in SnoHATS to rotate the flow into the mean wind direction. Datasets include quantities to characterize the anisotropy of the Reynolds stress tensor, such as eigenvalues, eigenvectors, and the angles between the eigenvectors and the streamline coordinate system, defined in the direction of the mean wind vector.</p> <p>1c: one-component Reynolds stress tensor</p> <p>2c: two-component axisymmetric Reynolds stress tensor</p> <p>3c: isotropic Reynolds stress tensor</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.