Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
9,863
datasets available to search
ShareScore release 0.7.1
Dataset results
9,863 results for “Buildings”
Relevance of "Building BioData.pt" indicators identified in ESFRI (E), OECD (O) and RI-PATHS (R) impact assessment frameworks
<p>The relevance of the indicators maintained by the "Building BioData.pt" project was assessed against the objectives of different organizations/initiatives: 1) Strategic objectives of BioData.pt; 2) Objectives of the Portuguese Roadmap for Research Infrastructures; 3) Objectives of ELIXIR; 4) Objectives of EOSC; 5) Sustainable Development Goals of the United Nations.</p>
Stormwater runoff pollution of an existing catchment consisting essentially of apartment buildings in Braunschweig/Germany.
<p>This dataset includes stormwater runoff concentrations of pollutants (COD, TP, DP, NO<sub>3</sub>, NH<sub>4</sub>, TSS) taken from a sampling point in Braunschweig, Germany. The total catchment size is 5 ha with approximately 1.8 ha total imperviousness consisting essentially of apartment buildings from the 1970s and roads. This catchment was chosen for comparatively clear delimitation to different land uses considering location-independent results. Samples were taken as part of the research project TransMiT (https://www.transmit-zukunftsstadt.de/). More information including sampling and analytical methods are detailed in the corresponding journal paper "Dynamization of Urban Runoff Pollution and Quantity" and supplementary data, submitted to the MDPI-journal Water.</p> <p>Description of fields:</p> <p>- SamplingPoint:<br> - storm sewer: stormwater runoff sampled during individual rain events in a storm water sewer manhole receiving runoff from the entire catchment<br> - SamplingTime: time of sampling during individual rain event (CET)<br> - COD: measured value of chemical oxygen demand (COD)<br> - TP: measured value of total phosphorus (TP)<br> - DP: measured value of dissovled phosphorus (COD)<br> - NO<sub>3</sub>: measured value of nitrate (NO3)<br> - NH<sub>4</sub>: measured value of ammonium (NH4)<br> - TSS: measured value of total suspended solids (TSS)<br> - N/A: data not available<br> - UnitsAbbreviation: all data given in milligram per litre (mg/L)</p> <p>The data file is provided in comma separated format ("TransMiT.csv") and contains concentrations of all samples.</p> <p> </p>
Brazilian theses and dissertations on Building Information Modeling
<p>Results of a survey of Brazilian dissertations and theses on the topic of Building Information Modeling from the period of 2012 to 2019. </p> <p>The spreadsheet (xls and csv) is organized with the following fields:</p> <p>TYPE: can be thesis or dissertation;</p> <p>AUTHOR: full name of the author in capital letters;</p> <p>TITLE: title of academic work;</p> <p>YEAR: year of defense;</p> <p>UNIVERSITY: Higher Education Institution;</p> <p>LINK: url to the IES repository or to the CAPES SUCUPIRA website where it is possible to obtain the complete work or its cataloging form:</p> <p>RELATED TO EDUCATION AND TEACHING?: YES, when the work is about teaching BIM and NO when it is about BIM unrelated to teaching;</p> <p>ABSTRACT: Abstract of academic work.</p> <p> </p>
Brazilian papers on Building Information Modeling published until 2021
<p>Characterization of Brazilian research in BIM through articles published in Brazilian journals: PARC Research in Architecture and Construction, Project Management and Technology and Built Environment. It covered articles published until 07/2021.</p> <p>The characterization of the articles is by the fields:<br> ITEM: article number in the survey;<br> JOURNAL: Name of the Brazilian journal between Ambiente Construéido, Gestão & Tecnologia de Projetos and PARC Research in Architecture and Construction;<br> TITLE: Title of the article;<br> YEAR: year of publication;<br> INSTITUTION: name by extension of the Institution of the main author;<br> TYPE OF INSTITUTION: values between Public Education Institution, Private Education Institution, Private Company, State Public Company …;<br> LOCATION: Federative Unit of the main author's institution;<br> USE CATEGORY: chosen from MODEL USE SERIES in Category II as in https://bimexcellence.org/wp-content/uploads/211in-Model-Uses-Table.pdf ;<br> SUBCLASSIFICATION: chosen from MODEL USE in Category II as in https://bimexcellence.org/wp-content/uploads/211in-Model-Uses-Table.pdf;<br> BIBLIOGRAPHIC REFERENCE: ABNT format ;<br> LINK: url to the full article.</p>
Data associated with the Tectonics manuscript "Building a Young Mountain Range: Insight into the Growth of the Greater Caucasus Mountains from Detrital Zircon (U-Th)/He Thermochronology and 10Be Erosion Rates"
<p>U-Pb and U-Th/He ages of zircons from a suite of detrital catchments reported in the manuscript "Building a Young Mountain Range: Insight into the Growth of the Greater Caucasus Mountains from Detrital Zircon (U-Th)/He Thermochronology and 10Be Erosion Rates" submitted to Tectonics. Repository includes sample locations and DEMs of each sampled catchment.</p>
UK Administrative Shapefiles clipped to buildings (simplified at 100m)
<p>This dataset includes a series of modified UK administrative boundary shapefiles based on the 2011 census which are intended for use in more accurate visualisation of UK geospatial data analysis. There are two key features of these shapefiles: (1) administrative shapes have been clipped to the Ordnance Survey buildings shapefile, so that in choropleth visualisations relating to demographic data filled spaces represent populated areas of the UK rather than large undifferentiated blocks. (2) Shapefiles have been simplified to reduce loading and processing time, in the case of this repository at 100m. After testing, we have settled on a procedure to render buildings layer visually comprehensible at high zoom levels, by adding a small buffer, dissolving (so that individual overlapping shapes combine into a single more easily visualised shape) and then simplifying at 150m. It is important to emphasise that because of the use of simplification (using a Ramer–Douglas–Peucker algorithm), these shapefiles are not suitable for analysis as boundaries may not be suitably precise or accurate. For users interested in the process used to generate these files you can consult the codebase deposited on <a href="https://github.com/kidwellj/uk_census_shapes_clipped">github</a>.</p> <p>Many thanks to colleagues including Alasdair Rae for recommendations on technique used here. Computations were performed using the University of Birmingham's BEAR Cloud service, which provides flexible resource for intensive computational work to the University's research community. See <a href="http://www.birmingham.ac.uk/bear">http://www.birmingham.ac.uk/bear</a> for more details. Given the massive size of datasets involved (including the district buildings vector shapefile which is 1.4gb and consists of hundreds of thousands of individual shapes), this work would have been impossible without this invaluable resource. I hope that these files will be of use to colleagues who may not have access to similar large computational arrays and make the process of visualising UK boundary and census data more accurate and efficient.</p> <p>Original files are under OGLv3 licenses. Derived data files, where possible are licensed for use under CC BY 4.0.</p> <p>Files include the following:</p> <p><em>Original unmodified data:</em></p> <ul> <li>infuse_ctry_2011.zip - original country level shapes, based on 2011 census, downloaded from https://borders.ukdataservice.ac.uk/ukborders/easy_download</li> <li>infuse_dist_lyr_2011.zip - original local authority shapes, based on 2011 census, downloaded from https://borders.ukdataservice.ac.uk/ukborders/easy_download</li> <li>TermsAndConditions.html - UK Data Service license details (OGLv3), applies to all the above</li> <li> GB_Postcodes.zip - UK postcode district shapes, prepared by Addy Pope, https://datashare.ed.ac.uk/handle/10283/2597</li> </ul> <p>Derived data files:</p> <ul> <li>OS_Open_Zoomstack_district_buildings.zip - buildings layer extracted from <a href="https://www.ordnancesurvey.co.uk/business-government/products/open-zoomstack">Ordnance Survey Zoomstack package</a>, licensed under <a href="https://www.nationalarchives.gov.uk/doc/open-government-licence/version/3/">OGLv3</a> and exported to gpkg format.</li> <li>*_simplified_100m.gpkg - Administrative shapes from above, simplified in R at a resolution of 100 metres.</li> <li>*_simplified_100m_buildings_overlay_simplified.gpkg - Administrative shapes from above, simplified in R at a resolution of 100 metres, and then clipped to the buildings layer.</li> <li>*_simplified_100m_buildings_overlay_simplified.gpkg - Administrative shapes from above, simplified in R at a resolution of 100 metres, and then run against the buildings layer as a difference layer. Suitable for using as an overlay as the shapes are inverse.</li> </ul> <p>Users who wish to use these shapefiles in a reproducible research context may want to download individual files directly from this repository. To do so, you could use the following R code:</p> <pre><code># load packages require(sf) # load simplefeature data class, supercedes sp() and used for st_read # given the size and complexity even of simplified files here, ragg is highly recommended # for users on macos given inefficiencies in default R graphics device require(ragg) # create paths as needed if (dir.exists("data") == FALSE) { dir.create("data") } # download data files only if they aren't already present if (file.exists("data/infuse_dist_lyr_2011.shp") == FALSE) { download.file("https://borders.ukdataservice.ac.uk/ukborders/easy_download/prebuilt/shape/infuse_dist_lyr_2011.zip", destfile = "data/infuse_dist_lyr_2011.zip") unzip("infuse_dist_lyr_2011.zip", exdir = "data")} local_authorities <- st_read("data/infuse_dist_lyr_2011.shp")</code></pre> <p> </p>
Code and data associated with: Searching the web builds fuller picture of arachnid trade
<p>Data and code used in the paper: Searching the web builds fuller picture of arachnid trade. Throughout the methods we have indicated the stage of analysis each data component was used and the code script connected. We have numbered to code and data supplements to reflect as closely as possible the order in which data generation and summary was undertaken. The following provide additional details linked to each of the data files.</p> <p>Data S1 - Website data: lang = language of the search engine used, ad hoc websites had language described after discovery; engine = the search engine used; page = the page on which the website appeared from the search engine; searchdate = search date in YYYY-mm-dd HH:MM:SS; link = link to the webpage, redacted to protect website identity; reviewdate = date revewied for arachnids being sold and search strategy; sells = whether the website sells arachnids (1 == sells); allow = whether the site explcicilt forbids automated searching (1 == allows, NA when search method was not fully automated, e.g., single page); type = the type of the website (e.g., trade, classified ads); order = whether arachnids where organised in a particular ways; target = a refined target URL to start search; method = the search method chosen, see methods for details; refine = any refinement or filter than could constrain the scope of the website to be searched; spages = the number of pages required to cycle through to cover the entire stock (also separated by ; if multiple cycles where needed or multiple single pages could be easily collected); prelimCheck = whether the website passed initial checks for arachnid selling; notes = any details that might need special attention during searches; webID = code used for subsequent data summary.</p> <p>Data S2 - Raw keyword searches outputs: species keywords. sp = the modern species or genus that a keyword is associated with; page = the number of the page the keyword was detected on; keyw = the exact keyword that was detected; spORgen = whether the keyword was a species binomial or just genus; termsSurrounding = the words surrounding a genus keyword detection (only applies to Data S3); webID = the website ID.</p> <p>Data S3 – Raw keyword searches outputs: genus keywords. sp = the modern species or genus that a keyword is associated with; page = the number of the page the keyword was detected on; keyw = the exact keyword that was detected; spORgen = whether the keyword was a species binomial or just genus; termsSurrounding = the words surrounding a genus keyword detection (multiple detections separated by ;); webID = the website ID.</p> <p>Data S4 - Raw keyword search outputs: temporal sample. sp = the modern species or genus that a keyword is associated with; page = the number of the page the keyword was detected on; keyw = the exact keyword that was detected; spORgen = whether the keyword was a species binomial or just genus; termsSurrounding = the words surrounding a genus keyword detection (multiple detections separated by ;); webID = the website ID; timestamp.parse = the timestamp extracted from the archived web page; year = a simplified timestamp including only the year.</p> <p>Data S5 - LEMIS data used. An arachnid filtered version of <sup>74,75</sup>.</p> <p>Data S6 - CITES trade database data used <sup>76</sup>.</p> <p>Data S7 - CITES appendices data used <sup>77</sup>.</p> <p>Data S8 - IUCN Redlist data used <sup>78</sup>.</p> <p>Data S9 - Compiled final dataset, with data deriving from WSC, Scorpion files, ITIS, WAM and the data collection process. speciesId = a numeric code, one per species; clade = the clade the species belongs to; family = the family the species belongs to; genus = the genus of the species; species = the species epithet; author = the species authority name; year = the species authority year; parentheses = whether parentheses are needed with the authority; distribution = WSC original distribution descriptions; invalid = whether the species is considered valid; source = the species source, either World Spider Catalogue, Scorpion files, ITIS or WAM; accName = the species binomial being used as our accepted name; allNames = the accepted species binomial and all synonyms; allGenera = the accepted genus, and all other genera the species has belonged to at one point; onlineTradeSnap = whether the species was detected via a match to the accName in the snapshot data; onlineTradeSnap_Any = whether the species was detected via any synonym in the snapshot data; onlineTradeSnap_genus = whether the genus was detected via a match to the genus in the snapshot data; onlineTradeSnap_genusAny = whether the genus was detected via any synonym in the snapshot data; onlineTradeTemp = whether the species was detected via a match to the accName in the temporal data; onlineTradeTemp_Any = whether the species was detected via any synonym in the temporal data; onlineTradeTemp_genus = whether the genus was detected via a match to the genus in the temporal data; onlineTradeTemp_genusAny = whether the genus was detected via any synonym in the temporal data; onlineTradeEither = whether the species was detected via a match to the accName in the temporal data or snapshot data; onlineTradeEither_Any = whether the species was detected via any synonym in the temporal data or snapshot data; LEMIStrade = whether the species was detected via a match to the accName in the LEMIS data; LEMIStrade_Any = whether the species was detected via any synonym in the LEMIS data; LEMIStrade_genus = whether the genus was detected via any synonym in the LEMIS data; LEMIStrade_genusAny = whether the genus was detected via any synonym in the LEMIS data; CITEStrade = whether the species was detected via a match to the accName in the CITES trade database data; CITEStrade_Any = whether the species was detected via any synonym in the CITES trade database data; CITEStrade_genus = whether the genus was detected via any synonym in the CITES trade database data; CITEStrade_genusAny = whether the genus was detected via any synonym in the CITES trade database data; CITESapp = the CITES appendix the species is listed under using an exact match to the accName; CITESapp_Any = the CITES appendix the species is listed under using any match to any of the species’ synonyms; redlist = the IUCN Redlist category the species is listed under using an exact match to the accName; redlist_Any = the IUCN Redlist category the species is listed under using any match to any of the species’ synonyms; extactMatchTraded = the species is detected in any of the trade sources via a match to the accName; anyMatchTraded = the species is detected in any of the trade sources via a match to any species’ synonym.</p> <p>Data S10 - Forum listings of “What species are you currently keeping” from an online fora posted between 9th September 2021 and 9th October 2021, to provide an idea of online discussions. Each user with a separate list is provided in a separate tab. Morph_collector is the same as poster1, but the potential cryptic species or morphs are noted separately to make them clearer.</p> <p>Data S11 – Distribution information for spiders. Only two columns used in summaries: accName = the accepted name used throughout summaries; NAME = the country name the spider occurs in.</p> <p>Data S12 - Distribution information for scorpions. species = the accepted name used throughout summaries; NAME = the country name the scorpions occurs in.</p> <p>Code S1 - Search URL Extract.R</p> <p>Code S2 - Retrieve web data.R</p> <p>Code S3 - Temporal Classified Ads.R</p> <p>Code S4 - Keyword Generation.R</p> <p>Code S5 - Keyword Search.R</p> <p>Code S6 - LEMIS filter and summary.R</p> <p>Code S7 - Compiling results.R</p> <p>Code S8 - Summary Figures.R</p> <p>Code S9 - Temporal Figures.R</p> <p>Code S10 - New description figure.R</p> <p>Code S11 - Term exploration.R</p> <p>Code S12 - LEMIS summary and mapping.R</p>
Dataset of EnergyPlus models to evaluate the impact of modeling the hysteresis phenomenon of phase change materials on the building energy performance
<p>This dataset is the research data generated to evaluate the impact of modeling the hysteresis phenomenon of phase change materials (PCM) on the building performance simulation, which includes:<br> - A series of EnergyPlus models representing the medium office of the Prototype Building Models developed by DOE. These are the original model without PCM (Baseline), and four models with different PCM modeling approaches (melting-curve, solidification-curve, mean-curve, hysteresis-model).<br> - The typical meteorological year (TMY) for Frankfurt city that was used to obtain the results, which is freely provided by Climate.One.Building.Org repository (https://climate.onebuilding.org/).</p>
Building status collected during AWOPS tests
<p>The dataset is about building status collected during the validation of the Automated Work Planning Services (AWOPS). The dataset includes:</p> <ul> <li>1 csv file regarding renovation works' activities carried out from March 7<sup>th</sup> to April 11<sup>th</sup>;</li> <li>9 json files regarding progress data collected during the following days: <ul> <li>March 14<sup>th</sup> 2022 (day 2);</li> <li>March 17<sup>th</sup> 2022 (day 5);</li> <li>March 21<sup>st</sup> 2022 (day 9);</li> <li>March 24<sup>th</sup> 2022 (day 12);</li> <li>March 28<sup>th</sup> 2022 (day 16);</li> <li>March 31<sup>st</sup> 2022 (day 19);</li> <li>April 4<sup>th</sup> 2022 (day 23);</li> <li>April 7<sup>th</sup> 2022 (day 26);</li> <li>April 11<sup>th</sup> 2022(day 30).</li> </ul> </li> </ul>
Building and characterizing a fluorescence setup to measure very low concentrations of analytes/biomarkers
<p>This training report is the result of my internship in the B-Phot Brussels Photonics team of the VUB<br> that took place between February 3 and April 3, 2020. The project of this internship nds its context<br> in the European SensApp project which regroups several European research institutes and universities,<br> including the VUB. The goal of this project is to develop a method to diagnose the Alzheimer's disease<br> in a faster and non-invasive manner, simply through a blood test, which is currently not possible because<br> the concentration of biomarkers of the Alzheimer's disease in the blood is too low. During this internship,<br> I was lead to build, align and calibrate a uorescence detection setup. Using this setup, I made mea-<br> surements of the uorescence intensity of low concentrations of dye solutions. From those measurements,<br> I performed calculations of the signal to noise ratio in order to determine the limit of detection of the<br> setup. Finally, I studied the kinetics of photobleaching in order to get to a better understanding of its<br> impact on the measurements.</p>
BIM4EEB Demonstration building BIM Model
<p>BIM4EEB Demonstration building BIM Model. The model contains only external perimeter walls, structural geometry, and common public spaces, The internal layout of the apartments is undesclosed for privacy reasons.</p>
Historic building's interior (iPhone LiDAR scan)
<p>This model shows a LiDAR-scan of selected room interiors of a historic building located in Lower Austria (AUT) recorded during building-archaeological measures by the archaeological company <strong><a href="https://www.ardig.at/">ARDIG</a></strong>, triggered by recent remodeling work in the course of house renovation. Scanning was done in two rounds using 3dScannerApp on an iPhone 13 Pro. Each round took approx. < 5 min of capturing and another < 5 min of on-board processing time. After the first room was fully captured in 3D and with textures in the first round, the scan was (nearly) seamlessly extended by another room in a second round using the corresponding app function. Therefore, <strong>within less than approx. 20 minutes, both rooms could be fully documented in 3D by a scaled and textured 3D model</strong>. Considering the extreme flexibility and fast acquisition and processing time, the method holds tremendous future potential for archaeological work, even if still several minor errors occur in the final product.</p>
A full-year data regarding a smart building
<p>The smart building represented by the dataset is divided into five zones, where each zone has 2 to 3 offices. The building has photovoltaic panels with a generation peak of 7.5 kW, inside sensors, light intensity control, and all consumption measured by load type. The data was collected in 5 minutes periods.</p> <p>By zones, the buildings have a distribution of 6 researchers in Zone#1, 5 researchers in Zone#2, 5 researchers in Zone#3, 3 researchers in Zone#4, and 5 researchers in Zone#5. Zone#1 includes a meeting room, and Zone#2 includes a server room. Regarding the server room, its HVAC unit was measured in HVAC#2, however, near the end of the year, the unit was removed from the monitoring system.</p> <p>The light intensity values represent the current state of the lamps and have a linear correlation with the lamp’s consumption.</p> <p>The dataset represents raw data without any treatment, this means that it is possible to find errors. The data do not have missing reading periods, but they can have a fixed zero (0) value, indicating a failure in the system. You can assume that periods with zero voltage represent an error in that zone at that period.</p> <p> </p> <p><br>We would be grateful if you could acknowledge the use of this dataset in your publications. Please use the Zenodo publication to cite this work.</p>
Estimated height of the OpenStreetMap buildings of 24 French communes using the GeoClimate Software (version 0.0.1)
<p>This repository contains:</p> <ul> <li>a folder called "_toReproduceResults" containing data, script and methodology to reproduce most of the work described in the research manuscript,</li> <li>the main output of the research work: 24 folders (each of them corresponding to a French city), containing building geometry footprints and their corresponding building height as well as averaged building height value aggregated at rectangular grid cell (100 m by 100 m). The footprint geometries comes from the OpenStreetMap project and the building height has been estimated using a RandomForest algorithm using as independent variables indicators describing the building size and shape and the building environment. The data has been produced using the GeoClimate Software (version 0.0.1).</li> </ul> <p>A more detailed description of the content can be used in the file "Metadata.csv".</p>
Xanthene[n]arenes: Exceptionally Large, Bowl-Shaped Macrocyclic Building Blocks Suitable for Self-Assembly
<p>Data underlying the figures in the publication “Xanthene[<em>n</em>]arenes: Exceptionally Large, Bowl-Shaped Macrocyclic Building Blocks Suitable for Self-Assembly”, published in <em>J</em><em>ACS Au</em> 2021, 1, 11, 1885–1891. <a href="https://doi.org/10.1021/jacsau.1c00343">https://doi.org/10.1021/jacsau.1c00343</a></p>
A private house in 'Marea'/Philoxenite transformed into a monastic institution and other Christian hybrid buildings in the Mareotis region - additional illustrations
<p>Set of three documents on potsherds (ostraca) has been found in Philoxenite (Egypt). All three of the ostraca, M200249 (= Document A), M200250 (= Document B), and M200251 (= Document C) are preserved in their entirety, in the sense that the sherds, meant to serve as the canvas for the documents, were not broken when they were thrown out into the rubbish pit in the piscina.</p> <p>The photos uploaded here provide supplementary material for an article discussing the content of these ostraca. </p> <p> </p>
Dataset of multi-objective optimization results for a new latent energy storage approach in buildings based on several phase change materials with different melting temperatures
<p>This dataset comprises the multi-objective optimization results obtained for a new latent energy storage approach based on several phase change materials (PCMs) with different melting temperatures in buildings. The results were obtained for a small office building in eight climate-representative locations according to the ASHRAE 169-2020 climate classification and within the WMO Region VI (Europe).</p> <p>The dataset contains:</p> <p>- The EnergyPlus baseline models employed as a case study for each climate.</p> <p>- The Pareto fronts obtained after the multi-objective optimization in each climate.</p> <p>- The EnergyPlus models for the best designs achieved on the Pareto fronts in terms of annual total load reductions.</p>
PheKnowLator Human Disease Knowledge Graphs - Build Data (Processed)
<p><strong>RELEASE V2.1.0 KNOWLEDGE GRAPH: PROCESSED DATA SOURCES </strong></p> <p><strong>Release:</strong> <a href="https://github.com/callahantiff/PheKnowLator/wiki/v2.0.0">v2.1.0 </a></p> <p>The goal of this build was to create a knowledge graph that represented human disease mechanisms and included the central dogma. The data sources utilized in this release include many of the sources used in the initial release, as well as some new data made available by the <a href="https://ctdbase.org/">Comparative Toxicogenomics Database</a> and experimental data from the <a href="https://www.proteinatlas.org/">Human Protein Atlas</a>.</p> <p>Data sources are listed by type (Ontology and Data not represented in an ontology [Database Sources]). Additional details are provided for each data source below. Please see documentation on the primary release (<a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources">https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources</a>) for additional details on each data source as well as citation information.</p> <p><strong>Data Access:</strong></p> <ul> <li><a href="https://console.cloud.google.com/storage/browser/pheknowlator/archived_builds/release_v2.1.0/build_01MAY2021?project=pheknowlator">https://console.cloud.google.com/storage/browser/pheknowlator/archived_builds/release_v2.1.0/build_01MAY2021</a></li> </ul> <p> </p> <p><strong>ONTOLOGIES</strong></p> <ul> <li>Cell Ontology</li> <li>Cell Line Ontology</li> <li>Chemical Entities of Biological Interest (ChEBI) Ontology</li> <li>Gene Ontology</li> <li>Human Phenotype Ontology</li> <li>Mondo Disease Ontology</li> <li>Pathway Ontology</li> <li>Protein Ontology</li> <li>Relations Ontology</li> <li>Sequence Ontology</li> <li>Uber-Anatomy Ontology</li> <li>Vaccine Ontology</li> </ul> <p> </p> <p><strong>Cell Ontology (CL)</strong></p> <p><strong>Homepage:</strong> <strong><a href="https://github.com/obophenotype/cell-ontology"><code>GitHub</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Bard J, Rhee SY, Ashburner M. <a href="https://genomebiology.biomedcentral.com/articles/10.1186/gb-2005-6-2-r21">An ontology for cell types</a>. Genome Biology. 2005;6(2):R21</p> </blockquote> <p><strong>Usage:</strong> Utilized to connect <code>transcripts</code> and <code>proteins</code> to <code>cells</code>. Additionally, the edges between this ontology and its dependencies are utilized:</p> <ul> <li><strong><a href="https://www.ebi.ac.uk/chebi/"><code>ChEBI</code></a></strong></li> <li><strong><a href="http://geneontology.org/"><code>GO</code></a></strong></li> <li><strong><a href="https://github.com/pato-ontology/pato/"><code>PATO</code></a></strong></li> <li><strong><a href="https://proconsortium.org/"><code>PRO</code></a></strong></li> <li><strong><a href="https://github.com/oborel/obo-relations/"><code>RO</code></a></strong></li> <li><strong><a href="https://uberon.github.io/"><code>UBERON</code></a></strong></li> </ul> <p> </p> <p><strong>Cell Line Ontology (CLO)</strong></p> <p><strong>Homepage:</strong> <strong><a href="http://www.clo-ontology.org/"><code>http://www.clo-ontology.org/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Sarntivijai S, Lin Y, Xiang Z, Meehan TF, Diehl AD, Vempati UD, Schürer SC, Pang C, Malone J, Parkinson H, Liu Y. <a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4387853/">CLO: the cell line ontology</a>. Journal of Biomedical Semantics. 2014;5(1):37</p> </blockquote> <p><strong>Usage:</strong> Utilized this ontology to map <code>cell lines</code> to <code>transcripts</code> and <code>proteins</code>. Additionally, the edges between this ontology and its dependencies are utilized:</p> <ul> <li><strong><a href="http://www.obofoundry.org/ontology/cl.html"><code>CL</code></a></strong></li> <li><strong><a href="http://disease-ontology.org/"><code>DOID</code></a></strong></li> <li><strong><a href="https://github.com/obophenotype/ncbitaxon"><code>NCBITaxon</code></a></strong></li> <li><strong><a href="https://uberon.github.io/"><code>UBERON</code></a></strong></li> </ul> <p> </p> <p><strong>Chemical Entities of Biological Interest (ChEBI)</strong></p> <p><strong>Homepage:</strong> <strong><a href="https://www.ebi.ac.uk/chebi/"><code>https://www.ebi.ac.uk/chebi/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Hastings J, Owen G, Dekker A, Ennis M, Kale N, Muthukrishnan V, Turner S, Swainston N, Mendes P, Steinbeck C. <a href="https://academic.oup.com/nar/article-abstract/44/D1/D1214/2502583">ChEBI in 2016: Improved services and an expanding collection of metabolites</a>. Nucleic Acids Research. 2015;44(D1):D1214-9</p> </blockquote> <p><strong>Usage:</strong> Utilized to connect <code>chemicals</code> to <code>complexes</code>, <code>diseases</code>, <code>genes</code>, <code>GO biological processes</code>, <code>GO cellular components</code>, <code>GO molecular functions</code>, <code>pathways</code>, <code>phenotypes</code>, <code>reactions</code>, and <code>transcripts</code>.</p> <p> </p> <p><strong>Gene Ontology (GO)</strong></p> <p><strong>Homepage:</strong> <strong><a href="http://geneontology.org/"><code>http://geneontology.org/</code></a></strong><br> <strong>Citations:</strong></p> <blockquote> <p>Ashburner M, Ball CA, Blake JA, Botstein D, Butler H, Cherry JM, Davis AP, Dolinski K, Dwight SS, Eppig JT, Harris MA. <a href="https://www.nature.com/articles/ng0500_25">Gene ontology: tool for the unification of biology</a>. Nature Genetics. 2000;25(1):25</p> <p>The Gene Ontology Consortium. <a href="https://academic.oup.com/nar/article/47/D1/D330/5160994">The Gene Ontology Resource: 20 years and still GOing strong</a>. Nucleic Acids Research. 2018;47(D1):D330-8</p> </blockquote> <p><strong>Usage:</strong> Utilized to connect <code>biological processes</code>, <code>cellular components</code>, and <code>molecular functions</code> to <code>chemicals</code>, <code>pathways</code>, and <code>proteins</code>. Additionally, the edges between this ontology and its dependencies are utilized:</p> <ul> <li><strong><a href="http://www.obofoundry.org/ontology/cl.html"><code>CL</code></a></strong></li> <li><strong><a href="https://github.com/obophenotype/ncbitaxon"><code>NCBITaxon</code></a></strong></li> <li><strong><a href="https://github.com/oborel/obo-relations/"><code>RO</code></a></strong></li> <li><strong><a href="https://uberon.github.io/"><code>UBERON</code></a></strong></li> </ul> <p><strong>Other Gene Ontology Data Used:</strong> <a href="http://geneontology.org/gene-associations/goa_human.gaf.gz"><code>goa_human.gaf.gz</code></a></p> <p> </p> <p><strong>Human Phenotype Ontology (HPO)</strong></p> <p><strong>Homepage:</strong> <strong><a href="https://hpo.jax.org/"><code>https://hpo.jax.org/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Köhler S, Carmody L, Vasilevsky N, Jacobsen JO, Danis D, Gourdine JP, Gargano M, Harris NL, Matentzoglu N, McMurry JA, Osumi-Sutherland D. <a href="https://academic.oup.com/nar/article-abstract/47/D1/D1018/5198478">Expansion of the Human Phenotype Ontology (HPO) knowledge base and resources</a>. Nucleic Acids Research. 2018;47(D1):D1018-27</p> </blockquote> <p><strong>Usage:</strong> Utilized to connect <code>phenotypes</code> to <code>chemicals</code>, <code>diseases</code>, <code>genes</code>, and <code>variants</code>. Additionally, the edges between this ontology and its dependencies are utilized:</p> <ul> <li><strong><a href="http://www.obofoundry.org/ontology/cl.html"><code>CL</code></a></strong></li> <li><strong><a href="https://www.ebi.ac.uk/chebi/"><code>ChEBI</code></a></strong></li> <li><strong><a href="http://geneontology.org/"><code>GO</code></a></strong></li> <li><strong><a href="https://uberon.github.io/"><code>UBERON</code></a></strong></li> </ul> <p><strong>Files</strong></p> <ul> <li>Other Human Phenotype Ontology Data Used: <a href="http://purl.obolibrary.org/obo/hp/hpoa/phenotype.hpoa"><code>phenotype.hpoa</code></a></li> </ul> <p> </p> <p><strong>Mondo Disease Ontology (Mondo)</strong></p> <p><strong>Homepage:</strong> <strong><a href="https://mondo.monarchinitiative.org/"><code>https://mondo.monarchinitiative.org/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Mungall CJ, McMurry JA, Köhler S, Balhoff JP, Borromeo C, Brush M, Carbon S, Conlin T, Dunn N, Engelstad M, Foster E. <a href="https://academic.oup.com/nar/article-abstract/45/D1/D712/2605791">The Monarch Initiative: an integrative data and analytic platform connecting phenotypes to genotypes across species</a>. Nucleic Acids Research. 2017;45(D1):D712-22</p> </blockquote> <p><strong>Usage:</strong> Utilized to connect <code>diseases</code> to <code>chemicals</code>, <code>phenotypes</code>, <code>genes</code>, and <code>variants</code>. Additionally, the edges between this ontology and its dependencies are utilized:</p> <ul> <li><strong><a href="http://www.obofoundry.org/ontology/cl.html"><code>CL</code></a></strong></li> <li><strong><a href="https://www.ncbi.nlm.nih.gov/taxonomy"><code>NCBITaxon</code></a></strong></li> <li><strong><a href="http://geneontology.org/"><code>GO</code></a></strong></li> <li><strong><a href="https://hpo.jax.org/"><code>HPO</code></a></strong></li> <li><strong><a href="https://uberon.github.io/"><code>UBERON</code></a></strong></li> </ul> <p> </p> <p><strong>Pathway Ontology (PW)</strong></p> <p><strong>Homepage:</strong> <strong><a href="https://rgd.mcw.edu/wg/home/pathway2/"><code>rgd.mcw.edu</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Petri V, Jayaraman P, Tutaj M, Hayman GT, Smith JR, De Pons J, Laulederkind SJ, Lowry TF, Nigam R, Wang SJ, Shimoyama M. <a href="https://www.ncbi.nlm.nih.gov/pubmed/24499703">The pathway ontology–updates and applications</a>. Journal of Biomedical Semantics. 2014;5(1):7.</p> </blockquote> <p><strong>Usage:</strong> Utilized to connect <code>pathways</code> to <code>GO biological processes</code>, <code>GO cellular components</code>, <code>GO molecular functions</code>, <code>Reactome pathways</code>. Several steps are taken in order to connect <code>Pathway Ontology</code> identifiers to <code>Reactome</code> pathways and <code>GO biological processes</code>. To connect <code>Pathway Ontology</code> identifiers to <code>Reactome</code> pathways, we use <a href="https://github.com/ComPath/resources/tree/master/mappings">ComPath Pathway Database Mappings</a> developed by Daniel Domingo-Fernández (<a href="https://www.ncbi.nlm.nih.gov/pubmed/30564458">PMID:30564458</a>).</p> <p><strong>Files</strong></p> <ul> <li>Downloaded Mapping Data <ul> <li><a href="http://compath.scai.fraunhofer.de/export_mappings"><code>curated_mappings.txt</code></a></li> <li><a href="https://github.com/ComPath/resources/blob/master/mappings/kegg_reactome.csv"><code>kegg_reactome.csv</code></a></li> </ul> </li> <li>Generated Mapping Data <ul> <li><a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/REACTOME_PW_GO_MAPPINGS.txt"><code>REACTOME_PW_GO_MAPPINGS.txt</code></a></li> </ul> </li> </ul> <p> </p> <p><strong>Protein Ontology (PRO)</strong></p> <p><strong>Homepage:</strong> <strong><a href="https://proconsortium.org/"><code>https://proconsortium.org/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Natale DA, Arighi CN, Barker WC, Blake JA, Bult CJ, Caudy M, Drabkin HJ, D’Eustachio P, Evsikov AV, Huang H, Nchoutmboube J. <a href="https://academic.oup.com/nar/article-abstract/39/suppl_1/D539/2508558">The Protein Ontology: a structured representation of protein forms and complexes</a>. Nucleic Acids Research. 2010;39(suppl_1):D539-45</p> </blockquote> <p><strong>Usage:</strong> Utilized to connect <code>proteins</code> to <code>chemicals</code>, <code>genes</code>, <code>anatomy</code>, <code>catalysts</code>, <code>cell lines</code>, <code>cofactors</code>, <code>complexes</code>, <code>GO biological processes</code>, <code>GO cellular components</code>, <code>GO molecular functions</code>, <code>pathways</code>, <code>proteins</code>, <code>reactions</code>, and <code>transcripts</code>. Additionally, the edges between this ontology and its dependencies are utilized:</p> <ul> <li><strong><a href="https://www.ebi.ac.uk/chebi/"><code>ChEBI</code></a></strong></li> <li><strong><a href="http://disease-ontology.org/"><code>DOID</code></a></strong></li> <li><strong><a href="http://geneontology.org/"><code>GO</code></a></strong></li> </ul> <p><strong>Notes:</strong> A partial, human-only version of this ontology was used. Details on how this version of the ontology was generated can be found under the Protein Ontology section of the <a href="https://github.com/callahantiff/PheKnowLator/blob/master/notebooks/Data_Preparation.ipynb"><code>Data_Preparation.ipynb</code></a> Jupyter Notebook.</p> <p><strong>Files</strong></p> <ul> <li> <p>Generated Human Version Protein Ontology (PRO)</p> <ul> <li><a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/human_pro.owl"><code>human_pro.owl</code></a> (closed with <a href="http://www.hermit-reasoner.com/">hermit reasoner</a>)</li> </ul> </li> <li> <p>Other PRO Data Used: <a href="https://proconsortium.org/download/current/promapping.txt"><code>promapping.txt</code></a></p> </li> <li> <p>Generated Mapping Data</p> <ul> <li>Merged Gene, RNA, Protein Map: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/Merged_gene_rna_protein_identifiers.pkl"><code>Merged_gene_rna_protein_identifiers.pkl</code></a></li> <li>Ensembl Transcript-PRO Identifier Mapping: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/ENTREZ_GENE_ENSEMBL_TRANSCRIPT_MAP.txt"><code>ENSEMBL_TRANSCRIPT_PROTEIN_ONTOLOGY_MAP.txt</code></a></li> <li>Entrez Gene-PRO Identifier Mapping: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/ENTREZ_GENE_PRO_ONTOLOGY_MAP.txt"><code>ENTREZ_GENE_PRO_ONTOLOGY_MAP.txt</code></a></li> <li>UniProt Accession-PRO Identifier Mapping: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/UNIPROT_ACCESSION_PRO_ONTOLOGY_MAP.txt"><code>UNIPROT_ACCESSION_PRO_ONTOLOGY_MAP.txt</code></a></li> <li>STRING-PRO Identifier Mapping: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/STRING_PRO_ONTOLOGY_MAP.txt"><code>STRING_PRO_ONTOLOGY_MAP.txt</code></a></li> </ul> </li> </ul> <p> </p> <p><strong>Relations Ontology (RO)</strong></p> <p><strong>Homepage:</strong> <strong><a href="https://github.com/oborel/obo-relations/"><code>GitHub</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Smith B, Ceusters W, Klagges B, Köhler J, Kumar A, Lomax J, Mungall C, Neuhaus F, Rector AL, Rosse C. <a href="https://genomebiology.biomedcentral.com/articles/10.1186/gb-2005-6-5-r46">Relations in biomedical ontologies</a>. Genome Biology. 2005;6(5):R46.</p> </blockquote> <p><strong>Usage:</strong> Utilizing this ontology to connect all data sources in knowledge graph. Additionally, the ontology is queried prior to building the knowledge graph to identify all relations, their inverse properties, and their labels.</p> <p><strong>Files</strong></p> <ul> <li>Generated RO Data <ul> <li><a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/INVERSE_RELATIONS.txt"><code>INVERSE_RELATIONS.txt</code></a></li> <li><a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/RELATIONS_LABELS.txt"><code>RELATIONS_LABELS.txt</code></a></li> </ul> </li> </ul> <p> </p> <p><strong>Sequence Ontology (SO)</strong></p> <p><strong>Homepage:</strong> <strong><a href="https://github.com/The-Sequence-Ontology/SO-Ontologies"><code>GitHub</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Eilbeck K, Lewis SE, Mungall CJ, Yandell M, Stein L, Durbin R, Ashburner M. <a href="https://link.springer.com/article/10.1186/gb-2005-6-5-r44">The Sequence Ontology: a tool for the unification of genome annotations</a>. Genome Biology. 2005;6(5):R44</p> </blockquote> <p><strong>Usage:</strong> Utilized to connect <code>transcripts</code> and other genomic material like <code>genes</code> and <code>variants</code>.</p> <p><strong>Files</strong></p> <ul> <li>Generated Mapping Data <ul> <li><a href="https://storage.googleapis.com/pheknowlator/curated_data/genomic_sequence_ontology_mappings.xlsx"><code>genomic_sequence_ontology_mappings.xlsx</code></a></li> <li><a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/SO_GENE_TRANSCRIPT_VARIANT_TYPE_MAPPING.txt"><code>SO_GENE_TRANSCRIPT_VARIANT_TYPE_MAPPING.txt</code></a></li> </ul> </li> </ul> <p> </p> <p><strong>Uber-Anatomy Ontology (Uberon)</strong></p> <p><strong>Homepage:</strong> <strong><a href="https://uberon.github.io/"><code>GitHub</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Mungall CJ, Torniai C, Gkoutos GV, Lewis SE, Haendel MA. <a href="https://genomebiology.biomedcentral.com/articles/10.1186/gb-2012-13-1-r5">Uberon, an integrative multi-species anatomy ontology</a>. Genome Biology. 2012;13(1):R5</p> </blockquote> <p><strong>Usage:</strong> Utilized to connect <code>tissues</code>, <code>fluids</code>, and <code>cells</code> to <code>proteins</code> and <code>transcripts</code>. Additionally, the edges between this ontology and its dependencies are utilized:</p> <ul> <li><strong><a href="https://www.ebi.ac.uk/chebi/"><code>ChEBI</code></a></strong></li> <li><strong><a href="http://www.obofoundry.org/ontology/cl.html"><code>CL</code></a></strong></li> <li><strong><a href="http://geneontology.org/"><code>GO</code></a></strong></li> <li><strong><a href="https://proconsortium.org/"><code>PRO</code></a></strong></li> </ul> <p> </p> <p><strong>Vaccine Ontology (VO)</strong></p> <p><strong>Homepage:</strong> <strong><a href="http://www.violinet.org/vaccineontology/"><code>http://www.violinet.org/vaccineontology/</code></a></strong><br> <strong>Citations:</strong></p> <blockquote> <p>He Y, Racz R, Sayers S, Lin Y, Todd T, Hur J, Li X, Patel M, Zhao B, Chung M, Ostrow J. <a href="https://academic.oup.com/nar/article-abstract/42/D1/D1124/1053128">Updates on the web-based VIOLIN vaccine database and analysis system</a>. Nucleic Acids Research. 2013;42(D1):D1124-32</p> <p>Xiang Z, Todd T, Ku KP, Kovacic BL, Larson CB, Chen F, Hodges AP, Tian Y, Olenzek EA, Zhao B, Colby LA. <a href="https://academic.oup.com/nar/article-abstract/36/suppl_1/D923/2505793">VIOLIN: vaccine investigation and online information network</a>. Nucleic Acids Research. 2007;36(suppl_1):D923-8</p> </blockquote> <p><strong>Usage:</strong> Utilized the edges between this ontology and its dependencies:</p> <ul> <li><strong><a href="https://www.ebi.ac.uk/chebi/"><code>ChEBI</code></a></strong></li> <li><strong><a href="http://disease-ontology.org/"><code>DOID</code></a></strong></li> <li><strong><a href="http://geneontology.org/"><code>GO</code></a></strong></li> <li><strong><a href="https://proconsortium.org/"><code>PRO</code></a></strong></li> <li><strong><a href="https://uberon.github.io/"><code>UBERON</code></a></strong></li> </ul> <p> </p> <p><strong>DATABASE SOURCES</strong></p> <ul> <li>BioPortal</li> <li>ClinVar</li> <li>Comparative Toxicogenomics Database</li> <li>DisGeNET</li> <li>Ensembl</li> <li>GeneMANIA</li> <li>Genotype-Tissue Expression Project</li> <li>Human Genome Organisation Gene Nomenclature Committee</li> <li>Human Protein Atlas</li> <li>National Center for Biotechnology Information Gene</li> <li>Reactome Pathway Database</li> <li>Search Tool for Recurring Instances of Neighbouring Genes Database</li> <li>Universal Protein Resource Knowledgebase</li> </ul> <p> </p> <p><strong>BioPortal</strong></p> <p><strong>Homepage:</strong> <strong><a href="https://bioportal.bioontology.org/"><code>BioPortal</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>BioPortal. <a href="https://www.bioontology.org/wiki/LOOM">Lexical OWL Ontology Matcher (LOOM)</a></p> <p>Ghazvinian A, Noy NF, Musen MA. <a href="https://www.ncbi.nlm.nih.gov/pmc/articles/pmc2815474/">Creating mappings for ontologies in biomedicine: simple methods work</a>. In AMIA Annual Symposium Proceedings 2009 (Vol. 2009, p. 198). American Medical Informatics Association</p> </blockquote> <p><strong>Usage:</strong> BioPortal was utilized to obtain mappings between <code>MeSH identifiers</code> and <code>ChEBI identifiers</code> for <code>chemicals-diseases</code>, <code>chemicals-genes</code>, <code>chemical-GO biological processes</code>, <code>chemicals-GO cellular components</code>, <code>chemicals-GO molecular functions</code>, <code>chemicals-phenotypes</code>, <code>chemicals-proteins</code>, and <code>chemicals-transcripts</code>. Additional information on how this data was processed can be obtained from the <a href="https://gist.github.com/callahantiff/a28fb3160782f42f104e9ec41553af0d"><code>NCBO_rest_api.py</code></a> GitHub Gist script.</p> <p>⭐ <strong>ALTERNATIVE METHOD</strong>⭐ Since the above approach can take over two days to process, we have developed an alternative solution that downloads the <a><code>mesh2021.nt</code></a> data file directly from MeSH and the <a><code>Flat_file_tab_delimited/names.tsv.gz</code></a> file directly from ChEBI. Using these files, we have recapitulated the <a href="https://www.bioontology.org/wiki/BioPortal_Mappings"><code>LOOM</code></a> algorithm implemented by BioPortal when creating mappings between these resources. The procedure is relatively straightforward and utilizes the following information from each resource:</p> <ul> <li>For all MeSH <code>SCR Chemicals</code>, obtain the following information: <ul> <li>Identifiers: MeSH identifiers</li> <li>Labels: string labels using the <code>RDFS:label</code> object property</li> <li>Synonyms: track down all synonyms using the <code>vocab:concept</code> and <code>vocab:preferredConcept</code> object properties</li> </ul> </li> <li>For all ChEBI classes, obtain the following information: <ul> <li>Labels: string labels using the <code>RDFS:label</code> object property</li> <li>Synonyms: track down all synonyms using all <code>synonym</code> object properties</li> </ul> </li> </ul> <p><strong>Files</strong></p> <ul> <li>Generated Data: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/MESH_CHEBI_MAP.txt"><code>MESH_CHEBI_MAP.txt</code></a></li> </ul> <p> </p> <p><strong>ClinVar</strong></p> <p><strong>Homepage:</strong> <strong><a href="https://www.ncbi.nlm.nih.gov/clinvar/"><code>https://www.ncbi.nlm.nih.gov/clinvar/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Landrum MJ, Lee JM, Benson M, Brown GR, Chao C, Chitipiralla S, Gu B, Hart J, Hoffman D, Jang W, Karapetyan K. <a href="https://academic.oup.com/nar/article-abstract/46/D1/D1062/4641904">ClinVar: improving access to variant interpretations and supporting evidence</a>. Nucleic Acids Research. 2017;46(D1):D1062-7</p> </blockquote> <p><strong>Usage:</strong> ClinVar was utilized to create <code>variant-gene</code>, <code>variant-disease</code>, and <code>variant-phenotype</code> edges. The original data is filtered such that only records meeting the following criteria were included:</p> <ul> <li> <p><code>Assembly</code> = "GRCh38"</p> </li> <li> <p><code>ClinSigSimple</code> = <code>1</code></p> <ul> <li> <blockquote> <p>1 = at least one current record submitted with an interpretation of Likely pathogenic or Pathogenic (independent of whether that record includes assertion criteria and evidence)"</p> </blockquote> </li> </ul> </li> <li> <p><code>ReviewStatus</code> in ["criteria provided, multiple submitters, no conflicts", "reviewed by expert panel", "practice guideline"]</p> </li> </ul> <p><strong>Files</strong></p> <ul> <li> <p>Downloaded Data</p> <ul> <li><a href="https://ftp.ncbi.nlm.nih.gov/pub/clinvar/tab_delimited/variant_summary.txt.gz"><code>variant_summary.txt.gz</code></a></li> <li><a href="https://ftp.ncbi.nlm.nih.gov/pub/clinvar/tab_delimited/var_citations.txt"><code>var_citations.txt</code></a></li> <li><a href="https://ftp.ncbi.nlm.nih.gov/pub/clinvar/tab_delimited/allele_gene.txt.gz"><code>allele_gene.txt.gz</code></a></li> </ul> </li> <li> <p>Generated Edge Data: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/CLINVAR_VARIANT_GENE_DISEASE_PHENOTYPE_EDGES.txt"><code>CLINVAR_VARIANT_GENE_DISEASE_PHENOTYPE_EDGES.txt</code></a></p> </li> </ul> <p> </p> <p><strong>Comparative Toxicogenomics Database (CTD)</strong></p> <p><strong>Homepage:</strong> <strong><a href="http://ctdbase.org/"><code>http://ctdbase.org/</code></a></strong><br> <strong>Citations:</strong></p> <blockquote> <p>Curated [chemical–gene interactions|chemical-go interactions|chemical–disease interactions|gene–pathway interactions] data were retrieved from the Comparative Toxicogenomics Database (CTD), MDI Biological Laboratory, Salisbury Cove, Maine, and NC State University, Raleigh, North Carolina. World Wide Web (URL: <a href="http://ctdbase.org/">http://ctdbase.org/</a>)</p> <p>Davis AP, Grondin CJ, Johnson RJ, Sciaky D, McMorran R, Wiegers J, Wiegers TC, Mattingly CJ. <a href="https://academic.oup.com/nar/article-abstract/47/D1/D948/5106145">The comparative toxicogenomics database: update 2019</a>. Nucleic Acids Research. 2018;47(D1):D948-54</p> </blockquote> <p>Usage: Comparative Toxicogenomics Database (CTD) was utilized to create <code>chemical-disease</code>, <code>chemical-gene</code>, <code>chemical-GO biological process</code>, <code>chemical-GO cellular components</code>, <code>chemical-GO molecular functions</code>, <code>chemical-phenotype</code>, <code>chemical-protein</code>, <code>chemical-rna</code>, and <code>gene-pathway</code> edges. The original data is filtered such that only records meeting the following criteria were included:</p> <ul> <li><code>chemical-disease</code>: <code>DirectEvidence</code> != ""</li> <li><code>chemical-gene</code>: <code>Organism</code> == "Homo sapiens", <code>GeneForms</code> == "gene", and affects not in <code>InteractionActions</code></li> <li><code>chemical-GO biological process</code>: <code>PhenotypeName</code> == "Biological Process" and <code>Interaction</code> <= "1.04e-47" (10th percentile)</li> <li><code>chemical-GO cellular components</code>: <code>PhenotypeName</code> == "Cellular Component" and <code>Interaction</code> <= "1.04e-47" (10th percentile)</li> <li><code>chemical-GO molecular functions</code>: <code>PhenotypeName</code> == "Molecular Function" and <code>Interaction</code> <= "1.04e-47" (10th percentile)</li> <li><code>chemical-phenotype</code>: <code>DirectEvidence</code> != ""</li> <li><code>chemical-protein</code>: <code>Organism</code> == "Homo sapiens", <code>GeneForms</code> == "protein", and affects not in <code>InteractionActions</code></li> <li><code>chemical-rna</code>: <code>Organism</code> == "Homo sapiens", <code>GeneForms</code> == "mRNA", and affects and activity not in <code>InteractionActions</code></li> <li><code>gene-pathway edges</code>: <code>PathwayName</code> == R-HSA-</li> </ul> <p><strong>Files</strong></p> <ul> <li>Downloaded Data <ul> <li>Chemical-Gene Relations: <a href="http://ctdbase.org/reports/CTD_chem_gene_ixns.tsv.gz"><code>CTD_chem_gene_ixns.tsv.gz</code></a></li> <li>Chemical-Disease/Phenotype Relations: <a href="http://ctdbase.org/reports/CTD_chemicals_diseases.tsv.gz"><code>CTD_chemicals_diseases.tsv.gz</code></a></li> <li>Chemical-GO Relations: <a href="http://ctdbase.org/reports/CTD_chem_go_enriched.tsv.gz"><code>CTD_chem_go_enriched.tsv.gz</code></a></li> <li>Gene-Pathway Relations: <a href="http://ctdbase.org/reports/CTD_genes_pathways.tsv.gz"><code>CTD_genes_pathways.tsv.gz</code></a></li> </ul> </li> </ul> <p> </p> <p><strong>DisGeNET</strong></p> <p><strong>Homepage:</strong> <strong><a href="https://www.disgenet.org/"><code>https://www.disgenet.org/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Gene-disease association data retrieved from DisGeNET v6.0 (<a href="http://www.disgenet.org/">http://www.disgenet.org/</a>), Integrative Biomedical Informatics Group GRIB/IMIM/UPF. [December, 2019].</p> <p>Piñero J, Ramírez-Anguita JM, Saüch-Pitarch J, Ronzano F, Centeno E, Sanz F, Furlong LI. <a href="https://academic.oup.com/nar/advance-article-abstract/doi/10.1093/nar/gkz1021/5611674">The DisGeNET knowledge platform for disease genomics: 2019 update</a>. Nucleic Acids Research. 2019.</p> </blockquote> <p><strong>Usage:</strong> DisGeNET was utilized to create <code>gene-disease</code>, and <code>gene-phenotype</code> edges. The original data is filtered such that only records meeting the following criteria were included: <code>EI</code> >= "1.0" (90th percentile). Additionally, data from this source was used to create mappings between different types of disease and phenotype identifiers, including:</p> <ul> <li>OMIM, ORPHA, UMLS, ICD ➞ DOID</li> <li>OMIM, ORPHA, UMLS, ICD ➞ HPO</li> </ul> <p><strong>Files</strong></p> <ul> <li> <p>Downloaded Data</p> <ul> <li>Disease/Phenotype-Gene Relations: <a href="https://www.disgenet.org/static/disgenet_ap1/files/downloads/curated_gene_disease_associations.tsv.gz"><code>curated_gene_disease_associations.tsv.gz</code></a></li> <li>Disease Identifier Mapping: <a href="https://www.disgenet.org/static/disgenet_ap1/files/downloads/disease_mappings.tsv.gz"><code>disease_mappings.tsv.gz</code></a></li> </ul> </li> <li> <p>Generated Mapping Data</p> <ul> <li>Disease Identifier Mapping: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/PHENOTYPE_HPO_MAP.txt"><code>PHENOTPYE_HPO_MAP.txt</code></a></li> <li>Phenotype Identifier Mapping: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/DISEASE_MONDO_MAP.txt"><code>DISEASE_DOID_MAP.txt</code></a></li> </ul> </li> </ul> <p> </p> <p><strong>Ensembl</strong></p> <p><strong>Homepage:</strong> <strong><a href="https://uswest.ensembl.org/index.html"><code>https://uswest.ensembl.org/index.html</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Zerbino DR, Achuthan P, Akanni W, Amode MR, Barrell D, Bhai J, Billis K, Cummins C, Gall A, Girón CG, Gil L. <a href="https://academic.oup.com/nar/article/46/D1/D754/4634002">Ensembl 2018</a>. Nucleic Acids Research. 2017;46(D1):D754-61</p> </blockquote> <p><strong>Usage:</strong> Ensembl data was utilized to create mappings between Ensembl genes, transcripts, and proteins with <a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#national-center-for-biotechnology-information-gene">NCBI Gene identifiers</a>, <a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#human-genome-organisation-gene-nomenclature-committee">HUGO gene symbols</a>, <a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#universal-protein-resource-knowledgebase">UniProt Accession identifiers</a>, and <a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#protein-ontology">Protein Ontology identifiers</a> in the knowledge graph (for additional details on the processing of these data, see <a href="https://github.com/callahantiff/PheKnowLator/blob/master/notebooks/Data_Preparation.ipynb"><code>Data_Preparation.ipynb</code></a>):</p> <ul> <li>Ensembl Transcript IDs ➞ PRO IDs</li> <li>Gene Ensembl IDs ➞ Entrez Gene IDs</li> <li>Gene Ensembl IDs ➞ PRO IDs</li> <li>Gene Symbols ➞ Transcript Ensembl IDs</li> <li>Entrez Gene IDs ➞ Transcript Ensembl IDs</li> <li>Entrez Gene IDs ➞ PRO IDs</li> <li>Protein Ensembl IDs ➞ UniProt Protein Accession</li> <li>STRING IDs ➞ PRO IDs</li> <li>UniProt Protein Accession ➞ Entrez Gene IDs</li> </ul> <p><strong>Files</strong></p> <ul> <li> <p>Downloaded Data</p> <ul> <li><a><code>Homo_sapiens.GRCh38.102.gtf</code></a></li> <li><a><code>Homo_sapiens.GRCh38.102.uniprot.tsv.gz</code></a></li> <li><a><code>Homo_sapiens.GRCh38.102.entrez.tsv.gz</code></a></li> </ul> </li> <li> <p>Generated Mapping Data</p> <ul> <li>Cleaned Ensembl Gene Set: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/ensembl_identifier_data_cleaned.txt"><code>ensembl_identifier_data_cleaned.txt</code></a></li> <li>Merged Gene, RNA, Protein Map: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/Merged_gene_rna_protein_identifiers.pkl"><code>Merged_gene_rna_protein_identifiers.pkl</code></a></li> <li>Ensembl Transcript-PRO Identifier Mapping: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/ENSEMBL_TRANSCRIPT_PROTEIN_ONTOLOGY_MAP.txt"><code>ENSEMBL_TRANSCRIPT_PROTEIN_ONTOLOGY_MAP.txt</code></a></li> <li>Gene Symbol-Ensembl Transcript Identifier Mapping: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/GENE_SYMBOL_ENSEMBL_TRANSCRIPT_MAP.txt"><code>GENE_SYMBOL_ENSEMBL_TRANSCRIPT_MAP.txt</code></a></li> <li>Entrez Gene-Ensembl Transcript Identifier Mapping: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/ENTREZ_GENE_ENSEMBL_TRANSCRIPT_MAP.txt"><code>ENTREZ_GENE_ENSEMBL_TRANSCRIPT_MAP.txt</code></a></li> <li>Entrez Gene-PRO Identifier Mapping: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/ENTREZ_GENE_PRO_ONTOLOGY_MAP.txt"><code>ENTREZ_GENE_PRO_ONTOLOGY_MAP.txt</code></a></li> <li>Ensembl Gene-Entrez Gene Identifier Mapping: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/ENSEMBL_GENE_ENTREZ_GENE_MAP.txt"><code>ENSEMBL_GENE_ENTREZ_GENE_MAP.txt</code></a></li> </ul> </li> </ul> <p> </p> <p><strong>GeneMANIA</strong></p> <p><strong>Homepage:</strong> <strong><a href="https://genemania.org/"><code>https://genemania.org/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Warde-Farley D, Donaldson SL, Comes O, Zuberi K, Badrawi R, Chao P, Franz M, Grouios C, Kazi F, Lopes CT, Maitland A. <a href="https://academic.oup.com/nar/article-abstract/38/suppl_2/W214/1126704">The GeneMANIA prediction server: biological network integration for gene prioritization and predicting gene function</a>. Nucleic Acids Research. 2010;38(suppl_2):W214-20</p> </blockquote> <p><strong>Usage:</strong> GeneMANIA was utilized to create <code>gene-gene</code> edges.</p> <p><strong>Files</strong></p> <ul> <li>Downloaded Data: <a href="http://genemania.org/data/current/Homo_sapiens.COMBINED/COMBINED.DEFAULT_NETWORKS.BP_COMBINING.txt"><code>COMBINED.DEFAULT_NETWORKS.BP_COMBINING.txt</code></a></li> </ul> <p> </p> <p><strong>Genotype-Tissue Expression Project (GTEx)</strong></p> <p><strong>Homepage:</strong> <strong><a href="https://gtexportal.org/home/"><code>https://gtexportal.org/home/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Lonsdale J, Thomas J, Salvatore M, Phillips R, Lo E, Shad S, Hasz R, Walters G, Garcia F, Young N, Foster B. <a href="http://www.nature.com/ng/journal/v45/n6/full/ng.2653.html">The genotype-tissue expression (GTEx) project</a>. Nature Genetics. 2013;45(6):580</p> </blockquote> <p><strong>Usage:</strong> The Genotype-Tissue Expression (GTEx) Project was utilized to create edges between <code>protein-cell</code>, <code>protein-anatomy</code>, <code>rna-cell</code> and <code>rna-anatomy</code> entities. The original data were filtered such that only those edges where the median TPM was >=<code>1.0</code> and genes were of any type other than protein-coding were included. It should also be noted that we chose to use the RNASeQC file over the RSEM file as advised by the GTEx website.</p> <blockquote> <p>The RSEM estimates are based on combining isoform-level estimates, which adds uncertainty to the resulting gene-level values (the isoform-level estimates are highly inaccurate in some cases).</p> </blockquote> <p>The file contains <code>54</code> unique tissue and/or cell types. GTEx provides mappings from tissue types to UBERON and EFO. These provided <a href="https://gtexportal.org/home/samplingSitePage">mappings</a> were verified and extended, such that all samples which referenced a cell type were also mapped to the Cell and the Cell Line ontologies. This resulted in a total of <code>56</code> mappings (<code>1.04</code> mappings/concepts).</p> <p><strong>Files</strong></p> <ul> <li> <p>Downloaded Data: <a href="https://storage.googleapis.com/gtex_analysis_v8/rna_seq_data/GTEx_Analysis_2017-06-05_v8_RNASeQCv1.1.9_gene_median_tpm.gct.gz"><code>GTEx_Analysis_2017-06-05_v8_RNASeQCv1.1.9_gene_median_tpm.gct</code></a></p> </li> <li> <p>Mapping Results: <a href="https://storage.googleapis.com/pheknowlator/curated_data/zooma_tissue_cell_mapping_04JAN2020.xlsx"><code>zooma_tissue_cell_mapping_04JAN2020.xlsx</code></a></p> </li> <li> <p>Generated Data<br> The final mapping set was combined with terms from the <a href="https://www.proteinatlas.org/">Human Protein Atlas</a>, see <a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources/t#human-protein-atlas">here</a> for more information.</p> <ul> <li>All HPA tissue and cell type strings: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/HPA_tissues.txt"><code>HPA_tissues.txt</code></a></li> <li>Final Term Mapping: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/HPA_GTEx_TISSUE_CELL_MAP.txt"><code>HPA_GTEx_TISSUE_CELL_MAP.txt</code></a></li> <li>Final RNA, Gene, Protein-Tissues and Cell Types Relations: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/HPA_GTEX_RNA_GENE_PROTEIN_EDGES.txt"><code>HPA_GTEX_RNA_GENE_PROTEIN_EDGES.txt</code></a></li> </ul> </li> </ul> <p> </p> <p><strong>Human Genome Organisation Gene Nomenclature Committee (HUGO)</strong></p> <p><strong>Homepage:</strong> <strong><a href="https://www.genenames.org/"><code>https://www.genenames.org/</code></a></strong><br> <strong>Citations:</strong></p> <blockquote> <p>HGNC Database, HUGO Gene Nomenclature Committee (HGNC), European Molecular Biology Laboratory, European Bioinformatics Institute (EMBL-EBI), Wellcome Genome Campus, Hinxton, Cambridge CB10 1SD, United Kingdom <a href="https://www.genenames.org/">www.genenames.org</a></p> <p>Yates B, Braschi B, Gray K, Seal R, Tweedie S, Bruford E. <a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5210531/">Genenames.org: the HGNC and VGNC Resources in 2017</a>. Nucleic Acids Research. 2017;45(D1):D619-625</p> </blockquote> <p><strong>Usage:</strong> The Human Genome Organisation (HUGO) data was utilized to obtain mappings between <a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#national-center-for-biotechnology-information-gene">NCBI Gene identifiers</a>, HUGO gene symbols, <a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#universal-protein-resource-knowledgebase">UniProt Accession identifiers</a>, and <a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#protein-ontology">Protein Ontology identifiers</a>. For additional details on the processing of these data, see <a href="https://github.com/callahantiff/PheKnowLator/blob/master/notebooks/Data_Preparation.ipynb"><code>Data_Preparation.ipynb</code></a>:</p> <ul> <li>Ensembl Transcript IDs ➞ PRO IDs</li> <li>Gene Ensembl IDs ➞ Entrez Gene IDs</li> <li>Gene Ensembl IDs ➞ PRO IDs</li> <li>Gene Symbols ➞ Transcript Ensembl IDs</li> <li>Entrez Gene IDs ➞ Transcript Ensembl IDs</li> <li>Entrez Gene IDs ➞ PRO IDs</li> <li>Protein Ensembl IDs ➞ UniProt Protein Accession</li> <li>STRING IDs ➞ PRO IDs</li> <li>UniProt Protein Accession ➞ Entrez Gene IDs</li> </ul> <p><strong>Files</strong></p> <ul> <li> <p>Downloaded Data: <a href="http://ftp.ebi.ac.uk/pub/databases/genenames/hgnc/tsv/hgnc_complete_set.txt"><code>hgnc_complete_set.txt</code></a></p> </li> <li> <p>Generated Data</p> <ul> <li>Merged Gene, RNA, Protein Map: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/Merged_gene_rna_protein_identifiers.pkl"><code>Merged_gene_rna_protein_identifiers.pkl</code></a></li> <li>Gene Symbol-Ensembl Transcript Identifier Mapping: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/GENE_SYMBOL_ENSEMBL_TRANSCRIPT_MAP.txt"><code>GENE_SYMBOL_ENSEMBL_TRANSCRIPT_MAP.txt</code></a></li> </ul> </li> </ul> <p> </p> <p><strong>Human Protein Atlas (HPA)</strong></p> <p><strong>Homepage:</strong> <strong><a href="https://www.proteinatlas.org/"><code>https://www.proteinatlas.org/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Uhlén M, Fagerberg L, Hallström BM, Lindskog C, Oksvold P, Mardinoglu A, Sivertsson Å, Kampf C, Sjöstedt E, Asplund A, Olsson I. <a href="https://science.sciencemag.org/content/347/6220/1260419.short">Tissue-based map of the human proteome</a>. Science. 2015;347(6220):1260419</p> </blockquote> <p><strong>Usage:</strong> The Human Protein Atlas (HPA) was utilized to create <code>rna-cell</code>, <code>rna-anatomy</code>, <code>protein-cell</code>, and <code>protein-anatomy</code> edges. Evidence between gene and RNA expression in specific tissue types was derived by HPA, such that the <a href="https://www.proteinatlas.org/about/assays+annotation#normalization_rna">consensus normalized expression</a> was >=<code>1.0</code>. Zooma was utilized to automatically annotate the <code>153</code> unique tissues and cell types from Human Protein Atlas for all human protein-coding genes in the <a href="https://www.proteinatlas.org/humanproteome">Human Proteome</a> to the Cell Ontology, Cell Line Ontology, and the Uber-Anatomy Ontology. To best represent each concept, the automatic mappings from Zooma were extend through manual mapping efforts to ensure each concept cell type was matched to a Cell Ontology, Cell Line Ontology, and UBERON ontology term. This resulted in a total of <code>281</code> mappings (<code>1.84</code> mappings/concepts).</p> <p><strong>Files</strong></p> <ul> <li> <p>Downloaded Data: <a href="https://www.proteinatlas.org/api/search_download.php?search=&columns=g,eg,up,pe,rnatsm,rnaclsm,rnacasm,rnabrsm,rnabcsm,rnablsm,scl,t_RNA_adipose_tissue,t_RNA_adrenal_gland,t_RNA_amygdala,t_RNA_appendix,t_RNA_basal_ganglia,t_RNA_bone_marrow,t_RNA_breast,t_RNA_cerebellum,t_RNA_cerebral_cortex,t_RNA_cervix,_uterine,t_RNA_colon,t_RNA_corpus_callosum,t_RNA_ductus_deferens,t_RNA_duodenum,t_RNA_endometrium_1,t_RNA_epididymis,t_RNA_esophagus,t_RNA_fallopian_tube,t_RNA_gallbladder,t_RNA_heart_muscle,t_RNA_hippocampal_formation,t_RNA_hypothalamus,t_RNA_kidney,t_RNA_liver,t_RNA_lung,t_RNA_lymph_node,t_RNA_midbrain,t_RNA_olfactory_region,t_RNA_ovary,t_RNA_pancreas,t_RNA_parathyroid_gland,t_RNA_pituitary_gland,t_RNA_placenta,t_RNA_pons_and_medulla,t_RNA_prostate,t_RNA_rectum,t_RNA_retina,t_RNA_salivary_gland,t_RNA_seminal_vesicle,t_RNA_skeletal_muscle,t_RNA_skin_1,t_RNA_small_intestine,t_RNA_smooth_muscle,t_RNA_spinal_cord,t_RNA_spleen,t_RNA_stomach_1,t_RNA_testis,t_RNA_thalamus,t_RNA_thymus,t_RNA_thyroid_gland,t_RNA_tongue,t_RNA_tonsil,t_RNA_urinary_bladder,t_RNA_vagina,t_RNA_B-cells,t_RNA_dendritic_cells,t_RNA_granulocytes,t_RNA_monocytes,t_RNA_NK-cells,t_RNA_T-cells,t_RNA_total_PBMC,cell_RNA_A-431,cell_RNA_A549,cell_RNA_AF22,cell_RNA_AN3-CA,cell_RNA_ASC_diff,cell_RNA_ASC_TERT1,cell_RNA_BEWO,cell_RNA_BJ,cell_RNA_BJ_hTERT+,cell_RNA_BJ_hTERT+_SV40_Large_T+,cell_RNA_BJ_hTERT+_SV40_Large_T+_RasG12V,cell_RNA_CACO-2,cell_RNA_CAPAN-2,cell_RNA_Daudi,cell_RNA_EFO-21,cell_RNA_fHDF/TERT166,cell_RNA_HaCaT,cell_RNA_HAP1,cell_RNA_HBEC3-KT,cell_RNA_HBF_TERT88,cell_RNA_HDLM-2,cell_RNA_HEK_293,cell_RNA_HEL,cell_RNA_HeLa,cell_RNA_Hep_G2,cell_RNA_HHSteC,cell_RNA_HL-60,cell_RNA_HMC-1,cell_RNA_HSkMC,cell_RNA_hTCEpi,cell_RNA_hTEC/SVTERT24-B,cell_RNA_hTERT-HME1,cell_RNA_HUVEC_TERT2,cell_RNA_K-562,cell_RNA_Karpas-707,cell_RNA_LHCN-M2,cell_RNA_MCF7,cell_RNA_MOLT-4,cell_RNA_NB-4,cell_RNA_NTERA-2,cell_RNA_PC-3,cell_RNA_REH,cell_RNA_RH-30,cell_RNA_RPMI-8226,cell_RNA_RPTEC_TERT1,cell_RNA_RT4,cell_RNA_SCLC-21H,cell_RNA_SH-SY5Y,cell_RNA_SiHa,cell_RNA_SK-BR-3,cell_RNA_SK-MEL-30,cell_RNA_T-47d,cell_RNA_THP-1,cell_RNA_TIME,cell_RNA_U-138_MG,cell_RNA_U-2_OS,cell_RNA_U-2197,cell_RNA_U-251_MG,cell_RNA_U-266/70,cell_RNA_U-266/84,cell_RNA_U-698,cell_RNA_U-87_MG,cell_RNA_U-937,cell_RNA_WM-115,blood_RNA_basophil,blood_RNA_classical_monocyte,blood_RNA_eosinophil,blood_RNA_gdT-cell,blood_RNA_intermediate_monocyte,blood_RNA_MAIT_T-cell,blood_RNA_memory_B-cell,blood_RNA_memory_CD4_T-cell,blood_RNA_memory_CD8_T-cell,blood_RNA_myeloid_DC,blood_RNA_naive_B-cell,blood_RNA_naive_CD4_T-cell,blood_RNA_naive_CD8_T-cell,blood_RNA_neutrophil,blood_RNA_NK-cell,blood_RNA_non-classical_monocyte,blood_RNA_plasmacytoid_DC,blood_RNA_T-reg,blood_RNA_total_PBMC,brain_RNA_amygdala,brain_RNA_basal_ganglia,brain_RNA_cerebellum,brain_RNA_cerebral_cortex,brain_RNA_hippocampal_formation,brain_RNA_hypothalamus,brain_RNA_midbrain,brain_RNA_olfactory_region,brain_RNA_pons_and_medulla,brain_RNA_thalamus&format=tsv"><code>proteinatlas_search.tsv</code></a></p> </li> <li> <p>Mapping Results: <a href="https://storage.googleapis.com/pheknowlator/curated_data/zooma_tissue_cell_mapping_04JAN2020.xlsx"><code>zooma_tissue_cell_mapping_04JAN2020.xlsx</code></a></p> </li> <li> <p>Generated Data</p> <ul> <li>Final Term Mapping: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/HPA_GTEx_TISSUE_CELL_MAP.txt"><code>HPA_GTEx_TISSUE_CELL_MAP.txt</code></a></li> <li>Final RNA, Gene, Protein-Tissues and Cell Types Relations: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/HPA_GTEX_RNA_GENE_PROTEIN_EDGES.txt"><code>HPA_GTEX_RNA_GENE_PROTEIN_EDGES.txt</code></a></li> </ul> </li> </ul> <p> </p> <p><strong>National Center for Biotechnology Information (NCBI) Entrez Gene</strong></p> <p><strong>Homepage:</strong> <strong><a href="https://www.ncbi.nlm.nih.gov/gene/"><code>https://www.ncbi.nlm.nih.gov/gene/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Maglott D, Ostell J, Pruitt KD, Tatusova T. <a href="https://academic.oup.com/nar/article-abstract/33/suppl_1/D54/2505255">Entrez Gene: gene-centered information at NCBI</a>. Nucleic Acids Research. 2005;33(suppl_1):D54-8.</p> </blockquote> <p><strong>Usage:</strong> The National Center for Biotechnology Information (NCBI) Gene data was utilized to obtain mappings between <a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#ncbi-gene">NCBI Gene identifiers</a>, <a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#hugo-gene-nomenclature-committee">HUGO gene symbols</a>, <a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#uniprot-knowledgebase">UniProt Accession identifiers</a>, and <a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#protein-ontology">Protein Ontology identifiers</a>. For additional details on the processing of these data, see <a href="https://github.com/callahantiff/PheKnowLator/blob/master/notebooks/Data_Preparation.ipynb"><code>Data_Preparation.ipynb</code></a>:</p> <ul> <li>Ensembl Transcript IDs ➞ PRO IDs</li> <li>Gene Ensembl IDs ➞ Entrez Gene IDs</li> <li>Gene Ensembl IDs ➞ PRO IDs</li> <li>Gene Symbols ➞ Transcript Ensembl IDs</li> <li>Entrez Gene IDs ➞ Transcript Ensembl IDs</li> <li>Entrez Gene IDs ➞ PRO IDs</li> <li>Protein Ensembl IDs ➞ UniProt Protein Accession</li> <li>STRING IDs ➞ PRO IDs</li> <li>UniProt Protein Accession ➞ Entrez Gene IDs</li> </ul> <p><strong>Files</strong></p> <ul> <li> <p>Downloaded Data: <a href="https://ftp.ncbi.nih.gov/gene/DATA/GENE_INFO/Mammalia/Homo_sapiens.gene_info.gz"><code>Homo_sapiens.gene_info.gz</code></a></p> </li> <li> <p>Generated Data</p> <ul> <li>Merged Gene, RNA, Protein Map: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/Merged_gene_rna_protein_identifiers.pkl"><code>Merged_gene_rna_protein_identifiers.pkl</code></a></li> <li>Entrez Gene-Ensembl Transcript Identifier Mapping: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/ENTREZ_GENE_ENSEMBL_TRANSCRIPT_MAP.txt"><code>ENTREZ_GENE_ENSEMBL_TRANSCRIPT_MAP.txt</code></a></li> <li>Entrez Gene-PRO Identifier Mapping: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/ENTREZ_GENE_PRO_ONTOLOGY_MAP.txt"><code>ENTREZ_GENE_PRO_ONTOLOGY_MAP.txt</code></a></li> <li>Ensembl Gene-Entrez Gene Identifier Mapping: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/ENSEMBL_GENE_ENTREZ_GENE_MAP.txt"><code>ENSEMBL_GENE_ENTREZ_GENE_MAP.txt</code></a></li> <li>Uniprot Accession-Entrez Gene Identifier Mapping: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/UNIPROT_ACCESSION_ENTREZ_GENE_MAP.txt"><code>UNIPROT_ACCESSION_ENTREZ_GENE_MAP.txt</code></a></li> </ul> </li> </ul> <p> </p> <p><strong>Reactome Pathway Database</strong></p> <p><strong>Homepage:</strong> <strong><a href="https://reactome.org/"><code>https://reactome.org/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Fabregat A, Jupe S, Matthews L, Sidiropoulos K, Gillespie M, Garapati P, Haw R, Jassal B, Korninger F, May B, Milacic M. <a href="https://academic.oup.com/nar/article-abstract/46/D1/D649/4626770">The reactome pathway knowledgebase</a>. Nucleic Acids Research. 2017;46(D1):D649-55</p> </blockquote> <p><strong>Usage:</strong> The Reactome Database was utilized to create <code>chemical-pathway</code>, <code>GO Biological process-pathway</code>, <code>pathway-GO Cellular component</code>, <code>GO Molecular function-pathway</code>, and <code>protein-pathway</code> edges. The original data is filtered such that only records meeting the following criteria were included:</p> <ul> <li><code>chemical-pathway</code>: column[5] == "Homo sapiens"</li> <li><code>GO Biological process-pathway</code>: column[5] startswith "REACTOME", column[8] == "P", and column[12] == "taxon:9606"</li> <li><code>pathway-GO Cellular component</code>: column[5] startswith "REACTOME", column[8] == "C", and column[12] == "taxon:9606"</li> <li><code>GO Molecular function-pathway</code>: column[5] startswith "REACTOME", column[8] == "F", and column[12] == "taxon:9606"</li> <li><code>protein-pathway</code>: column[5] == "Homo sapiens"</li> </ul> <p><strong>Files</strong></p> <ul> <li>Downloaded Data <ul> <li>Chemical-Pathway Relations: <a href="https://reactome.org/download/current/ChEBI2Reactome_All_Levels.txt"><code>ChEBI2Reactome_All_Levels.txt</code></a></li> <li>Pathway-GO Relations: <a href="https://reactome.org/download/current/gene_association.reactome"><code>gene_association.reactome</code></a></li> <li>Protein-Pathway Relations: <a href="https://reactome.org/download/current/UniProt2Reactome_All_Levels.txt"><code>UniProt2Reactome_All_Levels.txt</code></a></li> </ul> </li> </ul> <p> </p> <p><strong>Search Tool for Recurring Instances of Neighbouring Genes (STRING) Database</strong></p> <p><strong>Homepage:</strong> <strong><a href="https://string-db.org/"><code>string-db.org</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>Szklarczyk D, Gable AL, Lyon D, Junge A, Wyder S, Huerta-Cepas J, Simonovic M, Doncheva NT, Morris JH, Bork P, Jensen LJ. <a href="https://academic.oup.com/nar/article-abstract/47/D1/D607/5198476">STRING v11: protein–protein association networks with increased coverage, supporting functional discovery in genome-wide experimental datasets</a>. Nucleic Acids Research. 2018;47(D1):D607-13</p> </blockquote> <p><strong>Usage:</strong> The Search Tool for Recurring Instances of Neighbouring Genes (STRING) Database was utilized to create <code>protein-protein</code> edges. The original data is filtered such that only records meeting the following criteria were included: <code>combined_score</code> >= "700" (>90th percentile).</p> <p><strong>Files</strong></p> <ul> <li> <p>Downloaded Data: <a href="https://stringdb-static.org/download/protein.links.v11.0/9606.protein.links.v11.0.txt.gz"><code>9606.protein.links.v11.0.txt.gz</code></a></p> </li> <li> <p>Generated Data: STRING-PRO Identifier Mapping: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/STRING_PRO_ONTOLOGY_MAP.txt"><code>STRING_PRO_ONTOLOGY_MAP.txt</code></a></p> </li> </ul> <p> </p> <p><strong>Universal Protein Resource (UniProt) Knowledgebase</strong></p> <p><strong>Homepage:</strong> <strong><a href="https://www.uniprot.org/"><code>https://www.uniprot.org/</code></a></strong><br> <strong>Citation:</strong></p> <blockquote> <p>UniProt Consortium. <a href="https://academic.oup.com/nar/article-abstract/47/D1/D506/5160987">UniProt: a worldwide hub of protein knowledge</a>. Nucleic acids research. 2018;47(D1):D506-15</p> </blockquote> <p><strong>Usage:</strong> The Universal Protein Resource (UniProt) Knowledgebase was utilized to obtain <code>cofactor</code>/<code>catalyst</code>-<code>protein</code> and <code>protein-coding gene</code>-<code>protein</code> edges as well as mappings between <a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#national-center-for-biotechnology-information-gene">NCBI Gene identifiers</a>, <a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#human-genome-organisation-gene-nomenclature-committee">HUGO gene symbols</a>, <a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#universal-protein-resource-knowledgebase">Universal Protein Resource (UniProt) Accession identifiers</a>, and <a href="https://github.com/callahantiff/PheKnowLator/wiki/v2-Data-Sources#protein-ontology">Protein Ontology identifiers</a>. For additional details on the processing of these data, see <a href="https://github.com/callahantiff/PheKnowLator/blob/master/notebooks/Data_Preparation.ipynb"><code>Data_Preparation.ipynb</code></a>:</p> <ul> <li>Ensembl Transcript IDs ➞ PRO IDs</li> <li>Gene Ensembl IDs ➞ Entrez Gene IDs</li> <li>Gene Ensembl IDs ➞ PRO IDs</li> <li>Gene Symbols ➞ Transcript Ensembl IDs</li> <li>Entrez Gene IDs ➞ Transcript Ensembl IDs</li> <li>Entrez Gene IDs ➞ PRO IDs</li> <li>Protein Ensembl IDs ➞ UniProt Protein Accession</li> <li>STRING IDs ➞ PRO IDs</li> <li>UniProt Protein Accession ➞ Entrez Gene IDs</li> </ul> <p><strong>Files</strong></p> <ul> <li> <p>Downloaded Data</p> <ul> <li>Cofactor and Catalyst relations: <a href="https://www.uniprot.org/uniprot/?query=&fil=organism%3A%22Homo%20sapiens%20(Human)%20%5B9606%5D%22&columns=id%2Centry%20name%2Creviewed%2Cdatabase(PRO)%2Cchebi(Cofactor)%2Cchebi(Catalytic%20activity)"><code>Cofactor/Catalyst Query Results</code></a></li> <li>UniProt Identifier Mapping: <a href="https://www.uniprot.org/uniprot/?query=&fil=organism%3A%22Homo%20sapiens%20(Human)%20%5B9606%5D%22&columns=id%2Cdatabase(GeneID)%2Cdatabase(Ensembl)%2Cdatabase(HGNC)%2Cgenes(PREFERRED)%2Cgenes(ALTERNATIVE)"><code>UniProt Identifier Query Results</code></a></li> </ul> </li> <li> <p>Generated Data</p> <ul> <li>Merged Gene, RNA, Protein Map: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/Merged_gene_rna_protein_identifiers.pkl"><code>Merged_gene_rna_protein_identifiers.pkl</code></a></li> <li>Protein-Cofactor Relations: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/UNIPROT_PROTEIN_COFACTOR.txt"><code>UNIPROT_PROTEIN_COFACTOR.txt</code></a></li> <li>Protein-Catalyst Relations: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/UNIPROT_PROTEIN_CATALYST.txt"><code>UNIPROT_PROTEIN_CATALYST.txt</code></a></li> <li>UniProt Accession-PRO Identifier Mapping: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/UNIPROT_ACCESSION_PRO_ONTOLOGY_MAP.txt"><code>UNIPROT_ACCESSION_PRO_ONTOLOGY_MAP.txt</code></a></li> <li>UniProt Accession-Entrez Gene Identifier Mapping: <a href="https://storage.googleapis.com/pheknowlator/current_build/data/processed_data/UNIPROT_ACCESSION_ENTREZ_GENE_MAP.txt"><code>UNIPROT_ACCESSION_ENTREZ_GENE_MAP.txt</code></a></li> </ul> </li> </ul> <p> </p> <p>This project is licensed under Apache License 2.0 - see the <strong><a href="https://github.com/callahantiff/PheKnowLator/blob/master/LICENSE"><code>LICENSE.md</code></a></strong> file for details. If you intend to use any of the information on this Wiki, please provide the appropriate attribution by citing this repository:</p> <pre><code>@misc{callahan_tj_2019_3401437, author = {Callahan, TJ}, title = {PheKnowLator}, month = mar, year = 2019, doi = {10.5281/zenodo.3401437}, url = {https://doi.org/10.5281/zenodo.3401437} }</code></pre>
Demonstration Cases - Simulation data of energy consumption of residential building typologies
<p>The dataset is about the energy analysis for retrofit strategies of 5 building typologies and the EDEA project located in 3 climates zones in Europe: South (Madrid), Central (Berlin) and North (Helsinki).<br> The dataset includes:<br> (1) Open Document Spreadsheet (.ods) file with the results of Heating Consumption (kWh/m2·year) and Cooling Consumption (kWh/m2·year) for the five buildings, in three locations and for several scenarios:<br> - Locating external new insulation in walls and roof.<br> - Replacing Windows.<br> - Combination strategies: locating new insulation layers and replacing the existing windows.<br> - Installing solar protection devices.</p>
Three-dimensional building and mobility infrastructure of the CONUS
<p>Humanity's role in changing the face of the earth is a long-standing concern, as is the human domination of ecosystems. Geologists are debating the introduction of a new geological epoch, the 'anthropocene', as humans are 'overwhelming the great forces of nature'. In this context, the accumulation of artefacts, i.e., human-made physical objects, is a pervasive phenomenon. Variously dubbed 'manufactured capital', 'technomass', 'human-made mass', 'in-use stocks' or 'socioeconomic material stocks', they have become a major focus of sustainability sciences in the last decade. Globally, the mass of socioeconomic material stocks now exceeds 10e14 kg, which is roughly equal to the dry-matter equivalent of all biomass on earth. It is doubling roughly every 20 years, almost perfectly in line with 'real' (i.e. inflation-adjusted) GDP. In terms of mass, buildings and infrastructures (here collectively called 'built structures') represent the overwhelming majority of all socioeconomic material stocks.</p><p>This dataset features intermediate mapping results for estimating material stocks in the CONUS (see related identifiers) on a 10m grid based on high resolution Earth Observation data (Sentinel-1 + Sentinel-2), Microsoft building footprints, NLCD Impervious data, and crowd-sourced geodata (OSM). These data may also be useful on their own.</p><p><strong>Provided layers @10m resolution</strong><br>- Building height<br>- Building type<br>- Building area<br>- Impervious fraction<br>- street, and rail area<br>- Building and street climate zones<br>- County zones<br>- State masks<br>- EQUI7 correction factors</p><p><strong>Spatial extent</strong><br>This dataset covers the whole CONUS. </p><p><strong>Temporal extent</strong><br>The maps are representative for ca. 2018.</p><p><strong>Data format</strong><br>The data are organized in 100km x 100km tiles (EQUI7 grid), and mosaics are provided.</p><p><strong>Further information</strong><br>For further information, please see the main publication.<br>A web-visualization of the resulting dataset is available <a href="https://ows.geo.hu-berlin.de/webviewer/us-stocks/">here</a>.<br>Visit our <a href="https://boku.ac.at/understanding-the-role-of-material-stock-patterns-for-the-transformation-to-a-sustainable-society-mat-stocks">website</a> to learn more about our project MAT_STOCKS - Understanding the Role of Material Stock Patterns for the Transformation to a Sustainable Society.</p><p><strong>Publication</strong><br>D. Frantz, F. Schug, D. Wiedenhofer, A. Baumgart, D. Virág, S. Cooper, C. Gómez-Medina, F. Lehmann, T. Udelhoven, S. van der Linden, P. Hostert, and H. Haberl (2023): Unveiling patterns in human dominated landscapes through mapping the mass of US built structures. <i>Nature Communications</i> <strong>14</strong>, 8014. <a href="https://doi.org/10.1038/s41467-023-43755-5">https://doi.org/10.1038/s41467-023-43755-5</a></p><p><strong>Funding</strong><br>This research was primarly funded by the European Research Council (ERC) under the European Union's Horizon 2020 research and innovation programme (MAT_STOCKS, grant agreement No 741950). </p><p><strong>Acknowledgments</strong><br>We thank the European Space Agency and the European Commission for freely and openly sharing Sentinel imagery; USGS for the National Land Cover Database; Microsoft for Building Footprints; Geofabrik and all contributors for OpenStreetMap.This dataset was partly produced on <a href="https://eodc.eu/">EODC</a> - we thank Clement Atzberger for supporting the generation of this dataset by sharing disc space on EODC, and Wolfgang Wagner for granting access to preprocessed Sentinel-1 data.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.