Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
25
datasets available to search
ShareScore release 0.9.0
Dataset results
25 results for “Ontology, Phenotype”
Gold standard corpus, ontologies, and Entity-Quality ontology annotations for evolutionary phenotypes
<p>This data set includes a gold-standard corpus of evolutionary phenotype descriptions (in the form of character state descriptions pulled from a variety of phylogenetic systematics studies), and their corresponding expert-curated annotations with ontology terms in the form of Entity-Quality (EQ) statements. EQ annotatons allow machine-reasoning (through the semantics encoded in the requisite ontologies from which the ontology terms are drawn), and machine-reasoning in turn enables computing metrics for quantifying the semantic similarity between different phenotype descriptions as represented by their EQ annotations.</p> <p>Also included are the ontologies, and the human expert-generated and Semantic Charaparser (i.e., machine) generated EQ annotations used to assess Semantic Charaparser performance relative to inter-curator variation and to the effect of having access to external knowledge. The ontologies include those used as input, the "augmented" ontologies created by human curators in each experiment round, and the merged ontology used to maximize Semantic Charaparser's performance.</p> <p>The production of the gold standard corpus, annotation experiments, and evaluation of the results are described in detail in the following manuscript:</p> <blockquote> <p>Dahdul et al (2018) Annotation of phenotypes using ontologies: a Gold Standard for the training and evaluation of natural language processing systems. BioRxiv https://doi.org/10.1101/322156. Submitted to Database.</p> </blockquote> <p>The analysis code for evaluating the gold standard corpus (and the input data and ontologies for that) are available separately from the following:</p> <blockquote> <p>Manda et al (2018) Code and data for analysis of evolutionary phenotype ontology annotations and gold standard corpus. Zenodo. https://doi.org/10.5281/zenodo.1218010</p> </blockquote> <p>In comparison to the previous version (v1.0.0), this record includes a file of MD5 checksums of the Gold Standard data files. The data files themselves are unchanged.</p>
Ontology based text mining of gene-phenotype associations: application to candidate gene prediction
<p>Gene-phenotype associations play an important role in understanding<br> the disease mechanisms which is a requirement for treatment<br> development. A portion of gene-phenotype associations are observed<br> mainly experimentally and made publicly available through several<br> standard resources such as MGI. However, there is still a vast<br> amount of gene--phenotype associations buried in the biomedical<br> literature. Given the large amount of literature data, we need<br> automated text mining tools to alleviate the burden in manual<br> curation of gene-phenotype associations and to develop<br> comprehensive resources. We developed an ontology based<br> approach in combination with statistical methods to text mine<br> gene-phenotype associations from literature. Our method achieved<br> AUC values of 0.90 and 0.75 in recovering known gene-phenotype<br> associations from HPO and MGI respectively. We posit that candidate<br> genes and their relevant diseases should be expressed with similar<br> phenotypes in publications. Thus, we demonstrate the utility of our<br> approach by predicting disease candidate genes based on the semantic<br> similarities of phenotypes associated with genes and diseases. We evaluated our disease candidate prediction model on<br> the gene-disease associations from MGI. Our model achieved AUC<br> values of 0.90 and 0.87 on OMIM (human) and MGI (mouse) datasets of<br> gene-disease associations respectively. Our manual analysis on the<br> text mined data revealed that, our method can accurately extract<br> gene-phenotype associations which are not currently covered by the<br> existing public gene-phenotype resources. Overall, results indicate<br> that our method can precisely extract known as well as new<br> gene-phenotype associations from literature. This released dataset at Zenodo covers our gene-phenotype extracts from the literature. All the methods used to extract the data are available at https://github.com/bio-ontology-research-group/genepheno.</p>
Mapping between Human Phenotype Ontology and phecode terminologies
Open the record for dataset details and reuse information.
Data and software associated with PHENOstruct: Prediction of human phenotype ontology terms using heterogeneous data sources
<p>Data and software associated with the paper:</p> <p>PHENOstruct: Prediction of human phenotype ontology terms using heterogeneous data sources</p>
Zebrafish Phenotype Ontology
<p>Current release of Zebrafish Phenotype Ontology</p>
The Unified Phenotype Ontology (uPheno) - Version 2, Alpha release
<p>This is the Alpha release of the re-designed Unified Phenotype Ontology (uPheno), referred to internally as uPheno 2. The effort responsible for the content of uPheno 2 is the <a href="https://github.com/obophenotype/upheno/wiki/Phenotype-Ontologies-Reconciliation-Effort">Phenotype Reconciliation Effort</a>. The pipeline that generates uPheno 2 can be found here: <a href="https://github.com/obophenotype/upheno-dev">https://github.com/obophenotype/upheno-dev</a>.</p>
Using ontologies to extract disease--phenotype associations from literature
<p>This dataset contains disease-phenotype associations. We developed and used a text-mining system which utilizes semantic relations in the phenotype ontologies and statistical methods to extract disease-phenotype associations from the literature.</p>
Results of Neuron Phenotype Ontology competency queries
<p>Full list of neurons with human readable labels returned from competency queries against the Neuron Phenotype Ontology as described in Gillespie et al. (2020). Three .CSV files are included with the number of competency query (Q1, Q2 and Q3) appended. The DL query used for the query is given in the column header of column A. The numbers provided for Q1 and Q2 in columns B-E represent the count of neurons returned for each of the 3 evidence based models (EBMs) described in the paper and the Common Usage Types (CUTs) The total number of neurons returned across all categories is shown in column F.</p>
Data from: Integration of anatomy ontologies and evo-devo using structured Markov models suggests a new framework for modeling discrete phenotypic traits
Modeling discrete phenotypic traits for either ancestral character state reconstruction or morphology-based phylogenetic inference suffers from ambiguities of character coding, homology assessment, dependencies, and selection of adequate models. These drawbacks occur because trait evolution is driven by two key processes – hierarchical and hidden – which are not accommodated simultaneously by the available phylogenetic methods. The hierarchical process refers to the dependencies between anatomical body parts, while the hidden process refers to the evolution of gene regulatory networks underlying trait development. Herein, I demonstrate that these processes can be efficiently modeled using structured Markov models equipped with hidden states, which resolves the majority of the problems associated with discrete traits. Integration of structured Markov models with anatomy ontologies can adequately incorporate the hierarchical dependencies, while the use of the hidden states accommodates hidden evolution of gene regulatory networks and substitution rate heterogeneity. I assess the new models using simulations and theoretical synthesis. The new approach solves the long-standing "tail color problem," in which the trait is scored for species with tails of different colors or no tails. It also presents a previously unknown issue called the "two-scientist paradox," in which the nature of coding the trait and the hidden processes driving the trait's evolution are confounded; failing to account for the hidden process may result in a bias, which can be avoided by using hidden state models. All this provides a clear guideline for coding traits into characters. This paper gives practical examples of using the new framework for phylogenetic inference and comparative analysis.
Human Phenotype Ontology: Standardized Framework Cataloging Comprehensive Phenotypic Information
<p><strong>ABSTRACT:</strong></p> <p>The Human Phenotype Ontology (HPO) is a standardized and comprehensive framework that catalogs and organizes phenotypic information related to human diseases and genetic disorders. It provides a structured vocabulary of terms to describe observable traits and abnormalities associated with various medical conditions. This dataset showcases the Human Phenotype Ontology (HPO) by including HPO IDs, names, descriptions, and additional references. The HPO IDs serve as unique identifiers for specific phenotypes. The names provide concise labels for the phenotypic traits, while the descriptions offer detailed explanations of their characteristics. In addition to the essential information, the dataset includes references that further describe and support the understanding of diverse phenotypes. These references can be used to explore the scientific literature, studies, or clinical resources associated with each phenotype. By combining HPO IDs, names, descriptions, and references, the dataset offers a comprehensive overview of different phenotypes, facilitating research, diagnostics, and the exploration of genotype-phenotype relationships.</p> <p><strong>Instructions</strong>:</p> <p>The data was downloaded as an .obo file then converted to a .csv file. Unrelated columns were removed and a references column was added. </p> <p><strong>Acknowledgements</strong>:</p> <p><a>Sebastian Köhler</a>, <a>Michael Gargano</a>, <a>Nicolas Matentzoglu</a>, <a>Leigh C Carmody</a>, <a>David Lewis-Smith</a>, <a>Nicole A Vasilevsky</a>, <a>Daniel Danis</a>, <a>Ganna Balagura</a>, <a>Gareth Baynam</a>, <a>Amy M Brower</a>, </p> <p><a>Tiffany J Callahan</a>, <a>Christopher G Chute</a>, <a>Johanna L Est</a>, <a>Peter D Galer</a>, <a>Shiva Ganesan</a>, <a>Matthias Griese</a>, <a>Matthias Haimel</a>, <a>Julia Pazmandi</a>, <a>Marc Hanauer</a>, <a>Nomi L Harris</a>, <a>Michael J Hartnett</a>, <a>Maximilian Hastreiter</a>, <a>Fabian Hauck</a>, <a>Yongqun He</a>, <a>Tim Jeske</a>, <a>Hugh Kearney</a>, <a>Gerhard Kindle</a>, <a>Christoph Klein</a>, <a>Katrin Knoflach</a>, <a>Roland Krause</a>, <a>David Lagorce</a>, <a>Julie A McMurry</a>, <a>Jillian A Miller</a>, <a>Monica C Munoz-Torres</a>, <a>Rebecca L Peters</a>, <a>Christina K Rapp</a>, <a>Ana M Rath</a>, <a>Shahmir A Rind</a>, <a>Avi Z Rosenberg</a>, <a>Michael M Segal</a>, <a>Markus G Seidel</a>, <a>Damian Smedley</a>, <a>Tomer Talmy</a>, <a>Yarlalu Thomas</a>, <a>Samuel A Wiafe</a>, <a>Julie Xian</a>, <a>Zafer Yüksel</a>, <a>Ingo Helbig</a>, <a>Christopher J Mungall</a>, <a>Melissa A Haendel</a>, <a>Peter N Robinson</a></p> <p><em>Nucleic Acids Research</em>, Volume 49, Issue D1, 8 January 2021, Pages D1207–D1217, <a href="https://doi.org/10.1093/nar/gkaa1043">https://doi.org/10.1093/nar/gkaa1043</a></p> <p><strong>Human Phenotype Ontology downloads</strong> <strong>page:</strong> https://hpo.jax.org/app/data/ontology</p> <p><strong>U-BRITE last update: </strong>7/6/23</p> <p> </p>
Data from: Integration of anatomy ontologies and evo-devo using structured Markov models suggests a new framework for modeling discrete phenotypic traits
Open the record for dataset details and reuse information.
Data from: Toward synthesizing our knowledge of morphology: using ontologies and machine reasoning to extract presence/absence evolutionary phenotypes across studies
The reality of larger and larger molecular databases and the need to integrate data scalably have presented a major challenge for the use of phenotypic data. Morphology is currently primarily described in discrete publications, entrenched in noncomputer readable text, and requires enormous investments of time and resources to integrate across large numbers of taxa and studies. Here we present a new methodology, using ontology-based reasoning systems working with the Phenoscape Knowledgebase (KB; kb.phenoscape.org), to automatically integrate large amounts of evolutionary character state descriptions into a synthetic character matrix of neomorphic (presence/absence) data. Using the KB, which includes more than 55 studies of sarcopterygian taxa, we generated a synthetic supermatrix of 639 variable characters scored for 1051 taxa, resulting in over 145,000 populated cells. Of these characters, over 76% were made variable through the addition of inferred presence/absence states derived by machine reasoning over the formal semantics of the source ontologies. Inferred data reduced the missing data in the variable character-subset from 98.5% to 78.2%. Machine reasoning also enables the isolation of conflicts in the data, that is, cells where both presence and absence are indicated; reports regarding conflicting data provenance can be generated automatically. Further, reasoning enables quantification and new visualizations of the data, here for example, allowing identification of character space that has been undersampled across the fin-to-limb transition. The approach and methods demonstrated here to compute synthetic presence/absence supermatrices are applicable to any taxonomic and phenotypic slice across the tree of life, providing the data are semantically annotated. Because such data can also be linked to model organism genetics through computational scoring of phenotypic similarity, they open a rich set of future research questions into phenotype-to-genome relationships.
Data from: A revision of Evaniscus (Hymenoptera, Evaniidae) using ontology-based semantic phenotype annotation
The Neotropical evaniid genus Evaniscus Szépligeti currently includes six species. Two new species are described, Evaniscus lansdownei Mullins, sp. n. from Colombia and Brazil and Evaniscus rafaeli Kawada, sp. n. from Brazil. Evaniscus sulcigenis Roman, syn. n., is synonymized under Evaniscus rufithorax Enderlein. An identification key to species of Evaniscus is provided. Thirty-five parsimony informative morphological characters are analyzed for six ingroup and four outgroup taxa. A topology resulting in a monophyletic Evaniscus is presented with Evaniscus tibialis and Evaniscus rafaeli as sister to the remaining Evaniscus species. The Hymenoptera Anatomy Ontology and other relevant biomedical ontologies are employed to create semantic phenotype statements in Entity-Quality (EQ) format for species descriptions. This approach is an early effort to formalize species descriptions and to make descriptive data available to other domains.
Figure 31 from: Mullins P, Kawada R, Balhoff J, Deans A (2012) A revision of Evaniscus (Hymenoptera, Evaniidae) using ontology-based semantic phenotype annotation. ZooKeys 223: 1-38. https://doi.org/10.3897/zookeys.223.3572
Figure 31 - Most parsimonious tree from exhaustive search in PAUP*. Numbers above nodes show bootstrap support and numbers below nodes show jackknife support from the maximum parsimony analysis.
Figures 13-18 from: Mullins P, Kawada R, Balhoff J, Deans A (2012) A revision of Evaniscus (Hymenoptera, Evaniidae) using ontology-based semantic phenotype annotation. ZooKeys 223: 1-38. https://doi.org/10.3897/zookeys.223.3572
Figures 13-18 - Brightfield images of Evaniscus rufithorax Enderlein. 13, 14 Lateral habitus 15, 16 Dorsal habitus 17 Anterior oblique 18 Anterior face.
Figures 7-12 from: Mullins P, Kawada R, Balhoff J, Deans A (2012) A revision of Evaniscus (Hymenoptera, Evaniidae) using ontology-based semantic phenotype annotation. ZooKeys 223: 1-38. https://doi.org/10.3897/zookeys.223.3572
Figures 7-12 - Brightfield images of Evaniscus rafaeli Kawada sp. n. 7, 8 Lateral habitus 9, 10 Dorsal habitus 11 Anterior oblique 12 Anterior face.
Figures 1-6 from: Mullins P, Kawada R, Balhoff J, Deans A (2012) A revision of Evaniscus (Hymenoptera, Evaniidae) using ontology-based semantic phenotype annotation. ZooKeys 223: 1-38. https://doi.org/10.3897/zookeys.223.3572
Figures 1-6 - Brightfield images of Evaniscus lansdownei Mullins sp. n. 1, 2 Lateral habitus 3, 4 Dorsal habitus 5 Anterior oblique 6 Anterior face.
Figures 25-30 from: Mullins P, Kawada R, Balhoff J, Deans A (2012) A revision of Evaniscus (Hymenoptera, Evaniidae) using ontology-based semantic phenotype annotation. ZooKeys 223: 1-38. https://doi.org/10.3897/zookeys.223.3572
Figures 25-30 - Brightfield images of Evaniscus tibialis Szépligeti. 25, 26 Lateral habitus 27, 28 Dorsal habitus 29 Anterior oblique 30 Anterior face.
Figures 32-33 from: Mullins P, Kawada R, Balhoff J, Deans A (2012) A revision of Evaniscus (Hymenoptera, Evaniidae) using ontology-based semantic phenotype annotation. ZooKeys 223: 1-38. https://doi.org/10.3897/zookeys.223.3572
Figures 32-33 - Brightfield images of Evaniscus rufithorax . 32 Male specimen; arrow points to visible lower tubular sclerite 33 Female specimen; arrow shows where lower tubular sclerite is not visible.
Figures 19-24 from: Mullins P, Kawada R, Balhoff J, Deans A (2012) A revision of Evaniscus (Hymenoptera, Evaniidae) using ontology-based semantic phenotype annotation. ZooKeys 223: 1-38. https://doi.org/10.3897/zookeys.223.3572
Figures 19-24 - Brightfield images of Evaniscus marginatus Cameron. 19, 20 Lateral habitus 21, 22 Dorsal habitus 23 Anterior oblique 24 Anterior face.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.