Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
270
datasets available to search
ShareScore release 0.9.0
Dataset results
270 results for “Disease phenotype”
Data from : From early development to maturity: A phenotypic analysis of the Townes sickle cell disease mice
Open the record for dataset details and reuse information.
DYRK1a inhibitor mediated rescue of Drosophila models of Alzheimer’s disease-Down Syndrome phenotypes
Open the record for dataset details and reuse information.
Using ontologies to extract disease--phenotype associations from literature
<p>This dataset contains disease-phenotype associations. We developed and used a text-mining system which utilizes semantic relations in the phenotype ontologies and statistical methods to extract disease-phenotype associations from the literature.</p>
Data from: ApoE is a correlate of phenotypic heterogeneity in Alzheimer's disease in a national cohort
Objective: To compare the proportion of APOEε4 genotype carriers in aphasic versus amnestic variants of Alzheimer's disease (AD). Method: The proportion of APOEε4 carriers was compared among 3 groups. 1) Forty-two patients with primary progressive aphasia (PPA) and AD pathology (PPA/AD) enrolled in the Northwestern Alzheimer Disease Center Clinical Core. 2) 1,418 patients with autopsy confirmed AD and amnestic dementia of the Alzheimer-type (DAT/AD); 3) 2,608 cognitively normal controls (NC). The latter two groups were compiled from the National Alzheimer Coordinating Center (NACC) database. Logistic regression models analyzed the relationship between groups and APOEε4 carrier status, adjusting for age of onset and sex as needed. Results: Using NC as the reference and adjusting for sex and age, the DAT/AD group was 3.97 times more likely to be APOEε4 carriers. Adjusting for sex and age at symptom onset, the DAT/AD group was 2.46 times as likely to be carriers compared to PPA/AD. There was no significant difference in the proportion of APOEε4 carriers for PPA/AD compared to NC. PPA subtypes included 24 logopenic, 10 agrammatic nonfluent, and eight either mixed (n=5) or too severe (n=3) to subtype. The proportion of carriers and non carriers was similar for logopenic and agrammatic subtypes, both having fewer carriers. Conclusion: The proportion of APOEε4 carriers was elevated in amnestic but not aphasic manifestations of AD. These results suggest that APOEε4 is an anatomically selective risk factor that preferentially increases the vulnerability to AD pathology of memory-related medial temporal areas rather that language-related neocortices.
Data associated with 'Metformin rescues Parkinson's disease phenotypes caused by hyperactive mitochondria'
<p>Metabolic dysfunction occurs in many age-related neurodegenerative diseases, yet its role in disease etiology remains poorly understood. We recently discovered a potential causal link between the branched-chain amino acid transferase, <i>BCAT-1,</i> and the neurodegenerative movement disorder, Parkinson's disease (PD). RNAi-mediated knockdown of <i>C. elegans bcat-1</i> recapitulates PD-like features, including progressive motor deficits and neurodegeneration with age, yet the underlying mechanisms have remained unknown. Using transcriptomic, metabolomic, and imaging approaches, we show here that <i>bcat-1 </i>knockdown increases mitochondrial respiration and induces oxidative damage in neurons through mTOR-independent mechanisms. Increased mitochondrial respiration, or 'mitochondrial hyperactivity,' is required for <i>bcat-1(RNAi)</i> neurotoxicity. Moreover, we show that post-disease onset administration of the type 2 diabetes medication, metformin, reduces mitochondrial respiration to control levels and significantly improves both motor function and neuronal viability. Together, our findings suggest that mitochondrial hyperactivity may be an early event in PD pathogenesis, and strategies aimed at reducing mitochondrial respiration may constitute a surprising new avenue for PD treatment.</p>
Natural history, phenotypic spectrum, and discriminative features of multisystemic RFC1-disease
<p>Objective: To delineate the full phenotypic spectrum, discriminative features, piloting longitudinal progression data, and sample size calculations of RFC1-repeat expansions, recently identified as causing cerebellar ataxia, neuropathy, vestibular areflexia syndrome (CANVAS).</p> <p>Methods: Multimodal RFC1 repeat screening (PCR, southern blot, whole-exome/genome (WES/WGS)-based approaches) combined with cross-sectional and longitudinal deep-phenotyping in (i) cross-European cohort A (70 families) with ≥2 features of CANVAS and/or ataxia-with-chronic-cough (ACC); and (ii) Turkish cohort B (105 families) with unselected late-onset ataxia.</p> <p>Results: Prevalence of RFC1-disease was 67% in cohort A, 14% in unselected cohort B, 68% in clinical CANVAS, and 100% in ACC. RFC1-disease was also identified in Western and Eastern Asians, and even by WES. Visual compensation, sensory symptoms, and cough were strong positive discriminative predictors (>90%) against RFC1-negative patients. The phenotype across 70 RFC1-positive patients was mostly multisystemic (69%), including dysautonomia (62%) and bradykinesia (28%) (=overlap with cerebellar-type multiple system atrophy [MSA-C]), postural instability (49%), slow vertical saccades (17%), and chorea and/or dystonia (11%). Ataxia progression was ~1.3 SARA points/year (32 cross-sectional, 17 longitudinal assessments, follow-up ≤9 years [mean 3.1]), but also included early falls, variable non-linear phases of MSA-C-like progression (SARA 2.5-5.5/year), and premature death. Treatment trials require 330 (1-year-trial) and 132 (2-year-trial) patients in total to detect 50% reduced progression.</p> <p>Conclusions: RFC1-disease is frequent and occurs across continents, with CANVAS and ACC as highly diagnostic phenotypes, yet as variable, overlapping clusters along a continuous multisystemic disease spectrum, including MSA-C-overlap. Our natural history data help to inform future RFC1-treatment trials.</p>
PheneBank: Processed Medline Abstracts and PMC full articles + Phenotype-Disease Associations
<p><strong>The PheneBank project:</strong></p> <p>Free text scientific literature has the potential to be an incredibly valuable source of data for uncovering the often hidden relationships between genes, diseases and phenotypes. Phenotypic descriptions cover abnormalities in anatomical structures, processes and behaviours. For example 'growth delay' and 'body weight loss'. Such descriptions form the basis for determining the existence and treatment of a disease but, because of their inherent complexity, have previously received less attention by the text mining community. In recent years, significant effort has been spent by a small number of expert curators to create coding systems for phenotypes (called "ontologies"), such as the Human Phenotype Ontology (HP) and the Mammalian Phenotype Ontology (MP). The PheneBank project proposes to support and speed up curation using terms discovered directly from the literature and to automatically integrate them with such standard ontologies. <br> <br> The project seeks to harness texts for extracting statistically significant associations between phenotypes, diseases and genes. Earlier approaches have suffered from not providing deep semantic representations of the phenotypes they tried to target. Our deep learning-based approach is an attempt to overcome this issue by reducing the uncertainty between textual and ontological forms of phenotypes. Specifically, the model treats multitoken named entities as a single token which allows more reliable handling of multiword expressions. The approach builds on ground breaking research at the European Bininformatics Institute by the PI (Nigel Collier) and the Co-investigator (Damian Smedley, Queen Mary University London), including terminology alignment of phenotypes using pairwise scoring of the conceptual elements that make up the phenotype. </p> <p><a href="http://www.phenebank.org">http://www.phenebank.org</a></p> <p><br> <strong>The dataset:</strong></p> <p>As an output of the PheneBank project, we release the set of 24 million MEDLINE abstracts as well as 3.8M open-access PMC full articles annotated with 9 classes of entity: Phenotype, Disease, Anatomy, Cell, Cell_line, GPR, Gene_variant, Molecule, and Pathway. The entities have been mapped to five major ontologies: SNOMED, HPO, MeSH, PRO, and FMA.</p> <p>In addition, we release the phenotype-disease associations that are automatically extracted based on co-occurrences statistics in Medline abstracts. Among different statistical measures we evaluated, the Fisher test best corresponded to the known tuples available from the curated associations available from the Monarch Initiative (https://monarchinitiative.org).</p> <p><br> <strong>Processing:</strong></p> <p>The NER tagging has been done using a BiLSTM-CRF neural model (<a href="https://github.com/pilehvar/phenebank">https://github.com/pilehvar/phenebank</a>) trained on expert-annotated data (to be released for research). The grounding to ontologies relies on semantic embedding of concepts and entities in a unified semantic space.</p> <p><br> <strong>Data format:</strong></p> <p><strong>PheneBank_Processed_PubMed.part[x].tar.gz </strong>contains 24,359,010 .txt files that are classified into 812 directories. Each .txt file is named with a PubMed article ID and contains the corresponding article's abstract and its annotations. The dataset is split into four (unequal) parts based on PubMed's structure:<br> part1: medline16n00* medline16n01* medline16n02* [299 directories, 2.8GB]<br> part2: medline16n03* medline16n04* [200 directories, 4.7GB]<br> part3: medline16n05* medline16n06* [200 directories, 5.3GB]<br> part4: medline16n07* medline16n08* [113 directories, 3.1GB]</p> <p>The <strong>PheneBank_Processed_PMC.tar.gz</strong> files has 6,180 directories which are named after the journal titles from which the articles have been drawn. There are three files per each article (i.e., 3 .txt files for the 3,751,770 distinct articles), containing text from different parts of the article: .title.txt, .abstract.txt, and .body.txt. </p> <p>Each line starts with a word; for those words that are identified as entities, entity type and mapping information are followed in the same line (tab separated), with the following format:</p> <p>word <TAB> ::: <TAB> entity_type <TAB> entity_concept_ID_1##confidence_score_1 entity_concept_ID_2##confidence_score_2 ...</p> <p>Note that the concepts are sorted according to their mapping confidence scores.</p> <p><br> As for the <strong>PheneBank_Associations.tsv</strong> file, there are ten columns that correspond to the following (left to right):</p> <p>- Disease Name<br> - Disease (MONDO) ID<br> - Phenotype Name<br> - Phenotype (HPO) ID<br> - Co-occurrence Frequency<br> - Disease Frequency<br> - Phenotype Frequency<br> - Fisher (log)<br> - Dice<br> - Normalized PMI</p> <p> </p> <p> </p>
Dynamic loading of human engineered heart tissue enhances contractile function and drives a desmosome-linked disease phenotype (TEM data)
<p>This is the TEM imaging data for the desmosome analysis as reported in the manuscript titled "Dynamic loading of human engineered heart tissue enhances contractile function and drives a desmosome-linked disease phenotype."</p>
Study of T Cell Phenotype Activation Pathway in Human Alcoholic Liver Disease
ClinicalTrials.gov study NCT00610597. IPD Sharing: Not stated. Countries: 1. Publications: 1.
Phenotyping Disease Severity in Asthma
ClinicalTrials.gov study NCT05078021. IPD Sharing: NO. Countries: 1. Publications: 18.
Towards Understanding the Phenotype of Cardiovascular Disease in CKD - TRUE-Type-CKD Study
ClinicalTrials.gov study NCT03749551. IPD Sharing: NO. Countries: 1. Publications: 4.
Phenotypes and Vascular Damage in Chronic Obstructive Pulmonary Disease (COPD)
ClinicalTrials.gov study NCT01527773. IPD Sharing: Not stated. Countries: 1. Publications: 8.
Peri-implant Phenotype, Calprotectin and Mmp-8 Levels in Cases Diagnosed With Peri-implant Disease
ClinicalTrials.gov study NCT06173739. IPD Sharing: NO. Countries: 1. Publications: 1.
Phenotypic and Genotypic Characterization of Malassezia Species Isolated From Malassezia Associated Skin Diseases
ClinicalTrials.gov study NCT05476731. IPD Sharing: UNDECIDED. Countries: 1. Publications: 4.
Investigating a Von Willebrand Factor (VWF) Functional Screening Assay for Assigning the Phenotypic Variants of Von Willebrand Disease (VWD)
ClinicalTrials.gov study NCT02466789. IPD Sharing: Not stated. Countries: 1. Publications: 10.
Phenotyping the Chronic Respiratory Diseases (CRD) in Ho Chi Minh City, Vietnam
ClinicalTrials.gov study NCT02517983. IPD Sharing: Not stated. Countries: 1. Publications: 4.
Superinfection and Hyperinflammatory Phenotype in COVID-19 (Coronavirus Disease 2019) Pneumonia Patients
ClinicalTrials.gov study NCT04867161. IPD Sharing: UNDECIDED. Countries: 1. Publications: 5.
Relationship Between Metabolic Profile and Clinical Phenotype in Chronic Obstructive Pulmonary Disease
ClinicalTrials.gov study NCT03310177. IPD Sharing: Not stated. Countries: 1. Publications: 1.
Molecular Bases of Response to Copper Treatment in Menkes Disease, Related Phenotypes, and Unexplained Copper Deficiency
ClinicalTrials.gov study NCT00811785. IPD Sharing: Not stated. Countries: 1. Publications: 3.
Genetic Markers as Predictors of Phenotypes in Pediatric Onset Crohn's Disease
ClinicalTrials.gov study NCT00783575. IPD Sharing: Not stated. Countries: 1. Publications: 23.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.