Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
152
datasets available to search
ShareScore release 0.9.0
Dataset results
152 results for “polygenes”
Data from: Recent natural selection causes adaptive evolution of an avian polygenic trait
We used extensive data from a long-term study of great tits (Parus major) in the United Kingdom and Netherlands to better understand how genetic signatures of selection translate into variation in fitness and phenotypes. We found that genomic regions under differential selection contained candidate genes for bill morphology and used genetic architecture analyses to confirm that these genes, especially the collagen gene COL4A5, explained variation in bill length. COL4A5 variation was associated with reproductive success, which, combined with spatiotemporal patterns of bill length, suggested ongoing selection for longer bills in the United Kingdom. Last, bill length and COL4A5 variation were associated with usage of feeders, suggesting that longer bills may have evolved in the United Kingdom as a response to supplementary feeding.
Data from: Genetic redundancy fuels polygenic adaptation in Drosophila
The genetic architecture of adaptive traits is of key importance to predict evolutionary responses. Most adaptive traits are polygenic—i.e., result from selection on a large number of genetic loci—but most molecularly characterized traits have a simple genetic basis. This discrepancy is best explained by the difficulty in detecting small allele frequency changes (AFCs) across many contributing loci. To resolve this, we use laboratory natural selection to detect signatures for selective sweeps and polygenic adaptation. We exposed 10 replicates of a Drosophila simulans population to a new temperature regime and uncovered a polygenic architecture of an adaptive trait with high genetic redundancy among beneficial alleles. We observed convergent responses for several phenotypes—e.g., fitness, metabolic rate, and fat content—and a strong polygenic response (99 selected alleles; mean s = 0.059). However, each of these selected alleles increased in frequency only in a subset of the evolving replicates. We discerned different evolutionary paradigms based on the heterogeneous genomic patterns among replicates. Redundancy and quantitative trait (QT) paradigms fitted the experimental data better than simulations assuming independent selective sweeps. Our results show that natural D. simulans populations harbor a vast reservoir of adaptive variation facilitating rapid evolutionary responses using multiple alternative genetic pathways converging at a new phenotypic optimum. This key property of beneficial alleles requires the modification of testing strategies in natural populations beyond the search for convergence on the molecular level.
Integrative genomics reveals the polygenic basis of seedless in grapevine (Vitis vinifera L.)
<p>The publicly accessible genome data comprises:</p> <p>1) Haplotype genomes of 'Thompson Seedless' (TS) and 'Black Monukka' (BM) (.fa files) along with annotations (.gff files),</p> <p>2) Cytoplasmic genomes (mitochondria and chloroplasts),</p> <p>3) Annotations for tandem repeats,</p> <p>4) EDTA annotations (available at https://github.com/oushujun/EDTA),</p> <p>5) panTE (Transposable element) annotations (available at https://github.com/unavailable-2374/TE_Detective-Annotation).</p> <p>Additionally,</p> <p>6) Several genomes, such as 'Cabernet Sauvignon', 'Black Corinth Seedless', and 'Black Corinth Seeded', were restructured based on PN_T2T genome similarity using RagTag (https://github.com/malonge/RagTag).</p> <p>Among these resources,</p> <p>7) The 'Cabernet Sauvignon' genome was annotated using PN_T2T annotation through Liftoff (https://github.com/agshumate/Liftoff), and it maintains consistent gene IDs between the 'Cabernet Sauvignon' genome and PN_T2T.</p> <p>Furthermore, we have retained the original data for the protein files of 'Black Corinth Seedless' and 'Black Corinth Seeded'.<br> <br> If you have any questions, please let me know: 571720850@qq.com</p>
Barcoded Bulk QTL mapping reveals highly polygenic and epistatic architecture of complex traits in yeast
<p>Mapping the genetic basis of complex traits is critical to uncovering the biological mechanisms that underlie disease and other phenotypes. Genome-wide association studies (GWAS) in humans and quantitative trait locus (QTL) mapping in model organisms can now explain much of the observed heritability in many traits, allowing us to predict phenotype from genotype. However, constraints on power due to statistical confounders in large GWAS and smaller sample sizes in QTL studies still limit our ability to resolve numerous small-effect variants, map them to causal genes, identify pleiotropic effects across multiple traits, and infer non-additive interactions between loci (epistasis). Here, we introduce barcoded bulk quantitative trait locus (BB-QTL) mapping, which allows us to construct, genotype, and phenotype 100,000 offspring of a budding yeast cross, two orders of magnitude larger than the previous state of the art. We use this panel to map the genetic basis of eighteen complex traits, finding that the genetic architecture of these traits involves hundreds of small-effect loci densely spaced throughout the genome, many with widespread pleiotropic effects across multiple traits. Epistasis plays a central role, with thousands of interactions that provide insight into genetic networks. By dramatically increasing sample size, BB-QTL mapping demonstrates the potential of natural variants in high-powered QTL studies to reveal the highly polygenic, pleiotropic, and epistatic architecture of complex traits.</p>
Incorporating family history of disease improves polygenic risk scores in diverse populations
<p>Code relevant to Hujoel et al. "Incorporating family history of disease improves polygenic risk scores in diverse populations"</p>
Supplemental data for: Development and validation of a polygenic risk score for stroke in the Chinese population
<div class="WordSection1"> <div class="WordSection1"> <strong>Objective</strong>: To construct a polygenic risk score (PRS) for stroke and evaluate its utility in risk stratification and primary prevention for stroke.</div> <div class="WordSection1"> </div> <div class="WordSection1"> <strong>Methods</strong>: Using meta-analytic approach and large genome-wide association results for stroke and stroke-related traits in East Asians, we generated a combined PRS (metaPRS) by incorporating 534 genetic variants in a training set of 2,872 patients with stroke and 2,494 controls. We then validated its association with incident stroke using Cox regression models in large Chinese population-based prospective cohorts comprising 41,006 individuals.</div> <div class="WordSection1"> </div> <div class="WordSection1"> <strong>Results</strong>: During a total of 367,750 person-years (mean follow-up 9.0 years), 1,227 participants developed stroke before age of 80 years. Individuals with high polygenic risk had an about 2-fold higher risk of incident stroke compared with those with low polygenic risk (HR: 1.99, 95% CI: 1.66-2.38), with the lifetime risk of stroke being 25.2% (95% CI: 22.5%-27.7%) and 13.6% (95% CI: 11.6%-15.5%), respectively. Individuals with both high polygenic risk and family history displayed the lifetime risk as high as 41.1% (95% CI: 31.4%-49.5%). Moreover, individuals with high polygenic risk achieved greater benefits in terms of absolute risk reductions from adherence to ideal fasting blood glucose and total cholesterol than those with low polygenic risk. Maintaining favorable cardiovascular health (CVH) profile could substantially mitigate the increased risk conferred by high polygenic risk to the level of the low polygenic risk (from 34.6 % to 13.2%).</div> <div class="WordSection1"> </div> <div class="WordSection1"> <strong>Conclusions</strong>: Our metaPRS has great potential for risk stratification of stroke and identification of individuals who may benefit more from maintaining ideal CVH. </div> <div class="WordSection1"> </div> <div class="WordSection1"> <strong>Classification of Evidence</strong>: This study provides Class I evidence that a meta-polygenic risk score is predictive of stroke risk.</div> <p> </p> </div>
Data from: Genome-wide Polygenic Risk Scores Predict Risk of Glioma and Molecular Subtypes
<div> <div> <div> <p><strong>Background</strong>: Polygenic risk scores (PRS) aggregate the contribution of many risk variants to provide a personalized genetic susceptibility profile. Since sample sizes of glioma genome-wide association studies (GWAS) remain modest, there is a need to efficiently capture genetic risk using available data.</p> <p><strong>Methods</strong>: We applied a method based on continuous shrinkage priors (PRS-CS) to model the joint effects of over 1 million common variants on disease risk and compared this to an approach (PRS-CT) that only selects a limited set of independent variants that reach genome-wide significance (P<5×10-8). PRS models were trained using GWAS stratified by histological (10,346 cases, 14,687 controls) and molecular subtype (2,632 cases, 2,445 controls), and validated in two independent cohorts.</p> <p><strong>Results</strong>: PRS-CS was generally more predictive than PRS-CT with a median increase in explained variance (R2) of 24% (interquartile range=11-30%) across glioma subtypes. Improvements were pronounced for glioblastoma (GBM), with PRS-CS yielding larger odds ratios (OR) per standard deviation (OR=1.93, P=2.0×10-54 vs. OR=1.83, P=9.4×10-50) and higher explained variance (R2=2.82% vs. R2=2.56%). Individuals in the 80th percentile of the PRS- CS distribution had significantly higher risk of GBM (0.107%) at age 60 compared to those with average PRS (0.046%, P=2.4×10-12). Lifetime absolute risk reached 1.18% for glioma and 0.76% for IDH wildtype tumors for individuals in the 95th PRS percentile. PRS-CS augmented the classification of IDH mutation status in cases when added to demographic factors (AUC=0.839 vs. AUC=0.895, P=6.8×10-9).</p> <p><strong>Conclusions</strong>: Genome-wide PRS has potential to enhance the detection of high-risk individuals and help distinguish between prognostic glioma subtypes.</p> <p><strong>Citation</strong>: Nakase T, Guerra GA, Ostrom QT, et al. Genome-wide Polygenic Risk Scores Predict Risk of Glioma and Molecular Subtypes. <em>Neuro-Oncology</em>. Published online June 25, 2024:noae112. doi:10.1093/neuonc/noae112</p> </div> </div> </div>
Source data for manuscript "Polygenic burden of short tandem repeat expansions promote risk of Alzheimer's disease"
<p>Included are the source data and scripts used to make all plots for the mansucript "Polygenic burden of short tandem repeat expansions promote risk of Alzheimer's disease"</p>
Data from: The Red Death meets the abdominal bristle: polygenic mutation for susceptibility to a bacterial pathogen in Caenorhabditis elegans
Understanding the genetic basis of susceptibility to pathogens is an important goal of medicine and of evolutionary biology. A key first step toward understanding the genetics and evolution of any phenotypic trait is characterizing the role of mutation. However, the rate at which mutation introduces genetic variance for pathogen susceptibility in any organism is essentially unknown. Here we quantify the per-generation input of genetic variance by mutation (VM) for susceptibility of Caenorhabditis elegans to the pathogenic bacterium Pseudomonas aeruginosa (defined as the median time of death, LT50). VM for LT50 is slightly less than VM for a variety of life-history and morphological traits in this strain of C. elegans, but is well within the range of reported values in a variety of organisms. Mean LT50 did not change significantly over 250 generations of mutation accumulation. Comparison of VM to the standing genetic variance (VG) implies a strength of selection against new mutations of a few tenths of a percent. These results suggest that the substantial standing genetic variation for susceptibility of C. elegans to P. aeruginosa can be explained by polygenic mutation coupled with purifying selection.
Genome-wide association study of polygenic risk score-defined phenotype suffers from inflated test-statistics
<p>Simulation results from running the following script 100 times: https://github.com/euffelmann/paper-ad_prs_extremes/blob/main/scripts/ad_prs_extremes_simulation.R.</p> <p>These files can be used to reproduce tables and figures in: https://github.com/euffelmann/paper-ad_prs_extremes</p>
Polygenic risk score in Africa populations: progress and challenges
<p>Polygenic Risk Score (PRS) analysis is a method that predicts the genetic risk of an individual towards targeted traits. Even when there are no significant markers, it gives evidence of a genetic effect beyond the results of Genome-Wide Association Studies (GWAS). Moreover, it selects SNPs that contribute to the disease with low effect size making it more precise at individual level risk prediction. PRS analysis addresses the shortfall of GWAS by taking into account the SNPs/alleles with low effect size but play an indispensable role to the observed phenotypic/trait variance. PRS analysis has application which investigate the genetic basis of several traits which includes rare diseases. However, the accuracy of PRS analysis depends on the genomic data of the underlying population. For instance, several studies show that obtaining higher prediction power of PRS analysis is challenging for non-Europeans. In this manuscript, we reviewed the conventional PRS methods and their application to Sub-Saharan African communities. We concluded that lack of sufficient GWAS data and tools is the limiting factor of applying PRS analysis to Sub-Saharan populations. We recommend developing Africa-specific PRS methods and tools for estimating, and analyzing Africa population data for clinical evaluation of PRSs of interest and predicting rare diseases.</p>
Application of Polygenic Methylation Markers in Postoperative Recurrence Monitoring of Colorectal Cancer
ClinicalTrials.gov study NCT05444491. IPD Sharing: NO. Countries: 1. Publications: 2.
Polygenic Risk Score to Predict Weight Loss Intervention in Children With Obesity
ClinicalTrials.gov study NCT05466097. IPD Sharing: UNDECIDED. Countries: 1. Publications: 13.
Polygen Defi-Alpha: Genetic Polymorphisms Study in Children With Alpha-1 Antitrypsin Deficiency, Included in the DEFI-ALPHA Cohort
ClinicalTrials.gov study NCT01862211. IPD Sharing: Not stated. Countries: 1. Publications: 2.
Food Supplementation With Eufortyn Colesterolo Plus for LDL Modulation in Subjects With Polygenic Hypercholesterolemia
ClinicalTrials.gov study NCT04574505. IPD Sharing: NO. Countries: 1. Publications: 1.
EValuation Of poLygenic Scores and CT imAging In Risk Factor Modification in Patients With diabEtes
ClinicalTrials.gov study NCT07091162. IPD Sharing: NO. Countries: 1. Publications: 0.
Polygenic Risk Score Implementation and Stratification for Managing Blood Pressure
ClinicalTrials.gov study NCT06962488. IPD Sharing: UNDECIDED. Countries: 1. Publications: 2.
Polygenic Risk-based Detection of Subclinical Coronary Atherosclerosis and Intervention With Statin and Colchicine
ClinicalTrials.gov study NCT05850091. IPD Sharing: YES. Countries: 1. Publications: 0.
High Polygenic Risk and Health Behavior
ClinicalTrials.gov study NCT05603663. IPD Sharing: NO. Countries: 1. Publications: 0.
PheWAS of a Polygenic Predictor of Thyroid Function
ClinicalTrials.gov study NCT03597659. IPD Sharing: UNDECIDED. Countries: 1. Publications: 1.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.