Skip to main content
zenodoopen

GWAS summary statistics for 9 quantitative phenotypes from the UK Biobank (5-fold cross-validation)

<p>This dataset contains GWAS summary statistics for 9 quantitative phenotypes from the UK Biobank.</p> <p>The dataset is designed to enable systematic PRS analyses with 5-fold cross validation. For each phenotype and fold, we provide GWAS summary statistics for the training, validation, and test sets. The validation summary statistics can be used for model selection/tuning. The test summary statistics can be used to evaluate PRS models via pseudo-validation metrics. Association testing for all phenotypes and samples was done with <strong>plink2</strong>.</p> <p>&nbsp;</p> <p>The&nbsp;<strong>phenotypes</strong> included in this dataset are:</p> <ul> <li><a href="https://biobank.ndph.ox.ac.uk/showcase/field.cgi?id=50">HEIGHT</a>: Standing height (Data-Field: 50)</li> <li><a href="https://biobank.ndph.ox.ac.uk/showcase/field.cgi?id=21001">BMI</a>: Body mass index (Data-Field: 21001)</li> <li><a href="https://biobank.ndph.ox.ac.uk/showcase/field.cgi?id=48">WC</a>: Waist circumference (Data-Field: 48)</li> <li><a href="https://biobank.ndph.ox.ac.uk/showcase/field.cgi?id=49">HC</a>: Hip circumference (Data-Field: 49)</li> <li><a href="https://biobank.ndph.ox.ac.uk/showcase/field.cgi?id=20022">BW</a>: Birth weight (Data-Field: 20022)</li> <li><a href="https://biobank.ndph.ox.ac.uk/showcase/field.cgi?id=3062">FVC</a>: Forced vital capacity (Data-Field: 3062)</li> <li><a href="https://biobank.ndph.ox.ac.uk/showcase/field.cgi?id=3063">FEV1</a>: Forced expiratory volume in 1-second (Data-Field: 3063)</li> <li><a href="https://biobank.ndph.ox.ac.uk/showcase/field.cgi?id=30760">HDL</a>: HDL cholesterol (Data-Field: 30760)</li> <li><a href="https://biobank.ndph.ox.ac.uk/showcase/field.cgi?id=30780">LDL</a>: LDL cholesterol (Data-Field: 30780)</li> </ul> <p>&nbsp;</p> <p>To allow users to assess PRS performance as a function of sample size, we also provide <strong>subsampled training GWAS summary statistics</strong>. This is done by taking the training samples and randomly selecting (without replacement) a subset of them for conducting association testing. The training sample sizes are:</p> <ul> <li>N = 5000</li> <li>N = 10000</li> <li>N = 20000</li> <li>N = 40000</li> <li>N = 80000</li> <li>N = 160000</li> <li>Full training set (sample size varies by phenotype).</li> </ul> <p><strong>NOTE</strong>: Due to the smaller overall sample size for the Birth weight phenotype, we do not include training data for the `N=160000` setting.<br><br></p> <p>The <strong>folder structure</strong> of the GWAS data for each phenotype is as follows:</p> <ul> <li><code>train</code> <ul> <li><code>N_5000</code> <ul> <li><code>&nbsp;fold_1</code> <ul> <li><code>chr_1.PHENO1.glm.linear</code></li> <li><code>chr_2.PHENO1.glm.linear</code></li> <li><code>...</code></li> </ul> </li> <li><code>fold_2</code></li> <li><code>fold_3</code></li> <li><code>...</code></li> </ul> </li> <li><code>N_10000</code></li> <li><code>N_20000</code></li> <li><code>N_40000</code></li> <li><code>N_80000</code></li> <li><code>N_160000</code></li> <li><code>full</code></li> </ul> </li> <li><code>validation</code> <ul> <li><code>fold_1</code> <ul> <li><code>chr_1.PHENO1.glm.linear</code></li> <li><code>chr_2.PHENO1.glm.linear</code></li> <li><code>...</code></li> </ul> </li> <li><code>fold_2</code></li> <li><code>fold_3</code></li> <li><code>...</code></li> </ul> </li> <li><code>test</code> <ul> <li><code>fold_1</code></li> <li><code>fold_2</code></li> <li><code>fold_3</code></li> <li><code>...</code></li> </ul> </li> </ul> <p>For more details about the GWAS study, Quality Control (QC) criteria, or other information, please consult our publication:</p> <p>Zabad, S., Gravel, S., &amp; Li, Y. (2023).&nbsp;<strong>Fast and accurate Bayesian polygenic risk modeling with variational inference.</strong>&nbsp;The American Journal of Human Genetics, 110(5), 741&ndash;761.&nbsp;<a href="https://doi.org/10.1016/j.ajhg.2023.03.009" rel="nofollow">https://doi.org/10.1016/j.ajhg.2023.03.009</a></p> <p>If you use this data in your work, please cite the publication above.</p> <p>&nbsp;</p>

ShareScore

32/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
4
Harmonization
4
Access
16
Reuse readiness
8
Engagement
0