Skip to main content
zenodoopen

Datasets used in the benchmarking study of MR methods

<p>We conducted a benchmarking analysis of 16 summary-level data-based MR methods for causal inference with five real-world genetic datasets, focusing on three key aspects: type I error control, the accuracy of causal effect estimates, replicability, and power.</p> <p>The datasets used in the MR benchmarking study can be downloaded here:</p> <ol> <li>"dataset-GWASATLAS-negativecontrol.zip":&nbsp; the GWASATLAS dataset for evaluation of type I error control in confounding scenario (a): Population stratification</li> <li>"dataset-NealeLab-negativecontrol.zip": the Neale Lab dataset for evaluation of type I error control in confounding scenario (a): Population stratification;</li> <li>"dataset-PanUKBB-negativecontrol.zip": the Pan UKBB dataset for evaluation of type I error control in confounding scenario (a): Population stratification;</li> <li>"dataset-Pleiotropy-negativecontrol": the dataset&nbsp; used for evaluation of type I error control in confounding scenario (b): Pleiotropy;</li> <li>"dataset-familylevelconf-negativecontrol.zip": the dataset used for evaluation of type I error control in confounding scenario (c): Family-level confounders;</li> <li>"dataset_ukb-ukb.zip": the dataset used for evaluation of the accuracy of causal effect estimates;</li> <li>"dataset-LDL-CAD_clumped.zip": the dataset used for evaluation of replicability and power;</li> </ol> <p>Each of the datasets contains the following files:</p> <ol> <li>&nbsp;"Tested Trait pairs": the exposure-outcome trait pairs to be analyzed;</li> <li>"MRdat" refers to the summary statistics after performing IV selection (p-value &lt; 5e-05) and PLINK LD clumping with a clumping window size of 1000kb and an r^2 threshold of 0.001.</li> <li>"bg_paras" are the estimated background parameters "Omega" and "C" which will be used for MR estimation in MR-APSS.</li> </ol> <p>Note:</p> <ol> <li>The formatted dataset after quality control can be accessible at our GitHub website (https://github.com/YangLabHKUST/MRbenchmarking).</li> <li>The details on quality control of GWAS summary statistics, formatting GWASs, and LD clumping for IV selection can be found on the MR-APSS software tutorial on the MR-APSS&nbsp;&nbsp;website (https://github.com/YangLabHKUST/MR-APSS).</li> <li>R code for running MR methods is also available at https://github.com/YangLabHKUST/MRbenchmarking.</li> </ol>

ShareScore

44/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
12
Harmonization
8
Access
16
Reuse readiness
8
Engagement
0

Topics