Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

260

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

260 results for “Biobank”

Learn how ShareScore rates datasets ↗
zenodo48/100

Genome-wide association study suggests that variation at the RCOR1 locus is associated with tinnitus in UK Biobank

<p>The dataset contains results of a genome-wide association studies for age-related hearing impairment (ARHI)-related traits as described in the following publication:<br> Wells, H.R.R., Abidin, F.N.Z., Freidin, M.B. et al. Genome-wide association study suggests that variation at the RCOR1 locus is associated with tinnitus in UK Biobank. Sci Rep 11, 6470 (2021). https://doi.org/10.1038/s41598-021-85871-6</p>

opencc-by-4.0Jun 2022View details →
zenodo48/100

Voxel-level summary statistics of hippocampus shape, white matter microstructure, and cortical surface curvature in UK Biobank (n=33,324)

<p>This deposit hosts GWAS summary statistics of hippocampus shape (n=33,324), white matter microstructure (n=33,324), and cortical surface curvature (n=15,752) using UKB unrelated white subjects. The data was generated by using the highly efficient imaging genetics (<a href="https://github.com/Zhiwen-Owen-Jiang/heig">HEIG v1.1.0</a>) framework where only the triplets - summary statistics of low-dimensional representations (LDRs), the functional bases, and the variance-covariance matrix LDRs - are shared, which is sufficient to recover all voxel-variant pairs as well as to conduct voxel-level heritability and (cross-trait) genetic correlation analysis. Check the <a href="https://github.com/Zhiwen-Owen-Jiang/heig/wiki">tutorial</a> and&nbsp;the <a href="../records/13770930">example data</a> used in the tutorial.&nbsp;</p> <p>The shared data includes:</p> <p>1. Triplets for hippocampus shape measured by the radial distance from the medial model for each vertex. The original images contain 30,000 vertices while the shared data contains 49 LDRs. Left and right hemispheres were analyzed separately, each with 15,000 vertices.</p> <p>2. Triplets for 21 white matter tracts measured by fractional anisotropy. The original images contain 32,217 voxels and each tract contains 88 ~ 3503 voxels while the shared data contains 1,034 LDRs. Tracts were analyzed separately.</p> <p>3. Triplets for cortical surface curvature. The original images contain 59,412 vertices while the shared data contains 1,750 LDRs. The entire brain was analyzed as a whole.</p> <p>4. LD matrix and its inverse for 22 chromosomes including 460k genotyped SNPs. LD matrix and its inverse were estimated by using two separate datasets each containing 8.4k white unrelated subjects in UKB. Two regularization levels are provided: {85%, 80%} for heritability and genetic correlations within images and {75%, 70%} for cross-trait genetic correlations.</p> <p>5. LD matrix and its inverse for 22 chromosomes including 1.2 million imputed HapMap3 SNPs. &nbsp;LD matrix and its inverse were estimated by using two separate datasets each containing 42k white unrelated subjects in UKB. Two regularization levels are provided: {98%, 95%} for heritability and genetic correlations within images and {90%, 85%} for cross-trait genetic correlations.</p>

opencc-by-4.0Sep 2024View details →
zenodo44/100

Efficient and accurate framework for genome-wide gene-environment interaction analysis in large-scale biobanks

<p>Gene-environment interaction (GxE) analysis elucidates the interplay between genetic predispositions and environmental influences, offering significant potential for precision medicine. With the increasing use of electronic health records (EHR) linked to genetic data in large-scale biobanks, genome-wide association studies (GWAS) have expanded to encompass complex traits with intricate structures, such as time-to-event and ordinal categorical traits. Although these complex traits convey more phenotypic information, most existing scalable genome-wide GxE analysis approaches only focus on quantitative or binary traits. In this work, we propose a scalable and accurate analysis framework, SPAGxE<sub>CCT</sub>, that is applicable to a wide variety of trait types. We extend SPAGxE to SPAGxE+, which can account for sample relatedness. In addition, we extend SPAGxE<sub>CCT</sub> to SPAGxEmix<sub>CCT</sub>, which accounts for population stratification and is applicable to include individuals from multiple ancestries or admixed populations. We applied SPAGxE<sub>CCT</sub>, SPAGxE+, and SPAGxEmix<sub>CCT</sub> to analyze time-to-event traits in UK Biobank. For the SPAGxE<sub>CCT</sub> analyses, 281,149 White British individuals were included. For the SPAGxE+ analyses, 337,367 WB individuals with sample relatedness were included.&nbsp; For the SPAGxEmix<sub>CCT</sub> analyses, 338,044 individuals from all ancestries were included. SPAGxE<sub>CCT</sub>, SPAGxE+, and SPAGxEmix<sub>CCT</sub> are computationally efficient to analyze large datasets with hundreds of thousands of individuals, can accurately control type I error rates while remaining powerful to identify novel GxE findings.</p>

opencc-by-4.0Jun 2024View details →
zenodo44/100

UK Biobank release and systematic evaluation of optimised polygenic risk scores for 53 diseases and quantitative traits

<p>Summary-level GWAS data for 53 traits generated by <a href="https://www.genomicsplc.com/">Genomics plc</a> as presented in:</p> <p>Thompson D. et al. UK Biobank release and systematic evaluation of optimised polygenic risk scores for 53 diseases and quantitative traits (<a href="https://doi.org/10.1101/2022.06.16.22276246">https://doi.org/10.1101/2022.06.16.22276246</a>)</p> <p>If you have any questions or comments regarding these files, please contact Genomics plc at <a href="mailto:research@genomicsplc.com">research@genomicsplc.com</a></p> <p><strong>NOTES</strong></p> <p>These analyses were carried out using the full UK Biobank (UKB) imputation data release (v3b). After removal of exclusions and withdrawals, a subset of 337,151 UKB individuals, the White British Unrelated (WBU) subgroup, was defined as the intersection of two sample groups created by Bycroft et al 2018 (Nature 562, 203-209): the &lsquo;White British ancestry&rsquo; group (UKB Data Field 22006) and the &lsquo;used in genetic principal components&rsquo; group (UKB Data Field 22020), the latter being high quality samples that were filtered to avoid closely related individuals. All GWAS analyses were performed on the WBU subgroup.</p> <p>Phenotypes were defined as described in Supplementary Table 1 &lsquo;Phenotype definitions&rsquo; using a combination of Hospital Episode Statistics, Cancer Registry reports (where applicable) and self-report responses, with the exception of coronary artery disease (CAD). GWAS data was generated for both a &ldquo;narrow&rdquo; and a &ldquo;broad&rdquo; definition of CAD. The former was used as part of the training data for the Enhanced CAD PRS, the latter was used as part of the training data for the Enhanced CVD PRS. The phenotype definitions for &ldquo;narrow&rdquo; and a &ldquo;broad&rdquo; CAD are as follows:</p> <table> <tbody> <tr> <td>Narrow CAD<br> (includes angina)</td> <td>ICD10 codes (where .X indicates all subcodes) from both hospital and death records: I21, I22, I23, I24.1, I25.2, I20.X. ICD9 codes: 410-412, 42979, 413.X. OPCS-4 codes (K40.1&ndash;40.4, K41.1&ndash;41.4, K45.1&ndash;45.5,K49.1&ndash;49.2, K49.8&ndash;49.9, K50.2, K75.1&ndash;75.4, K75.8&ndash;75.9), self-reported heart attack (UKB codes 1075 in field 20002; code 1 in field 6150), self-reported coronary angioplasty (ptca) or coronary artery bypass graft (UKB codes 1070 and 1095 &nbsp;in field 20004), self-reported angina.</td> </tr> <tr> <td>Broad CAD<br> (includes angina and all ischaemic heart disease)</td> <td>As for Narrow CAD, plus ICD10 codes I24.X, I25X, and ICD9 codes 414.X (where .X indicates all subcodes).</td> </tr> </tbody> </table> <p>Note that there is no GWAS for cardiovascular disease (CVD) per se. This is because the UKB training data for the Enhanced CVD PRS consisted of separate GWASs for &ldquo;narrow&rdquo; CAD and ischaemic stroke.</p> <p>All analyses included Age at assessment, sex (for non-sex specific traits), genotyping chip, and 10 principal components as covariates.</p> <p>GWAS summary statistics for each trait were generated by applying PLINK 2.0 to the WBU subgroup, using a logistic regression for disease traits, and a linear regression model for quantitative traits. For chromosome X variants males were treated as having 0 or 2 alternative alleles.</p> <p>The results are not adjusted for genomic control.</p> <p><strong>DATA FILE CONTENT DESCRIPTION (DISEASE TRAITS)</strong></p> <table> <tbody> <tr> <td>cpra</td> <td>Variant ID in &lsquo;CPRA&rsquo; format. Position reflects position in b37</td> </tr> <tr> <td>chrom</td> <td>Chromosome</td> </tr> <tr> <td>pos</td> <td>Position in base pairs (b37, 1-based)</td> </tr> <tr> <td>alt</td> <td>Alternative allele (effect allele)</td> </tr> <tr> <td>beta</td> <td>Effect size (log odds ratio)</td> </tr> <tr> <td>standard_error</td> <td>Standard error of beta</td> </tr> <tr> <td>minus_log10_p</td> <td>Minus log(base 10) of P-value</td> </tr> <tr> <td>ref</td> <td>Reference allele (non-effect allele)</td> </tr> <tr> <td>ncase</td> <td>Number of cases</td> </tr> <tr> <td>ncontrol</td> <td>Number of controls</td> </tr> </tbody> </table> <p><strong>DATA FILE CONTENT DESCRIPTION (QUANTITATIVE TRAITS)</strong></p> <table> <tbody> <tr> <td>cpra</td> <td>Variant ID in &lsquo;CPRA&rsquo; format. Position reflects position in b37</td> </tr> <tr> <td>chrom</td> <td>Chromosome</td> </tr> <tr> <td>pos</td> <td>Position in base pairs (b37, 1-based)</td> </tr> <tr> <td>alt</td> <td>Alternative allele (effect allele)</td> </tr> <tr> <td>beta</td> <td>Effect size</td> </tr> <tr> <td>standard_error</td> <td>Standard error of beta</td> </tr> <tr> <td>minus_log10_p</td> <td>Minus log(base 10) of P-value</td> </tr> <tr> <td>ref</td> <td>Reference allele (non-effect allele)</td> </tr> <tr> <td>ntotal</td> <td>Total sample size</td> </tr> </tbody> </table> <p><strong>FILE NAMES</strong></p> <p>The following is a list of traits and their corresponding file names.</p> <p><em><strong>DISEASE TRAITS</strong></em></p> <table> <tbody> <tr> <td>Age-related macular degeneration</td> <td>amd_strict_UKB_WBU.csv.gz</td> </tr> <tr> <td>Alzheimer&#39;s disease</td> <td>alzheimers_disease_UKB_WBU.csv.gz</td> </tr> <tr> <td>Asthma</td> <td>asthma_UKB_WBU.csv.gz</td> </tr> <tr> <td>Atrial fibrillation</td> <td>atrial_fibrillation_UKB_WBU.csv.gz</td> </tr> <tr> <td>Bipolar disorder</td> <td>bipolar_disorder_UKB_WBU.csv.gz</td> </tr> <tr> <td>Bowel cancer</td> <td>CRC_UKB_WBU.csv.gz</td> </tr> <tr> <td>Breast cancer</td> <td>BC_UKB_WBU_women.csv.gz</td> </tr> <tr> <td>Coeliac disease</td> <td>celiac_disease_UKB_WBU.csv.gz</td> </tr> <tr> <td>Narrow coronary artery disease</td> <td>NARROW_CAD_UKB_WBU.csv.gz</td> </tr> <tr> <td>Broad coronary artery disease</td> <td>BROAD_CAD_UKB_WBU.csv.gz</td> </tr> <tr> <td>Crohn&#39;s disease</td> <td>crohns_disease_UKB_WBU.csv.gz</td> </tr> <tr> <td>Epithelial ovarian cancer</td> <td>OC_UKB_WBU.csv.gz</td> </tr> <tr> <td>Hypertension</td> <td>HT_UKB_WBU.csv.gz</td> </tr> <tr> <td>Ischaemic stroke</td> <td>IS_stroke_UKB_WBU.csv.gz</td> </tr> <tr> <td>Melanoma</td> <td>melanoma_UKB_WBU.csv.gz</td> </tr> <tr> <td>Multiple sclerosis</td> <td>multiple_sclerosis_UKB_WBU.csv.gz</td> </tr> <tr> <td>Osteoporosis</td> <td>OP_WBU_training.csv.gz</td> </tr> <tr> <td>Prostate cancer</td> <td>PC_UKB_WBU.csv.gz</td> </tr> <tr> <td>Parkinson&#39;s disease</td> <td>parkinsons_disease_UKB_WBU.csv.gz</td> </tr> <tr> <td>Primary open angle glaucoma</td> <td>POAG_WBU_training.csv.gz</td> </tr> <tr> <td>Psoriasis</td> <td>psoriasis_UKB_WBU.csv.gz</td> </tr> <tr> <td>Rheumatoid arthritis</td> <td>rheumatoid_arthritis_UKB_WBU.csv.gz</td> </tr> <tr> <td>Schizophrenia</td> <td>schizophrenia_UKB_WBU.csv.gz</td> </tr> <tr> <td>Systemic lupus erythematosus</td> <td>lupus_UKB_WBU.csv.gz</td> </tr> <tr> <td>Type 1 diabetes</td> <td>t1d_UKB_WBU.csv.gz</td> </tr> <tr> <td>Type 2 diabetes</td> <td>T2D_UKB_WBU.csv.gz</td> </tr> <tr> <td>Ulcerative colitis</td> <td>ulcerative_colitis_UKB_WBU.csv.gz</td> </tr> <tr> <td>Venous thromboembolic disease</td> <td>VTE_UKB_WBU.csv.gz</td> </tr> </tbody> </table> <p><em><strong>QUANTITATIVE TRAITS</strong></em></p> <table> <tbody> <tr> <td>Age at menopause</td> <td>age_at_menopause_UKB_WBU.csv.gz</td> </tr> <tr> <td>Apolipoprotein A1</td> <td>apolipoprotein_a1_UKB_WBU.csv.gz</td> </tr> <tr> <td>Apolipoprotein B</td> <td>apolipoprotein_b_UKB_WBU.csv.gz</td> </tr> <tr> <td>Body mass index</td> <td>bmi_UKB_WBU.csv.gz</td> </tr> <tr> <td>Calcium</td> <td>calcium_UKB_WBU.csv.gz</td> </tr> <tr> <td>Docosahexaenoic acid</td> <td>docosahexaenoic_acid_UKB_WBU.csv.gz</td> </tr> <tr> <td>Estimated bone mineral density T-score</td> <td>BMD_WBU_training.csv.gz</td> </tr> <tr> <td>Estimated glomerular filtration rate (creatinine based)</td> <td>egfr_UKB_WBU.csv.gz</td> </tr> <tr> <td>Estimated glomerular filtration rate (cystatin based)</td> <td>egfr_cys_UKB_WBU.csv.gz</td> </tr> <tr> <td>Glycated haemoglobin</td> <td>hba1c_UKB_WBU_nodiabetes.csv.gz</td> </tr> <tr> <td>High density lipoprotein cholesterol</td> <td>hdl_cholesterol_UKB_WBU.csv.gz</td> </tr> <tr> <td>Height</td> <td>height_UKB_WBU.csv.gz</td> </tr> <tr> <td>Intraocular pressure</td> <td>iop_WBU_training.csv.gz</td> </tr> <tr> <td>Low density lipoprotein cholesterol</td> <td>ldl_UKB_WBU_nostatins.csv.gz</td> </tr> <tr> <td>Omega-6 fatty acids</td> <td>omega_6_fatty_acids_UKB_WBU.csv.gz</td> </tr> <tr> <td>Omega-3 fatty acids</td> <td>omega_3_fatty_acids_UKB_WBU.csv.gz</td> </tr> <tr> <td>Phosphatidylcholines</td> <td>phosphatidylcholines_UKB_WBU.csv.gz</td> </tr> <tr> <td>Phosphoglycerides</td> <td>phosphoglycerides_UKB_WBU.csv.gz</td> </tr> <tr> <td>Polyunsaturated fatty acids</td> <td>polyunsaturated_fatty_acids_UKB_WBU.csv.gz</td> </tr> <tr> <td>Resting heart rate</td> <td>resting_heart_rate_UKB_WBU.csv.gz</td> </tr> <tr> <td>Remnant cholesterol (Non-HDL, Non-LDL cholesterol)</td> <td>remnant_cholesterol__UKB_WBU.csv.gz</td> </tr> <tr> <td>Sphingomyelins</td> <td>sphingomyelins_UKB_WBU.csv.gz</td> </tr> <tr> <td>Total cholesterol</td> <td>total_cholesterol_UKB_WBU.csv.gz</td> </tr> <tr> <td>Total fatty acids</td> <td>total_fatty_acids_UKB_WBU.csv.gz</td> </tr> <tr> <td>Total triglycerides</td> <td>total_triglycerides_UKB_WBU.csv.gz</td> </tr> </tbody> </table>

opencc-by-4.0Jun 2022View details →
zenodo44/100

GWAS on self-reported hearing difficulty in the UK Biobank

<p>The dataset contains results of two genome-wide association studies for age-related hearing impairment (ARHI)-related traits as described in the following publication</p> <p><a href="https://www.sciencedirect.com/science/article/pii/S0002929719303477#!"><em>Wells HRR,&nbsp;Freidin MB, Zainul Abidin FN, Payton A, Dawes P, Munro KJ, Morton CC, Moore DR, Dawson SJ, Williams FMK. GWAS Identifies 44 Independent Associated Genomic Loci for Self-Reported Adult Hearing Difficulty in UK Biobank.&nbsp;Am J Hum Genet. 2019 Oct 3;105(4):788-802. doi: 10.1016/j.ajhg.2019.09.008. Epub 2019 Sep 26.</em></a>&nbsp;</p> <p>Please cite the article if using this dataset.</p> <p>Two files provide summary statistics for discovery analysis of <em><strong>Hearing difficulty (HD)&nbsp;</strong></em>and <em><strong>Hearing aid use (HAID)</strong></em> phenotypes for individuals of European descent from <a href="https://www.ukbiobank.ac.uk/">UK Biobank</a>.</p> <p><strong>Acknowledgements</strong></p> <p>The research was carried out using the UK Biobank Resource under application number 11516. H.R.R.W. is funded by a PhD Studentship Grant, S44, from&nbsp;Action on Hearing Loss. The study was also supported by funding from&nbsp;NIHR UCLH BRC Deafness and Hearing&nbsp;Problems Theme, a grant from&nbsp;MED_EL, and the&nbsp;NIHR Manchester Biomedical Research Centre. The English Longitudinal Study of Aging is jointly run by University College London, Institute for Fiscal Studies, University of Manchester, and National Centre for Social Research. Genetic analyses have been carried out by UCL Genomics and funded by the&nbsp;Economic and Social Research Council&nbsp;and the&nbsp;National Institute on Aging. Data governance was provided by the METADAC data access committee, funded by&nbsp;ESRC,&nbsp;Wellcome, and&nbsp;MRC&nbsp;(2015-2018: Grant Number&nbsp;MR/N01104X/1&nbsp;2018-2020: Grant Number&nbsp;ES/S008349/1). TwinsUK&nbsp;is funded by the&nbsp;Wellcome Trust,&nbsp;Medical Research Council,&nbsp;European Union, the National Institute for Health Research (NIHR)-funded&nbsp;BioResource,&nbsp;Clinical Research Facility, and&nbsp;Biomedical Research Centre&nbsp;based at Guy&rsquo;s and St Thomas&rsquo; NHS Foundation Trust in partnership with King&rsquo;s College London. We would like to thank all the participants of UK Biobank, English Longitudinal Study of Aging, and TwinsUK.</p> <p><strong>Column headers:</strong></p> <p>SNP, SNP rsID</p> <p>CHR, chromosome</p> <p>BP, genomic position&nbsp;(GRCh37 build)&nbsp;</p> <p>ALLELE1, effect allele (coded as &quot;1&quot;)</p> <p>ALLELE0, reference allele (coded as &quot;0&quot;)&nbsp;</p> <p>A1FREQ, effect allele frequency</p> <p>INFO, imputation quality</p> <p>BETA, effect size of effect allele</p> <p>SE: standard error of effect size</p> <p>P, P-value of association (without GC correction)</p>

opencc-by-4.0Oct 2019View details →
zenodo44/100

Neural Protein Associations with Parkinson's, Stroke, and Alzheimer's: Insights from UK Biobank Data

<p>该数据集来自英国生物样本库 (UKB) 和英国生物样本库制药蛋白质组学项目 (UKB-PPP),包含来自 54,219 名参与者的全面蛋白质组学和人口统计信息。该数据集包括 2,941 种蛋白质分析物的测量值,代表 2,923 种独特蛋白质。选择了一组 217 种神经学相关蛋白和 10 种人口统计学和生活方式协变量进行分析。该研究侧重于帕金森病、中风和阿尔茨海默病,使用 ICD-10 代码确定病例。该数据集是研究蛋白质组学生物标志物与神经系统疾病之间关系的宝贵资源,可用于复制和进一步研究。</p>

opencc-by-4.0Sep 2024View details →
zenodo44/100

A common NFKB1 variant detected through antibody analysis in UK Biobank predicts risk of infection and allergy: Summary statistics - Health records

<p>Infectious agents contribute significantly to the global burden of diseases, through both acute infection and their chronic sequelae. We leveraged the UK Biobank to identify genetic loci that influence humoral immune response to multiple infections. From 45 genome-wide association studies in 9,611 participants from UK Biobank, we identified NFKB1 as a locus associated with quantitative antibody responses to multiple pathogens including those from the herpes, retro- and polyoma-virus families. An insertion-deletion variant thought to affect NFKB1 expression (rs28362491), was mapped as the likely causal variant. This variant has persisted throughout hominid evolution and could play a key role in regulation of the immune response. Using 121 infection and inflammation related traits in 487,297 UK Biobank participants, we show that the deletion allele was associated with an increased risk of infection from diverse pathogens but had a protective effect against allergic disease. We propose that altered expression of NFKB1, as a result of the deletion, modulates haematopoietic pathways, and likely impacts cell survival, antibody production, and inflammation. Taken together, we show that disruptions to the tightly regulated immune processes may tip the balance between exacerbated immune responses and allergy, or increased risk of infection and impaired resolution of inflammation.&nbsp;</p> <p>-------------------------------------------------------------------------------------</p> <p>This dataset contains GWAS summary statistics for infection, inflammation, and allergy related traits in&nbsp;487,297 individuals</p>

opencc-by-4.0Nov 2022View details →
zenodo44/100

A common NFKB1 variant detected through antibody analysis in UK Biobank predicts risk of infection and allergy: Summary statistics - Serology

<p>Infectious agents contribute significantly to the global burden of diseases, through both acute infection and their chronic sequelae. We leveraged the UK Biobank to identify genetic loci that influence humoral immune response to multiple infections. From 45 genome-wide association studies in 9,611 participants from UK Biobank, we identified NFKB1 as a locus associated with quantitative antibody responses to multiple pathogens including those from the herpes, retro- and polyoma-virus families. An insertion-deletion variant thought to affect NFKB1 expression (rs28362491), was mapped as the likely causal variant. This variant has persisted throughout hominid evolution and could play a key role in regulation of the immune response. Using 121 infection and inflammation related traits in 487,297 UK Biobank participants, we show that the deletion allele was associated with an increased risk of infection from diverse pathogens but had a protective effect against allergic disease. We propose that altered expression of NFKB1, as a result of the deletion, modulates haematopoietic pathways, and likely impacts cell survival, antibody production, and inflammation. Taken together, we show that disruptions to the tightly regulated immune processes may tip the balance between exacerbated immune responses and allergy, or increased risk of infection and impaired resolution of inflammation.&nbsp;</p> <p>-------------------------------------------------------------------------------------</p> <p>This dataset contains GWAS summary statistics for quantitative antibody responses&nbsp;in 9611 individuals and results for a&nbsp;meta-analysis of UK Biobank and CoLaus/PsyCoLaus antibody responses.</p> <p>&nbsp;</p>

opencc-by-4.0Nov 2022View details →
zenodo40/100

Results for 2,230 UK Biobank binary and continuous traits

<p>Results for 2,230 UK Biobank binary and continuous&nbsp;traits.&nbsp;</p> <p>We applied the gene-based tests (Gene1D, Gene3D, GeneScan1D and GeneScan3D) to 1,403 UK Biobank binary phecodes and&nbsp;827 continuous phenotypes (797 continuous traits + 30 biomarkers) using GWAS summary statistics on 28 million imputed variants.&nbsp;</p> <p>The results are&nbsp;in 3 different&nbsp;zipped folders:&nbsp;&#39;GeneScan3D_UKBB_1403binary_results.zip&#39;, &#39;GeneScan3D_UKBB_797continuous_results.zip&#39; and &#39;GeneScan3D_UKBB_30biomarkers_results.zip&#39;. A list of all 2,230 binary and continuous phenotypes&nbsp;is available in excel file&nbsp;&#39;UKBB_phenotype_description.xlsx&#39;.</p> <p>Reference: Ma, S., Dalgleish, J. L ., Lee, J., Wang, C., Liu, L., Gill, R., Buxbaum, J. D., Chung, W., Aschard, H., Silverman, E. K., Cho, M. H., He, Z. and Ionita-Laza, I. &quot;Improved gene-based testing by integrating long-range chromatin interactions and knockoff statistics&quot;, 2021</p>

opencc-by-4.0Jan 2021View details →
zenodo40/100

LD matrices from the White British cohort in the UK Biobank in Zarr format

<p>This dataset contains the Linkage Disequilibrium (LD) matrices that were used in the analyses described in the manuscript:</p> <p><strong>Fast and Accurate Bayesian Polygenic Risk Modeling with Variational Inference</strong><br> Shadi Zabad, Simon Gravel, Yue Li<br> McGill University</p> <p>LD matrices record the SNP-by-SNP correlations in a given sample of individuals from a general population. In this case, we threshold the matrices so that we only record the correlations between SNPs that are at most 3 centi Morgan apart. These matrices record the SNP correlations in a random sample of 50,000 individuals&nbsp;from the White British cohort in the UK Biobank dataset. There is one matrix per autosomal chromosome (chr_1, chr_2, ..., chr_22). The matrices are stored in <a href="https://zarr.readthedocs.io/en/stable/">Zarr</a> format, a chunked on-disk array storage format that allows for multi-threaded read and write access.</p> <p>To access these matrices, consult the codebase of <a href="https://github.com/shz9/magenpy"><strong>magenpy</strong></a>, our custom python package with special data structures for processing these LD matrices.</p> <p>UPDATE (03/09/2022): We updated the matrices to add the reference allele attribute (A2) and we also now have one tar archive per chromosome.<br> &nbsp;</p>

opencc-by-4.0May 2022View details →
zenodo40/100

QR GWAS summary statistics for 39 quantitative traits in the UK Biobank

<p>Quantile regression (QR) GWAS summary statistics from the study "Genome-wide discovery for biomarkers using quantile regression at biobank scale". The preprint is available at <a href="https://doi.org/10.1101/2023.06.05.543699" target="_blank" rel="noopener">https://doi.org/10.1101/2023.06.05.543699</a>.&nbsp;</p> <p><strong>List of traits</strong></p> <p>A comma-delimited text file, QRGWAS.Traits_n39.csv, includes the list of 39 quantitative traits from the UK Biobank reported in the QR GWAS analyses above.</p> <p><strong>Summary statistics</strong></p> <p>The tab-delimited text files are QR GWAS summary statistics, which are bgzip compressed (.tsv.gz files) and tabix indexed (.tbi files).</p> <ul> <li>Column "CHR": chromosome</li> <li>Column "POS": based pair position</li> <li>Column "ID": variant ID</li> <li>Column "REF": non-effect allele</li> <li>Column "ALT": effect allele tested in GWAS</li> <li>Column "EAF": frequency of the effect allele</li> <li>Column "N": sample size</li> <li>Column "P_QR": integrated p-value of the quantile regression (QR) model across multiple quantile levels.</li> <li>Column "P_LR": p-value of the linear regression (LR) association statistic</li> <li>Columns from "P_Q10" to "P_Q90": quantile-specific QR p-value for the quantile levels 0.1, 0.2, ..., 0.9 (10th, 20th, ..., 90th quantiles).</li> </ul>

opencc-by-4.0Apr 2024View details →
zenodo40/100

COJO ARG variants from "Biobank-scale inference of ancestral recombination graphs enables genealogical analysis of complex traits"

<p>These are&nbsp;COJO ARG variants accompanying the manuscript&nbsp;&quot;Biobank-scale inference of ancestral recombination graphs enables genealogical analysis of complex traits&quot;. For more details, view the README.md file and refer to our manuscript.</p>

opencc-by-4.0Dec 2022View details →
zenodo40/100

A biobank of patients with Primary Immune Deficiencies (PID)

<pre>A biobank of patients with Primary Immune Deficiencies (PID).</pre>

opencc-by-4.0Oct 2023View details →
zenodo36/100

Samples and data accessibility in research biobanks

<p>This dataset contains answers at a questionnaire relative to modes of sample and data accessibility in research Biobanks</p>

opencc-zeroApr 2015View details →
zenodo36/100

Minimal dataset for the manuscript "Better together against genetic heterogeneity: a sex-combined joint main and interaction analysis of 290 quantitative traits in the UK Biobank".

<p>Dataset "lin2024-sex_combined_interaction-association_signifincant_in_one_or_more_tests-summary.txt" is a minimal dataset to reproduce the figures and tables in the manuscript "Better together against genetic heterogeneity: a sex-combined joint main and interaction analysis of 290 quantitative traits in the UK Biobank".&nbsp;</p> <p><br>To generate this dataset, see "https://github.com/BoxiLin/t2meta" Steps 0, 1.</p> <p>This dataset is the input for Steps 2, 3, 4, 5 to generate Figures 1-3 and Table 2-3.</p> <p>&nbsp;</p> <p>##### Column information ########################</p> <p>The following columns are annotations on each variant in the GWAS, calculated across the analysis subset of 361,194 samples by the Neale lab:</p> <p>code: Phenotype identifier in the form of "[UKB Data field]_raw"<br>variant: Unique variant identifier in the form "chr:pos:ref:alt", where "ref" is aligned to the forward strand.<br>chr: Chromosome of the variant.<br>pos: Position of the variant in GRCh37 coordinates.<br>rsid: rs ID<br>ref: Reference allele on the forward strand.<br>alt: Alternate allele (not necessarily minor allele).<br>p_hwe: Hardy-Weinberg p-value.<br>info: Imputation INFO score as provided by UK Biobank.</p> <p>&nbsp;</p> <p>The following columns are sex-stratified test statistics calculated by the Neale lab:</p> <p>minor_allele.x: Minor allele (AF &lt; 0.5) in the female GWAS&nbsp;<br>minor_AF.x: Minor allele frequency in the female GWAS&nbsp;<br>beta.x: Estimated effect size of alt allele in the female GWAS&nbsp;<br>se.x: Estimated standard error of beta in the female GWAS<br>tstat.x: t-statistic of beta estimate (= beta/se) in the female GWAS&nbsp;<br>pval.x: p-value of beta significance test in the female GWAS&nbsp;</p> <p>minor_allele.y: Minor allele (AF &lt; 0.5) in the male GWAS&nbsp;<br>minor_AF.y: Minor allele frequency in the male GWAS&nbsp;<br>beta.y: Estimated effect size of alt allele in the male GWAS&nbsp;<br>se.y: Estimated standard error of beta in the male GWAS&nbsp;<br>tstat.y: t-statistic of beta estimate (= beta/se) in the male GWAS&nbsp;<br>pval.y: p-value of beta significance test in the male GWAS&nbsp;</p> <p>&nbsp;</p> <p><br>The following columns are sex-combined test statistics calculated in our analysis:</p> <p>T.I: test statsitic for interaction effect-only&nbsp;<br>p.T.I: &nbsp;p-value of the interaction effect-only test&nbsp;<br>TSG.L: &nbsp;test statsitic for inverse variance weighted meta-analysis<br>p.TSG.L: p-value of the inverse variance weighted meta-analysis<br>TSG.Q: test statsitic for the omnibus meta-analysis<br>p.TSG.Q: p-value for the omnibus meta-analysis</p>

opencc-by-4.0Mar 2024View details →
zenodo36/100

Fine-mapping gene-based associations via knockoff analysis of biobank-scale data with applications to UK Biobank

<p>The results of BIGKnock analyses of manuscript&nbsp;&#39;&#39;Fine-mapping gene-based associations via knockoff analysis of biobank-scale data with applications to UK Biobank&#39;&#39;</p>

opencc-by-4.0May 2022View details →
dryad36/100

Improving genome-wide association discovery and genomic prediction accuracy in biobank data

<p>Genetically informed, deep-phenotyped biobanks are an important research resource and it is imperative that the most powerful, versatile, and efficient analysis approaches are used. Here, we apply our recently developed Bayesian grouped mixture of regressions model (GMRM) in the UK and Estonian Biobanks and obtain the highest genomic prediction accuracy reported to date across 21 heritable traits. When compared to other approaches, GMRM accuracy was greater than annotation prediction models run in the LDAK or LDPred-funct software by 15% (SE 7%) and 14% (SE 2%), respectively, and was 18% (SE 3%) greater than a baseline BayesR model without single-nucleotide polymorphism (SNP) markers grouped into minor allele frequency–linkage disequilibrium (MAF-LD) annotation categories. For height, the prediction accuracy R 2 was 47% in a UK Biobank holdout sample, which was 76% of the estimated h SNP 2 . We then extend our GMRM prediction model to provide mixed-linear model association (MLMA) SNP marker estimates for genome-wide association (GWAS) discovery, which increased the independent loci detected to 16,162 in unrelated UK Biobank individuals, compared to 10,550 from BoltLMM and 10,095 from Regenie, a 62 and 65% increase, respectively. The average χ<sup>2</sup> value of the leading markers increased by 15.24 (SE 0.41) for every 1% increase in prediction accuracy gained over a baseline BayesR model across the traits. Thus, we show that modeling genetic associations accounting for MAF and LD differences among SNP markers, and incorporating prior knowledge of genomic function, is important for both genomic prediction and discovery in large-scale individual-level studies.</p>

opencc-zeroSep 2022View details →
dryad36/100

Pleiotropy of UK Biobank metabolites

<p>Pleiotropy and genetic correlation are widespread features in GWAS, but they are often difficult to interpret at the molecular level. Here, we perform GWAS of 16 metabolites clustered at the intersection of amino acid catabolism, glycolysis, and ketone body metabolism in a subset of UK Biobank. We utilize the well-documented biochemistry jointly impacting these metabolites to analyze pleiotropic effects in the context of their pathways. Among the 213 lead GWAS hits, we find a strong enrichment for genes encoding pathway-relevant enzymes and transporters. We demonstrate that the effect directions of variants acting on biology between metabolite pairs often contrast with those of upstream or downstream variants as well as the polygenic background. Thus, we find that these outlier variants often reflect biology local to the traits. Finally, we explore the implications for interpreting disease GWAS, underscoring the potential of unifying biochemistry with dense metabolomics data to understand the molecular basis of pleiotropy in complex traits and diseases.</p>

opencc-zeroOct 2022View details →
zenodo36/100

Large-scale UK Biobank Whole Body Atlases

<p>Population atlases are commonly utilised in medical imaging to facilitate the investigation of variability across populations.&nbsp;<br>Such atlases enable the mapping of medical images into a common coordinate system, promoting comparability and enabling the study of inter-subject differences.&nbsp;<br>Constructing such atlases becomes particularly challenging when working with highly heterogeneous datasets, such as whole-body images, where subjects show significant anatomical variations.<br>In this work, we propose a pipeline for generating a standardised whole-body atlas for a highly heterogeneous population by partitioning the population into anatomically meaningful subgroups.&nbsp;<br>Using magnetic resonance (MR) images from the UK Biobank dataset, we create six whole-body atlases representing a healthy population average.&nbsp;<br>We furthermore unbias them, and this way obtain a realistic representation of the population.<br>In addition to the anatomical atlases, we generate probabilistic atlases that capture the distributions of abdominal fat (visceral and subcutaneous) and five abdominal organs across the population (liver, spleen, pancreas, left and right kidneys).<br>We demonstrate a clinical application of these atlases, using the differences between subjects with medical conditions such as diabetes and cardiovascular diseases and healthy subjects from the atlas space.<br>With this work, we make the constructed anatomical and label atlases publically available and anticipate them to support medical research conducted on whole-body MR images.</p>

opencc-by-4.0Jul 2024View details →
zenodo36/100

Summary-level data from meta-analysis of fat distribution phenotypes in UK Biobank and GIANT

<p>~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~</p> <p>Summary-level data as presented in:</p> <p>&quot;Meta-analysis of genome-wide association studies for body fat distribution in 694,649 individuals of European ancestry.&quot; Pulit, SL et al. bioRxiv, 2018. https://www.biorxiv.org/content/early/2018/04/18/304030</p> <p>**If you use these data, please cite the above preprint.</p> <p>If you have any questions or comments regarding these files, please contact me:</p> <p>Sara L Pulit<br> spulit@well.ox.ac.uk or s.l.pulit@umcutrecht.nl</p> <p>~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~</p> <p><strong>(1) Data files</strong></p> <p><em>i. whradjbmi.giant-ukbb.meta-analysis.combined.23May2018.txt</em><br> Meta-analysis of waist-to-hip ratio adjusted for body mass index (whradjbmi) in UK Biobank and GIANT data. Combined set of samples, max N = 694,649.</p> <p><em>ii. whradjbmi.giant-ukbb.meta-analysis.females.23May2018.txt</em><br> Meta-analysis of whradjbmi in UK Biobank and GIANT data. Female samples only, max N = 379,501.</p> <p><em>iii. whradjbmi.giant-ukbb.meta-analysis.males.23May2018.txt</em><br> Meta-analysis of whradjbmi in UK Biobank and GIANT data. Male samples only, max N = 315,284.</p> <p><em>iv. whr.giant-ukbb.meta-analysis.combined.23May2018.txt</em><br> Meta-analysis of waist-to-hip ratio (whr) in UK Biobank and GIANT data. Combined set of samples, max N = 697,734.</p> <p><em>v. whr.giant-ukbb.meta-analysis.females.23May2018.txt</em><br> Meta-analysis of whr in UK Biobank and GIANT data. Female samples only, max N = 381,152.</p> <p><em>vi. whr.giant-ukbb.meta-analysis.males.23May2018.txt</em><br> Meta-analysis of whr in UK Biobank and GIANT data. Male samples only, max N = 316,772.</p> <p><em>vii. bmi.giant-ukbb.meta-analysis.combined.23May2018.txt</em><br> Meta-analysis of body mass index (bmi) in UK Biobank and GIANT data. Combined set of samples, max N = 806,834.</p> <p><em>viii. bmi.giant-ukbb.meta-analysis.females.23May2018.txt</em><br> Meta-analysis of bmi in UK Biobank and GIANT data. Female samples only, max N = 434,794.</p> <p><em>ix. bmi.giant-ukbb.meta-analysis.males.23May2018.txt</em><br> Meta-analysis of bmi in UK Biobank and GIANT data. Male samples only, max N = 374,756.</p> <p><strong>(2) Data file format</strong></p> <p>CHR:&nbsp;Chromosome</p> <p>POS:&nbsp;Chromosomal position of the SNP, build hg19</p> <p>SNP: the dbSNP151 identifier of the SNP, followed by the first allele and second allele of the SNP, delimited with a colon. A small number of SNPs (&lt;9,000) from the GIANT data had no dbSNP151 identifier, and are left as just an rsID. Note that these SNPs are also missing chromosome and position information (not provided in the GIANT data).</p> <p>Tested_Allele: the allele for which all association statistics are reported</p> <p>Other_Allele: the other allele at the SNP</p> <p>Freq_Tested_Allele:&nbsp;frequency of the tested allele</p> <p>BETA: the effect size of the tested allele</p> <p>SE: the standard error of the beta</p> <p>P:&nbsp;the p-value of the SNP, as reported from the inverse variance-weighted fixed effects meta-analysis</p> <p>N:&nbsp;the total sample size for this SNP</p> <p>INFO: the imputation quality (info score) of the SNP, as reported by UK Biobank. A number between 0 and 1 indicating quality of imputation (0, poor quality; 1, high quality or genotyped). Note that the summary-level GIANT data does not report info score, so SNPs appearing only in the GIANT analysis do not have info scores.</p>

opencc-by-4.0May 2018View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record