Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
260
datasets available to search
ShareScore release 0.9.0
Dataset results
260 results for “Biobank”
UK Biobank GWAS of WHRadjBMI in combined, women only, and men only
<p>GWAS of WHRadjBMI in UK Biobank. GWAS was performed in BOLT-LMM and full details can be found here: <a href="https://www.ncbi.nlm.nih.gov/pubmed/30239722">https://www.ncbi.nlm.nih.gov/pubmed/30239722</a> (the paper is open access).</p> <p>Columns:</p> <ul> <li>SNP, SNP identifier</li> <li>CHR, chromosome</li> <li>BP, position (hg19)</li> <li>GENPOS, genetic position</li> <li>ALLELE1, tested allele</li> <li>ALLELE0, other allele</li> <li>A1FREQ, frequency of the tested allele</li> <li>INFO, imputation quality (info score)</li> <li>CHISQ_LINREG, chi-square from linear regression</li> <li>P_LINREG, p-value from linear regression</li> <li>BETA, beta (effect) from the linear mixed model</li> <li>SE, standard error from the linear mixed model</li> <li>CHISQ_BOLT_LMM_INF, chi-square from linear mixed model</li> <li>P_BOLT_LMM_INF, p-value from the linear mixed model</li> </ul> <p> </p>
UK Biobank GWAS of WHR in combined, women only, and men only
<p>GWAS of WHR in UK Biobank. GWAS was performed in BOLT-LMM and full details can be found here: <a href="https://www.ncbi.nlm.nih.gov/pubmed/30239722">https://www.ncbi.nlm.nih.gov/pubmed/30239722</a> (the paper is open access).</p> <p>Columns:</p> <ul> <li>SNP, SNP identifier</li> <li>CHR, chromosome</li> <li>BP, position (hg19)</li> <li>GENPOS, genetic position</li> <li>ALLELE1, tested allele</li> <li>ALLELE0, other allele</li> <li>A1FREQ, frequency of the tested allele</li> <li>INFO, imputation quality (info score)</li> <li>CHISQ_LINREG, chi-square from linear regression</li> <li>P_LINREG, p-value from linear regression</li> <li>BETA, beta (effect) from the linear mixed model</li> <li>SE, standard error from the linear mixed model</li> <li>CHISQ_BOLT_LMM_INF, chi-square from linear mixed model</li> <li>P_BOLT_LMM_INF, p-value from the linear mixed model</li> </ul>
UK Biobank GWAS of WHRadjBMI and WHR in combined, women only, and men only (unrelated, self-identified white British)
<p>GWAS of WHRadjBMI and WHR in UK Biobank, using unrelated self-identified white British samples only.</p> <p>GWAS was performed in BOLT-LMM and full details can be found here: <a href="https://www.ncbi.nlm.nih.gov/pubmed/30239722">https://www.ncbi.nlm.nih.gov/pubmed/30239722</a> (the paper is open access).</p> <p>Columns:</p> <ul> <li>SNP, SNP identifier</li> <li>CHR, chromosome</li> <li>BP, position (hg19)</li> <li>GENPOS, genetic position</li> <li>ALLELE1, tested allele</li> <li>ALLELE0, other allele</li> <li>A1FREQ, frequency of the tested allele</li> <li>INFO, imputation quality (info score)</li> <li>CHISQ_LINREG, chi-square from linear regression</li> <li>P_LINREG, p-value from linear regression</li> <li>BETA, beta (effect) from the linear mixed model</li> <li>SE, standard error from the linear mixed model</li> <li>CHISQ_BOLT_LMM_INF, chi-square from linear mixed model</li> <li>P_BOLT_LMM_INF, p-value from the linear mixed model</li> </ul>
GCTB sparse shrunk LD matrices from 2.8M common variants from the UK Biobank - Part AA - START HERE
<p>GCTB sparse shrunk LD matrices from 2.8M common variants from the UK Biobank.</p> <p>Part <strong>AA </strong>of AA, AB, AC, AD and AE.</p> <p><strong>TO JOIN AND UNZIP THESE MATRICES</strong></p> <p><strong>Download all parts to one folder from:</strong></p> <p> PartAA - 10.5281/zenodo.3375373</p> <p> PartAB - 10.5281/zenodo.3376357</p> <p> Part AC - 10.5281/zenodo.3376456</p> <p> Parts AD and AE - 10.5281/zenodo.3376628</p> <p><strong>Use cat to join </strong></p> <p>cat <a href="https://zenodo.org/api/files/8ffd1abc-07ef-4b8a-a6cf-779a73ee03c2/ukb_50k_bigset_2.8M.zip.partaa?versionId=102ded84-352a-494d-8da3-cebb63a20aee">ukb_50k_bigset_2.8M.zip.part* > ukb_50k_bigset_2.8M.zip </a></p> <p>Then unzip. See README for further details.</p> <p>unzip <a href="https://zenodo.org/api/files/8ffd1abc-07ef-4b8a-a6cf-779a73ee03c2/ukb_50k_bigset_2.8M.zip.partaa?versionId=102ded84-352a-494d-8da3-cebb63a20aee">ukb_50k_bigset_2.8M.zip </a></p>
GWAS summary statistics for corneal resistance factor in 72,301 unrelated UK Biobank white-British participants
<p>The dataset contains summary statistics of the genome-wide association study (GWAS) for corneal resistance factor (CRF) of 72,301 unrelated UK-biobank white-British individuals, originally generated by Jiang X, et al (2020<a href="https://doi.org/10.1038/s42003-020-01497-w">)</a> (doi: <a href="https://doi.org/10.1038%2Fs42003-020-01497-w">10.1038/s42003-020-01497-w</a>).</p> <p><strong>Generation of the dataset: </strong></p> <ul> <li><strong>Phynotypic filtering:</strong></li> </ul> <p>The CRF analysis was performed on the average of the left and right eye measurements. Any outliers (i.e. CRF greater than the population mean difference + 3 standard deviations) were removed. Further phenotypic filtering was applied by removing samples linked to, or self-reporting, ocular conditions that could affect the measurements accuracy such as eye surgery, refractive laser surgery, cataract surgery, glaucoma high pressure surgery or laser treatment, corneal graft surgery, eye injury, keratoconus or cornea disorders.</p> <ul> <li><strong>Population selection:</strong></li> </ul> <p>Using the genetic quality control of the UK Biobank, white British participants with imputed data failing heterozygosity or/and missingness, or having a mismatch between self-reported and genotype-derived gender or showing putative sex chromosome aneuploidy as well as individuals who have withdrawn from the study at the time of analysis were removed. <em>This GWAS is restricted to individuals who were <strong>not closely related</strong> (pairwise kinship coefficient greater than 0.025 calculated by KING).</em> Finally, a total of 72,301 individuals were included in the GWAS.</p> <ul> <li><strong>GWAS:</strong></li> </ul> <p>The GWAS was performed using the generalized linear model (--glm) in PLINK2, common and low-frequency (MAF > 0.5%) well-imputed (INFO > 0.6) variants. Covariates fitted in the model were: age, sex, genotyping array, assessment center, genotyping batch, and the 20 first principal components of ancestry provided by the UK Biobank.</p> <ul> <li>UK Biobank application number: 19655</li> </ul> <p> </p> <p>Methods:</p> <p>Please see Jiang X, et al (2020<a href="https://doi.org/10.1038/s42003-020-01497-w">)</a> (doi: <a href="https://doi.org/10.1038%2Fs42003-020-01497-w">10.1038/s42003-020-01497-w</a>) for more details:</p> <p> </p> <p><strong>Header Description:</strong></p> <ul> <li>CHROM: chromosome,</li> <li>POS: hg19/GRCh37 position</li> <li>ID: variant Rsid</li> <li>A1: effective allele, minor allele</li> <li>A2: alternative allele, major allele</li> <li>BETA: effect size</li> <li>SE: standard error</li> <li>P: P value</li> <li>INFO: imputation quality</li> <li>MAF: Minor allele frequency</li> </ul> <p> </p> <p><strong>Some related analysis using this dataset: </strong></p> <p>Jiang X, et al (2020<a href="https://doi.org/10.1038/s42003-020-01497-w">)</a> (doi: <a href="https://doi.org/10.1038%2Fs42003-020-01497-w">10.1038/s42003-020-01497-w</a>): Fine-mapped CRF GWAS loci using FINEMAP and in-sample linkage disequilibrium data </p> <p>Jiang X, et al (2023) (doi: <a href="https://doi.org/10.3389%2Ffgene.2023.1171217">10.3389/fgene.2023.1171217</a>): Two colocalisation analysis using fine-mapped data: 1) CRF - cis-e/sQTL (GTEx v8); (2) CRF-keratoconus</p>
DeepRVAT gene-trait association testing results on the 470k UK Biobank WES dataset
<p>Association testing results from DeepRVAT on the 470k UK Biobank WES dataset, covering all tested genes and traits. Tests were performed on the full dataset and on Caucasian individuals ("cohort" column). "Significant" indicates significance after multiple testing correction (FWER < 5%).</p>
Genetic fine-mapping results for 56 NMR metabolites measured in 246,683 UK Biobank participants
<p>Fine-mapping credible sets for 56 metabolites measured in 246,683 UK Biobank participants using the Nightingale Health platform. Fine-mapping was performed using the https://github.com/AlasooLab/reGSusie workflow.<br><br>The 56_metabolites_finemapping_credible_sets.tsv file contains the fine-mapped credible sets for all 56 metabolites. The *_coloc5_final.tsv.gz files contain the log Bayes factors for each metabolite in each fine-mapped region. </p>
GWAS summary statistics for 56 NMR metabolites measured in 246,683 UK Biobank participants
<p>GWAS summary statistics for 56 metabolites measured in 246,683 UK Biobank participants using the Nightingale Health platform. GWAS was performed using the https://github.com/AlasooLab/reGSusie workflow.</p>
Nationwide genomic biobank in Mexico unravels demographic history and complex trait architecture from 6,057 individuals: GWAS summary statistics
<p>Latin America continues to be severely underrepresented in genomics research, and fine-scale genetic histories as well as complex trait architectures remain hidden due to the lack of Big Data. To fill this gap, the Mexican Biobank project genotyped 1.8 million markers in 6,057 individuals from 32 states and 898 sampling localities across Mexico with linked complex trait and disease information creating a valuable nationwide genotype-phenotype database. Through a suite of state-of-the-art methods for ancestry deconvolution and inference of identity-by-descent (IBD) segments, we inferred detailed ancestral histories for the last 200 generations in different Mesoamerican regions, unravelling native and colonial/post-colonial demographic dynamics. We observed large variations in runs of homozygosity (ROH) among genomic regions with different ancestral origins reflecting their demographic histories, which also affect the distribution of rare deleterious variants across Mexico. We analysed a range of biomedical complex traits and identified significant genetic and environmental factors explaining their variation, such as ROH found to be significant predictors for trait variation in BMI and triglycerides.<br> ======================================</p> <p>This dataset contains GWAS summary statistics for the Mexico Biobank Project. Summary statistics for 22 binary and quantitative traits are provided from the full cohort of 5721 individuals from across Mexico, and a subset of 1061 individuals inferred to have more than 90% Native American ancestry.</p>
DOCUMENTED HUMAN OSTEOLOGICAL COLLECTIONS AS BIOBANKS: RELEVANCE FOR RARE DISEASES IDENTIFICATION IN THE PAST
<p><em><strong>Presented at: 23rd Paleopathology Association European Meeting, Vilnius, Lituânia, 25-29 Agosto. Paleopathology Association European (Vilnius, Lituânia)</strong></em></p> <p>Disease identification in paleopathology relies on the exercise of differential diagnosis, and interpretation. Only a few diseases leave macroscopic pathognomonic traits in bone, and even in cases where microscopic, biochemical and biomolecular analyses are used, diagnosis is invariably inconclusive. Additionally, bone response to a variety of etiologies tends to be homogenous, with mosaic pattern(s) of bone formation and destruction. Therefore, access to pathological cases from human remains of Documented Human Osteological Collections (DHOC) is an exceptional approach. The access to biographical data of the individuals incorporated into the DHOC includes the cause of death, ancestry, sex, age, clinical data and other information akin to clinical data allowing for the possibility of hypothesis-driven research in which bones changes correlate with causes of death - hence providing tested and informed differential diagnosis. In this sense, DHOC may be viewed as a biobank equivalent, i.e. biorepository that stores biological samples for research in the identification of bone changes related to diseases associated with clinical and personal data. This paper will explore known cases of diseases’ diagnoses, such as lepra, neoplasias, tuberculosis, syphilis, and diffuse idiopathic skeletal hyperostosis that have used DHOC as diagnostic testing grounds, to explore bone changes and methodological advancements. The paper also introduces the idea of DHOC as biobanks dedicated to the study of rare diseases, as rarely reposted diseases, in paleopathology. </p> <p><strong>Keywords: </strong>Health, biorepository, biobanks, DHOC, differential diagnosis</p>
PLINK association test statistics of UK Biobank blood traits
<p>PLINK association test statistics of UK Biobank blood traits generated for mvSuSiE fine-mapping analyses; see https://www.biorxiv.org/content/10.1101/2023.04.14.536893.</p>
Evaluation of polygenic scoring methods in five biobanks
<p>Raw experimental data to investigate the performance of polygenic score development methods across five biobanks</p>
NeuroPsyBiT-BD Omics: Genomic & Epigenomic Biobank of Bipolar Disorder
ClinicalTrials.gov study NCT07173842. IPD Sharing: YES. Countries: 1. Publications: 16.
Neuroimmunology Registry and Biobank
ClinicalTrials.gov study NCT06958341. IPD Sharing: NO. Countries: 1. Publications: 9.
Pleiotropy of UK Biobank metabolites
Open the record for dataset details and reuse information.
Improving genome-wide association discovery and genomic prediction accuracy in biobank data
Open the record for dataset details and reuse information.
Heritability analysis of the 38 cancers in the UK Biobank
<p>The dataset includes genome-wide and gene-level heritability analyses of the 38 cancers in the UK Biobank, using the Bayesian Gene HERitability Analysis (BAGHERA) method. The dataset provides also enrichment analyses results for cancer heritability genes with respect to Gene Ontology terms and cancer signatures. </p> <p>A detailed description of the dataset is provided in the MANIFEST.md file.</p>
Data from: Metabolomics - Emory Cardiovascular Biobank
<p>Untargeted high-resolution plasma metabolomic profiling among patients with coronary artery disease. Patients recruited from the Emory Cardiovascular Biobank into independent discovery and validation cohorts. </p>
GWAS summary statistics for 9 quantitative phenotypes from the UK Biobank (5-fold cross-validation)
<p>This dataset contains GWAS summary statistics for 9 quantitative phenotypes from the UK Biobank.</p> <p>The dataset is designed to enable systematic PRS analyses with 5-fold cross validation. For each phenotype and fold, we provide GWAS summary statistics for the training, validation, and test sets. The validation summary statistics can be used for model selection/tuning. The test summary statistics can be used to evaluate PRS models via pseudo-validation metrics. Association testing for all phenotypes and samples was done with <strong>plink2</strong>.</p> <p> </p> <p>The <strong>phenotypes</strong> included in this dataset are:</p> <ul> <li><a href="https://biobank.ndph.ox.ac.uk/showcase/field.cgi?id=50">HEIGHT</a>: Standing height (Data-Field: 50)</li> <li><a href="https://biobank.ndph.ox.ac.uk/showcase/field.cgi?id=21001">BMI</a>: Body mass index (Data-Field: 21001)</li> <li><a href="https://biobank.ndph.ox.ac.uk/showcase/field.cgi?id=48">WC</a>: Waist circumference (Data-Field: 48)</li> <li><a href="https://biobank.ndph.ox.ac.uk/showcase/field.cgi?id=49">HC</a>: Hip circumference (Data-Field: 49)</li> <li><a href="https://biobank.ndph.ox.ac.uk/showcase/field.cgi?id=20022">BW</a>: Birth weight (Data-Field: 20022)</li> <li><a href="https://biobank.ndph.ox.ac.uk/showcase/field.cgi?id=3062">FVC</a>: Forced vital capacity (Data-Field: 3062)</li> <li><a href="https://biobank.ndph.ox.ac.uk/showcase/field.cgi?id=3063">FEV1</a>: Forced expiratory volume in 1-second (Data-Field: 3063)</li> <li><a href="https://biobank.ndph.ox.ac.uk/showcase/field.cgi?id=30760">HDL</a>: HDL cholesterol (Data-Field: 30760)</li> <li><a href="https://biobank.ndph.ox.ac.uk/showcase/field.cgi?id=30780">LDL</a>: LDL cholesterol (Data-Field: 30780)</li> </ul> <p> </p> <p>To allow users to assess PRS performance as a function of sample size, we also provide <strong>subsampled training GWAS summary statistics</strong>. This is done by taking the training samples and randomly selecting (without replacement) a subset of them for conducting association testing. The training sample sizes are:</p> <ul> <li>N = 5000</li> <li>N = 10000</li> <li>N = 20000</li> <li>N = 40000</li> <li>N = 80000</li> <li>N = 160000</li> <li>Full training set (sample size varies by phenotype).</li> </ul> <p><strong>NOTE</strong>: Due to the smaller overall sample size for the Birth weight phenotype, we do not include training data for the `N=160000` setting.<br><br></p> <p>The <strong>folder structure</strong> of the GWAS data for each phenotype is as follows:</p> <ul> <li><code>train</code> <ul> <li><code>N_5000</code> <ul> <li><code> fold_1</code> <ul> <li><code>chr_1.PHENO1.glm.linear</code></li> <li><code>chr_2.PHENO1.glm.linear</code></li> <li><code>...</code></li> </ul> </li> <li><code>fold_2</code></li> <li><code>fold_3</code></li> <li><code>...</code></li> </ul> </li> <li><code>N_10000</code></li> <li><code>N_20000</code></li> <li><code>N_40000</code></li> <li><code>N_80000</code></li> <li><code>N_160000</code></li> <li><code>full</code></li> </ul> </li> <li><code>validation</code> <ul> <li><code>fold_1</code> <ul> <li><code>chr_1.PHENO1.glm.linear</code></li> <li><code>chr_2.PHENO1.glm.linear</code></li> <li><code>...</code></li> </ul> </li> <li><code>fold_2</code></li> <li><code>fold_3</code></li> <li><code>...</code></li> </ul> </li> <li><code>test</code> <ul> <li><code>fold_1</code></li> <li><code>fold_2</code></li> <li><code>fold_3</code></li> <li><code>...</code></li> </ul> </li> </ul> <p>For more details about the GWAS study, Quality Control (QC) criteria, or other information, please consult our publication:</p> <p>Zabad, S., Gravel, S., & Li, Y. (2023). <strong>Fast and accurate Bayesian polygenic risk modeling with variational inference.</strong> The American Journal of Human Genetics, 110(5), 741–761. <a href="https://doi.org/10.1016/j.ajhg.2023.03.009" rel="nofollow">https://doi.org/10.1016/j.ajhg.2023.03.009</a></p> <p>If you use this data in your work, please cite the publication above.</p> <p> </p>
Genome-wide study on 72,298 individuals in Korean biobank data for 76 traits
<p>Summary statistics relevant to Nam et al. "Genome-wide study on 72,298 individuals in Korean biobank data for 76 traits".</p> <p>See README.txt for details.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.