Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
252
datasets available to search
ShareScore release 0.7.1
Dataset results
252 results for “GWAS”
Formatting hemiclone Drosophila melanogaster genotype data for GWAS
<p>Data and code for generating filtering and formatting of Drosophila melanogaster genotype data, from the Sussex LHM hemiclone population sample.</p>
GWAS Summary Statistics from "Sex and statin-related genetic associations at the PCSK9 gene locus – results of genome-wide association meta-analysis"
<p>GWAMA summary statistics of PCSK9 levels stratified by sex and statin useage in Europeans.</p> <p>When using this data, please cite:</p> <p>Pott, J., Kheirkhah, A., Gadin, J.R. <em>et al.</em> Sex and statin-related genetic associations at the <em>PCSK9</em> gene locus: results of genome-wide association meta-analysis. <em>Biol Sex Differ</em> <strong>15</strong>, 26 (2024). https://doi.org/10.1186/s13293-024-00602-6</p> <p>All txt files contain the following columns:</p> <ul> <li>markername (unique SNP ID)</li> <li>chr</li> <li>bp_hg19 (base position according to hg19)</li> <li>EA (effect allele)</li> <li>OA (other allele)</li> <li>EAF (effect allele frequency)</li> <li>info (minimal info score across all used studies)</li> <li>nSamples (sample size per SNP)</li> <li>nStudies (in case of double-stratified data: number of studies; in case of single-stratified data: 2, as it is a meta-analysis of the two double-stratified data sets)</li> <li>beta (effect estimate)</li> <li>SE (standard error)</li> <li>pval (p-value)</li> <li>I2 (SNP heterogeneity across studies)</li> <li>invalidAssoc (TRUE/FALSE flag if this variant was excluded in our analysis)</li> <li>reason4exclusion (reason why this SNP was excluded)</li> <li>phenotype (phenotyp setting)</li> </ul>
GWAS summary statistics for waist-to-hip ratio and body principal components
<p>This dataset contains genome-wide association summary statistics for waist-to-hip ratio (WHR), as well as those for body principal components (PCs). A subset of 387,139 unrelated, white British individuals were analyzed for WHR. PCs were combined from the summary statistics for WHR and 13 other anthropometric traits (body mass index, standing height, weight, hip circumference, waist circumference, arm lean mass (left), arm fat mass (left), leg lean mass (left), leg fat mass (left), trunk lean mass, trunk fat mass, body fat percentage, basal metabolic rate) provided by the Neale lab (http://www.nealelab.is/uk-biobank). All traits were inverse-rank normal transformed (by the Neale lab or ourselves for WHR).</p> <p>All effect sizes, including those for PCs, are standardized, i.e. they represent the effects on a trait with variance 1.</p> <p>The zip files contain the data to run the sample pipeline and the shiny app, both available from <a href="https://github.com/JonSulc/PCA_Cross-sex_MR">https://github.com/JonSulc/PCA_Cross-sex_MR</a>.</p>
Phenotype data for Sussex LHM Drosophila melanogaster reproductive fitness GWAS
<p>Input data, code, logs, graphs and output data for the Sussex LHM Drosophila melanogaster hemiclones.</p> <p>Aim is to generate single, standardised values of female and male reproductive fitness for each hemiclone genome, for using in genome-wide association test using Plink software.</p> <p>Notes on how to run are provided in the code.</p>
Bivariate GWAS for female and male fitness in Drosophila melanogaster (Sussex, LHm)
<p>Code, logs, results and graphs for genome-wide association study of reproductive fitness in D.melanogaster hemiclone lines, using the R package 'Mulitphen'.</p>
Summary statistic of a Trans ancestry multi-trait GWAS
<p>Trans ancestry multi-trait GWAS by adapting the omnibus test to the trans ancestry setting. </p><p>Genome wide summary statistics for 19 blood count traits were retrieved from Chen et al paper and were downloaded from the GWAS Catalog (<a href="https://www.ebi.ac.uk/gwas/publications/32888493#study_panel">https://www.ebi.ac.uk/gwas/publications/32888493#study_panel)</a></p><p>They curated using the JASS (Joint Analysis of Summary Statistics) pipeline https://gitlab.pasteur.fr/statistical-genetics/jass_suite_pipeline</p><p>See Troubat et al preprint for all details on the obtention of this dataset https://doi.org/10.1101/2023.06.23.546248</p>
GWAS summary stats in "Genome-wide association meta-analysis identifies two novel loci associated with dental caries."
<p>Summary stats of the genome-wide meta-analysis for dental caries and periodontal diseases in our study (population A and B).</p> <p>Article "Genome-wide association meta-analysis identifies two novel loci associated with dental caries."</p> <p>https://doi.org/10.1186/s12903-024-04799-1<br><br></p>
Collider Bias Correction for Multiple Covariates in GWAS Using Robust Multivariable Mendelian Randomization
<p>This repository contains the data underlying the figures in paper "Collider Bias Correction for Multiple Covariates in GWAS<br>Using Robust Multivariable Mendelian Randomization".</p> <p> </p> <p> </p> <p>The file names and sheet names in the xlsx file indicate the corresponding figures of data. </p> <p><br>The underlying data of manhattan plots and QQ plots are in text file. For other figures, the underlying data are in the spreadsheet.</p> <p>In each file, column names indicate the MVMR method used to obtain the result. </p> <p>For example: </p> <p>In text files:</p> <p>The abbreviation "mPC" refers to metabolomic principle components.</p> <p>beta_no_correction: the SNP effect estimate without bias correction.</p> <p>beta_cml or beta_MVMR_cml: the standard error of SNP effect estimate after the bias correction of MVMR-cML.</p> <p>SE_UVMR_cml: the standard error of SNP effect estimate after the bias correction of UVMR-cML.</p> <p>p_value_Egger or p_value_MVMR_Egger: the p-value of SNP effect estimate after the bias correction of MVMR-Egger regression.</p> <p><br>In the spreadsheet, column names follow the same style. </p> <p>The GWAS data is also available. The column names follows the plink output file. The detailed explanations are available at https://www.cog-genomics.org/plink/2.0/formats#glm_linear</p>
GWAS Summary Statistics for Publication: Identifying novel genetic and phenotypic associations to genomic features by leveraging off-target reads in exome sequencing data
<p>This dataset contains summary statistics for genome-wide association studies (GWAS) conducted on genomic features derived from off-target reads in whole-exome sequencing (WES) data. The study utilized tools like Seeing Beyond the Target (SBT) and ImReP to construct novel phenotypic features from unmapped reads in ~50,000 participants in the UK Biobank. Features include mitochondrial DNA (mtDNA) copy number, ribosomal DNA (rDNA) copy number (5S, 18S, 28S), immune repertoire metrics (e.g., T-cell receptor alpha diversity), and microvial genome load (viral and fungal).</p> <p>Summary statistics can be used for replication studies, meta-analyses, or further exploration of these phenotypes.</p>
Summary Statistics from "Meta-GWAS of PCSK9 levels detects two novel loci at APOB and TM6SF2"
<p>GWAMA summary statistics of PCSK9 levels using fixed-effect model. Genome-wide data is given for Europeans with statin adjustment and Europeans without statin treatment only (subset of the population). In addition, locus-wide data of the PCSK9 gene locus for African-Americans without statin treatment is listed.</p> <p>When using this data, please cite: Pott J, Gadin J, Theusch E, et al.. Meta-GWAS of PCSK9 levels detects two novel loci at APOB and TM6SF2. Hum Mol Genet. 2021 Sep 30:ddab279. doi: 10.1093/hmg/ddab279. PMID: 34590679</p> <p>All txt files contain the following columns:</p> <ul> <li>markername</li> <li>chr</li> <li>bp_hg19 (base position according to hg19)</li> <li>ea (effect allele)</li> <li>oa (other allele)</li> <li>eaf (effect allele frequency)</li> <li>info (minimal info score across all used studies)</li> <li>nSamples (sample size per SNP)</li> <li>nStudies (number of studies)</li> <li>beta (effect estimate)</li> <li>se (standard error)</li> <li>p (p-value)</li> <li>I2 (SNP heterogeneity across studies)</li> <li>phenotype (phenotyp setting)</li> </ul>
Separate-sex GWAS for reproductive fitness in Drosophila melanogaster (Sussex LHM sample)
<p>Code, data, logs, and graphs for GWAS on seperate-sex reproductive fitness in Drosophila melanogaster, Sussex LHM population sample.</p> <p>The shell script, code_drive_basic_gwas.sh, downloads input data files from the internet, drives Plink to select LD-independent SNPs, and then perform a genome-wide association test against female and male fitness, separately. Plink is also used to assign functions and gene names to SNPs. Bash/Unix code is used for formatting/compatibility adjustments, and also to add NCBI-dbSNP IDs to results. The shell script starts an R script that generates basic diagnostic graphs. This updated version differs from the first in that three large unconfirmed snRNA genes have been omitted to improve assignment of SNPs to genes.</p> <p>See https://f1000research.com/articles/5-2644/v3 and http://www.sussex.ac.uk/lifesci/morrowlab/</p>
GWAS on self-reported hearing difficulty in the UK Biobank
<p>The dataset contains results of two genome-wide association studies for age-related hearing impairment (ARHI)-related traits as described in the following publication</p> <p><a href="https://www.sciencedirect.com/science/article/pii/S0002929719303477#!"><em>Wells HRR, Freidin MB, Zainul Abidin FN, Payton A, Dawes P, Munro KJ, Morton CC, Moore DR, Dawson SJ, Williams FMK. GWAS Identifies 44 Independent Associated Genomic Loci for Self-Reported Adult Hearing Difficulty in UK Biobank. Am J Hum Genet. 2019 Oct 3;105(4):788-802. doi: 10.1016/j.ajhg.2019.09.008. Epub 2019 Sep 26.</em></a> </p> <p>Please cite the article if using this dataset.</p> <p>Two files provide summary statistics for discovery analysis of <em><strong>Hearing difficulty (HD) </strong></em>and <em><strong>Hearing aid use (HAID)</strong></em> phenotypes for individuals of European descent from <a href="https://www.ukbiobank.ac.uk/">UK Biobank</a>.</p> <p><strong>Acknowledgements</strong></p> <p>The research was carried out using the UK Biobank Resource under application number 11516. H.R.R.W. is funded by a PhD Studentship Grant, S44, from Action on Hearing Loss. The study was also supported by funding from NIHR UCLH BRC Deafness and Hearing Problems Theme, a grant from MED_EL, and the NIHR Manchester Biomedical Research Centre. The English Longitudinal Study of Aging is jointly run by University College London, Institute for Fiscal Studies, University of Manchester, and National Centre for Social Research. Genetic analyses have been carried out by UCL Genomics and funded by the Economic and Social Research Council and the National Institute on Aging. Data governance was provided by the METADAC data access committee, funded by ESRC, Wellcome, and MRC (2015-2018: Grant Number MR/N01104X/1 2018-2020: Grant Number ES/S008349/1). TwinsUK is funded by the Wellcome Trust, Medical Research Council, European Union, the National Institute for Health Research (NIHR)-funded BioResource, Clinical Research Facility, and Biomedical Research Centre based at Guy’s and St Thomas’ NHS Foundation Trust in partnership with King’s College London. We would like to thank all the participants of UK Biobank, English Longitudinal Study of Aging, and TwinsUK.</p> <p><strong>Column headers:</strong></p> <p>SNP, SNP rsID</p> <p>CHR, chromosome</p> <p>BP, genomic position (GRCh37 build) </p> <p>ALLELE1, effect allele (coded as "1")</p> <p>ALLELE0, reference allele (coded as "0") </p> <p>A1FREQ, effect allele frequency</p> <p>INFO, imputation quality</p> <p>BETA, effect size of effect allele</p> <p>SE: standard error of effect size</p> <p>P, P-value of association (without GC correction)</p>
Database for GWAS SVatalog: a visualization tool to aid fine-mapping of GWAS loci with structural variations.
<p>GWAS SVatalog is a novel visualization tool and database for structural variants (SV) found in a predominantly European population of 101 individuals with Cystic Fibrosis (CF). Aside from the CF-causing variants on chromosome 7 and the LD block in which they lie, the remainder of the genome is comparable to a the 1000 Genomes healthy European population. This data is a collection of SV calls and their linkage disequilibrium (LD) statistics with GWAS-significant SNPs reported in the GWAS Catalog.</p> <p> </p> <p>The goal of this project is to provide a resource to aid fine mapping of GWAS loci using SVs. GWAS loci are generally identified by SNPs which account for an incomplete proportion of genetic variation and phenotypic heritability. Their relevance to the phenotype might be limited, tagging other polymorphisms, such as SVs, that could be the cause of the association signal. To leverage this data to its full potential, visit the <a href="https://svatalog.research.sickkids.ca/" target="_blank" rel="noopener">GWAS SVatalog</a> web tool. Here, interactive visualizations can illustrate SVs identified in high LD with GWAS-significant SNPs, suggesting putative causal variation that could guide additional functional investigation.</p> <p> </p> <p>For more information on how to use GWAS SVatalog, visit the<a href="https://gwas-svatalog-docs.readthedocs.io/en/latest/index.html" target="_blank" rel="noopener noreferrer"> documentation</a>.</p> <p> </p> <p>This project was accomplished in collaboration with the <a href="https://lab.research.sickkids.ca/strug/" target="_blank" rel="noopener">Strug Lab</a> at <a href="https://www.sickkids.ca/en/" target="_blank" rel="noopener">The Hospital for Sick Children (SickKids)</a>, <a href="https://www.tcag.ca/" target="_blank" rel="noopener">The Center for Applied Genomics (TCAG)</a>, and <a href="https://www.utoronto.ca/" target="_blank" rel="noopener">University of Toronto</a>.</p>
GWAS to single cell: Intersecting single-cell transcriptomics and genome wide association studies identifies crucial cell-populations and candidate genes for atherosclerosis.
<p><strong>Background</strong></p> <p>Genome-wide association studies (GWAS) have discovered hundreds of common genetic variants for atherosclerotic disease and cardiovascular risk factors. The translation of susceptibility loci into biological mechanisms and targets for drug discovery remains challenging. Intersecting genetic and gene expression data has led to identification of candidate genes. However, the assayed tissues are often non-diseased and heterogeneous in cell composition confounding the candidate prioritization. We collected single-cell transcriptomics (scRNA-seq) from atherosclerotic plaques and aimed to identify cell-type-specific expression of disease-associated genes. </p> <p> </p> <p><strong>Methods and Results</strong></p> <p>To identify disease-associated candidate genes, we applied gene-based analyses using GWAS summary statistics from 46 atherosclerotic, cardiometabolic, and other traits. Next we intersected these candidates with single-cell transcriptomics (scRNA-seq) to identify those genes that are specifically expressed in individual cell (sub)populations of atherosclerotic plaques. We derive an enrichment score and show that loci that associated with coronary artery disease demonstrated a prominent substrate in plaque smooth muscle cells (<em>SKI</em>, <em>KANK2</em>, <em>SORT1</em>), endothelial cells (<em>SLC44A1</em>, <em>ATP2B1</em>), and macrophages (<em>APOE</em>, <em>HNRNPUL1</em>). Further sub clustering of SMC-subtypes revealed genes in risk loci for coronary calcification specifically enriched in a synthetic cluster of SMCs. To verify the robustness of our approach, we used liver-derived scRNAseq-data and showed enrichment of circulating lipids-associated loci in hepatocytes.</p> <p><br> <strong>Conclusion</strong></p> <p>We confirm known gene-cell pairs relevant for atherosclerotic disease, and discovered novel pairs pointing to new biological mechanisms amenable for therapy. We present an intuitive single-cell transcriptomics driven workflow rooted in human large-scale genetic studies to identify putative candidate genes and affected cells associated with cardiovascular traits.</p> <p> </p>
Curated GWAS summary statistics on African ancestry on 19 blood count traits and glycemic traits (hg38)
<p>Genome wide curated summary statistics on 19 blood count traits and glycemic traits</p> <p>File format is the inittable format intended to be used with the Joint Analysis of Summary Statistics (JASS), which allows to perform multi-trait GWAS:</p> <p>https://gitlab.pasteur.fr/statistical-genetics/jass</p> <p>GWAS of hematological traits originate from Chen et al paper and were downloaded from the GWAS Catalog (<a href="https://www.ebi.ac.uk/gwas/publications/32888493#study_panel">https://www.ebi.ac.uk/gwas/publications/32888493#study_panel</a>). GWAS of glycemic traits come from the <a href="https://www.zotero.org/google-docs/?S1MIfx">(18)</a> study downloadable from GWAS Catalog (<a href="https://www.ebi.ac.uk/gwas/publications/34059833">https://www.ebi.ac.uk/gwas/publications/34059833</a>).</p> <p> </p>
Curated GWAS summary statistics on East Asian ancestry on 19 blood count traits and glycemic traits
<p>Genome wide curated summary statistics on 19 blood count traits and glycemic traits</p> <p>File format is the inittable format intended to be used with the Joint Analysis of Summary Statistics (JASS), which allows to perform multi-trait GWAS:</p> <p>https://gitlab.pasteur.fr/statistical-genetics/jass</p> <p>GWAS of hematological traits originate from Chen et al paper and were downloaded from the GWAS Catalog (<a href="https://www.ebi.ac.uk/gwas/publications/32888493#study_panel">https://www.ebi.ac.uk/gwas/publications/32888493#study_panel</a>). GWAS of glycemic traits come from the <a href="https://www.zotero.org/google-docs/?S1MIfx">(18)</a> study downloadable from GWAS Catalog (<a href="https://www.ebi.ac.uk/gwas/publications/34059833">https://www.ebi.ac.uk/gwas/publications/34059833</a>).</p> <p>Full description of the method used to derive this dataset can be found in </p>
Common factor GWAS and TWAS output for nociplastic type pain
<p>file: GSEM_commonFactorGWAS_COPC_6trait_30MAY2023.csv.gz</p> <p>description: Common factor GWAS output for GenomicSEM analyses of 6 COPC traits (see doi: https://doi.org/10.1101/2023.06.27.23291959)</p> <p>columns:</p> <p>SNP = rsID SNP identifier </p> <p>CHR = chromosome</p> <p>BP = base pair position</p> <p>MAF = minor allele frequency</p> <p>A1 = effect allele</p> <p>A2 = other allele</p> <p>i = index (1 - n SNPs)</p> <p>lhs = left hand side of equation </p> <p>op = equation operator (lavaan syntax)</p> <p>rhs = right hand side of equation</p> <p>est = effect size (beta)</p> <p>se_c = standard error of effect estimate</p> <p>Z_Estimate = Z value</p> <p>Pval_Estimate = p value of effect</p> <p>Q = Q (heterogeneity) value</p> <p>Q_df = degrees of freedom for Q</p> <p>Q_pval = Q p value</p> <p>fail = GSEM fail message if applicable</p> <p>warning = GSEM warning message if applicable </p> <p>Z_smooth = smoothing parameter if applicable </p> <p>N_estimate = N estimate </p> <p> </p> <p>file: GSEM_commonFactorTWAS_COPC_6trait_30MAY2023.csv.gz</p> <p>description: Common factor TWAS output for GenomicSEM analyses of 6 COPC traits (see doi: https://doi.org/10.1101/2023.06.27.23291959)</p> <p>columns:</p> <p>Gene = ensembl gene ID </p> <p>Panel = which model (tissue+gene) i.e. reference weights file</p> <p>HSQ = gene heritability </p> <p>i = index (1 - n gene-tissue models)</p> <p>lhs = equation left hand side</p> <p>op = operator (lavaan syntax)</p> <p>rhs = equation right hand side</p> <p>est = association estimate (beta)</p> <p>se_c = standard error of beta</p> <p>Z_Estimate = Z value </p> <p>Pval_Estimate = p value of association test</p> <p>Q = Q (heterogeneity) value</p> <p>Q_df = degrees of freedom for Q</p> <p>Q_pval = p value for Q</p> <p>fail = GSEM fail message if applicable </p> <p>warning = GSEM warning message if applicable </p> <p>tissue = tissue</p> <p>p_bonf_tissue = adjusted p value - bonferroni adjustment within tissue </p> <p>p_fdr_tissue = adjusted p value - false discovery rate adjustment within tissue</p> <p>threshold_bonf_tissue = p value threshold for bonferroni adjustment within tissue </p> <p>p_bonf_experiment = adjusted p value - bonferroni adjustment experiment-wide</p> <p>p_threshold_bonf_experiment = p value threshold (bonferroni, experiment-wide)</p> <p>Q_bonf_tissue = adjusted p value for Q, bonferroni within-tissue </p> <p> </p>
G2G-EBV GWAS summary statistics
<p>G2G results from manuscript titled:</p> <p><strong>The influence of human genetic variation on Epstein-Barr virus sequence diversity</strong></p> <p>The GWAS result files (*.mlma) contains:<br> chromosome, SNP, physical position, reference allele (the coded effect allele), the other allele, frequency of the reference allele, SNP effect, standard error and p-value (<a href="https://cnsgenomics.com/software/gcta/#MLMA">GCTA-MLMA</a>).<br> <br> The scripts used to generate the GWAS results, are available here: <a href="https://github.com/sinarueeger/G2G-EBV-manuscript">github.com/sinarueeger/G2G-EBV-manuscript</a>.</p>
EMMA-X GWAS results on S. pimpinellifolium RSA traits - part 2
<p>The results of GWAS run through EMMA-X pipeline, on root system architecture traits of S. pimpinellifolim (wild tomato) exposed to control and salt stress conditions on agar plates. Traits are abbreviated as: aLRG for average lateral root growth rate, aLRL for average lateral root length, aLRLpTRS for the ratio between average lateral root length and total root size. The treatments are abbreviated with C or S for control or 100 mM NaCl treatments. The numbers in file name (0-4) indicate the days post transfer to agar plates containing treatment. </p>
EMMA-X GWAS results on S. pimpinellifolium RSA traits - part 1
<p>The results of GWAS run through EMMA-X pipeline, on root system architecture traits of S. pimpinellifolim (wild tomato) exposed to control and salt stress conditions on agar plates. Traits are abbreviated as: aLRG for average lateral root growth rate, aLRL for average lateral root length, aLRLpTRS for the ratio between average lateral root length and total root size. The treatments are abbreviated with C or S for control or 100 mM NaCl treatments. The numbers in file name (0-4) indicate the days post transfer to agar plates containing treatment. </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.