Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
186
datasets available to search
ShareScore release 0.9.0
Dataset results
186 results for “genome wide association study”
Genome-wide association study in quinoa reveals selection pattern typical for crops with a short breeding history
Open the record for dataset details and reuse information.
Combining genome-wide studies of breast, prostate, ovarian and endometrial cancers maps cross-cancer susceptibility loci and identifies new genetic associations
<p>Data set linked to the paper, "Combining genome-wide studies of breast, prostate, ovarian and endometrial cancers maps cross-cancer susceptibility loci and identifies new genetic associations". Pre-print of the paper is here: <a href="https://doi.org/10.1101/2020.06.16.146803">https://doi.org/10.1101/2020.06.16.146803</a>.</p> <p> </p> <p>cross_cancer_sum_stats.txt.gz contains summary genome-wide association statistics for susceptibility to single cancers (breast (BR), prostate (PR), ovarian (OV), endometrial (EN), estrogen receptor (ER)-positive breast (POS), ER-negative breast (NEG), and high-grade serous ovarian (HGS) cancers) and from the cross-cancer meta-analysis (main [main] and subtype-focused [sub]). EA in the header refers to the effect allele, OA is the other allele, EAF is the effect allele frequency in the largest of the single cancer data sets (BR), IMPR2 is the imputation quality in the largest of the single cancer data sets (BR), SE is the standard error, PVAL is the P-value, RE2Cs1 is the RE2C statistic mean effect part, RE2Cs2 is the RE2C statistic heterogeneity part, RE2Cp* is the RE2C* P-value. More on RE2Cp* can be found here: <a href="http://software.buhmhan.com/RE2C/index.php?mid=contact&act=dispBoardWrite">http://software.buhmhan.com/RE2C/index.php?mid=contact&act=dispBoardWrite</a> and in <a href="https://academic.oup.com/bioinformatics/article/33/14/i379/3953957">https://academic.oup.com/bioinformatics/article/33/14/i379/3953957</a> SNP names in cross_cancer_sum_stats.txt.gz include the chromosome and build 37 position.</p> <p> </p> <p>main_tetrachoric_corr_matrix.txt and subtype_tetrachoric_corr_matrix.txt provide the tetrachoric correlation matrices used in the main and subtype-focused meta-analyses. These were also used to specify the cryptic.cor argument of the exh.abf function of MetABF. More on MetABF can be found here: <a href="https://github.com/trochet/metabf">https://github.com/trochet/metabf</a> and in <a href="https://onlinelibrary.wiley.com/doi/abs/10.1002/gepi.22202">https://onlinelibrary.wiley.com/doi/abs/10.1002/gepi.22202</a></p> <p> </p> <p>prior_sigmas_for_metabf.txt contains the values used to specify the prior.sigma argument of the exh.abf function in MetABF.</p> <p> </p> <p>The breast cancer data used are described in <a href="https://pubmed.ncbi.nlm.nih.gov/29059683/"><strong>PMID 29059683</strong></a> and can be downloaded from <a href="http://bcac.ccge.medschl.cam.ac.uk/bcacdata/oncoarray/oncoarray-and-combined-summary-result/gwas- summary-results-breast-cancer-risk-2017/">http://bcac.ccge.medschl.cam.ac.uk/bcacdata/oncoarray/oncoarray-and-combined-summary-result/gwas- summary-results-breast-cancer-risk-2017/</a> (this link also includes acknowledgements). The prostate cancer data are described in <a href="https://pubmed.ncbi.nlm.nih.gov/29892016/"><strong>PMID 29892016</strong></a> and can be downloaded from: <a href="http://practical.icr.ac.uk/blog/?page_id=8164">http://practical.icr.ac.uk/blog/?page_id=8164</a> (this link also includes acknowledgements). The ovarian cancer data used are described in <a href="https://pubmed.ncbi.nlm.nih.gov/28346442/"><strong>PMID 28346442</strong></a> and can be downloaded from <a href="https://www.ebi.ac.uk/gwas/studies/GCST004415">https://www.ebi.ac.uk/gwas/studies/GCST004415</a>. The endometrial cancer data are described in <a href="https://pubmed.ncbi.nlm.nih.gov/30093612/"><strong>PMID 30093612</strong></a> and can be downloaded from <a href="https://www.ebi.ac.uk/gwas/studies/GCST006464">https://www.ebi.ac.uk/gwas/studies/GCST006464</a>. These links point to the same data that form the basis of the cross_cancer_sum_stats.txt.gz file.</p> <p> </p> <p><strong>The sample size and precision of the data presented should preclude identification of any individual study participant. However, in downloading these data, you undertake not to attempt to identify individual study participant and not to re-post these data to a third-party website. Please cite the PMIDs highlighted above along with the appropriate acknowledements if you use the cross_cancer_sum_stats.txt.gz file.</strong></p> <p> </p> <p>If you have any questions about this repository, please email Siddhartha Kar at siddhartha dot kar at bristol dot ac dot uk</p>
Association of common genetic variants with brain microbleeds: A genome-wide association study
<p><strong>Objective:</strong> To identify common genetic variants associated with the presence of brain microbleeds (BMB).</p> <p><strong>Methods:</strong> We performed genome-wide association studies in 11 population-based cohort studies and 3 case-control or case-only stroke cohorts. Genotypes were imputed to the Haplotype Reference Consortium or 1000 Genomes reference panel. BMB were rated on susceptibility-weighted or T2*-weighted gradient echo magnetic resonance imaging sequences, and further classified as lobar, or mixed (including strictly deep and infratentorial, possibly with lobar BMB). In a subset, we assessed the effects of <em>APOE</em> ε2 and ε4 alleles on BMB counts. We also related previously identified cerebral small vessel disease variants to BMB.</p> <p><strong>Results: </strong>BMB were detected in 3,556 of the 25,862 participants, of which 2,179 were strictly lobar and 1,293 mixed. One locus in the <em>APOE</em> region reached genome-wide significance for its association with BMB (lead SNP rs769449; OR<sub>any BMB</sub> (95% CI)=1.33 (1.21-1.45); p=2.5x10-10). <em>APOE</em> ε4 alleles were associated with strictly lobar (OR (95% CI)=1.34 (1.19- 1.50); p=1.0x10-6) but not with mixed BMB counts (OR (95% CI)=1.04 (0.86-1.25); p=0.68). <em>APOE</em> ε2 alleles did not show associations with BMB counts. Variants previously related to deep intracerebral hemorrhage and lacunar stroke, and a risk score of cerebral white matter hyperintensity variants, were associated with BMB.</p> <p><strong>Conclusions: </strong>Genetic variants in the <em>APOE</em> region are associated with the presence of BMB, most likely due to the <em>APOE</em> ε4 allele count related to a higher number of strictly lobar BMB. Genetic predisposition to small vessel disease confers risk of BMB, indicating genetic overlap with other cerebral small vessel disease markers.</p>
Data from: Combining high-throughput micro-CT-RGB phenotyping and genome-wide association study to dissect the genetic architecture of tiller growth in rice
Manual phenotyping of rice tillers is time consuming and labor intensive and lags behind the rapid development of rice functional genomics. Thus, automated, non-destructive phenotyping of rice tiller traits at a high spatial resolution and high-throughput for large-scale assessment of rice accessions is urgently needed. In this study, we developed a high-throughput micro-CT-RGB (HCR) imaging system to non-destructively extract 730 traits from 234 rice accessions at 9 time points. We could explain 30% of the grain yield variance from 2 tiller traits assessed in the early growth stages. A total of 402 significantly associated loci were identified by GWAS, and dynamic and static genetic components were found across the nine time points. A major locus associated with tiller angle was detected at nine time points, which contained a major gene TAC1. Significant variants associated with tiller angle were enriched in the 3'-UTR of TAC1. Three haplotypes for the gene were found and rice accessions containing haplotype H3 displayed much smaller tiller angles. Further, we found two loci contained associations with both vigor-related HCR traits and yield. The superior alleles would be beneficial for breeding of high yield and dense planting.
Data from: Genome-wide association study in Arabidopsis thaliana of natural variation in seed oil melting point, a widespread adaptive trait in plants
Seed oil melting point is an adaptive, quantitative trait determined by the relative proportions of the fatty acids that compose the oil. Micro- and macro-evolutionary evidence suggests selection has changed the melting point of seed oils to covary with germination temperatures because of a trade-off between total energy stores and the rate of energy acquisition during germination under competition. The seed oil compositions of 391 natural accessions of Arabidopsis thaliana, grown under common-garden conditions, were used to assess whether seed oil melting point within a species varied with germination temperature. In support of the adaptive explanation, long-term monthly spring and fall field temperatures of the accession collection sites significantly predicted their seed oil melting points. In addition, a genome-wide association study (GWAS) was performed to determine which genes were most likely responsible for the natural variation in seed oil melting point. The GWAS found a single highly significant association within the coding region of FAD2, which encodes a fatty acid desaturase central to the oil biosynthesis pathway. In a separate analysis of fifteen a priori oil synthesis candidate genes, two (FAD2 and FATB) were located near significant SNPs associated with seed oil melting point. These results comport with others' molecular work showing that lines with alterations in these genes affect seed oil melting point as expected. Our results suggest natural selection has acted on a small number of loci to alter a quantitative trait in response to local environmental conditions.
Data from: Penalized Multi-Marker versus Single-Marker Regression methods for genome-wide association studies of quantitative traits
The data from genome-wide association studies (GWAS) in humans are still predominantly analyzed using single marker association methods. As an alternative to Single Marker Analysis (SMA), all or subsets of markers can be tested simultaneously. This approach requires a form of Penalized Regression (PR) as the number of SNPs is much larger than the sample size. Here we review PR methods in the context of GWAS, extend them to perform penalty parameter and SNP selection by False Discovery Rate (FDR) control, and assess their performance in comparison with SMA. PR methods were compared with SMA using realistically simulated GWAS data with a continuous phenotype and real data. Based on these comparisons our analytic FDR criterion may currently be the best approach to SNP selection using PR for GWAS. We found that PR with FDR control provides substantially more power than SMA with genome-wide type-I error control but somewhat less power than SMA with Benjamini-Hochberg FDR control (SMA-BH). PR with FDR based penalty parameter selection controlled the FDR somewhat conservatively while SMA-BH may not achieve FDR control in all situations. Differences among PR methods seem quite small when the focus is on SNP selection with FDR control. Incorporating linkage disequilibrium into the penalization by adapting penalties developed for covariates measured on graphs can improve power but also generate more false positives or wider regions for follow-up. We recommend the Elastic Net with a mixing weight for the Lasso penalty near 0.5 as the best method.
Data from: Genome-wide association study identifies vitamin B5 biosynthesis as a host specificity factor in Campylobacter
Genome-wide association studies have the potential to identify causal genetic factors underlying important phenotypes but have rarely been performed in bacteria. We present an association mapping method that takes into account the clonal population structure of bacteria and is applicable to both core and accessory genome variation. Campylobacter is a common cause of human gastroenteritis as a consequence of its proliferation in multiple farm animal species and its transmission via contaminated meat and poultry. We applied our association mapping method to identify the factors responsible for adaptation to cattle and chickens among 192 Campylobacter isolates from these and other host sources. Phylogenetic analysis implied frequent host switching but also showed that some lineages were strongly associated with particular hosts. A seven-gene region with a host association signal was found. Genes in this region were almost universally present in cattle but were frequently absent in isolates from chickens and wild birds. Three of the seven genes encoded vitamin B5 biosynthesis. We found that isolates from cattle were better able to grow in vitamin B5-depleted media and propose that this difference may be an adaptation to host diet.
Data from: Genome-wide association study of a Varroa-specific defense behavior in honeybees (Apis mellifera)
Honey bees are exposed to many damaging pathogens and parasites. The most devastating is Varroa destructor, which mainly affects the brood. A promising approach for preventing its spread is to breed Varroa-resistant honey bees. One trait that has been shown to provide significant resistance against the Varroa mite is hygienic behavior, which is a behavioral response of honeybee workers to brood diseases in general. Here we report the use of an Affymetrix 44K SNP array to analyze SNPs associated with detection and uncapping of Varroa-parasitized brood by individual worker bees (Apis mellifera). For this study, 22,000 individually labeled bees were video-monitored and a sample of 122 cases and 122 controls was collected and analyzed to determine the dependence / independence of SNP genotypes from hygienic and non-hygienic behavior on a genome-wide scale. After false-discovery rate correction of the p-values, six SNP markers had highly significant associations with the trait investigated (alpha < 0.01). Inspection of the genomic regions around these SNPs led to the discovery of putative candidate genes.
Data from: Genome-wide association studies in apple reveal loci of large effect controlling apple polyphenols
Apples are a nutritious food source with significant amounts of polyphenols that contribute to human health and wellbeing, primarily as dietary antioxidants. Although numerous pre- and post-harvest factors can affect the composition of polyphenols in apples, genetics is presumed to play a major role because polyphenol concentration varies dramatically among apple cultivars. Here we investigated the genetic architecture of apple polyphenols by combining high performance liquid chromatography (HPLC) data with ~100,000 single nucleotide polymorphisms (SNPs) from two diverse apple populations. We found that polyphenols can vary in concentration by up to two orders of magnitude across cultivars, and that this dramatic variation was often predictable using genetic markers and frequently controlled by a small number of large effect genetic loci. Using GWAS, we identified candidate genes for the production of quercitrin, epicatechin, catechin, chlorogenic acid, 4-O-caffeoylquinic acid and procyanidins B1, B2, and C1. Our observation that a relatively simple genetic architecture underlies the dramatic variation of key polyphenols in apples suggests that breeders may be able to improve the nutritional value of apples through marker-assisted breeding or gene editing.
Data from: Genome-wide association study of insect bite hypersensitivity in Swedish-born Icelandic horses
Insect bite hypersensitivity (IBH) is the most common allergic skin disease in horses and is caused by biting midges, mainly of the genus Culicoides. The disease predominantly comprises a type I hypersensitivity reaction, causing severe itching and discomfort that reduce the welfare and commercial value of the horse. It is a multifactorial disorder influenced by both genetic and environmental factors, with heritability ranging from 0.16 to 0.27 in various horse breeds. The worldwide prevalence in different horse breeds ranges from 3% to 60%; it is more than 50% in Icelandic horses exported to the European continent and approximately 8% in Swedish-born Icelandic horses. To minimize the influence of environmental effects, we analyzed Swedish-born Icelandic horses to identify genomic regions that regulate susceptibility to IBH. We performed a genome-wide association (GWA) study on 104 affected and 105 unaffected Icelandic horses genotyped using Illumina® EquineSNP50 Genotyping BeadChip. Quality control and population stratification analyses were performed with the GenABEL package in R (λ = 0.81). The association analysis was performed using the Bayesian variable selection method, Bayes C, implemented in GenSel software. The highest percentage of genetic variance was explained by the windows on X chromosomes (0.51% and 0.36% by 73 and 74 mb), 17 (0.34% by 77 mb), and 18 (0.34% by 26 mb). Overlapping regions with previous GWA studies were observed on chromosomes 7, 9, and 17. The windows identified in our study on chromosomes 7, 10, and 17 harbored immune system genes and are priorities for further investigation.
Petal size in rapeseed: novel QTL and candidate genes detected through genome-wide association study and transcriptome comparison
<p>Petal size determines the value of ornamental plants, and thus their economic worth. However, the molecular mechanisms controlling petal size remain unclear in most non-model species. To identify quantitative trait loci and candidate genes regulating petal size in rapeseed (<i>Brassica napus</i>), we performed a genome-wide association study (GWAS) using data from 588 accessions over three consecutive years. We detected 17 significant single nucleotide polymorphisms (SNPs) associated with petal size, with the most significant SNPs located on chromosomes A05 and C06. A combination of GWAS and transcriptomic sequencing based on two accessions with extreme differences in petal size identified 11 differentially expressed genes (DEGs) that may control petal size variation in rapeseed. In particular, <i>BnaA05</i><i>.</i><i>RAP2.2</i> homologous to <i>RAP2.2</i> in rapeseed may be a critical gene negatively influencing petal size through the ethylene signaling pathway. In addition, a comparison of petal epidermal cells indicated that petal size differences between the two extreme accessions were determined mainly by cell number differences. Finally, we propose a preliminary model for the control of petal size in rapeseed. Our results provide insights into the genetic mechanisms regulating petal size, and also lay the foundation for a better understanding of petal development in plants.</p>
Code for manuscript A genome-wide association study of neonatal metabolites
Open the record for dataset details and reuse information.
Structural equation models to interpret genome-wide association studies for morphological and productive traits in soybean [Glycine max (L.)]
Open the record for dataset details and reuse information.
Data from: Polygenic adaptation on height is overestimated due to uncorrected stratification in genome-wide association studies
Genetic predictions of height differ among human populations and these differences have been interpreted as evidence of polygenic adaptation. These differences were first detected using SNPs genome-wide significantly associated with height, and shown to grow stronger when large numbers of sub-significant SNPs were included, leading to excitement about the prospect of analyzing large fractions of the genome to detect polygenic adaptation for multiple traits. Previous studies of height have been based on SNP effect size measurements in the GIANT Consortium meta-analysis. Here we repeat the analyses in the UK Biobank, a much more homogeneously designed study. We show that polygenic adaptation signals based on large numbers of SNPs below genome-wide significance are extremely sensitive to biases due to uncorrected population structure. More generally, our results imply that typical constructions of polygenic scores are sensitive to population structure and that population-level differences should be interpreted with caution.
Association of SUMOlation pathway genes with stroke in a genome-wide association study in India
<p><strong>Objective:</strong> To undertake a genome-wide association study (GWAS) to identify genetic variants for stroke in Indians.</p> <p><strong>Methods:</strong> In a hospital-based case-control study, eight teaching hospitals in India recruited 4,088 subjects, including 1,609 stroke cases. Imputed genetic variants were tested for association with stroke subtypes using both single-marker and gene-based tests. Association with vascular risk factors was performed using logistic regression. Various databases were searched for replication, functional annotation, and association with related traits. Status of candidate genes previously reported in the Indian population was also checked.</p> <p><strong>Results:</strong> Association of vascular risk factors with stroke were similar to previous reports, and show modifiable risk factors like hypertension, smoking, and alcohol consumption having the highest effect. Single-marker based association revealed two loci for cardioembolic stroke (1p21 and 16q24), two for small vessel disease stroke (3p26 and 16p13), and four for hemorrhagic stroke (3q24, 5q33, 6q13, and 19q13) at P<5×10-8. The index SNP of 1p21 is an eQTL (Plowest=1.74×10-58) for RWDD3 involved in SUMOlation and is associated with platelet distribution width (1.15×10-9) and 18-carbon fatty acid metabolism (P=7.36×10-12). In gene-based analysis we identified three genes (SLC17A2, FAM73A & OR52L1) at P<2.7×10-6. 11 of 32 candidate gene loci studied in Indians replicated (P<0.05), and 21 of 32 loci identified through previous GWAS replicated based on directionality of effect.</p> <p><strong>Conclusions:</strong> This first GWAS of stroke in Indians identified novel loci and replicated previously known loci. For the first time, genetic variants in the SUMOlation pathway which has been implicated in brain ischemia were identified.</p>
Data from: Molecular insights into genome-wide association studies of chronic kidney disease-defining traits
Open the record for dataset details and reuse information.
Data from: Genome-wide association study identifies vitamin B5 biosynthesis as a host specificity factor in Campylobacter
Open the record for dataset details and reuse information.
Data from: Genome-wide association study of behavioral, physiological and gene expression traits in outbred CFW mice
Open the record for dataset details and reuse information.
Data from: Genome-wide association study of Arabidopsis thaliana identifies determinants of natural variation in seed oil composition
Open the record for dataset details and reuse information.
Data from: Genome-wide association study of a Varroa-specific defense behavior in honeybees (Apis mellifera)
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.