Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
76
datasets available to search
ShareScore release 0.9.0
Dataset results
76 results for “Genetic association study”
A scalable, accurate, and universal analysis framework using individual-level allele frequency for large-scale genetic association studies in an admixed population
<p>Inclusion of individuals with diverse or admixed genetic ancestries is crucial to discover novel findings that may be missed by genomics analyses rooted solely in Caucasian population. Here, we present an analysis framework, SPAmix, which is scalable to a large-scale biobank data analysis including hundreds of thousands of admixed individuals and is universally applicable to various types of complex traits including binary trait, quantitative trait, time-to-event trait, longitudinal traits, etc. For each genetic variant, SPAmix uses genotype data and genetic principal components (PCs) to estimate individual-level allele frequency, which is subsequently used to calibrate p values via a retrospective analysis. A hybrid strategy including saddlepoint approximation (SPA) can greatly increase the accuracy to analyze rare genetic variants, especially if the phenotypic distribution is unbalanced or extremely unbalanced. Compared to Tractor, SPAmix does not require local ancestry information and can be straightforwardly applicable to a multi-way admixed population. Meanwhile, SPAmix can also be extended to SPAmix<sub>local</sub> in which the local ancestry can be incorporated if available. In addition, we propose SPAmix<sub>CCT</sub> to combine the p values of SPAmix and SPAmix<sub>local</sub> via Cauchy combination (CCT). SPAmix<sub>local</sub> performs close to Tractor when analyzing quantitative traits and is more accurate when analyzing binary traits with an unbalanced case-control ratio. And SPAmix<sub>CCT </sub>is an optimal unified approach for various cross-ancestry genetic architectures. Extensive simulation studies and real data analyses of 369,314 UK Biobank individuals from multiple ancestries demonstrated that SPAmix is scalable and can discover novel hits while controlling type I error rates well.</p>
Genetic Architecture Reconciles Linkage and Association Studies of Complex Traits
<p>This (zipped) folder contains 3 sub-folders:</p> <p>#**********************************************************************************************************<br>The "bin" folder contains fuctions and gentic maps needed for analyes<br>bin \<br> predLink.R - function to predict linkage <br> sibREML_v0.1.1.R - function to run SibREML<br> sim-sib-array.R - script to simulate sib-pairs from parental haplotypes<br> Summarised_genetic_map_bcf.txt - genetic map per 0.5-cM long segments, based on map from bcftools <br> (BCFtools: https://samtools.github.io/bcftools/bcftools.html)<br> Summarised_genetic_map_OMNI.txt - genetic map per 0.5-cM long segments, based on OMNI map <br> (https://github.com/joepickrell/1000-genomes-genetic-maps/tree/master/interpolated_OMNI)<br>#**********************************************************************************************************</p> <p> </p> <p>#**********************************************************************************************************<br>The "SIM" folder contains the simulation pipeline (scripts 01-15) as well as IBD sharing and simulated phenotypes for Simulated sib-pairs.<br>SIM \<br> 01_sim-sib-array.sh *pre-run*<br> 02_bed_recode_bcf_map.sh *pre-run*<br> 03_make_merlin.R *pre-run*<br> 04_error_merlin.sh *pre-run*<br> 05_merlin_IBD.sh *pre-run*<br> 06_sample_causal_snps.R *pre-run*<br> 07_simulate_pheno.sh *pre-run*<br> 08_bhat_gwas.R *can be run using provided data* <br> 09_Linkage_VH.R *can be run using provided data* <br> 10_predLink.R *can be run using provided data*<br> 11_phi_hat.R *can be run using provided data*<br> 12_IBD_Mb.R *can be run using provided data*<br> 13_IBD_cM_recombrate_stratified.R *can be run using provided data*<br> 14_SibREML.R *can be run using provided data*<br> 15_SibREML_stratified_Q4.R *can be run using provided data*<br> causal_snps \ *provided causal SNPs*<br> IBD_results \ *provided IBD-probabilities for 1000 simulated sib-pairs*<br> Linkage_VH_results \ <br> pheno \ *provided simulated phenotypes (h2=1) for 8 genetic architectures*<br> Phi_hat_results.txt<br> predicted \<br> README<br> SibREML_results.txt<br> SibREML_stratified_Q4.txt</p> <p>The data can be used to run Linkage analysis, predict linkage, estimate phi_hat, <br>as well as estimate non-stratified and recombination rate stratified sib-heritability (h2_FS and c).<br>The README is provided within the folder. <br>#**********************************************************************************************************</p> <p> </p> <p>#**********************************************************************************************************<br>The "HT_BMI" folder contains data and scripts to predict linkage and estimate phi_hat for height and BMI.<br>HT_BMI \<br> 01_predLink_HT_BMI.R<br> 02_phi_hat_HT_BMI.R<br> gws_sumstats \ *provided summary GWAS summary statistics to predict linkage for height and BMI*<br> Linkage_results \ *provided linkage meta-analysis results for height and BMI from this study*<br> Phi_hat_results_HT_BMI.txt<br> PREDLINK_bmi.txt<br> PREDLINK_height.txt<br> README<br>The README is provided within the folder.<br>#**********************************************************************************************************</p> <p><strong> </strong></p>
Summary statistics from "Genetic Association Study of Eight Steroid Hormones and Implications for Sexual Dimorphism of Coronary Artery Disease"
<p>GWAMA summary statistics of four steroid hormone levels using fixed-effect model and GWAS summary statistics of four other steroid hormones.</p> <p>When using this data, please cite: Pott J, Bae YJ, Horn K, et al.. Genetic Association Study of Eight Steroid Hormones and Implications for Sexual Dimorphism of Coronary Artery Disease. <em>J Clin Endocrinol Metab</em> <strong>2019</strong> Nov 1;104(11):5008-5023. doi: 10.1210/jc.2019-00757</p> <p>All txt files contain the following columns:</p> <ul> <li>markername</li> <li>chr</li> <li>bp_hg19 (base position according to hg19)</li> <li>effect_allele</li> <li>other_allele</li> <li>effect_allele_freq</li> <li>min_info (minimal info score across all used studies)</li> <li>n (sample size per SNP)</li> <li>beta (effect estimate)</li> <li>se (standard error)</li> <li>p (p-value)</li> <li>CochransQ (only in GWAMA; SNP heterogeneity across studies)</li> <li>pCochransQ (only in GWAMA; p-value of Cochrans Q value)</li> </ul>
Genetic variants beyond amyloid and tau associated cognitive decline: a cohort study
<p>Objective: To identify single nucleotide polymorphisms (SNPs) associated with cognitive decline independent of amyloid &[beta] (A&[beta]) and tau pathology in Alzheimer's disease (AD). Methods: Discovery and replication datasets consisting of 414 subjects (94 cognitively normal control [CN), 185 with mild cognitive impairment [MCI], and 135 AD) and 72 subjects (22 CN, 39 MCI, and 11 AD), respectively, were obtained from the Alzheimer's Disease Neuroimaging Initiative database. Genome-wide association analysis was conducted to identify SNPs associated with individual cognitive function (measured using the MMSE and ADAS-cog) while controlling for the level of A&[beta] and tau (measured as CSF p-tau/A&[beta]1-42). Gene ontology analysis was performed on SNP associated genes.</p> <p>Results: We identified one significant (rs55906536, &[beta]=-1.91,standard error 0.34, P =4.07×10<sup>-8</sup>) and four suggestive variants on chromosome 6, which were associated with poorer cognitive function. Congruent results were found in the replication data. A structural equation model showed that the identified SNP deteriorated cognitive function partially through cortical thinning of the brain in a region-specific manner. Furthermore, a bioinformatics analysis showed that the identified SNPs were associated with genes related to glutathione metabolism.</p> <p>Conclusions: In this study, we identified SNPs related to cognitive decline, in a manner which could not be explained by A&[beta] and tau levels. Our findings provide insight into the complexity of AD pathogenesis and support the growing literature on the role of glutathione in AD. This study suggests anti-oxidative agents may serve therapeutic for AD subjects with the identified SNPs.</p>
A 4,302-patient cohort study of association of rare genetic alterations with endocrine disorders
<p><span>Endocrine pathologies including disorders such as diabetes and dysfunctions of endocrine glands, are frequently associated with genetic predisposition. This study investigated the association of endocrine diseases with genetic variants, copy number variations (CNVs), and mutational load of molecular pathways in 4302 patients with 409 ICD-10 diagnoses who underwent DNA testing using next-generation sequencing at the National Medical Research Center for Endocrinology (Moscow) from November 2017 </span><span>till</span><span> January 2024. We analyzed rare protein-altering genetic variants using three control cohorts (gnomAD3, RUSeq healthy, experimental).<span> </span>We identified 143 associated variants for diabetes mellitus and 188 genetic variants across other 18 different ICD-10 groups of diagnoses, including 25% and 30% of previously undescribed variants, respectively. In addition, we investigated the aggregation of genetic variants across individual genes and their functional ensembles (molecular pathways) and identified 105 and 101 associations with ICD-10 diagnoses, respectively. In addition, we identified 35 pathogenic and 91 likely pathogenic CNVs in 925 patients with whole exome sequencing profiles. Among them, 9 and 44 CNVs, respectively, were not previously described. Totally, we found statistically significant associations between CNVs and endocrine pathologies for 168 genes. These results expand our understanding of endocrine disease mechanisms and may indicate new potential therapeutic targets.</span></p>
Data from: Genome-wide association studies across environmental and genetic contexts reveal complex genetic architecture of symbiotic extended phenotypes
<p>A goal of modern biology is to develop the genotype-phenotype (G→P) map, a predictive understanding of how genomic information generates trait variation that forms the basis of both natural and managed communities. As microbiome research advances, however, it has become clear that many of these traits are symbiotic extended phenotypes, being governed by genetic variation encoded not only by the host's own genome, but also by the genomes of myriad cryptic symbionts. Building a reliable G→P map therefore requires accounting for the multitude of interacting genes and even genomes involved in symbiosis. Here we use naturally-occurring genetic variation in 191 strains of the model microbial symbiont <em>Sinorhizobium meliloti</em> paired with two genotypes of the host <em>Medicago truncatula</em> in four genome-wide association studies (GWAS) to determine the genomic architecture of a key symbiotic extended phenotype – partner quality, or the fitness benefit conferred to a host by a particular symbiont genotype, within and across environmental contexts and host genotypes. We define three novel categories of loci in rhizobium genomes that must be accounted for if we want to build a reliable G→P map of partner quality; namely, 1) loci whose identities depend on the environment, 2) those that depend on the host genotype with which rhizobia interact, and 3) universal loci that are likely important in all or most environments.</p> <p><span>IMPORTANCE:</span><strong> </strong>Given the rapid rise of research on how microbiomes can be harnessed to improve host health, understanding the contribution of microbial genetic variation to host phenotypic variation is pressing, and will better enable us to predict the evolution of (and select more precisely for) symbiotic extended phenotypes that impact host health. We uncover extensive context-dependency in both the identity and functions of symbiont loci that control host growth, which makes predicting the genes and pathways important for determining symbiotic outcomes under different conditions more challenging. Despite this context-dependency, we also resolve a core set of universal loci that are likely important in all or most environments, and thus, serve as excellent targets both for genetic engineering and future coevolutionary studies of symbiosis.</p>
Data for: Dissecting the genetic architecture of leaf morphology traits in mungbean (Vigna radiata (L.) Wizcek) using genome‐wide association study
<p><span>Mungbean (<em>Vigna radiata</em> (L) Wizcek) is an important pulse crop, increasingly used as a source of protein, fiber, low fat, carbohydrates, minerals, and bioactive compounds in human diets. Mungbean is a dicot plant with trifoliate leaves. Leaves are central to various plant processes like photosynthesis, light interception, and overall canopy structure. The objectives were to study leaf morphological traits, use image analysis to extract leaf traits from images from the Iowa Mungbean Diversity (IMD) panel, develop a regression model for the prediction of leaflet area, and conduct association mapping for leaf morphological traits. We collected more than 5000 leaf images of the IMD panel consisting of 484 accessions over two years (2020 and 2021) with two replications per experiment. Leaf traits were extracted using image analysis, analyzed, and used for association mapping. Morphological diversity included leaflet type (oval or lobed), leaflet size (small, medium, large), lobed angle (shallow, deep), and vein coloration (green, purple). A regression model was developed to predict each ovate leaflet's area (adjusted R<sup>2</sup> = 0.97; residual standard errors of <= 1.10). The candidate genes <em>Vradi01g07560</em>, <em>Vradi05g01240</em>, <em>Vradi02g05730</em>, and <em>Vradi03g00440</em>, are associated with multiple traits (length, width, perimeter, and area) across the leaflets (left, terminal, and right). These are suitable candidate genes for further investigation in their role in leaf development, growth, and function. Future studies will be needed to correlate the observed traits discussed here with yield or important agronomic traits for use as phenotypic or genotypic markers in marker-aided selection methods for mungbean crop improvement.</span></p>
Association of Host Genetics With Vaccine Efficacy and Study of Immune Correlates of Risk From a Tetravalent Dengue Vaccine
ClinicalTrials.gov study NCT02827162. IPD Sharing: YES. Countries: 1. Publications: 2.
Data from: Genome-wide association studies across environmental and genetic contexts reveal complex genetic architecture of symbiotic extended phenotypes
Open the record for dataset details and reuse information.
Genetic variants beyond amyloid and tau associated cognitive decline: a cohort study
Open the record for dataset details and reuse information.
Data for: Dissecting the genetic architecture of leaf morphology traits in mungbean (Vigna radiata (L.) Wizcek) using genome‐wide association study
Open the record for dataset details and reuse information.
Data from: New insights into the dynamics between reef corals and their associated dinoflagellate endosymbionts from population genetic studies.
The mutualistic symbioses between reef-building corals and micro-algae form the basis of coral reef ecosystems, yet recent environmental changes threaten their survival. Diversity in host-symbiont pairings on the sub-species level could be an unrecognized source of functional variation in response to stress. The Caribbean elkhorn coral, Acropora palmata, associates predominantly with one symbiont species (Symbiodinium 'fitti'), facilitating investigations of individual-level (genotype) interactions. Individual genotypes of both host and symbiont were resolved across the entire range of the species. Most colonies of a particular animal genotype were dominated by one symbiont genotype (or strain) that may persist in the host for decades or more. While Symbiodinium are primarily clonal, the occurrence of recombinant genotypes indicates sexual recombination is the source of this genetic variation, and some evidence suggests this happens within the host. When these data are examined at spatial scales spanning the entire distribution of A. palmata, gene flow among animal populations was an order of magnitude greater than among populations of the symbiont. This suggests that independent micro-evolutionary processes created dissimilar population genetic structures between host and symbiont. The lower effective dispersal exhibited by the dinoflagellate raises questions regarding the extent to which populations of host and symbiont can co-evolve during times of rapid and substantial climate change. However, these findings also support a growing body of evidence suggesting that genotype by genotype interactions may provide significant physiological variation; influencing the adaptive potential of symbiotic reef corals to severe selection.
Data from: Genetic dissection of grain iron and zinc, and thousand kernel weight in wheat (Triticum aestivum L.) using genome-wide association study
<p>The study material in GWAS panel with 280 common bread wheat genotypes was selected from All India Coordinated Research Project on Wheat and Barley to map the genomic regions responsible for enhanced Grain Zinc Content (GZnC), Grain Iron Content (GZnC) and Thousand Kernel weight (TKW).</p> <p><strong>Phenotypic data:</strong></p> <p>The GWAS panel was evaluated at five different environments: E1-University of Agricultural Sciences, research farm, Dharwad (15°29'20.71"N, 74°59'3.35"E, 750m AMSL), E2-ICAR- Indian Agricultural Research Institute, New Delhi (28°38′30.5″N, 77°09′58.2″E, 228 m AMSL), E3-Indian Agricultural Research Institute, Jharkhand (24°16'58.4"N, 85°21'16.1"E, 651m AMSL), E4-ICAR-Indian Institute of Wheat and Barley, Karnal (29°41'8.2644''N, 76°59'25.9692''E, 250m AMSL), and E5-Punjab Agricultural University, Ludhiana (30o54' N, 75o48'E, 247m AMSL). Around 20 g of grain sample from each genotype were used for phenotyping GFeC and GZnC through high-throughput Energy Dispersive X-ray Fluorescence (ED-XRF) machine (model X-Supreme 8000; Oxford Instruments plc, Abingdon, United Kingdom) calibrated with glass beads-based values. To record TKW, the Numigral grain counter was used to count the grain number, the reading was set at 1000 grains and the weight of the grains was recorded in grams with an electronic balance. The GFeC, GZnC were expressed as milligram per kilogram (mg/kg), GPC in percentage (%), TKW in grams (gms).</p> <p><strong>Genotypic data:</strong></p> <p>Genomic DNA of the GWAS panel was extracted from the leaves of 21 days-old seedlings by Cetyl Trimethyl Ammonium Bromide (CTAB) method. The panel was genotyped using Axiom Wheat Breeder's Genotyping Array (Affymetrix, Santa Clara, CA, United States) having 35,143 genome-wide SNPs. The monomorphic, markers with minor allele frequency (MAF) of <5%, missing data of >20%, and heterozygote frequency >25% were removed from the analysis. The remaining set of 14,790 high-quality SNPs was used in GWAS analysis. The detailed information of the methods and software used, data analysis and GWAS is available at DOI: 10.1038/s41598-022-15992-z.</p>
Association between a glucagon-like peptide 1 receptor genetic polymorphism and therapeutic response to sitagliptin in a sample of type 2 diabetic patients: an observational study
<p>Association between a glucagon-like peptide 1 receptor genetic polymorphism and therapeutic response to sitagliptin in a sample of type 2 diabetic patients: an observational study</p>
Effectiveness Study of Vilazodone to Treat Depression and to Discover Genetic Markers Associated With Response
ClinicalTrials.gov study NCT00285376. IPD Sharing: Not stated. Countries: 1. Publications: 6.
Genetic Association Study Between GAD1 and Reelin Polymorphisms and GABA/Glutamate MRS in Bipolar Disorder Type 1 and Healthy Controls: SPECGENE PROJECT
ClinicalTrials.gov study NCT01237158. IPD Sharing: Not stated. Countries: 1. Publications: 5.
Clinical and Genetic Studies of VACTERL Association
ClinicalTrials.gov study NCT00766571. IPD Sharing: Not stated. Countries: 1. Publications: 3.
Identification of Genetic Polymorphisms Related to Propofol Requirement and Recovery Through Genome-wide Association Study (GWAS) in Total Intravenous Anesthesia for Clipping of Unruptured Cerebral An
ClinicalTrials.gov study NCT03087383. IPD Sharing: NO. Countries: 1. Publications: 1.
Genetic Studies of Strabismus, Congenital Cranial Dysinnervation Disorders (CCDDs), and Their Associated Anomalies
ClinicalTrials.gov study NCT03059420. IPD Sharing: NO. Countries: 1. Publications: 18.
Creation of a Prospective Cohort of Healthy and Sick Subjects and of a Collection of Associated Biological Resources, for the Study of the Immune System and of Its Genetic and Environmental Determinan
ClinicalTrials.gov study NCT03925272. IPD Sharing: NO. Countries: 1. Publications: 16.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.