Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
76
datasets available to search
ShareScore release 0.9.0
Dataset results
76 results for “Association studies in genetics”
Study of Factors of Genetic Susceptibility Associated to Severe Caries Phenotype
ClinicalTrials.gov study NCT00541060. IPD Sharing: Not stated. Countries: 1. Publications: 0.
Study on the Association Between SXCI and RM and the Possible Genetic Mechanism
ClinicalTrials.gov study NCT02504281. IPD Sharing: Not stated. Countries: 1. Publications: 1.
Study to Identify the Genetic Variations Associated With Phantom Limb Pain
ClinicalTrials.gov study NCT01462448. IPD Sharing: NO. Countries: 1. Publications: 6.
Genetic Association Study Between Single Nucleotide Polymorphisms (SNPs) and Cognitive Performance in Young Bipolar Type I Patients: LICAVALGENE
ClinicalTrials.gov study NCT00969930. IPD Sharing: Not stated. Countries: 1. Publications: 1.
Data from: New insights into the dynamics between reef corals and their associated dinoflagellate endosymbionts from population genetic studies.
Open the record for dataset details and reuse information.
Genetic variation in GC and CYP2R1 affects 25-hydroxyvitamin D concentration and skeletal parameters: a genome-wide association study in 24-month-old Finnish children
Open the record for dataset details and reuse information.
Data from: Genome-wide association studies dissect the genetic architecture of seed and yield component traits in cowpea (Vigna unguiculata L. Walp)
Open the record for dataset details and reuse information.
Data from: Genetic dissection of grain iron and zinc, and thousand kernel weight in wheat (Triticum aestivum L.) using genome-wide association study
Open the record for dataset details and reuse information.
Combining genome-wide studies of breast, prostate, ovarian and endometrial cancers maps cross-cancer susceptibility loci and identifies new genetic associations
<p>Data set linked to the paper, "Combining genome-wide studies of breast, prostate, ovarian and endometrial cancers maps cross-cancer susceptibility loci and identifies new genetic associations". Pre-print of the paper is here: <a href="https://doi.org/10.1101/2020.06.16.146803">https://doi.org/10.1101/2020.06.16.146803</a>.</p> <p> </p> <p>cross_cancer_sum_stats.txt.gz contains summary genome-wide association statistics for susceptibility to single cancers (breast (BR), prostate (PR), ovarian (OV), endometrial (EN), estrogen receptor (ER)-positive breast (POS), ER-negative breast (NEG), and high-grade serous ovarian (HGS) cancers) and from the cross-cancer meta-analysis (main [main] and subtype-focused [sub]). EA in the header refers to the effect allele, OA is the other allele, EAF is the effect allele frequency in the largest of the single cancer data sets (BR), IMPR2 is the imputation quality in the largest of the single cancer data sets (BR), SE is the standard error, PVAL is the P-value, RE2Cs1 is the RE2C statistic mean effect part, RE2Cs2 is the RE2C statistic heterogeneity part, RE2Cp* is the RE2C* P-value. More on RE2Cp* can be found here: <a href="http://software.buhmhan.com/RE2C/index.php?mid=contact&act=dispBoardWrite">http://software.buhmhan.com/RE2C/index.php?mid=contact&act=dispBoardWrite</a> and in <a href="https://academic.oup.com/bioinformatics/article/33/14/i379/3953957">https://academic.oup.com/bioinformatics/article/33/14/i379/3953957</a> SNP names in cross_cancer_sum_stats.txt.gz include the chromosome and build 37 position.</p> <p> </p> <p>main_tetrachoric_corr_matrix.txt and subtype_tetrachoric_corr_matrix.txt provide the tetrachoric correlation matrices used in the main and subtype-focused meta-analyses. These were also used to specify the cryptic.cor argument of the exh.abf function of MetABF. More on MetABF can be found here: <a href="https://github.com/trochet/metabf">https://github.com/trochet/metabf</a> and in <a href="https://onlinelibrary.wiley.com/doi/abs/10.1002/gepi.22202">https://onlinelibrary.wiley.com/doi/abs/10.1002/gepi.22202</a></p> <p> </p> <p>prior_sigmas_for_metabf.txt contains the values used to specify the prior.sigma argument of the exh.abf function in MetABF.</p> <p> </p> <p>The breast cancer data used are described in <a href="https://pubmed.ncbi.nlm.nih.gov/29059683/"><strong>PMID 29059683</strong></a> and can be downloaded from <a href="http://bcac.ccge.medschl.cam.ac.uk/bcacdata/oncoarray/oncoarray-and-combined-summary-result/gwas- summary-results-breast-cancer-risk-2017/">http://bcac.ccge.medschl.cam.ac.uk/bcacdata/oncoarray/oncoarray-and-combined-summary-result/gwas- summary-results-breast-cancer-risk-2017/</a> (this link also includes acknowledgements). The prostate cancer data are described in <a href="https://pubmed.ncbi.nlm.nih.gov/29892016/"><strong>PMID 29892016</strong></a> and can be downloaded from: <a href="http://practical.icr.ac.uk/blog/?page_id=8164">http://practical.icr.ac.uk/blog/?page_id=8164</a> (this link also includes acknowledgements). The ovarian cancer data used are described in <a href="https://pubmed.ncbi.nlm.nih.gov/28346442/"><strong>PMID 28346442</strong></a> and can be downloaded from <a href="https://www.ebi.ac.uk/gwas/studies/GCST004415">https://www.ebi.ac.uk/gwas/studies/GCST004415</a>. The endometrial cancer data are described in <a href="https://pubmed.ncbi.nlm.nih.gov/30093612/"><strong>PMID 30093612</strong></a> and can be downloaded from <a href="https://www.ebi.ac.uk/gwas/studies/GCST006464">https://www.ebi.ac.uk/gwas/studies/GCST006464</a>. These links point to the same data that form the basis of the cross_cancer_sum_stats.txt.gz file.</p> <p> </p> <p><strong>The sample size and precision of the data presented should preclude identification of any individual study participant. However, in downloading these data, you undertake not to attempt to identify individual study participant and not to re-post these data to a third-party website. Please cite the PMIDs highlighted above along with the appropriate acknowledements if you use the cross_cancer_sum_stats.txt.gz file.</strong></p> <p> </p> <p>If you have any questions about this repository, please email Siddhartha Kar at siddhartha dot kar at bristol dot ac dot uk</p>
Association of common genetic variants with brain microbleeds: A genome-wide association study
<p><strong>Objective:</strong> To identify common genetic variants associated with the presence of brain microbleeds (BMB).</p> <p><strong>Methods:</strong> We performed genome-wide association studies in 11 population-based cohort studies and 3 case-control or case-only stroke cohorts. Genotypes were imputed to the Haplotype Reference Consortium or 1000 Genomes reference panel. BMB were rated on susceptibility-weighted or T2*-weighted gradient echo magnetic resonance imaging sequences, and further classified as lobar, or mixed (including strictly deep and infratentorial, possibly with lobar BMB). In a subset, we assessed the effects of <em>APOE</em> ε2 and ε4 alleles on BMB counts. We also related previously identified cerebral small vessel disease variants to BMB.</p> <p><strong>Results: </strong>BMB were detected in 3,556 of the 25,862 participants, of which 2,179 were strictly lobar and 1,293 mixed. One locus in the <em>APOE</em> region reached genome-wide significance for its association with BMB (lead SNP rs769449; OR<sub>any BMB</sub> (95% CI)=1.33 (1.21-1.45); p=2.5x10-10). <em>APOE</em> ε4 alleles were associated with strictly lobar (OR (95% CI)=1.34 (1.19- 1.50); p=1.0x10-6) but not with mixed BMB counts (OR (95% CI)=1.04 (0.86-1.25); p=0.68). <em>APOE</em> ε2 alleles did not show associations with BMB counts. Variants previously related to deep intracerebral hemorrhage and lacunar stroke, and a risk score of cerebral white matter hyperintensity variants, were associated with BMB.</p> <p><strong>Conclusions: </strong>Genetic variants in the <em>APOE</em> region are associated with the presence of BMB, most likely due to the <em>APOE</em> ε4 allele count related to a higher number of strictly lobar BMB. Genetic predisposition to small vessel disease confers risk of BMB, indicating genetic overlap with other cerebral small vessel disease markers.</p>
Data from: Combining high-throughput micro-CT-RGB phenotyping and genome-wide association study to dissect the genetic architecture of tiller growth in rice
Manual phenotyping of rice tillers is time consuming and labor intensive and lags behind the rapid development of rice functional genomics. Thus, automated, non-destructive phenotyping of rice tiller traits at a high spatial resolution and high-throughput for large-scale assessment of rice accessions is urgently needed. In this study, we developed a high-throughput micro-CT-RGB (HCR) imaging system to non-destructively extract 730 traits from 234 rice accessions at 9 time points. We could explain 30% of the grain yield variance from 2 tiller traits assessed in the early growth stages. A total of 402 significantly associated loci were identified by GWAS, and dynamic and static genetic components were found across the nine time points. A major locus associated with tiller angle was detected at nine time points, which contained a major gene TAC1. Significant variants associated with tiller angle were enriched in the 3'-UTR of TAC1. Three haplotypes for the gene were found and rice accessions containing haplotype H3 displayed much smaller tiller angles. Further, we found two loci contained associations with both vigor-related HCR traits and yield. The superior alleles would be beneficial for breeding of high yield and dense planting.
Data from: Genetic distance as an alternative to physical distance for definition of gene units in association studies
Background: Some association studies, as the implemented in VEGAS, ALIGATOR, i-GSEA4GWAS, GSA-SNP and other software tools, use genes as the unit of analysis. These genes include the coding sequence plus flanking sequences. Polymorphisms in the flanking sequences are of interest because they involve cis-regulatory elements or they inform on untyped genetic variants trough linkage disequilibrium. Gene extensions have customarily been defined as ± 50 Kb. This approach is not fully satisfactory because genetic relationships between neighbouring sequences are a function of genetic distances, which are only poorly replaced by physical distances. Results: Standardized recombination rates (SRR) from the deCODE recombination map were used as units of genetic distances. We searched for a SRR producing flanking sequences near the ± 50 Kb offset that has been common in previous studies. A SRR ≥ 2 was selected because it led to gene extensions with median length = 45.3 Kb and the simplicity of an integer value. As expected, boundaries of the genes defined with the ± 50 Kb and with the SRR ≥2 rules were rarely concordant. The impact of these differences was illustrated with the interpretation of top association signals from two large studies including many hits and their detailed analysis based in different criteria. The definition based in genetic distance was more concordant with the results of these studies than the based in physical distance. In the analysis of 18 top disease associated loci form the first study, the SRR ≥2 genes led to a fully concordant interpretation in 17 loci; the ± 50 Kb genes only in 6. Interpretation of the 43 putative functional genes of the second study based in the SRR ≥2 definition only missed 4 of the genes, whereas the based in the ± 50 Kb definition missed 10 genes. Conclusions: A gene definition based on genetic distance led to results more concordant with expert detailed analyses than the commonly used based in physical distance. The genome coordinates for each gene are provided to maintain a simple use of the new definitions.
Data from: A practical introduction to random forest for genetic association studies in ecology and evolution
Large genomic studies are becoming increasingly common with advances in sequencing technology, and our ability to understand how genomic variation influences phenotypic variation between individuals has never been greater. The exploration of such relationships first requires the identification of associations between molecular markers and phenotypes. Here we explore the use of Random Forest (RF), a powerful machine learning algorithm, in genomic studies to discern loci underlying both discrete and quantitative traits, particularly when studying wild or non-model organisms. RF is becoming increasingly used in ecological and population genetics because, unlike traditional methods, it can efficiently analyze thousands of loci simultaneously and account for non-additive interactions. However, understanding both the power and limitations of Random Forest is important for its proper implementation and the interpretation of results. We therefore provide a practical introduction to the algorithm and its use for identifying associations between molecular markers and phenotypes, discussing such topics as data limitations, algorithm initiation and optimization, as well as interpretation. We also provide short R tutorials as examples, with the aim of providing a guide to the implementation of the algorithm. Topics discussed here are intended to serve as an entry point for molecular ecologists interested in employing Random Forest to identify trait associations in genomic data sets.
Figure 3 from: Al-Musawi ZJMA, Al-Juhaishi AMR (2024) Association study between D2 receptor A-241G, rs1799978 genetic variation and olanzapine efficacy in Iraqi schizophrenic patients. Pharmacia 71: 1-6. https://doi.org/10.3897/pharmacia.71.e111984
Figure 3 The Prevalence of D2 receptor alleles A-241G (rs1799978) among both male and female volunteers.
First Year Growth Response Associated Genetic Markers Validation Phase IV Open-label Study in Growth Hormone Deficient and Turner Syndrome Pre-pubertal Children: the PREDICT Pharmacogenetics Validatio
ClinicalTrials.gov study NCT01419249. IPD Sharing: Not stated. Countries: 9. Publications: 0.
Diet Quality and Genetic Association With Body Mass Index: Results From Three Observational Studies
ClinicalTrials.gov study NCT03577639. IPD Sharing: UNDECIDED. Countries: 0. Publications: 1.
Data from: Combining high-throughput micro-CT-RGB phenotyping and genome-wide association study to dissect the genetic architecture of tiller growth in rice
Open the record for dataset details and reuse information.
Data from: A practical introduction to random forest for genetic association studies in ecology and evolution
Open the record for dataset details and reuse information.
Association of common genetic variants with brain microbleeds: A genome-wide association study
Open the record for dataset details and reuse information.
Data from: Genetic distance as an alternative to physical distance for definition of gene units in association studies
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.