Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

172

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

172 results for “Genome-wide association studies”

Learn how ShareScore rates datasets ↗
dryad32/100

Data from: Genome-wide association study of meat quality traits in a White Duroc × Erhualian F2 intercross and Chinese Sutai pigs

Open the record for dataset details and reuse information.

publicJun 2013View details →
dryad32/100

Data from: Genomic predictions and genome-wide association study of resistance against Piscirickettsia salmonis in coho salmon (Oncorhynchus kisutch) using ddRAD sequencing

Open the record for dataset details and reuse information.

publicFeb 2019View details →
dryad32/100

Data from: Genome-wide association study for grain yield and component traits in wheat (Triticum aestivum L.)

Open the record for dataset details and reuse information.

publicJul 2022View details →
dryad32/100

Data from: Enhancing genomic prediction with genome-wide association studies in multiparental maize populations

Open the record for dataset details and reuse information.

publicJan 2017View details →
dryad32/100

Data from: A genome-wide scan study identifies a single nucleotide substitution in ASIP associated with white versus non-white coat-colour variation in sheep (Ovis aries)

Open the record for dataset details and reuse information.

publicAug 2013View details →
dryad32/100

UK dogs data from: Genome-wide association studies for canine hip dysplasia in single and multiple populations – implications and potential novel risk loci

Open the record for dataset details and reuse information.

publicAug 2021View details →
dryad32/100

Data from: Imputation of canine genotype array data using 365 whole-genome sequences improves power of genome-wide association studies

Open the record for dataset details and reuse information.

publicOct 2019View details →
dryad32/100

Genetic variation in GC and CYP2R1 affects 25-hydroxyvitamin D concentration and skeletal parameters: a genome-wide association study in 24-month-old Finnish children

Open the record for dataset details and reuse information.

publicDec 2019View details →
dryad32/100

Summary statistics from a genome-wide association study of narcolepsy

Open the record for dataset details and reuse information.

publicJan 2023View details →
dryad32/100

Data from: Multiple-trait genome-wide association study based on principal component analysis for residual covariance matrix

Open the record for dataset details and reuse information.

publicMay 2014View details →
dryad32/100

Multi-population genome-wide association studies involving four distinct barley populations

Open the record for dataset details and reuse information.

publicFeb 2025View details →
dryad32/100

Data from: Genome-wide association studies dissect the genetic architecture of seed and yield component traits in cowpea (Vigna unguiculata L. Walp)

Open the record for dataset details and reuse information.

publicFeb 2025View details →
dryad32/100

Data from: Genome-wide association study of an unusual dolphin mortality event reveals candidate genes for susceptibility and resistance to cetacean morbillivirus

Open the record for dataset details and reuse information.

publicNov 2018View details →
dryad32/100

Data from: Genetic dissection of grain iron and zinc, and thousand kernel weight in wheat (Triticum aestivum L.) using genome-wide association study

Open the record for dataset details and reuse information.

publicJul 2022View details →
dryad32/100

Genome-wide association study in quinoa reveals selection pattern typical for crops with a short breeding history

Open the record for dataset details and reuse information.

publicJul 2022View details →
zenodo28/100

Combining genome-wide studies of breast, prostate, ovarian and endometrial cancers maps cross-cancer susceptibility loci and identifies new genetic associations

<p>Data set linked to the paper, &quot;Combining genome-wide studies of breast, prostate, ovarian and endometrial cancers maps cross-cancer susceptibility loci and identifies new genetic associations&quot;.&nbsp; Pre-print of the paper is here: <a href="https://doi.org/10.1101/2020.06.16.146803">https://doi.org/10.1101/2020.06.16.146803</a>.</p> <p>&nbsp;</p> <p>cross_cancer_sum_stats.txt.gz contains summary genome-wide association statistics for susceptibility to single cancers (breast (BR), prostate (PR), ovarian (OV), endometrial (EN), estrogen receptor (ER)-positive breast (POS), ER-negative breast (NEG), and high-grade serous ovarian (HGS) cancers) and from the cross-cancer meta-analysis (main [main] and subtype-focused [sub]). EA in the header refers to the effect allele, OA is the other allele, EAF is the effect allele frequency in the largest of the single cancer data sets (BR), IMPR2 is the imputation quality in the largest of the single cancer data sets (BR), SE is the standard error, PVAL is the P-value, RE2Cs1 is the&nbsp; RE2C statistic mean effect part, RE2Cs2 is the RE2C statistic heterogeneity part, RE2Cp* is the RE2C* P-value.&nbsp; More on RE2Cp* can be found here: <a href="http://software.buhmhan.com/RE2C/index.php?mid=contact&amp;act=dispBoardWrite">http://software.buhmhan.com/RE2C/index.php?mid=contact&amp;act=dispBoardWrite</a> and in&nbsp;&nbsp;&nbsp;&nbsp; <a href="https://academic.oup.com/bioinformatics/article/33/14/i379/3953957">https://academic.oup.com/bioinformatics/article/33/14/i379/3953957</a> SNP names in&nbsp;cross_cancer_sum_stats.txt.gz include the chromosome and build 37 position.</p> <p>&nbsp;</p> <p>main_tetrachoric_corr_matrix.txt and subtype_tetrachoric_corr_matrix.txt provide the tetrachoric correlation matrices used in the main and subtype-focused meta-analyses.&nbsp; These were also used to specify the cryptic.cor argument of the exh.abf function of MetABF.&nbsp; More on MetABF can be found here: <a href="https://github.com/trochet/metabf">https://github.com/trochet/metabf</a> and in <a href="https://onlinelibrary.wiley.com/doi/abs/10.1002/gepi.22202">https://onlinelibrary.wiley.com/doi/abs/10.1002/gepi.22202</a></p> <p>&nbsp;</p> <p>prior_sigmas_for_metabf.txt contains the values used to specify the prior.sigma argument of the exh.abf function in MetABF.</p> <p>&nbsp;</p> <p>The breast cancer data used are described in <a href="https://pubmed.ncbi.nlm.nih.gov/29059683/"><strong>PMID 29059683</strong></a> and can be downloaded from <a href="http://bcac.ccge.medschl.cam.ac.uk/bcacdata/oncoarray/oncoarray-and-combined-summary-result/gwas- summary-results-breast-cancer-risk-2017/">http://bcac.ccge.medschl.cam.ac.uk/bcacdata/oncoarray/oncoarray-and-combined-summary-result/gwas- summary-results-breast-cancer-risk-2017/</a> (this link also includes acknowledgements).&nbsp; The prostate cancer data are described in <a href="https://pubmed.ncbi.nlm.nih.gov/29892016/"><strong>PMID 29892016</strong></a> and can be downloaded from: <a href="http://practical.icr.ac.uk/blog/?page_id=8164">http://practical.icr.ac.uk/blog/?page_id=8164</a> (this link also includes acknowledgements).&nbsp; The ovarian cancer data used are described in <a href="https://pubmed.ncbi.nlm.nih.gov/28346442/"><strong>PMID 28346442</strong></a> and can be downloaded from <a href="https://www.ebi.ac.uk/gwas/studies/GCST004415">https://www.ebi.ac.uk/gwas/studies/GCST004415</a>.&nbsp; The endometrial cancer data are described in <a href="https://pubmed.ncbi.nlm.nih.gov/30093612/"><strong>PMID 30093612</strong></a> and can be downloaded from <a href="https://www.ebi.ac.uk/gwas/studies/GCST006464">https://www.ebi.ac.uk/gwas/studies/GCST006464</a>.&nbsp; These links point to the same data that form the basis of the cross_cancer_sum_stats.txt.gz file.</p> <p>&nbsp;</p> <p><strong>The sample size and precision of the data presented should preclude identification of any individual study participant.&nbsp; However, in downloading these data, you undertake not to attempt to identify individual study participant and not to re-post these data to a third-party website.&nbsp; Please cite the PMIDs highlighted above along with the appropriate acknowledements if you use the cross_cancer_sum_stats.txt.gz file.</strong></p> <p>&nbsp;</p> <p>If you have any questions about this repository, please email Siddhartha Kar at siddhartha dot kar at bristol dot ac dot uk</p>

opencc-by-4.0Jun 2020View details →
dryad28/100

Association of common genetic variants with brain microbleeds: A genome-wide association study

<p><strong>Objective:</strong> To identify common genetic variants associated with the presence of brain microbleeds (BMB).</p> <p><strong>Methods:</strong> We performed genome-wide association studies in 11 population-based cohort studies and 3 case-control or case-only stroke cohorts. Genotypes were imputed to the Haplotype Reference Consortium or 1000 Genomes reference panel. BMB were rated on susceptibility-weighted or T2*-weighted gradient echo magnetic resonance imaging sequences, and further classified as lobar, or mixed (including strictly deep and infratentorial, possibly with lobar BMB). In a subset, we assessed the effects of <em>APOE</em> ε2 and ε4 alleles on BMB counts. We also related previously identified cerebral small vessel disease variants to BMB.</p> <p><strong>Results: </strong>BMB were detected in 3,556 of the 25,862 participants, of which 2,179 were strictly lobar and 1,293 mixed. One locus in the <em>APOE</em> region reached genome-wide significance for its association with BMB (lead SNP rs769449; OR<sub>any BMB</sub> (95% CI)=1.33 (1.21-1.45); p=2.5x10-10). <em>APOE</em> ε4 alleles were associated with strictly lobar (OR (95% CI)=1.34 (1.19- 1.50); p=1.0x10-6) but not with mixed BMB counts (OR (95% CI)=1.04 (0.86-1.25); p=0.68). <em>APOE</em> ε2 alleles did not show associations with BMB counts. Variants previously related to deep intracerebral hemorrhage and lacunar stroke, and a risk score of cerebral white matter hyperintensity variants, were associated with BMB.</p> <p><strong>Conclusions: </strong>Genetic variants in the <em>APOE</em> region are associated with the presence of BMB, most likely due to the <em>APOE</em> ε4 allele count related to a higher number of strictly lobar BMB. Genetic predisposition to small vessel disease confers risk of BMB, indicating genetic overlap with other cerebral small vessel disease markers.</p>

opencc-zeroAug 2021View details →
dryad28/100

Data from: Combining high-throughput micro-CT-RGB phenotyping and genome-wide association study to dissect the genetic architecture of tiller growth in rice

Manual phenotyping of rice tillers is time consuming and labor intensive and lags behind the rapid development of rice functional genomics. Thus, automated, non-destructive phenotyping of rice tiller traits at a high spatial resolution and high-throughput for large-scale assessment of rice accessions is urgently needed. In this study, we developed a high-throughput micro-CT-RGB (HCR) imaging system to non-destructively extract 730 traits from 234 rice accessions at 9 time points. We could explain 30% of the grain yield variance from 2 tiller traits assessed in the early growth stages. A total of 402 significantly associated loci were identified by GWAS, and dynamic and static genetic components were found across the nine time points. A major locus associated with tiller angle was detected at nine time points, which contained a major gene TAC1. Significant variants associated with tiller angle were enriched in the 3'-UTR of TAC1. Three haplotypes for the gene were found and rice accessions containing haplotype H3 displayed much smaller tiller angles. Further, we found two loci contained associations with both vigor-related HCR traits and yield. The superior alleles would be beneficial for breeding of high yield and dense planting.

opencc-zeroDec 2018View details →
dryad28/100

Data from: Genome-wide association study in Arabidopsis thaliana of natural variation in seed oil melting point, a widespread adaptive trait in plants

Seed oil melting point is an adaptive, quantitative trait determined by the relative proportions of the fatty acids that compose the oil. Micro- and macro-evolutionary evidence suggests selection has changed the melting point of seed oils to covary with germination temperatures because of a trade-off between total energy stores and the rate of energy acquisition during germination under competition. The seed oil compositions of 391 natural accessions of Arabidopsis thaliana, grown under common-garden conditions, were used to assess whether seed oil melting point within a species varied with germination temperature. In support of the adaptive explanation, long-term monthly spring and fall field temperatures of the accession collection sites significantly predicted their seed oil melting points. In addition, a genome-wide association study (GWAS) was performed to determine which genes were most likely responsible for the natural variation in seed oil melting point. The GWAS found a single highly significant association within the coding region of FAD2, which encodes a fatty acid desaturase central to the oil biosynthesis pathway. In a separate analysis of fifteen a priori oil synthesis candidate genes, two (FAD2 and FATB) were located near significant SNPs associated with seed oil melting point. These results comport with others' molecular work showing that lines with alterations in these genes affect seed oil melting point as expected. Our results suggest natural selection has acted on a small number of loci to alter a quantitative trait in response to local environmental conditions.

opencc-zeroDec 2015View details →
dryad28/100

Data from: Penalized Multi-Marker versus Single-Marker Regression methods for genome-wide association studies of quantitative traits

The data from genome-wide association studies (GWAS) in humans are still predominantly analyzed using single marker association methods. As an alternative to Single Marker Analysis (SMA), all or subsets of markers can be tested simultaneously. This approach requires a form of Penalized Regression (PR) as the number of SNPs is much larger than the sample size. Here we review PR methods in the context of GWAS, extend them to perform penalty parameter and SNP selection by False Discovery Rate (FDR) control, and assess their performance in comparison with SMA. PR methods were compared with SMA using realistically simulated GWAS data with a continuous phenotype and real data. Based on these comparisons our analytic FDR criterion may currently be the best approach to SNP selection using PR for GWAS. We found that PR with FDR control provides substantially more power than SMA with genome-wide type-I error control but somewhat less power than SMA with Benjamini-Hochberg FDR control (SMA-BH). PR with FDR based penalty parameter selection controlled the FDR somewhat conservatively while SMA-BH may not achieve FDR control in all situations. Differences among PR methods seem quite small when the focus is on SNP selection with FDR control. Incorporating linkage disequilibrium into the penalization by adapting penalties developed for covariates measured on graphs can improve power but also generate more false positives or wider regions for follow-up. We recommend the Elastic Net with a mixing weight for the Lasso penalty near 0.5 as the best method.

opencc-zeroDec 2013View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record