Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

186

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

186 results for “genome wide association study”

Learn how ShareScore rates datasets ↗
dryad36/100

Large scale across-breed genome-wide association study reveals a variant in HMGA2 associated with inguinal cryptorchidism risk in dogs

Open the record for dataset details and reuse information.

publicMay 2022View details →
dryad36/100

Supplemental material for: Genome-wide association study and fine-mapping using imputed sequences to prioritize candidate genes for 30 complex traits in 50,309 Holstein bulls

Open the record for dataset details and reuse information.

publicSep 2025View details →
dryad36/100

Supplementary information for: Redundancy analysis, genome-wide association studies, and the pigmentation of brown trout (Salmo trutta L.)

Open the record for dataset details and reuse information.

publicOct 2022View details →
dryad36/100

Data for: Dissecting the genetic architecture of leaf morphology traits in mungbean (Vigna radiata (L.) Wizcek) using genome‐wide association study

Open the record for dataset details and reuse information.

publicFeb 2023View details →
dryad32/100

Data from: Imputation of canine genotype array data using 365 whole-genome sequences improves power of genome-wide association studies

Genomic resources for the domestic dog have improved with the widespread adoption of a 173k SNP array platform and updated reference genome. SNP arrays of this density are sufficient for detecting genetic associations within breeds but are underpowered for finding associations across multiple breeds or in mixed-breed dogs, where linkage disequilibrium rapidly decays between markers, even though such studies would hold particular promise for mapping complex diseases and traits. Here we introduce an imputation reference panel, consisting of 365 diverse, whole-genome sequenced dogs and wolves, which increases the number of markers that can be queried in genome-wide association studies approximately 130-fold. Using previously genotyped dogs, we show the utility of this reference panel in identifying potentially novel associations, including a locus on CFA20 significantly associated with cranial cruciate ligament disease, and fine-mapping for canine body size and blood phenotypes, even when causal loci are not in strong linkage disequilibrium with any single array marker. This reference panel resource will improve future genome-wide association studies for canine complex diseases and other phenotypes.

opencc-zeroAug 2020View details →
dryad32/100

Data from: Genome-wide association study of outcrossing in cytoplasmic male sterile lines of rice

Stigma exsertion and panicle enclosure of male sterile lines are two key determinants of outcrossing in hybrid rice seed production. Based on 43,394 single nucleotide polymorphism markers, 217 cytoplasmic male sterile lines were assigned into two subpopulations and a mixed-group where the LD decay distance varied from 975 to 2,690 kb. Genome-wide association studies (GWAS) were performed for stigma exsertion rate (SE), panicle enclosure rate (PE) and seed-setting rate (SSR). A total of 154 significant association signals (P < 0.001) were identified. They were situated in 27 quantitative trait loci (QTLs), including 11 for SE, 6 for PE, and 10 for SSR. It was shown that six of the ten QTLs for SSR were tightly linked to QTLs for SE or/and PE with the expected allelic direction. These QTL clusters could be targeted to improve the outcrossing of female parents in hybrid rice breeding. Our study also indicates that GWAS-base QTL mapping can complement and enhance previous QTL information for understanding the genetic relationship between outcrossing and its related traits.

opencc-zeroOct 2019View details →
dryad32/100

Data from: Genome-wide association study of an unusual dolphin mortality event reveals candidate genes for susceptibility and resistance to cetacean morbillivirus

Infectious diseases are significant demographic and evolutionary drivers of populations, but studies about the genetic basis of disease resistance and susceptibility are scarce in wildlife populations. Cetacean morbillivirus (CeMV) is a highly contagious disease that is increasing in both geographic distribution and incidence, causing unusual mortality events (UME) and killing tens of thousands of individuals across multiple cetacean species worldwide since the late 1980's. The largest CeMV outbreak in the Southern Hemisphere reported to date occurred in Australia in 2013, where it was a major factor in a UME, killing mainly young Indo-Pacific bottlenose dolphins (Tursiops aduncus). Using cases (non-survivors) and controls (putative survivors) from the most affected population, we carried out a genome-wide association study to identify candidate genes for resistance and susceptibility to CeMV. The genomic dataset consisted of 278,147,988 sequence reads and 35,493 high quality SNPs genotyped across 38 individuals. Association analyses found highly significant differences in allele and genotype frequencies amongst cases and controls at 65 SNPs, and Random Forests conservatively identified eight as candidates. Annotation of these SNPs identified five candidate genes (MAPK8, FBXW11, INADL, ANK3, and ACOX3) with functions associated with stress, pain and immune responses. Our findings provide the first insights into the genetic basis of host defence to this highly contagious disease, enabling the development of an applied evolutionary framework to monitor CeMV resistance across cetacean species. Biomarkers could now be established to assess potential risk factors associated with these genes in other CeMV affected cetacean populations and species. These results could also possibly aid in the advancement of vaccines against morbilliviruses.

opencc-zeroDec 2017View details →
dryad32/100

Data from: Genomic predictions and genome-wide association study of resistance against Piscirickettsia salmonis in coho salmon (Oncorhynchus kisutch) using ddRAD sequencing

Piscirickettsia salmonis is one of the main infectious diseases affecting coho salmon (Oncorhynchus kisutch) farming, and current treatments have been ineffective for the control of this disease. Genetic improvement for P. salmonis resistance has been proposed as a feasible alternative for the control of this infectious disease in farmed fish. Genotyping by sequencing (GBS) strategies allow genotyping of hundreds of individuals with thousands of single nucleotide polymorphisms (SNPs), which can be used to perform genome wide association studies (GWAS) and predict genetic values using genome-wide information. We used double-digest restriction-site associated DNA (ddRAD) sequencing to dissect the genetic architecture of resistance against P. salmonis in a farmed coho salmon population and to identify molecular markers associated with the trait. We also evaluated genomic selection (GS) models in order to determine the potential to accelerate the genetic improvement of this trait by means of using genome-wide molecular information. A total of 764 individuals from 33 full-sib families (17 highly resistant and 16 highly susceptible) were experimentally challenged against P. salmonis and their genotypes were assayed using ddRAD sequencing. A total of 9,389 SNPs markers were identified in the population. These markers were used to test genomic selection models and compare different GWAS methodologies for resistance measured as day of death (DD) and binary survival (BIN). Genomic selection models showed higher accuracies than the traditional pedigree-based best linear unbiased prediction (PBLUP) method, for both DD and BIN. The models showed an improvement of up to 95% and 155% respectively over PBLUP. One SNP related with B-cell development was identified as a potential functional candidate associated with resistance to P. salmonis defined as DD.

opencc-zeroDec 2017View details →
dryad32/100

Data from: Genome-wide association study for body weight in cattle populations from Siberia

Body weight is a complex trait in cattle associated with commonly used commercial breeding measurements related to growth. Although many quantitative trait loci (QTL) for body weight have been identified in cattle so far, searching for genetic determinants in different breeds or environments is promising. Therefore, we carried out a genome‐wide association study (GWAS) in two cattle populations from the Russian Federation (Siberian region) using the GGP HD150K array containing 139 376 single nucleotide polymorphism (SNP) markers. Association tests for 107 550 SNPs left after filtering revealed five statistically significant SNPs on BTA5, considering a false discovery rate of less than 0.05. The chromosomal region containing these five SNPs contains the CCND2 gene, which was previously associated with average daily weight gain and body mass index in US beef cattle populations and in humans respectively. Our study is the first GWAS for body weight in beef cattle populations from the Russian Federation. The results provided here suggest that, despite the existence of breed‐ and species‐specific QTL, the genetic architecture of body weight could be evolutionarily conserved in mammals.

opencc-zeroDec 2018View details →
dryad32/100

Data from: A genome-wide association study identifies a region strongly associated with symmetrical onychomadesis on chromosome 12 in dogs

Symmetrical onychomadesis causes periodic loss of claws in otherwise healthy dogs. Genome-wide association analysis in 225 Gordon Setters identified a single region associated with symmetrical onychomadesis on chromosome 12 (spanning about 3.3 mb). A meta-analysis including also English Setters indicated that this genomic region predisposes for symmetrical onychomadesis in English Setters as well. The associated region spans most of the major histocompatibility complex and nearly 1 Mb downstream. Like many other autoimmune diseases, associations of symmetrical onychomadesis with DLA class II alleles have been reported. In this study, no associated markers were revealed within any of the DLA-DRB1, -DQA1 or -DQB1 genes, and the odds for symmetrical onychomadesis in the Gordon Setters were much higher, carrying significant single nucleotide polymorphisms compared to the odds of any of the recorded DLA-DRB1/DQA1/DQB1 haplotypes. We noticed that some of the associated DLA haplotypes were different between the English Setters and the Gordon Setters. Interestingly, associated SNP chip markers showed a more consistent pattern of allelic variants related to cases or controls regardless of breed. In conclusion, the associated genetic markers identified in this study hold the potential to aid in selection of breeding animals to reduce the frequency of symmetrical onychomadesis in the dog.

opencc-zeroDec 2015View details →
dryad32/100

Phenotypic variation and genome-wide association studies of main culm panicle node number, maximum node production rate, and degree-days to heading in rice

<p>To understand the genetic basis of main culm panicle node number, maximum node production rate, and degree-days to heading in rice (Oryza sativa), we conducted genome-wide association studies using a diversity panel of 220 rice accessions and 854,832 SNP markers generated using genotyping-by-sequencing (GBS), with 1X coverage. The raw genotype data was filtered, selecting single nucleotide polymorphisms (SNPs) having less than 50% missing data and minimum allele frequency (MAF) &gt;5%. After initial filtering, imputation was conducted using BEAGLE V4.0 in 1,075,302 SNP markers. After imputation, the dataset was filtered a second time by removing SNPs with less than 5% MAF and more than 5% missing data. A total of 854,832 SNPs were used in the genome-wide association analyses. The dataset representing the genotype data of 854,832 SNP markers by 220 rice accessions is presented here.</p>

opencc-zeroDec 2021View details →
dryad32/100

Data from: Genetic dissection of grain iron and zinc, and thousand kernel weight in wheat (Triticum aestivum L.) using genome-wide association study

<p>The study material in GWAS panel with 280 common bread wheat genotypes was selected from All India Coordinated Research Project on Wheat and Barley to map the genomic regions responsible for enhanced Grain Zinc Content (GZnC), Grain Iron Content (GZnC) and Thousand Kernel weight (TKW).</p> <p><strong>Phenotypic data:</strong></p> <p>The GWAS panel was evaluated at five different environments: E1-University of Agricultural Sciences, research farm, Dharwad (15°29'20.71"N, 74°59'3.35"E, 750m AMSL), E2-ICAR- Indian Agricultural Research Institute, New Delhi (28°38′30.5″N, 77°09′58.2″E, 228 m AMSL), E3-Indian Agricultural Research Institute, Jharkhand (24°16'58.4"N, 85°21'16.1"E, 651m AMSL), E4-ICAR-Indian Institute of Wheat and Barley, Karnal (29°41'8.2644''N, 76°59'25.9692''E,  250m AMSL), and E5-Punjab Agricultural University, Ludhiana (30o54' N, 75o48'E, 247m AMSL). Around 20 g of grain sample from each genotype were used for phenotyping GFeC and GZnC through high-throughput Energy Dispersive X-ray Fluorescence (ED-XRF) machine (model X-Supreme 8000; Oxford Instruments plc, Abingdon, United Kingdom) calibrated with glass beads-based values. To record TKW, the Numigral grain counter was used to count the grain number, the reading was set at 1000 grains and the weight of the grains was recorded in grams with an electronic balance. The GFeC, GZnC were expressed as milligram per kilogram (mg/kg), GPC in percentage (%), TKW in grams (gms).</p> <p><strong>Genotypic data:</strong></p> <p>Genomic DNA of the GWAS panel was extracted from the leaves of 21 days-old seedlings by Cetyl Trimethyl Ammonium Bromide (CTAB) method. The panel was genotyped using Axiom Wheat Breeder's Genotyping Array (Affymetrix, Santa Clara, CA, United States) having 35,143 genome-wide SNPs. The monomorphic, markers with minor allele frequency (MAF) of &lt;5%, missing data of &gt;20%, and heterozygote frequency &gt;25% were removed from the analysis. The remaining set of 14,790 high-quality SNPs was used in GWAS analysis. The detailed information of the methods and software used, data analysis and GWAS is available at DOI: 10.1038/s41598-022-15992-z.</p>

opencc-zeroJul 2022View details →
dryad32/100

Data from: Genome-wide association study for grain yield and component traits in wheat (Triticum aestivum L.)

<p>The study material in GWAS panel with 280 common bread wheat genotypes was selected from All India Coordinated Research Project on Wheat and Barley to map the genomic regions governing days to heading (DH), grain filling duration (GFD), grain number per spike (GNPS), grain weight per spike (GWPS), plant height (PH), and grain yield (GY).</p> <p><strong>Phenotypic data:</strong></p> <p>The GWAS panel was evaluated at five different environments during the 2020-21 <em>Rabi</em> (winter) season: E1-University of Agricultural Sciences, research farm, Dharwad (15°29'20.71"N, 74°59'3.35"E, 750m AMSL), E2-ICAR-Indian Agricultural Research Institute, New Delhi (28°38′30.5″N, 77°09′58.2″E, 228 m AMSL), E3-Indian Agricultural Research Institute, Jharkhand (24°16'58.4"N, 85°21'16.1"E, 651m AMSL), E4-ICAR-Indian Institute of Wheat and Barley, Karnal (29°41'8.2644''N, 76°59'25.9692''E,  250m AMSL), and E5-Punjab Agricultural University, Ludhiana (30o54' N, 75o48'E, 247m AMSL). The genotypes were planted in an augmented block design along with repeated checks (DBW187, MACS6222, WH1124, and WH1142). All the genotypes of a GWAS panel were phenotyped for six quantitative traits i.e. GWPS (gm), GY (gm), PH (cm) at five locations, GFD (days), DH (days) at four locations and GNPS (number) at two locations. Phenotypic data were analyzed using the R package 'augmentedRCBD'</p> <p><strong>Genotypic data:</strong></p> <p>Genomic DNA of the GWAS panel was extracted from the leaves of 21 days-old seedlings by Cetyl Trimethyl Ammonium Bromide (CTAB) method. The panel was genotyped using Axiom Wheat Breeder's Genotyping Array (Affymetrix, Santa Clara, CA, United States) having 35,143 genome-wide SNPs. The monomorphic, markers with minor allele frequency (MAF) of &lt;5%, missing data of &gt;20%, and heterozygote frequency &gt;25% were removed from the analysis. The remaining set of 14,790 high-quality SNPs was used in GWAS analysis. </p>

opencc-zeroJul 2022View details →
dryad32/100

UK dogs data from: Genome-wide association studies for canine hip dysplasia in single and multiple populations – implications and potential novel risk loci

<p>Background: <span>Association mapping studies of quantitative trait loci (QTL) for canine hip dysplasia (CHD) </span><span>can contribute to the understanding of the genetic background of this common and debilitating disease and might contribute to its genetic improvement. The power of association studies for CHD is limited by relatively small sample numbers for CHD records within countries, suggesting potential benefits of joining data across countries. However, this is complicated due to the use of different scoring systems across countries. In this study, we incorporated routinely assessed CHD records and genotype data of German Shepherd dogs from </span><span><span>two</span></span><span> countries </span><span><span>(UK and Sweden)</span></span><span> to perform </span><span>genome-wide association stud</span><span>ies (GWAS) within populations using different variations of CHD phenotypes. As phenotypes, dogs were either classified into cases and controls based on the </span><i><span>Fédération Cynologique Internationale</span></i><span> (FCI) five-level grading of the worst hip or the FCI grade was treated as an ordinal trait. </span><span><span>In a subsequent meta-analysis, we added publicly available data from a Finnish population and performed the GWAS across all populations.</span></span><span> Genetic associations for the CHD phenotypes were evaluated in a linear mixed model using 62,089 SNPs.</span></p> <p>Results: <span><span>Multiple SNPs with genome-wide significant and suggestive</span></span><span><span> associations</span></span><span> </span><span><span>were detected in single-population GWAS and the meta-analysis.</span></span><span> Few of these SNPs overlapped between populations </span><span><span>or between single-population GWAS and the meta-analysis</span></span><span>, suggesting that many CHD-related QTL are population-specific. More significant or suggestive SNPs were identified when FCI grades were used as phenotypes in comparison to the case-control approach. </span><i><span>MED13</span></i><span> (Chr 9) and </span><i><span>PLEKHA7</span></i><span> (Chr 21) emerged as novel positional candidate genes associated with hip dysplasia.</span></p> <p>Conclusions: <span>Our findings confirm the complex genetic nature of hip dysplasia in dogs, with multiple loci associated with the trait, </span><span><span>most</span></span><span> of which are population-specific. Routinely assessed CHD information collected across countries provide an opportunity to increase sample sizes and statistical power for association studies. While the lack of standardisation of CHD assessment schemes across countries poses a challenge, we showed that conversion of traits can be utilised to overcome this obstacle.</span></p>

opencc-zeroAug 2021View details →
zenodo32/100

Genome-wide association study of polygenic risk score-defined phenotype suffers from inflated test-statistics

<p>Simulation results from running the following&nbsp;script 100&nbsp;times:&nbsp;https://github.com/euffelmann/paper-ad_prs_extremes/blob/main/scripts/ad_prs_extremes_simulation.R.</p> <p>These files can be used to reproduce tables and figures in: https://github.com/euffelmann/paper-ad_prs_extremes</p>

opencc-by-4.0Jul 2022View details →
dryad32/100

Summary statistics from a genome-wide association study of narcolepsy

<p>Type 1 narcolepsy (T1N) is a neurological condition, in which the death of hypocretin-producing neurons in the lateral hypothalamus leads to excessive daytime sleepiness and symptoms of abnormal Rapid Eye Movement (REM) sleep. Known triggers for narcolepsy are influenza-A infection and associated immunization during the 2009 H1N1 influenza pandemic. Here, we genotyped all remaining consented narcolepsy cases worldwide and assembled this with the existing genotyped individuals. We used this multi-ethnic sample in genome wide association study (GWAS) to dissect disease mechanisms and interactions with environmental triggers (5,339 cases and 20,518 controls). Overall, we found significant associations with HLA (2 GWA significant subloci) and 11 other loci. Six of these other loci have been previously reported (<em>TRA</em>, <em>TRB</em>, <em>CTSH</em>, <em>IFNAR1</em>, <em>ZNF365</em> and <em>P2RY11</em>) and five are new (<em>PRF1</em>, <em>CD207</em>, <em>SIRPG</em>, <em>IL27</em> and <em>ZFAND2A</em>). Strikingly, in vaccination-related cases, GWA significant effects were found in <em>HLA</em>, <em>TRA</em>, and in a novel variant near <em>SIRPB1</em>. Furthermore, <em>IFNAR1</em>-associated polymorphisms regulated dendritic cell response to influenza-A infection in vitro (p-value =1.92*10<sup>-25</sup>). A partitioned heritability analysis indicated specific enrichment of functional elements active in cytotoxic and helper T cells. Furthermore, functional analysis showed the genetic variants in <em>TRA</em> and <em>TRB</em> loci act as remarkably strong chain usage QTLs for <em>TRAJ*24</em> (p-value = 0.0017), <em>TRAJ*28</em> (p-value = 1.36*10<sup>-10</sup>) and <em>TRBV*4-2</em> (p-value = 3.71*10-<sup>117</sup>). This was further validated in TCR sequencing of 60 narcolepsy cases and 60 DQB1*06:02 positive controls, where chain usage effects were further accentuated. Together these findings show that the autoimmune component in narcolepsy is defined by antigen presentation, mediated through specific T cell receptor chains, and modulated by influenza-A as a critical trigger.</p>

opencc-zeroDec 2022View details →
zenodo32/100

Genome-Wide Association Studies meta-analysis uncovers NOJO and SGS3 novel genes involved in Arabidopsis thaliana primary root development and plasticity

<p>Postembryonic primary root growth relies on meristems that harbour multipotent stem cells that produce new cells that will duplicate and provide all the different root cell types. <em>Arabidopsis thaliana</em> primary root growth has become a model for evo-devo studies due to its simplicity and facility to record cell proliferation and differentiation. To identify new genetic components relevant to primary root growth, we used a Genome-Wide Association Studies (GWAS) meta-analysis approach using data published in the last decade. In this work, we performed intra and inter-studies analyses to discover new genetic components that could participate in primary root growth. We used 639 accessions from nine different studies and performed different GWAS tests ranging from single studies and pairwise analysis with high correlation associations, analyzing the same number of accessions in different studies to using the daily data of the root growth kinetic of the same research. We found that primary root growth changes were associated with 41 genomic loci, of which six (14.6%) have been previously described as inhibitors or promoters of primary root growth. The knockdown of genes associated with two of these loci: a gene that participates in Trans-acting siRNAs (tasiRNAs) processing <em>Suppressor of Gene Silencing</em> (<em>SGS3</em>) and a gene with a Sterile Alpha Motif (SAM) confirmed their participation as repressors of primary root growth. As none has been shown to participate in this developmental process before, our GWAS analysis identified new genes that participate in primary root growth. Overall, our findings provide novel insights into the genomic basis of root development and further demonstrate the usefulness of GWAS meta-analyses in non-human species.</p>

opencc-byJul 2023View details →
zenodo32/100

GWAS summary stats in "A web-based genome-wide association study reveals the susceptibility loci of common adverse events following COVID-19 vaccination in the Japanese population."

<p>Summary stats of the genome-wide meta-analysis with METAL software in the article "A web-based genome-wide association study reveals the susceptibility loci of common adverse events following COVID-19 vaccination in the Japanese population."</p><p>https://doi.org/10.1038/s41598-023-47632-5<br>Due to the absence of cases in either population, we were unable to perform GWAS for the following conditions.<br>constipation at BNT162b1 1st dose<br>dyspnea at BNT162b1 1st dose<br>eczema (long-term rash) at BNT162b1 1st dose<br>dyspnea at BNT162b1 2nd dose<br>sneeze at mRNA-1273 1st dose<br>dyspnea at mRNA-1273 1st dose</p>

opencc-by-4.0Sep 2023View details →
ClinicalTrials.gov32/100

Identification of Genetic Polymorphisms Related to Propofol Requirement and Recovery Through Genome-wide Association Study (GWAS) in Total Intravenous Anesthesia for Clipping of Unruptured Cerebral An

ClinicalTrials.gov study NCT03087383. IPD Sharing: NO. Countries: 1. Publications: 1.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov32/100

A Genome-Wide Association Study for Neonatal Diseases

ClinicalTrials.gov study NCT04074824. IPD Sharing: Not stated. Countries: 1. Publications: 7.

restrictedIPD-UNDECIDEDFeb 2026View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record