Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
505
datasets available to search
ShareScore release 0.9.0
Dataset results
505 results for “genome-wide association”
Genome-wide association mapping to identify genetic loci for cold tolerance and cold recovery during germination in rice
Open the record for dataset details and reuse information.
Data from: Genome-wide association analysis of type 2 diabetes in the EPIC-InterAct study
Open the record for dataset details and reuse information.
Dataset for: Identification of genomic regions of wheat associated with grain Fe and Zn content under drought and heat stress using genome-wide association study
Open the record for dataset details and reuse information.
Sequence-based genome-wide association study of individual milk mid-infrared wavenumbers in mixed-breed dairy cattle
Open the record for dataset details and reuse information.
Genome-wide association results from: Transcriptomic stratification of late-onset Alzheimer’s cases reveals novel genetic modifiers of disease pathology
Open the record for dataset details and reuse information.
A genome-wide association study of deafness in three canine breeds
Open the record for dataset details and reuse information.
Data from: Genome-wide association analyses in the model rhizobium Ensifer meliloti
Open the record for dataset details and reuse information.
Large scale across-breed genome-wide association study reveals a variant in HMGA2 associated with inguinal cryptorchidism risk in dogs
Open the record for dataset details and reuse information.
Supplemental material for: Genome-wide association study and fine-mapping using imputed sequences to prioritize candidate genes for 30 complex traits in 50,309 Holstein bulls
Open the record for dataset details and reuse information.
Data from: Genome-wide analysis reveals associations between climate and regional patterns of adaptive divergence and dispersal in American pikas
Open the record for dataset details and reuse information.
Supplementary information for: Redundancy analysis, genome-wide association studies, and the pigmentation of brown trout (Salmo trutta L.)
Open the record for dataset details and reuse information.
Data from: Genome-wide SNP identification and association mapping for seed mineral concentration in Mung bean (Vigna radiata L.)
<p><span><span>Mung bean (<i>Vigna radiata</i> L.) quality is dependent on seed chemical composition, which in turn determines the benefits of mung bean consumption for human health. While rich in a range of nutritional components, such as protein, macro- and micro- nutrients, carbohydrates and vitamins, mung bean remains less well studied than other legume crops. Mung bean genomics and genetic resources are relatively sparse. To further improve nutritional levels of mung bean grain requires genome-wide marker system tools. The objectives of this research were to develop these tools and conduct nutrient analysis in order to 1) identify single nucleotide polymorphisms (SNPs) using genotyping by sequencing (GBS) and to 2) perform genome-wide association studies (GWAS) for levels of calcium, iron, potassium, manganese, phosphorous, sulfur, and zinc in mung bean grain produced over two years of field experiment. A total of 112 GWAS models were explored using 6,486 high quality SNPs discovered in 92 cultivated mung bean accessions chosen from USDA core collection that represented 13 countries. The data obtained allowed for the identification of 43 associated SNPs and 20 main genomic regions that explained on average 22 % of the overall variation in seed macro- and micro- nutrients concentration on the basis of a multiple-year analysis. Most of the regions discovered in this study provide valuable candidate gene to use in future breeding of new varieties of mung bean with novel nutritional properties. Identification of the <a>underlying genes</a> will help to reveal the genetic control of mung bean seed nutritional property. Other SNPs identified in this study will serve as important resources to enable marker-assisted selection (MAS) in the species <i>V</i>. <i>radiata</i>, including wide and narrow crosses with / between cultivated and wild mung bean.</span></span></p>
Data from: Genome-wide association mapping of resistance to Septoria nodorum leaf blotch in a Nordic spring wheat collection
Parastagonospora nodorum is the causal agent of septoria nodorum blotch (SNB) in wheat. It is the most important leaf blotch pathogen in Norwegian spring wheat. Several quantitative trait loci (QTL) for SNB susceptibility have been identified. Some of these QTL are the result of underlying gene-for-gene interactions involving necrotrophic effectors (NEs) and corresponding sensitivity (Snn) genes. A collection of diverse spring wheat lines was evaluated for SNB resistance/susceptibility over seven growing seasons in the field. In addition, wheat seedlings were inoculated and infiltrated with culture filtrates (CFs) from four single spore isolates and infiltrated with semi-purified NEs (SnToxA, SnTox1 and SnTox3) under greenhouse conditions. In adult plants, the most stable SNB resistance QTL were located on 2B, 2D, 4A, 4B, 5A, 6B, 7A and 7B. The QTL on 2D was effective most years in the field. At the seedling stage, the most significant QTL after inoculation were located on 1A, 1B, 3A, 4B, 5B, 6B, 7A and 7B. The QTL on 3A and 6B were significant both after inoculation and CF infiltration, indicating the presence of novel NE-Snn interactions. The QTL on 4B and 7A were significant in both seedlings and adult plants. Correlations between SnToxA sensitivity and disease severity in the field were significant. To our knowledge, this is the first genome wide association mapping study (GWAS) to investigate SNB resistance at the adult plant stage under field conditions.
Data from: Imputation of canine genotype array data using 365 whole-genome sequences improves power of genome-wide association studies
Genomic resources for the domestic dog have improved with the widespread adoption of a 173k SNP array platform and updated reference genome. SNP arrays of this density are sufficient for detecting genetic associations within breeds but are underpowered for finding associations across multiple breeds or in mixed-breed dogs, where linkage disequilibrium rapidly decays between markers, even though such studies would hold particular promise for mapping complex diseases and traits. Here we introduce an imputation reference panel, consisting of 365 diverse, whole-genome sequenced dogs and wolves, which increases the number of markers that can be queried in genome-wide association studies approximately 130-fold. Using previously genotyped dogs, we show the utility of this reference panel in identifying potentially novel associations, including a locus on CFA20 significantly associated with cranial cruciate ligament disease, and fine-mapping for canine body size and blood phenotypes, even when causal loci are not in strong linkage disequilibrium with any single array marker. This reference panel resource will improve future genome-wide association studies for canine complex diseases and other phenotypes.
Data from: Genome-wide association mapping of phenotypic traits subject to a range of intensities of natural selection in Timema cristinae
The genetic architecture of adaptive traits can reflect the evolutionary history of populations and also shape divergence among populations. Despite this central role in evolution, relatively little is known regarding the genetic architecture of adaptive traits in nature, particularly for traits subject to known selection intensities. Here we quantitatively describe the genetic architecture of traits that are subject to known intensities of differential selection between host plant species in Timema cristinae stick insects. Specifically, we used phenotypic measurements of 10 traits and 211,004 single-nucleotide polymorphisms (SNPs) to conduct multilocus genome-wide association mapping. We identified a modest number of SNPs that were associated with traits and sometimes explained a large proportion of trait variation. These SNPs varied in their strength of association with traits, and both major and minor effect loci were discovered. However, we found no relationship between variation in levels of divergence among traits in nature and variation in parameters describing the genetic architecture of those same traits. Our results provide a first step toward identifying loci underlying adaptation in T. cristinae. Future studies will examine the genomic location, population differentiation, and response to selection of the trait-associated SNPs described here.
Data from: Genome-wide exon-capture approach identifies genetic variants of Norway spruce genes associated with susceptibility to Heterobasidion parviporum infection
Root and butt rot caused by members of the Heterobasidion annosum species complex is the most economically important disease of conifer trees in boreal forests. Wood decay in the infected trees dramatically decreases their value and causes considerable losses to forest owners. Trees vary in their susceptibility to Heterobasidion infection, but the genetic determinants underlying the variation in the susceptibility are not well-understood. We performed the identification of Norway spruce genes associated with the resistance to Heterobasidion parviporum infection using genome-wide exon-capture approach. Sixty-four clonal Norway spruce lines were phenotyped, and their responses to H. parviporum inoculation were determined by lesion length measurements. Afterwards, the spruce lines were genotyped by targeted resequencing and identification of genetic variants (SNPs). Genome-wide association analysis identified 10 SNPs located within 8 genes as significantly associated with the larger necrotic lesions in response to H. parviporum inoculation. The genetic variants identified in our analysis are potential marker candidates for future screening programs aiming at the differentiation of disease-susceptible and resistant trees.
Data from: Genome-wide association study of outcrossing in cytoplasmic male sterile lines of rice
Stigma exsertion and panicle enclosure of male sterile lines are two key determinants of outcrossing in hybrid rice seed production. Based on 43,394 single nucleotide polymorphism markers, 217 cytoplasmic male sterile lines were assigned into two subpopulations and a mixed-group where the LD decay distance varied from 975 to 2,690 kb. Genome-wide association studies (GWAS) were performed for stigma exsertion rate (SE), panicle enclosure rate (PE) and seed-setting rate (SSR). A total of 154 significant association signals (P < 0.001) were identified. They were situated in 27 quantitative trait loci (QTLs), including 11 for SE, 6 for PE, and 10 for SSR. It was shown that six of the ten QTLs for SSR were tightly linked to QTLs for SE or/and PE with the expected allelic direction. These QTL clusters could be targeted to improve the outcrossing of female parents in hybrid rice breeding. Our study also indicates that GWAS-base QTL mapping can complement and enhance previous QTL information for understanding the genetic relationship between outcrossing and its related traits.
Data from: Genome-wide association study of an unusual dolphin mortality event reveals candidate genes for susceptibility and resistance to cetacean morbillivirus
Infectious diseases are significant demographic and evolutionary drivers of populations, but studies about the genetic basis of disease resistance and susceptibility are scarce in wildlife populations. Cetacean morbillivirus (CeMV) is a highly contagious disease that is increasing in both geographic distribution and incidence, causing unusual mortality events (UME) and killing tens of thousands of individuals across multiple cetacean species worldwide since the late 1980's. The largest CeMV outbreak in the Southern Hemisphere reported to date occurred in Australia in 2013, where it was a major factor in a UME, killing mainly young Indo-Pacific bottlenose dolphins (Tursiops aduncus). Using cases (non-survivors) and controls (putative survivors) from the most affected population, we carried out a genome-wide association study to identify candidate genes for resistance and susceptibility to CeMV. The genomic dataset consisted of 278,147,988 sequence reads and 35,493 high quality SNPs genotyped across 38 individuals. Association analyses found highly significant differences in allele and genotype frequencies amongst cases and controls at 65 SNPs, and Random Forests conservatively identified eight as candidates. Annotation of these SNPs identified five candidate genes (MAPK8, FBXW11, INADL, ANK3, and ACOX3) with functions associated with stress, pain and immune responses. Our findings provide the first insights into the genetic basis of host defence to this highly contagious disease, enabling the development of an applied evolutionary framework to monitor CeMV resistance across cetacean species. Biomarkers could now be established to assess potential risk factors associated with these genes in other CeMV affected cetacean populations and species. These results could also possibly aid in the advancement of vaccines against morbilliviruses.
Data from: Validating genome-wide association candidates controlling quantitative variation in nodulation
Genome-wide association (GWA) studies offer the opportunity to identify genes that contribute to naturally occurring variation in quantitative traits. However, GWA relies exclusively on statistical association, so functional validation is necessary to make strong claims about gene function. We used a combination of gene-disruption platforms (Tnt1 retrotransposons, hairpin RNA-interference constructs, and CRISPR/Cas9 nucleases) together with randomized, well-replicated experiments to evaluate the function of genes that an earlier GWA study in Medicago truncatula had identified as candidates contributing to variation in the symbiosis between legumes and rhizobia. We evaluated ten candidate genes found in six clusters of strongly associated single nucleotide polymorphisms, selected on the basis of their strength of statistical association, proximity to annotated gene models, and root or nodule expression. We found statistically significant effects on nodule production for three candidate genes, each validated in two independent mutants. Annotated functions of these three genes suggest their contributions to quantitative variation in nodule production occur through processes not previously connected to nodulation, including phosphorous supply and salicylic acid-related defense response. These results demonstrate the utility of GWA combined with reverse mutagenesis technologies to discover and validate genes contributing to naturally occurring variation in quantitative traits. The results highlight the potential for GWA to complement forward genetics in identifying the genetic basis of ecologically and economically important traits.
Data from: Genomic predictions and genome-wide association study of resistance against Piscirickettsia salmonis in coho salmon (Oncorhynchus kisutch) using ddRAD sequencing
Piscirickettsia salmonis is one of the main infectious diseases affecting coho salmon (Oncorhynchus kisutch) farming, and current treatments have been ineffective for the control of this disease. Genetic improvement for P. salmonis resistance has been proposed as a feasible alternative for the control of this infectious disease in farmed fish. Genotyping by sequencing (GBS) strategies allow genotyping of hundreds of individuals with thousands of single nucleotide polymorphisms (SNPs), which can be used to perform genome wide association studies (GWAS) and predict genetic values using genome-wide information. We used double-digest restriction-site associated DNA (ddRAD) sequencing to dissect the genetic architecture of resistance against P. salmonis in a farmed coho salmon population and to identify molecular markers associated with the trait. We also evaluated genomic selection (GS) models in order to determine the potential to accelerate the genetic improvement of this trait by means of using genome-wide molecular information. A total of 764 individuals from 33 full-sib families (17 highly resistant and 16 highly susceptible) were experimentally challenged against P. salmonis and their genotypes were assayed using ddRAD sequencing. A total of 9,389 SNPs markers were identified in the population. These markers were used to test genomic selection models and compare different GWAS methodologies for resistance measured as day of death (DD) and binary survival (BIN). Genomic selection models showed higher accuracies than the traditional pedigree-based best linear unbiased prediction (PBLUP) method, for both DD and BIN. The models showed an improvement of up to 95% and 155% respectively over PBLUP. One SNP related with B-cell development was identified as a potential functional candidate associated with resistance to P. salmonis defined as DD.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.