Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
606
datasets available to search
ShareScore release 0.9.0
Dataset results
606 results for “association genetics”
GWAS Summary Statistics from "Sex and statin-related genetic associations at the PCSK9 gene locus – results of genome-wide association meta-analysis"
<p>GWAMA summary statistics of PCSK9 levels stratified by sex and statin useage in Europeans.</p> <p>When using this data, please cite:</p> <p>Pott, J., Kheirkhah, A., Gadin, J.R. <em>et al.</em> Sex and statin-related genetic associations at the <em>PCSK9</em> gene locus: results of genome-wide association meta-analysis. <em>Biol Sex Differ</em> <strong>15</strong>, 26 (2024). https://doi.org/10.1186/s13293-024-00602-6</p> <p>All txt files contain the following columns:</p> <ul> <li>markername (unique SNP ID)</li> <li>chr</li> <li>bp_hg19 (base position according to hg19)</li> <li>EA (effect allele)</li> <li>OA (other allele)</li> <li>EAF (effect allele frequency)</li> <li>info (minimal info score across all used studies)</li> <li>nSamples (sample size per SNP)</li> <li>nStudies (in case of double-stratified data: number of studies; in case of single-stratified data: 2, as it is a meta-analysis of the two double-stratified data sets)</li> <li>beta (effect estimate)</li> <li>SE (standard error)</li> <li>pval (p-value)</li> <li>I2 (SNP heterogeneity across studies)</li> <li>invalidAssoc (TRUE/FALSE flag if this variant was excluded in our analysis)</li> <li>reason4exclusion (reason why this SNP was excluded)</li> <li>phenotype (phenotyp setting)</li> </ul>
Genetic association analysis of anti-VEGF treatment response in neovascular age-related macular degeneration
<p>Summary statisics of an association study of 6,908,005 genetic variants with anti-VEGF nAMD treatment response in 179 treatment-naïve nAMD probands. This dataset supplements the publication "Genetic Association Analysis of Anti-VEGF Treatment Response in Neovascular Age-Related Macular Degeneration" (DOI: 10.3390/ijms23116094). Details regarding the methods and version numbers can be found in the corresponding manuscript.</p>
Metabomatching: Using Genetic Association to Identify Metabolites in Proton NMR Spectroscopy. CoLaus Pseudospectra.
<p>Summary statistics between urine NMR metabolome features and genotypes in the CoLaus cohort. Used as test pseudospectra for metabomatching, a method for metabolite identification using genetic spiking.</p>
Metabomatching: Using Genetic Association to Identify Metabolites in Proton NMR Spectroscopy. SHIP Pseudospectra.
<p>Summary statistics between urine NMR metabolome features and genotypes in the SHIP cohort. Used as test pseudospectra for metabomatching, a method for metabolite identification using genetic spiking.</p>
Supplementary data: Agro-morphological and molecular characterization reveal deep insights in promising genetic diversity and marker-trait associations in Fagopyrum esculentum and F. tataricum
<p>Our study focuses on the global/European buckwheat germplasm collected as part of the ECOBREDD project. The potential of this highly diverse collection for organic buckwheat breeding was evaluated at two complementary levels: phenotypic and genetic. Here, we characterized the phenotypic and genetic diversity of a global collection of the two cultivated buckwheat species <em>Fagopyrum esculentum</em> and <em>F. tataricum</em> (190 and 51 accessions, respectively) using 37 agro-morphological traits and 24 SSR markers (Simple Sequence Repeats) (see publication and info sheet of the data).</p>
Genome-wide association summary statistics of chronic musculoskeletal pain at four anatomic sites and their genetically independent components
<p>The dataset contains results of a genome-wide association study of distinct chronic musculoskeletal pain conditions: back pain, knee pain, neck pain, and hip pain. Additionally, there are genome-wide association summary statistics for four genetically independent components of pain conditions, listed above. For more details, please, read the paper XXX.</p> <p>All files contain association summary statistics for genome-wide association meta-analysis of the 265,000 white British individuals from the UK Biobank and additional 191,580 individuals of European Ancestry from the UK biobank (total N = 456,580). Cases and controls were defined based on questionnaire responses. First, participants responded to “Pain type(s) experienced in the last months” followed by questions inquiring if the specific pain had been present for more than 3 months. Those who reported back, neck or shoulder, hip, or knee pain lasting more than 3 months were considered chronic back, neck/shoulder, hip, and knee pain cases, respectively. Participants reporting no such pain lasting longer than 3 months were considered controls (regardless of whether they had another regional chronic pain, such as abdominal pain, or not). Individuals who preferred not to answer were excluded from the study. Besides this, we excluded individuals who reported more than 3 months of pain all over the body.</p> <p>The data are provided on an "AS-IS" basis, without warranty of any type, expressed or implied, including but not limited to any warranty as to their performance, merchantability, or fitness for any particular purpose. If investigators use these data, any and all consequences are entirely their responsibility. By downloading and using these data, you agree that you will cite the appropriate publication in any communications or publications arising directly or indirectly from these data; for utilization of data available prior to publication, you agree to respect the requested responsibilities of resource users under 2003 Fort Lauderdale principles; you agree that you will never attempt to identify any participant. This research has been conducted using the UK Biobank Resource and the use of the data is guided by the principles formulated by the UK Biobank.</p> <p><strong>When using downloaded data, please cite the corresponding paper and this repository:</strong></p> <ol> <li>Tsepilov et al 2020</li> </ol> <p><strong>Funding:</strong></p> <p>The work of YSA and SZS was supported by the Russian Ministry of Education and Science under the 5-100 Excellence Programme and by the Federal Agency of Scientific Organizations via the Institute of Cytology and Genetics (project 0324-2019-0040). The work of YAT, ASSh, and EEE was supported by the Russian Foundation for Basic Research (project 19-015-00151). The contribution of LСK was funded by PolyOmica. Dr. Suri was supported by VA Career Development Award # 1IK2RX001515 from the United States (U.S.) Department of Veterans Affairs Rehabilitation Research and Development (RR&D) Service. Dr. Suri is a Staff Physician at the VA Puget Sound Health Care System. The contents of this work do not represent the views of the U.S. Department of Veterans Affairs or the United States Government.</p> <p><strong>List of files:</strong></p> <ol> <li>Back_output_done.csv: GWAS summary statistics for the chronic back pain</li> <li>gpc1_output_done.csv: GWAS summary statistics for the GIP1</li> <li>gpc2_output_done.csv: GWAS summary statistics for the GIP2</li> <li>gpc3_output_done.csv: GWAS summary statistics for the GIP3</li> <li>gpc4_output_done.csv: GWAS summary statistics for the GIP4</li> <li>Hip_output_done.csv: GWAS summary statistics for the chronic hip pain</li> <li>Knee_output_done.csv: GWAS summary statistics for the chronic knee pain</li> <li>Neck_output_done.csv: GWAS summary statistics for the chronic neck pain</li> </ol> <p><strong>Column headers:</strong></p> <ol> <li>gwas_id: uninformative field</li> <li>rs_id: dbSNP rsID (GRCh37 build) </li> <li>snp_num: uninformative field</li> <li>chr: chromosome (GRCh37 build) </li> <li>bp: position (GRCh37 build) </li> <li>ea: effect allele (coded as "1")</li> <li>ra: reference allele (coded as "0")</li> <li>eaf: effect allele frequency</li> <li>af_ref: uninformative field</li> <li>beta: effect size of effect allele</li> <li>se: standard error of effect size</li> <li>p: P-value of association (without GC correction)</li> <li>n:Total sample size</li> <li>z: Z-statistic of association</li> <li>info: uninformative field</li> <li>af_outlier: uninformative field</li> <li>pz_outlier: uninformative field</li> </ol>
GWAS Summary Statistics for Publication: Identifying novel genetic and phenotypic associations to genomic features by leveraging off-target reads in exome sequencing data
<p>This dataset contains summary statistics for genome-wide association studies (GWAS) conducted on genomic features derived from off-target reads in whole-exome sequencing (WES) data. The study utilized tools like Seeing Beyond the Target (SBT) and ImReP to construct novel phenotypic features from unmapped reads in ~50,000 participants in the UK Biobank. Features include mitochondrial DNA (mtDNA) copy number, ribosomal DNA (rDNA) copy number (5S, 18S, 28S), immune repertoire metrics (e.g., T-cell receptor alpha diversity), and microvial genome load (viral and fungal).</p> <p>Summary statistics can be used for replication studies, meta-analyses, or further exploration of these phenotypes.</p>
A scalable, accurate, and universal analysis framework using individual-level allele frequency for large-scale genetic association studies in an admixed population
<p>Inclusion of individuals with diverse or admixed genetic ancestries is crucial to discover novel findings that may be missed by genomics analyses rooted solely in Caucasian population. Here, we present an analysis framework, SPAmix, which is scalable to a large-scale biobank data analysis including hundreds of thousands of admixed individuals and is universally applicable to various types of complex traits including binary trait, quantitative trait, time-to-event trait, longitudinal traits, etc. For each genetic variant, SPAmix uses genotype data and genetic principal components (PCs) to estimate individual-level allele frequency, which is subsequently used to calibrate p values via a retrospective analysis. A hybrid strategy including saddlepoint approximation (SPA) can greatly increase the accuracy to analyze rare genetic variants, especially if the phenotypic distribution is unbalanced or extremely unbalanced. Compared to Tractor, SPAmix does not require local ancestry information and can be straightforwardly applicable to a multi-way admixed population. Meanwhile, SPAmix can also be extended to SPAmix<sub>local</sub> in which the local ancestry can be incorporated if available. In addition, we propose SPAmix<sub>CCT</sub> to combine the p values of SPAmix and SPAmix<sub>local</sub> via Cauchy combination (CCT). SPAmix<sub>local</sub> performs close to Tractor when analyzing quantitative traits and is more accurate when analyzing binary traits with an unbalanced case-control ratio. And SPAmix<sub>CCT </sub>is an optimal unified approach for various cross-ancestry genetic architectures. Extensive simulation studies and real data analyses of 369,314 UK Biobank individuals from multiple ancestries demonstrated that SPAmix is scalable and can discover novel hits while controlling type I error rates well.</p>
Data from: Using genetic relatedness to understand heterogeneous distributions of urban rat-associated pathogens
<p>Urban Norway rats (<i>Rattus norvegicus</i>) carry several pathogens transmissible to people. However, pathogen prevalence can vary across fine spatial scales (i.e., by city block). Using a population genomics approach, we sought to describe rat movement patterns across an urban landscape, and to evaluate whether these patterns align with pathogen distributions. We genotyped 605 rats from a single neighborhood in Vancouver, Canada and used 1,495 genome-wide single nucleotide polymorphisms to identify parent-offspring and sibling relationships using pedigree analysis. We resolved 1,246 pairs of relatives, of which only 1% of pairs were captured in different city blocks. Relatives were primarily caught within 33 meters of each other leading to a highly leptokurtic distribution of dispersal distances. Using binomial generalized linear mixed models we evaluated whether family relationships influenced rat pathogen status with the bacterial pathogens <i>Leptospira interrogans</i>, <i>Bartonella tribocorum</i>, and <i>Clostridium difficile</i>, and found that an individual's pathogen status was not predicted any better by including disease status of related rats. The spatial clustering of related rats and their pathogens lends support to the hypothesis that spatially restricted movement promotes the heterogeneous patterns of pathogen prevalence evidenced in this population. <span>Our findings also highlight the utility of evolutionary tools to understand movement and rat-associated health risks in urban landscapes.</span></p>
Data from: Association genetics of growth and adaptive traits in loblolly pine (Pinus taeda L.) using whole-exome-discovered polymorphisms
In the United States, forest genetics research began over 100 years ago and loblolly pine breeding programs were established in the 1950s. However, the genetics underlying complex traits of loblolly pine remains to be discovered. To address this, adaptive and growth traits were measured and analyzed in a clonally tested loblolly pine (Pinus taeda L.) population. Over 2.8 million single nucleotide polymorphism (SNP) markers detected from exome sequencing were used to test for single locus associations, SNP-SNP interactions and correlation of individual heterozygosity with phenotypic traits. A total of 36 SNP-trait associations were found for specific leaf area (5 SNPs), branch angle (2), crown width (3), stem diameter (4), total height (9), carbon isotope discrimination (4), nitrogen concentration (2), and pitch canker resistance traits (7). Eleven SNP-SNP interactions were found to be associated with branch angle (1 SNP-SNP interaction), crown width (2), total height (2), carbon isotope discrimination (2), nitrogen concentration (1), and pitch canker resistance (3). Non-additive effects imposed by dominance and epistasis account for a large fraction of the genetic variance for the quantitative traits. Genes that contain the identified SNPs have a wide spectrum of functions. Individual heterozygosity positively correlated with water use efficiency and nitrogen concentration. In conclusion, multiple effects identified in this study influence the performance of loblolly pines, provide resources for understanding the genetic control of complex traits, and have potential value for assessing with breeding through marker assisted selection and genomic selection.
Common Genetic Variants in FOXP2 are Not Associated with Individual Differences in Language Development
<p>Three data sets used in” Common Genetic Variants <em>in FOXP2</em> Are Not Associated with Individual Differences in Language Development” are provided. The discovery data set was comprised of 834 children who were members of a Longitudinal sample and children who were members of a School sample. Both samples are contained in the Iowa data set. The Iowa data set contains a quantitative variable LCOMP that represents a composite z-score representing oral language ability. The data set also identifies which sample the children belonged to and the allele calls for 13 tag SNPs located across <em>FOXP2</em>. A second data file, ELVS, contains data from a separate sample of children who were used to test for replication of inconsistent evidence of an association between language and the SNP rs1916988. The ELVS file provides a composite oral language score scaled in standard score units (mean=100, SD=15) and the genotype calls for the SNP rs1916988.</p>
Data and modeling results for publication: Landscape genetics indicate recently increased habitat fragmentation in African forest-associated chafers
<ul> <li>DNA sequences: <em>cox1</em> and ITS1 alignments</li> <li>spatial records (in hypervolume archive)</li> <li>spatial principal component 1-3 used for <em>hypervolume</em> models (in hypervolume archive)</li> <li>Present and past species distribution models (SDMs): <ul> <li><em>biomod2</em> ensemble SDMs <ul> <li>Present</li> <li>Holocene Altithermal</li> <li>Last Glacial Maximum</li> </ul> </li> <li><em>biomod2</em> SDMs for single PMIP3 models <ul> <li>Present</li> <li>Holocene Altithermal</li> <li>Last Glacial Maximum</li> </ul> </li> <li><em>hypervolume</em> SDMs</li> </ul> </li> <li>landscape connectivity models <ul> <li>circuitscape (for F0, F1, and F2)</li> <li>least cost corridors and paths (for F0, F1, and F2)</li> </ul> </li> </ul>
Data from: The genetic basis of traits associated with the evolution of serpentine endemism in monkeyflowers
<p>The floras on chemically and physically challenging soils, such as gypsum, shale, and serpentine, are characterized by narrowly endemic species. The evolution of edaphic endemics may be facilitated or constrained by genetic correlations among traits contributing to adaptation and reproductive isolation across soil boundaries. The yellow monkeyflowers in the <em>Mimulus guttatus</em> species complex are an ideal system in which to examine these evolutionary patterns. To determine the genetic basis of adaptive and prezygotic isolating traits, we performed genetic mapping experiments with F2 hybrids derived from a cross between a serpentine endemic, <em>M. nudatus</em>, and its close relative <em>M. guttatus</em>. Few large effect and many small effect QTL contribute to interspecific divergence in life history, floral and leaf traits, and a history of directional selection contributed to trait divergence. Loci contributing to adaptive traits and prezygotic reproductive isolation overlap, and their allelic effects are largely in the direction of species divergence. These loci contain promising candidate genes regulating flowering time and plant organ size. Together our results suggest that genetic correlations among traits can facilitate the evolution of adaptation and speciation and may be a common feature of the genetic architecture of divergence between edaphic endemics and their widespread relatives.</p>
Genetic Architecture Reconciles Linkage and Association Studies of Complex Traits
<p>This (zipped) folder contains 3 sub-folders:</p> <p>#**********************************************************************************************************<br>The "bin" folder contains fuctions and gentic maps needed for analyes<br>bin \<br> predLink.R - function to predict linkage <br> sibREML_v0.1.1.R - function to run SibREML<br> sim-sib-array.R - script to simulate sib-pairs from parental haplotypes<br> Summarised_genetic_map_bcf.txt - genetic map per 0.5-cM long segments, based on map from bcftools <br> (BCFtools: https://samtools.github.io/bcftools/bcftools.html)<br> Summarised_genetic_map_OMNI.txt - genetic map per 0.5-cM long segments, based on OMNI map <br> (https://github.com/joepickrell/1000-genomes-genetic-maps/tree/master/interpolated_OMNI)<br>#**********************************************************************************************************</p> <p> </p> <p>#**********************************************************************************************************<br>The "SIM" folder contains the simulation pipeline (scripts 01-15) as well as IBD sharing and simulated phenotypes for Simulated sib-pairs.<br>SIM \<br> 01_sim-sib-array.sh *pre-run*<br> 02_bed_recode_bcf_map.sh *pre-run*<br> 03_make_merlin.R *pre-run*<br> 04_error_merlin.sh *pre-run*<br> 05_merlin_IBD.sh *pre-run*<br> 06_sample_causal_snps.R *pre-run*<br> 07_simulate_pheno.sh *pre-run*<br> 08_bhat_gwas.R *can be run using provided data* <br> 09_Linkage_VH.R *can be run using provided data* <br> 10_predLink.R *can be run using provided data*<br> 11_phi_hat.R *can be run using provided data*<br> 12_IBD_Mb.R *can be run using provided data*<br> 13_IBD_cM_recombrate_stratified.R *can be run using provided data*<br> 14_SibREML.R *can be run using provided data*<br> 15_SibREML_stratified_Q4.R *can be run using provided data*<br> causal_snps \ *provided causal SNPs*<br> IBD_results \ *provided IBD-probabilities for 1000 simulated sib-pairs*<br> Linkage_VH_results \ <br> pheno \ *provided simulated phenotypes (h2=1) for 8 genetic architectures*<br> Phi_hat_results.txt<br> predicted \<br> README<br> SibREML_results.txt<br> SibREML_stratified_Q4.txt</p> <p>The data can be used to run Linkage analysis, predict linkage, estimate phi_hat, <br>as well as estimate non-stratified and recombination rate stratified sib-heritability (h2_FS and c).<br>The README is provided within the folder. <br>#**********************************************************************************************************</p> <p> </p> <p>#**********************************************************************************************************<br>The "HT_BMI" folder contains data and scripts to predict linkage and estimate phi_hat for height and BMI.<br>HT_BMI \<br> 01_predLink_HT_BMI.R<br> 02_phi_hat_HT_BMI.R<br> gws_sumstats \ *provided summary GWAS summary statistics to predict linkage for height and BMI*<br> Linkage_results \ *provided linkage meta-analysis results for height and BMI from this study*<br> Phi_hat_results_HT_BMI.txt<br> PREDLINK_bmi.txt<br> PREDLINK_height.txt<br> README<br>The README is provided within the folder.<br>#**********************************************************************************************************</p> <p><strong> </strong></p>
Summary statistics from "Genetic Association Study of Eight Steroid Hormones and Implications for Sexual Dimorphism of Coronary Artery Disease"
<p>GWAMA summary statistics of four steroid hormone levels using fixed-effect model and GWAS summary statistics of four other steroid hormones.</p> <p>When using this data, please cite: Pott J, Bae YJ, Horn K, et al.. Genetic Association Study of Eight Steroid Hormones and Implications for Sexual Dimorphism of Coronary Artery Disease. <em>J Clin Endocrinol Metab</em> <strong>2019</strong> Nov 1;104(11):5008-5023. doi: 10.1210/jc.2019-00757</p> <p>All txt files contain the following columns:</p> <ul> <li>markername</li> <li>chr</li> <li>bp_hg19 (base position according to hg19)</li> <li>effect_allele</li> <li>other_allele</li> <li>effect_allele_freq</li> <li>min_info (minimal info score across all used studies)</li> <li>n (sample size per SNP)</li> <li>beta (effect estimate)</li> <li>se (standard error)</li> <li>p (p-value)</li> <li>CochransQ (only in GWAMA; SNP heterogeneity across studies)</li> <li>pCochransQ (only in GWAMA; p-value of Cochrans Q value)</li> </ul>
Summary statistics for association tests between human and Plasmodium falciparum genetic variants in 3,346 severe malaria cases from The Gambia and Kenya
<p>This dataset contains summary statistics for association tests between human and<br> <em>Plasmodium falciparum</em> malaria parasite genetic variants, using data from 3,346 severe malaria cases from The Gambia and Kenya. These results underlie the analysis described in our paper:</p> <p><strong>"Malaria protection due to sickle haemoglobin depends on parasite genotype"</strong></p> <p>Gavin Band, Ellen M. Leffler, Muminatou Jallow, Fatoumatta Sisay-Joof, Carolyne M. Ndila, Alexander W. Macharia, Christina Hubbart, Anna E. Jeffreys, Kate Rowlands, Thuy Nguyen, Sónia M. Gonçalves, Cristina V. Ariani, Jim Stalker, Richard D. Pearson, Roberto Amato, Eleanor Drury, Giorgio Sirugo, Umberto d'Alessandro, Kalifa A. Bojang, Kevin Marsh, Norbert Peshu, Joseph W. Saelens, Mahamadou Diakité, Steve M. Taylor, David J. Conway, Thomas N. Williams, Kirk A. Rockett, Dominic P. Kwiatkowski</p> <p>Nature (2021) doi: <a href="https://doi.org/10.1038/s41586-021-04288-3">10.1038/s41586-021-04288-3</a> <strong>bioRxiv link</strong>:: <a href="https://doi.org/10.1101/2021.03.30.437659">doi.org/10.1101/2021.03.30.437659</a><br> <br> The genotype data underlying these summary statistics has also been deposited on Zenodo<br> (<a href="https://zenodo.org/record/4973477">doi:10.5281/zenodo.4973477</a>). The <a href="https://www.well.ox.ac.uk/~gav/hptest)">HPTEST software</a> used to generate these results has also been deposited (<a href="https://doi.org/10.5281/zenodo.5685580">doi:10.5281/zenodo.5685580</a>). Please see the <a href="https://www.malariagen.net/resource/32">MalariaGEN website</a> for a full list of datasets which have been released with this manuscript.</p> <p><strong>Data contents.</strong></p> <p>The dataset consists of a single <a href="http://sqlite.org">sqlite database file</a> containing the results, and an accompanying README file in markdown and html format. Please see the README file for full details of the data contents.</p> <p> </p>
Deep-sequencing of viral genomes from treatment-naive HIV-infected persons shows positive association between intrahost genetic diversity and viral load
<p><strong><span>Background:</span></strong><span> Infection with human immunodeficiency virus type 1 (HIV) typically results from transmission of a small and genetically uniform viral population. Following transmission, the virus population becomes more diverse because of recombination and acquired mutations through genetic drift and selection. Viral intrahost genetic diversity remains a major obstacle to the cure of HIV; however, there is a disagreement whether intrahost viral genetic diversification associates positively or negatively with disease progression and progression markers. Viral load is a key progression marker and understanding its relationship to viral intrahost genetic diversity could help design future strategies for HIV monitoring and treatment.</span></p> <p><span><strong>Methods:</strong> </span><span>We analyzed deep-sequenced viral genomes from 2,650 treatment-naive HIV-infected persons to measure the intrahost genetic diversity of 2,447 genomic codon positions as calculated by Shannon entropy. We tested for associations between viral load (VL) and amino acid (AA) entropy accounting for sex, age, race, duration of infection, and HIV population structure.</span></p> <p><strong><span>Results:</span></strong><span><strong> </strong>We confirmed that the intrahost genetic diversity is highest in the <em>env</em> gene. Furthermore, we showed that mean Shannon entropy is significantly associated with VL, especially in infections of >24 months duration. We identified 16 significant associations between VL (p-value<2.0x10<sup>-5</sup>) and Shannon entropy at AA positions which in our association analysis explained 13% of the variance in VL.</span></p> <p><strong><span>Conclusions: </span></strong><span>Our results elucidate that viral intrahost genetic diversity is associated with VL and could be used as a better disease progression marker than HIV consensus sequence variants, especially in infections of longer duration. We emphasize that viral intrahost diversity should be considered when studying viral genomes and infection outcomes.</span></p>
Data from: Genome-wide association mapping within a local Arabidopsis thaliana population more fully reveals the genetic architecture for defensive metabolite diversity
<p>A paradoxical finding from genome-wide association studies (GWAS) in plants is that variation in metabolite profiles typically maps to a small number of loci, despite the complexity of underlying biosynthetic pathways. This discrepancy may partially arise from limitations presented by geographically diverse mapping panels. Properties of metabolic pathways that impede GWAS by diluting the additive effect of a causal variant, such as allelic and genic heterogeneity and epistasis, would be expected to increase in severity with the geographic range of the mapping panel. We hypothesized that a population from a single locality would reveal an expanded set of associated loci. We tested this in a French <em>Arabidopsis thaliana</em> population (< 1 km transect) by profiling and conducting GWAS for glucosinolates, a suite of defensive metabolites that have been studied in depth through functional and genetic mapping approaches. For two distinct classes of glucosinolates, we discovered more associations at biosynthetic loci than previous GWAS with continental-scale mapping panels. Candidate genes underlying novel associations were supported by concordance between their observed effects in the TOU-A population and previous functional genetic and biochemical characterization. Local populations complement geographically diverse mapping panels to reveal a more complete genetic architecture for metabolic traits.</p>
FIGURE 4 in One step closer but still far from solving the puzzle - The phylogeny of marine associated mites (Acari, Oribatida, Ameronothroidea) inferred from morphological and molecular genetic data
FIGURE 4 Bayesian inference topology based on 66 morphological traits of 102 oribatid mite species. Posterior probability values are shown near nodes. Photographs of selected species are given to provide an insight into the basic morphology of each larger group. *Photograph shows Tegeocranellus knysnaensis, this species was not used for the analyses but is given here to visualize the typical habitus of Tegeocranellus species.
FIGURE 3 in One step closer but still far from solving the puzzle - The phylogeny of marine associated mites (Acari, Oribatida, Ameronothroidea) inferred from morphological and molecular genetic data
FIGURE 3 One of 14 most parsimonious trees based on 66 characters or character states of 98 ameronothroid and four terrestrial oribatid mite species. Bootstrap values are shown near nodes. Colours refer to different families and are the same as in preceding figures.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.