Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
152
datasets available to search
ShareScore release 0.9.0
Dataset results
152 results for “polygenes”
Expression-based polygenic score from the amygdala 5HTT gene network
<p>This pipeline intends to facilitate the calculation of biologically informed polygenic scores from collected genomic data. This template can be adapted to create other expression-based polygenic risk scores. Data is 1) step by step description and 2) a list of genes that compose the gene network.</p>
An integrated polygenic tool substantially enhances coronary artery disease prediction
<p>Summary-level CAD GWAS data generated by Genomics plc as presented in:</p> <p>Riveros-Mckay F. et al. An integrated polygenic tool substantially enhances coronary artery disease prediction. Circulation: Genomics and Precision Medicine (in press). </p> <p>If you have any questions or comments regarding these files, please contact Genomics plc at research@genomicsplc.com</p> <p> </p> <p>NOTES<br> -----------------------------<br> These analyses were carried out using the full UK Biobank imputation data release (v3b). Analyses were restricted to a subset of UK Biobank, described as “Group I” in the published paper. Group I, “no PCE/QRISK3 available”, included 114,196 European-ancestry individuals with missing data that prevented PCE or QRISK3 calculation.</p> <p>CAD case phenotypes were defined as described in the “Phenotype definitions” section of the paper’s Supplementary Materials, using both prevalent (pre-baseline) and incident (post-baseline) events.</p> <p>All analyses included Age at assessment, sex, genotyping chip, and 10 principal components as covariates. </p> <p>We used plink2.0 logistic regression. For chromosome X variants males were treated as having 0 or 2 alternative alleles. </p> <p>The results are not adjusted for genomic control.</p> <p> </p> <p>DATA FILE CONTENT DESCRIPTION<br> -----------------------------<br> cpra Variant ID in ‘CPRA’ format. Position reflects position in b37. <br> chrom Chromosome<br> pos Position in base pairs (b37, 1-based)<br> alt Alternative allele (effect allele)<br> beta Effect size (log odds ratio)<br> standard_error Standard error of beta <br> minus_log10_p Minus log(base 10) of P-value<br> ref Reference allele (non-effect allele)<br> ncase Number of cases<br> ncontrol Number of controls</p>
UK Biobank release and systematic evaluation of optimised polygenic risk scores for 53 diseases and quantitative traits
<p>Summary-level GWAS data for 53 traits generated by <a href="https://www.genomicsplc.com/">Genomics plc</a> as presented in:</p> <p>Thompson D. et al. UK Biobank release and systematic evaluation of optimised polygenic risk scores for 53 diseases and quantitative traits (<a href="https://doi.org/10.1101/2022.06.16.22276246">https://doi.org/10.1101/2022.06.16.22276246</a>)</p> <p>If you have any questions or comments regarding these files, please contact Genomics plc at <a href="mailto:research@genomicsplc.com">research@genomicsplc.com</a></p> <p><strong>NOTES</strong></p> <p>These analyses were carried out using the full UK Biobank (UKB) imputation data release (v3b). After removal of exclusions and withdrawals, a subset of 337,151 UKB individuals, the White British Unrelated (WBU) subgroup, was defined as the intersection of two sample groups created by Bycroft et al 2018 (Nature 562, 203-209): the ‘White British ancestry’ group (UKB Data Field 22006) and the ‘used in genetic principal components’ group (UKB Data Field 22020), the latter being high quality samples that were filtered to avoid closely related individuals. All GWAS analyses were performed on the WBU subgroup.</p> <p>Phenotypes were defined as described in Supplementary Table 1 ‘Phenotype definitions’ using a combination of Hospital Episode Statistics, Cancer Registry reports (where applicable) and self-report responses, with the exception of coronary artery disease (CAD). GWAS data was generated for both a “narrow” and a “broad” definition of CAD. The former was used as part of the training data for the Enhanced CAD PRS, the latter was used as part of the training data for the Enhanced CVD PRS. The phenotype definitions for “narrow” and a “broad” CAD are as follows:</p> <table> <tbody> <tr> <td>Narrow CAD<br> (includes angina)</td> <td>ICD10 codes (where .X indicates all subcodes) from both hospital and death records: I21, I22, I23, I24.1, I25.2, I20.X. ICD9 codes: 410-412, 42979, 413.X. OPCS-4 codes (K40.1–40.4, K41.1–41.4, K45.1–45.5,K49.1–49.2, K49.8–49.9, K50.2, K75.1–75.4, K75.8–75.9), self-reported heart attack (UKB codes 1075 in field 20002; code 1 in field 6150), self-reported coronary angioplasty (ptca) or coronary artery bypass graft (UKB codes 1070 and 1095 in field 20004), self-reported angina.</td> </tr> <tr> <td>Broad CAD<br> (includes angina and all ischaemic heart disease)</td> <td>As for Narrow CAD, plus ICD10 codes I24.X, I25X, and ICD9 codes 414.X (where .X indicates all subcodes).</td> </tr> </tbody> </table> <p>Note that there is no GWAS for cardiovascular disease (CVD) per se. This is because the UKB training data for the Enhanced CVD PRS consisted of separate GWASs for “narrow” CAD and ischaemic stroke.</p> <p>All analyses included Age at assessment, sex (for non-sex specific traits), genotyping chip, and 10 principal components as covariates.</p> <p>GWAS summary statistics for each trait were generated by applying PLINK 2.0 to the WBU subgroup, using a logistic regression for disease traits, and a linear regression model for quantitative traits. For chromosome X variants males were treated as having 0 or 2 alternative alleles.</p> <p>The results are not adjusted for genomic control.</p> <p><strong>DATA FILE CONTENT DESCRIPTION (DISEASE TRAITS)</strong></p> <table> <tbody> <tr> <td>cpra</td> <td>Variant ID in ‘CPRA’ format. Position reflects position in b37</td> </tr> <tr> <td>chrom</td> <td>Chromosome</td> </tr> <tr> <td>pos</td> <td>Position in base pairs (b37, 1-based)</td> </tr> <tr> <td>alt</td> <td>Alternative allele (effect allele)</td> </tr> <tr> <td>beta</td> <td>Effect size (log odds ratio)</td> </tr> <tr> <td>standard_error</td> <td>Standard error of beta</td> </tr> <tr> <td>minus_log10_p</td> <td>Minus log(base 10) of P-value</td> </tr> <tr> <td>ref</td> <td>Reference allele (non-effect allele)</td> </tr> <tr> <td>ncase</td> <td>Number of cases</td> </tr> <tr> <td>ncontrol</td> <td>Number of controls</td> </tr> </tbody> </table> <p><strong>DATA FILE CONTENT DESCRIPTION (QUANTITATIVE TRAITS)</strong></p> <table> <tbody> <tr> <td>cpra</td> <td>Variant ID in ‘CPRA’ format. Position reflects position in b37</td> </tr> <tr> <td>chrom</td> <td>Chromosome</td> </tr> <tr> <td>pos</td> <td>Position in base pairs (b37, 1-based)</td> </tr> <tr> <td>alt</td> <td>Alternative allele (effect allele)</td> </tr> <tr> <td>beta</td> <td>Effect size</td> </tr> <tr> <td>standard_error</td> <td>Standard error of beta</td> </tr> <tr> <td>minus_log10_p</td> <td>Minus log(base 10) of P-value</td> </tr> <tr> <td>ref</td> <td>Reference allele (non-effect allele)</td> </tr> <tr> <td>ntotal</td> <td>Total sample size</td> </tr> </tbody> </table> <p><strong>FILE NAMES</strong></p> <p>The following is a list of traits and their corresponding file names.</p> <p><em><strong>DISEASE TRAITS</strong></em></p> <table> <tbody> <tr> <td>Age-related macular degeneration</td> <td>amd_strict_UKB_WBU.csv.gz</td> </tr> <tr> <td>Alzheimer's disease</td> <td>alzheimers_disease_UKB_WBU.csv.gz</td> </tr> <tr> <td>Asthma</td> <td>asthma_UKB_WBU.csv.gz</td> </tr> <tr> <td>Atrial fibrillation</td> <td>atrial_fibrillation_UKB_WBU.csv.gz</td> </tr> <tr> <td>Bipolar disorder</td> <td>bipolar_disorder_UKB_WBU.csv.gz</td> </tr> <tr> <td>Bowel cancer</td> <td>CRC_UKB_WBU.csv.gz</td> </tr> <tr> <td>Breast cancer</td> <td>BC_UKB_WBU_women.csv.gz</td> </tr> <tr> <td>Coeliac disease</td> <td>celiac_disease_UKB_WBU.csv.gz</td> </tr> <tr> <td>Narrow coronary artery disease</td> <td>NARROW_CAD_UKB_WBU.csv.gz</td> </tr> <tr> <td>Broad coronary artery disease</td> <td>BROAD_CAD_UKB_WBU.csv.gz</td> </tr> <tr> <td>Crohn's disease</td> <td>crohns_disease_UKB_WBU.csv.gz</td> </tr> <tr> <td>Epithelial ovarian cancer</td> <td>OC_UKB_WBU.csv.gz</td> </tr> <tr> <td>Hypertension</td> <td>HT_UKB_WBU.csv.gz</td> </tr> <tr> <td>Ischaemic stroke</td> <td>IS_stroke_UKB_WBU.csv.gz</td> </tr> <tr> <td>Melanoma</td> <td>melanoma_UKB_WBU.csv.gz</td> </tr> <tr> <td>Multiple sclerosis</td> <td>multiple_sclerosis_UKB_WBU.csv.gz</td> </tr> <tr> <td>Osteoporosis</td> <td>OP_WBU_training.csv.gz</td> </tr> <tr> <td>Prostate cancer</td> <td>PC_UKB_WBU.csv.gz</td> </tr> <tr> <td>Parkinson's disease</td> <td>parkinsons_disease_UKB_WBU.csv.gz</td> </tr> <tr> <td>Primary open angle glaucoma</td> <td>POAG_WBU_training.csv.gz</td> </tr> <tr> <td>Psoriasis</td> <td>psoriasis_UKB_WBU.csv.gz</td> </tr> <tr> <td>Rheumatoid arthritis</td> <td>rheumatoid_arthritis_UKB_WBU.csv.gz</td> </tr> <tr> <td>Schizophrenia</td> <td>schizophrenia_UKB_WBU.csv.gz</td> </tr> <tr> <td>Systemic lupus erythematosus</td> <td>lupus_UKB_WBU.csv.gz</td> </tr> <tr> <td>Type 1 diabetes</td> <td>t1d_UKB_WBU.csv.gz</td> </tr> <tr> <td>Type 2 diabetes</td> <td>T2D_UKB_WBU.csv.gz</td> </tr> <tr> <td>Ulcerative colitis</td> <td>ulcerative_colitis_UKB_WBU.csv.gz</td> </tr> <tr> <td>Venous thromboembolic disease</td> <td>VTE_UKB_WBU.csv.gz</td> </tr> </tbody> </table> <p><em><strong>QUANTITATIVE TRAITS</strong></em></p> <table> <tbody> <tr> <td>Age at menopause</td> <td>age_at_menopause_UKB_WBU.csv.gz</td> </tr> <tr> <td>Apolipoprotein A1</td> <td>apolipoprotein_a1_UKB_WBU.csv.gz</td> </tr> <tr> <td>Apolipoprotein B</td> <td>apolipoprotein_b_UKB_WBU.csv.gz</td> </tr> <tr> <td>Body mass index</td> <td>bmi_UKB_WBU.csv.gz</td> </tr> <tr> <td>Calcium</td> <td>calcium_UKB_WBU.csv.gz</td> </tr> <tr> <td>Docosahexaenoic acid</td> <td>docosahexaenoic_acid_UKB_WBU.csv.gz</td> </tr> <tr> <td>Estimated bone mineral density T-score</td> <td>BMD_WBU_training.csv.gz</td> </tr> <tr> <td>Estimated glomerular filtration rate (creatinine based)</td> <td>egfr_UKB_WBU.csv.gz</td> </tr> <tr> <td>Estimated glomerular filtration rate (cystatin based)</td> <td>egfr_cys_UKB_WBU.csv.gz</td> </tr> <tr> <td>Glycated haemoglobin</td> <td>hba1c_UKB_WBU_nodiabetes.csv.gz</td> </tr> <tr> <td>High density lipoprotein cholesterol</td> <td>hdl_cholesterol_UKB_WBU.csv.gz</td> </tr> <tr> <td>Height</td> <td>height_UKB_WBU.csv.gz</td> </tr> <tr> <td>Intraocular pressure</td> <td>iop_WBU_training.csv.gz</td> </tr> <tr> <td>Low density lipoprotein cholesterol</td> <td>ldl_UKB_WBU_nostatins.csv.gz</td> </tr> <tr> <td>Omega-6 fatty acids</td> <td>omega_6_fatty_acids_UKB_WBU.csv.gz</td> </tr> <tr> <td>Omega-3 fatty acids</td> <td>omega_3_fatty_acids_UKB_WBU.csv.gz</td> </tr> <tr> <td>Phosphatidylcholines</td> <td>phosphatidylcholines_UKB_WBU.csv.gz</td> </tr> <tr> <td>Phosphoglycerides</td> <td>phosphoglycerides_UKB_WBU.csv.gz</td> </tr> <tr> <td>Polyunsaturated fatty acids</td> <td>polyunsaturated_fatty_acids_UKB_WBU.csv.gz</td> </tr> <tr> <td>Resting heart rate</td> <td>resting_heart_rate_UKB_WBU.csv.gz</td> </tr> <tr> <td>Remnant cholesterol (Non-HDL, Non-LDL cholesterol)</td> <td>remnant_cholesterol__UKB_WBU.csv.gz</td> </tr> <tr> <td>Sphingomyelins</td> <td>sphingomyelins_UKB_WBU.csv.gz</td> </tr> <tr> <td>Total cholesterol</td> <td>total_cholesterol_UKB_WBU.csv.gz</td> </tr> <tr> <td>Total fatty acids</td> <td>total_fatty_acids_UKB_WBU.csv.gz</td> </tr> <tr> <td>Total triglycerides</td> <td>total_triglycerides_UKB_WBU.csv.gz</td> </tr> </tbody> </table>
Data From: Powerful detection of polygenic selection and environmental adaptation in US beef cattle
<p>GEMMA output containing summary statistics for generation proxy selection mapping (GPSM) and environmental GWAS (envGWAS) selection analyses from <br> Rowan et al. "Powerful detection of polygenic selection and environmental adaptation in US beef cattle" 2021<br> https://doi.org/10.1101/2020.03.11.988121 </p> <p>File names identify the analysis run, for example<br> "Gelbvieh_envgwas_desert_summary_stats.txt.gz"<br> Is the Gelbvieh dataset analyzed using the Desert ecoregion as the dependent variable <br> in a univariate envGWAS model. </p> <p>Files are formated according to GEMMA output.</p>
Datasets for polygenic mechanisms of hybrid incompatibility in butterflies
<p><strong>Version 1.2 includes data that are missing in the previous versions.</strong></p> <p> </p> <p>Note: relevant scripts can also be found at</p> <p>https://github.com/tzxiong/2022_Papilio_HybridIncompatibilityMapping</p> <p>======================================================<br>Description of source data and scripts for all figures<br>======================================================</p> <p>==== MAIN FIGURES ====</p> <p>Fig. 1</p> <p> - Panel A<br> * Schematic figure, no source data are provided</p> <p> - Panel B<br> * Schematic figure, no source data are provided</p> <p> - Panel C<br> * Source data folder(s):<br> SourceData/Fig1/Fig1C<br> * Source code:<br> SourceData/Code.JupyterLab/Code02_Mapping.ipynb: Section 5</p> <p> - Panel D<br> * Source data folder(s):<br> SourceData/Fig1/Fig1D<br> * Source code:<br> SourceData/Code.JupyterLab/Code02_Mapping.ipynb: Section 5</p> <p> - Panel E<br> * Source data folder(s):<br> SourceData/Fig1/Fig1E<br> * Source code:<br> SourceData/Code.JupyterLab/Code01_SequencingData.ipynb: Section 2</p> <p> - Panel F<br> * Source data folder(s):<br> SourceData/Fig1/Fig1F<br> * Source code:<br> SourceData/Code.JupyterLab/Code01_SequencingData.ipynb: Section 2</p> <p>Fig. 2</p> <p> - Panels A-G<br> * Source data folder(s): <br> SourceData/Fig2+S1toS2 <br> * The "Raw" folder contains unedited images.<br> * Two edited images used in Fig2 is also included for each subfigure.</p> <p> - Panels H-L<br> * Source data folder(s): <br> SourceData/Fig2+S1toS2 <br> * The "Raw" folder (unzipped) contains unedited confocal data in .czi format.<br> * Edited images are included with both monochrome and merged versions.</p> <p>Fig. 3</p> <p> - Panel A<br> * Source data folder(s):<br> SourceData/Fig3+S8toS10/Fig3A_3B_3C_3D_3E<br> * Source code:<br> D(DB): SourceData/Code.JupyterLab/Code02_Mapping.ipynb: Section 3.1<br> B(BD): SourceData/Code.JupyterLab/Code02_Mapping.ipynb: Section 2.3</p> <p> - Panel B<br> * Source data folder(s): <br> SourceData/Fig3+S8toS10/Fig3A_3B_3C_3D_3E<br> * Source code:<br> D(DB): SourceData/Code.JupyterLab/Code02_Mapping.ipynb: Section 4.1<br> B(BD): SourceData/Code.JupyterLab/Code02_Mapping.ipynb: Section 4.2</p> <p> - Panel C<br> * Source data folder(s): <br> SourceData/Fig3+S8toS10/Fig3A_3B_3C_3D_3E<br> * Source code:<br> D(DB): SourceData/Code.JupyterLab/Code02_Mapping.ipynb: Section 4.1<br> B(BD): SourceData/Code.JupyterLab/Code02_Mapping.ipynb: Section 4.2</p> <p> - Panel D<br> * Source data folder(s): <br> SourceData/Fig3+S8toS10/Fig3A_3B_3C_3D_3E<br> * Source code:<br> SourceData/Code.JupyterLab/Code02_Mapping.ipynb: Run all of Sections 4.1 and 4.2</p> <p> - Panel E<br> * Source data folder(s): <br> SourceData/Fig3+S8toS10/Fig3A_3B_3C_3D_3E<br> * Source code:<br> SourceData/Code.JupyterLab/Code02_Mapping.ipynb: Run all of Sections 4.1 and 4.2</p> <p>Fig. 4</p> <p> - Note 1: For Heliconius analysis, all data are from SourceData/Fig4-Heliconius+S11C/dat.4.qtl.lumped.csv. This file contains Heliconius ovary dysgenesis data from https://doi.org/10.1111/mec.16272</p> <p> - Panel A (Heliconius)<br> * Source data folder(s): <br> SourceData/Fig4-Heliconius+S11C<br> * Source code:<br> SourceData/Code.JupyterLab/Code02_Mapping.ipynb: Section 6.1<br> <br> - Panel A (Papilio) <br> * Source data folder(s): <br> SourceData/Fig4-Papilio+S13<br> * Source code:<br> SourceData/Code.JupyterLab/Code02_Mapping.ipynb: Section 3.2</p> <p> - Panel B (Heliconius)<br> * Source data folder(s): <br> SourceData/Fig4-Heliconius+S11C<br> * Source code:<br> SourceData/Code.JupyterLab/Code02_Mapping.ipynb: Section 6.1<br> <br> - Panel B (Papilio)<br> * Source data folder(s): <br> SourceData/Fig4-Papilio+S13<br> * Source code:<br> SourceData/Code.JupyterLab/Code02_Mapping.ipynb: Section 3.2</p> <p> - Panel C (Heliconius) <br> * Source data folder(s): <br> SourceData/Fig4-Heliconius+S11C<br> * Source code:<br> SourceData/Code.JupyterLab/Code02_Mapping.ipynb: Section 6.2<br> <br> - Panel C (Papilio) <br> * Source data folder(s):<br> SourceData/Fig4-Papilio+S13<br> * Source code:<br> SourceData/Code.JupyterLab/Code02_Mapping.ipynb: Section 3.3</p> <p> - Panel D<br> * Schematic figure, no source data are provided<br> <br> - Panel E (Heliconius) <br> * Source data folder(s): <br> SourceData/Fig4-Heliconius+S11C<br> * Source code:<br> SourceData/Code.JupyterLab/Code02_Mapping.ipynb: Section 6.3<br> <br> - Panel E (Papilio) <br> * Source data folder(s): <br> SourceData/Fig4-Papilio+S13<br> * Source code:<br> SourceData/Code.JupyterLab/Code02_Mapping.ipynb: Section 3.3<br> <br> - Panel F (Heliconius)<br> * Source data folder(s): <br> SourceData/Fig4-Heliconius+S11C<br> * Source code:<br> SourceData/Code.JupyterLab/Code02_Mapping.ipynb: Section 6.3<br> <br> - Panel F (Papilio) <br> * Source data folder(s): <br> SourceData/Fig4-Papilio+S13<br> * Source code:<br> SourceData/Code.JupyterLab/Code02_Mapping.ipynb: Section 3.3<br> <br> <br> <br>==== SUPPLEMENTARY FIGURES ====</p> <p>Fig. S1-S2</p> <p> * Source data folder(s): <br> SourceData/Fig2+S1toS2 <br> * The "Raw" folder (unzipped) contains unedited confocal data in .czi format.<br> * Edited images are included with both monochrome and merged versions.</p> <p>Fig. S3</p> <p> * Source data folder(s): <br> SourceData/FigS3<br> * Source code:<br> SourceData/Code.JupyterLab/Code01_SequencingData.ipynb: Section 1<br> * Note: Source data file 04.0_IBD.NgsRelate.zip contains results from the NGSRelate software.<br> </p> <p>Fig. S4</p> <p> * Source data folder(s): <br> SourceData/FigS4<br> * Source code:<br> SourceData/Code.JupyterLab/Code01_SequencingData.ipynb: Section 2<br> * Note 1: Source data file CorrectedReferenceGenome.zip is the corrected reference genome used for all analyses. It is in .fasta format.<br> * Note 2: Source data file DenovoMarkerOrder_on_CorrectedRefGenome.zip contains all outputs from the LepMap3/OrderMarkers2 module that uses genotype likelihoods and pedigree information to generate a new marker order. Use script "OrderMarkers2_ReOrder.sh" from the script repo.<br> </p> <p>Fig. S5-S7</p> <p> * Source data folder(s): <br> SourceData/FigS5toS7<br> * Source code:<br> SourceData/Code.JupyterLab/Code01_SequencingData.ipynb: Section 2<br> * Note: The source data file PedigreeAncestryInGrandparentalPhase.zip contains all outputs from the LepMap3/OrderMarkers2 module that uses genotype likelihoods and pedigree information to impute ancestry at each marker. Ancestry is phased according to the sex of grandparents. Use script "OrderMarkers2.sh" from the GitHub repo.</p> <p>Fig. S8</p> <p> - Panel A <br> * Source data folder(s): <br> SourceData/Fig3+S8toS10<br> * Source code:<br> SourceData/Code.JupyterLab/Code02_Mapping.ipynb: Section 3.1</p> <p> - Panel B <br> * Source data folder(s): <br> SourceData/Fig3+S8toS10<br> * Source code:<br> SourceData/Code.JupyterLab/Code02_Mapping.ipynb: Section 2.3</p> <p> - Panel C<br> * Source data folder(s): <br> SourceData/Fig3+S8toS10<br> * Source code:<br> See previous two panels</p> <p>Fig. S9</p> <p> - Panel A<br> * Source data folder(s): <br> SourceData/Fig3+S8toS10<br> * Source code:<br> SourceData/Code.JupyterLab/Code02_Mapping.ipynb: Section 4.1</p> <p> - Panel B<br> * Source data folder(s): <br> SourceData/Fig3+S8toS10<br> * Source code:<br> SourceData/Code.JupyterLab/Code02_Mapping.ipynb: Section 4.2</p> <p>Fig. S10</p> <p> - Panel A<br> * Source data folder(s):<br> SourceData/Fig3+S8toS10<br> * Source code:<br> SourceData/Code.JupyterLab/Code02_Mapping.ipynb: Section 4.1</p> <p> - Panel B<br> * Source data folder(s): <br> SourceData/Fig3+S8toS10<br> * Source code:<br> SourceData/Code.JupyterLab/Code02_Mapping.ipynb: Section 4.2</p> <p>Fig. S11</p> <p> - Panel A<br> * Source data folder(s): <br> SourceData/FigS11AB<br> * Source code:<br> Data in SpeciesAncestry.B0D1.zip can be directly visualized to get the figure</p> <p> - Panel B<br> * Source data folder(s): <br> SourceData/FigS11AB<br> * Source code:<br> Data in SpeciesAncestry.B0D1.zip can be directly visualized to get the figure<br> * Note: This file contains ancestry at each marker phased according to species (bianor=0, dehaanii=1). These data are directly transformed from files in SourceData/FigS5toS7/PedigreeAncestryInGrandparentalPhase.zip.</p> <p> - Panel C<br> * Source data folder(s): <br> SourceData/Fig4-Heliconius+S11C<br> * Source code:<br> SourceData/Code.JupyterLab/Code02_Mapping.ipynb: Section 6.4</p> <p>Fig. S12</p> <p> * Source data folder(s): <br> No source data are needed<br> * Source code:<br> SourceData/Code03_PolygenicGhostQTL.ipynb</p> <p>Fig. S13</p> <p> - Panel A<br> * Source data folder(s): <br> SourceData/Fig4-Papilio+S13<br> * Source code:<br> SourceData/Code.JupyterLab/Code02_Mapping.ipynb: Section 3.3</p> <p> - Panel B<br> * Source data folder(s): <br> SourceData/Fig4-Papilio+S13<br> * Source code:<br> SourceData/Code.JupyterLab/Code02_Mapping.ipynb: Section 3.2</p> <p>Fig. S14</p> <p> * Source data folder(s): <br> No source data are needed<br> * Source code:<br> SourceData/Code03_PolygenicGhostQTL.ipynb</p> <p>Fig. S15</p> <p> - Panel A<br> * Source data folder(s): <br> SourceData/FigS15/FigS15A<br> * Source code:<br> SourceData/Code.JupyterLab/Code02_Mapping.ipynb: Section 2.1</p> <p> - Panel B<br> * Source data folder(s): <br> SourceData/FigS15/FigS15B<br> * Source code:<br> SourceData/Code.JupyterLab/Code02_Mapping.ipynb: Section 2.2</p> <p>Fig. S16</p> <p> * Source data folder(s): <br> SourceData/FigS16<br> * Source code:<br> SourceData/Code.JupyterLab/Code01_SequencingData.ipynb: Section 3<br> </p> <p>==== OTHER SOURCE DATA & SUMMARY OF SOURCE CODE FOLDERS====</p> <p>LepMap3-SourceData</p> <p> * LepMap3_SourceData-Family_Info_Finalized_withPseudoGrandParents_transposed.txt<br> <br> This file is the pedigree file ready-to-use in LepMap3. Note that it contains pseudo grandparents for families missing grandparents in sequencing. Pseudo grandparents are simply created from fixed SNPs in all existing grandparents and adding them to the original vcf files containing genotype likelihoods.</p> <p> * LepMap3_SourceData-vcf_files.zip<br> <br> The vcf files containing genotype likelihoods for LepMap3 to use. Note that it contains the aforementioned pseudo grandparents.</p> <p><br>Code.NGSRelate</p> <p> * Code used for inferring kinship from low-coverage sequencing data</p> <p>Code.LepMap3<br> <br> * Code used for all LepMap3 analysis</p> <p>Code.JupyterLab</p> <p> * Code used for all Julia and R analysis in .ipynb format</p> <p><br> </p>
Common Ancestry of the Id Locus: Chromosomal Rearrangement and Polygenic Possibilities
<h2>The Id locus, with a potential polygenic nature, is inverted alongside ZARU1 at the distal end of the q-arm of chromosome Z, indicating a shared ancestry among BBC breeds.</h2>
Figure 2 in Adaptations in wild radish (Raphanus raphanistrum) flowering time, Part 1: Individual-based modeling of a polygenic trait
Figure 2. The relationship between the number of semidominant larger M1 alleles and days to first flower (DFF), needed to model the long-day selections. Note that six smallereffect M2 genes will add another 0 to 18 days to DFF. If each M1 allele added the same amount in the long-day selections, the relationship would be linear.
Figure 1 in Adaptations in wild radish (Raphanus raphanistrum) flowering time, Part 1: Individual-based modeling of a polygenic trait
Figure 1. Goodness of fit of the modeled data (A) when compared with the recorded glasshouse data (B). Graphs show cumulative days to first flowering (DFF) adaptations in Raphanus raphanistrum populations as a result of repeated early and late days to flowering selection. In both graphs, the basal population is the black solid line in the center,the darker lines are the matched generations of early flowering (EF1, EF3, FE4,and EF5; colored orange), and late flowering (LF2 and LF3; colored purple), with and late flowering (LF2 and LF3),with the two unmatched generations (EF2 and LF2) shown in lighter tones.The far-left population (early flowering EF5) displays very little phenotypic variability, whereas the far-right population (late flowering 3) is very diverse.
Polygenic risk scores validated in patient-derived cells stratify for mitochondrial subtypes of Parkinson's disease
<p><strong>Background</strong> Parkinson’s disease (PD) is the fastest growing neurodegenerative disorder, with affected individuals expected to double during the next 20 years. This raises the urgent need to better understand the genetic architecture and downstream cellular alterations underlying PD pathogenesis, in order to identify more focused therapeutic targets. While only ∼10% of PD cases can be clearly attributed to monogenic causes, there is mounting evidence that additional genetic factors could play a role in idiopathic PD (iPD). In particular, common variants with low to moderate effect size in multiple genes regulating key neuroprotective activities may act as risk factors for PD. In light of the well-established involvement of mitochondrial dysfunction in PD, we hypothesized that a fraction of iPD cases may harbour a pathogenic combination of common variants in nuclear-encoded mitochondrial genes, ultimately resulting in neurodegeneration.</p> <p><strong>Methods</strong> to capture this mitochondria-related “missing heritability”, we leveraged on existing data from previous genome-wide association studies (GWAS) – i.e., the large PD GWAS from Nalls and colleagues. We then used computational approaches based on mitochondria-specific polygenic risk scores (mitoPRSs) for imputing the genotype data obtained from different iPD case-control datasets worldwide, including the Luxembourg Parkinson’s Study (412 iPD patients and 576 healthy controls) and the COURAGE-PD cohorts (7270 iPD cases and 6819 healthy controls).</p> <p><strong>Results</strong> applying this approach to gene sets controlling mitochondrial pathways potentially relevant for neurodegeneration in PD, we demonstrated that common variants in genes regulating <em>Oxidative Phosphorylation (OXPHOS</em>-PRS<em>)</em> were significantly associated with a higher PD risk both in the Luxembourg Parkinson’s Study (odds ratio, OR=1.31[1.14-1.50], <em>p</em>=5.4e-04) and in COURAGE-PD (OR=1.23[1.18-1.27], <em>p</em>=1.5e-29). Functional analyses in primary skin fibroblasts and in the corresponding induced pluripotent stem cells-derived neuronal progenitor cells from Luxembourg Parkinson’s Study iPD patients stratified according to the <em>OXPHOS</em>-PRS, revealed significant differences in mitochondrial respiration between high and low risk groups (<em>p</em> < 0.05). Finally, we also demonstrated that iPD patients with high <em>OXPHOS</em>-PRS have a significantly earlier age at disease onset compared to low-risk patients.</p> <p><strong>Conclusions</strong> our findings suggest that OXPHOS-PRS may represent a promising strategy to stratify iPD patients into pathogenic subgroups – in which the underlying neurodegeneration is due to a genetically defined mitochondrial burden – potentially eligible for future, more tailored mitochondrially targeted treatments.</p>
Data From: Polygenic basis and the role of genome duplication in adaptation to similar selective environments
Open the record for dataset details and reuse information.
Data from: Whole-genome resequencing reveals polygenic signatures of directional and balancing selection on alternative migratory life histories
Open the record for dataset details and reuse information.
Data from: The polygenic strategies of host-specific and general virulence of Botrytis cinerea across diverse eudicot hosts
Open the record for dataset details and reuse information.
Data from: The genetic architecture of recombination rates is polygenic and differs between the sexes in wild house sparrows (Passer domesticus)
Open the record for dataset details and reuse information.
Data and Code: No support for the genetic hypothesis of the Black-white achievement gap using polygenic scores and tests for divergent selection
<p>Data and Code for article "No support for the genetic hypothesis of the Black-white achievement gap using polygenic scores and tests for divergent selection"</p>
Data from: Population genomics of rapid evolution in natural populations: polygenic selection in response to power station thermal effluents
Background: Examples of rapid evolution are common in nature but difficult to account for with the standard population genetic model of adaptation. Instead, selection from the standing genetic variation permits rapid adaptation via soft sweeps or polygenic adaptation. Empirical evidence of this process in nature is currently limited but accumulating. Results: We provide genome-wide analyses of rapid evolution in two Fundulus heteroclitus populations subjected to recently elevated temperatures due to coastal power station thermal effluents. Bayesian and multivariate analyses of population genomic structure reveal a substantial portion of genetic variation that is most parsimoniously explained by selection at the site of thermal effluents. An FST outlier approach in conjunction with additional conservative requirements identify significant allele frequency differentiation that exceeds neutral expectations among exposed and closely related reference populations. Genomic variation patterns near these candidate loci reveal that individuals living near thermal effluents have rapidly evolved from the standing genetic variation through small allele frequency changes at many loci in a pattern consistent with polygenic selection on the standing genetic variation. Conclusions: While the ultimate trajectory of selection in these populations is unknown, our findings suggest that polygenic models of adaptation may play important roles in large, natural populations experiencing recent selection due to environmental changes that cause broad physiological impacts.
Genetic insight into a polygenic trait using a novel Genome Wide Association approach in a wild amphibian population
<p>Body size variation is central in the evolution of life history traits in amphibians, but the underlying genetic architecture of this complex trait is still largely unknown. Herein, we studied the genetic basis of body size and fecundity of the alternative morphotypes in a wild population of the Greek smooth newt (<em>Lissotriton graecus</em>). By combining a Genome-wide association approach with linkage disequilibrium network analysis, we were able to identify clusters of highly correlated loci thus maximizing sequence data for downstream analysis. The putatively associated variants explained 12.8% to 44.5% of the total phenotypic variation in body size and were mapped to genes with functional roles in the regulation of gene expression and cell cycle processes. Our study is the first to provide insights into the genetic basis of complex traits in newts and provides a useful tool to identify loci potentially involved in fitness related traits in small data sets from natural populations in non-model species.</p>
Dataset for Rapid polygenic adaptation in a wild population of ash trees under a novel fungal epidemic
<p><strong>Dataset for Rapid polygenic adaptation in a wild population of ash trees under a novel fungal epidemic</strong></p> <p>Code used for plotting of main figures and quantifying allelic shifts attached in the GitHub repository CareyMetheringham/MardenPark. Additional files for analysis of GEBV shifts between adults and juveniles, estimation of heritability, trends in green up and simulations of allelic shifts included as:</p> <ul> <li>GEBV-regression.Rmd</li> <li>heritabilityEst.R</li> <li>GreeningAnalysis.R</li> <li>Simulated_selection_analysis.Rmd</li> <li>Distinguishing the effects of selection from genetic drift.pdf</li> </ul> <p>Data files:</p> <ul> <li>S1 - phenotypic measurements for adult and juvenile trees</li> <li>S2 - Effect sizes for SNPs used to calculate GEBV</li> <li>S3 - Allelic frequencies of sites used to calculate GEBV</li> <li>related_trees.csv - Predicted parentage of trees</li> <li>gebv_model_df.csv - Data used for greenup calculations in GreeningAnalysis.R</li> <li>maf1.pass2.miss25.snps.only.LD.vcf - Filtered file of high MAF SNPs used for parentage estimation</li> <li>unlinked_sites.csv - unlinked sites used in GEBV-regression.Rmd</li> <li>ebv_table_10000_250 - table of estimated breeding values in field trial populatio, used for heritability estimation</li> <li>MP_eefects_MIA_and_MAA.csv - Data for plotting Figure 3 - Estimated effect size of the major and minor allele, plus standard error on the estimate</li> </ul>
A polygenic architecture with habitat-dependent effects underlies ecological differentiation in Silene
<p><span>Ecological differentiation can drive speciation but it is unclear how the genetic architecture of habitat-dependent fitness contributes to lineage divergence. We investigated the genetic architecture of cumulative flowering, a fitness component, in second-generation hybrids between <em>Silene dioica</em> and <em>S. latifolia</em> transplanted into the natural habitat of each species.</span></p> <p><span>We used reduced-representation sequencing and Bayesian Sparse Linear Mixed Models (BSLMMs) to analyze the genetic control of cumulative flowering in each habitat.</span></p> <p><span>Our results point to a polygenic architecture of cumulative flowering. Allelic effects were mostly beneficial or deleterious in one habitat and neutral in the other. Positive-effect alleles were often derived from the native species, whereas negative-effect alleles, at other loci, tended to originate from the non-native species.</span></p> <p><span>We conclude that ecological differentiation is governed and maintained by many loci with small, habitat-dependent effects consistent with conditional neutrality. This pattern may result from differences in selection targets in the two habitats and from environmentally-dependent deleterious load. Our results further suggest that selection for native alleles and against non-native alleles acts as a barrier to gene flow between species.</span></p>
Data from: Signature of altered retinal microstructures and electrophysiology in schizophrenia spectrum disorders is associated with disease severity and polygenic risk
<p>This dataset contains supporting data for the publication: </p> <p>Boudriot, E.<em> et al.</em> Signature of altered retinal microstructures and electrophysiology in schizophrenia spectrum disorders is associated with disease severity and polygenic risk. <em>Biological Psychiatry</em><span> </span><a href="https://doi.org/10.1016/j.biopsych.2024.04.014">https://doi.org/10.1016/j.biopsych.2024.04.014</a></p> <p> </p> <p>Files:</p> <ul> <li><em>clinical.csv </em>contains data from clinical assessment and polygenic risk scores for schizophrenia</li> <li><em>ophthalmic_examination.csv </em>contains data on spherical equivalent, intraocular pressure and visual acuity</li> <li><em>oct.csv </em>contains segmentation output from Iowa Reference Algorithms</li> <li><em>erg.csv </em>contains pre-processed ERG data for the four ERG conditions</li> <li><em>mri.csv </em>contains ICV-corrected MRI volumes</li> </ul>
simulated datasets for evaluating polygenic detection methods
<p>This dataset contains simulation files corresponding to a combination of each demographic model (1/2/3), environment (linear/quadratic), selection duration (200/400/600/800/1000), and simulation replicate(1-20). This resulted in 600 simulation files with 600 unique combinations of demographic models, environments, selection durations, and simulation replicates. For each individual in the genotype data file, we have the files containing the values of selective pressure(linear and quadratic environment) in the metadata folder. </p> <p>The variant position are 1-based which is default SLiM output. To compare the results with the causal loci user must make the positions 0-based (i.e. POS-1). The details are provided in a github tutorial.</p> <p>Please refer to the documentation for a detailed description of the files and folder structure.</p> <p>The article describing the simulated data and its application is accepted for publication in Nucleic Acids Research (https://doi.org/10.1093/nar/gkae1027).</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.