Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
57
datasets available to search
ShareScore release 0.9.0
Dataset results
57 results for “Linkage disequilibrium”
An updated map of GRCh38 linkage disequilibrium blocks based on European ancestry data
<p>A map of approximately independent linkage disequilibrium (LD) blocks has many uses in statistical genetics. Current publicly available LD block maps are based on sparse recombination maps and are only available for GRCh37 (hg19) and prior genome assemblies. We generated LD blocks in GRCh38 for European (EUR) ancestry populations using a recent recombination map based on more than 115,000 individuals. This new map consists of 1,361 independent LD blocks across the 22 autosomal chromosomes and can be accessed at https://github.com/jmacdon/LDblocks_GRCh38</p>
Partitioned linkage disequilibrium scores for active regulatory elements in ROADMAP datasets
<p>Partitioned linkage disequilibrium scores for active regulatory elements in ROADMAP epigenomics datasets, to accompany paper Lynall et al 2021</p> <p>Accompanying code available at https://github.com/maryellenlynall/psychimmgen2021</p> <p>Active regulatory elements annotations are a union of the following IDEAS annotations, representing enhancers and active promoters (see http://bx.psu.edu/~yuzhang/Roadmap_ideas/trackDb_test.txt for IDEAS track hubs): </p> <p>4_Enh<br> 6_EnhG<br> 8_TssAFlnk<br> 10_TssA<br> 14_TssWk<br> 17_EnhGA</p> <p>tissues.txt provides the list of ROADMAP tissues </p> <p>The partitioned_LD_scores folder contains partitioned LD scores in a format suitable for stratified LDSC analysis for European participants</p>
Genetic diversity, population structure, and linkage disequilibrium among tropical quality protein maize (QPM) lines assessed with high-density SNP markers
<p>The study of genetic diversity (GD), population structure, and linkage disequilibrium (LD) provides a better understanding of the genetic relationships between individuals in a population which can be utilized in crop research and improvement. Genotyping-by-sequencing (GBS) was used to detect and genotype single nucleotide polymorphisms (SNPs) in a collection of 74 quality protein maize (QPM) lines and further to characterize their genetic diversity, population structure, and linkage disequilibrium. A total of 235,214 high-quality SNPs were used for different genetic analyses except for structure analysis where 11,950 SNPs were used. Analysis of molecular variance (AMOVA) based on these SNPs revealed high genetic heterozygosity among the five populations with 1% of the total genetic variation present among the subpopulations and 99% of the variation among individuals within the populations. Population structure analysis using Bayesian-based clustering revealed that the 74 lines could be clustered into four groups. However, neighbor-joining trees indicate the lines are grouped into three major clusters. Further analysis using principal component analyses (PCA) clustered the genotypes into five groups which are concordant with the groups based on pedigree information. Higher genetic diversity was detected in population 1 with a GD value of 0.484 and the lowest in population 5 (0.396) and overall, with a mean of 0.434. The LD pattern in the quality protein maize was investigated and we observed a relatively rapid LD decay of 3.53kb and 10.66kb at r<sup>2</sup> =0.2 and r<sup>2</sup>= 0.1, respectively. Our findings provide important information for future Linkage mapping studies, genome-wide association analyses, and marker-assisted selective breeding of maize as well as genomic prediction-based selection in tropical germplasm.</p>
Effect of variable thresholds on calculating linkage disequilibrium and population structure, using Plink 1.9
<p>The figure presented here shows how the number of LD-independent SNPs and the apparent population structure can change drastically, depending on what input thresholds are used for the calculations. The population sample consists of 220 fruit-fly (Drosophila melanogaster). Most of the population structure plots indicate four subpopulations, which on further investigation using Fst indicate that this is caused by defined trans-centromeric regions, without evidence for genotyping error, and probably reflective of historic admixture. In the plots of population structure, the number in each box indicates the number of independent SNPs which were used in the IBD calculatations. Points are coloured by order in which each fly was sequenced, and some error is noticable for the beige points in the top-right plots.</p> <p>The Plink program provides a useful method for selecting single-nucleotide polymorphisms (SNPs) which are independent of linkage disequilibrium (LD), and also of visualising the genetic relatedness between individuals in a population sample, using identity-by-descent analysis (IBD). The selection of LD-independent SNPs requires three user-specfied paramaters, alongside the genotype data: i. Window-size, in kilobases (Kb) within which all pairwise comparisons between SNPs will be made, ii. Step-size, in number of SNPs, iii. r2 threshold between any two SNPs, below which they are considered to be independent (fixed here at 0.5).</p>
Genome-wide estimation of linkage disequilibrium-independent SNPs in Drosophila melanogaster (Sussex LHM).
<p>Uses R to create SNP density across each chromosome arm. Uses Plink 1.9 to select independent SNPs with step sizes corresponding to chromosome density. Output data is combined_chromosomes_lhm_indep.txt a list of SNP IDs.<br> </p>
Supplementary material for: "Assessment of linkage disequilibrium patterns between structural variants and single nucleotide polymorphisms in three commercial chicken populations"
<p>Supplementary material for the publication "Assessment of linkage disequilibrium patterns between structural variants and single nucleotide polymorphisms in three commercial chicken populations"</p> <p>The realated preprint can be found at Research Square (<a href="https://doi.org/10.21203/rs.3.rs-861830/v1">https://doi.org/10.21203/rs.3.rs-861830/v1</a>)</p> <p>Supplementary file 1: Supplementary results, tables and figures.</p> <p>Supplementary file 2: MultiQC report.</p> <p>Supplementary file 3: Observer concordance of the visual filtering step.</p> <p>Supplementary file 4: Snakemake workflow and scripts.</p> <p> </p>
Fig. 2 in Genetic Differentiation And Linkage Disequilibrium In A Spatially Fragmented Population Of Cheilosia Vernalis (Diptera: Syrphidae) From The Balkan Peninsula
Fig. 2. Standardized variance of allelic frequencies FST (open symbols) and genetic distance D (NEI 1978) (filled symbols) plotted against corresponding geographic distance between subpopulation pairs of Cheilosia vernalis: Durmitor-Morinj (75 km), Fruška Gora- Durmitor (240 km), and Fruška Gora-Morinj (306 km). Pearson correlation coefficients between geographic distance and FST and D
Fig. 1 in Genetic Differentiation And Linkage Disequilibrium In A Spatially Fragmented Population Of Cheilosia Vernalis (Diptera: Syrphidae) From The Balkan Peninsula
Fig. 1. Map of Serbia and Montenegro showing sampling sites for the studied subpopulations of Chelosia vernalis, and genotype distribution at the Pgm locus. The Pgm locus was the most variable locus in the surveyed subpopulations, and along with differences of allele frequency variances at the
Negative linkage disequilibrium between amino acid changing variants reveals interference among deleterious mutations in the human genome
Open the record for dataset details and reuse information.
GCTB SBayesR shrunk sparse linkage disequilibrium matrices for HM3 variants, summary statistics and predictors generated from "Improved polygenic prediction by Bayesian multiple regression on summary statistics" by Lloyd-Jones, Zeng et al. 2019.
<p>GCTB LD matrices and results for HapMap 3 variants and 2.8M variants, which were used for</p> <p>simulation, cross-validation and across biobank analyses in the manuscript "Improved polygenic</p> <p>prediction by Bayesian multiple regression on summary statistics" by Lloyd-Jones, Zeng et al.</p> <p>2019.</p> <p>Unzip and see README for further details.</p>
Data for Extremely sparse models of linkage disequilibrium in ancestrally diverse association studies
<p>Data from <em>Extremely sparse models of linkage disequilibrium in ancestrally diverse association studies </em>(2023). This includes linkage disequilibrium graphical models (LDGMs) created from <a href="https://www.biorxiv.org/content/10.1101/2021.02.06.430068v2">high-coverage 1000 Genomes Project sequencing data</a>. This dataset consists of LDGM precision matrices, LDGM graphical models of SNPs, and lists of SNPs, all split into <a href="https://www.biorxiv.org/content/10.1101/2022.03.04.483057v1">1,361 approximately independent LD blocks across the genome</a>. The dataset additionally contains genotype information from chromosomes 21 and 22, and inferred tree sequences of high coverage 1000 Genomes Project Data, summary statistics from four traits in the UK Biobank, and UK biobank correlation matrices from chromosomes 21 and 22. All genomic data is in the GRCh38 build.</p> <p>The data can be cited as follows:</p> <p>Pouria Salehi Nowbandegani, Anthony Wilder Wohns, Jenna L. Ballard, Eric S. Lander, Alex Bloemendal, Benjamin M. Neale, and Luke J. O’Connor. Extremely sparse models of linkage disequilibrium in ancestrally diverse association studies. Nat Genet. (2023) DOI: 10.1038/s41588-023-01487-8</p> <p> </p> <p>The directory contains `.tar.gz` files, which can be extracted and unzipped with:</p> <pre><code class="language-bash">$ tar -xvf FILENAME.tar.gz</code></pre> <p>All LD block files are named by chromosome and start/end basepair coordinates.</p> <ul> <li> <p>1kg_nygc_trios_removed_All_pops_geno_ids_pops.csv: The file contains 5008 rows, 2 for each individual in the 1000 Genomes Project. Each row contains the individual ID of the 1000 genomes individual, and the ancestry group and continental ancestry group that individual was assigned to. Rows correspond to columns in `.genos` files. </p> </li> <li><em>AFR/AMR/EAS/EUR/SAS.precision.tar.gz</em>: Precision matrices for the relevant ancestry group for each LD block. Edge lists contain one row for each non-zero entry of the precision matrix. There are no column names.</li> <li><em>genos_chr21_22.tar.gz</em>: for the 40 LD blocks on chromosomes 21-22, .genos files are 0/1 matrices, with dimension number-of-SNPs by number-of-samples . Each LD matrix contains one column for each row in the SNP list files, and one row for each row in the sample ID files.</li> <li><em>ldgms.tar.gz:</em> 1361 LDGMs (*.edgelist files). Edge lists contain one row for each non-zero entry of the LDGM adjacency matrix. There is one LDGM edge list for each LD block. Each row represents an edge, as a tuple (index_1, index_2, entry). For the LDGM adjacency matrices, the entry is the edge weight, where 0 represents a strong dependency and e.g. 6 represents a weak dependency.</li> <li><em>snplists_GRch38positions.tar.gz</em>: 1361 *.snplist files, each of which contains information on the SNPs in each LD block. Each SNP list is an <em>n</em> x 11 table (<em>n </em>=<em> </em>number of SNPs<em>)</em>, one for each LD block. The columns are: <ul> <li> <p>index: these non-unique indices, starting at zero, correspond to rows and columns of the LDGMs. There can be multiple SNPs for a single index, which occurs when the corresponding mutations occur on the same brick of the bricked tree sequence. SNPs with the same index have high (nearly perfect) LD.</p> </li> <li> <p>anc_alleles: ancestral allele</p> </li> <li> <p>deriv_alleles: derived allele</p> </li> <li> <p>EUR: allele frequency of derived allele in EUR samples</p> </li> <li> <p>EAS: allele frequency of derived allele in EAS samples</p> </li> <li> <p>AMR: allele frequency of derived allele in AMR samples</p> </li> <li> <p>SAS: allele frequency of derived allele in SAS samples</p> </li> <li> <p>AFR: allele frequency of derived allele in AFR samples</p> </li> <li> <p>site_ids: unique identifier of each SNP, mostly as RSIDs</p> </li> <li> <p>position: GRCh38 position of SNP</p> </li> <li> <p>swap: indicates strandness swap</p> </li> </ul> </li> <li> <p><em>ukb.tar</em>: Correlation matrices and SNP lists for SNPs in the UK Biobank.</p> <ul> <li> <p>correlation_matrices/: Correlation matrices for SNPs in the UK biobank, computed by Weissbrod et al. 2020 Nat Genet and can be downloaded by following the instructions <a href="https://alkesgroup.broadinstitute.org/UKBB_LD">here</a>.</p> </li> <li> <p>snplists/: List of SNPs in the *.snplist format included in the UK Biobank</p> </li> </ul> </li> <li> <p><em>tree_seqs.tar</em>: contains 22 tree sequences inferred by <a href="https://tsinfer.readthedocs.io">tsinfer</a> from the <a href="https://www.biorxiv.org/content/10.1101/2021.02.06.430068v2">30x 1000 Genomes Project Data</a>. Tree sequences can be unzipped with <a href="https://tszip.readthedocs.io/en/latest/">tszip</a>.</p> </li> <li> <p>Summary statistics: there are four summary statistics files, obtained from <a href="https://alkesgroup.broadinstitute.org/UKBB/">https://alkesgroup.broadinstitute.org/UKBB/</a>, and computed by Loh et al. 2018 Nat Genet.</p> </li> </ul> <table> <tbody> <tr> <td> <p>Phenotype</p> </td> <td> <p>Heritability estimate </p> </td> <td> <p>Effective sample size</p> </td> <td> <p>Number of SNPs</p> </td> </tr> <tr> <td> <p>Height</p> </td> <td> <p>0.570</p> </td> <td> <p>650K</p> </td> <td> <p>12 Million</p> </td> </tr> <tr> <td> <p>Body mass index</p> </td> <td> <p>0.303</p> </td> <td> <p>500K</p> </td> <td> <p>12 Million</p> </td> </tr> <tr> <td> <p>Cardiovascular disease</p> </td> <td> <p>0.155</p> </td> <td> <p>450K</p> </td> <td> <p>12 Million</p> </td> </tr> <tr> <td> <p>Type 2 diabetes</p> </td> <td> <p>0.073</p> </td> <td> <p>450K</p> </td> <td> <p>12 Million</p> </td> </tr> </tbody> </table>
Founder effects shape linkage disequilibrium and genomic diversity of a partially clonal invader
Open the record for dataset details and reuse information.
Data from: A method for detecting recent changes in contemporary effective population size from linkage disequilibrium at linked and unlinked loci
Estimation of contemporary effective population size (Ne) from linkage disequilibrium (LD) between unlinked pairs of genetic markers has become an important tool in the field of population and conservation genetics. If data pertaining to physical linkage or genomic position are available for genetic markers, estimates of recombination rate between loci can be combined with LD data to estimate contemporary Ne at various times in the past. We extend the well-known, LD-based method of estimating contemporary Ne to include linkage information and show via simulation that even relatively small, recent changes in Ne can be detected reliably with a modest number of SNP loci. We explore several issues important to interpretation of the results and quantify the bias in estimates of contemporary Ne associated with the assumption that all loci in a large SNP dataset are unlinked. The approach is applied to an empirical dataset of SNP genotypes from a population of a marine fish where a recent, temporary decline in Ne is known to have occurred.
Data from: Population structure, genetic variation and linkage disequilibrium in perennial ryegrass populations divergently selected for freezing tolerance
Low temperature is one of the abiotic stresses seriously affecting the growth of perennial ryegrass (Lolium perenne L. Understanding the genetic control of freezing tolerance would aid in the development of cultivars of perennial ryegrass with improved adaptation to frost. A total number of 80 individuals (24 of High frost [HF]; 29 of Low frost [LF] and 27 of Unselected [US]) from the second generation of the two divergently selected populations and an unselected control population were genotyped using 278 genome-wide SNPs derived from Lolium perenne L. transcriptome sequence. Our studies showed that the HF and LF populations are very divergent after selection for freezing tolerance, whereas the HF and US populations are more similar. Linkage disequilibrium (LD) decay varied across the seven chromosomes and the conspicuous pattern of LD between the HF and LF population confirmed their divergence in freezing tolerance. Furthermore, two Fst outlier methods; finite island model (fdist) by LOSITAN and hierarchical structure model using ARLEQUIN detected six loci under directional selection. These outlier loci are most probably linked to genes involved in freezing tolerance, cold adaptation and abiotic stress and might be the potential marker resources for breeding perennial ryegrass cultivars with improved freezing tolerance.
Data from: Whole-genome patterns of linkage disequilibrium across flycatcher populations clarify the causes and consequences of fine-scale recombination rate variation in birds
Recombination rate is heterogeneous across the genome of various species, and so are genetic diversity and differentiation as a consequence of linked selection. However, we still lack a clear picture of the underlying mechanisms for regulating recombination. Here we estimated fine-scale population recombination rate based on the patterns of linkage disequilibrium (LD) across the genomes of multiple populations of two closely related flycatcher species (Ficedula albicollis and F. hypoleuca). This revealed an overall conservation of the recombination landscape between these species at the scale of 200-kb, but we also identified differences in the local rate of recombination despite their recent divergence (<1 million years). Genetic diversity and differentiation were associated with recombination rate in a lineage-specific manner, indicating differences in the extent of linked selection between species. We detected 400-3,085 recombination hotspots per population. Location of hotspots was conserved between species, but the intensity of hotspot activity varied between species. Recombination hotspots were primarily associated with CpG islands (CGIs), regardless of whether CGIs were at promoter regions or away from genes. Recombination hotspots were also associated with specific transposable elements (TEs), but this association appears indirect due to shared preferences of the transposition machinery and the recombination machinery for accessible open chromatin regions. Our results suggest that CGIs are a major determinant of the localization of recombination hotspots, and we propose that both the distribution of TEs and fine-scale variation in recombination rate may be associated with the evolution of the epigenetic landscape.
Data from: Patterns of cyto-nuclear linkage disequilibrium in Silene latifolia: genomic heterogeneity and temporal stability
Non-random association of alleles in the nucleus and cytoplasmic organelles, or cyto-nuclear linkage disequilibrium (LD), is both an important component of a number of evolutionary processes and a statistical indicator of others. The evolutionary significance of cyto-nuclear LD will depend on both its magnitude and how stable those associations are through time. Here, we use a longitudinal population genetic data set to explore the magnitude and temporal dynamics of cyto-nuclear disequilibria through time. We genotyped 135 and 170 individuals from 16 and 17 patches of the plant species Silene latifolia in Southwestern VA, sampled in 1993 and 2008, respectively. Individuals were genotyped at 14 highly polymorphic microsatellite markers and a single-nucleotide polymorphism (SNP) in the mitochondrial gene, atp1. Normalized LD (D′) between nuclear and cytoplasmic loci varied considerably depending on which nuclear locus was considered (ranging from 0.005–0.632). Four of the 14 cyto-nuclear associations showed a statistically significant shift over approximately seven generations. However, the overall magnitude of this disequilibrium was largely stable over time. The observed origin and stability of cyto-nuclear LD is most likely caused by the slow admixture between anciently diverged lineages within the species' newly invaded range, and the local spatial structure and metapopulation dynamics that are known to structure genetic variation in this system.
Data from: Genetic diversity, linkage disequilibrium and selection signatures in Chinese and Western pigs revealed by genome-wide SNP markers
To investigate population structure, linkage disequilibrium (LD) pattern and selection signature at the genome level in Chinese and Western pigs, we genotyped 304 unrelated animals from 18 diverse populations using porcine 60 K SNP chips. We confirmed the divergent evolution between Chinese and Western pigs and showed distinct topological structures of the tested populations. We acquired the evidence for the introgression of Western pigs into two Chinese pig breeds. Analysis of runs of homozygosity revealed that historical inbreeding reduced genetic variability in several Chinese breeds. We found that intrapopulation LD extents are roughly comparable between Chinese and Western pigs. However, interpopulation LD is much longer in Western pigs compared with Chinese pigs with average r20.3 values of 125 kb for Western pigs and only 10.5 kb for Chinese pigs. The finding indicates that higher-density markers are required to capture LD with causal variants in genome-wide association studies and genomic selection on Chinese pigs. Further, we looked across the genome to identify candidate loci under selection using FST outlier tests on two contrast samples: Tibetan pigs versus lowland pigs and belted pigs against non-belted pigs. Interestingly, we highlighted several genes including ADAMTS12, SIM1 and NOS1 that show signatures of natural selection in Tibetan pigs and are likely important for genetic adaptation to high altitude. Comparison of our findings with previous reports indicates that the underlying genetic basis for high-altitude adaptation in Tibetan pigs, Tibetan peoples and yaks is likely distinct from one another. Moreover, we identified the strongest signal of directional selection at the EDNRB loci in Chinese belted pigs, supporting EDNRB as a promising candidate gene for the white belt coat color in Chinese pigs. Altogether, our findings advance the understanding of the genome biology of Chinese and Western pigs.
Biomass data to accompany Inter-chromosomal linkage disequilibrium and linked fitness cost loci associated with selection for herbicide resistance
<ul> <li>The adaptation of weeds to herbicide is both a significant problem in agriculture and a model of rapid adaptation. However, significant gaps remain in our knowledge of resistance controlled by many loci and the evolutionary factors that influence the maintenance of resistance.</li> <li>Here, using herbicide-resistant populations of the common morning glory (<em>Ipomoea</em> <em>purpurea</em>), we perform a multi-level analysis of the genome and transcriptome to uncover putative loci involved in nontarget-site herbicide resistance (NTSR) and to examine evolutionary forces underlying the maintenance of resistance in natural populations.</li> <li>We found loci involved in herbicide detoxification and stress sensing to be under selection and confirmed that detoxification is responsible for glyphosate resistance using a functional assay. We identified interchromosomal linkage disequilibrium (ILD) among loci under selection reflecting either historical processes or additive effects leading to the resistance phenotype. We further identified potential fitness cost loci that were strongly linked to resistance alleles, indicating the role of genetic hitchhiking in maintaining the cost.</li> <li>Overall, our work suggests that NTSR glyphosate resistance in<em> I. purpurea</em> is conferred by multiple genes which are potentially maintained through generations <em>via</em> ILD and that the fitness cost associated with resistance in this species is likely a by-product of genetic-hitchhiking.</li> </ul>
Mapping of End Stage Renal Disease Genetic Susceptibility in African Americans by Admixture Linkage Disequilibrium
ClinicalTrials.gov study NCT00559767. IPD Sharing: Not stated. Countries: 1. Publications: 3.
Data from: Genome-wide linkage disequilibrium and past effective population size in three Korean cattle breeds
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.