Skip to main content
zenodoopen

Relate-estimated coalescence rates, allele ages, and selection p-values for the 1000 Genomes Project

<p><strong>Overview</strong></p> <p>Coalescence rates, allele ages, and p-values for evidence of positive selection calculated for 2478&nbsp;samples of the&nbsp;1000 Genomes Project&nbsp;using Relate.</p> <p>We estimated the joint genealogy of all 1000 GP populations and then extracted the embedded genealogy for each population.<br> For the genealogy of each population, we jointly estimated the population size history and branch lengths.&nbsp;<br> Variants segregating in more than one&nbsp;population&nbsp;therefore have&nbsp;correlated but different allele ages in each population.</p> <p>Please refer to&nbsp;<a href="https://www.nature.com/articles/s41588-019-0484-x">Speidel et al.&nbsp;Nature Genetics (2019)</a>&nbsp;for more details or email leo.speidel@outlook.com for any queries.</p> <p><strong>Coalescence rates</strong></p> <p>The zipped directory&nbsp;coalescence_rates.zip&nbsp;contains coalescence rates for 26 populations in the 1000 Genomes Project data set.</p> <ul> <li>The .coal files show the haploid coalescence rates, please refer to the&nbsp;<a href="https://myersgroup.github.io/relate/modules.html#PopulationSizeScript_FileFormats">Relate documentation</a>&nbsp;for the file format.</li> <li>The popsize.RData file is an R data frame storing the diploid population sizes (0.5/coalescence rate) calculated using the .coal files. The columns of this data frame, named &quot;pop_size&quot;,&nbsp;are <ul> <li>gens_ago: Time in generations at which epoch starts. (To get years from generations, we multiply by 28.)</li> <li>population_size: Diploid population size in this epoch.</li> <li>population: Name of population&nbsp;</li> <li>region: Name of region (AFR, AMR, EAS, EUR, SAS)</li> </ul> </li> </ul> <p><strong>Allele ages and selection p-values</strong></p> <p>The zipped directories&nbsp;allele_ages_*.zip&nbsp;contain&nbsp;R&nbsp;data frames for each 1000GP population storing allele ages and selection p-values.<br> Please note that only mutations that segregate in the population and map to a unique branch in the Relate-estimated marginal trees are included. Selection p-values are only provided for mutations of DAF &gt; 2 that pass quality filters (see Speidel et al., 2019).&nbsp;</p> <p>To get an age estimate for a neutral mutation, use&nbsp;0.5*(lower_age + upper_age). To get years from generations, we multiply by 28.</p> <p>The columns of these&nbsp;data frames, named &quot;allele_ages&quot;,&nbsp;are</p> <ul> <li>CHR: chromosome index</li> <li>BP: base-pair position (GRCh37)</li> <li>ID: id of SNP</li> <li>lower_age: Age in generations of coalescence event at the lower end of the branch onto which the mutation maps</li> <li>upper_age: Age in generations of coalescence event at the upper end of the branch onto which the mutation maps</li> <li>ancestral/derived: Ancestral/derived allele</li> <li>upstream: Upstream (5&#39;) allele</li> <li>downstream: Downstream (3&#39;) allele</li> <li>DAF: Derived-allele frequency</li> <li>pvalue: log10 p-value for selection evidence</li> </ul>

ShareScore

44/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
8
Harmonization
8
Access
16
Reuse readiness
8
Engagement
4

Topics