Relate-estimated coalescence rates, allele ages, and selection p-values for the 1000 Genomes Project
<p><strong>Overview</strong></p> <p>Coalescence rates, allele ages, and p-values for evidence of positive selection calculated for 2478 samples of the 1000 Genomes Project using Relate.</p> <p>We estimated the joint genealogy of all 1000 GP populations and then extracted the embedded genealogy for each population.<br> For the genealogy of each population, we jointly estimated the population size history and branch lengths. <br> Variants segregating in more than one population therefore have correlated but different allele ages in each population.</p> <p>Please refer to <a href="https://www.nature.com/articles/s41588-019-0484-x">Speidel et al. Nature Genetics (2019)</a> for more details or email leo.speidel@outlook.com for any queries.</p> <p><strong>Coalescence rates</strong></p> <p>The zipped directory coalescence_rates.zip contains coalescence rates for 26 populations in the 1000 Genomes Project data set.</p> <ul> <li>The .coal files show the haploid coalescence rates, please refer to the <a href="https://myersgroup.github.io/relate/modules.html#PopulationSizeScript_FileFormats">Relate documentation</a> for the file format.</li> <li>The popsize.RData file is an R data frame storing the diploid population sizes (0.5/coalescence rate) calculated using the .coal files. The columns of this data frame, named "pop_size", are <ul> <li>gens_ago: Time in generations at which epoch starts. (To get years from generations, we multiply by 28.)</li> <li>population_size: Diploid population size in this epoch.</li> <li>population: Name of population </li> <li>region: Name of region (AFR, AMR, EAS, EUR, SAS)</li> </ul> </li> </ul> <p><strong>Allele ages and selection p-values</strong></p> <p>The zipped directories allele_ages_*.zip contain R data frames for each 1000GP population storing allele ages and selection p-values.<br> Please note that only mutations that segregate in the population and map to a unique branch in the Relate-estimated marginal trees are included. Selection p-values are only provided for mutations of DAF > 2 that pass quality filters (see Speidel et al., 2019). </p> <p>To get an age estimate for a neutral mutation, use 0.5*(lower_age + upper_age). To get years from generations, we multiply by 28.</p> <p>The columns of these data frames, named "allele_ages", are</p> <ul> <li>CHR: chromosome index</li> <li>BP: base-pair position (GRCh37)</li> <li>ID: id of SNP</li> <li>lower_age: Age in generations of coalescence event at the lower end of the branch onto which the mutation maps</li> <li>upper_age: Age in generations of coalescence event at the upper end of the branch onto which the mutation maps</li> <li>ancestral/derived: Ancestral/derived allele</li> <li>upstream: Upstream (5') allele</li> <li>downstream: Downstream (3') allele</li> <li>DAF: Derived-allele frequency</li> <li>pvalue: log10 p-value for selection evidence</li> </ul>
ShareScore
44/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 8
- Harmonization
- 8
- Access
- 16
- Reuse readiness
- 8
- Engagement
- 4