ReMM score
<p>The <strong>Re</strong>gulatory <strong>M</strong>endelian <strong>M</strong>utation (ReMM) score was created for relevance prediction of non-coding variations (SNVs and small InDels) in the human genome (hg19/hg38) in terms of Mendelian diseases.</p> <p> </p> <p><strong>Usage</strong></p> <p>The ReMM score is genome position wise (nucleotide changes are neglected). We precomputed all positions in the human genome (hg19 and hg38 release) and stored the values in a tabix file (1-based). The scores ranging from 0 (non-deleterious) to 1 (deleterious).</p> <p>If you want to use the ReMM score together with the Genomiser, please have a look at the <a href="https://exomiser.github.io/Exomiser/">Exomiser framework manual</a></p> <p>For more information, direct VCF file scoring, and an API please have a look at our website: <a href="https://remm.bihealth.org">https://remm.bihealth.org</a></p> <p><strong>ReMM score changelog</strong></p> <p>0.4:</p> <ul> <li>Features: <ul> <li>For missing values using genome mean of feature for sequence and conservaton features. 1 for p-value. All other features have zero as missing value</li> <li>Updating DGVCount to 02/25/2020 on hg19/hg38</li> <li>Update dbVARCount to 10/20/2021 on hg19/hg38</li> <li>Update ISCApath to 11/03/2021 on hg19/hg38</li> <li>Replace tfbsConsSites with UCSC table encRegTfbsClustered on hg19/hg38</li> </ul> </li> <li>Software: <ul> <li>Using parSMURF for training</li> </ul> </li> <li>Complete retraining of hg19 and hg38 builds (hg19: AUROC=0.993; AUPRC=0.394; hg38: AUROC=0.996; AUPRC=0.610)</li> </ul> <p>0.3.1.post1:</p> <ul> <li>New hg38 release. Completely retrained on the new genome build. <ul> <li>Training data: <ul> <li>Liftover positives. No change in size.</li> <li>Negatives used from CADD v1.4 GRCh38 (human derived), filtered as described in the original paper (<a href="https://doi.org/10.1016/j.ajhg.2016.07.005">https://doi.org/10.1016/j.ajhg.2016.07.005</a>). Size slightly different (hg38: 13,902,234; hg19: 14,755,199) .</li> </ul> </li> <li>Features <ul> <li>Same size as in hg19: 26 features.</li> <li>We tried to use the same features as in hg19. Sometimes new versions of data have to be used (e.g. DGV, ISCA, dbVAR).</li> </ul> </li> <li>Training was done with the parSMURF implementation of hyperSMURF. <ul> <li>Same hg19 parameters are used.</li> </ul> </li> <li>Metrics via 10-fold cytoband cross-validation (same cytoband to fold map): <ul> <li>Area under the ROC curve: 0.996 (hg19: 0.989, see <a href="https://doi.org/10.1016/j.ajhg.2016.07.005">https://doi.org/10.1016/j.ajhg.2016.07.005</a>)</li> <li>Area under the precision recall curve: 0.548 (hg19: 0.441, see <a href="https://doi.org/10.1016/j.ajhg.2016.07.005">https://doi.org/10.1016/j.ajhg.2016.07.005</a>)</li> </ul> </li> </ul> </li> <li>Scores for hg19 in this release are the same as version 0.3.1. Only the files have been renamed.</li> </ul> <p>0.3.1:</p> <ul> <li>Bugfix of region chr17:79759050-81195210. Region is missing in older versions.</li> </ul> <p>0.3:</p> <ul> <li>First official public version.</li> <li>Values for positions in training data are computed by cytoband-aware 10 fold cross-validation.</li> <li>Other position scores are compted by a generalized model of all training data.</li> <li>This version was used in the Genomiser publication (Smeley et.al. A Whole-Genome Analysis Framework for Effective Identification of Pathogenic Regulatory Variants in Mendelian Disease. AHJG. 2016)</li> </ul>
ShareScore
48/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 8
- Harmonization
- 8
- Access
- 20
- Reuse readiness
- 8
- Engagement
- 4