Skip to main content
zenodoopen

Supplementary Data: Uncovering DFG-out sequence propensity determinants of kinases with machine learning

<h2>General description</h2><p>This submission accompanies the paper "Uncovering DFG-out sequence propensity determinants of kinases with machine learning" and covers:</p><ul><li><strong>Trained models (train/models_latest)</strong>. A subset of these is also published as a part of the <a href="https://github.com/edikedik/kinactive">KinActive tool</a>; each model is of `KinactiveClassifier` type and can be loaded using this tool as well via `kinactive.io.load()` function).</li><li><strong>Patched sequences (patched_seqs)</strong>. PDB sequences, with missing regions patched by UniProt sequences. This is a collection of <a href="https://github.com/edikedik/lXtractor">lXtractor</a> ChainSequence objects.</li><li><strong>All labels, variables, and datasets</strong> (datasets and labels).</li><li><strong>Initial chains and predictions for SwissProt proteins</strong> (SP_predictions).</li></ul><h2>Note on model abbreviations</h2><p>The main text emphasized datasets that yielded more interpretable results. These were constructed from domain sequences, labeled as apo, inactive, DFG-in, or DFG-out, and further divided into TK and STK subsets. We refer to these datasets with the abbreviation <strong>AAIO</strong> (Apo All In or Out) to distinguish them from additional datasets.</p><p>In addition to <strong>AAIO</strong>, we explored two alternative labeling strategies:</p><ul><li><strong>AHAO</strong> (Apo/Holo Any Out): This dataset includes sequences from ligand-bound entries. All sequences within 95\% identity clusters are labeled as DFG-out if the cluster contains at least one sequence in this state. All others are labeled as DFG-in.</li><li><strong>AAO</strong> (Apo Any Out): This dataset excludes sequences corresponding to ligand-bound entries but includes those with conflicting conformational tendencies within 95\% identity clusters.</li></ul><p>Each of these datasets had two versions:</p><ol><li>A seed version, denoted by a "*" symbol.</li><li>A version enriched with orthologous sequences (no special designation).</li></ol><p>Together with the <strong>TkST</strong> datasets (encompassing TK and STK labels) used for testing the methodology, a total of 14 datasets were used, and both <i>RF</i> and <i>XGB</i> models were applied to each, using the same initial settings, resulting in 28 different models. This additional information is provided for completeness.</p>

ShareScore

32/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
8
Harmonization
4
Access
16
Reuse readiness
0
Engagement
4