Representations and associated fitness values of protein sequences
<p>We studied the ability to predict protein fitness from sequence using our method <a href="https://github.com/amillig/MERGE" target="_blank" rel="noopener">MERGE</a> and other methods (i.e., ECNet, eUniRep, EVmutation, One-Hot, and UniRep). The following files are included in this repository:</p> <ul> <li>The folder <em>ECNet</em> contains the braw, csv, and fasta files required to run <a href="https://github.com/luoyunan/ECNet" target="_blank" rel="noopener">ECNet</a>.</li> <li>The folders <em>eUniRep, </em><em>MERGE, One-Hot</em>, and <em>UniRep</em> contain csv files including the names and fitness values of protein variants as well as a numerical representation of their sequence. The folder <em>MERGE </em>also contains representations of the wild type sequences as npy files in the <em>wts</em> folder.</li> <li>The folder <em>Params</em> contains params files generated with <a href="https://github.com/debbiemarkslab/plmc" target="_blank" rel="noopener">PLMC</a>.</li> <li>The script <em>get_performances.py</em> enables to generate models using different methods (i.e., eUniRep, EVmutation, MERGE, One-Hot, Pure_ML, and UniRep) and to determine their performance for predicting the fitness of protein variants from sequence.</li> </ul> <p><strong>References</strong></p> <table> <tbody> <tr> <td><a href="https://doi.org/10.1038/s41467-021-25976-8" target="_blank" rel="noopener">ECNet</a></td> <td> <p>Luo, Y., Jiang, G., Yu, T. et al. ECNet is an evolutionary context-integrated deep learning framework for protein engineering. Nat Commun 12, 5743 (2021).</p> </td> </tr> <tr> <td><a href="https://doi.org/10.1038/s41592-021-01100-y" target="_blank" rel="noopener">eUniRep</a></td> <td>Biswas, S., Khimulya, G., Alley, E.C. et al. Low-N protein engineering with data-efficient deep learning. Nat Methods 18, 389–396 (2021)</td> </tr> <tr> <td><a href="https://doi.org/10.1038/nbt.3769" target="_blank" rel="noopener">EVmutation</a></td> <td>Hopf, T., Ingraham, J., Poelwijk, F. et al. Mutation effects predicted from sequence co-variation. Nat Biotechnol 35, 128–135 (2017).</td> </tr> <tr> <td><a href="https://doi.org/10.1038/s41592-019-0598-1" target="_blank" rel="noopener">UniRep</a></td> <td>Alley, E.C., Khimulya, G., Biswas, S. et al. Unified rational protein engineering with sequence-based deep representation learning. Nat Methods 16, 1315–1322 (2019).</td> </tr> </tbody> </table> <p> </p>
ShareScore
24/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 4
- Harmonization
- 4
- Access
- 8
- Reuse readiness
- 8
- Engagement
- 0