Skip to main content
zenodoopen

Two ChEMBL-34 subsets (lead-like and drug-like molecules)

<p>Two subsets of molecules from ChEMBL-34[1].</p> <p>Those molecular datasets might be useful to people training molecular generators.</p> <p>After decompression, you will get:<br>chembl34_stable_ES_OA_LL.smi: 585,272 molecules.<br>chembl34_stable_ES_OA_DL.smi: 756,420 molecules.</p> <p>stable=non-reactive molecules (filtered-out reactive functional groups from [5]).<br><a href="https://github.com/UnixJunkie/molenc/blob/master/bin/molenc_stable.py">https://github.com/UnixJunkie/molenc/blob/master/bin/molenc_stable.py</a></p> <p>ES=Easy Synthesis (SAscore &lt;= 3.0) [2].<br><a href="https://github.com/UnixJunkie/molenc/blob/master/bin/molenc_SA.py">https://github.com/UnixJunkie/molenc/blob/master/bin/molenc_SA.py</a></p> <p>OA=Orally Available (according to a classifier trained on the dataset from [6]).</p> <p>LL=Lead-Like (almost the definition from [3]).<br><a href="https://github.com/UnixJunkie/molenc/blob/master/bin/molenc_lead.py">https://github.com/UnixJunkie/molenc/blob/master/bin/molenc_lead.py</a></p> <p>DL=Drug-Like (definition from [4]).<br><a href="https://github.com/UnixJunkie/molenc/blob/master/bin/molenc_drug.py">https://github.com/UnixJunkie/molenc/blob/master/bin/molenc_drug.py</a></p> <div> <h1>Bibliography</h1> <a href="https://github.com/UnixJunkie/chembl34_subsets#bibliography"></a></div> <ol> <li> <p>Zdrazil, B., Felix, E., Hunter, F., Manners, E. J., Blackshaw, J., Corbett, S., ... &amp; Leach, A. R. (2024). The ChEMBL Database in 2023: a drug discovery platform spanning multiple bioactivity data types and time periods. Nucleic acids research, 52(D1), D1180-D1192. <a href="https://doi.org/10.1093/nar/gkad1004" rel="nofollow">https://doi.org/10.1093/nar/gkad1004</a></p> </li> <li> <p>Ertl, P., &amp; Schuffenhauer, A. (2009). Estimation of synthetic accessibility score of drug-like molecules based on molecular complexity and fragment contributions. Journal of cheminformatics, 1, 1-11. <a href="https://jcheminf.biomedcentral.com/articles/10.1186/1758-2946-1-8" rel="nofollow">https://jcheminf.biomedcentral.com/articles/10.1186/1758-2946-1-8</a></p> </li> <li> <p>Hann, M. M., &amp; Oprea, T. I. (2004). Pursuing the leadlikeness concept in pharmaceutical research. Current opinion in chemical biology, 8(3), 255-263. <a href="https://doi.org/10.1016/j.cbpa.2004.04.003" rel="nofollow">https://doi.org/10.1016/j.cbpa.2004.04.003</a></p> </li> <li> <p>Tran-Nguyen, V. K., Jacquemard, C., &amp; Rognan, D. (2020). LIT-PCBA: an unbiased data set for machine learning and virtual screening. Journal of chemical information and modeling, 60(9), 4263-4273. <a href="https://pubs.acs.org/doi/10.1021/acs.jcim.0c00155" rel="nofollow">https://pubs.acs.org/doi/10.1021/acs.jcim.0c00155</a></p> </li> <li> <p>Lisurek, M., Rupp, B., Wichard, J., Neuenschwander, M., von Kries, J. P., Frank, R., ... &amp; K&uuml;hne, R. (2010) Design of chemical libraries with potentially bioactive molecules applying a maximum common substructure concept. Molecular diversity, 14, 401-408. <a href="https://link.springer.com/article/10.1007/s11030-009-9187-z" rel="nofollow">https://link.springer.com/article/10.1007/s11030-009-9187-z</a></p> </li> <li> <p>Falcon-Cano, G., Molina, C., &amp; Cabrera-Perez, M. A. (2020). ADME prediction with KNIME: development and validation of a publicly available workflow for the prediction of human oral bioavailability. Journal of chemical information and modeling, 60(6), 2660-2667. <a href="https://pubs.acs.org/doi/10.1021/acs.jcim.0c00019" rel="nofollow">https://pubs.acs.org/doi/10.1021/acs.jcim.0c00019</a></p> </li> </ol>

ShareScore

36/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
4
Harmonization
8
Access
16
Reuse readiness
8
Engagement
0