Skip to main content
zenodoopen

Datasets for understanding the importance of conformation in property prediction models

<p>Descriptor and conformer data sets for molecular property and reaction selectivity prediction tasks. The PQC data set was created based on a part of the PubChemQC PM6 dataset (J. Chem. Inf. Model. 2020, 60, 12, 5891&ndash;5899), which contains two- and three-dimensional descriptors and conformers. The APTC data sets are based on the data sets for asymmetric phase transfer catalysts with enantio-selectivity (<a href="https://github.com/Laboratoire-de-Chemoinformatique/3D-MIL-QSSR/tree/main/datasets" target="_blank" rel="noopener">https://github.com/Laboratoire-de-Chemoinformatique/3D-MIL-QSSR/tree/main/datasets</a>). The melting point data set was created from the Jean-Claude Bradley Double Plus Good (Highly Curated and Validated) Melting Points Dataset (<a href="https://doi.org/10.6084/m9.figshare.1031638.v1">https://doi.org/10.6084/m9.figshare.1031638.v1</a>).</p> <p>They contained descriptors and conformers to train and validate machine learning models.</p> <p>Detailed explanations on how to use these datasets are found in the Github repository: <a href="https://github.com/YuHamakawa/Conformation-Importance-ML-Models">https://github.com/YuHamakawa/Conformation-Importance-ML-Models</a>.&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p>

ShareScore

32/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
4
Harmonization
4
Access
16
Reuse readiness
8
Engagement
0