Datasets for understanding the importance of conformation in property prediction models
<p>Descriptor and conformer data sets for molecular property and reaction selectivity prediction tasks. The PQC data set was created based on a part of the PubChemQC PM6 dataset (J. Chem. Inf. Model. 2020, 60, 12, 5891–5899), which contains two- and three-dimensional descriptors and conformers. The APTC data sets are based on the data sets for asymmetric phase transfer catalysts with enantio-selectivity (<a href="https://github.com/Laboratoire-de-Chemoinformatique/3D-MIL-QSSR/tree/main/datasets" target="_blank" rel="noopener">https://github.com/Laboratoire-de-Chemoinformatique/3D-MIL-QSSR/tree/main/datasets</a>). The melting point data set was created from the Jean-Claude Bradley Double Plus Good (Highly Curated and Validated) Melting Points Dataset (<a href="https://doi.org/10.6084/m9.figshare.1031638.v1">https://doi.org/10.6084/m9.figshare.1031638.v1</a>).</p> <p>They contained descriptors and conformers to train and validate machine learning models.</p> <p>Detailed explanations on how to use these datasets are found in the Github repository: <a href="https://github.com/YuHamakawa/Conformation-Importance-ML-Models">https://github.com/YuHamakawa/Conformation-Importance-ML-Models</a>. </p> <p> </p> <p> </p>
ShareScore
32/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 4
- Harmonization
- 4
- Access
- 16
- Reuse readiness
- 8
- Engagement
- 0