XTB2-MolData : Dataset of 12 Million Molecules
<p>This dataset is an open chemistry database containing optimized molecular geometries and electronic properties calculated by the GFN2-xTB method (<a href="https://doi.org/10.1002/wcms.1493">C. Bannwarth et al.</a>) for 12.6 million organic molecules contained C, H, O, and N atoms.</p> <p>The initial geometries, before optimization by GFN2-xTB method, are taken from PubChem PM6 (<a href="https://doi.org/10.1021/acs.jcim.0c00740">Shimazaki et al.</a>) database.</p> <p>We also include our python code to manage a large molecule database. This code includes scripts to generate input files for Gaussian software, to read Gaussian output files, to create a small reduced dataset based on clustering algorithm, and many scripts to analyze the molecular properties included in the database.</p> <p>This code can be also taken from github: <a href="https://github.com/Castaneche/MolDataFW">https://github.com/Castaneche/MolDataFW</a>.</p>
ShareScore
44/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 8
- Harmonization
- 8
- Access
- 16
- Reuse readiness
- 8
- Engagement
- 4