Skip to main content
zenodoopen

XTB2-MolData : Dataset of 12 Million Molecules

<p>This dataset is an open chemistry database containing optimized molecular geometries and electronic properties calculated by the GFN2-xTB method (<a href="https://doi.org/10.1002/wcms.1493">C. Bannwarth et al.</a>) for 12.6 million organic molecules contained C, H, O, and N atoms.</p> <p>The initial geometries, before optimization by GFN2-xTB method, are taken from PubChem PM6 (<a href="https://doi.org/10.1021/acs.jcim.0c00740">Shimazaki et al.</a>) database.</p> <p>We also include our python code to manage a large molecule database. This code includes scripts to generate input files for Gaussian software, to read Gaussian output files, to create a small reduced dataset based on clustering algorithm, and many scripts to analyze the molecular properties included in the database.</p> <p>This code can be also taken from github:&nbsp; <a href="https://github.com/Castaneche/MolDataFW">https://github.com/Castaneche/MolDataFW</a>.</p>

ShareScore

44/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
8
Harmonization
8
Access
16
Reuse readiness
8
Engagement
4

Topics