Skip to main content
zenodoopen

SDDF Energy Dataset

<p>This conformational energy dataset, developed as part of the Smart Distributed Data Factory (SDDF) project, contains over 2.75 million molecular conformations based on drug-like molecules sourced from the <strong>ENAMINE</strong> database. Energies were calculated using&nbsp;<strong>DFT</strong> with the &omega;B97x density functional and the 6&ndash;31G(d) basis set. The conformations were generated from SMILES using RDKit, MMFF94 optimization, and molecular dynamics (MD) simulations, providing a diverse set of molecular structures and energy states.</p> <ul> <li><strong>RDKit Conformations:</strong> 1,123,693</li> <li><strong>RDKit + MMFF94 Optimized:</strong> 1,151,936</li> <li><strong>MD-Generated:</strong> 483,279</li> </ul> <p>This dataset serves as a benchmark for energy prediction models, with training (638,617 examples), validation (134,732 examples), and test subsets (24,890 examples) created using a strict scaffold-based split to ensure no overlap and less than 70% similarity between the training and test sets.</p> <p>Dataset contents:</p> <ul> <li><em>data.tar.gz</em>: contains the conformations in Structured Data File format, grouped into separate folders based on the molecule ID. Each conformation's label is provided within its SDF file as a property named "energy".</li> <li><em>INDEX.smi</em>: specifies the molecule IDs and their corresponding SMILES.</li> <li><em>SOURCES.csv</em>: specifies the conformation generation method for each conformation.</li> <li><em>SDDF_train.tsv</em>, <em>SDDF_validation.tsv</em>, and <em>SDDF_test.tsv</em>&nbsp;specify the molecule IDs and conformations for each subset of the benchmark.</li> </ul> <p>A detailed description is provided in the accompanying paper.</p>

ShareScore

44/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
8
Harmonization
4
Access
20
Reuse readiness
8
Engagement
4

Topics