SDDF Energy Dataset
<p>This conformational energy dataset, developed as part of the Smart Distributed Data Factory (SDDF) project, contains over 2.75 million molecular conformations based on drug-like molecules sourced from the <strong>ENAMINE</strong> database. Energies were calculated using <strong>DFT</strong> with the ωB97x density functional and the 6–31G(d) basis set. The conformations were generated from SMILES using RDKit, MMFF94 optimization, and molecular dynamics (MD) simulations, providing a diverse set of molecular structures and energy states.</p> <ul> <li><strong>RDKit Conformations:</strong> 1,123,693</li> <li><strong>RDKit + MMFF94 Optimized:</strong> 1,151,936</li> <li><strong>MD-Generated:</strong> 483,279</li> </ul> <p>This dataset serves as a benchmark for energy prediction models, with training (638,617 examples), validation (134,732 examples), and test subsets (24,890 examples) created using a strict scaffold-based split to ensure no overlap and less than 70% similarity between the training and test sets.</p> <p>Dataset contents:</p> <ul> <li><em>data.tar.gz</em>: contains the conformations in Structured Data File format, grouped into separate folders based on the molecule ID. Each conformation's label is provided within its SDF file as a property named "energy".</li> <li><em>INDEX.smi</em>: specifies the molecule IDs and their corresponding SMILES.</li> <li><em>SOURCES.csv</em>: specifies the conformation generation method for each conformation.</li> <li><em>SDDF_train.tsv</em>, <em>SDDF_validation.tsv</em>, and <em>SDDF_test.tsv</em> specify the molecule IDs and conformations for each subset of the benchmark.</li> </ul> <p>A detailed description is provided in the accompanying paper.</p>
ShareScore
44/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 8
- Harmonization
- 4
- Access
- 20
- Reuse readiness
- 8
- Engagement
- 4