Skip to main content
zenodoopen

31 ChEMBL data sets for regression modeling

<p>From ChEMBL version 17, 31 compound data sets have been selected for regression modeling. Compounds had to be active against human targets in a direct inhibition/binding assay with highest ChEMBL confidence score and Ki values below 100 micromolar.&nbsp;Multiple Ki values for the same compound were averaged if they fell into the same order of magnitude, or else they were disregarded. Duplicates,&nbsp;known pan-assay interference, and other reactive molecules were removed.&nbsp;Only sets with at least 500 compounds were considered.</p> <p>&nbsp;</p> <p>Note:&nbsp;The SD files contain a field &quot;pKi&quot;; note however that this field contains the Ki value in nM units, not the logarithmic value.</p>

ShareScore

36/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
8
Harmonization
4
Access
16
Reuse readiness
8
Engagement
0

Topics