Benchmarks & Property Prediction

ADMET property prediction on the TDC benchmarks and evaluation of CHEESE search against state-of-the-art methods.

Suggest an edit

We used the learned molecule representations to train ADMET property prediction models. The newly trained models enable real-time property prediction. We also evaluated our fine-tuned models for property prediction using the TDC (Therapeutics Data Commons) ADMET Benchmarks, and they stand fairly well in most of the tasks.

TDC ADMET benchmarks

BenchmarkMetricOur scoreBest score
CYP2C9_VeithAUPRC ↑0.5890.794
CYP2D6_VeithAUPRC ↑0.4300.721
CYP3A4_VeithAUPRC ↑0.6930.882
CYP2D6_substrateAUPRC ↑0.4670.711
CYP3A4_substrateAUPRC ↑0.6130.680
CYP2C9_substrateAUPRC ↑0.2820.437
Bioavailability_MaAUROC ↑0.6310.748
HIA_HouAUROC ↑0.8700.988
Pgp_BroccatelliAUROC ↑0.8250.994
BBB_MartinsAUROC ↑0.7630.923
hERGAUROC ↑0.7730.875
AMESAUROC ↑0.7190.865
DILIAUROC ↑0.8550.937
Caco2_WangMAE ↓0.3770.285
Lipophilicity_AstraZenecaMAE ↓0.5950.533
Solubility_AqSolDBMAE ↓0.8830.727
PPBR_AZMAE ↓9.7398.251
LD50_ZhuMAE ↓0.6400.588

Evaluation against SOTA

Comparison of our models against state-of-the-art molecular search on enumerative databases called Smallworld (used in ZINC22) and random molecule retrieval. 100 search queries with 30 results were performed using each method.

Our models are clearly better than random search since they are able to retrieve much more similar molecules. They are also beating Smallworld on 2D Morgan and 3D Electrostatic search.

On this page