Benchmarks & Property Prediction
ADMET property prediction on the TDC benchmarks and evaluation of CHEESE search against state-of-the-art methods.
We used the learned molecule representations to train ADMET property prediction models. The newly trained models enable real-time property prediction. We also evaluated our fine-tuned models for property prediction using the TDC (Therapeutics Data Commons) ADMET Benchmarks, and they stand fairly well in most of the tasks.
TDC ADMET benchmarks
| Benchmark | Metric | Our score | Best score |
|---|---|---|---|
| CYP2C9_Veith | AUPRC ↑ | 0.589 | 0.794 |
| CYP2D6_Veith | AUPRC ↑ | 0.430 | 0.721 |
| CYP3A4_Veith | AUPRC ↑ | 0.693 | 0.882 |
| CYP2D6_substrate | AUPRC ↑ | 0.467 | 0.711 |
| CYP3A4_substrate | AUPRC ↑ | 0.613 | 0.680 |
| CYP2C9_substrate | AUPRC ↑ | 0.282 | 0.437 |
| Bioavailability_Ma | AUROC ↑ | 0.631 | 0.748 |
| HIA_Hou | AUROC ↑ | 0.870 | 0.988 |
| Pgp_Broccatelli | AUROC ↑ | 0.825 | 0.994 |
| BBB_Martins | AUROC ↑ | 0.763 | 0.923 |
| hERG | AUROC ↑ | 0.773 | 0.875 |
| AMES | AUROC ↑ | 0.719 | 0.865 |
| DILI | AUROC ↑ | 0.855 | 0.937 |
| Caco2_Wang | MAE ↓ | 0.377 | 0.285 |
| Lipophilicity_AstraZeneca | MAE ↓ | 0.595 | 0.533 |
| Solubility_AqSolDB | MAE ↓ | 0.883 | 0.727 |
| PPBR_AZ | MAE ↓ | 9.739 | 8.251 |
| LD50_Zhu | MAE ↓ | 0.640 | 0.588 |
Evaluation against SOTA
Comparison of our models against state-of-the-art molecular search on enumerative databases called Smallworld (used in ZINC22) and random molecule retrieval. 100 search queries with 30 results were performed using each method.
Our models are clearly better than random search since they are able to retrieve much more similar molecules. They are also beating Smallworld on 2D Morgan and 3D Electrostatic search.