Binding Affinity Prediction with 3D Machine Learning: Training Data and Challenging External Testing

Jose Carlos Gómez-Tamayo
Lili Cao
Mazen Ahmad
Gary Tresadern

Read the full article

Discuss this preprint

Start a discussion What are Sciety discussions?

Listed in

This article is not in any list yet, why not save it to one of your lists.

Abstract

Protein-ligand binding affinity prediction is one of the major challenges in computational assisted drug discovery. An active area of research uses machine learning (ML) models trained on 3D structures of protein ligand complexes to predict binding modes, discriminate active and inactives, or predict affinity. Methodological advances in deep learning, and artificial intelligence along with increased experimental data (3D structures and bioactivities) has led to many studies using different architectures, representation, and features. Unfortunately, many models do not learn details of interactions or the underlying physics that drive protein-ligand affinity, but instead just memorize patterns in the available training data with poor generalizability and future use. In this work we incorporate “dense”, feature rich datasets that contain up to several thousand analogue molecules per drug discovery target. For the training set, PDBbind dataset is used with enrichment from 8 internal lead optimization (LO) datasets and inactive and decoy poses in a variety of combinations. A variety of different model architectures was used and the model performance was validated using the binding affinity for 12 internal LO and 6 ChEMBL external test sets. Results show a significant improvement in the performance and generalization power, especially for virtual screening and suggest promise for the future of ML protein-ligand affinity prediction with a greater emphasis on training using datasets that capture the rich details of the affinity landscape.

Version published to 10.21203/rs.3.rs-3969529/v1 on Research Square
Mar 5, 2024

Feature-Optimized Machine Learning Benchmarking for Protein Interface Prediction in Permanent Homodimer Complexes with Distinct Structural Features

This article has 4 authors:
1. Tayyip Topuz
2. Zeki Erdem
3. Halil Bisgin
4. E. Demet Akten
This article has no evaluationsLatest version Feb 2, 2026
Multi-Modal Ensemble Learning for TLR4 Binding Prediction: Addressing Data Scarcity and Leakage in Small Molecule Drug Discovery

This article has 3 authors:
1. Brandon Yee
2. Maximilian Rutkowski
3. Wilson Collins
This article has no evaluationsLatest version Jan 28, 2026
Integrating Evolutionary and Compositional Features with ML and DL for Robust and Interpretable Druggable Protein Prediction

This article has 5 authors:
1. Mujeebu Rehman
2. Qinghua Liu
3. Muhammad Javed
4. Ali Ghulam
5. Teerath Kumar
This article has no evaluationsLatest version Dec 11, 2025

Discuss this preprint

Listed in

Abstract

Article activity feed

Related articles

Feature-Optimized Machine Learning Benchmarking for Protein Interface Prediction in Permanent Homodimer Complexes with Distinct Structural Features

Multi-Modal Ensemble Learning for TLR4 Binding Prediction: Addressing Data Scarcity and Leakage in Small Molecule Drug Discovery

Integrating Evolutionary and Compositional Features with ML and DL for Robust and Interpretable Druggable Protein Prediction