MolXProt: A Cross-Attention Transformer-Based Graph Neural Network for Protein-Ligand Binding Affinity Prediction
Clinical Snapshot
PICO Framework
| P — Population | Protein-ligand pairs from mixed datasets (>100,000 pairs) |
| I — Intervention | MolXProt transformer-based graph neural network architecture |
| C — Comparator | Traditional docking and empirical scoring methods |
| O — Outcomes | Binding affinity prediction accuracy (±1.0 and ±2.0 kcal/mol thresholds), computational efficiency, interpretability of binding mechanisms |
Bottom Line
MolXProt represents a significant advancement in computational drug discovery, offering a transformer-based approach that achieves 80% accuracy within ±2.0 kcal/mol for protein-ligand binding affinity prediction. The model's key innovation lies in its bidirectional cross-attention mechanism that integrates graph-based ligand representations with protein language models, enabling interpretation of binding mechanisms while maintaining computational efficiency. The architecture successfully scales to large datasets and provides insights into key binding pocket residues, though it requires post-hoc calibration to address prediction bias. For Australian pharmaceutical research, this tool could accelerate early-stage drug discovery by providing rapid, accurate binding predictions that complement traditional experimental approaches. The computational efficiency makes it particularly valuable for high-throughput screening applications, potentially reducing time and costs in drug development pipelines.
Key Findings
P Value: Not reported
Effect Size: 50% of predictions within ±1.0 kcal/mol, 80% within ±2.0 kcal/mol
Primary Outcome: Binding affinity prediction accuracy within specified error thresholds
Nnt Or Sensitivity: 80% sensitivity for predictions within ±2.0 kcal/mol threshold
Confidence Interval: Not explicitly reported
Clinical Application
High feasibility due to computational efficiency and scalability Relevant for Australian pharmaceutical research institutions and TGA drug evaluation processes, supporting local drug discovery initiatives Pharmaceutical researchers and computational biologists involved in drug discovery
Abstract
Accurate and fast prediction of drug-target binding affinities (DTAs) is key for drug discovery; however, many methods, such as docking and empirical scoring, fail when generalizing to unseen cases. In this study, we introduced the MolXProt architecture, a novel transformer-based graph neural network that integrates graph ligand representations with protein language models using bidirectional multihead cross-attention. Our model is shown to be scalable to over 100000 protein-ligand pairs of mixed data sets, achieving 50% of predictions within ±1.0 kcal/mol and 80% within ±2.0 kcal/mol. We show that the architecture can explicitly learn residue-atom interactions while being computationally friendly via protein-token compression. By mapping token-residue interactions, we demonstrated that the model learns key binding pocket residues in benchmark complexes, such as CDK2-Staurosporine and DHFR-Methotrexate, but under-represents the hydrogen-bonding networks. Our calibration bias analysis revealed that the model overpredicted strong binders and underpredicted weak binders, which are linked to data imbalances and heteroscedastic noise. A simple posthoc isotonic correction partially mitigated the bias. Latent space analysis showed that the model learned continuous binding affinity manifolds without split leakage. Our work highlights a novel architecture that offers unique insights into binding mechanisms via transformer-based cross-attention and is computationally inexpensive.
References
- 1.Cucco, B. (2026). MolXProt: A Cross-Attention Transformer-Based Graph Neural Network for Protein-Ligand Binding Affinity Prediction. Journal of Chemical Theory and Computation. https://doi.org/10.1021/acs.jctc.6c00026
Related Research
Current opinion in chemical biology
Prediction of protein-protein interactions and co-complex models with deep learning
3 Aug 2026
Journal of chemical information and modeling
ProphDR: An Interpretable Deep Learning Model for Predicting Cancer Drug Response via Multi-Omics and Cross-Attention Mechanisms
28 July 2026
PloS one
Multi-view graph-regularized deep metric subspace clustering network
26 July 2026
This content is for educational purposes for healthcare professionals only and does not constitute clinical advice. Clinical decisions should be based on individual patient assessment, current guidelines, and appropriate specialist consultation. Editorial Standards · Privacy Policy · Terms of Service