Research Appraisalobservational

Protein Language Model-Based Fitness Estimates Facilitate Resistance Mutation Identification

Journal of chemical information and modelingSchwarz, Dominik, Giese, Sven H, Gupta, Akansha et al.27 July 2026DOI

Clinical Snapshot

55CEBM
Evidence: Weakobservational

PICO Framework

P — PopulationCancer-relevant protein targets (ERK2 and EGFR exon20 mutant) subjected to deep mutational scanning (DMS) experiments; by extension, oncology drug development pipelines targeting kinase inhibitor resistance
I — InterventionCombined computational pipeline integrating physics-based free energy perturbation (FEP) affinity estimates with protein language model (PLM)-based protein fitness estimates to identify resistance-conferring amino acid mutations in silico
C — ComparatorFEP affinity estimates alone (without PLM-based fitness filtering) as the baseline in silico resistance prediction approach
O — OutcomesAccuracy of in silico identification of resistance mutations validated against deep mutational scanning (DMS) experimental data; reduction of false-positive resistance predictions (mutations flagged as affinity-decreasing by FEP but non-resistant in DMS due to fitness cost)

Bottom Line

This computational study from Bayer AG researchers proposes combining physics-based free energy perturbation (FEP) with protein language model (PLM)-based fitness scoring to improve in silico identification of drug resistance mutations in cancer targets. The core insight is methodologically sound: FEP can flag mutations that reduce drug binding affinity, but some of these mutations are biologically non-viable because they also impair protein function — PLM fitness scores can filter these false positives efficiently. Validation against deep mutational scanning data for ERK2 and an EGFR exon20 mutant demonstrates proof-of-concept. However, the study has significant limitations: only two benchmark systems are tested, one dataset is proprietary to Bayer, no formal statistical performance metrics or confidence intervals are reported, and there is no prospective clinical validation. The tool is a drug discovery research instrument, not a clinical decision support system. For Australian oncologists, the EGFR exon20 context is clinically relevant given PBS-listed EGFR inhibitor therapies, but this study does not yet provide actionable clinical guidance. Independent replication across diverse targets, rigorous statistical benchmarking, and prospective validation in patient-derived samples are required before this approach can influence clinical trial design or resistance monitoring practice.

Evidence: Weak

Key Findings

  • P Value: Not reported

  • Effect Size: Not quantified in abstract; direction of effect is positive (improved specificity of resistance identification when PLM filter is applied to FEP predictions)

  • Primary Outcome: PLM-based protein fitness estimates successfully filtered out mutations flagged as affinity-decreasing by FEP (and thus potentially resistant) that were confirmed as non-resistant in DMS experiments, reducing false-positive resistance predictions

  • Nnt Or Sensitivity: Sensitivity, specificity, and positive/negative predictive values for resistance mutation identification are not reported in the abstract; no NNT applicable (computational tool validation study)

  • Confidence Interval: Not reported

Clinical Application

The pipeline requires access to protein structural data (for FEP), computational infrastructure for FEP simulations, and PLM inference capability. PLM-based fitness scoring is described as computationally efficient relative to explicit stability modelling. Implementation in routine clinical practice is not currently feasible; this is a drug discovery research tool. Prospective clinical validation would be required before any clinical translation. ERK2 inhibitors are not currently PBS-listed in Australia. EGFR-targeted therapies (e.g., osimertinib, afatinib, erlotinib) are PBS-listed for NSCLC with specific EGFR mutations (PBS items vary by line of therapy and mutation status). The EGFR exon20 insertion mutation space is clinically relevant in Australia, with amivantamab and mobocertinib under TGA/PBS evaluation pathways. RACGP and Cancer Australia guidelines emphasise molecular profiling of NSCLC prior to systemic therapy. If validated prospectively, this computational approach could inform companion diagnostic development and resistance surveillance protocols aligned with Australian Genomics and the Australian Tumour Mutation Consortium frameworks. No immediate PBS or TGA regulatory implications arise from this computational study. Primarily applicable to pharmaceutical drug discovery teams developing small molecule kinase inhibitors where structural binding mode data are available; indirectly relevant to oncologists and clinical trialists designing resistance monitoring strategies for ERK2- or EGFR-targeted therapies in solid tumours

Abstract

Drug resistance is a major challenge in cancer therapy. Cancer cells with pre-existing or acquired mutations that confer resistance to a given drug treatment outgrow the susceptible cell population and cause cancer recurrence after an initial successful treatment response. Knowledge about resistance mutations before they occur in the clinic could prevent unnecessary patient treatment with ineffective drugs, in clinical trials as well as clinical practice, or potentially speed up the development of follow-up compounds. Here, we focused on on-target amino acid mutations that confer resistance to an inhibitor compound with a known binding mode. We evaluated whether a combination of physics-based free energy perturbation (FEP) affinity estimates and protein language model-based protein fitness estimates could improve the in silico identification of resistance mutations. Validation was done with data from deep mutational scanning (DMS) experiments that tested for resistance to single amino acid mutations. A public data set testing ERK2 resistance against the inhibitor SCH772984 and an internal data set testing resistance of an EGFR_exon20 mutant against a Bayer small molecule inhibitor were used. Our results show that protein fitness estimates can facilitate the identification of resistance mutations by filtering mutations with a low estimated fitness. Even though FEP has flagged such mutations as affinity-decreasing and thus potentially resistant, they were not resistant according to the DMS experiment and therefore correctly filtered out. This indicates that protein language model-based protein fitness estimates could be a computationally efficient method to filter mutations without having to model the negative impact of mutations on native function or protein stability, which is error-prone and computationally expensive.

References

  1. 1.Schwarz, D., Giese, S. H., Gupta, A., Dziubańska-Kusibab, P. J., Villalba, S. D., Cherniack, A., Yang, X., Root, D., Karsli-Uzunbas, G., Greulich, H., Siegel, F., Kamburov, A., Nevedomskaya, E., Christ, C., Aldeghi, M., & Mortier, J. (2026). Protein language model-based fitness estimates facilitate resistance mutation identification. Journal of Chemical Information and Modeling. https://doi.org/10.1021/acs.jcim.6c00768
Share:XLinkedIn

This content is for educational purposes for healthcare professionals only and does not constitute clinical advice. Clinical decisions should be based on individual patient assessment, current guidelines, and appropriate specialist consultation. Editorial Standards · Privacy Policy · Terms of Service