Research Appraisalother

Explainable machine learning in healthcare: methods, interpretation, and applications for clinical research

Journal of the American Medical Informatics Association : JAMIAPadmanabhan, Krishna, Lu, Minxin, Feng, Dai et al.1 Aug 2026DOI

Clinical Snapshot

45CEBM
Evidence: Moderateother

PICO Framework

P — PopulationClinical researchers, data scientists, and healthcare professionals seeking to apply or interpret machine learning models in clinical research and decision support contexts
I — InterventionExplainable machine learning (XML) methodologies including SHAP, LIME, Partial Dependence Plots (PDP), and Individual Conditional Expectation (ICE) plots applied to clinical datasets
C — ComparatorTraditional (black-box) predictive machine learning models without interpretability frameworks; no formal active comparator group
O — OutcomesInterpretability, transparency, and clinical applicability of ML model outputs; ability to identify nonlinear relationships, interaction effects, and patient-level heterogeneity in predicted risk

Bottom Line

This structured primer from Padmanabhan and colleagues provides a clinically accessible introduction to explainable machine learning (XML) methods — specifically SHAP, LIME, PDP, and ICE plots — for healthcare researchers seeking to interpret and communicate ML model outputs. The paper's principal contribution is educational: it bridges a genuine gap between data science methodology and clinical applicability, offering structured guidance on when and how to use each XML approach. Worked examples using the UCI Heart Disease dataset effectively illustrate global versus local interpretability concepts. However, the paper has important limitations that senior clinicians should recognise. It is a narrative primer, not a systematic review, and lacks a reproducible search strategy. The underlying ML model for worked examples is incompletely specified, and no uncertainty estimates are provided for XML outputs — a significant omission given known instability issues with methods such as LIME. Critically, the paper does not adequately address the risk of misinterpreting XML feature importance as causal evidence, which could lead to flawed clinical reasoning. Industry affiliations among authors warrant transparency. For Australian clinicians and researchers, this paper is a useful starting point for understanding XML in the context of TGA SaMD frameworks and digital health governance, but should be supplemented with more rigorous methodological literature before clinical deployment of XML-informed decision tools.

Evidence: Moderate

Key Findings

  • P Value: Not reported — no formal hypothesis testing conducted

  • Effect Size: Not applicable — no formal effect sizes reported; results are illustrative and qualitative in nature

  • Primary Outcome: Demonstration that XML techniques (SHAP, LIME, PDP, ICE) provide interpretable visual and quantitative insights into ML model predictions using the UCI Heart Disease dataset, distinguishing global (population-level) from local (patient-level) feature contributions

  • Nnt Or Sensitivity: Not applicable — this is a methodological primer; no diagnostic accuracy, therapeutic efficacy, or prognostic hazard metrics are reported

  • Confidence Interval: Not reported — no uncertainty estimates or confidence intervals provided for any XML output

Clinical Application

The XML methods described (SHAP, LIME, PDP, ICE) are implemented in widely available open-source Python and R libraries (e.g., shap, lime, pdpbox, iml), making adoption technically feasible for research teams with data science capacity. However, clinical implementation at the point of care requires additional infrastructure, validation, and governance frameworks not addressed in this paper. Computational cost of SHAP for large datasets or complex ensemble models may be prohibitive without optimisation. In Australia, the Therapeutic Goods Administration (TGA) has issued guidance on Software as a Medical Device (SaMD), which encompasses ML-based clinical decision support tools. XML methods are directly relevant to TGA requirements for transparency and explainability in regulated ML applications. The Australian Digital Health Agency's National Digital Health Strategy emphasises trustworthy AI in healthcare, aligning with the paper's objectives. RACGP and specialist colleges have not yet issued specific guidance on XML in clinical practice, representing a gap this paper could help address. The Medical Board of Australia's position on AI-assisted clinical decision-making underscores the need for clinician understanding of model outputs — precisely the gap XML tools aim to fill. PBS implications are indirect but relevant: ML models informing prescribing or resource allocation decisions would benefit from XML transparency to satisfy PBAC evidentiary standards. Clinical researchers, biostatisticians, and clinician-scientists working with ML-based predictive models in any clinical domain, particularly those seeking to communicate model outputs to clinical audiences or regulatory bodies. Most directly applicable to researchers in cardiology, oncology, critical care, and chronic disease management where ML risk models are increasingly deployed.

Abstract

OBJECTIVES: To provide a practical and methodologically grounded overview of explainable machine learning (XML) approaches in healthcare, with emphasis on their interpretation and application in clinical research and decision support. By moving beyond traditional predictive models, this primer aims to foster trust, transparency, and informed clinical decision-making, ultimately bridging the gap between data science and medical practice. MATERIALS AND METHODS: We present a structured review of commonly used XML methodologies, including global and local interpretability tools such as SHapley Additive exPlanations (SHAP), Local Interpretable Model-Agnostic Explanations (LIME), Partial Dependence Plots (PDP), and Individual Conditional Expectation (ICE) plots. For each method, we explain the underlying mechanism at a high level, visualize representative outputs, and provide structured guidance on interpretation, appropriate use, and limitations, illustrated using the publicly available Heart Disease dataset. RESULTS: XML techniques provided intuitive visual and quantitative insights into how predictors influence model predictions. Global methods characterized population-level feature effects, whereas local methods revealed patient-level contributions useful for individualized interpretation. Our worked examples demonstrate how XML outputs can identify nonlinear relationships, detect interaction effects, and reveal heterogeneity in predicted risk across patients, addressing key challenges in translating ML predictions into interpretable outputs for clinical research. DISCUSSION AND CONCLUSION: XML tools offer valuable interpretability for ML models and support more transparent and accountable ML applications in clinical research. By providing a methodologically grounded overview alongside practical implementation examples and structured guidance on each method's strengths and limitations, this primer helps bridge the gap between advanced ML methodology and clinical applicability. Thoughtful adoption of XML approaches may facilitate better understanding, communication, and critical evaluation of ML predictions in healthcare research, ultimately supporting evidence-based clinical decision-making.

References

  1. 1.Padmanabhan, K., Lu, M., Feng, D., Kan-Dobrosky, N., Konduri, S., Litman, H. J., & Livieratos, A. (2026). Explainable machine learning in healthcare: methods, interpretation, and applications for clinical research. Journal of the American Medical Informatics Association: JAMIA. https://doi.org/10.1093/jamia/ocag077
Share:XLinkedIn

This content is for educational purposes for healthcare professionals only and does not constitute clinical advice. Clinical decisions should be based on individual patient assessment, current guidelines, and appropriate specialist consultation. Editorial Standards · Privacy Policy · Terms of Service