Artificial intelligence to revolutionize surgical decision-making is still just around the corner.
Clinical Snapshot
PICO Framework
| P — Population | Surgical patients and clinical settings where AI-based clinical decision support systems (CDSS) are deployed or considered for deployment |
| I — Intervention | Artificial intelligence tools, including large language models (LLMs) and machine learning-based clinical decision support systems, applied to surgical diagnosis, perioperative decision-making, and patient safety |
| C — Comparator | Historical AI/CDSS approaches (e.g., Bayesian systems such as AAPHelp), traditional clinical decision-making without AI augmentation |
| O — Outcomes | Diagnostic accuracy, perioperative decision-making quality, patient safety, postoperative complication prediction, clinical documentation efficiency, transparency, bias, reliability, and workflow integration |
Bottom Line
This narrative review from Harvard Medical School traces the arc of AI-based clinical decision support in surgery from early Bayesian systems through to contemporary large language models. While it provides a readable historical synthesis and appropriately acknowledges persistent challenges — including algorithmic bias, opacity, reliability concerns, and integration difficulties — it falls short of the methodological rigour expected of evidence-based clinical guidance. No systematic search strategy, quality appraisal of primary studies, or quantitative synthesis is presented. The optimistic framing, while intellectually honest in acknowledging limitations, risks overstating readiness for clinical deployment. For Australian surgical teams, this review serves as useful background reading rather than actionable evidence. Before adopting any AI CDSS tool in perioperative practice, clinicians should seek prospective validation studies in comparable healthcare settings, ensure TGA regulatory compliance for SaMD classification, and apply institutional governance frameworks. The field is advancing rapidly, but the evidence base for LLMs specifically in surgical decision-making remains nascent. Senior clinicians should approach vendor claims with scepticism and demand local validation data before implementation.
Key Findings
P Value: Not reported
Effect Size: Not reported — no quantitative effect sizes presented in this narrative review
Primary Outcome: Narrative synthesis of the historical evolution, current capabilities, and persistent limitations of AI-based clinical decision support systems in surgical practice, with particular focus on large language models
Nnt Or Sensitivity: Not applicable — no diagnostic accuracy metrics, NNT, hazard ratios, or sensitivity/specificity values are reported in the abstract; the review is qualitative in nature
Confidence Interval: Not reported
Clinical Application
The review does not provide implementation guidance or feasibility data. Adoption of AI CDSS in surgical practice requires institutional EHR integration, clinician training, governance frameworks, and ongoing performance monitoring — none of which are operationally detailed in this publication In Australia, the TGA regulates AI-based Software as a Medical Device (SaMD) under the Therapeutic Goods (Medical Devices) Regulations 2002, with updated guidance on AI/ML-enabled devices issued from 2023 onward. The Australian Commission on Safety and Quality in Health Care (ACSQHC) has published frameworks on clinical decision support. RACGP and RACS have not yet issued formal position statements on LLM use in surgical decision-making as of mid-2025. PBS implications are indirect — AI tools may influence prescribing and procedural decisions but are not themselves PBS-listed. Australian public hospital EHR heterogeneity (varying systems across states) presents a significant barrier to uniform AI CDSS deployment. The review's US-centric perspective limits direct applicability without local validation studies. Surgical clinicians, perioperative teams, hospital administrators, and health informaticians considering adoption or evaluation of AI-based CDSS tools in surgical settings
Abstract
Medical artificial intelligence, especially large language models, has engendered both excitement and unease across the medical community, promising improved surgical diagnostic accuracy, perioperative decision-making, and patient safety amidst concerns of considerable bias and its implications on the future of human expertise. The evolution of artificial intelligence clinical decision support originated in the surgical field through early systems like AAPHelp, which leveraged Bayesian reasoning to diagnose acute abdominal pain. Despite the initial promise of artificial intelligence clinical decision support, attempts to develop more robust diagnostic tools largely fell short in the 1980s, with instruments unable to adequately diagnose and manage complex presentations, a lack of transparency in their decision-making processes, and limitations in scope. It was not until the 2010s that clinical decision support began to show renewed potential as a meaningful adjunct in surgical practice, largely driven by advancements in machine learning techniques and wider access to large clinical data sets. New artificial intelligence models, particularly those utilizing large language models, now demonstrate impressive capabilities, from predicting the risk of postoperative complications for individual patients to streamlining clinical documentation and beyond. However, challenges persist regarding transparency, bias, reliability, and integration into clinical workflows. Despite these hurdles, artificial intelligence-based tools like large language models represent an exciting new chapter in surgical decision-making. Built on a long history of clinical decision support systems in surgery, these technologies hold great promise to meaningfully augment surgical practice and improve care for patients.
References
- 1.Zhang, S. K., & Rodman, A. (2026). Artificial intelligence to revolutionize surgical decision-making is still just around the corner. Surgery. https://doi.org/10.1016/j.surg.2026.110326
Related Research
Minimally invasive therapy & allied technologies : MITAT : official journal of the Society for Minimally Invasive Therapy
Ventral hernia repair in emergency settings. A machine learning model to predict post-operative complications.
3 Aug 2026
Journal of pediatric surgery
Interpretable deep learning model for pediatric strangulated small bowel obstruction on CT: A multicenter study
3 Aug 2026
European journal of radiology
A CT-based deep learning model to differentiate between benign and malignant adrenal lesions
3 Aug 2026
This content is for educational purposes for healthcare professionals only and does not constitute clinical advice. Clinical decisions should be based on individual patient assessment, current guidelines, and appropriate specialist consultation. Editorial Standards · Privacy Policy · Terms of Service