Research AppraisalSystematic Review

Artificial intelligence language models for medical text analysis: A systematic review

Artificial intelligence in medicineSorayaie Azar, Amir, Bagherzadeh Mohasefi, Jamshid, Wiil, Uffe Kock et al.1 Aug 2026DOI

Clinical Snapshot

20CEBM
Evidence: WeakSystematic Review

PICO Framework

P — PopulationMedical text records and clinical documentation datasets (electronic health records, clinical notes, discharge summaries, and other medical textual data sources)
I — InterventionArtificial intelligence language models, including transformer-based architectures (BERT, GPT) and conventional NLP/ML approaches applied to medical text analysis
C — ComparatorConventional NLP and machine learning approaches (e.g., rule-based systems, traditional statistical models, non-transformer ML classifiers)
O — OutcomesPerformance in disease classification, automated clinical documentation, predictive analytics, and clinical decision support — assessed via model accuracy, F1-score, AUC, and related metrics

Bottom Line

This systematic review synthesises 22 studies examining AI language models — particularly BERT and GPT architectures — for medical text analysis tasks including disease classification, clinical documentation, and predictive analytics. The overarching finding is that transformer-based models outperform conventional NLP approaches, but this conclusion rests on a narrow, heterogeneous evidence base without quantitative pooling, formal risk of bias assessment, or GRADE certainty grading. The review's methodological rigour is insufficient to support practice-changing recommendations. Critical barriers to clinical adoption — including limited external validation, non-standardised preprocessing, and absent interpretability frameworks — are acknowledged by the authors themselves. For senior clinicians and health system leaders, this review is best understood as a landscape mapping exercise rather than definitive evidence for implementation. Institutions considering AI-assisted clinical text analysis should require locally validated performance data, regulatory compliance under TGA SaMD frameworks, and robust clinical governance structures before deployment. The field is evolving rapidly, and higher-quality prospective evaluations with patient-centred outcomes are urgently needed before AI language models can be confidently integrated into routine clinical workflows.

Evidence: Weak

Key Findings

  • P Value: Not reported

  • Effect Size: Not quantified; narrative synthesis only — no pooled effect size reported

  • Primary Outcome: Transformer-based AI language models (BERT, GPT) consistently outperform conventional NLP and ML approaches across medical text analysis tasks including disease classification, automated clinical documentation, and predictive analytics

  • Nnt Or Sensitivity: No NNT, sensitivity, specificity, or AUC values are reported in the abstract; individual study metrics are not summarised quantitatively

  • Confidence Interval: Not reported

Clinical Application

Implementation feasibility is currently limited by the barriers identified in the review itself: lack of externally validated, clinically representative datasets; absence of interpretable AI frameworks acceptable to clinicians and regulators; and variable data preprocessing standards. Deployment in routine clinical settings would require substantial local validation, governance frameworks, and clinician training. In Australia, AI language models applied to medical text must navigate TGA regulatory pathways for software as a medical device (SaMD) under the Digital Health Strategy. The Australian Digital Health Agency's national EHR infrastructure (My Health Record) uses Australian-specific clinical terminology and ICD-10-AM coding, which may not align with training corpora of US- or European-derived models. RACGP and ACHS standards for clinical documentation would need to be incorporated into any locally deployed system. PBS implications are indirect but relevant if AI-assisted diagnostics influence prescribing decisions. No Australian-specific studies are identifiable from the abstract, limiting direct applicability. The Australian Commission on Safety and Quality in Health Care's guidance on clinical AI governance would be essential for any implementation pathway. Healthcare systems seeking to implement AI-assisted clinical documentation, disease classification from EHR text, or NLP-based predictive analytics. Most directly applicable to tertiary hospital settings with structured EHR infrastructure and informatics capacity.

Abstract

Medical text records serve as essential repositories of patient information, providing a foundation for informed clinical decision-making, accurate diagnosis, reliable prognosis, and effective treatment planning. Recent advancements in Artificial Intelligence (AI), particularly in Natural Language Processing (NLP) and Machine Learning (ML), have positioned AI-driven language models as powerful tools for analyzing, classifying, and generating medical textual data. In this systematic literature review, an initial search retrieved 548 records published between 1 January 2000 and 1 July 2024. After rigorous screening based on predefined inclusion and exclusion criteria, 22 original research articles were included. The review highlights substantial progress in applying advanced architectures such as Bidirectional Encoder Representations from Transformers (BERT) and Generative Pre-trained Transformers (GPT) to medical text processing tasks. These models consistently outperform conventional NLP and ML approaches, achieving superior results in disease classification, automated clinical documentation, and predictive analytics. However, critical challenges persist, including the limited availability of clinically validated datasets, variability in data preprocessing protocols, insufficient external validation, and the lack of interpretable AI frameworks, all of which collectively hinder clinical trust and large-scale adoption. Future research should prioritize the development of hybrid AI systems that integrate multimodal data sources (text, imaging, and structured records), incorporate explainable AI mechanisms, and adhere to standardized reporting frameworks. Addressing these methodological gaps will be pivotal in enhancing the reliability, clinical applicability, and impact of AI language models, thereby advancing evidence-based medicine, personalized treatment strategies, and overall patient care.

References

  1. 1.Sorayaie Azar, A., Bagherzadeh Mohasefi, J., Wiil, U. K., Naemi, A., & Ebrahimi, A. (2026). Artificial intelligence language models for medical text analysis: A systematic review. Artificial Intelligence in Medicine. https://doi.org/10.1016/j.artmed.2026.103441
Share:XLinkedIn

This content is for educational purposes for healthcare professionals only and does not constitute clinical advice. Clinical decisions should be based on individual patient assessment, current guidelines, and appropriate specialist consultation. Editorial Standards · Privacy Policy · Terms of Service