Research Appraisalobservational

Comparing clinical decision-making between colposcopists and large language models in cervical dysplasia management: a pilot prospective multicenter study

Archives of gynecology and obstetricsStalp, Jan Lennart, Schneider, Juliane Alexandra, Steinkasserer, Lena et al.18 June 2026DOI

Clinical Snapshot

55CEBM
Evidence: Weakobservational

PICO Framework

P — PopulationBoard-certified colposcopists and two commercially available large language models (ChatGPT-4o and ChatGPT-5) evaluated against 23 anonymised real-life patient cases of cervical dysplasia
I — InterventionClinical decision-making by ChatGPT-4o and ChatGPT-5 (LLMs) responding to multiple-choice treatment decision questions
C — ComparatorClinical decision-making by ten board-certified colposcopists, with a gold standard defined by two guideline authors
O — OutcomesConcordance rates with guideline-defined gold standard answers across all cases and histopathological subgroups (CIN, unspecific histopathology, cervical cancer)

Bottom Line

This small pilot study (23 cases, 10 colposcopists) found broadly similar overall concordance rates between board-certified colposcopists and ChatGPT-4o (both 69.6%) against a guideline-defined gold standard in cervical dysplasia management, with ChatGPT-5 performing slightly lower (65.2%). Subgroup analyses suggest LLMs may perform comparably to clinicians in straightforward precancerous lesion cases but fall short in complex or ambiguous scenarios requiring nuanced clinical judgement. Clinicians showed a tendency to overtreat CIN I, which may represent a genuine quality improvement opportunity. However, these findings must be interpreted with extreme caution: the sample is too small for reliable conclusions, no statistical testing was performed, confidence intervals are absent, and the multiple-choice format likely inflates LLM performance relative to real-world open-ended clinical reasoning. The risk of LLM training data contamination with the guideline content used as gold standard is unaddressed. For Australian clinicians, the study's German guideline basis limits direct applicability. This is a hypothesis-generating pilot only. LLMs should not be used as clinical decision support in cervical dysplasia management without prospective validation against local guidelines, regulatory clearance, and appropriate clinical governance.

Evidence: Weak

Key Findings

  • P Value: Not reported

  • Effect Size: Overall concordance: clinicians 69.6%, ChatGPT-4o 69.6%, ChatGPT-5 65.2%. Subgroup: ChatGPT-5 outperformed clinicians in precancerous lesions (81.8% vs. 66.4%); clinicians outperformed LLMs in unspecific histopathology cases (86% vs. 60%)

  • Primary Outcome: Concordance rate with guideline-defined gold standard treatment decisions across 23 cervical dysplasia cases

  • Nnt Or Sensitivity: Not applicable to this comparative accuracy design; no sensitivity, specificity, or diagnostic accuracy metrics reported

  • Confidence Interval: Not reported

Clinical Application

LLMs are commercially available and accessible, but integration into clinical workflows requires validation against local guidelines, regulatory approval, and robust governance frameworks. Multiple-choice prompting as tested here does not reflect real-world clinical complexity. Australian cervical screening follows the National Cervical Screening Program (NCSP) renewed in 2017, with management guided by NHMRC-endorsed guidelines that differ from German protocols used in this study. The Therapeutic Goods Administration (TGA) has not approved any LLM as a clinical decision support device in this context. The RACGP and RANZCOG would require prospective validation against Australian guidelines before any clinical implementation. The finding that clinicians overtreated CIN I is relevant to Australian practice, where conservative surveillance of low-grade lesions is guideline-recommended. PBS does not fund LLM-based decision support tools. Board-certified colposcopists managing patients with screen-detected cervical abnormalities, particularly in settings where decision support tools for low-grade lesion management are being considered

Abstract

PURPOSE: This prospective multicenter study aimed to compare the decision-making abilities of board-certified colposcopists and two commercially available large language models (LLM), ChatGPT-4o and ChatGPT-5, in cervical dysplasia management. METHODS: Twenty-three anonymized real-life patient cases with multiple-choice (MC) questions regarding treatment decisions were used to assess answer quality. Ten board-certified colposcopists and the two LLMs addressed the MC questions. The gold standard was defined by two guideline authors. LLMs were prompted to justify their responses. Concordance rates were calculated and compared across all questions and histopathological subgroups, including cervical intraepithelial neoplasia (CIN), unspecific histopathological results, and cervical cancer cases. RESULTS: Clinicians and LLMs achieved similar overall concordance rates compared to the gold standard (69.6% for clinicians, 69.6% for ChatGPT-4o, and 65.2% for ChatGPT-5). ChatGPT-5 outperformed clinicians in precancerous lesions (81.8% vs. 66.4%), while clinicians excelled in complex cases with unspecific histopathology (86% vs. 60%). Clinicians showed a tendency to overtreat low-grade lesions (CIN I), opting for more intensive surveillance. ChatGPT-4o performed better than ChatGPT-5 in cervical cancer cases, though both models struggled with these scenarios. CONCLUSION: This study highlights the potential of LLMs as decision support tools in cervical dysplasia management, particularly for straightforward cases like precancerous lesions. However, clinicians remain superior in handling complex or ambiguous cases. The tendency of clinicians to overtreat low-grade lesions may offer the potential to test the implementation of a decision support tool for those cases. While LLMs show promise, exploring open-ended clinical scenarios and integrating retrieval-augmented generation could enhance their practical application.

References

  1. 1.Stalp, J. L., Schneider, J. A., Steinkasserer, L., Hachenberg, J., Jentschke, M., Hillemanns, P., Wolff, D., & Denecke, A. (2026). Comparing clinical decision-making between colposcopists and large language models in cervical dysplasia management: a pilot prospective multicenter study. Archives of Gynecology and Obstetrics. https://doi.org/10.1007/s00404-025-08123-4
Share:XLinkedIn

This content is for educational purposes for healthcare professionals only and does not constitute clinical advice. Clinical decisions should be based on individual patient assessment, current guidelines, and appropriate specialist consultation. Editorial Standards · Privacy Policy · Terms of Service