Comparing clinical decision-making between colposcopists and large language models in cervical dysplasia management: a pilot prospective multicenter study
Clinical Snapshot
PICO Framework
| P — Population | Board-certified colposcopists and two commercially available large language models (ChatGPT-4o and ChatGPT-5) evaluated against 23 anonymised real-life patient cases of cervical dysplasia |
| I — Intervention | Clinical decision-making by ChatGPT-4o and ChatGPT-5 (LLMs) responding to multiple-choice treatment decision questions |
| C — Comparator | Clinical decision-making by ten board-certified colposcopists, with a gold standard defined by two guideline authors |
| O — Outcomes | Concordance rates with guideline-defined gold standard answers across all cases and histopathological subgroups (CIN, unspecific histopathology, cervical cancer) |
Bottom Line
This small pilot study (23 cases, 10 colposcopists) found broadly similar overall concordance rates between board-certified colposcopists and ChatGPT-4o (both 69.6%) against a guideline-defined gold standard in cervical dysplasia management, with ChatGPT-5 performing slightly lower (65.2%). Subgroup analyses suggest LLMs may perform comparably to clinicians in straightforward precancerous lesion cases but fall short in complex or ambiguous scenarios requiring nuanced clinical judgement. Clinicians showed a tendency to overtreat CIN I, which may represent a genuine quality improvement opportunity. However, these findings must be interpreted with extreme caution: the sample is too small for reliable conclusions, no statistical testing was performed, confidence intervals are absent, and the multiple-choice format likely inflates LLM performance relative to real-world open-ended clinical reasoning. The risk of LLM training data contamination with the guideline content used as gold standard is unaddressed. For Australian clinicians, the study's German guideline basis limits direct applicability. This is a hypothesis-generating pilot only. LLMs should not be used as clinical decision support in cervical dysplasia management without prospective validation against local guidelines, regulatory clearance, and appropriate clinical governance.
Key Findings
P Value: Not reported
Effect Size: Overall concordance: clinicians 69.6%, ChatGPT-4o 69.6%, ChatGPT-5 65.2%. Subgroup: ChatGPT-5 outperformed clinicians in precancerous lesions (81.8% vs. 66.4%); clinicians outperformed LLMs in unspecific histopathology cases (86% vs. 60%)
Primary Outcome: Concordance rate with guideline-defined gold standard treatment decisions across 23 cervical dysplasia cases
Nnt Or Sensitivity: Not applicable to this comparative accuracy design; no sensitivity, specificity, or diagnostic accuracy metrics reported
Confidence Interval: Not reported
Clinical Application
LLMs are commercially available and accessible, but integration into clinical workflows requires validation against local guidelines, regulatory approval, and robust governance frameworks. Multiple-choice prompting as tested here does not reflect real-world clinical complexity. Australian cervical screening follows the National Cervical Screening Program (NCSP) renewed in 2017, with management guided by NHMRC-endorsed guidelines that differ from German protocols used in this study. The Therapeutic Goods Administration (TGA) has not approved any LLM as a clinical decision support device in this context. The RACGP and RANZCOG would require prospective validation against Australian guidelines before any clinical implementation. The finding that clinicians overtreated CIN I is relevant to Australian practice, where conservative surveillance of low-grade lesions is guideline-recommended. PBS does not fund LLM-based decision support tools. Board-certified colposcopists managing patients with screen-detected cervical abnormalities, particularly in settings where decision support tools for low-grade lesion management are being considered
Abstract
PURPOSE: This prospective multicenter study aimed to compare the decision-making abilities of board-certified colposcopists and two commercially available large language models (LLM), ChatGPT-4o and ChatGPT-5, in cervical dysplasia management. METHODS: Twenty-three anonymized real-life patient cases with multiple-choice (MC) questions regarding treatment decisions were used to assess answer quality. Ten board-certified colposcopists and the two LLMs addressed the MC questions. The gold standard was defined by two guideline authors. LLMs were prompted to justify their responses. Concordance rates were calculated and compared across all questions and histopathological subgroups, including cervical intraepithelial neoplasia (CIN), unspecific histopathological results, and cervical cancer cases. RESULTS: Clinicians and LLMs achieved similar overall concordance rates compared to the gold standard (69.6% for clinicians, 69.6% for ChatGPT-4o, and 65.2% for ChatGPT-5). ChatGPT-5 outperformed clinicians in precancerous lesions (81.8% vs. 66.4%), while clinicians excelled in complex cases with unspecific histopathology (86% vs. 60%). Clinicians showed a tendency to overtreat low-grade lesions (CIN I), opting for more intensive surveillance. ChatGPT-4o performed better than ChatGPT-5 in cervical cancer cases, though both models struggled with these scenarios. CONCLUSION: This study highlights the potential of LLMs as decision support tools in cervical dysplasia management, particularly for straightforward cases like precancerous lesions. However, clinicians remain superior in handling complex or ambiguous cases. The tendency of clinicians to overtreat low-grade lesions may offer the potential to test the implementation of a decision support tool for those cases. While LLMs show promise, exploring open-ended clinical scenarios and integrating retrieval-augmented generation could enhance their practical application.
References
- 1.Stalp, J. L., Schneider, J. A., Steinkasserer, L., Hachenberg, J., Jentschke, M., Hillemanns, P., Wolff, D., & Denecke, A. (2026). Comparing clinical decision-making between colposcopists and large language models in cervical dysplasia management: a pilot prospective multicenter study. Archives of Gynecology and Obstetrics. https://doi.org/10.1007/s00404-025-08123-4
Related Research
Histochemistry and cell biology
Molecular plasticity of LAMA3 across the disease spectrum: pathogenic mechanisms and clinical translation.
29 July 2026
Journal of robotic surgery
Robotic versus laparoscopic and open surgery for endometrial cancer: a systematic review of randomized trials and pooled analysis of conversion rates
21 July 2026
Journal of robotic surgery
Mapping the evolution of deep learning and computer vision in robotic surgery: a bibliometric analysis of surgical video intelligence, instrument perception, and clinical translation.
21 July 2026
This content is for educational purposes for healthcare professionals only and does not constitute clinical advice. Clinical decisions should be based on individual patient assessment, current guidelines, and appropriate specialist consultation. Editorial Standards · Privacy Policy · Terms of Service