ChatGPT in urogynecology: Comparing large language model responses to human experts
Clinical Snapshot
PICO Framework
| P — Population | Women attending urogynecology outpatient clinics (n=203, median age 56 years, IQR 46–66) |
| I — Intervention | ChatGPT-generated responses to six common urogynecology patient questions |
| C — Comparator | Responses to the same six questions provided by a single consultant urogynecologist |
| O — Outcomes | Patient-rated understandability, helpfulness, and reassurance scored on a 5-point Likert scale per domain (maximum total score 15 per response); preference across six question pairs |
Bottom Line
This exploratory cross-sectional survey of 203 Irish urogynecology patients found that ChatGPT-generated responses to six common clinical questions were rated statistically higher than those from a single consultant urogynecologist across understandability, helpfulness, and reassurance domains. The median score difference was modest (4 points on a 15-point scale) and of uncertain clinical significance. The study carries substantial methodological limitations: a single clinician comparator, unspecified ChatGPT version and prompting methodology, absence of validated outcome instruments, uncertain participant blinding to response source, and no confidence intervals reported. The finding likely reflects differences in response length, readability, and tone rather than superior clinical accuracy. Clinical accuracy was independently verified, which is reassuring, but verification relied on a single reviewer without inter-rater reliability data. These results should not be interpreted as endorsing LLM use as a substitute for clinical consultation. Rather, they suggest a potential adjunctive role for LLMs in patient education within urogynecology — a specialty where stigma limits help-seeking — provided robust governance, version control, and clinician oversight are in place. Replication with validated instruments, multiple clinician comparators, and blinded study designs is needed before practice recommendations can be made.
Key Findings
P Value: p < 0.01 for total score and all three individual domains (understandability, helpfulness, reassurance)
Effect Size: ChatGPT median total score 76 (IQR 67–85) versus consultant median 72 (IQR 63–80); 4-point median difference on a 15-point scale. ChatGPT preferred in 4 of 6 questions; one question showed no difference; one favoured the consultant
Primary Outcome: Patient-rated total quality score (understandability, helpfulness, reassurance) for ChatGPT versus consultant urogynecologist responses across six common urogynecology questions
Nnt Or Sensitivity: Not applicable for this survey-based comparative study; no NNT, sensitivity, specificity, or hazard ratio reported. The absolute score difference of 4 points (26.7% of maximum 15-point scale) is of uncertain clinical significance without a defined MCID for this instrument
Confidence Interval: Not reported
Clinical Application
ChatGPT and similar LLMs are freely or low-cost accessible to patients and clinicians. However, clinical implementation requires governance frameworks addressing accuracy verification, version control, liability, informed consent for AI-assisted information, and integration with existing patient education resources. The non-deterministic nature of LLM outputs means quality cannot be guaranteed without human oversight. In Australia, urogynecology services are provided through public hospital specialist clinics and private practice, with significant wait times in the public sector. The Continence Foundation of Australia and RANZCOG provide patient education resources. TGA does not currently regulate patient-facing LLM health information tools as therapeutic goods, though this regulatory landscape is evolving. RACGP guidelines on digital health emphasise that AI-generated health information must be supplementary to, not a replacement for, clinical consultation. PBS-listed treatments for conditions such as overactive bladder (e.g., oxybutynin, solifenacin, mirabegron) and stress urinary incontinence management are not affected by this study, but patient education quality may influence treatment adherence and help-seeking behaviour. Australian clinicians should note that LLM responses may not reflect TGA-approved indications or Australian clinical guidelines. Women attending specialist urogynecology outpatient services seeking written health information about common urogynecological conditions. Findings are most directly applicable to literate, English-speaking patients in tertiary care settings.
Abstract
INTRODUCTION: Large language models (LLMs) are increasingly used in healthcare, including urogynecology, where stigma may limit open discussion. LLM-based chat platforms may provide a less intimidating and more accessible way for patients to obtain information, but their reliability requires evaluation. This study compared the quality of ChatGPT-generated responses in urogynecology with those provided by a consultant urogynecologist, focusing on understandability, helpfulness, and reassurance. MATERIAL AND METHODS: A cross-sectional survey was conducted among urogynecology patients. After informed consent, participants reviewed responses to six common questions, each answered by ChatGPT and a single consultant. A blinded third-party consultant verified clinical accuracy. Patients rated responses using a 5-point Likert scale across three domains (maximum score 15 per response). Wilcoxon signed-rank tests were used for comparison. RESULTS: A total of 203 patients participated (median age 56 years, interquartile range 46-66). ChatGPT responses received higher total ratings than consultant responses (76 [67-85] vs. 72 [63-80], p < 0.01). Scores were higher for understandability, helpfulness, and reassurance (all p < 0.01). ChatGPT was preferred in four of six questions, one showed no difference, and one favored the consultant. Subgroup analyses showed no significant variation based on patient characteristics. CONCLUSIONS: In this exploratory study, women rated ChatGPT's responses as clearer and more reassuring than consultant answers. These findings reflect patient perceptions in a limited setting and should be interpreted with caution. While LLMs may have a supportive role in patient education, their use must remain secondary to expert clinical care and subject to careful oversight.
References
- 1.Rotem, R., Simon, C., Rottenstreich, M., Misgav, O., O'Reilly, B. A., Weintraub, A. Y., & O'Sullivan, O. E. (2025). ChatGPT in urogynecology: Comparing large language model responses to human experts. Acta Obstetricia et Gynecologica Scandinavica. https://doi.org/10.1186/s12916-025-04076-0
This content is for educational purposes for healthcare professionals only and does not constitute clinical advice. Clinical decisions should be based on individual patient assessment, current guidelines, and appropriate specialist consultation. Editorial Standards · Privacy Policy · Terms of Service