Research Appraisalobservational

Reliability and readability of AI chatbot responses to patient questions about robot-assisted radical cystectomy

Journal of robotic surgeryLiu, Yang, Li, Rongkang, Yu, Peng et al.22 July 2026DOI

Clinical Snapshot

45CEBM
Evidence: Weakobservational

PICO Framework

P — PopulationSimulated patient-education context: 20 core patient questions about robot-assisted radical cystectomy (RARC), evaluated by three senior urologic experts
I — InterventionFour contemporary large-language-model chatbots: ChatGPT-5, DeepSeek-V4, Claude Sonnet 4.6, and Gemini 3.5 Pro
C — ComparatorCross-chatbot comparative evaluation; no human clinician or validated patient-education material used as a reference standard
O — OutcomesReliability (DISCERN, EQIP, GQS, JAMA benchmark criteria) and readability (ARI, Coleman-Liau Index, Flesch-Kincaid Grade Level, Flesch Reading Ease, Gunning Fog Index, SMOG) of chatbot-generated responses

Bottom Line

This cross-sectional comparative evaluation assessed four large-language-model chatbots on the reliability and readability of their responses to 20 standardised patient-education questions about robot-assisted radical cystectomy. Using multiple validated instruments, DeepSeek-V4 outperformed peers on reliability metrics, while Gemini 3.5 Pro scored lowest. All models produced text exceeding the recommended sixth-grade reading level, and Flesch Reading Ease scores fell below recommended thresholds across the board. The study is hypothesis-generating but carries significant methodological limitations: factual accuracy was not assessed, inter-rater reliability is unreported, rater blinding to chatbot identity is not described, and the 20-question sample was developed without patient involvement or formal consensus methodology. The stochastic nature of chatbot outputs means findings are not reproducible. For Australian urologists and urology nurses, the key clinical message is clear: current chatbot tools produce variable-quality, often difficult-to-read responses, and — critically — their factual accuracy for complex surgical patient education remains unvalidated. These tools should not replace individualised clinician counselling or replace validated patient-education resources. Any supplementary use warrants clinician review for factual accuracy before patient distribution.

Evidence: Weak

Key Findings

  • P Value: Statistical significance claimed for DISCERN, EQIP, and GQS differences; specific p-values not provided in abstract

  • Effect Size: Not reported numerically in the abstract; described as 'differed significantly' across models without quantification of magnitude

  • Primary Outcome: Reliability (DISCERN, EQIP, GQS, JAMA benchmark) and readability (six indices) of chatbot responses to 20 RARC patient-education questions. DeepSeek-V4 achieved highest mean scores for DISCERN, EQIP, and GQS; ChatGPT-5 and DeepSeek-V4 showed strongest JAMA benchmark performance; Gemini 3.5 Pro had lowest reliability and transparency scores and generated most complex text.

  • Nnt Or Sensitivity: Not applicable to this descriptive comparative study; all models exceeded recommended sixth-grade reading level threshold; Flesch Reading Ease scores remained below recommended threshold across all models

  • Confidence Interval: Not reported

Clinical Application

Chatbot tools evaluated are freely or commercially accessible. However, the study's failure to assess factual accuracy means clinicians cannot safely recommend any of these tools for unsupervised patient use based on this evidence alone. Any implementation would require clinician review of chatbot-generated content for factual correctness before distribution. In Australia, RARC is performed at major tertiary referral centres and is not universally PBS-subsidised as a standalone item; robotic surgical access varies by state and institution. The TGA does not currently regulate large-language-model chatbots as therapeutic goods unless embedded in a regulated software-as-a-medical-device (SaMD) framework. RACGP and Urological Society of Australia and New Zealand (USANZ) guidelines emphasise shared decision-making and informed consent for complex urological procedures. Australian patients using chatbots for RARC information face the same readability and accuracy risks identified in this study. The finding that all models exceeded sixth-grade reading level is particularly relevant given Australian adult health literacy data, where approximately 60% of adults have health literacy below the level needed to adequately understand health information. Clinicians should not direct patients to chatbot platforms as primary information sources without validated, locally contextualised patient-education materials as a supplement. Urologists, urology nurses, and patient-education coordinators considering chatbot-generated materials as supplementary resources for patients undergoing or considering robot-assisted radical cystectomy for bladder cancer

Abstract

Robot-assisted radical cystectomy (RARC) is a complex procedure that requires patients to understand surgical indications, urinary diversion, perioperative treatment, complications, recovery, and long-term functional outcomes. Although artificial intelligence (AI) chatbots are increasingly used to obtain medical information, their suitability for RARC patient education remains unclear. We conducted a cross-sectional comparative evaluation of four contemporary AI chatbots: ChatGPT-5, DeepSeek-V4, Claude Sonnet 4.6, and Gemini 3.5 Pro. A set of 20 core patient-education questions on RARC was developed by three senior urologic experts. Chatbot responses were assessed using DISCERN, the Ensuring Quality Information for Patients tool, the Global Quality Scale, and JAMA benchmark criteria. Readability was evaluated using the Automated Readability Index, Coleman-Liau Index, Flesch-Kincaid Grade Level, Flesch Reading Ease, Gunning Fog Index, and SMOG. Reliability scores differed significantly across models for DISCERN, EQIP, and GQS, while JAMA benchmark criteria were summarized descriptively as transparency signals. DeepSeek-V4 achieved the highest mean scores for DISCERN, EQIP, and GQS, while ChatGPT-5 and DeepSeek-V4 showed the strongest JAMA benchmark performance. Gemini 3.5 Pro generally had the lowest reliability and transparency scores. Readability also varied across models. DeepSeek-V4 produced the most readable responses overall, whereas Gemini 3.5 Pro generated the most complex text. However, all models exceeded the recommended sixth-grade reading level, and FRES scores remained below the recommended threshold. Contemporary AI chatbots generated responses with variable presentation quality, transparency, and readability for common RARC patient-education questions. Because factual accuracy was not directly assessed, these tools should not be interpreted as validated sources of clinical guidance and should not replace individualized counseling by urologists.

References

  1. 1.Liu, Y., Li, R., Yu, P., Zhang, Y., Zhao, A., Xiao, X., Li, Z., Liang, R., Peng, L., & Dong, Z. (2026). Reliability and readability of AI chatbot responses to patient questions about robot-assisted radical cystectomy. Journal of Robotic Surgery. https://doi.org/10.3233/shti230562
Share:XLinkedIn

This content is for educational purposes for healthcare professionals only and does not constitute clinical advice. Clinical decisions should be based on individual patient assessment, current guidelines, and appropriate specialist consultation. Editorial Standards · Privacy Policy · Terms of Service