Evidence-Based Medicine

Research Appraisals

Evidence-based critical appraisals of the latest medical research, systematically evaluated using Oxford CEBM methodology.

Showing 2 appraisals

observationalEvidence: Weak
45CEBM

Medicine

Performance of DeepSeek V3 and ChatGPT-4o in answering esophageal cancer-related questions

Esophageal cancer remains a significant global health issue. ChatGPT-4o and DeepSeek V3 can provide the public with health-related knowledge about esophageal cancer. This study aimed to evaluate the accuracy of DeepSeek V3 and ChatGPT-4o in responding to health knowledge questions related to esophageal cancer. Fifty-two questions related to esophageal cancer were classified into themes of basic knowledge, diagnosis and molecular biology, management of local and locoregional diseases, management of advanced and metastatic diseases, clinical case analysis and patient frequently asked questions (FAQs). These questions were entered into DeepSeek V3 and ChatGPT-4o to obtain responses, and 2 experienced gastroenterologists independently evaluated the accuracy and temporal stability of each response. Overall, the scores of DeepSeek V3 and ChatGPT-4o on all questions were 4 (3-4), and there was no statistically significant difference between the 2 groups. The final scores of DeepSeek V3 in basic knowledge, diagnosis and molecular biology, management of local and locoregional diseases, management of advanced and metastatic diseases, clinical case analysis, and FAQs were 4 (3-4), 4 (3-4), 4 (3-4), 3 (3-4), 4 (4-4), and 4 (4-4), respectively, while the scores of ChatGPT-4o were 4 (3-4), 3 (2-4), 4 (3-4), 3 (3-4), 4 (4-4), and 4 (4-4), respectively. For temporal stability across 2 independent test runs, DeepSeek V3 presented inconsistent responses on 2 questions, and ChatGPT-4o on 1 question; no statistically significant differences were found in overall and subgroup scores between the 2 runs for both models (all P > .05). ChatGPT-4o and DeepSeek V3 showed favorable accuracy and comprehensive responses to most of the 52 esophageal cancer-related questions in this study, but our findings do not confirm their general reliability for esophageal cancer health information in routine clinical or public use.

27 July 2026

Read appraisal →
observationalEvidence: Weak
45CEBM

Journal of robotic surgery

Reliability and readability of AI chatbot responses to patient questions about robot-assisted radical cystectomy

Robot-assisted radical cystectomy (RARC) is a complex procedure that requires patients to understand surgical indications, urinary diversion, perioperative treatment, complications, recovery, and long-term functional outcomes. Although artificial intelligence (AI) chatbots are increasingly used to obtain medical information, their suitability for RARC patient education remains unclear. We conducted a cross-sectional comparative evaluation of four contemporary AI chatbots: ChatGPT-5, DeepSeek-V4, Claude Sonnet 4.6, and Gemini 3.5 Pro. A set of 20 core patient-education questions on RARC was developed by three senior urologic experts. Chatbot responses were assessed using DISCERN, the Ensuring Quality Information for Patients tool, the Global Quality Scale, and JAMA benchmark criteria. Readability was evaluated using the Automated Readability Index, Coleman-Liau Index, Flesch-Kincaid Grade Level, Flesch Reading Ease, Gunning Fog Index, and SMOG. Reliability scores differed significantly across models for DISCERN, EQIP, and GQS, while JAMA benchmark criteria were summarized descriptively as transparency signals. DeepSeek-V4 achieved the highest mean scores for DISCERN, EQIP, and GQS, while ChatGPT-5 and DeepSeek-V4 showed the strongest JAMA benchmark performance. Gemini 3.5 Pro generally had the lowest reliability and transparency scores. Readability also varied across models. DeepSeek-V4 produced the most readable responses overall, whereas Gemini 3.5 Pro generated the most complex text. However, all models exceeded the recommended sixth-grade reading level, and FRES scores remained below the recommended threshold. Contemporary AI chatbots generated responses with variable presentation quality, transparency, and readability for common RARC patient-education questions. Because factual accuracy was not directly assessed, these tools should not be interpreted as validated sources of clinical guidance and should not replace individualized counseling by urologists.

23 July 2026

Read appraisal →