Performance of DeepSeek V3 and ChatGPT-4o in answering esophageal cancer-related questions
Clinical Snapshot
PICO Framework
| P — Population | Esophageal cancer-related health knowledge questions (n=52), evaluated by two experienced gastroenterologists at a single Chinese tertiary centre |
| I — Intervention | DeepSeek V3 large language model responses to esophageal cancer questions |
| C — Comparator | ChatGPT-4o large language model responses to the same esophageal cancer questions |
| O — Outcomes | Accuracy of responses (scored by two independent gastroenterologists on an unspecified Likert-type scale), temporal stability of responses across two independent test runs, and subgroup performance across six thematic question categories |
Bottom Line
This cross-sectional comparative evaluation found that both DeepSeek V3 and ChatGPT-4o achieved similar median accuracy scores (4 out of a maximum 4, IQR 3–4) across 52 esophageal cancer questions assessed by two gastroenterologists, with no statistically significant differences between models overall or across six thematic subgroups. Temporal stability was high, with only 2–3 inconsistent responses per model across repeat testing. However, the study is substantially limited by an unvalidated scoring instrument with undescribed criteria, absence of inter-rater reliability statistics, unreported prompt standardisation, unspecified model query dates and language, and no confidence intervals. The question set was not derived from a validated source and may not reflect real-world patient information needs. For Australian clinicians, the findings are further limited by the Chinese clinical context, where squamous cell carcinoma predominates — unlike the adenocarcinoma-predominant Australian epidemiology. The authors themselves appropriately conclude that findings do not confirm general reliability for routine clinical or public use. Senior clinicians should treat this study as preliminary signal-generating work only, and should not use it to endorse either model for patient-facing esophageal cancer information without substantially more rigorous evaluation.
Key Findings
P Value: All comparisons P > 0.05 (both overall and subgroup scores, and temporal stability across two test runs for both models)
Effect Size: No effect size measure reported; median scores were identical for both models overall. Subgroup differences were numerically small: ChatGPT-4o scored lower in diagnosis and molecular biology (median 3 [IQR 2–4]) and management of advanced/metastatic disease (median 3 [IQR 3–4]) compared to DeepSeek V3 (median 4 [IQR 3–4] for both subgroups), but these differences were not statistically significant
Primary Outcome: Overall accuracy scores for both DeepSeek V3 and ChatGPT-4o across 52 esophageal cancer questions were identical: median 4 (IQR 3–4), with no statistically significant difference between models (P > 0.05)
Nnt Or Sensitivity: Not applicable to this study design. Temporal instability: DeepSeek V3 produced inconsistent responses on 2/52 questions (3.8%); ChatGPT-4o on 1/52 questions (1.9%) — no formal reliability metric reported
Confidence Interval: Not reported for any comparison
Clinical Application
Both models are freely or readily accessible to patients and clinicians. However, the study's methodological limitations mean that clinicians cannot use these findings to confidently endorse either model for patient-facing esophageal cancer information. The absence of validated accuracy benchmarks, inter-rater reliability data, and prompt standardisation means that the reported 'favorable accuracy' cannot be independently verified or operationalised in clinical practice Esophageal cancer in Australia is predominantly adenocarcinoma (arising from Barrett's oesophagus), in contrast to the squamous cell carcinoma predominance in China where this study was conducted — this epidemiological difference significantly limits the applicability of the question set and findings to Australian patients. Australian clinical practice is guided by Cancer Council Australia Clinical Practice Guidelines, COSA recommendations, and PBS-listed systemic therapies (including nivolumab, pembrolizumab for advanced disease). TGA-approved indications and PBS subsidy criteria for esophageal cancer treatments differ from Chinese practice. RACGP guidelines emphasise shared decision-making and patient health literacy. Australian clinicians should note that neither model has been evaluated against Australian-specific guidelines, and patients using these tools may receive information inconsistent with local practice. The Therapeutic Goods Administration has not approved or regulated large language models as medical devices for clinical decision support in Australia. Patients seeking health information about esophageal cancer and clinicians considering the use of large language model tools for patient education or clinical decision support in gastroenterology and oncology settings
Abstract
Esophageal cancer remains a significant global health issue. ChatGPT-4o and DeepSeek V3 can provide the public with health-related knowledge about esophageal cancer. This study aimed to evaluate the accuracy of DeepSeek V3 and ChatGPT-4o in responding to health knowledge questions related to esophageal cancer. Fifty-two questions related to esophageal cancer were classified into themes of basic knowledge, diagnosis and molecular biology, management of local and locoregional diseases, management of advanced and metastatic diseases, clinical case analysis and patient frequently asked questions (FAQs). These questions were entered into DeepSeek V3 and ChatGPT-4o to obtain responses, and 2 experienced gastroenterologists independently evaluated the accuracy and temporal stability of each response. Overall, the scores of DeepSeek V3 and ChatGPT-4o on all questions were 4 (3-4), and there was no statistically significant difference between the 2 groups. The final scores of DeepSeek V3 in basic knowledge, diagnosis and molecular biology, management of local and locoregional diseases, management of advanced and metastatic diseases, clinical case analysis, and FAQs were 4 (3-4), 4 (3-4), 4 (3-4), 3 (3-4), 4 (4-4), and 4 (4-4), respectively, while the scores of ChatGPT-4o were 4 (3-4), 3 (2-4), 4 (3-4), 3 (3-4), 4 (4-4), and 4 (4-4), respectively. For temporal stability across 2 independent test runs, DeepSeek V3 presented inconsistent responses on 2 questions, and ChatGPT-4o on 1 question; no statistically significant differences were found in overall and subgroup scores between the 2 runs for both models (all P > .05). ChatGPT-4o and DeepSeek V3 showed favorable accuracy and comprehensive responses to most of the 52 esophageal cancer-related questions in this study, but our findings do not confirm their general reliability for esophageal cancer health information in routine clinical or public use.
References
- 1.Yang, Q., Yi, J., Gong, A., Li, Y., Feng, Q., Fu, A., Li, J., & Zhan, Y. (2026). Performance of DeepSeek V3 and ChatGPT-4o in answering esophageal cancer-related questions. Medicine. https://doi.org/10.1097/MD.0000000000049896
Related Research
The American journal of pathology
An Integrated Machine Learning and Genomic Framework for Precise Detection of Gastric Cancer
3 Aug 2026
Clinical gastroenterology and hepatology : the official clinical practice journal of the American Gastroenterological Association
Artificial Intelligence Tools for Gastrointestinal Research: A Practical Guide
2 Aug 2026
Current opinion in critical care
Smart feeding: the role of artificial intelligence and integrated nutrition platforms in the ICU
2 Aug 2026
This content is for educational purposes for healthcare professionals only and does not constitute clinical advice. Clinical decisions should be based on individual patient assessment, current guidelines, and appropriate specialist consultation. Editorial Standards · Privacy Policy · Terms of Service