Research AppraisalSystematic Review

Accuracy of Large Language Models in Answering Dental Examination Questions: A Systematic Review and Meta-Analysis

International dental journalDashti, Mahmood, Khosraviani, Farshad, Meyari, Atieh et al.1 Aug 2026DOI

Clinical Snapshot

55CEBM
Evidence: WeakSystematic Review

PICO Framework

P — PopulationDental examination questions (multiple-choice and other formats) drawn from dental licensing, board, and educational assessments across various dental specialties and educational levels
I — InterventionLarge language models (LLMs) including ChatGPT-3.5, ChatGPT-4, Microsoft Copilot (GPT-based), Google Gemini, and other non-GPT systems evaluated on dental examination question sets
C — ComparatorComparison between LLM versions (e.g., GPT-4 vs GPT-3.5), between different LLM platforms, and implicitly against human passing thresholds for dental examinations
O — OutcomesAccuracy rate (proportion of correctly answered dental examination questions) as the primary outcome; subgroup analyses by LLM type/version, question format, and dental subdiscipline

Bottom Line

This systematic review and meta-analysis of 39 studies reports a pooled LLM accuracy of 63.7% (95% CI: 60.3%–67.1%) on dental examination questions, with ChatGPT-4 and Microsoft Copilot achieving approximately 73% and 75% accuracy respectively in subgroup analyses. While these figures place advanced LLMs at or above typical dental examination passing thresholds, the extremely high heterogeneity (I² = 91.5%) substantially limits confidence in the pooled estimate. The absence of a GRADE assessment, unspecified risk of bias methodology, and lack of prediction intervals are notable methodological gaps. The authors appropriately conclude that LLMs are insufficient for autonomous clinical decision-making in dentistry. For Australian dental educators and practitioners, these findings support cautious, supervised use of LLMs as adjuncts in examination preparation and self-directed learning, but do not justify their deployment as clinical decision support tools. The rapid evolution of LLM technology means this evidence base will require frequent updating. Clinicians and educators should critically evaluate LLM outputs, particularly for complex clinical reasoning tasks where hallucination and miscalibration remain unquantified risks in this review.

Evidence: Weak

Key Findings

  • P Value: Not reported in abstract; ChatGPT-4 significantly outperformed earlier versions and some competitor models on direct comparison

  • Effect Size: Pooled accuracy: 63.7%; ChatGPT-4 subgroup: ~73%; Microsoft Copilot subgroup: ~75%

  • Primary Outcome: Pooled accuracy of LLMs in answering dental examination questions across 39 included studies

  • Nnt Or Sensitivity: Not applicable as a performance accuracy metric; no NNT calculable. Passing threshold context: most dental board examinations require 60–70% accuracy, placing overall LLM performance at the lower margin of passing thresholds and GPT-4/Copilot above typical passing benchmarks

  • Confidence Interval: Overall: 95% CI 60.3%–67.1%

Clinical Application

LLMs such as ChatGPT-4 and Microsoft Copilot are freely or affordably accessible to dental students and practitioners globally. Integration into examination preparation workflows is technically feasible. However, the high heterogeneity of performance across question types and dental subdisciplines means that reliability cannot be assumed across all use cases. Structured prompting and retrieval-augmented generation approaches are identified as warranting further investigation. In Australia, dental graduates must pass the Australian Dental Council (ADC) written examination for overseas-trained dentists, and domestic graduates complete university-based assessments aligned with ADC competency standards. The included studies do not appear to specifically evaluate ADC examination formats. The TGA does not currently regulate LLMs as medical devices for clinical decision support in dentistry. RACGP and ADC guidelines do not yet formally address LLM use in dental education or examination preparation. Australian dental educators should treat these findings as preliminary and context-dependent. The PBS does not apply to LLM tools. Dental schools considering LLM integration into curricula should note that a 63.7% pooled accuracy — with I² of 91.5% — does not support uncritical adoption, but does not preclude supervised, educationally scaffolded use. Dental students, dental educators, and dental examination bodies considering LLM integration into examination preparation, self-directed learning, or formative assessment tools. Not applicable to clinical patient care decision-making based on current evidence.

Abstract

INTRODUCTION: Large language models (LLMs), including OpenAI's GPT family accessed via interfaces such as ChatGPT and Microsoft Copilot, as well as non-GPT systems such as Google Gemini, are increasingly applied in healthcare and dental education. However, the accuracy of these systems in specialized tasks such as answering dental examination questions remains unclear. METHODS: This systematic review and meta-analysis evaluated LLM performance in answering dental questions. Databases searched were PubMed, Embase, Scopus, and Web of Science. Data on question type and number, LLM versions, and accuracy rates were extracted. Pooled accuracy was estimated using a random-effects model; heterogeneity and publication bias were assessed. RESULTS: A total of 39 studies were included, with ChatGPT-4 being the most frequently evaluated model. The pooled accuracy for LLMs was 63.7% (95% CI: 60.3%-67.1%), with high heterogeneity (I² = 91.5%). Subgroup analysis revealed ChatGPT-4 and Copilot (a GPT-based interface) achieved the highest pooled accuracies (∼73% and ∼75%, respectively). Direct comparisons confirmed ChatGPT-4 significantly outperformed earlier versions and some competitor models. Sensitivity analyses supported the robustness of findings. CONCLUSION: LLMs demonstrate moderate accuracy in answering dental examination questions and are currently insufficient for autonomous clinical decision-making. When their limitations are explicitly recognized, however, these systems may serve as valuable adjuncts in dental education and examination preparation. Methodological strategies such as structured prompting and retrieval-augmented approaches warrant further investigation but were not the primary focus of the present analysis.

References

  1. 1.Dashti, M., Khosraviani, F., Meyari, A., Amirzade-Iranaq, M. H., Chaurasia, A., Hefzi, D., Ghadimi, N., Tichy, A., Khurshid, Z., & Schwendicke, F. (2026). Accuracy of large language models in answering dental examination questions: A systematic review and meta-analysis. International Dental Journal. Advance online publication. https://doi.org/10.1002/jrsm.1564
Share:XLinkedIn

This content is for educational purposes for healthcare professionals only and does not constitute clinical advice. Clinical decisions should be based on individual patient assessment, current guidelines, and appropriate specialist consultation. Editorial Standards · Privacy Policy · Terms of Service