Research Appraisalother

Large Language Models in German Continuing Medical Education Assessments: Protocol for a Fully Crossed Experimental Study

JMIR research protocolsÖzmen, Leyla, Burisch, Christian, Gödde, Daniel et al.28 July 2026DOI

Clinical Snapshot

60CEBM
Evidence: Weakother

PICO Framework

P — PopulationGerman continuing medical education (CME) assessments — specifically 18 expired CME articles from 3 major German publishers across 6 medical specialties (no human participants)
I — InterventionFour current large language models (GPT-5, Claude Sonnet 4, Grok-4, Gemini 3) processing CME test questions across four document format conditions (searchable PDF, protected PDF, raster PDF, vector PDF)
C — ComparatorPerformance compared across 16 model-format combinations; primary benchmark is the CME passing threshold (typically 70% correct answers per module)
O — OutcomesPrimary: proportion of correctly answered CME questions per model-format combination. Secondary: pass/fail rate per model-format combination

Bottom Line

This is a well-conceived and transparently registered experimental protocol addressing a genuinely important and timely question: can simple document format manipulation reduce the ability of large language models to pass continuing medical education assessments? The fully crossed design is methodologically sound, ethics approval has been obtained, and the use of expired materials avoids harm to active assessment integrity. However, as a protocol paper, it offers no empirical findings — all conclusions remain prospective. Critical methodological gaps include the absence of a specified prompt engineering strategy, inadequate attention to training data contamination, and a vague statistical analysis plan. The three-run repetition per condition may also underestimate stochastic output variance. Clinicians and CME administrators should note that any format-based safeguard is likely to be a temporary measure: current LLMs already possess multimodal vision capabilities that can process raster images, and model capabilities are advancing rapidly. Australian CPD providers and the RACGP should monitor these findings as part of broader assessment integrity reviews, while simultaneously exploring more robust long-term solutions such as adaptive testing, clinical vignette complexity, and supervised assessment modalities.

Evidence: Weak

Key Findings

  • P Value: Not yet available

  • Effect Size: Not yet available

  • Primary Outcome: Proportion of correctly answered CME questions per LLM-format combination — not yet available (protocol paper)

  • Nnt Or Sensitivity: Not applicable at protocol stage; the operationally relevant benchmark is the 70% passing threshold specified per CME module

  • Confidence Interval: Not yet available

Clinical Application

If results demonstrate that non-searchable PDF formats (raster, protected, vector) meaningfully reduce LLM pass rates, implementation would require only minor technical changes to CME delivery platforms. This is a low-cost, low-infrastructure intervention. However, accessibility considerations for physicians with disabilities must be weighed, and any format-based safeguard is likely to be a temporary measure given ongoing LLM capability advances, particularly multimodal vision capabilities. Australian CPD requirements are governed by the RACGP (triennium-based CPD framework), ACRRM, and specialist college requirements linked to AHPRA registration. The Medical Board of Australia mandates CPD participation as a condition of registration. Many Australian CPD providers, including those offering online modules through platforms such as gplearning and Healthed, use PDF-based or web-based MCQ assessments structurally similar to German CME formats. The integrity concerns raised by this study are directly relevant to Australian CPD governance. The TGA and RACGP have not yet issued formal guidance on LLM use in CPD assessments. Australian providers may wish to monitor these findings when reviewing their assessment security policies. PBS and TGA regulatory implications are indirect but relevant insofar as CME/CPD underpins prescribing competency and safe medication management. CME providers, medical regulators, postgraduate medical education administrators, and physicians engaged in online CME assessments — particularly in jurisdictions using PDF-based online testing formats

Abstract

BACKGROUND: Continuing medical education (CME) is a legal and ethical obligation for physicians in Germany. The rapid rise of large language models (LLMs) such as ChatGPT, Gemini, Claude, and Grok raises concerns about the integrity of CME assessments, as LLMs can already pass German CME tests. OBJECTIVE: This study aims to determine whether the choice of document format (searchable PDF, protected PDF, raster PDF, or vector PDF) and LLM influences the ability of LLMs to solve CME test questions at rates exceeding the passing threshold specified for each CME module (typically 70%). METHODS: In a fully crossed within-subjects repeated-measures design, 18 expired CME articles from 3 major German publishers across 6 specialties will be converted into 3 cheating-impeding PDF formats and processed alongside the original PDF files by 4 current LLMs (GPT-5, Claude Sonnet 4, Grok-4, and Gemini 3). This results in 16 model-format combinations. Each model will answer every article 3 times per file-format condition, with outcomes derived from aggregated run-level results. The primary outcome is the proportion of correctly answered questions; the secondary outcome is the pass/fail rate. RESULTS: The study has been approved by the Witten/Herdecke University Ethics Committee (S-260/2025; dated August 10, 2025) and is preregistered at the Open Science Framework. The study is supported by internal departmental resources only, and no external funding was received. Because this protocol evaluates LLMs using expired CME materials, no human participants are being recruited. Data collection is planned to begin in June 2026 and is expected to last approximately 4 weeks. At the time of manuscript submission, no data have been collected or analyzed. Results are expected to be available after the completion of data collection and statistical analysis in 2026. The analyses will quantify performance differences across document formats; these findings may inform the feasibility of nonsearchable document formats as a temporary measure to reduce LLM-enabled cheating risks in CME contexts. CONCLUSIONS: By quantifying how document format constrains LLM performance, this study aims to evaluate simple technical safeguards that may reduce artificial intelligence-assisted manipulation of CME tests and inform regulators and CME providers about how to balance assessment validity, accessibility, and responsible LLM integration into postgraduate medical education.

References

  1. 1.Özmen, L., Burisch, C., Gödde, D., Breuckmann, F., Ehlers, J., & Sellmann, T. (2026). Large language models in German continuing medical education assessments: Protocol for a fully crossed experimental study. JMIR Research Protocols. https://doi.org/10.2196/72356
Share:XLinkedIn

This content is for educational purposes for healthcare professionals only and does not constitute clinical advice. Clinical decisions should be based on individual patient assessment, current guidelines, and appropriate specialist consultation. Editorial Standards · Privacy Policy · Terms of Service