Applications of Large Language Models in Ovarian Cancer Management: Protocol for a Systematic Review and Meta-Analysis
Clinical Snapshot
PICO Framework
| P — Population | Patients with ovarian cancer (all stages and histological subtypes) and clinicians involved in ovarian cancer management |
| I — Intervention | Large language models (LLMs) including GPT-4, Claude, Google Gemini, and comparable systems applied to ovarian cancer care tasks (diagnosis, prognosis, treatment planning, patient communication, report generation) |
| C — Comparator | Standard clinical practice, conventional diagnostic or prognostic tools, or other AI/non-AI comparators as reported in included studies |
| O — Outcomes | Primary: model performance metrics (accuracy, sensitivity, specificity, AUC, F1-score); Secondary: clinical process impacts, safety concerns, usability, and patient engagement outcomes |
Bottom Line
This paper is a pre-registered systematic review protocol — not a completed evidence synthesis — and must be evaluated accordingly. The authors plan to conduct the first comprehensive meta-analysis of large language model applications across the ovarian cancer management pathway, encompassing diagnosis, prognosis, treatment planning, and patient communication. The methodological framework is generally sound: PRISMA-P compliance, PROSPERO registration, multi-database searching including technical and Chinese-language sources, and a sophisticated multi-tool risk-of-bias approach matched to anticipated study design heterogeneity are all commendable. However, several important gaps exist: grey literature and preprint sources are not explicitly included despite their prominence in AI research; GRADE certainty of evidence assessment is absent; patient-reported and health economic outcomes are not planned; and the critical challenge of LLM version heterogeneity across included studies is unaddressed. For Australian gynaecological oncology clinicians, this review — when completed — may provide useful evidence to inform institutional governance decisions around LLM adoption. However, no practice change is warranted based on this protocol alone. Senior clinicians should await the completed review, anticipated in late 2026, before drawing any conclusions about LLM utility in ovarian cancer care.
Key Findings
P Value: Not yet available
Effect Size: Not yet available
Primary Outcome: No results available — this is a protocol paper. Planned primary outcomes are pooled model performance metrics including accuracy, sensitivity, specificity, AUC, and F1-score for LLMs applied to ovarian cancer management tasks.
Nnt Or Sensitivity: Not yet available; planned analyses include bivariate modelling of sensitivity and specificity for diagnostic accuracy studies using the mada package in R
Confidence Interval: Not yet available
Clinical Application
Clinical application cannot be assessed at protocol stage. The completed review will need to address LLM integration into existing electronic medical record systems, clinician training requirements, medicolegal accountability frameworks, and the practical limitations of LLM hallucination in high-stakes oncological decision-making Ovarian cancer remains a significant burden in Australia, with approximately 1,800 new diagnoses annually (AIHW). The TGA has not yet approved any LLM-based clinical decision support tool specifically for ovarian cancer management. RACGP and RANZCOG have not issued formal guidance on LLM use in gynaecological oncology as of 2025. PBS-listed treatments for ovarian cancer (including PARP inhibitors such as olaparib and niraparib) involve complex eligibility criteria where LLM-assisted decision support could theoretically reduce prescribing errors, but this remains unvalidated. Australian clinicians should note that LLMs trained predominantly on non-Australian datasets may not reflect local epidemiology, treatment guidelines, or PBS restrictions. The Australian Digital Health Agency's framework for AI in healthcare will be relevant to any future implementation guidance derived from this review. Women with suspected or confirmed ovarian cancer across all stages and histological subtypes; clinicians involved in gynaecological oncology including medical oncologists, gynaecological surgeons, radiation oncologists, and allied health professionals
Abstract
BACKGROUND: Ovarian cancer (OC) is a highly fatal gynecologic malignancy with complex management challenges and limited long-term survival for advanced stages. Large language models (LLMs)-including systems such as GPT-4, Claude, Google Gemini, and others-are emerging artificial intelligence (AI) tools capable of performing health care-related tasks such as diagnostic support, treatment planning, report generation, and patient communication. However, their applications in OC care have not yet been comprehensively assessed. OBJECTIVE: This protocol outlines a systematic review and meta-analysis aimed at evaluating the use, performance, and clinical impact of LLMs in OC management. We will examine how LLMs have been applied across various domains (eg, diagnosis, prognosis, treatment planning, and patient engagement), the metrics used to assess their performance (eg, accuracy, sensitivity, and area under the curve), and their strengths and limitations. METHODS: This review will be conducted in accordance with PRISMA-P (Preferred Reporting Items for Systematic Reviews and Meta-Analyses Protocols) guidelines. A comprehensive search strategy will be implemented across biomedical, technical, and Chinese-language databases (eg, PubMed, Embase, Web of Science, IEEE Xplore, and China National Knowledge Infrastructure) from inception to December 31, 2025. Eligible studies include clinical evaluations, validation studies, and real-world implementation reports involving LLMs in OC care. Two independent reviewers will perform screening, data extraction, and quality appraisal using validated tools (eg, version 2 of the Cochrane risk-of-bias tool for randomized trials, Risk of Bias in Nonrandomized Studies of Interventions, Quality Assessment of Diagnostic Accuracy Studies 2, and Prediction Model Study Risk of Bias Assessment Tool+AI). Outcomes of interest include model performance metrics, clinical process impacts, safety concerns, and usability. Meta-analyses will be conducted where feasible using random-effects models in R (meta, metafor, and mada packages), including bivariate models for sensitivity and specificity. RESULTS: The review is currently in progress. The PROSPERO registration has been completed, and the literature search and selection process is underway. Study selection, data extraction, and quality assessment are expected to be completed by mid-2026. Final results will include pooled performance metrics (eg, accuracy, F1-score, and area under the curve), qualitative insights into clinical integration, and identification of limitations such as reporting bias or insufficient external validation. CONCLUSIONS: This systematic review will provide the first comprehensive synthesis of evidence on the application of LLMs in OC care. It will identify promising use cases, highlight safety and reporting challenges, and inform future research directions. The findings are expected to support evidence-based integration of LLMs into gynecologic oncology workflows while promoting transparency and methodological rigor in AI evaluation.
References
- 1.Wang, Y., Yao, J., Tian, J., Wang, Y., & Yang, Y. (2026). Applications of large language models in ovarian cancer management: Protocol for a systematic review and meta-analysis. JMIR Research Protocols. https://doi.org/10.2196/88163
Related Research
Histochemistry and cell biology
Molecular plasticity of LAMA3 across the disease spectrum: pathogenic mechanisms and clinical translation.
29 July 2026
Journal of robotic surgery
Robotic versus laparoscopic and open surgery for endometrial cancer: a systematic review of randomized trials and pooled analysis of conversion rates
21 July 2026
Journal of robotic surgery
Mapping the evolution of deep learning and computer vision in robotic surgery: a bibliometric analysis of surgical video intelligence, instrument perception, and clinical translation.
21 July 2026
This content is for educational purposes for healthcare professionals only and does not constitute clinical advice. Clinical decisions should be based on individual patient assessment, current guidelines, and appropriate specialist consultation. Editorial Standards · Privacy Policy · Terms of Service