Optimizing Treatment Strategies in the Bipolar Disorder Spectrum With Classical AI Approaches: Systematic Review of Performance, Bias, and Clinical Applicability
Clinical Snapshot
PICO Framework
| P — Population | Adult patients with bipolar disorder (BD) spectrum conditions |
| I — Intervention | Classical artificial intelligence approaches (machine learning, regression-based predictive models, rule-based systems, digital phenotyping models) applied to treatment optimisation |
| C — Comparator | Implicit comparator: standard clinical decision-making without AI-assisted prediction; no active comparator arm in included studies |
| O — Outcomes | Predictive model performance (AUC, accuracy) across five outcome categories: acute symptomatic response, long-term maintenance response, relapse and readmission risk, safety and dose optimisation, and brain aging/phenotyping |
Bottom Line
This systematic review of 35 studies synthesises evidence on classical AI predictive models for treatment optimisation in bipolar disorder. Pooled discriminative performance ranges from modest (AUC 0.68 for acute symptomatic response) to moderate (AUC 0.80 for long-term maintenance), with select biomarker and digital phenotyping models reporting higher accuracy. However, the PROBAST+AI assessment reveals high risk of bias in the majority of included studies, driven by small sample sizes, overfitting, and near-universal absence of external validation. Three studies show implausibly high performance likely attributable to these limitations. Confidence intervals around pooled estimates are not reported, and formal heterogeneity statistics are absent, limiting interpretation. No GRADE assessment was conducted. The authors appropriately conclude that current AI tools in bipolar disorder remain exploratory and are not ready for clinical deployment. For Australian psychiatrists, this review reinforces that AI-assisted treatment decision support in bipolar disorder is a promising but immature field. Clinicians should maintain critical appraisal of commercially promoted AI tools and await prospective, externally validated studies before incorporating such approaches into practice. Lithium's apparent neuroprotective signal in brain aging studies warrants further investigation.
Key Findings
P Value: Not reported
Effect Size: Pooled AUC: acute symptomatic response 0.68; long-term maintenance response 0.80; relapse and readmission prediction 0.71. Safety/dose optimisation accuracy 85–97%. Imaging-enhanced acute response models AUC 0.74–0.77. Digital phenotyping/rule-based relapse models AUC 0.85–0.88. Biomarker/cellular maintenance models accuracy 96–99%.
Primary Outcome: Predictive model performance (AUC or accuracy) across five BD treatment outcome categories
Nnt Or Sensitivity: Sensitivity and specificity not reported in the abstract. No NNT calculable from available data. AUC of 0.68 for acute response corresponds to modest discriminative ability marginally above chance (AUC 0.50). AUC of 0.80 for maintenance response represents moderate-to-good discrimination by conventional thresholds.
Confidence Interval: Not reported in the abstract
Clinical Application
Currently not feasible for routine clinical implementation. Most models lack external validation, are derived from small and potentially non-representative samples, and have not been prospectively tested in clinical workflows. Integration into electronic health record systems or clinical decision support tools would require substantial further development, regulatory approval, and prospective validation. The high risk of bias identified by PROBAST+AI across most studies means that reported performance metrics should not be taken at face value for clinical planning. In the Australian context, bipolar disorder management is guided by the Royal Australian and New Zealand College of Psychiatrists (RANZCP) Clinical Practice Guidelines for Mood Disorders (2020). Lithium remains a first-line mood stabiliser and is PBS-listed; the finding that lithium partially mitigates accelerated brain aging in BD is of particular interest to Australian prescribers. The TGA has not approved any AI-based clinical decision support tool specifically for BD treatment optimisation. The Australian Commission on Safety and Quality in Health Care's framework for AI in healthcare would require prospective clinical validation before any such tool could be recommended. Digital phenotyping approaches (AUC 0.85–0.88 for relapse prediction) may have future relevance to Australian telehealth and remote monitoring contexts, particularly for rural and Indigenous populations with limited specialist access, but this remains speculative pending robust validation studies. Adults with bipolar disorder spectrum conditions being considered for pharmacological treatment optimisation, particularly those with treatment-refractory courses or complex polypharmacy. Models incorporating biomarkers, neuroimaging, or digital phenotyping data may be most applicable to specialist psychiatric settings with access to these data streams.
Abstract
BACKGROUND: Bipolar disorder (BD) is a complex and heterogeneous psychiatric condition, characterized by fluctuating clinical courses that affect approximately 1%-2% of the global population in their lifetime. Despite pharmacological advances, treatment response varies significantly among patients, making the identification of individualized treatment strategies a major challenge. Artificial Intelligence (AI), through its classical approaches, has emerged as a powerful tool in precision psychiatry to identify subtle patterns in complex data and inform personalized clinical decisions. OBJECTIVE: The present systematic review aimed to examine the current evidence on classical AI-supported treatment optimization in the BD spectrum. METHODS: The review was conducted in accordance with the PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) 2020 guidelines. Four databases (PubMed, Web of Science, Scopus, and Embase) were searched for original studies published after 2015 on the application of classical AI in the treatment of BD in adult patients. Publication bias was evaluated by visual inspection of a funnel plot. The methodological quality, risk of bias, and clinical applicability of the predictive models were assessed using the Prediction Model Risk Of Bias Assessment Tool for prediction models using regression or AI methods (PROBAST+AI; PROBAST+AI Working Group) tool. RESULTS: A total of 35 studies were included and classified into 5 outcome-based categories, including acute symptomatic response, long-term maintenance response, relapse and readmission risk, safety and dose optimization, and brain aging and phenotyping. Acute symptomatic response models performed modestly (pooled area under the curve [AUC] 0.68), while imaging improved accuracy (74%-77%). Long-term maintenance response models showed moderate-to-high performance (pooled AUC 0.80), with biomarker- and cellular-based models reaching 96%-99% accuracy. Relapse and readmission prediction achieved a pooled AUC of 0.71, with digital phenotyping and rule-based methods performing best (AUC 0.85-0.88). Safety and dose optimization models achieved 85%-97% accuracy. Brain aging and phenotyping studies highlighted accelerated brain aging in BD, partially mitigated by lithium, and revealed novel data-driven subgroups. However, 3 studies were considered at high risk of bias due to small sample sizes associated with disproportionately high-performance estimates. An additional study was identified as potentially biased because it lay markedly distant from the funnel plot's confidence line. Finally, the PROBAST+AI assessment revealed a high risk of bias in most studies, primarily due to data analysis limitations, small sample sizes, and lack of external validation. CONCLUSIONS: The adoption of classical AI tools in BD serves as a driver for therapeutic optimization, although current AI tools in BD should still be considered exploratory rather than ready for clinical use. Effective implementation in real-world clinical scenarios requires more robust, transparent, and externally validated models to ensure reliability and generalizability.
References
- 1.De Francesco, S., Archetti, D., Baronio, C. M., Demaria, C., Boccali, A., Crema, C., Tura, G. B., & Redolfi, A. (2026). Optimizing treatment strategies in the bipolar disorder spectrum with classical AI approaches: Systematic review of performance, bias, and clinical applicability. JMIR Mental Health. https://doi.org/10.2196/93307
Related Research
Journal of psychiatric research
Systematic review of machine learning and deep learning models for EEG-based detection of depression.
2 Aug 2026
Sleep medicine reviews
Artificial intelligence in sleep medicine I: Diagnosis, treatment, care, and research
2 Aug 2026
International journal of yoga therapy
Randomized Controlled Trial on the Effect of Yoga on Anxiety Severity, Quality of Life, and Biomarkers in Individuals with Anxiety Disorder
31 July 2026
This content is for educational purposes for healthcare professionals only and does not constitute clinical advice. Clinical decisions should be based on individual patient assessment, current guidelines, and appropriate specialist consultation. Editorial Standards · Privacy Policy · Terms of Service