Research AppraisalSystematic Review

Optimizing Treatment Strategies in the Bipolar Disorder Spectrum With Classical AI Approaches: Systematic Review of Performance, Bias, and Clinical Applicability

JMIR mental healthDe Francesco, Silvia, Archetti, Damiano, Baronio, Cesare Michele et al.21 July 2026DOI

Clinical Snapshot

60CEBM
Evidence: ModerateSystematic Review

PICO Framework

P — PopulationAdult patients with bipolar disorder (BD) spectrum conditions
I — InterventionClassical artificial intelligence approaches (machine learning, regression-based predictive models, rule-based systems, digital phenotyping models) applied to treatment optimisation
C — ComparatorImplicit comparator: standard clinical decision-making without AI-assisted prediction; no active comparator arm in included studies
O — OutcomesPredictive model performance (AUC, accuracy) across five outcome categories: acute symptomatic response, long-term maintenance response, relapse and readmission risk, safety and dose optimisation, and brain aging/phenotyping

Bottom Line

This systematic review of 35 studies synthesises evidence on classical AI predictive models for treatment optimisation in bipolar disorder. Pooled discriminative performance ranges from modest (AUC 0.68 for acute symptomatic response) to moderate (AUC 0.80 for long-term maintenance), with select biomarker and digital phenotyping models reporting higher accuracy. However, the PROBAST+AI assessment reveals high risk of bias in the majority of included studies, driven by small sample sizes, overfitting, and near-universal absence of external validation. Three studies show implausibly high performance likely attributable to these limitations. Confidence intervals around pooled estimates are not reported, and formal heterogeneity statistics are absent, limiting interpretation. No GRADE assessment was conducted. The authors appropriately conclude that current AI tools in bipolar disorder remain exploratory and are not ready for clinical deployment. For Australian psychiatrists, this review reinforces that AI-assisted treatment decision support in bipolar disorder is a promising but immature field. Clinicians should maintain critical appraisal of commercially promoted AI tools and await prospective, externally validated studies before incorporating such approaches into practice. Lithium's apparent neuroprotective signal in brain aging studies warrants further investigation.

Evidence: Moderate

Key Findings

  • P Value: Not reported

  • Effect Size: Pooled AUC: acute symptomatic response 0.68; long-term maintenance response 0.80; relapse and readmission prediction 0.71. Safety/dose optimisation accuracy 85–97%. Imaging-enhanced acute response models AUC 0.74–0.77. Digital phenotyping/rule-based relapse models AUC 0.85–0.88. Biomarker/cellular maintenance models accuracy 96–99%.

  • Primary Outcome: Predictive model performance (AUC or accuracy) across five BD treatment outcome categories

  • Nnt Or Sensitivity: Sensitivity and specificity not reported in the abstract. No NNT calculable from available data. AUC of 0.68 for acute response corresponds to modest discriminative ability marginally above chance (AUC 0.50). AUC of 0.80 for maintenance response represents moderate-to-good discrimination by conventional thresholds.

  • Confidence Interval: Not reported in the abstract

Clinical Application

Currently not feasible for routine clinical implementation. Most models lack external validation, are derived from small and potentially non-representative samples, and have not been prospectively tested in clinical workflows. Integration into electronic health record systems or clinical decision support tools would require substantial further development, regulatory approval, and prospective validation. The high risk of bias identified by PROBAST+AI across most studies means that reported performance metrics should not be taken at face value for clinical planning. In the Australian context, bipolar disorder management is guided by the Royal Australian and New Zealand College of Psychiatrists (RANZCP) Clinical Practice Guidelines for Mood Disorders (2020). Lithium remains a first-line mood stabiliser and is PBS-listed; the finding that lithium partially mitigates accelerated brain aging in BD is of particular interest to Australian prescribers. The TGA has not approved any AI-based clinical decision support tool specifically for BD treatment optimisation. The Australian Commission on Safety and Quality in Health Care's framework for AI in healthcare would require prospective clinical validation before any such tool could be recommended. Digital phenotyping approaches (AUC 0.85–0.88 for relapse prediction) may have future relevance to Australian telehealth and remote monitoring contexts, particularly for rural and Indigenous populations with limited specialist access, but this remains speculative pending robust validation studies. Adults with bipolar disorder spectrum conditions being considered for pharmacological treatment optimisation, particularly those with treatment-refractory courses or complex polypharmacy. Models incorporating biomarkers, neuroimaging, or digital phenotyping data may be most applicable to specialist psychiatric settings with access to these data streams.

Abstract

BACKGROUND: Bipolar disorder (BD) is a complex and heterogeneous psychiatric condition, characterized by fluctuating clinical courses that affect approximately 1%-2% of the global population in their lifetime. Despite pharmacological advances, treatment response varies significantly among patients, making the identification of individualized treatment strategies a major challenge. Artificial Intelligence (AI), through its classical approaches, has emerged as a powerful tool in precision psychiatry to identify subtle patterns in complex data and inform personalized clinical decisions. OBJECTIVE: The present systematic review aimed to examine the current evidence on classical AI-supported treatment optimization in the BD spectrum. METHODS: The review was conducted in accordance with the PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) 2020 guidelines. Four databases (PubMed, Web of Science, Scopus, and Embase) were searched for original studies published after 2015 on the application of classical AI in the treatment of BD in adult patients. Publication bias was evaluated by visual inspection of a funnel plot. The methodological quality, risk of bias, and clinical applicability of the predictive models were assessed using the Prediction Model Risk Of Bias Assessment Tool for prediction models using regression or AI methods (PROBAST+AI; PROBAST+AI Working Group) tool. RESULTS: A total of 35 studies were included and classified into 5 outcome-based categories, including acute symptomatic response, long-term maintenance response, relapse and readmission risk, safety and dose optimization, and brain aging and phenotyping. Acute symptomatic response models performed modestly (pooled area under the curve [AUC] 0.68), while imaging improved accuracy (74%-77%). Long-term maintenance response models showed moderate-to-high performance (pooled AUC 0.80), with biomarker- and cellular-based models reaching 96%-99% accuracy. Relapse and readmission prediction achieved a pooled AUC of 0.71, with digital phenotyping and rule-based methods performing best (AUC 0.85-0.88). Safety and dose optimization models achieved 85%-97% accuracy. Brain aging and phenotyping studies highlighted accelerated brain aging in BD, partially mitigated by lithium, and revealed novel data-driven subgroups. However, 3 studies were considered at high risk of bias due to small sample sizes associated with disproportionately high-performance estimates. An additional study was identified as potentially biased because it lay markedly distant from the funnel plot's confidence line. Finally, the PROBAST+AI assessment revealed a high risk of bias in most studies, primarily due to data analysis limitations, small sample sizes, and lack of external validation. CONCLUSIONS: The adoption of classical AI tools in BD serves as a driver for therapeutic optimization, although current AI tools in BD should still be considered exploratory rather than ready for clinical use. Effective implementation in real-world clinical scenarios requires more robust, transparent, and externally validated models to ensure reliability and generalizability.

References

  1. 1.De Francesco, S., Archetti, D., Baronio, C. M., Demaria, C., Boccali, A., Crema, C., Tura, G. B., & Redolfi, A. (2026). Optimizing treatment strategies in the bipolar disorder spectrum with classical AI approaches: Systematic review of performance, bias, and clinical applicability. JMIR Mental Health. https://doi.org/10.2196/93307
Share:XLinkedIn

This content is for educational purposes for healthcare professionals only and does not constitute clinical advice. Clinical decisions should be based on individual patient assessment, current guidelines, and appropriate specialist consultation. Editorial Standards · Privacy Policy · Terms of Service