Machine Learning-Based Prediction of Poor Outcomes in Intracerebral Hemorrhage: A Systematic Review and Meta-Analysis
Clinical Snapshot
PICO Framework
| P — Population | Adults with spontaneous intracerebral hemorrhage (ICH) |
| I — Intervention | Machine learning (ML) models incorporating clinical features, radiomics features, or both for prognostic prediction |
| C — Comparator | Comparisons between ML model subtypes (e.g., logistic regression vs. complex ML algorithms; clinical-only vs. clinical-radiomics models); internal vs. external validation performance |
| O — Outcomes | Discriminative performance (concordance index [C-index]), sensitivity, and specificity for three key adverse outcomes: hematoma expansion (HE), poor functional outcome (modified Rankin Scale 3–6), and mortality |
Bottom Line
This systematic review and meta-analysis of 83 studies (≥136,840 patients) reports strong pooled discriminative performance for ML-based prognostic models in spontaneous ICH, with C-indices of 0.82–0.86 for hematoma expansion, poor functional outcome, and mortality prediction. Integrated clinical-radiomics models performed best across all outcome domains. However, these results must be interpreted with considerable caution. The pooled estimates derive predominantly from internal validation sets, which are known to systematically overestimate real-world performance. True external validation remains sparse. Heterogeneity statistics are not reported despite the use of random-effects models, and no GRADE certainty assessment is provided. Publication bias and optimistic reporting in the primary ML literature are not formally evaluated. Notably, logistic regression performed comparably to more complex algorithms — a finding with important implications for implementation, favouring transparent and auditable models. For Australian clinicians, no ML-based ICH prognostic tool is currently TGA-approved or endorsed by the Stroke Foundation Clinical Guidelines. The evidence base supports continued research investment but does not yet justify routine clinical deployment. Prospective, multicentre external validation studies — ideally including Australian cohorts — are the essential next step before these models can meaningfully inform ICH triage, goals-of-care discussions, or treatment escalation decisions.
Key Findings
P Value: Not reported in abstract
Effect Size: Clinical-radiomics integrated models: HE C-index 0.822; poor functional outcome C-index 0.850; mortality C-index 0.860. Logistic regression performance was comparable to more complex ML algorithms across outcome domains.
Primary Outcome: Discriminative performance (pooled C-index) of ML models for three ICH outcomes: hematoma expansion, poor functional outcome (mRS 3–6), and mortality
Nnt Or Sensitivity: Pooled sensitivity and specificity calculated (bivariate model) but specific values not reported in abstract. C-index values in the 0.82–0.86 range indicate good-to-strong discrimination, though these are predominantly internal validation estimates and likely overstate external performance.
Confidence Interval: HE: 95% CI 0.789–0.855; Poor functional outcome: 95% CI 0.830–0.869; Mortality: 95% CI 0.809–0.911
Clinical Application
Clinical-only ML models (including logistic regression) are immediately feasible in most acute settings using routinely collected variables. Radiomics-integrated models require dedicated imaging post-processing infrastructure and expertise, limiting near-term implementation to tertiary centres. The comparable performance of logistic regression to complex ML algorithms suggests that simpler, more transparent models may be preferable for clinical deployment pending prospective validation. In Australia, ICH management follows RACGP and Stroke Foundation Clinical Guidelines for Stroke Management, which emphasise early neurosurgical review, blood pressure control, and reversal of anticoagulation. Validated prognostic tools could support triage decisions, ICU admission thresholds, and goals-of-care conversations. However, no ML-based ICH prognostic model is currently TGA-approved as a clinical decision support device in Australia. Radiomics platforms are not universally available across Australian public hospitals, and PBS does not fund specific radiomics analyses. Any future implementation would require prospective validation in Australian cohorts, given differences in case-mix, Indigenous health burden, and healthcare system structure. The Stroke Foundation's 2023 Clinical Guidelines do not currently endorse ML-based prognostic tools for routine ICH management. Adults presenting with spontaneous (non-traumatic) intracerebral hemorrhage in acute hospital settings. Most applicable to centres with access to CT imaging and radiomics analysis capability. Generalisability to community hospitals, resource-limited settings, or specific ICH subtypes (e.g., lobar vs. deep) requires further study.
Abstract
BACKGROUND: Spontaneous intracerebral hemorrhage (ICH) is associated with high risks of mortality and disability, yet early and accurate outcome prediction remains challenging. This study systematically evaluated the performance of machine learning (ML) models in predicting key adverse outcomes (hematoma expansion [HE], poor functional outcome, mortality) in ICH, aiming to provide consolidated evidence for future research and clinical translation. METHODS: We systematically searched PubMed, Embase, Web of Science, and Cochrane Library up to September 2025. Studies developing and validating ML models for predicting HE, poor functional outcome (modified Rankin Scale 3-6), or mortality in adults with spontaneous ICH were included. Pooled concordance index (C-index), sensitivity, and specificity were calculated using a random-effects or bivariate model. RESULTS: Eighty-three studies (involving at least 136,840 patients) were included. Meta-analysis of model performance, derived predominantly from internal validation set, demonstrated that models integrating both clinical and radiomics features achieved the highest discriminative performance across key prognostic prediction tasks: predicting HE (pooled C-index 0.822, 95% confidence interval [CI] 0.789-0.855), poor functional outcome (C-index 0.850, 95% CI 0.830-0.869), and mortality (C-index 0.860, 95% CI 0.809-0.911). These findings should be interpreted with caution, as true external validation remains sparse. Logistic regression exhibited performance comparable to more complex ML algorithms. CONCLUSIONS: ML models, particularly integrated clinical-radiomics models, demonstrate strong performance for the prediction of outcomes in ICH. They hold significant potential to enhance risk stratification and guide personalized management, pending further validation in diverse cohorts.
References
- 1.Deng, Q., Cheng, W., Lu, H., Fei, N., He, R., Chen, Y., Bi, J., & Zhang, W. (2026). Machine learning-based prediction of poor outcomes in intracerebral hemorrhage: A systematic review and meta-analysis. Brain and Behavior. https://doi.org/10.1002/brb3.71572
Related Research
Sleep medicine reviews
Machine and deep learning in REM sleep behavior disorder: a scoping review and analysis of reporting quality
3 Aug 2026
Neurological research
Federated deep learning model for epilepsy seizure detection using electroencephalogram (EEG) signal
2 Aug 2026
Journal of psychiatric research
Systematic review of machine learning and deep learning models for EEG-based detection of depression.
2 Aug 2026
This content is for educational purposes for healthcare professionals only and does not constitute clinical advice. Clinical decisions should be based on individual patient assessment, current guidelines, and appropriate specialist consultation. Editorial Standards · Privacy Policy · Terms of Service