Systematic review of machine learning and deep learning models for EEG-based detection of depression.
Clinical Snapshot
PICO Framework
| P — Population | Human participants with depression (or suspected depression) undergoing EEG-based assessment, drawn from studies published 2020–2024 |
| I — Intervention | Machine learning (ML) or deep learning (DL) algorithms applied to quantitative EEG (QEEG) signals for depression detection |
| C — Comparator | Comparisons between ML-based and DL-based models; implicit comparison against conventional clinical diagnostic approaches (no active clinical comparator formally specified) |
| O — Outcomes | Classification accuracy, sensitivity, specificity, and other reported performance metrics (AUC, F1-score) for depression detection; risk of bias across included studies assessed via QUADAS-2 |
Bottom Line
This systematic review of 42 studies examines ML and DL approaches to EEG-based depression detection, reporting a descriptive accuracy range of 76%–100% and a non-significant trend toward higher mean accuracy in DL versus ML models (93.92% vs. 90.78%). However, the review's conclusions must be interpreted with considerable caution. The absence of meta-analytic pooling, GRADE assessment, and protocol pre-registration substantially limits the strength of inference. Critically, the authors themselves identify that near-perfect accuracy values cluster in studies with small samples and internal-only validation — a pattern strongly suggestive of overfitting rather than genuine diagnostic performance. QUADAS-2 assessment revealed recurrent high risk of bias in patient selection and index test conduct. For Australian clinicians, this technology remains firmly investigational: no TGA-approved EEG-based ML/DL diagnostic tool exists for depression, no MBS pathway is established, and RACGP guidelines continue to recommend validated clinical instruments. The review makes a useful contribution by cataloguing methodological pitfalls in this literature and providing a clear agenda for future research — larger samples, subject-independent validation, standardised reporting — but does not provide evidence sufficient to influence clinical practice.
Key Findings
P Value: Non-significant in exploratory non-parametric comparison between ML and DL accuracy distributions (exact statistic and p-value not specified in abstract)
Effect Size: DL mean reported accuracy: 93.92%; ML mean reported accuracy: 90.78%; overall accuracy range across all studies: approximately 76%–100%
Primary Outcome: Classification accuracy for EEG-based depression detection using ML or DL algorithms across 42 included studies (23 ML-based, 19 DL-based)
Nnt Or Sensitivity: Sensitivity and specificity not pooled; individual study values reported descriptively. No NNT calculable. SROC curve analysis not performed. QUADAS-2 identified high risk of bias in patient selection and index test domains across multiple studies.
Confidence Interval: Not reported — no pooled estimates or confidence intervals calculated; descriptive averages only
Clinical Application
EEG acquisition is feasible in specialist psychiatric, neurological, and sleep medicine settings, but is not routinely available in primary care. Clinical-grade QEEG systems with integrated ML/DL pipelines are not yet commercially validated or regulatory-approved for depression diagnosis. Implementation would require standardised EEG protocols, trained operators, and validated software — none of which are currently standardised across the included studies. In Australia, EEG is Medicare-rebatable (MBS item numbers 11000–11027) for neurological indications but is not currently indicated or rebated for psychiatric diagnosis of depression. The TGA has not approved any ML/DL-based EEG diagnostic device for depression. RACGP guidelines for depression management (including the RACGP Red Book and beyondblue clinical practice guidelines) rely on validated clinical instruments (PHQ-9, DASS-21) and structured clinical assessment — not biomarker-based tools. PBS-listed antidepressant prescribing does not require EEG confirmation. This technology remains investigational in the Australian context, and any future clinical translation would require TGA regulatory approval, MBS pathway development, and prospective validation in Australian populations. The review's findings do not alter current Australian clinical practice. Adults with suspected or confirmed depressive disorder in whom objective, scalable biomarker-based diagnostic adjuncts are being considered. The evidence base does not yet support application in any specific clinical subgroup (e.g., treatment-resistant depression, adolescents, elderly) given the heterogeneity and methodological limitations of included studies.
Abstract
OBJECTIVE: Depression is a leading cause of global disability, motivating the development of objective and scalable diagnostic approaches. Quantitative electroencephalography (QEEG) combined with machine learning (ML) and deep learning (DL) techniques has gained increasing attention for depression detection. This systematic review aimed to critically examine and descriptively compare the methodologies, performance metrics, and limitations of ML- and DL-based models applied to EEG data for depression detection. METHODS: A systematic review was conducted in accordance with PRISMA 2020 guidelines. Seven electronic databases (PubMed, Scopus, IEEE Xplore, ScienceDirect, Web of Science, SAGE Journals, and MDPI) were searched for peer-reviewed studies published between 2020 and 2024. Eligible studies included human participants, used EEG signals for depression detection, and applied ML or DL algorithms. Extracted information comprised algorithm type, sample size, EEG acquisition parameters, validation strategies, and reported performance metrics, which were synthesized descriptively across studies. Risk of bias was assessed using the Quality Assessment of Diagnostic Accuracy Studies-2 (QUADAS-2) tool. RESULTS: A total of 42 studies met the inclusion criteria, including 23 ML-based and 19 DL-based investigations. Reported classification accuracy ranged from approximately 76% to 100%. DL studies showed a higher mean reported accuracy than ML studies (93.92% vs. 90.78%); however, this difference was not statistically significant in the exploratory non-parametric comparison. Near-perfect performance values were frequently observed in studies with small sample sizes and subject-dependent or exclusively internal validation strategies, raising concerns regarding overfitting and limited generalizability. Studies relying on publicly available datasets tended to report more stable performance. QUADAS-2 assessment revealed recurrent risk-of-bias concerns, particularly in the domains of patient selection and index test conduct. CONCLUSIONS: Both ML and DL approaches demonstrate potential for EEG-based depression detection, but reported performance differences between them should be interpreted cautiously. Although DL studies tended to report higher accuracy values, this pattern was not statistically significant in exploratory analyses and was strongly influenced by sample size, validation strategy, and methodological design. Future research should prioritize larger and more diverse samples, subject-independent or external validation strategies, and standardized reporting frameworks to enhance methodological rigor and clinical applicability.
References
- 1.Luna-Guevara, G. R., Vargas-Huicochea, I., & Rodríguez-Machain, A. C. (2026). Systematic review of machine learning and deep learning models for EEG-based detection of depression. Journal of Psychiatric Research. https://doi.org/10.1016/j.jpsychires.2026.04.030
Related Research
Sleep medicine reviews
Artificial intelligence in sleep medicine I: Diagnosis, treatment, care, and research
2 Aug 2026
International journal of yoga therapy
Randomized Controlled Trial on the Effect of Yoga on Anxiety Severity, Quality of Life, and Biomarkers in Individuals with Anxiety Disorder
31 July 2026
BMC ophthalmology
The effect of lithium on the structure and function of the human retina: a systematic review
31 July 2026
This content is for educational purposes for healthcare professionals only and does not constitute clinical advice. Clinical decisions should be based on individual patient assessment, current guidelines, and appropriate specialist consultation. Editorial Standards · Privacy Policy · Terms of Service