Machine and deep learning in REM sleep behavior disorder: a scoping review and analysis of reporting quality
Clinical Snapshot
PICO Framework
| P — Population | Studies applying machine learning (ML) or deep learning (DL) models in the context of REM sleep behavior disorder (RBD), including isolated RBD (iRBD) as an early alpha-synucleinopathy |
| I — Intervention | Machine learning and deep learning models applied to RBD detection, phenoconversion prediction, and phenotyping |
| C — Comparator | No direct comparator (scoping review of methodological and reporting quality); implicit comparison against established reporting standards (APPRAISE-AI tool criteria) |
| O — Outcomes | Methodological and reporting quality of included studies as assessed by APPRAISE-AI scores; prevalence of specific deficiencies including data leakage, transparency in data/code sharing, hyperparameter reporting, bias assessment, and error analysis |
Bottom Line
This scoping review systematically maps 75 ML and deep learning studies applied to REM sleep behavior disorder and rigorously evaluates their methodological quality using the APPRAISE-AI tool. The findings are sobering: 80% of studies achieved only moderate quality scores, data leakage was present in nearly one-third of studies, and error analysis — a fundamental component of model evaluation — was reported in fewer than 1% of papers. Most studies relied on very small samples (20–100 participants), and critical reporting elements including data sharing, code availability, and bias assessment were routinely absent. These deficiencies collectively prevent meaningful cross-study comparison, replication, and clinical translation. The review does not endorse any specific ML/DL tool for clinical use; rather, it provides a clear-eyed account of why none currently meets the bar for implementation. For Australian sleep medicine clinicians and researchers, the message is unambiguous: existing AI tools for RBD detection and phenoconversion prediction should not be adopted into clinical practice without independent external validation on adequately powered, representative datasets. Future studies must embrace open science principles, rigorous validation frameworks, and transparent reporting to realise the genuine potential of AI in this clinically important prodromal neurodegenerative condition.
Key Findings
Effect Size: Not applicable (scoping review of reporting quality); descriptive findings: 73.3% of studies focused on RBD detection; 16% addressed phenoconversion prediction; polysomnography was the dominant data modality for detection; imaging data most used for phenoconversion
Primary Outcome: Methodological and reporting quality of 75 ML/DL studies in RBD assessed via APPRAISE-AI: 80% of studies achieved only moderate overall quality scores
Nnt Or Sensitivity: Key deficiency rates: data leakage present in 32% of studies; lack of data/code transparency in 23.3%; poor hyperparameter tuning reporting in 17.1%; inadequate bias assessment in 26.9%; error analysis reported in only 0.66% of studies
Confidence Interval: Not reported; point estimates only provided for deficiency rates
Clinical Application
The review does not introduce a new clinical tool but provides a critical framework for evaluating existing and future ML/DL tools in RBD. Clinicians and researchers can use the identified deficiency checklist (data leakage, reporting transparency, bias assessment) as a practical filter when assessing published AI tools for potential clinical adoption. Feasibility of implementing any specific ML/DL tool remains limited by the small sample sizes and methodological flaws identified across the literature. In Australia, RBD diagnosis currently relies on attended video-polysomnography per AASM criteria, with no TGA-approved or PBS-listed ML/DL diagnostic tool for RBD. The RACGP and Australasian Sleep Association (ASA) have not yet issued specific guidance on AI-assisted RBD diagnosis. This review reinforces that no currently published ML/DL tool meets the methodological standards required for TGA regulatory submission or clinical guideline endorsement. Australian sleep medicine researchers contributing to this field should adopt APPRAISE-AI reporting standards and open science practices. The identification of isolated RBD as a prodromal synucleinopathy is particularly relevant given Australia's growing interest in neuroprotective trial recruitment, where accurate early identification tools would have significant public health value. Sleep medicine clinicians, neurologists, and clinical researchers evaluating or developing ML/DL tools for RBD diagnosis, risk stratification, and phenoconversion prediction in patients with suspected or confirmed RBD or prodromal alpha-synucleinopathy
Abstract
Rapid eye movement (REM) sleep behavior disorder (RBD) is a parasomnia, and its isolated form is of particular interest, as it is an early phase alpha-synucleinopathy. Machine learning (ML) and deep learning (DL) models offer potential for automated detection, prediction of phenoconversion, and phenotyping. This scoping review identified 75 studies applying ML/DL in RBD and evaluated their methodological and reporting quality using the APPRAISE-AI tool. Most studies (73.3%) focused on RBD detection and mainly used polysomnographic data for this, while 16% addressed prediction of phenoconversion, with imaging data being the most employed modality. Sample sizes were generally small (most studies including only 20-100 individuals). According to APPRAISE-AI scores, 80% of studies had moderate overall methodological and reporting quality. Common deficiencies included lack of transparency in data and code sharing (23.3%), and poor reporting of hyperparameter tuning (17.1%), bias assessment (26.9%), and error analysis (0.66%). Data leakage was observed in 32% of studies. These issues hinder clinical translation and prevent incremental progress between research groups. Without transparent reporting and shared resources, replication and model comparisons become nearly impossible. Future work should adopt open science principles and rigorous validation to advance AI-based tools in sleep medicine.
References
- 1.Brink-Kjaer, A., Rechichi, I., Arnaldi, D., Cesarone, O., During, E., Feuerstein, S., Gunter, K. M., Högl, B., Ibrahim, A., Jennum, P., Mignot, E., Olmo, G., Roascio, M., Stefani, A., Tang, Q., & Cesari, M. (2026). Machine and deep learning in REM sleep behavior disorder: a scoping review and analysis of reporting quality. Sleep Medicine Reviews. https://doi.org/10.1016/j.smrv.2026.102299
Related Research
Sleep medicine
Concordance of wearable device sleep metrics with patient-reported sleep quality: A systematic review
3 Aug 2026
Sleep medicine reviews
Artificial intelligence in sleep medicine I: Diagnosis, treatment, care, and research
2 Aug 2026
JMIR mHealth and uHealth
Wearable Devices for Monitoring and Management of Comorbid Obstructive Sleep Apnea and Hypertension: Scoping Review
1 Aug 2026
This content is for educational purposes for healthcare professionals only and does not constitute clinical advice. Clinical decisions should be based on individual patient assessment, current guidelines, and appropriate specialist consultation. Editorial Standards · Privacy Policy · Terms of Service