Research AppraisalSystematic Review

Machine and deep learning in REM sleep behavior disorder: a scoping review and analysis of reporting quality

Sleep medicine reviewsBrink-Kjaer, Andreas, Rechichi, Irene, Arnaldi, Dario et al.1 Aug 2026DOI

Clinical Snapshot

60CEBM
Evidence: ModerateSystematic Review

PICO Framework

P — PopulationStudies applying machine learning (ML) or deep learning (DL) models in the context of REM sleep behavior disorder (RBD), including isolated RBD (iRBD) as an early alpha-synucleinopathy
I — InterventionMachine learning and deep learning models applied to RBD detection, phenoconversion prediction, and phenotyping
C — ComparatorNo direct comparator (scoping review of methodological and reporting quality); implicit comparison against established reporting standards (APPRAISE-AI tool criteria)
O — OutcomesMethodological and reporting quality of included studies as assessed by APPRAISE-AI scores; prevalence of specific deficiencies including data leakage, transparency in data/code sharing, hyperparameter reporting, bias assessment, and error analysis

Bottom Line

This scoping review systematically maps 75 ML and deep learning studies applied to REM sleep behavior disorder and rigorously evaluates their methodological quality using the APPRAISE-AI tool. The findings are sobering: 80% of studies achieved only moderate quality scores, data leakage was present in nearly one-third of studies, and error analysis — a fundamental component of model evaluation — was reported in fewer than 1% of papers. Most studies relied on very small samples (20–100 participants), and critical reporting elements including data sharing, code availability, and bias assessment were routinely absent. These deficiencies collectively prevent meaningful cross-study comparison, replication, and clinical translation. The review does not endorse any specific ML/DL tool for clinical use; rather, it provides a clear-eyed account of why none currently meets the bar for implementation. For Australian sleep medicine clinicians and researchers, the message is unambiguous: existing AI tools for RBD detection and phenoconversion prediction should not be adopted into clinical practice without independent external validation on adequately powered, representative datasets. Future studies must embrace open science principles, rigorous validation frameworks, and transparent reporting to realise the genuine potential of AI in this clinically important prodromal neurodegenerative condition.

Evidence: Moderate

Key Findings

  • Effect Size: Not applicable (scoping review of reporting quality); descriptive findings: 73.3% of studies focused on RBD detection; 16% addressed phenoconversion prediction; polysomnography was the dominant data modality for detection; imaging data most used for phenoconversion

  • Primary Outcome: Methodological and reporting quality of 75 ML/DL studies in RBD assessed via APPRAISE-AI: 80% of studies achieved only moderate overall quality scores

  • Nnt Or Sensitivity: Key deficiency rates: data leakage present in 32% of studies; lack of data/code transparency in 23.3%; poor hyperparameter tuning reporting in 17.1%; inadequate bias assessment in 26.9%; error analysis reported in only 0.66% of studies

  • Confidence Interval: Not reported; point estimates only provided for deficiency rates

Clinical Application

The review does not introduce a new clinical tool but provides a critical framework for evaluating existing and future ML/DL tools in RBD. Clinicians and researchers can use the identified deficiency checklist (data leakage, reporting transparency, bias assessment) as a practical filter when assessing published AI tools for potential clinical adoption. Feasibility of implementing any specific ML/DL tool remains limited by the small sample sizes and methodological flaws identified across the literature. In Australia, RBD diagnosis currently relies on attended video-polysomnography per AASM criteria, with no TGA-approved or PBS-listed ML/DL diagnostic tool for RBD. The RACGP and Australasian Sleep Association (ASA) have not yet issued specific guidance on AI-assisted RBD diagnosis. This review reinforces that no currently published ML/DL tool meets the methodological standards required for TGA regulatory submission or clinical guideline endorsement. Australian sleep medicine researchers contributing to this field should adopt APPRAISE-AI reporting standards and open science practices. The identification of isolated RBD as a prodromal synucleinopathy is particularly relevant given Australia's growing interest in neuroprotective trial recruitment, where accurate early identification tools would have significant public health value. Sleep medicine clinicians, neurologists, and clinical researchers evaluating or developing ML/DL tools for RBD diagnosis, risk stratification, and phenoconversion prediction in patients with suspected or confirmed RBD or prodromal alpha-synucleinopathy

Abstract

Rapid eye movement (REM) sleep behavior disorder (RBD) is a parasomnia, and its isolated form is of particular interest, as it is an early phase alpha-synucleinopathy. Machine learning (ML) and deep learning (DL) models offer potential for automated detection, prediction of phenoconversion, and phenotyping. This scoping review identified 75 studies applying ML/DL in RBD and evaluated their methodological and reporting quality using the APPRAISE-AI tool. Most studies (73.3%) focused on RBD detection and mainly used polysomnographic data for this, while 16% addressed prediction of phenoconversion, with imaging data being the most employed modality. Sample sizes were generally small (most studies including only 20-100 individuals). According to APPRAISE-AI scores, 80% of studies had moderate overall methodological and reporting quality. Common deficiencies included lack of transparency in data and code sharing (23.3%), and poor reporting of hyperparameter tuning (17.1%), bias assessment (26.9%), and error analysis (0.66%). Data leakage was observed in 32% of studies. These issues hinder clinical translation and prevent incremental progress between research groups. Without transparent reporting and shared resources, replication and model comparisons become nearly impossible. Future work should adopt open science principles and rigorous validation to advance AI-based tools in sleep medicine.

References

  1. 1.Brink-Kjaer, A., Rechichi, I., Arnaldi, D., Cesarone, O., During, E., Feuerstein, S., Gunter, K. M., Högl, B., Ibrahim, A., Jennum, P., Mignot, E., Olmo, G., Roascio, M., Stefani, A., Tang, Q., & Cesari, M. (2026). Machine and deep learning in REM sleep behavior disorder: a scoping review and analysis of reporting quality. Sleep Medicine Reviews. https://doi.org/10.1016/j.smrv.2026.102299
Share:XLinkedIn

This content is for educational purposes for healthcare professionals only and does not constitute clinical advice. Clinical decisions should be based on individual patient assessment, current guidelines, and appropriate specialist consultation. Editorial Standards · Privacy Policy · Terms of Service