Research AppraisalSystematic Review

Concordance of wearable device sleep metrics with patient-reported sleep quality: A systematic review

Sleep medicineSrivali, Narat, Cheungpasitporn, Wisit1 Aug 2026DOI

Clinical Snapshot

60CEBM
Evidence: ModerateSystematic Review

PICO Framework

P — PopulationAdults (≥18 years) across general and clinical populations, including good sleepers, insomnia patients, and older adults
I — InterventionConsumer wrist-worn actigraphy-based wearable devices (commercial sleep trackers) measuring objective sleep metrics
C — ComparatorValidated subjective sleep quality measures (e.g., Pittsburgh Sleep Quality Index, sleep diaries) and polysomnography as reference standard
O — OutcomesConcordance between wearable-derived sleep metrics (total sleep time, sleep efficiency, sleep onset latency, wake after sleep onset) and patient-reported sleep quality; agreement statistics including correlation coefficients, intraclass correlation coefficients, and variance explained

Bottom Line

This systematic review of five observational studies (n=2,006) found that commercial wrist-worn wearable devices demonstrate poor to moderate concordance with validated subjective sleep quality measures. Devices explained only 2.5–16.2% of variance in subjective sleep scores, with total sleep time showing only moderate correlation with sleep diaries (r=0.367). Against polysomnography, wearables systematically overestimated sleep efficiency (+1.75% to +7.9%) and underestimated wake after sleep onset (−7 to −30 minutes), with sleep onset latency correlation near zero (r=0.033). Critically, concordance was substantially worse in insomnia patients (39.4%) compared with good sleepers (82.4%), and similarly poor in older adults and clinical populations — precisely the groups where accurate sleep assessment matters most. GRADE certainty was low to moderate. The evidence base is small and no pooled meta-analytic estimates with confidence intervals were generated, limiting precision. For Australian clinicians, the practical message is clear: wearable sleep data should be treated as supplementary information only. Validated instruments — the PSQI, Epworth Sleepiness Scale, and structured sleep diaries — remain the appropriate clinical standard. Patients presenting wearable data showing apparently normal sleep should not have clinical concern dismissed on that basis alone.

Evidence: Moderate

Key Findings

  • P Value: Population-specific concordance difference: p = 0.006 (good sleepers vs. insomnia patients)

  • Effect Size: Wearable devices explained only 2.5–16.2% of variance in subjective sleep quality scores. Total sleep time correlated moderately with same-day sleep diaries (r = 0.367). Sleep efficiency ICC with PSG: 0.478–0.570 (fair). Sleep onset latency correlated poorly with PSG (r = 0.033). Concordance varied by population: 82.4% in good sleepers vs. 39.4% in insomnia patients

  • Primary Outcome: Concordance between commercial wrist-worn actigraphy-based wearable devices and validated subjective sleep quality measures in adults

  • Nnt Or Sensitivity: Diagnostic agreement metrics: Sleep efficiency ICC 0.478–0.570 (fair agreement); systematic overestimation of sleep efficiency +1.75% to +7.9%; WASO underestimation −7 to −30 minutes; SOL correlation r = 0.033 (poor). These represent clinically meaningful systematic biases rather than random error.

  • Confidence Interval: Not reported for individual metrics; ranges provided across studies

Clinical Application

Commercial wearable devices are widely accessible and increasingly used by patients independently of clinical recommendation. The findings are immediately applicable to clinical consultations where patients present wearable sleep data. No additional resources are required to implement the recommendation that wearable data should supplement — not replace — validated tools such as the PSQI or structured sleep diaries. In Australia, commercial wearable sleep trackers (Fitbit, Apple Watch, Garmin, Oura Ring) are widely used by the general population and are not subject to TGA regulatory oversight as medical devices in most configurations. The RACGP supports the use of validated subjective tools such as the PSQI and Epworth Sleepiness Scale in primary care sleep assessment. This review reinforces RACGP guidance that wearable consumer devices should not substitute for validated clinical assessment instruments. Sleep studies (polysomnography) are Medicare-rebatable under specific criteria (MBS items 12203–12215), and the systematic biases identified — particularly WASO underestimation and sleep efficiency overestimation — are relevant to triage decisions for PSG referral. Clinicians should be cautious when patients present wearable data as evidence against a sleep disorder, particularly in insomnia and older adult populations where concordance is poorest. Adults across primary care, sleep medicine, and general medical settings who use or are considering commercial wearable sleep trackers. Particularly relevant for clinical populations including insomnia patients, older adults, and those with comorbid sleep disorders where device performance is demonstrably poorer.

Abstract

BACKGROUND: Commercial wearable devices increasingly monitor sleep, but their concordance with patient-reported sleep quality remains poorly characterized. This systematic review evaluates concordance between wearable sleep metrics and validated subjective measures. METHODS: Following PRISMA guidelines, we searched PubMed/MEDLINE, Embase, and Cochrane Library through October 2025 for studies comparing consumer wrist-worn actigraphy-based devices with validated subjective sleep quality measures in adults (≥18 years). Two reviewers independently performed screening, extraction, and QUADAS-2 quality assessment. GRADE criteria evaluated evidence certainty. RESULTS: Five observational studies (2006 participants) were included. Wearable devices showed poor to moderate agreement with subjective assessments, explaining only 2.5-16.2% variance. Total sleep time moderately correlated with same-day diaries (r = 0.367), but devices failed to capture Pittsburgh Sleep Quality Index scores. Agreement varied substantially by population: good sleepers showed 82.4% concordance versus 39.4% in insomnia patients (p = 0.006). Clinical populations and older adults demonstrated poor agreement. Polysomnography concordance was also poor: sleep efficiency showed fair intraclass correlation coefficient values (0.478-0.570) with systematic overestimation (+1.75% to +7.9%), sleep onset latency correlated poorly (r = 0.033), and wake after sleep onset was underestimated (-7 to -30 min). Evidence certainty ranged from low to moderate. CONCLUSIONS: Commercial wearable sleep trackers demonstrate poor to moderate agreement with validated subjective sleep quality measures, with significant population-specific variation. Device data should complement, not replace, validated subjective assessments, as current technology inadequately captures patient-reported sleep quality and shows systematic bias in objective parameters.

References

  1. 1.Srivali, N., & Cheungpasitporn, W. (2026). Concordance of wearable device sleep metrics with patient-reported sleep quality: A systematic review. Sleep Medicine. https://doi.org/10.1016/j.sleep.2026.108941
Share:XLinkedIn

This content is for educational purposes for healthcare professionals only and does not constitute clinical advice. Clinical decisions should be based on individual patient assessment, current guidelines, and appropriate specialist consultation. Editorial Standards · Privacy Policy · Terms of Service