What Your Sleep Tracker Can and Cannot Tell You
- Consumer wearables infer sleep from movement and peripheral signals; PSG scores stages from EEG and related channels the watch does not record.
- Devices are generally better at detecting sleep than quiet wake; insomnia patients often lose that argument with the wrist.
- Multi-night schedule trends are more trustworthy than nightly proprietary scores.
- AASM: consumer sleep technology should not diagnose or treat sleep disorders without validation and clinical context.[1]
- Orthosomnia: chasing perfect scores can worsen sleep anxiety; remove the score if bedtime became an exam.
A wearable can show schedule drift, short nights, irregular timing, and changes with travel or illness that memory misses. It cannot turn a wrist sensor into an EEG, diagnose a sleep disorder from a proprietary score, or prove the brain got exactly 47 minutes of deep sleep because an app painted a purple bar.
AASM makes the distinction clear: consumer sleep technology may support engagement, but it should not diagnose or treat sleep disorders without validation and clinical context.[1] Helpful and authoritative are different jobs. Most problems start when those jobs are collapsed.
Most consumer wearables infer sleep from movement, heart rate, HRV, temperature, and sometimes oxygen-related measures. Laboratory polysomnography records brain electrical activity, eye movements, muscle tone, respiratory effort, airflow, oxygen saturation, cardiac rhythm, and often leg movements. Stages are scored from signals consumer devices generally do not record. The wearable is solving a different problem by inference.
Algorithms can estimate sleep from movement and cardiovascular patterns. Quiet wakefulness may be labeled sleep. Restless sleep may be labeled wake. Stages may be misclassified because peripheral signals overlap. Software updates can "improve" deep sleep while biology is unchanged. Treat the tracker as a model of sleep, not a recording.
Devices tend to detect sleep better than wake. Comparing seven consumer devices with PSG, sleep sensitivity was generally high while wake specificity was more variable and often lower; quiet wake was frequently mistaken for sleep.[2] That asymmetry matters in insomnia. Lying motionless and awake for 45 minutes may be scored as sleep. Prefer the conscious report for that dispute.
Stage estimates need more humility. Some newer platforms show improved performance against ambulatory PSG in selected healthy adults;[3] funding source belongs in the weighing. Good validation does not mean every stage estimate is correct for every person every night. Performance shifts with disrupted sleep, movement, arrhythmia, contact, illness, and population.
The most useful wearable output is usually boring: bedtime drifted 90 minutes over a month; work nights average six hours and weekends seven and a half; resting heart rate rose with illness and recovered. Less useful: Tuesday scored 71 and Wednesday 76. Proprietary scores combine inputs through company formulas; a five-point difference usually lacks established clinical meaning. Precision of display is not precision of measurement.
Multi-night trends buffer ordinary variability (late dinner, exercise, menstrual symptoms, alcohol, stress, travel, chance). One unusual night should rarely drive a decision. When testing a change, compare enough nights on each side for signal to separate from noise.
A wearable may show fragmentation. It cannot say why. OSA, RLS/PLMD, insomnia, hot flashes, pain, alcohol, reflux, circadian misalignment, medications, and environment all fragment sleep and need different treatment. Devices can also miss important disease. Reassuring scores do not close a case with snoring, witnessed apnea, gasping, EDS, drowsy driving, persistent insomnia, morning headache, or unexplained fatigue. Alarming scores in a well-rested person do not open one.
Useful curiosity can become performance anxiety. Orthosomnia describes preoccupation with perfect wearable data and rising sleep anxiety.[4] Retroactively feeling worse after a bad score, staying in bed to improve recovery metrics, or changing supplements every three nights because a graph moved are opposite of the intended use. People with insomnia are especially vulnerable. If the tracker informs without increasing anxiety, keep it. If bedtime became an exam, hide the score or remove the device for a week.
Ask modest questions: when sleep occurs, schedule regularity, sustained trends, multi-night change. Do not ask the device to settle apnea diagnosis, REM adequacy, insomnia resolution, or whether last night was medically good because a number said 84. Persistent disagreement between device and experience plus significant symptoms means evaluate the human, not buy a second tracker for a vote.
For clinicians: deep diveMechanism, evidence, and clinical reasoning. Select to expand.
Position statement
Khosla et al. AASM: consumer sleep tech for engagement, not diagnosis/treatment without validation and context.[1]
Validation highlights
Chinoy et al.: seven devices vs PSG; high sleep sensitivity, weaker wake specificity.[2] Svensson et al.: Oura Gen3 vs multi-night ambulatory PSG in 96 participants; improved staging metrics in that sample; industry support disclosed.[3]
Clinical use
Use trends for schedule counseling and adherence narratives. Do not let scores override history for OSA/insomnia. Screen for orthosomnia when wearables increase arousal or bedtime effort.[4]
New-page note
No prior wakesleep KB article on trackers (inventory gap). Publish as new slug sleep-trackers-what-they-can-tell-you. Cross-link sleep diary, OSA overview, insomnia entry pages.
[1] Khosla S, Deak MC, Gault D, et al. Consumer sleep technology: an American Academy of Sleep Medicine position statement. J Clin Sleep Med. 2018;14(5):877-880. doi:10.5664/jcsm.7128. PMID: 29734997.
[2] Chinoy ED, Cuellar JA, Huwa KE, et al. Performance of seven consumer sleep-tracking devices compared with polysomnography. Sleep. 2021;44(5). doi:10.1093/sleep/zsaa291. PMID: 33378539.
[3] Svensson T, Madhawa K, Nt H, Chung UI, Svensson AK. Validity and reliability of the Oura Ring Generation 3 with Oura sleep staging algorithm 2.0 when compared to multi-night ambulatory polysomnography: a validation study of 96 participants and 421,045 epochs. Sleep Med. 2024;115:251-263. doi:10.1016/j.sleep.2024.01.020. PMID: 38382312.
[4] Baron KG, Abbott S, Jao N, Manalo N, Mullen R. Orthosomnia: are some patients taking the quantified self too far? J Clin Sleep Med. 2017;13(2):351-354. doi:10.5664/jcsm.6472. PMID: 27855740.