Sleep tracker accuracy is strongest for how long you slept and when. Stage estimates are much weaker against overnight lab studies. Heart-rate trends can be usable. None of these devices is a diagnostic test.
If you only wanted the behavior answer — do they work well enough to change a night — start with do sleep trackers work. This page is the methods companion: how the studies are run, and what each device class got right.
It is tested by putting the gadget on someone who is also wired for polysomnography, then comparing the two nights.
Polysomnography (PSG) is the reference. It reads brain waves, eye movements, and muscle tone, then a trained scorer labels wake, N1, N2, N3 (deep), and REM. Consumer devices infer the same labels from motion and pulse. The paper then reports sensitivity (did the device catch sleep?), specificity (did it catch wake?), and often Cohen’s kappa for stage agreement.
Kappa near 1 is almost perfect agreement. Values around 0.4–0.6 are moderate. Below 0.4 is fair to poor. That is the scale to keep in your head when a marketing page says “clinical-grade staging.”
Even PSG is not a perfect oracle. Scorers agree only about 80% of the time. A wearable that disagrees with the lab is not automatically “wrong about how you feel.” It is wrong as a stage instrument. For the behavior question — do they work well enough to change a night — see do sleep trackers work.
Duration and timing are the most reliable outputs. Most modern wearables detect sleep about as well as research actigraphy, and sometimes a little better.
Chinoy and colleagues (2021) compared seven consumer devices with PSG. Sleep detection was high. Wake detection was the miss. That is the classic actigraphy pattern: lying still looks like sleep.
Kainec and colleagues (2024) put five commercial trackers next to both actigraphy and PSG. Total sleep time was comparable to actigraphy on most devices. Garmin Vivosmart was the total-sleep-time outlier. Wake-after-sleep-onset error was large on every device (MAPE 59–138%). Devices overestimated sleep on short-wake nights and underestimated it on long-wake nights. Light-sleep absolute bias was high.
A 2023 multicenter study of 11 wearable, nearable, and airable trackers — including Pixel Watch, Galaxy Watch 5, Fitbit Sense 2, Apple Watch 8, and Oura Ring 3 — found high proportional bias in sleep efficiency. Efficiency is duration divided by time in bed, so a wake-detection miss lands there too.
Daytime naps follow the same story. A 2023 home study of four devices, including Oura Ring Gen 2, found low bias for daytime time in bed. Timing still holds up outside a nocturnal lab night.
Not accurate enough to steer a night, a supplement, or a workout.
Robbins and colleagues (2024) tested Oura Ring Gen3, Fitbit Sense 2, and Apple Watch Series 8 against PSG in healthy adults. Sleep-versus-wake sensitivity was 95% or higher on all three. Stage sensitivity ranged from 50% to 86%. In that healthy-sleeper sample, Oura was not different from PSG on wake, light, deep, or REM duration. That is a minutes-in-bin result, not a claim that Oura scores every epoch correctly. Apple underestimated deep by 43 minutes and overestimated light by 45 minutes.
Do not generalize the Oura minutes result into “Oura stages accurately.” Epoch-by-epoch agreement is the harder test, and it is the one marketing pages skip.
Schyvens and colleagues (2025) scored six wrist devices against PSG: Fitbit Charge 5, Fitbit Sense, Withings Scanwatch, Garmin Vivosmart 4, Whoop 4.0, and Apple Watch Series 8. Sensitivity was above 90%. Specificity sat between 29% and 52%. Kappa for stages: Apple 0.53, Fitbit Sense 0.42, Charge 5 0.41, Whoop 0.37, Garmin 0.21.
A 2019 Fitbit meta-analysis had already shown staging models at 0.95–0.96 sensitivity and 0.58–0.69 specificity. The hardware got nicer. The stage problem did not disappear.
Numbers above are from the named papers, in the samples those papers enrolled — mostly healthy adults in a lab or a monitored night. They will not transfer cleanly to insomnia, apnea, shift work, or a restless partner. However, more research needs to be done in those groups.
“I haven’t seen anything correlating these scores to an independent assessment of sleep or the consequences of sleep,” says Dr. Jamie Zeitzer, co-director of the Center for Sleep and Circadian Sciences at Stanford University and a Rise Science scientific reviewer. “There is currently no requirement for sleep trackers to prove any claim.”
Overnight heart rate and HRV from a worn optical sensor are usually good enough to watch as a trend. They are not a diagnosis.
A well-seated ring or watch can track pulse reasonably through the night. That is a different task from staging the brain. de Zambotti’s 2019 review treats wearable HR as usable in research and clinic as a trend, with the same caveat the field still repeats in 2024: do not let a recovery score outrun the validation.
Use several nights, not Tuesday’s one low reading, before you change training or bedtime. Talk to your doctor about any sudden change that worries you. A wearable is not an ECG.

RISE makes it easy to improve your sleep and daily energy to reach your potential
The devices in the table above have published PSG comparisons in the last few years. That is not the same as “cleared” or “accurate enough to diagnose.”
The 2018 AASM position still holds: consumer sleep technology is not a medical device. Clinicians should treat the export as patient-generated data. A validation paper makes the error bars visible. It does not turn a ring into a sleep study.
If you already wear one of these, keep using it for bedtime and wake time. Import the nights if you want a second opinion on duration. Do not rewrite your life around a 12-minute REM change. And do not treat a “breathing disturbance” flag as apnea. That still needs a clinician.
Duration and timing accuracy matter. Stage accuracy mostly does not, because stages are not the lever.
Looking at data from 1.95 million RISE users aged 24 and up, sleep needs range from five hours to 11 hours 30 minutes. The question that changes energy is whether you met your number, not whether a wrist model split last night into the “right” colors. A slightly noisy bedtime still beats a beautiful, wrong hypnogram.
That is also why a 0–100 sleep score is a poor coach. Most scores fold the least reliable outputs — stages and interruptions — into the grade. An 82 can move because the model relabeled 20 minutes of N2, not because you slept better. The same error is why a smart alarm that claims to catch sleep inertia in “light sleep” cannot keep that promise. If mornings are the actual problem, how to wake up to an alarm and why you cannot wake up are the practical pages. Sleep debt is the number underneath both.
RISE is built around duration and timing. It estimates your need, subtracts the sleep you got, and rolls the difference over 14 nights. Optional wearables can feed those nights. The app does not natively stage-track, and it does not issue a composite score. For cost, limits, and who it is not for, see the RISE app review.
They are accurate enough for sleep duration and timing trends. They are much less accurate for sleep stages. Heart-rate trends overnight can be usable. No consumer tracker is accurate enough to diagnose a sleep disorder, and the American Academy of Sleep Medicine says the data should be treated as patient-generated, not diagnostic.
Researchers put the device on a sleeper who is also wired for polysomnography, the overnight lab study that reads brain waves, eye movements, and muscle tone. They then compare total sleep time, wake after sleep onset, and stage minutes. High sensitivity means the device catches sleep. Low specificity means it misses wake.
In recent lab papers, Oura Gen3, Fitbit Sense 2, and Apple Watch Series 8 all detected sleep versus wake at or above 95% sensitivity. Stage agreement was only fair. Whoop 4.0 scored a kappa of 0.37 for stages. Garmin Vivosmart 4 was a total-sleep-time outlier in one 2024 study and had the weakest stage kappa (0.21) in a 2025 set.
Wrist and finger devices infer sleep from motion and pulse, not from brain waves. Lying still looks like sleep. A 2021 seven-device study found high sensitivity for sleep and low specificity for wake. A 2024 five-device study found wake-after-sleep-onset error of 59% to 138%, and devices overestimated short-wake nights.
Overnight heart rate and HRV from a well-worn optical sensor are usually good enough to watch as a trend. They are not a diagnosis of recovery, overtraining, or illness. Use a several-night baseline, not a single reading, and talk to your doctor about any sudden change that worries you.
Duration and timing accuracy matter, because those are the inputs to sleep debt and circadian alignment. Stage accuracy does not, because you cannot usefully chase REM or deep-sleep minutes with a watch. A slightly noisy bedtime still beats a beautiful, wrong hypnogram.