trailwise
All field notes

Data

Reading sleep when the watch disagrees with you

Stage data is a model, not a measurement. How much weight it deserves, and when to ignore it.

5 minute readWritten by the Trailwise team

An empty single track path winding through sunlit woodland

You wake up feeling fine. Properly fine. Slept through, got up before the alarm, ready for the session.

Then you check the app. Thirty four minutes of deep sleep, well below your average. Sleep score 61. It suggests you consider taking it easy today.

Most people do one of two things here. They believe the watch and back off, which is sometimes right and frequently a waste of a good day. Or they decide the whole thing is nonsense and stop looking, which throws away the useful parts along with the part that is not.

The better answer comes from understanding what your watch can actually see.

What your watch can actually see

Your wrist device has three or four things to work with. Movement, from the accelerometer. Heart rate, from the optical sensor. Heart rate variability, derived from that. Sometimes skin temperature.

Notice what is not on that list. Brain activity. Eye movement. Muscle tone.

Those three are what a sleep lab measures, and they are what sleep stages are actually defined by. A polysomnogram puts electrodes on your scalp, next to your eyes and under your chin, and a trained technician scores the night in thirty second windows against those signals. Everything else is an attempt to approximate it.

So your watch is inferring stages it cannot see, from signals it can. It looks at how still you are, what your heart is doing, how variable the intervals between beats are, and it runs a model that says: this pattern usually corresponds to deep sleep, this one usually corresponds to REM.

The word doing the work in that sentence is *usually*.

Where the model is good, and where it is not

What follows is the direction of the finding rather than invented precision, because the numbers vary by device, by generation and by study.

Sleep versus wake is the easy problem, and modern devices are good at it. Working out whether you are asleep at all is largely a movement question, and accelerometers have been doing it since long before wearables were a category. When your watch tells you that you fell asleep around eleven and woke at half past six, believe it.

Total duration is close behind, because it inherits the accuracy of sleep versus wake. The usual failure mode is over counting: lie very still while reading and some devices will happily record you as asleep. If you are a fidgety sleeper the error runs the other way.

Stage split is where it gets shaky. Distinguishing light from deep from REM using only movement and heart rate is a much harder inference, and agreement with a lab is meaningfully worse than for duration. Deep sleep is the one to be most sceptical about, and it is also the number apps most like to put on the front screen, because it sounds the most consequential.

Continuity sits somewhere in between and is underrated. How fragmented the night was, how many times you surfaced. Your watch reads this reasonably well because it is closer to a movement question, and it tracks how you feel rather better than the stage pie chart does.

Two more things worth knowing. Your baseline is not a population baseline: when an app says your deep sleep is low, low usually means low for you, computed from a few weeks of your own nights, and if those weeks were unusual then so is the comparison. And the model changes under your feet. Firmware updates alter these algorithms, your deep sleep can shift by ten or fifteen minutes on average because the vendor improved something, and there is rarely a note in the release. If your sleep composition changes overnight and your life did not, look at your update history before you look at your habits.

How much weight each part deserves

Duration: high. If you are consistently getting less sleep than you were, that is real, it is important, and it sits upstream of almost everything else. This is the most useful number your device produces about your night.

Timing and regularity: high. Going to bed and getting up at wildly different times has real consequences, your device measures it well, and it is one of the few things on this list you can straightforwardly fix.

Continuity: medium. A night with six wake ups is genuinely different from a night with none, and worth noticing when it repeats.

Stage split: low, and only as a trend. One night of low deep sleep tells you close to nothing. Three weeks of your deep sleep sliding while everything else slides too is worth a look, because a consistent shift in a consistent model still carries information even when the absolute number is unreliable. The error is at least in the same direction every night.

The single sleep score: a headline, not a verdict. It is a compression of everything above, weighted by choices you cannot see. Same argument as readiness scores, which I have written about separately.

When to ignore it entirely

When you feel good and it says you should not. How you feel is a real measurement of the thing the watch is trying to approximate. If they disagree and you feel fine, start the session easy and reassess after fifteen minutes. Your legs will tell you more in a quarter of an hour than the model did overnight.

When you know why the number is odd. Late meal, warm room, glass of wine, travel, sleeping somewhere unfamiliar, a child in the bed. You have context the device does not, and the number is unsurprising rather than wrong.

When you slept badly for a reason that has now gone. Bad night before a flight, fine now. One night rarely dictates a training decision on its own.

What I would not ignore is the trend. One bad night is noise. Three weeks of duration sliding while your training load climbs is one of the clearest early warnings you will get, and it deserves more attention than any single morning’s score ever has.

What we do with it

Sleep contributes to our signals in a few places, and stage data is deliberately not the load bearing part.

Duration is compared against your own need rather than a fixed eight hours, because need genuinely varies. Continuity is read as its own fact. Stage split is present, shown when the device provides it, and honestly absent when it does not: some paths into our system carry no stage data at all, and rather than filling that gap with a plausible guess we show that it is not there.

That last decision is the one I would defend hardest. A missing stage split displayed as a missing stage split is mildly annoying. A missing stage split quietly replaced by an estimate is a number you will trust exactly as much as a real one, with no way to tell the difference.

Trailwise is not a medical device. Neither is it a replacement for a doctor or a physiotherapist. Trailwise flags what is worth a conversation. It does not diagnose, and it will tell you when it does not know.

Keep reading

Start with tomorrow.

Thirty days of Trailwise Pro, your history imported on day one, and a plan that answers to what your body is actually doing.

Get started