trailwise
All field notes

Method

Why we do not trust a single readiness score

Three devices, three answers, one body. What we do when the inputs disagree, and why the honest answer is sometimes that we do not know yet.

5 minute readWritten by the Trailwise team

A runner coming up a bend on a wooded trail

A watch, a ring, and a chest strap you slept in because you got curious. One night, one body.

The watch says your readiness is 34 and suggests you take it easy. The ring says 81 and tells you to go and enjoy yourself. The chest strap says your heart rate variability is the highest it has been in a fortnight.

Which one is right?

What a readiness score actually is

Every readiness score on the market does the same job. It takes a handful of overnight measurements, compares them against your own history, and squeezes the result into one number between 0 and 100 that you can read before your coffee has landed.

That squeeze is the product. It is also the problem.

Something inside it decided that last night’s HRV matters more than your sleep duration, or less. It decided how many days count as your normal. It decided whether a late dinner, a warm bedroom or a five hour flight are things it can see, and what to do when it cannot. It decided whether Sunday’s four hour run is still on the books today, and for how much.

Every one of those is a judgement call made by a product team you have never met, tuned on a population you are probably not part of, then handed to you as a number that looks like a measurement. An opinion with a decimal point.

Opinions are useful. But you would not take one from a stranger who refused to say what it was based on, and that is roughly the deal on offer here.

Why the three disagree

Nobody is being sloppy.

The sensors are different. A wrist optical sensor reads blood flow through skin that moves, sweats and changes temperature all night. A finger sensor sits on friendlier tissue and gets a cleaner read of the same rhythm. A chest strap reads the electrical signal directly and is about as close to ground truth as you get outside a lab.

They also sample at different moments. One takes HRV from a fixed window late in the night, another averages across your whole sleep, a third hunts for the lowest sustained stretch it can find. Ask three people to measure the temperature of a room and let each pick their own moment and their own corner. You will get three temperatures, and nobody lied.

They compare against different pasts. Your score is hardly ever about the absolute number. It is about today against your normal, and every vendor defines normal differently. Some use a rolling few weeks, some weight the last few days heavily, some quietly reset after a gap. Two devices can see the identical night and land forty points apart because they are measuring distance from different starting lines.

Then there is the model itself. What counts as recovered, what a bad night costs you, whether Sunday still matters on Wednesday. That is where the real product decisions live, it is proprietary everywhere, and it is doing more work than the sensor.

What to do with three answers

Averaging them gives you a fourth model, wrong in a new and harder to interpret way, and it throws away the most useful thing in front of you. The disagreement is the information.

When your devices agree, lean on it. When they disagree hard, your state this morning is genuinely ambiguous, and the response to ambiguity is to go and find out more, or to make a choice that survives being wrong.

That second one is underrated. If today is an easy hour, being wrong costs you almost nothing, so run it and see how the first fifteen minutes feel. If today is a hard interval set on legs that might already be cooked, being wrong is expensive, and expensive decisions deserve better evidence or a more careful choice. The stakes decide that. The score never did.

What we do instead

Trailwise does not produce a readiness score and hand it over as a verdict. It computes seven signals, keeps them apart, and tells you which ones are driving the answer.

Keeping them apart is the whole point. When readiness, load, sleep and resilience are four things rather than one, “you are at 62” becomes “your sleep is fine, your load is climbing faster than your recovery is keeping up, and that has been true for nine days”. The second is a sentence you can act on, argue with, or dismiss because you know something the model does not.

Each signal carries its own confidence, and we treat that confidence as an answer rather than as small print. If a signal is running on two days of data because your watch died on holiday, it says so. If a stream it depends on has gone quiet, it says that too. It is allowed to tell you it does not know yet.

Every product claims to show confidence, so here is the specific reason it matters. A model that always produces a number will produce one on the days it has nothing to go on, and that number looks exactly like the ones from the good days. You cannot tell them apart, so you learn to trust all of them equally, which means you have learned to trust the bad ones. After two or three of those you stop trusting any of it, and the whole thing becomes a widget you scroll past on the way to Strava.

The cost of that decision is real. “I do not know yet” is a worse product experience than 74. Nobody screenshots a confidence interval, and the honest version has more mornings where the answer is that there is not enough to say. We have argued about it internally more than once, and we keep landing in the same place: these signals are only worth anything on the days they change what you do, and a signal you have quietly learned to distrust will not change what you do on the day it matters.

Three habits, none of which need our product

Pick one device and stay with it. Not because it is the best one, but because comparing today against your own history on the same instrument is the only comparison that means anything, and switching devices resets your baseline whether the app admits it or not.

Read the direction rather than the number. One low morning is noise. Five days sliding the same way while your training has not eased off is worth reacting to. Almost nobody gets hurt because of one bad Tuesday.

When the numbers argue with how you feel, write down how you feel. Your own read on your own body is data, no device can collect it, and across a season it is often the best signal in the set. Measuring all this earns its keep by catching the fortnight where your own read has drifted and you have not noticed.

Trailwise is not a medical device. Neither is it a replacement for a doctor or a physiotherapist. Trailwise flags what is worth a conversation. It does not diagnose, and it will tell you when it does not know.

Keep reading

Start with tomorrow.

Thirty days of Trailwise Pro, your history imported on day one, and a plan that answers to what your body is actually doing.

Get started