Helix

Deep dive/Measurement

One Body,
Eight Scores

Eight products read the same night and disagree by thirty points. The sensors are fine. What differs is five design decisions that nobody publishes, and not one of the scores has ever been independently validated.

By The Helix team

About 32 minutes

14 chapters

How this was sourced

Disclosure, before anything else

Helix makes a product in the category this article assesses. That is a real conflict of interest and it shapes what we chose to examine. We have handled it by comparing design decisions rather than declaring winners, by putting our own algorithm in every chart alongside the others, and by devoting chapter twelve to where our bias most likely shows. Read it before you trust the rest.

Thirty days, one physiology, eight algorithms

Every line reads identical underlying data: the same HRV, the same resting heart rate, the same sleep, the same training load. Only the algorithm differs. Hover for any day, and click a name to isolate it.

0255075100Fourth hard day in a rowTwo short nightsOne heavy night outPeak of an infection161116212630day of the monthscore, 0 to 100

Mean disagreement across the month: 20 points.

A simulation, and the most heavily modelled chart in this article. The month of physiology is synthetic, built to contain recognisable events. Each algorithm is our model of that product’s publicly described inputs (references 2 to 6), expressed as weights plus a baseline window, a smoothing constant and a measurement window. No manufacturer discloses its coefficients, so these are informed reconstructions and not the real thing. What is defensible here is the shape of the disagreement, which follows from documented differences in inputs. The specific numbers are not measurements of any product.

20 pts
Average gap between the highest and lowest score, same day, same body
14
Days out of thirty where two products differ by twenty points or more
0
Composite readiness scores with independent peer-reviewed validation
01The opening

The morning eight products disagreed

Picture a person who, for reasons of curiosity rather than sanity, wears everything at once. A ring on one hand, a band on the other wrist, a watch above it, a second ring, and three apps reading the same phone. One body, one night, eight verdicts by breakfast.

On a Tuesday in the middle of the month, that person slept seven and a half hours, woke with an overnight HRV a little below their usual, a resting heart rate a beat above, and no symptoms of anything. A completely unremarkable night.

WHOOP said they were well recovered and should train. Garmin said they were running low and should not. Oura put them in the middle and mentioned their temperature. Apple declined to give a number at all and simply noted that nothing was outside its typical range.

All four were reading the same heart. None of them was broken.

A recovery score is not a measurement. It is an opinion, with a number on it, about what matters.

That sentence is the whole article, and everything that follows is an attempt to earn it. Because the interesting question is not which product is right, a question nobody can currently answer. The interesting question is what exactly are they disagreeing about, and it turns out to be five specific engineering decisions, each of which is defensible, each of which is invisible to you, and each of which moves your number by more than a bad night of sleep does.

We should start, though, with the part where they agree, because it is not where most people assume.

02The evidence

Where they actually agree

The instinctive explanation for eight products disagreeing is that the cheap sensors must be inaccurate. Optical heart rate from a light shining through your skin, on a moving body, in the dark. Of course it is noisy.

It is not, particularly. And this is the most useful thing to establish first, because it eliminates the obvious suspect.

In 2025 a group put five consumer devices against a single-lead ECG chest strap, across 536 nights of monitoring in 13 adults, and measured how closely each reproduced overnight resting heart rate and overnight HRV. The chest strap is the reference. The results are below, and they are good.

Error against a single-lead ECG

Two separate charts on purpose. Heart rate error and HRV error are different measures and do not belong on one axis.

Overnight resting heart rate, error against ECG

Oura Gen 41.94%
Oura Gen 31.67%
WHOOP 4.03.00%
Garmin Fenix 6not reported

Excluded from the resting heart rate analysis because its timestamp methodology is undocumented

Polar Grit X Pro2.71%

Lower is better. Mean absolute percentage error.

Overnight HRV (RMSSD), error against ECG

Oura Gen 45.96%
Oura Gen 37.15%
WHOOP 4.08.17%
Garmin Fenix 610.52%

Excluded from the resting heart rate analysis because its timestamp methodology is undocumented

Polar Grit X Pro16.32%

Lower is better. Mean absolute percentage error.

As reported: 13 adults, 536 nights, against a single-lead ECG
DeviceRHR CCCRHR MAPEHRV CCCHRV MAPE
Oura Gen 40.981.94%0.995.96%
Oura Gen 30.971.67%0.977.15%
WHOOP 4.00.913.00%0.948.17%
Garmin Fenix 6n/an/a0.8710.52%
Polar Grit X Pro0.862.71%0.8216.32%

Real published figures, reproduced as reportedfrom reference 1: Dial et al., Physiological Reports, 2025. 13 healthy adults, 536 nights, against a Polar H10 single-lead ECG processed in Kubios. Bar colours follow this article’s product palette; Polar has no assigned slot and is shown neutral. Note the sample size: 13 people is small, and the authors restrict their conclusions to apparently healthy adults.

An overnight heart rate within two to three percent of an ECG, from a ring. HRV within six to ten percent for the better devices. For a measurement taken by a light-emitting diode on a finger while you are unconscious, that is a genuinely impressive piece of engineering.

There is variation. The rings did best on both measures, the band was close behind, and the watch was the weakest on HRV. But the spread between best and worst on overnight heart rate is small enough that it cannot possibly explain a thirty point gap in a readiness score.

One detail from that paper worth pausing on

The Garmin Fenix 6 had to be excluded from the resting heart rate comparison entirely, because its timestamp methodology is undocumented. The researchers could not establish what period the device’s own reported figure referred to, so there was nothing to compare against the ECG.

That is a small procedural footnote and it is also a preview of the entire problem in this article. The sensor was fine. What was missing was the definition.

So the raw signals largely agree. Which means the disagreement has to be happening somewhere after the sensor, in the part nobody publishes.

03The gap

The number nobody has checked

Here is the fact that should reframe how you read every one of these products, including ours.

The underlying signals have been validated. Overnight heart rate and HRV from consumer wearables have been measured against ECG in peer-reviewed work, repeatedly, and you can look up the error bars.

The scoreshave not. Not one of the composite readiness, recovery, energy or body battery numbers has been independently validated in peer-reviewed literature against anything. The same 2025 paper notes plainly that manufacturers provide limited transparency, and that what feeds each device’s own readiness or recovery calculation is not disclosed.

Sit with the structure of that for a moment. The input is measurable and measured. The output is the thing you actually make decisions from, and it has never been checked by anyone outside the company that sells it, because it cannot be: there is no independent way to verify a proprietary blend against a ground truth that does not exist.

There is no gold standard for how recovered you are. There is no blood test for readiness. The score cannot be wrong, because there is nothing for it to be wrong about.

This is genuinely different from a step count, which has a true value you could in principle film and count. It is different again from blood pressure, where a cuff and a stethoscope define the answer. Readiness is a construct. Each company invents a definition, implements it, and shows you the result on a scale of one hundred that looks exactly like a measurement.

None of which makes the scores useless. A well-constructed index of things that genuinely matter can be informative even without a ground truth, in the same way a stock index or a cost of living index is informative. But it does mean you are buying an argument rather than an instrument, and it means the only sensible way to compare products is to compare their arguments.

So let us do that. There are five decisions, and every product in this article has taken a different position on most of them.

04Decision one

Which slice of the night

Start with the input that dominates almost every recovery score: overnight heart rate variability. One number, reported each morning, in milliseconds.

Except HRV is not one number. It is a quantity that changes continuously through the night, and it changes a lot. As you move into deep sleep, parasympathetic tone rises and HRV climbs. In REM it falls back. Across the night as a whole there is an upward drift as sleep pressure discharges. The trace is not noise, it is structure.

So a product has to decide which part of the night to report. And the products have decided differently. WHOOP takes its reading during the last slow wave sleep cycle. Ourasamples across the whole night and blends multiple windows. An app reading from a phone’s health store may only get whatever a single morning measurement recorded.

Here is one night. Pick a window.

The same night, four measurement windows

HRV through one night of sleep, with the sleep stage strip underneath. The horizontal line is the mean of the selected window.

304560759064.3 ms0h2h4h6h7hhours asleepHRV, ms
  • Deep
  • REM
  • Light
  • Awake
This window reads
64.3 mslast slow wave sleep cycle
Lowest window
55.0 mssame night
Highest window
69.8 mssame night
Difference
14.8 ms27% apart

Illustrative night, not a recording. The trace is constructed to show the documented structure of nocturnal HRV: higher in deep sleep, lower in REM, drifting upward across the night. Window definitions follow the published descriptions in reference 2. The point being demonstrated, that the same night yields materially different HRV values depending on the averaging window, follows from the measurement standards in reference 7, which require the window and metric to be held constant for values to be comparable.

The gap between the highest and lowest window on this single night is about 15 milliseconds. For context, that is a larger difference than most people see between a well recovered day and a poorly recovered one.

Which means two products can report your HRV as 54 and 68 on the same night, both correctly, because they are answering different questions. Neither is inaccurate. They have simply defined the measurement differently, and neither tells you which slice it used.

There is a defensible argument behind each choice. The last deep sleep cycle is arguably the cleanest look at autonomic state, least contaminated by the previous day. The whole-night average is more stable and less dependent on correctly identifying sleep stages, which consumer devices do imperfectly. A morning reading is worse on both counts but is the only thing available if you are reading someone else’s data.

Every one of those is a reasonable engineering decision. Together they guarantee that the same body produces different numbers.

05Decision two

What counts as normal

An HRV of 54 milliseconds means nothing on its own. Between-person variation in HRV is enormous, far larger than the within-person variation that actually carries information. A value of 54 could be exceptional for one person and alarming for another.

So every product compares you to yourself. Today’s value against your baseline. Which raises the question nobody asks: how long is a baseline?

This is not a detail. It is arguably the single most consequential parameter in the entire system, and it is never disclosed. A short baseline makes the score twitchy and quick to adapt. A long baseline makes it stable and slow to notice that you have changed. And critically, the two can give opposite answers about the same morning.

Take the month from the top of this article and move the baseline window.

The same morning, judged against different histories

Pick a morning, then change how much history counts as normal. Watch the verdict flip.

40506070baseline 59.0161116212630day of the month
Last night
46.0 ms
Baseline
59.0 ms
Deviation
-13.0 ms
Reads as
Below normal

Model, on the synthetic month described under the opening chart. The baseline is a simple rolling mean, which is the simplest defensible implementation; real products may use weighted or robust estimators they do not describe. The effect being shown, that baseline length changes whether a given morning reads as above or below normal, holds for any such estimator.

Try day 20 with a seven day baseline and then with a twenty-eight day one. In the first case you are recovering from a rough stretch and the recent days drag the baseline down, so today looks like an improvement. In the second, the baseline still remembers a better version of you, and the same morning reads as below normal.

Both are correct. They are answering different questions: better than this week versus better than this month. A product that picks the short window is optimistic during recovery and alarmist during a decline. One with a long window is the reverse.

The trap this creates for anyone changing their life

If you start training seriously after a sedentary year, your HRV will rise over weeks. A short baseline will keep resetting to your improving average, so your score stays around the middle and never reflects that you are getting fitter. A long baseline will show you consistently above normal for a month and then quietly recalibrate.

Neither is lying. But if you interpret the score as a fitness measurement rather than a deviation measurement, you will draw the wrong conclusion from both.

06Decision three

What goes into the blend

Now the part everyone assumes is the whole story: which inputs the score is made of, and how much each one counts.

From published descriptions, the inputs differ substantially. WHOOP uses HRV, resting heart rate, respiratory rate, blood oxygen, skin temperature, and sleep against a calculated sleep need, with HRV reportedly accounting for a little over half of the variance in the final score. Garmin weights training load and recovery time from your last hard session far more heavily, because it is modelling energy rather than autonomic state. Samsung pulls in weekly sleep and daily activity. Ultrahuman adds a stress rhythm component and keeps updating through the day.

Rather than describe that, here it is as a machine. The presets are each product’s published input mix, as best we can reconstruct it. Then take the sliders and build your own.

Build a recovery score

Same month, same physiology, your weighting. The dashed lines mark the hard training block, the short nights, the heavy night out and the infection.

050100WHOOP input mix, as published
Hangover day
18
Peak of infection
19
Fourth hard day
14
Best day
86

Model. The weights are our reconstructionof published input descriptions (references 2 to 6), not disclosed coefficients, and no manufacturer publishes theirs. Apple is omitted from the presets because it does not produce a composite score at all, which is the subject of chapter nine. Treat the presets as a fair reading of each product’s stated priorities rather than as a reimplementation of its algorithm.

The instructive experiment is the heavy night out on day 18. Load a preset that weights HRV heavily and the score collapses, which is arguably correct: your autonomic nervous system genuinely had a bad night. Load one that weights training load heavily and the score barely moves, which is also arguably correct: you did no training, so you are not fatigued in the sense that model cares about.

One product tells you to rest. The other tells you to train. Both are faithfully implementing a coherent theory of what recovery means. You are not choosing between accurate and inaccurate. You are choosing whose theory you find more useful.

And notice how much of the final number a single weighting decision controls. Moving HRV from twenty percent to fifty-six percent of the blend changes more days than any real physiological event in the month does.

07Decision four

How fast should it move

There is a slider in the previous chart that looks technical and is actually a philosophical position. Smoothing.

Every score has to decide how much of yesterday to carry into today. Set it to zero and the number reacts fully to each night’s data, which makes it responsive and also noisy: a single restless night sends it tumbling. Set it high and the score becomes a slow-moving trend that shrugs off individual nights but takes days to acknowledge that you are ill.

The products sit at visibly different points. WHOOP is built to react each morning, which suits an audience making a daily training decision. Garmin’s Body Battery, by design, drains and refills continuously and carries state across days, so it behaves more like a reservoir than a snapshot. The consequence is that on the day after one bad night, those two will not agree, and cannot, because one is answering about today and the other about the week.

This is where the trade-off becomes uncomfortable. A responsive score is more useful when something real happens and more annoying the rest of the time. A smooth score is calmer and misses the first day of an infection. There is no setting that is simply better, which is exactly why eight companies picked eight different settings.

Every complaint about a recovery score being too jumpy, and every complaint about it being too slow to notice anything, is a complaint about the same parameter set in opposite directions.

Go back and set smoothing to zero on the mixer, then to eighty percent, and watch what happens to the infection around day 23. At zero it is caught immediately and violently. At eighty it is caught two days late and gently. If you are an athlete deciding whether to do intervals, the first is better. If you are a person who does not want to be told they are broken every time they sleep badly, the second is better.

08Decision five

What is the score even for

The four decisions so far are technical. This one is not, and it explains more of the disagreement than the other four combined.

These products are not competing implementations of one idea. They are answering genuinely different questions, and the number on the screen is shaped by the question.

  1. Helix · Readiness

    What should you do today, and why.

    Interpretation layer · inputs: overnight hrv, resting heart rate, sleep, strain, blood markers · reference 08

  2. WHOOP · Recovery

    How recovered is your autonomic nervous system this morning.

    Own sensor · inputs: hrv during the last slow wave sleep cycle, resting heart rate, respiratory rate, blood oxygen, skin temperature, sleep against calculated sleep need · reference 02

  3. Oura · Readiness

    Are your multi-day trends where they usually are.

    Own sensor · inputs: hrv sampled across the whole night, body temperature, sleep, previous day activity · reference 02

  4. Garmin · Body Battery

    How much energy is left in the tank right now.

    Own sensor · inputs: continuous hrv and stress, sleep, recovery time from the last hard session, training load balance, acute fatigue · reference 02

  5. Apple · Vitals (no score)

    Is anything outside your typical range tonight.

    Platform · inputs: overnight heart rate, respiratory rate, wrist temperature, blood oxygen, sleep duration · reference 03

  6. Samsung · Energy Score

    How much physical and mental energy do you have.

    Own sensor · inputs: daily activity, weekly sleep, average sleeping heart rate, sleeping hrv · reference 04

  7. Ultrahuman · Dynamic Recovery

    How is your body adapting, updated through the day.

    Own sensor · inputs: sleep, stress rhythm score, temperature, resting heart rate, hrv · reference 05

  8. Bevel · Recovery

    What does the data your watch already collected mean.

    Interpretation layer · inputs: hrv from apple health, resting heart rate from apple health, sleep from apple health, strain from apple health · reference 06

Read that list as a set of answers to different exam questions and the disagreement stops being a scandal. A fuel gauge and a snapshot of autonomic state are not two attempts at the same measurement. They are two different instruments that happen to share a zero-to-one-hundred scale, which is the actual design sin in this whole category.

If one of them reported in litres and another in millivolts, nobody would expect them to match. Because they all report a percentage, everybody does.

The one thing that would clear most of this up

Not better sensors, and not better algorithms. A published definition. If each product stated which window it measures, how long its baseline is, what its inputs are weighted at, and how much smoothing it applies, most of the confusion in this article would evaporate. You would still get different numbers, but you would know why, and you could pick the definition that matched what you wanted to know.

09The outlier

The one that refuses to give you a number

One of the eight has taken a completely different position, and it deserves its own chapter because it is the most interesting design decision in the category.

Apple collects most of the same overnight measurements as everybody else: heart rate, respiratory rate, wrist temperature, blood oxygen, sleep duration. It has the sensors and the data to produce a readiness score.

It does not produce one. Instead it establishes a typical range for each metric individually, shows you where tonight sits inside that range, and notifies you only when several metrics are outside their usual range at once.

No composite. No percentage. No verdict. Just: these five things are typical for you, or these two are not.

It is worth being precise about what this buys and what it costs. What it buys is honesty. Every criticism in this article, the undisclosed weighting, the arbitrary baseline, the invented construct with no ground truth, applies to composite scores and not to per-metric outlier detection. Apple has declined to make the claim it cannot support.

What it costs is usefulness. Five numbers and a typical range do not tell you whether to do intervals this morning. The composite score, for all its arbitrariness, compresses a judgement into something you can act on before coffee. Refusing to make that judgement is more defensible and less helpful.

One product declines to guess and is criticised for being unhelpful. Seven products guess and are trusted because the guess arrives as a number.

Which of those trade-offs is right depends entirely on whether you want an instrument or an opinion. The category has mostly decided you want an opinion, and it is probably right about that. But the honest version of a recovery product would present the opinion as an opinion, with its reasoning attached.

10Structure

Sensor or interpreter

There is a structural split in this list that cuts across all five decisions, and it changes what a product is even capable of.

Some of these companies make hardware. WHOOP, Oura, Garmin, Samsung and Ultrahuman own the sensor. They control the sampling rate, they see the raw photoplethysmography waveform, they choose which parts of the night to analyse, and they can change any of it in firmware.

Others are interpretation layers. Bevelreads what your Apple Watch already recorded, through the phone’s health store, and does its analysis on top. It measures nothing itself. Helix works the same way, which we should say plainly here rather than in a footnote.

Each architecture has a real ceiling.

Owning the sensor means you can define the measurement window precisely, sample as often as your battery allows, and access signals the platform never exposes. It also means you have to sell people a second device, keep it charged, and persuade them to wear it forever. And it means your data is yours alone, which is convenient commercially and awkward for the user who wants to leave.

Being an interpretation layermeans you inherit whatever the platform decided to record, at whatever cadence it chose, already processed by somebody else’s algorithm. If the watch recorded one morning HRV reading, that is what you get. You cannot ask for the last slow wave cycle because nobody saved it. In exchange you require no new hardware, you work with the device someone already owns and already charges, and the user keeps their data in a store you do not control.

Which is why the honest comparison is not device against app

A ring with a purpose-built sensor will always have better raw access than an app reading a general-purpose watch. That is physics and platform design, not marketing.

What an interpretation layer can compete on is everything after the signal: the quality of the reasoning, how much of it is shown to you, and whether it connects data the hardware companies keep in separate silos. Nobody selling a ring is going to read your blood panel.

Architecture, and the number each product shows you
ProductShows youArchitectureInputs described
HelixReadinessInterpretation layer5
WHOOPRecoveryOwn sensor6
OuraReadinessOwn sensor4
GarminBody BatteryOwn sensor5
AppleVitals (no score)Platform5
SamsungEnergy ScoreOwn sensor4
UltrahumanDynamic RecoveryOwn sensor5
BevelRecoveryInterpretation layer4
11The result

How far apart they get

Put the five decisions together and you can ask the question this article started with properly. On a given morning, how far apart do eight reasonable algorithms get, reading one body?

Highest to lowest score, every day of the month

Each bar spans the full disagreement on that day. The dots are coloured by which product sat at each end.

0255075100161116212630day of the monthhighest to lowest score that day

Orange bars are days where at least two products disagree by twenty points or more. The dots at each end are coloured by which product was highest and which was lowest.

Derived from the simulation described under the opening chart, so it inherits every caveat there: the physiology is synthetic and the algorithms are reconstructions of published input descriptions rather than real implementations. The claim this chart supports is qualitative, that documented differences in inputs and baselines are sufficient to produce large disagreement. It is not a measurement of how far real products differ, which nobody has published.

The average gap is about 20 points, and on 14 of thirty days at least two products differ by twenty or more. The widest gaps do not appear on the chaotic days. They appear on the ambiguous ones.

That pattern is worth understanding. When something unmistakable happens, an infection with a temperature rise, a heart rate up nine beats, and HRV down twenty milliseconds, the algorithms converge. Everybody detects it, because every reasonable weighting detects it. Agreement is high exactly when you least need the help.

On the genuinely uncertain mornings, the ones where you actually want advice, the weighting decisions dominate and the products scatter. The score is least reliable precisely when it is most consulted.

They agree when it is obvious and disagree when it matters. Which is the opposite of what you would want from an instrument, and exactly what you would expect from an opinion.

12Disclosure

Our own conflict of interest

We build one of these. That should change how you read everything above, so here is our bias, as precisely as we can state it.

The obvious problem. We chose the framing. An article arguing that recovery scores are unvalidated opinions is convenient for a company whose pitch is that it shows you the reasoning. We did not select that thesis because it flatters us, but we would say that either way, and you should discount accordingly.

What we could not do honestly. We cannot claim Helix is more accurate. Nobody can, about any product in this category, because the thing being scored has no ground truth and no composite score has been independently validated, including ours. Any company telling you their readiness score is the accurate one is making a claim that is not currently checkable.

Where we are subject to the same criticisms. Helix picks a measurement window, a baseline length, a set of weights and a smoothing constant, exactly like everyone else. Those choices are just as arbitrary as anyone else’s. We are an interpretation layer, so we inherit whatever the platform recorded and we have less signal access than any ring in this article. Our line on the chart at the top is a reconstruction of our own approach, and it sits in the same scatter as the others because it is subject to the same decisions.

What we did to keep this defensible. We organised the article by design decision rather than by product, so no chapter is a hit piece. We used published input descriptions and cited them, marking the third-party summaries as secondary sources rather than passing them off as vendor specifications. We put real validation data in unmodified. And we labelled every simulated chart as a model, because the alternative, presenting reconstructions as measurements of competitors, would be the exact dishonesty this article criticises.

What we would want a sceptical reader to check. Our reconstruction of each competitor’s input weighting is the softest part of this piece. Vendor documentation changes, algorithms update silently, and our weights are inferred. If you find that we have mischaracterised how a product works, that is a defect and we want to hear about it.

The position we can actually defend

Not that our number is better. That a number without its reasoning is not worth much, and that a category built on unvalidated composites owes its users an explanation rather than a colour.

These deep dives are the argument for that position. They are not marketing for a score. They are what it looks like when a company shows its work, and you can judge whether that is worth anything by whether this article was useful to you even though it undermines its own product’s main claim.

13Practical

What would actually be worth buying

If the scores cannot be compared on accuracy, what should you compare them on? Six things, none of which appear on a spec sheet.

  1. Does it tell you what it measured

    Which window, which metric, how long the baseline is. A product that shows you the components behind the score, and lets you see the underlying HRV and heart rate rather than only the verdict, is giving you something you can reason with. One that shows a colour is asking for faith.

  2. Does its question match your question

    If you want to know whether to do intervals this morning, a responsive autonomic snapshot is the right instrument. If you want to manage energy across a week, a reservoir model is. Buying the wrong question and then complaining the number is wrong is the most common mistake in this category.

  3. Can you get your data out

    Your physiology is a multi-year asset and these products have a high failure rate as businesses. A device whose history you can export, or which writes to a store you control, is worth meaningfully more than one that holds your baseline hostage.

  4. Does it change its mind quietly

    Algorithms get updated, and a silent update can shift your scores by more than any real change in your body. Ask whether the company tells you when it has changed the definition. Most do not, which makes long-term comparison of your own history unreliable in a way nobody warns you about.

  5. Will you actually wear it

    The most accurate device you take off is worse than the mediocre one you never remove, because every method here depends on a continuous personal baseline. Gaps in the record damage the baseline, which damages every score computed against it.

  6. Does it connect things nobody else connects

    Most of these products see one slice of you. The interesting questions live between slices: how your training load relates to your sleep, how both relate to a blood panel. Breadth is a real differentiator, and unlike accuracy it is something you can actually verify before you buy.

14Conclusion

The number is an argument

Go back to the person wearing everything. Eight verdicts, one night, thirty points of disagreement, and not a broken sensor anywhere.

The sensors were within a few percent of an ECG. The disagreement came from five decisions made in offices: which slice of the night to trust, how much history counts as normal, what to put in the blend and at what weight, how much of yesterday to carry forward, and what question the number is answering in the first place.

Every one of those decisions is defensible. None of them is published. And the composite that comes out the other end has never been independently validated, for any product on the market, including the one we make.

Which means the right way to hold one of these numbers is as an argument rather than a reading. Somebody built a theory of what recovery means, implemented it, and is showing you the result. The useful questions are what the theory is, whether it matches what you want to know, and whether the company will tell you.

A score with its reasoning attached is worth having. A score without it is a colour, and you have no way of knowing what it means.

REFSources

References and sources

Links resolve by title search or to the primary page, so a misremembered identifier cannot send you to the wrong source. Where a source is a third-party summary rather than vendor documentation, it is marked, because this article makes claims about other companies’ products and the distinction matters.

Which charts are data and which are models

One chart in this article is real measured data: the accuracy comparison in chapter two, reproduced as published. Everything else is explicitly a model. The thirty day simulation, the disagreement spread, the window picker and the blend mixer all run on synthetic physiology and on reconstructions of published input descriptions. They are built to demonstrate that documented differences in design are sufficient to produce large disagreement. They are not measurements of any product’s behaviour, and no manufacturer’s coefficients are known to us or to anyone outside those companies.

  1. Validation of nocturnal resting heart rate and heart rate variability in consumer wearables

    Dial MB, Hollander ME, Vatne EA, Emerson AM, Edwards NA, Hagen JA. Physiological Reports, 2025. Find it

    Used for: The accuracy figures in chapter two, reproduced as reported: 13 adults, 536 nights, against a Polar H10 single-lead ECG processed in Kubios. Also the finding that the Garmin Fenix 6 had to be excluded from the resting heart rate analysis because its timestamp methodology is undocumented, and the observation that manufacturers do not disclose what feeds their readiness and recovery scores.

  2. Published breakdowns of how WHOOP Recovery, Oura Readiness and Garmin Body Battery are calculatedSecondary

    Independent comparison analyses of consumer recovery scores. Secondary sources, accessed 2026. Find it

    Used for: The input lists for WHOOP, Oura and Garmin used throughout, including that WHOOP samples HRV during the last slow wave sleep cycle while Oura samples across the whole night, and the reported figure that HRV alone accounts for roughly 56 percent of the variance in WHOOP Recovery. These are third-party summaries rather than vendor specifications, and should be checked against each company's own documentation before being relied on.

  3. Track your overnight vitals with Apple Watch

    Apple. Apple Support. Find it

    Used for: That the Vitals app reports overnight heart rate, respiratory rate, wrist temperature, blood oxygen and sleep duration, establishes a typical range for each, and notifies you when multiple metrics fall outside it, rather than producing a composite score.

  4. Energy Score on Galaxy Watch, and the University of Georgia collaboration

    Samsung. Samsung, accessed 2026. Find it

    Used for: That Energy Score is derived from daily activity, weekly sleep, average sleeping heart rate and sleeping heart rate variability, and that Samsung worked with the University of Georgia on defining and measuring energy.

  5. Ultrahuman Ring AIR: Dynamic Recovery

    Ultrahuman. Ultrahuman, accessed 2026. Find it

    Used for: That Dynamic Recovery is built from five factors, sleep, stress rhythm score, temperature, resting heart rate and HRV, and that the score updates through the day rather than being fixed each morning.

  6. Bevel: the connected health coach

    Bevel. Bevel, accessed 2026. Find it

    Used for: That Bevel reads data already collected through Apple Health rather than measuring anything itself, and produces its own Recovery, Sleep, Strain and Stress scores on top of it.

  7. Heart rate variability: standards of measurement, physiological interpretation, and clinical use

    Task Force of the European Society of Cardiology and NASPE. Circulation, 1996. Find it

    Used for: The measurement standards behind the point in chapter four that HRV values are only comparable when the recording window and the metric are held constant.

  8. Helix, and the editorial standards behind these deep dives

    Helix. This site. Find it

    Used for: The description of our own product in chapter twelve, and the conflict of interest disclosed there. We make a product in the category this article assesses.

Product names and trademarks belong to their respective owners. This article describes publicly documented behaviour and does not reproduce or reverse engineer any proprietary algorithm. Nothing here is medical advice, and no recovery score of any brand should be used to make a clinical decision.

← All deep dives