Helix

Deep dive/Signals

Your Resting Heart Rate
Knows First

Wearables really do detect an infection days before you feel it. Then the arithmetic arrives, and most of the alerts turn out to be wine, travel and a bad night of sleep.

ByThe Helix team

About 28 minutes

12 chapters

How this was sourced

One person, 365 nights, every alert the algorithm fired

Each bar is one night, measured as beats per minute above that person’s own rolling baseline. The dashed line is the published alert threshold. Two of these nights were the start of a real infection.

  • Ordinary night
  • Alert fired
alert at +3 bpm0510JanMarMayJulSepNovbpm above personal baseline, one night per bar

58 alerts fired. 4 of them were an infection.

Synthetic year, built from published effect sizes. This is not a recording of a real person. Night-to-night variation, alcohol at +3.0 bpm, vaccination peaking at +3.4 bpm, travel, hard training and short sleep are each set from the sources in references 3, 7 and 8, and two infections are inserted with a three-night rise before symptoms as described in reference 3. The alert logic is the published NightSignal rule. The year is calibrated so false alarm frequency lands near the 1.30 alert days per person per 21 days that Alavi et al. report for test-negative participants.

3 days
Median warning before symptoms begin, in the largest real-time trial
80%
Of infections flagged by that system
~1 in 9
Alerts that mean anything on an average night of an average year
CH 01The opening

The night the ring was right

Imagine you wake up on a Tuesday and your ring has left a note. Overnight heart rate elevated. Consider taking it easy today.

You feel completely fine. You go to work, you train in the evening, you forget about it. On Thursday your throat hurts. By Friday you are in bed with a box of tissues and a grudging respect for a piece of jewellery, telling colleagues that the thing knew two days before you did.

That story is true. It happens, it is documented in peer reviewed research on tens of thousands of people, and the physiology behind it is not mysterious. A wearable can see an infection coming before you can, reliably enough that during the pandemic serious research groups at Stanford, Scripps and UCSF built alerting systems on exactly this principle.

Here is the part that story leaves out. In the same year, that same ring, on that same finger, sent you 54 other alerts that meant nothing at all. You do not remember those, because nothing happened afterwards and there was no punchline to tell anyone. You remember the one that was followed by a sore throat.

This article is about how both of those things are true at the same time, and about the piece of arithmetic that decides which one you should believe on any particular Tuesday. It is not a debunking. The signal is real, the science is good, and the detection genuinely works. The problem is not the sensor. The problem is what happens when a very good test meets a very uncommon event, which is a problem no amount of better hardware can solve.

The device is not wrong when it alerts you. It is answering a different question from the one you think you asked.

Start with why your pulse changes at all.

CH 02Mechanism

Why an infection shows up in your pulse

Your resting heart rate is not really a measure of your heart. It is a readout of how hard your body is working to keep itself running while you are doing nothing, which makes it an accidental summary of almost everything else.

When a virus establishes itself, your immune system begins an expensive project. Signalling molecules circulate, temperature set points shift upward, and metabolic rate rises to fund the work. Higher metabolic rate needs more oxygen delivered and more waste carried away, which means more cardiac output. At rest, with stroke volume roughly fixed, the only lever available is rate. So your heart beats faster while you sleep.

There is a second contribution. Immune activation shifts autonomic balance toward the sympathetic side, the branch that raises heart rate and suppresses the vagal braking that normally slows you down overnight. This is why heart rate variability falls at the same time as resting heart rate rises. They are two views of the same shift.

Crucially, all of this begins before you notice anything. Your subjective sense of being unwell depends largely on symptoms that are themselves products of the immune response, and those arrive later, once the response is well underway. There is a window, measured in days, when your body is visibly fighting and your conscious experience has not been informed.

Sleep is the ideal time to look for this. You are still, you are not digesting much, you are not caffeinated, you are not stressed by a meeting. Overnight is the closest a consumer device gets to a controlled measurement, which is why almost every serious detection algorithm works on overnight resting heart rate rather than daytime numbers.

Why overnight, and not a spot reading

A single daytime heart rate tells you almost nothing, because it is dominated by what you were doing in the previous ten minutes. What the algorithms use instead is the average overnight resting rate, compared against a rolling baseline built from that individual’s own recent history. The comparison is always to you, never to a population, because resting heart rate varies enormously between healthy people.

Now watch what that looks like across the days around an infection, and pay attention to where the sensor and your own perception part company.

symptoms beginalert threshold-6-4-20+2+4+6days relative to symptom onsetwhat you notice: Nothing
  • Resting heart rate, bpm above baseline
  • Skin temperature, scaled

Day -6 to -4

Six days out, nothing at all

You have already been infected. Nothing has happened yet that either you or the sensor can see. Your overnight heart rate is doing what it always does.

Day -3 to -2

Three days out, the sensor moves first

The immune response begins in earnest. Metabolic rate rises, sympathetic drive increases, and overnight heart rate lifts a few beats above baseline. This is the window the algorithms are built to catch, and the median lead time reported is three days.

Day -1

The night before, you feel something and misread it

Now the deviation is large. You are also tired, and you attribute it to a hard week, a late night, a heavy meal. This is the moment the device has genuine information that you do not.

Day 0 to 2

Symptom onset, and the news is stale

The sore throat arrives. Your heart rate is roughly ten beats above baseline and your temperature has climbed most of a degree. Everything the sensor now reports, you already know.

Illustrative timeline, not a measured recording. The shape follows the presymptomatic trajectories described in references 2, 3 and 10: a rise beginning roughly three days before symptoms, a peak around onset, and resolution over about a week. The temperature track is scaled to share the axis and is shown for shape only, following reference 5. Individual trajectories vary widely, and some infections produce no detectable signal at all.

CH 03The evidence

What the studies actually found

The pandemic did something unusual for this field. It created a situation where hundreds of thousands of people were already wearing continuous physiological monitors, were highly motivated to report symptoms and test results, and were being infected by a single identifiable pathogen at a knowable time. That is close to an ideal natural experiment for presymptomatic detection, and several groups ran it.

The results were genuinely impressive, and they agree with each other.

81%

Mishra et al. 2020

About 5,300 enrolled, 32 infections

Of 32 infected people, 26 showed changes in heart rate, steps or sleep. Of the 25 with symptom dates, 22 were flagged before or at symptom onset, and four at least nine days early.

Reference 02

80%

Alavi et al. 2022

3,318 participants, 84 infections

A real-time overnight alerting system flagged 67 of 84 infections, at a median of three days before symptoms began. Reported specificity was 87.7%.

Reference 03

0.80

Quer et al. 2021

30,529 enrolled, 333 tested

Sensor data plus self-reported symptoms discriminated positive from negative cases with an area under the curve of 0.80, against 0.71 for symptoms alone.

Reference 04

38/50

Smarr et al. 2020

50 cases within a 65,000-person study

Continuous skin temperature from a ring identified elevated temperature in 38 of 50 cases before the participants reported symptoms that made them suspect illness.

Reference 05

Figures as reported in references 2, 3, 4 and 5. Each card names its cohort. Note that these are detection rates within confirmed infections, which is sensitivity, and that sensitivity alone says nothing about how often the same system alerts when nothing is wrong. That is the subject of chapter six.

The Stanford work is the one to look at most closely, because it is the only one of these built and run as a live alerting system rather than analysed after the fact. Alavi and colleagues put a real-time algorithm called NightSignal in front of 3,318 people. It watched overnight resting heart rate against a streaming baseline and raised a yellow alert when the rate sat about three beats per minute above that baseline, and a red alert at four or more across two consecutive nights.

Of 84 people who caught the virus during the study, the system flagged 67. That is a sensitivity of 80 percent. The median lead time was three days before symptoms began. It also caught 14 of 18 asymptomatic infections, which is a genuinely remarkable result, because those are people who would never have known they were infected at all.

Read only that paragraph and the conclusion writes itself. A cheap consumer device, worn on the wrist or the finger, catches four out of five infections three days early, including the ones that are otherwise invisible. If that were the whole story, this technology would have ended respiratory disease transmission in offices.

It did not, and the reason is in the same paper.

CH 04The turn

It works beautifully when nobody is looking at you

Before the awkward part, an important detour, because there is one application of this signal that works so well it is almost boring, and understanding why illuminates everything else.

In 2020, Jennifer Radin and colleagues at Scripps took Fitbit data from around 200,000 users, narrowed it to 47,249 people who wore their devices consistently in five states, and asked a different question. Not can we tell whether this person is getting sick, but can the share of people with elevated resting heart rate tell us what is happening to influenza rates across a whole state.

The answer was yes, emphatically. Adding wearable data improved state-level surveillance in every state tested, with final model correlations against reported illness rates ranging from 0.84 to 0.97. That is not a marginal signal. That is a public health instrument, and it reports faster than the traditional system because it does not wait for people to visit a doctor.

Here is the thing worth sitting with. The exact same measurement, from the exact same devices, is close to useless for one person and excellent for a population. Look at the two views next to each other.

The same signal, aggregated and alone

Above, the share of wearable users with elevated overnight heart rate against reported illness rates. Below, one person across the same weeks.

Everybody at once

52 weeks
  • Reported illness rate
  • Share of users with elevated overnight heart rate

Correlation with reported illness rates: 0.99. Radin et al. reported 0.84 to 0.97 across five states.

One person, the same 52 weeks

52 weeks

Correlation with reported illness rates: 0.13. The seasonal signal that is obvious in aggregate is not recoverable here.

Illustrative curves. The aggregate panel is constructed to reproduce the relationship reported in reference 1, where models reached correlations of 0.84 to 0.97 with CDC influenza-like illness rates across five states. It is not plotted from raw participant data. The individual panel is one synthetic person, generated the same way as the year strip at the top of this article.

Averaging is the whole trick. Every individual’s noise (their wine, their bad nights, their travel, their training) is essentially random with respect to everybody else’s. Add 47,000 people together and that noise cancels out, leaving only the thing they have in common, which is the virus moving through the population. The signal was always there. It was just buried under a much larger amount of personal chaos.

For you, alone, on a Tuesday, there is nothing to average over. You get the signal and all of your own chaos at once, with no way to separate them.

Forty-seven thousand people wearing a Fitbit can tell you what influenza is doing in Texas. One person wearing a Fitbit cannot reliably tell you what is happening to them.

CH 05The problem

Three beats per minute

Recall the threshold that NightSignal uses for a yellow alert. Overnight resting heart rate about three beats per minute above your baseline.

In 2025, a research group in Germany ran a controlled study on what alcohol does to overnight heart rate. Forty healthy adults, three alcohol-free baseline days, three days drinking a measured amount, three days of recovery, wearing smartwatches throughout. On the drinking nights, nocturnal resting heart rate rose from 63.6 to 66.6 beats per minute, then returned to baseline once they stopped.

That is an increase of three beats per minute.

Two or three drinks with dinner produces, to the decimal point, the signal that the algorithm is built to interpret as a possible incubating infection. Not something vaguely similar in the same direction. The same number.

And alcohol is only the most quotable example. Here is what else lives in the same neighbourhood as the alert line.

Everything that raises your overnight heart rate

Typical elevation in beats per minute above baseline, against the two published alert thresholds.

yellow alertred alertAn incubating infection+8.0Vaccination, second dose+3.4Two or three drinks+3.0Luteal phase+2.7A hard training session+2.0A short night of sleep+1.9Travel across time zones+2.40246810bpm above baseline, overnight

Mixed sources, stated per row. Alcohol at +3.0 bpm is measured (reference 7). Vaccination at +3.4 bpm follows the trajectory in reference 8. The luteal phase figure follows the direction established in reference 9 and cohort analyses of consumer wearable data. Hard training, short sleep and travel are modelled at magnitudes consistent with being named as alert triggers in reference 3, which does not publish per-cause effect sizes. The infection bar is an order of magnitude drawn from the detection literature, not a single reported value.

The uncomfortable shape of that chart is that the alert threshold does not sit above the ordinary business of being alive. It sits inside it. Vaccination clears it. Alcohol meets it exactly. For roughly half the population, the luteal phase of every single menstrual cycle brings resting heart rate up by something close to it, on a monthly schedule, forever.

None of this was hidden. The Stanford paper says it plainly: other respiratory infections, and events not associated with infection at all, such as stress, alcohol consumption and travel, could also trigger alerts. The researchers were not overselling anything. They reported the false alarms in the same breath as the detections.

What they reported, specifically, is that test-negative participants averaged 1.30 alert days per person per 21 days. Untested participants averaged 1.09. Do that arithmetic across a year and you get somewhere in the region of twenty alert days per person, in people who were not infected with the thing the system was looking for.

Twenty-something false alarms a year, against the two or three respiratory infections a typical adult actually gets. That ratio is the article. Everything from here is working out what it means.

CH 06The arithmetic

The base rate is the whole problem

There is a specific mistake almost everyone makes when reading a result like 80 percent sensitivity, and it is not a mistake about physiology. It is a mistake about which direction a conditional probability runs.

Sensitivity answers this question: given that you are getting sick, how likely is the alert to fire? The study says 80 percent, and that is genuinely good.

But when your ring buzzes on a Tuesday, that is not the question you have. Your question is the reverse: given that the alert fired, how likely is it that I am getting sick? Those two numbers are not the same, they are not close to the same, and converting between them requires one extra piece of information that no wearable can measure.

That piece is the base rate. How often, in truth, are you in the early days of an infection on a random night?

We can estimate it. The CDC puts the average adult at two to three colds a year. Add influenza and the various other things going around and call it three infections. Each has a detectable presymptomatic window of roughly three nights, going by the median lead time in the detection studies. Three infections times three nights is nine nights out of 365, which is about two and a half percent. Round it to two percent to be conservative.

So: roughly two percent of your nights are genuinely early infection. Ninety-eight percent are not. Now put the published numbers into that world and see what comes out.

One thousand nights

Sensitivity and specificity start at the published NightSignal values. Change the base rate and watch the answer to your actual question move.

You get an alert tonight. The chance you are actually getting sick

11.7%

7.5 false alarms for every real one. Averaged across a whole year, in and out of season.

1,000 nights

  • Alert, genuinely getting sick
  • Sick, no alert
  • Alert, nothing wrong
  • No alert, nothing wrong
Real, caught
16nights
False alarms
121nights
Missed
4sick, no alert
Correct silence
859nights

Why specificity dominates

Drag sensitivity from 50 to 100 percent and watch the headline number barely move. Now nudge specificity by two points. Because the well nights outnumber the sick ones by roughly fifty to one, a tiny false positive rate applied to a very large group produces more alerts than a high catch rate applied to a very small one.

Arithmetic, not a study.Sensitivity of 80 percent and specificity of 87.7 percent are as reported in reference 3, along with the CuSum and RHRAD comparison values. The base rate is derived on this page from reference 11 and is an estimate, which is exactly why it is a slider rather than a constant. Everything else follows from Bayes’ rule with no further assumptions.

At the default settings, an alert on an average night means roughly a one in nine chance that something is actually wrong. Eight times out of nine, it is your Saturday, your flight, your deadline, your cycle, or nothing identifiable at all.

This is not a flaw in the algorithm. Run the same numbers with a perfect 100 percent sensitivity and the answer barely improves, because sensitivity was never the constraint. The constraint is that a 12.3 percent false positive rate applied to 980 healthy nights produces about 120 false alerts, while an 80 percent catch rate applied to 20 genuine nights produces 16 real ones. The healthy nights outnumber the sick ones so heavily that they win on volume even while being individually much less likely to trigger anything.

Twenty sick nights and nine hundred and eighty well ones. The well nights only have to be wrong twelve percent of the time to bury the sick ones completely.

Now drag the base rate up to the household exposure setting. Your partner has tested positive, you share a bed, you are inside the window. Suddenly the probability that any given night is early infection is not two percent but something like thirty, and the same unchanged algorithm, with the same unchanged sensor, now gives an alert that is right about two times out of three.

Nothing about the device changed. What changed is the question it was asked. That is the single most useful idea in this article, and it is the reason the honest answer to does this work is that it depends entirely on what you already knew before you looked.

CH 07The trade

The threshold you cannot win

The natural engineering response to too many false alarms is to raise the bar. If three beats per minute is triggering on wine, use five. Or demand two consecutive nights instead of one.

This works, in the sense that false alarms fall. It also fails, in the sense that the thing you wanted goes with them. Every beat you add to the threshold removes an alert you did not want and, eventually, the warning you did. Worse, it costs you the thing that made the technology interesting in the first place, which is earliness. The earliest part of the signal is by definition the smallest part of it, so a higher bar does not just catch fewer infections, it catches them later, closer to the point where you would have noticed anyway.

Rather than take my word for it, take the dial.

Tune it yourself

The same synthetic year, with the alert rule under your control. Two real infections are hidden in there.

alert at +3 bpm0JanMarMayJulSepNovbpm above personal baseline, one night per bar
Infections caught
2of 2 in the year
False alarms
54separate episodes
Warning given
3.0dbefore symptoms, mean
Nights on alert
83out of 365

An alert every other week

54 false alarms in a year. Nobody keeps paying attention at this rate, which means the 2 real detections get ignored along with everything else.

Runs on the synthetic year described under the opening chart, so the specific counts belong to this constructed person and not to any population. The two preset buttons are the actual published thresholds from reference 3. The purpose here is the shape of the trade-off, which is general, rather than the individual numbers, which are not.

This is the standard shape of every detection problem ever built, and it has a name. You are moving along a receiver operating characteristic curve, trading sensitivity against specificity, and you cannot get both by tuning alone. The only way to genuinely improve is to find a better signal, or to add information the signal does not contain.

Notice, too, what happens if you tune until this particular year looks perfect. You will find settings where both infections are caught and almost nothing else fires. That feels like a triumph and it is actually the classic error: you have fitted your threshold to one person’s single year, including its accidents. Give the same settings to somebody who drinks more, travels more, or menstruates, and the false alarms come flooding back. This is precisely why the published studies report specificity across thousands of people rather than showing you one beautiful chart.

CH 08Consequence

What twenty false alarms a year do to a person

Suppose you accept the trade and leave the sensitivity high. You will get an alert roughly every other week. What happens next is not a technical problem, it is a human one, and it has been studied for decades in hospitals under the name alarm fatigue.

The pattern is consistent wherever it has been looked at. When most alarms are false, people stop responding to all of them, including the true ones. The response is not irrational. It is a correct adaptation to a signal that has been shown, repeatedly, to carry almost no information. A rational person who receives twenty meaningless alerts and two meaningful ones learns to ignore alerts, and there is no way for them to know in advance which category tonight belongs to.

So the false alarm rate does not merely add noise. It actively destroys the value of the true positives, which is a much worse failure than simply being unhelpful. A system with 80 percent sensitivity that nobody believes has an effective sensitivity of approximately zero.

There is a second cost, subtler and better documented in the sleep tracking literature than in the infection literature. Being told daily that your body might be failing has effects of its own. Anxiety raises resting heart rate. Anxiety damages sleep. Damaged sleep raises resting heart rate. It is not hard to see how a person who takes their alerts seriously can end up in a loop where worrying about the number moves the number, which produces more alerts, which produces more worry.

The asymmetry nobody designs for

A false negative costs you almost nothing, because you find out you are ill within a day or two anyway. A false positive costs you a cancelled training session, a skipped social commitment, or a day of low-grade health anxiety. Twenty times a year, that is a real amount of life.

Which means the sensible operating point for a consumer device is probably far more conservative than the one that maximises detection, and the studies were not designed to find it, because they were built to answer a research question rather than a product question.

CH 09Where it works

The situations where this genuinely earns its place

Everything above is an argument against one specific use, which is a general purpose illness alarm running on an ordinary person on an ordinary day. The base rate kills that use, and no engineering rescues it.

But the base rate is not a constant. It is a property of the situation, and there are plenty of situations where it is high enough to change the answer completely.

A known exposure. Somebody in your house is ill, or you spent an evening in a room with someone who tested positive two days later. Your prior probability is now high, the alert becomes genuinely informative, and the three days of lead time is actionable in a way it never is otherwise.

An outbreak in progress. During a severe seasonal wave, or the acute phase of a pandemic, population prevalence rises by an order of magnitude and so does the predictive value of every alert. This is precisely why these systems looked so good in 2020 and 2021 and rather less compelling afterwards. The algorithms did not degrade. The base rate fell, which is exactly the failure mode the systematic review warned about.

Populations where the cost of missing it is severe. Someone immunocompromised, undergoing chemotherapy, or recently transplanted has both a higher probability of infection and a vastly higher cost of finding out late. A false alarm that costs a healthy person an evening might cost that person nothing, while a true positive found three days early might genuinely matter.

Closed groups where transmission is expensive. Care homes, ships, military units, professional sports squads. Here the calculation changes because the cost of one person infecting thirty is enormous, and because these are settings where a positive alert can be followed immediately by an actual test rather than by guessing.

That last point is the general principle. The alert is not a diagnosis and was never capable of being one. It is a reason to perform a cheap test, and its value depends entirely on whether a cheap test is available and worth doing. In a world with a two-dollar rapid test in the bathroom cupboard, a wearable alert is a genuinely useful trigger. In a world without one, it is a feeling.

CH 10The literature

What happens when someone reviews all of it at once

Individual studies are enthusiastic. Systematic reviews, which exist to be unenthusiastic, tell a more measured story, and the honest version of this topic requires reading them.

In 2022, Mitratza and colleagues reviewed the wearable detection literature for The Lancet Digital Health. Twelve completed studies, with participant counts ranging from 29 to more than 32,000. Presymptomatic sensitivity across those studies ranged from 20 to 88 percent. Discrimination, measured as area under the curve, ranged from 0.52 to 0.92. The bottom of that range, 0.52, is very close to a coin toss.

That spread matters more than the headline figures from any single paper. When results vary that widely across studies, it usually means the performance depends heavily on the specifics: which device, which algorithm, which population, which variant, which season. It does not mean the effect is not real. It means you should be extremely cautious about carrying one study’s number into your own life.

The review made two other observations that are worth quoting more or less directly. The first is that none of the models detecting infection from physiological parameters had been tested or validated in real time. Retrospective analysis of a dataset where you already know who was infected is a fundamentally easier task than making a call tonight, in the dark, about tomorrow.

The second is the point this whole article turns on, and the reviewers got there first: the shifting prevalence of the disease could cause substantial overestimation of model performance. That is the base rate problem, stated in the passive voice of academic writing, sitting in a limitations section where nobody writing a headline was ever going to read it.

What would actually improve this

Not a better heart rate sensor. The heart rate signal is already good enough. What is missing is context, and context is information about your life rather than your body: whether you drank, whether you flew, whether you trained hard, where you are in your cycle, whether anybody near you is ill.

Every one of those, supplied to the algorithm, removes a chunk of the false positives without touching sensitivity, because it explains an elevation rather than raising the bar for it. That is a fundamentally different strategy from tuning a threshold, and it is the only one that escapes the trade-off in chapter seven.

CH 11Practical

How to read your own alert

So your device has flagged something this morning. Here is what the research above actually licenses you to conclude.

  1. Ask what you already knew

    Before the alert, was there any reason to think you might be getting ill? A sick child, a sick colleague, a wave moving through your city? If yes, the alert is meaningful and worth acting on. If genuinely nothing, it is probably noise, and the arithmetic in chapter six says so.

  2. Account for last night before blaming your immune system

    Alcohol, a late heavy meal, a hard session, a short night, a hot bedroom, a long flight, a stressful day, the luteal phase. Any one of these produces an elevation the same size as an incubating infection. If one of them applies, you have your explanation and it is not a virus.

  3. Two nights beats one

    Most confounders are single-night events. You drink on Saturday and you are back to baseline by Monday. An infection keeps climbing. A second consecutive elevated night, with no obvious cause, carries considerably more information than the first one, which is exactly why the published red alert requires two.

  4. Treat it as a prompt to test, not as a result

    The alert’s highest value use is telling you that this is a sensible morning to use a rapid test you already own, or to not visit your grandmother. It is a nudge toward cheap, reversible caution. It is not a diagnosis and it cannot be made into one.

  5. Do not cancel your life on one number

    At an average base rate, eight out of nine of these are nothing. Skipping a hard session on the back of a single unexplained alert, twenty times a year, costs more training than the two genuine infections ever will.

  6. Watch the trend, not the night

    A resting heart rate that has drifted up over three weeks is telling you something real about accumulated load, poor sleep or overreaching. That slow signal is considerably more reliable than any single night, and almost nobody looks at it because the alert is louder.

CH 12Conclusion

Both things are true

The headline at the top of this piece is accurate. Your resting heart rate does know before you do. In a well run trial on more than three thousand people, a consumer device caught four out of five infections at a median of three days before symptoms, and found asymptomatic cases that would otherwise have gone entirely unnoticed. That is a real scientific result and it deserves to be taken seriously.

It is also true that the same system, on the same wrist, will interrupt you roughly twenty times this year for no reason at all, and that on a random Tuesday with nothing going around, an alert carries about a one in nine chance of meaning anything.

Those two statements are not in conflict. They are the same statement, viewed from the two ends of a conditional probability. The gap between them is not a scandal, it is not marketing dishonesty, and it is not a sign that anyone did bad science. It is arithmetic, and it is the most reliably misunderstood arithmetic in consumer health.

The practical conclusion is smaller and more useful than either the enthusiastic version or the cynical one. A wearable alert is not information about your body. It is a slight update to a probability, and how much it should move you depends almost entirely on what you believed before it arrived. On a quiet week in June, it should barely move you at all. When your partner is coughing in the next room, it should move you a great deal.

The device cannot tell the difference between those two Tuesdays. You can.

REFSources

References and sources

Links resolve by title search rather than by a typed identifier, so a misremembered volume or page number cannot silently send you to the wrong paper. Where this article used a figure approximately, or built something from stated assumptions, it says so both here and underneath the chart itself.

Which charts are data and which are models

Two visuals in this piece present published figures: the four study cards in chapter three, and the reported performance values seeded into the Bayes calculator. The confounder chart mixes measured values with modelled ones and labels each row. Everything else is explicitly constructed: the year strip and the threshold engine run on a synthetic person built from published effect sizes, the aggregate curves are shaped to reproduce a published correlation rather than plotted from raw data, and the presymptomatic timeline is an illustrative trajectory. Nothing here is a measurement unless its source line says so.

  1. Harnessing wearable device data to improve state-level real-time surveillance of influenza-like illness in the USA: a population-based study

    Radin JM, Wineinger NE, Topol EJ, Steinhubl SR. The Lancet Digital Health, 2020. Find it

    Used for: The population surveillance result: 47,249 consistent Fitbit users across five states, more than 13 million resting heart rate and sleep measures, and final model correlations with reported illness rates of 0.84 to 0.97.

  2. Pre-symptomatic detection of COVID-19 from smartwatch data

    Mishra T, Wang M, Metwally AA, et al.. Nature Biomedical Engineering, 2020. Find it

    Used for: Of 32 infections in a cohort of about 5,300, 26 showed physiological changes. Of 25 with symptom dates, 22 were detected at or before symptom onset and four at least nine days early. Retrospectively, 63 percent could have been flagged before symptoms by a two-tier resting heart rate warning system.

  3. Real-time alerting system for COVID-19 and other stress events using wearable data

    Alavi A, Bogu GK, Wang M, et al.. Nature Medicine, 2022. Find it

    Used for: The central reference for this article. 3,318 participants, 84 infections, 67 detected (80 percent) at a median of three days before symptoms, specificity 87.7 percent. The NightSignal thresholds of 3 bpm for a yellow alert and 4 bpm across two nights for a red alert. Alert days per person per 21 days of 1.30 in test-negative and 1.09 in untested participants against 3.42 before infection. Stress, alcohol, travel and intense exercise named as triggers.

  4. Wearable sensor data and self-reported symptoms for COVID-19 detection

    Quer G, Radin JM, Gadaleta M, et al.. Nature Medicine, 2021. Find it

    Used for: The DETECT study. 30,529 enrolled, 54 test-positive and 279 test-negative symptomatic participants. Sensor plus symptom data reached an area under the curve of 0.80 against 0.71 for symptoms alone.

  5. Feasibility of continuous fever monitoring using wearable devices

    Smarr BL, Aschbacher K, Fisher SM, et al.. Scientific Reports, 2020. Find it

    Used for: The first published TemPredict result. Continuous ring temperature identified elevation in 38 of 50 cases before participants reported symptoms that made them suspect illness.

  6. The performance of wearable sensors in the detection of SARS-CoV-2 infection: a systematic review

    Mitratza M, Goodale BM, Shagadatova A, et al.. The Lancet Digital Health, 2022. Find it

    Used for: The limitations chapter. Twelve completed studies, participant counts from 29 to 32,198, presymptomatic sensitivity from 20 to 88 percent and area under the curve from 0.52 to 0.92. None of the physiological models were validated in real time, and the review warns that shifting prevalence can substantially overestimate model performance.

  7. The Impact of Alcohol on Sleep Physiology: A Prospective Observational Study on Nocturnal Resting Heart Rate Using Smartwatch Technology

    Guo H, Weingart M, Kolls JK, et al.. 2025. Find it

    Used for: The 3 bpm figure for alcohol. Forty healthy adults, three alcohol-free days, three exposure days at 60 g for men and 40 g for women, three recovery days. Nocturnal resting heart rate rose from 63.6 to 66.6 bpm and returned to baseline afterwards, Cohen's d 0.58.

  8. Inter-individual variation in objective measure of reactogenicity following COVID-19 vaccination via smartwatches and fitness bands

    Quer G, Gadaleta M, Radin JM, et al.. npj Digital Medicine, 2022. Find it

    Used for: The vaccination confounder. In more than 5,600 wearable users, resting heart rate rose the day after vaccination, peaked about two days after, and returned to normal by day four after a first dose and day six after a second.

  9. Menstrual cycle changes in vagally-mediated heart rate variability are associated with progesterone: evidence from two within-person studies

    Schmalenberger KM, Eisenlohr-Moul TA, Jarczok MN, et al.. Journal of Clinical Medicine, 2020. Find it

    Used for: The menstrual cycle confounder. Vagally mediated heart rate variability falls from the follicular to the luteal phase, tracking progesterone, with resting heart rate moving in the opposite direction.

  10. Assessment of physiological signs associated with COVID-19 measured using wearable devices

    Natarajan A, Su HW, Heneghan C. npj Digital Medicine, 2020. Find it

    Used for: Supporting evidence on the magnitude and timing of respiratory rate and heart rate changes around infection.

  11. About Common Cold

    Centers for Disease Control and Prevention. CDC, accessed 2026. Find it

    Used for: The base rate input: adults average two to three colds a year. This is what sets the roughly 2 percent of nights used as the default in the Bayes calculator.

  12. Assessment of the feasibility of using noninvasive wearable biometric monitoring sensors to detect influenza and the common cold before symptom onset

    Grzesiak E, Bent B, McClain MT, et al.. JAMA Network Open, 2021. Find it

    Used for: Evidence that the presymptomatic signal is not specific to one pathogen, and appears for influenza and common cold viruses in controlled challenge studies.

  13. Early adverse physiological event detection using commercial wearables: challenges and opportunities

    Breteler MJM, Numan L, Ruurda JP, et al.. npj Digital Medicine, 2024. Find it

    Used for: The alert burden and data quality problems that appear when detection systems are deployed rather than analysed retrospectively.

Written for general interest. Nothing here is medical advice. If a wearable alert coincides with symptoms that concern you, or you have a condition that affects your heart rate or immune function, talk to a clinician rather than to an algorithm.

Reported performance figures in this article, including sensitivity of 80 percent and specificity of 87.7 percent, come from a study of 3,318 participants of whom 84 were infected. The 54false alarms referenced in chapter one are from this article’s synthetic year, calibrated to the alert frequency that study reported.

← All deep dives