The Nutrition Dex

Dietary Assessment

Regression to the Mean

Also known as: statistical regression

The tendency for an extreme measurement to be followed by a less extreme one purely by chance, which manufactures the appearance of improvement where none occurred.

By James Oliver · Editor & Publisher ·

Key takeaways

  • Any measurement containing random error will, on average, be closer to the mean the second time it is taken.
  • It creates false evidence of effectiveness whenever a group is selected for being extreme and then re-measured.
  • In app testing it inflates apparent improvement between versions and between test rounds.
  • In personal tracking it is why a change made after an unusually bad week appears to work.
  • The defence is a control comparison or a pre-specified measurement schedule — not a larger sample.

Regression to the mean is a statistical certainty rather than a phenomenon: if a measurement contains any random error, an unusually extreme reading will on average be followed by a less extreme one, with no underlying change required.

It is the most reliable manufacturer of false evidence in measurement work, and it is invisible at the moment the mistake is made.

The mechanism

Any reading is a true value plus noise. An extreme reading is more likely to have had extreme noise pushing it out than to represent an extreme true value. Measure again and that particular noise is unlikely to repeat, so the reading moves back toward the centre.

Nothing changed. The instrument and the subject are identical.

Where it bites in instrument testing

  • Re-testing the worst cases. Pick the dishes an app handled worst, re-run them after a model update, and part of any improvement is regression. Without a control you cannot say how much.
  • Version-over-version comparisons. A version that scored unusually badly in one round tends to score better in the next regardless of engineering effort.
  • Recruiting the extreme. Any study enrolling the worst-performing users and re-measuring them shows improvement. It is a very common design.

Where it bites in personal tracking

You have an unusually bad week — weight up, adherence poor. You change something: a new app, a new split, a supplement. The following week is better.

Part of that is regression, because you selected the moment to intervene by the extremity of the reading. The bad week was partly noise, and noise does not repeat on schedule.

Interventions started after a bad week are systematically over-credited. Interventions started after a good week are systematically under-credited.

The defence is the ordinary one: judge from a rolling average across at least three weeks, and where possible decide when you will measure before you see the extreme reading.

What does not help

A larger sample. Regression to the mean arises from the selection of an extreme starting point, not from imprecision, so measuring more of the same selected group does not remove it. The corrective is a control comparison or a schedule fixed in advance.

Frequently asked

What is regression to the mean in plain terms?

If a measurement contains any random noise, an unusually extreme reading tends to be followed by a less extreme one even when nothing has changed. An extreme reading is more likely to have had extreme noise pushing it out than to reflect an extreme true value, and that particular noise is unlikely to repeat. The result looks exactly like improvement and is not.

How does it mislead people tracking their own weight?

Because people change something after a bad week, and the bad week was partly noise. The following week is better in part for reasons unrelated to the change, so interventions begun after a bad week are systematically over-credited and those begun after a good week under-credited. The defence is judging from a rolling average of at least three weeks, and where possible fixing when you will measure before you see the extreme reading.

Does a bigger sample size fix regression to the mean?

No. It arises from selecting an extreme starting point rather than from imprecision, so measuring more of the same selected group does not remove it. What removes it is a control comparison, or a measurement schedule pre-specified before the extreme value was observed. This is one of several places where the instinct to collect more data addresses the wrong problem.

References

  1. Barnett AG, van der Pols JC, Dobson AJ. "Regression to the mean: what it is and how to deal with it". International Journal of Epidemiology , 2005 .
  2. Bland JM, Altman DG. "Some examples of regression towards the mean". BMJ , 1994 .

Related terms