Accuracy

An extreme score is more likely to be followed by a middling one

If your first score was unusually high, the next is likely lower for purely statistical reasons, and people read that as the tool changing.

By Updated 3 min readAccuracy

Guides on Accuracy: There is no true score to be accurate against, Change one thing, hold the rest, repeat, Exposure that suits one skin tone hides another

Your retake probably scored worse because of regression to the mean, not because the tool changed its mind. When a measurement has noise in it, an unusually high result tends to be followed by one closer to average, even under near-identical conditions. The obvious reads - the tool is inconsistent, the first result was a mistake - are often wrong.

The mechanism, without the jargon

Every score a model gives you is made of two things: something real about the input, and some amount of noise from the sources covered elsewhere on this site - sampling, preprocessing, the exact framing that afternoon. When a score comes back unusually high, some of that extremeness is real and some of it is noise that happened to point upward that particular time. Noise, by definition, does not repeat in the same direction on the next attempt. So the second score tends to fall back toward the average, not because anything got worse, but because the lucky noise from the first attempt was never going to show up twice in a row.

This is not specific to AI scoring. Barnett, van der Pols and Dobson (2005) describe it as a statistical phenomenon that can make natural variation in repeated data look like real change. It is the same reason a sports team's best-ever performance is usually followed by a merely good one, and the reason a student's highest test score in a term rarely repeats itself the next time round. Extreme results are disproportionately made of extreme luck, and luck regresses.

Why it feels like the tool changed its mind

People naturally read a drop as causal - the lighting must have been worse, the angle must have been off, the model must have updated. Sometimes one of those things is genuinely true. But even holding everything constant as closely as a retake allows, the statistical pull toward the mean is enough on its own to produce a lower second score after an unusually high first one, and there is no need to invent a cause beyond ordinary variance. The same pull runs in the other direction too: an unusually low first score tends to be followed by a higher second one, for the identical reason, which is worth remembering before concluding a bad result was the "true" one either.

How to tell regression from a real change

The way to separate the two is to keep going past two attempts. A single retake cannot distinguish regression to the mean from an actual change, because both predict the same direction of movement. Several retakes can, because regression pulls scores toward the middle of the distribution and then they settle there, while a real change - a genuinely better photo, a genuinely different lighting setup - shifts the whole distribution and keeps the later scores away from where the first one sat. How many repeats that settling actually takes is covered separately, and it is more than two.

This is a different phenomenon from what happens at the extreme ends of a scale, where the model runs out of room to distinguish inputs at all - regression to the mean can happen anywhere on the scale, not only near the top. It is also distinct from the wider variance a single afternoon of scoring typically shows, which this pattern is one specific piece of, rather than the whole explanation.

The practical takeaway

Do not treat one unusually high score as your baseline, and do not treat the retake that follows it as a correction. Both are single draws from a spread, and the spread, not either individual point, is what tells you anything. Rate Cock reports separate axes rather than one blended total specifically because a single number invites exactly this kind of overreading, and a per-axis view makes it easier to notice when only one figure moved rather than the whole picture. The tape-measure equivalent exists too: a single reading can be an outlier for entirely mechanical reasons, and a repeatable method accounts for that the same way repeat AI scoring does. Reading any one result against the pattern rather than in isolation is the habit worth building, and it applies just as much to a human judge's opinion delivered on a single, particular day as it does to a machine score.

Read next

Full archive