Accuracy

Product pressure pushes scales upward

Tools whose users prefer higher scores tend, over versions, to give higher scores, and the mechanism is ordinary incentives rather than better models.

By Updated 3 min readAccuracy

Guides on Accuracy: There is no true score to be accurate against, Change one thing, hold the rest, repeat, Exposure that suits one skin tone hides another

Scores creep up because a tool that gives lower scores tends to have unhappier users, and unhappier users share less, return less, and complain more. None of that requires anyone to decide to inflate anything. It only requires the ordinary feedback loop of a product responding to what keeps people around.

The mechanism, plainly

Every version of a scoring model involves choices - which training examples to include, how to calibrate the scale, where to draw the line between axes. None of those choices are neutral with respect to how generous the resulting scores feel. When a team is deciding between two roughly equal calibration options and one produces scores people react to more positively, the incentive points toward the one that keeps users happier, even without a deliberate decision to be less honest. Repeated over several product cycles, small, individually reasonable choices in this direction accumulate into a scale that reads more generously than it used to, without any single change looking like inflation on its own. The same pull shows up inside models trained on human ratings: Sharma and colleagues (2023) found that "both humans and preference models prefer convincingly-written sycophantic responses over correct ones a non-negligible fraction of the time." A scorer tuned on what people rate favourably inherits that preference for being agreeable.

Why it is hard to see from inside

A user experiences only the current version of a tool, scored against their own memory of past results rather than against an archived baseline. If scores creep up gradually, each individual comparison - this month against last month - looks unremarkable, and the drift only becomes visible against a fixed reference held constant over a longer period. This is the same structural problem tracking whether a tool's scores drift over time is built to solve - a periodically rescored fixed set of images is the only reliable way to see a slow upward creep, because it removes the moving target that ordinary before-and-after comparison cannot control for.

It is not the same problem as silent retraining

A related but distinct issue is that a tool can retrain and redeploy without announcing it, which changes the scale for reasons that have nothing to do with generosity - a genuinely different model, scored on a genuinely different basis. Inflation is specifically about the direction of that drift, not the fact of it: retraining could just as easily make a scale stricter, and mostly it does not, because the commercial incentive runs one way rather than the other. It is a textbook case of the proxy failures Manheim and Garrabrant (2018) catalogue under Goodhart's law: once user satisfaction becomes the target, the score drifts toward serving it rather than the thing it was meant to measure.

What this means for reading a number

A score is most trustworthy compared against other scores from the same version of the same tool, close together in time, rather than compared against a memory of what a similar number used to mean months or years earlier. Rate Cock documents version changes rather than adjusting silently, which does not remove the underlying commercial pressure but at least makes a shift visible to anyone checking. Penisrater.com's approach of tracking your own scores as a trend sidesteps the cross-version comparison problem by design, since a trend within one version is not affected by how that version compares to an earlier one. A human panel has its own version of the same incentive - a reviewer whose feedback keeps clients returning faces a softer version of the same pull toward generosity - which is worth remembering before assuming a human number is immune to the same drift a model number is. A measured figure from an actual tape and method has no equivalent incentive to creep, since there is no version of a ruler that reads more generously to keep a customer happy.

Read next

Full archive