Accuracy

Optimising for a model stops measuring what the model meant

Once people adjust inputs to raise a score, the score tracks the adjustments rather than the thing it was built to reflect.

4 min readAccuracy

Goodhart's Law, in its common phrasing, is that when a measure becomes a target, it stops being a good measure. A scoring model built to reflect some underlying quality is only trustworthy while people are not actively trying to move the number. Once they are, the score starts tracking the strategies people found to move it, which is a narrower and less honest thing than what the score was originally built to reflect.

Why this happens to any proxy

A model score is a proxy - a stand-in, learned from training data, for a broader and fuzzier target that the model cannot access directly. The proxy and the target agree closely as long as the inputs it sees look like the training data it was built from. The moment someone starts adjusting an input specifically to move the proxy, rather than to change the underlying thing the proxy was meant to reflect, the two can come apart. This is not specific to AI, and it is not a flaw unique to any one tool - it is a property of any measurement that people can see and respond to, which is why the same pattern shows up in test scores optimised at the expense of learning, and in engagement metrics optimised at the expense of the content actually being good.

What it looks like with a scoring model specifically

A model trained to score photographs learned its sense of "good" from correlations in training data - lighting, angle, framing, all the presentation variables that move a score for reasons that have nothing to do with the subject. If someone learns which specific lighting setup, camera distance, or crop the model rewards most, and starts optimising directly for that combination rather than for what an actual viewer would find appealing, the score can rise while the photo, judged by a person with no stake in the number, does not improve in the same way, or improves less than the score suggests. This is a subtler failure than gaming a test with a known answer key, because nobody has to act in bad faith for it to happen - a person genuinely trying to get a good photo will naturally converge, through trial and error, on whatever the tool happens to reward, without ever intending to exploit anything.

Where the line to adversarial inputs sits

There is a harder, more deliberate version of this same idea, where someone makes a targeted change specifically designed to move a model's output regardless of what a person would perceive - that is its own subject, and it is a sharper, more mechanical failure than what Goodhart's Law describes. Goodhart's Law is closer to the ordinary case: nobody is trying to exploit a vulnerability, everyone is just responding rationally to visible feedback, and the aggregate effect is the same drift away from the original target regardless of intent. Both end with the score meaning less than it did before people started paying attention to it, but they get there by different routes and call for different responses.

Why this compounds with score inflation

A related effect shows up at the level of the whole tool rather than a single submission. If users who get lower scores are less likely to return, and the people making product decisions can see that pattern, there is a slow pressure toward scoring more generously over time even without anyone deciding to cheat the system - that drift is its own subject, and it is Goodhart's Law operating one level up, on the company rather than the individual user.

What this does not mean

None of this means a score is worthless once people know it exists, or that trying to improve toward a score is inherently dishonest. It means the score is a less reliable signal the more directly people are optimising for the number itself rather than for whatever the number was meant to stand in for, and that gap widens with scale and with how well-known a tool's specific preferences become. A score that has not yet been widely gamed is closer to its original meaning than one that has. There is no way to fully immunise a proxy measure against this - only ways to make the proxy harder to game, which is part of why a multi-axis breakdown is more resistant to this failure than a single blended total, since moving six independent numbers at once is a harder trial-and-error problem than moving one.

Reading a score with this in mind

Rate Cock reports on six separate axes rather than one aggregate for this exact reason - it is a harder target to optimise around than a single number, and a shift concentrated on one axis while the others stay flat is itself a visible signal something narrower than overall quality is being pushed. A human panel is not immune to a version of the same effect either, since people can learn what a specific panel or set of judges tends to reward just as readily as they can learn what a model rewards. Comparing your own submissions over time, the way penisrater.com frames the practical use of a score, is a more honest use of a proxy than chasing the highest number a tool will give, since it treats the score as a rough compass rather than as the actual destination. None of this touches a measured figure, which cannot be moved by learning what a model prefers, because there is no model in that path at all.

Read next

Full archive