Accuracy

The model is graded against people, and people disagree

When a tool says it matches human judgement, the claim depends on which humans, how many, and how much they agreed with each other.

By 5 min readAccuracy

Guides on Accuracy: There is no true score to be accurate against, Change one thing, hold the rest, repeat, Exposure that suits one skin tone hides another

"Matches human judgement" sounds like a single, checkable fact. It is closer to a small research design that has to be evaluated before the claim means anything. Which humans, how many, instructed how, and compared using which statistic - each choice changes what the resulting number is actually saying.

Correlation is not exact match

The most common way to report human agreement is a correlation coefficient between the model's scores and a panel's average score across a set of test images. A strong correlation means the model tends to rank images the same way the panel did, roughly. It does not mean the model's number and the human number were ever close on any individual photo. A model can correlate well while being systematically off by a point or two in one direction, which a correlation coefficient will not penalise the way an exact-match statistic would. Reading "correlates strongly with human ratings" as "gives the same number a person would give" is the most common misreading of this kind of claim, and it is worth checking which statistic was actually used before assuming either.

The panel is doing more work than the model

A model evaluated against five people who broadly share a taste will look highly accurate. The same model evaluated against five hundred strangers with a wider range of taste will look less accurate, even with no change to the model itself, because the target it is being compared to got noisier. This means panel size and panel composition are not incidental details of an accuracy report - they set the ceiling the model is being measured against, before anything about the model's own quality enters the picture. A small, self-selected panel is the easiest way to produce an impressive-looking agreement number, and it is also the least representative of what a general audience would say.

Instructions given to the panel matter just as much. A panel told to rate "overall appeal" and a panel told to rate "how likely you'd be to right-swipe on this" are answering different questions, even looking at the same photos, and a model's agreement with one tells you little about its agreement with the other. Anchoring a scale with reference examples shapes what a panel's numbers mean before a single test photo is shown to them, and an agreement statistic inherits whatever the anchors baked in.

The ceiling nobody mentions

Even with a large, well-instructed panel, humans do not agree with each other perfectly. When labellers disagree, the model cannot beat the ceiling their disagreement sets - if two independent human ratings of the same photo only agree within a point or two, on average, no model trained to predict "the" human rating can be expected to do noticeably better, because there was never a single target to hit. A reported accuracy number is only informative in context of this ceiling. A model matching humans 80% of the time sounds mediocre in isolation and sounds close to the practical limit if the humans themselves only agreed with each other 82% of the time. Almost no tool publishes the second number alongside the first, which makes the first one close to unreadable on its own. Raw percentages flatter everyone, too: as McHugh (2012) recounts, Jacob Cohen critiqued percent agreement in 1960 for "its inability to account for chance agreement," which is why chance-corrected statistics such as kappa exist.

What a well-reported agreement study looks like

A study worth trusting states the panel size, describes how the panel was recruited and instructed, reports the panel's own inter-rater agreement as the ceiling, and then reports the model's agreement against that ceiling rather than against an abstract maximum of 100%. This is a high bar, and most published claims about AI rating tools clear none of it - a single sentence, "matches expert ratings," with no numbers attached, is closer to marketing copy than to a result. It is not evidence the underlying claim is false. It is evidence the claim has not been made checkable, which is a different and more common failure than dishonesty.

Where this sits among the other candidates

Human agreement is one of several substitutes for a ground truth that does not otherwise exist for a subjective rating - the ground-truth problem generally covers the full set, including self-consistency and predictive validity, and each has different strengths. Agreement with humans is the most intuitive of the three because it is the one people assume "accurate" already means. It is also the one most sensitive to a methodological choice - the panel - that a headline number rarely discloses. At the extremes of a scale, model and human judgement tend to converge regardless of panel details, since clearly high and clearly low inputs are where agreement is easiest to find; the interesting cases, and the ones a panel-composition problem actually bites on, sit in the wide middle where most real photos land.

What to do with a stated agreement number

Treat "matches human raters" as an invitation to ask about the panel, not as a settled fact. Rate Cock reports its per-axis breakdown rather than a single blended agreement figure, which at least lets you see whether a model tracks people closely on the axes that are easiest to agree on, like symmetry, and less closely on the ones that are closer to taste. A panel of human judges brings its own version of the same problem, since a panel's shared taste is also a specific, unstated reference point rather than a neutral standard, and it is worth reading that side with the same scepticism. Comparing your own scores across submissions over time sidesteps the panel question almost entirely, which is the angle penisrater.com takes when it treats a score as something to track rather than something to validate against a reference group. None of this touches anything measured in centimetres, which has an actual reference to agree against rather than a panel's opinion, and that is a different kind of claim entirely.

Read next

Full archive