Accuracy

Averaging people vs trusting a machine

Several humans average out individual taste; one model averages nothing and is consistent about its bias. Which you want depends on the question.

By 3 min readAccuracy

Guides on Accuracy: There is no true score to be accurate against, Change one thing, hold the rest, repeat, Exposure that suits one skin tone hides another

Neither noise is simply worse. A panel of human raters averages away individual taste but drifts, while a single model is rock steady and consistent about whatever bias it learned. Which failure hurts more depends on whether you are comparing your own photos or asking an absolute question.

What a panel actually does

Put several independent raters on the same input and average their scores, and idiosyncratic taste - one rater who runs harsh, one who runs generous - tends to cancel out, leaving something closer to a shared centre. This works because the raters' individual quirks are at least partly uncorrelated: they diverge from each other for different, personal reasons. More raters narrows the spread further, with diminishing returns, and the panel's remaining disagreement is itself informative - a question the raters split on tells you something a unanimous one does not.

What a single model does instead

A model produces the same output for the same input every time, which looks like the panel's averaged stability but is not the same thing. There is no cancellation happening, because there is only one opinion. Whatever systematic bias the model picked up in training - toward certain body types, certain lighting, certain framing - is present in every single score it gives, applied with the same consistency that makes it useful for comparison. An ensemble of several models narrows the spread the way a human panel does, but only if the models disagree for independent reasons; if they share a training lineage or the same biased dataset, averaging them together does nothing to remove what they agree on incorrectly, which is the opposite of what averaging a human panel does.

The panel logic has been tested on model judges directly. Verga and colleagues (2024) found that a panel of diverse smaller language models outperformed a single large judge, "exhibits less intra-model bias", and cost over seven times less. Spreading the judgement across different models, rather than enlarging one, is the same logic a human panel runs on.

The actual trade-off

A human panel's noise is mostly random and partly cancellable; a single model's noise is mostly systematic and does not cancel by asking it twice, because asking it twice returns the same answer. This means the two approaches fail differently rather than by different amounts: a panel drifts session to session and rater to rater but is not stuck any particular direction, while a model is rock stable but stuck wherever its training put it.

Which is worse depends entirely on the question being asked. For comparing two of your own photos against each other, a fixed bias barely matters because it applies equally to both; consistency is the whole point. For asking whether a submission is unusually good in some absolute sense, a systematic model bias is a real problem a panel's averaged noise is less prone to.

Panel size changes this calculation too. A two-person panel barely averages anything - one strict rater and one generous one still shows up as a wide spread rather than a settled centre, and the benefit of averaging only really shows once there are enough independent raters for individual quirks to plausibly cancel rather than just offset each other by luck. A single model, by contrast, gets none of this benefit from scale in the ordinary sense: running the same model once or a hundred times on an unchanged input returns the same answer, because there is no independent variation between runs to average over in the first place. The nearest equivalent for a model is running several different models and averaging across those, not running one model repeatedly, and even that only helps if the models genuinely disagree for independent reasons rather than sharing the same blind spot.

Two products, two failure modes

This is not an argument that one approach is strictly better. Rate Cock uses the model approach and reports a breakdown by axis rather than pretending the number is bias-free. A commissioned panel of judges is the other route, and how that process is actually run is worth understanding on its own terms rather than as a simple upgrade. Reading either result with the failure mode in mind - not just the number - is the practical habit worth keeping, and it applies however many raters, human or otherwise, produced the score. A tape measure sidesteps both failure modes for the one axis it covers, which is a different method entirely and not a substitute for either kind of judgement, since neither a panel nor a model was ever measuring in the first place.

Read next

Full archive