Accuracy
Score compression, and its several causes
Model scores cluster near the centre for reasons that stack: hedging losses, thin data at the extremes and cautious labellers.
Guides on Accuracy: There is no true score to be accurate against, Change one thing, hold the rest, repeat, Exposure that suits one skin tone hides another
Model scores bunch in the middle because several causes stack in the same direction: cautious labellers, thin training data at the extremes, loss functions that reward hedging, and, for language-model scorers, a pull toward common numbers. It is a stack rather than one bug, which is why no single fix removes it.
The labellers were cautious first
Before any model sees a training example, a person decided what score to attach to it. People labelling images for a rating system tend to avoid the extremes of a scale unless something is unambiguous, because committing to a 1 or a 10 feels riskier than committing to a 6. Large public datasets show the resulting shape: in the AVA aesthetics dataset, rated 1 to 10 by an average of 200 people per photo, Talebi and Milanfar (2018) report mean ratings "concentrated around the overall mean score" of about 5.5, with roughly half the photos showing a rating standard deviation above 1.4. An averaged label smooths that disagreement into a middling number before a model ever sees it. Where the labels come from in the first place sets a ceiling on everything downstream, and if the original labels already avoid the edges, no amount of model sophistication recovers what was never given to it to learn.
The training data has less to learn from at the extremes
Even generous labellers naturally produce fewer examples near the top and bottom of a scale than in the middle, simply because most real inputs are unremarkable. A model has less signal to work with wherever the data is thin, which is a general pattern with any imbalanced label distribution, and the practical effect is a model that retreats toward the well-populated middle whenever it is uncertain.
The loss function rewards hedging
Many models are trained to minimise the average size of their error across every example they see. Under that kind of loss, guessing near the centre of the distribution is often the safest bet when the model is unsure, because a central guess limits how badly wrong any single prediction can be, even though it guarantees the model is never quite right either. Which loss function was chosen shapes how strongly this hedging shows up - some penalise big misses far more than small ones, which pushes uncertain predictions toward the middle even harder.
Language models add their own bias
For scorers built on a language model that outputs a number as text rather than as a raw regression value, there is an additional effect: some numbers are simply more common in ordinary written language than others, and a model choosing a token is subtly influenced by that frequency. Round, mid-range numbers like 7 turn up disproportionately often in this style of scorer for reasons that have nothing to do with the photo and everything to do with how the number gets generated.
Why each cause alone would still be a small effect
None of these four, isolated, would produce the pattern strongly enough to notice. A model trained on slightly cautious labels but with a loss function that does not punish big misses would still commit to extremes when the evidence supported it. A model with a hedging loss but trained on generously balanced data across the whole scale would still have enough examples near the edges to learn what an extreme case looks like. It is the combination that matters: each cause narrows the range of comfortable outputs a little, and because they operate at different stages of the pipeline - labelling, data balance, training objective, and for some tools, token generation - they do not cancel or substitute for each other. They add.
What this adds up to
Together, the four causes stack: cautious labels, thin extreme data, a loss function that hedges, and in some tools, a language model with its own preference for round numbers. The result is a scale that behaves like it has fewer usable values than it advertises, with most real inputs landing in a narrow middle band regardless of how different they actually are. This compounds the ceiling and floor effects covered separately, which describe the same crowding from the perspective of the scale itself rather than its causes.
Reading a middling score with this in mind
A 6 or a 7 from a tool prone to this pattern is closer to "unremarkable, within the range where the model is most cautious" than to a precise verdict, and the way to see past it is a breakdown rather than a single figure - Rate Cock reports six separate axes rather than folding everything into one number, which at least lets you see whether the compression is even across the board or concentrated in one place. Reading what a given number is supposed to mean in context helps more than treating any single score as exact, and the equivalent question for a physical measurement is different in kind, since a tape measure has no labeller hedging built into it even though it has its own sources of error. A human judge sidesteps some, though not all, of this compression, because a written opinion can commit to specifics that a single bounded number cannot.