Accuracy
Agreement is high at the ends and low in the middle
Model and human ratings tend to line up on clearly high and clearly low inputs and scatter across the middle, which is where most inputs sit.
A model and a human panel scoring the same photo do not disagree uniformly across the scale. They tend to converge at the extremes and diverge in the middle, and the middle is where nearly everything actually lands.
Why the extremes are easier
A photo that is clearly excellent or clearly poor on the properties either judge is reading tends to produce a strong, low-ambiguity signal, and both a trained model and a trained person pick up strong signals reliably. There is less room for taste to matter when the input is unambiguous in one direction. This is also where class imbalance in training data matters least in relative terms - even a model with thin extreme-score examples usually still knows an obvious case when it sees one, even if it hedges the exact number.
Why the middle scatters
Most real submissions are not extreme. They sit in the ordinary range where the properties being judged are mixed - decent on one axis, average on another - and that is exactly where individual taste, model bias, and genuine ambiguity in the input all have room to pull the score in different directions. A model and a person can both be "right" about the same middling photo and land on different numbers, because the judgement genuinely has more give in it there.
What this means for reading a score
A score near either end of the range is more likely to reflect something both methods would agree on if compared directly. A score in the middle is more likely to be the specific method's opinion rather than a shared fact about the photo, which is worth remembering before treating a 5 or a 6 as more precise than it is. Agreement with human raters as a metric usually gets reported as one aggregate correlation figure, which flattens this pattern away entirely - a tool can report solid overall agreement while still scattering badly across the range most people actually fall in.
This pattern is not unique to any one product; it follows from where each method's ambiguity is highest, and it is worth knowing before treating a number two points from the middle as a precise verdict rather than one read among several plausible ones. Where the two genuinely diverge is exactly the territory a human reviewer is best placed to actually explain, a repeated score is best placed to bound, and a tape measure has nothing to say about at all - it answers a completely different question. Rate Cock reports the breakdown by axis precisely so a middling total does not read as more settled than it is.