How it works
How the training objective shapes the numbers you get
Whether the model was penalised for big misses or for any miss at all changes whether it hedges toward the middle or commits to extremes.
Guides on How it works: Every score is a memory of somebody's opinion, An axis earns its place by being separately observable, What an image becomes before it is scored
The loss function decides whether a scorer hedges or commits. Squared error makes big misses expensive, so an uncertain model drifts toward the middle of the scale; absolute error does not, so a model trained on it uses the ends more readily. Someone makes that choice once, at training time, and it leaves a clear fingerprint on every score.
What a loss function actually does
During training, the model makes a guess, the guess is compared to the label, and the loss function converts the size of that error into a number the model is told to minimise. Different loss functions convert the same-sized error into very different penalties, and the model adapts to whichever penalty it was actually given, not to some abstract notion of correctness.
Mean squared error - the most common choice for a scorer that outputs a continuous number - squares the difference between guess and label. A miss of two points costs four times as much as a miss of one point. Mean absolute error treats every point of error the same regardless of size, so a miss of two costs exactly twice a miss of one. Cross-entropy, used when the model is choosing among discrete buckets rather than outputting a raw number, penalises confident wrong answers especially hard - being very sure of the wrong bucket is punished far more than being unsure and wrong.
The visible consequence: hedging
Squaring an error makes big misses expensive in a way that changes what the safest strategy looks like. If a model trained on squared error is unsure whether an image deserves a three or a nine, guessing six minimises its expected squared error better than committing to either extreme, because being wrong by a lot costs disproportionately more than being wrong by a little. The model learns to hedge toward the centre whenever it is uncertain, and uncertainty is common - most real photos are not obviously extreme cases.
Mean absolute error does not create the same incentive. Because every unit of error costs the same regardless of size, there is no extra penalty for committing to a confident extreme guess that turns out wrong, so models trained this way tend to be more willing to actually use the ends of the scale. The statistician Tilmann Gneiting set out the formal version in a 2009 paper on point forecasts: squared error is minimised by predicting the mean of the plausible outcomes, while piecewise linear losses such as absolute error are minimised by a quantile, the median in the absolute case. He warns that mixing them up "can lead to grossly misguided inferences". A mean over "three or nine" is six; a median can sit at either end if the evidence leans that way. This is a genuine trade-off rather than one option simply being better: a model trained on absolute error is more decisive and also more exposed to occasional large misses, while one trained on squared error is smoother and more cautious.
Why this matters more than it sounds
Two scorers can have identical accuracy on average and still feel completely different to use, purely because of which loss function trained them. One will rarely give you a two or a nine and will cluster heavily around the middle of its range even on photos that plausibly deserve an extreme. The other will use more of the scale and occasionally be more dramatically wrong for it. Neither is lying; each is doing exactly what it was optimised to do.
This sits downstream of the head architecture question - whether the model outputs a continuous value or picks among discrete buckets in the first place, which gets its own treatment here rather than being repeated in full. The loss function is the second, separate decision layered on top of whichever head design was chosen, and it is the one that decides how the model behaves when it is uncertain rather than how it structures its output.
What to take from this
A rating tool that seems to avoid the extremes of its own scale is not necessarily being coy. It may simply have been trained on an objective that makes hedging the mathematically safest move, in which case the clustering is a property of the training objective rather than a judgement about your photo specifically - a pattern that shows up independently in why almost everyone lands on a six or a seven. Rate Cock reports separate axes rather than one aggregate, and looking at the spread across axes is a better way to see whether a submission is genuinely middling or whether the total is smoothing over a real extreme on one dimension.
Loss functions are not the only place hedging comes from - thin label data at the extremes has the same effect through a completely different mechanism, and the two often compound rather than substitute for each other. A human reviewer does not have a loss function to hedge against in this sense at all, which is part of what Rate Penis covers about the difference between a trained scale and a person's direct opinion. None of this touches physical measurement, where Measure My Cock's method is the honest route when the number that matters is in centimetres rather than a rank. Reading what a given score is actually telling you, extremes included, is covered from the practical side at penisrater.com.