How it works

Token frequency leaks into the number

A language model picking a number is picking a token, and some numbers are simply more common in text, which shows up in the scores.

By 3 min readHow it works

Guides on How it works: Every score is a memory of somebody's opinion, An axis earns its place by being separately observable, What an image becomes before it is scored

Language-model scores cluster on 7 and on round numbers because the model is choosing a token, not computing a result. Numbers that appear often in ordinary text, and 7 in particular, are more probable completions whatever the image in front of the model shows.

Numbers are tokens like any other word

A model trained on huge volumes of internet text has seen "7" and "10" and "seven out of ten" an enormous number of times, in contexts ranging from ratings and reviews to casual conversation, and it has seen numbers like "37" or "6.4" comparatively rarely in the same kind of context. That imbalance shapes the model's underlying probability distribution over which token comes next, independent of anything about the specific image in front of it. Round numbers, and numbers that commonly appear as ratings in the text a model was trained on, are simply more probable completions in a way that has nothing to do with the visual evidence and everything to do with how language works.

This is separate from the well-documented tendency of ratings in general - human and machine - to bunch toward the middle of a scale, which has its own causes around hedging and cautious labelling and deserves its own treatment. The mechanism here is narrower and more specific to language models: even holding the underlying "true" judgement fixed, the token-selection step itself favours numbers that are common in text, on top of whatever tendency toward the middle already exists.

Why 7 specifically

The pull toward 7 has now been measured rather than just noticed. In a 2025 study of six language models asked for a random number, Javier Coronado-Blázquez ran 75,600 calls and found that in the 1-10 range "7 is the preferred value by far for every single model". Three of the models, GPT-4o-mini, Phi-4 and Gemini 2.0, chose 7 in roughly 80% of cases. That test asked for a random pick, not a rating, so it does not tell you how much a scoring prompt is skewed. It does show the habit lives in the models themselves, and the author attributes it to the models reproducing human number preferences from their training text. This tracks with "seven out of ten" being a familiar, well-worn phrase in the review and survey text these models learn from, in a way "six point three out of ten" is not.

What this means for reading a score

A single language-model-generated score sitting on a suspiciously round number is not necessarily evidence of anything wrong with the photo or the judgement - it may just be where the token distribution wanted to land. This is one more reason a single number from a language-model scorer is weaker evidence than several repeats aggregated together, and weaker again than a breakdown across separate axes, since averaging and decomposition both dilute a single token-level habit that a lone score cannot dilute on its own.

Rate Cock reports its axes on the same 1-10 scale users expect, and treats the underlying quirk as a known property of language-model output rather than something to be corrected invisibly. A physical measurement carries no such bias, because a figure in centimetres is read off a tape, not generated token by token from a learned distribution over text. Penis Rater's scores coverage is a reasonable place to read about interpreting a single number honestly. A human rater has round-number habits of their own for entirely different, more familiar reasons - social convention and rounding under time pressure rather than token frequency - and the two are worth telling apart even though the visible symptom looks similar.

Read next

Full archive