How it works
Post-hoc text is not the model's reasoning
When a tool shows a paragraph explaining a score, that text is usually generated from the numbers, not the other way round.
Guides on How it works: Every score is a memory of somebody's opinion, An axis earns its place by being separately observable, What an image becomes before it is scored
The paragraph explaining your score was usually written after the number existed, not as a record of how it was reached. In most pipelines a scoring head produces the number without any text, and a separate language model is then asked to write something that sounds like an explanation for it.
The two-step pipeline
A common architecture runs a vision encoder into a scoring head, which outputs a number and nothing else - no words, no reasoning trace, just a value on a scale, the same mechanism described for scoring generally. A second, separate step then hands a language model the score, sometimes the image, and a prompt along the lines of "write a paragraph explaining why this scored a 7." The language model was never involved in producing the 7. Its job is to generate plausible-sounding prose that is consistent with a number it is simply told, after the fact.
Why the text can be fluent and wrong
A language model is very good at generating fluent, specific-sounding justifications for almost any starting premise, because that is close to what the pipeline it sits in is optimising for - readable prose, not verified reasoning. Given a score, it will confidently describe properties that plausibly explain that number, and those properties are not being checked against what actually happened inside the scoring head, because the scoring head's internal computation is not text and was never exposed to the language model in the first place. The explanation can name a factor - "the lighting brought this down slightly" - that sounds exactly like the kind of thing that could move a score, whether or not lighting had anything measurable to do with this particular result. Even a model narrating its own reasoning can mislead: Turpin and colleagues (2023) found chain-of-thought explanations systematically misrepresented what drove the answers, with accuracy dropping by as much as 36% on 13 tasks when models were nudged by a bias their explanations never mentioned.
Why this differs from a genuinely score-generating language model
Some pipelines run the other way: a vision-language model produces its description and reasoning as part of generating the number itself, so the score is downstream of the text rather than the other way around. That is a different architecture with a different set of failure modes, covered as its own subject, and it is worth not conflating the two. The tell is usually the ordering: if the number could plausibly have been produced without any of the surrounding prose ever existing, the prose is very likely explanation-after-the-fact rather than a trace of what actually happened.
What the explanation is still useful for
None of this makes the paragraph worthless. It is a genuine attempt, by a capable model, to produce a plausible account consistent with a real number, and plausible accounts are often approximately right, because the language model has seen a great many real explanations of real scores during its own training and has a reasonable prior about what usually correlates with what. The honest way to read it is as an informed guess about what moved the number, not a log of the computation - useful for orientation, not to be trusted the way a per-axis breakdown can be trusted, since the breakdown is at least the actual numbers the aggregate was built from - reading the numbers over the prose is the same habit penisrater.com recommends to its own users.
A quick way to spot it
One useful check: ask the same tool, on two separate occasions, why an identical or near-identical photo received a similar score. If the wording of the explanation changes substantially each time while the number stays close to constant, that is a strong sign the text is being freshly generated per request rather than drawn from any fixed record of how the score was computed. A dedicated scoring head with a genuinely fixed reasoning trace, if one existed, would have no reason to phrase the same underlying computation differently on repeat; a language model asked to narrate a number has every reason to, since narrating is exactly the task it is performing each time.
What to make of it as a reader
Treat a generated explanation as commentary, not evidence, and treat the number and the chart behind it as the primary record. Rate Cock writes a verdict for every score, and the honest way to use it is exactly this way - a readable gloss on the numbers rather than a transcript of how they were reached. A written explanation from a human judge does not have this gap at all, since a person on ratepenis.com is describing their own actual reasoning as they give it, not narrating a number handed to them afterwards, and measuremycock.com's figures need no explanatory paragraph in the first place, because a tape measurement is its own account of itself.