How it works

Reasoning text before the number: help or theatre

Asking a model to reason before scoring can improve consistency, but the reasoning is not guaranteed to be what produced the number.

By 4 min readHow it works

Guides on How it works: Every score is a memory of somebody's opinion, An axis earns its place by being separately observable, What an image becomes before it is scored

Sometimes: asking a model to reason before scoring tends to make its outputs more consistent, but nothing guarantees the written reasoning is what produced the number. The first half is well replicated and now routine in scoring prompts. The second half is still open, and nobody outside the lab that built the model can be fully sure.

What chain-of-thought prompting does

The technique is simple to describe: instead of asking directly for a number, the prompt asks the model to describe what it observes, weigh it against the rubric, and only then state a score. Wei et al. (Google, 2022) showed that this kind of intermediate reasoning "improves performance on a range of arithmetic, commonsense, and symbolic reasoning tasks" compared with asking for a direct answer. The mechanism proposed is that generating intermediate text gives the model more computation to work with before committing to an answer, since each token it produces can condition the next one.

Applied to scoring, the hope is straightforward: a model that writes "the proportions are close to typical, the surface condition is clean, overall this reads as slightly above average" before saying "7" should be less erratic than one that jumps straight to "7", because the written reasoning constrains what number can plausibly follow.

Where the evidence is thinner than the technique's popularity suggests

The direction of the effect - reasoning before answering tends to help - is reasonably well supported for problems with a checkable right answer, like arithmetic or logic puzzles. Whether it helps to the same degree for a subjective visual judgement, where there is no single correct number to converge on, is far less studied, and claiming a specific size of improvement for scoring tasks specifically would be stating a number nobody has actually measured. What can be said honestly is directional: prompting for reasoning tends to reduce the number of wildly inconsistent outputs, without any guarantee about how much, or about any particular tool.

The faithfulness problem

The harder issue is not whether reasoning helps but whether the reasoning shown is the reasoning that happened. A language model generates its explanation the same way it generates everything else: one token at a time, predicting what a plausible-sounding justification would say next. There is no architectural guarantee that this text traces the actual computation behind the eventual number, rather than being a fluent post-hoc justification for a number the model was already leaning toward for other reasons. Researchers call this the faithfulness question, and it remains genuinely open. Turpin et al. (2023) nudged models toward wrong answers with biasing features in the prompt, saw accuracy drop "by as much as 36%" across 13 BIG-Bench Hard tasks, and found the explanations systematically failed to mention the bias that had actually moved the answer.

This matters for scoring because a written explanation is persuasive. A paragraph that walks through proportion, texture and symmetry before landing on a number feels like evidence the process was careful, and it may be, but the text itself is not proof of that - it is one more thing the model generated, using the same mechanism it used for the number.

Where this fits with the rest of the pipeline

This is distinct from the case where a tool generates an explanation purely as a presentation layer after the number already exists, without any pretence that the text influenced the score - a separate and more common pattern worth knowing apart from this one. It is also worth distinguishing from grounding, where a model ties a specific claim to a specific region of the image rather than writing free-form prose - that approach gives you something closer to checkable evidence than reasoning text does, because a region can be verified against the photo directly.

What a careful tool does with this uncertainty

The reasonable response is not to distrust every reasoning-based scorer, but to treat the written explanation as a hint rather than a proof. Rate Cock shows its per-axis breakdown alongside any written verdict specifically so the numbers can be checked against the prose, rather than asking a user to take the paragraph on faith. A trained human reviewer does not have this problem in the same form - a person can be asked a direct follow-up about their own reasoning and give a real answer, rather than generating a second plausible-sounding paragraph. If what you actually want is a stable number rather than a persuasive paragraph, a measured figure sidesteps the whole question, since there is no reasoning step to distrust when the number comes from a tape rather than an inference. Penis Rater's tool coverage looks at which scoring approaches are worth trusting from the outside, which is the practical mirror of the mechanism covered here.

Read next

Full archive