Accuracy
The interval around a number is the honest result
Ten repeats give a range; the range is the finding, and a single number quoted without it is an anecdote dressed as a measurement.
Quote one AI score and you have told a story. Quote a range built from several repeats and you have reported a finding. The difference is not about how many decimal places either version has - it is about whether the number you are showing someone accounts for the noise that produced it.
Why a single number misleads by omission
Every score, on its own, looks precise. A 7.2 reads like a measurement to the nearest tenth, and the number itself gives no hint of the spread it came from. But the same input scored repeatedly rarely returns the identical value every time, and a single 7.2 could sit anywhere inside a distribution that mostly ranges from 6.5 to 7.9, or one that barely moves at all - the number alone cannot tell you which, and quoting it without the range implies a precision the process never actually delivered.
What a range gives you that a point does not
A range - five, ten, however many repeats you have taken - tells you how much confidence to put in any single reading from it. A tight range around 7 and a wide one around 7 mean very different things, even though a single draw from either could produce the exact same headline number. How many repeats are worth taking before the range settles is a separate question with its own diminishing returns, but even a handful is enough to turn a bare number into an honest one, since the first few repeats do most of the work of revealing whether the spread is narrow or wide.
How to actually report one
Reporting a range does not require statistical software. Take several repeats of the identical input, note the lowest and the highest, and state both alongside whatever central value you quote - "mostly 6.5 to 7.5, centred around 7" is a complete, honest report in a way that "7.2" on its own is not. If the repeats cluster tightly, say so, because a tight range is itself informative - it tells the reader the single number is close to reliable. If they scatter, say that too, rather than picking the most flattering result from the batch and presenting it as the score.
Where this matters most
This habit matters most exactly where people are least likely to apply it: comparing two things. A single score of 7 beating a single score of 6 looks like a clear result. If both numbers came from distributions that overlap once you account for their spread, the "win" may not be real at all, and only reporting the range on both sides reveals that. Rate Cock breaking a result into separate axes rather than one total makes this easier to practise, since you can watch the spread on each axis rather than a single blended figure that hides where the noise actually sits. The same discipline applies wherever a single reading gets treated as gospel - a tape-measure result quoted without acknowledging measurement error has the identical problem, and comparing tools on a single run each rather than several inherits it too. A human judge's single written opinion is a different kind of artefact entirely, not a sample from a repeatable distribution in the same sense, which is worth remembering before applying this exact method to a person's review rather than a model's score.