Accuracy

One score is an anecdote; the spread is the data

The number of repeats you need depends on how wide the spread is, and you cannot know the spread until you have taken several.

4 min readAccuracy

One AI score tells you almost nothing about how stable that score is. It is a single draw from a distribution whose width you have not seen yet. The honest answer to "how many repeats do I need" is that it depends on how wide that distribution turns out to be, which is a slightly circular problem, and it is worth working through rather than skipping.

Why one number cannot answer its own question

Run a photo through a scorer once and get a 7. That 7 could be sitting at the dead centre of a distribution that never moves more than a tenth of a point either way, or it could be one point drawn from a distribution that regularly swings two or three points. Test-retest reliability is the name for how tight that swing is, and different tools, and different axes within the same tool, are reliable to different degrees. You cannot read reliability off a single result any more than you can read the shape of a die by rolling it once.

The standard error intuition

Statisticians have a formula for how much confidence grows as you add repeats, and the formula itself is less useful here than its shape. The core idea: your uncertainty about the "true" average score shrinks with the square root of how many times you have sampled it, not in a straight line. Going from one repeat to four cuts your uncertainty roughly in half. Going from four to sixteen cuts it in half again. That square-root relationship is why the first few repeats buy you the most, and every additional repeat after that buys progressively less.

This is not a figure to memorise, and no fixed count applies to every tool or every photo - the honest version of this advice avoids naming a specific number of repeats as universal, because the right count depends on a spread that is different for every scorer and every axis. What the shape tells you is directional: a handful of repeats moves you from "no idea" to "rough sense" fastest, and each repeat after that earns you diminishing certainty for the same effort.

Diminishing returns, and when to stop

There is a point past which more repeats stop being worth the time. If your first three or four scores already sit close together, you have probably found a tight distribution, and further repeats will mostly confirm what you already know. If they are scattered widely, more repeats help you see the actual range, but no amount of repeating turns a genuinely noisy instrument into a precise one - it only lets you describe the noise more accurately. Recognising which situation you are in - narrow and settled, or wide and still moving - is more useful than chasing a specific repeat count.

There is also a ceiling worth knowing about separately: near the top and bottom of any bounded scale, the model runs out of room to distinguish between inputs regardless of how many times you ask it, and no number of repeats fixes a compressed scale.

What to do with a handful of scores

Once you have several scores from the same unchanged input, the useful output is not their average - it is the range. Reporting the spread rather than the single number is the difference between an honest result and a cherry-picked one, and it costs nothing beyond running the tool a few extra times. If you are also changing something between attempts - the lighting, the angle, the crop - that is a different exercise entirely, closer to a controlled comparison than to a reliability check, and the two should not be mixed in the same set of runs. Spacing repeats out rather than firing them back to back also matters more than it looks, since order and caching effects across a repeat run can quietly narrow the spread you are trying to measure.

The same logic shows up wherever repeat measurement matters. Measure My Cock treats a single tape reading the same way - one pull of the tape is a data point, not a conclusion, and their method write-up covers how many pulls a careful measurement actually takes. Comparing scores across different apps runs into the same problem from another angle, since a single result from each tool tells you less than a handful from the one you actually plan to use. A commissioned human review sidesteps the repeat-count question differently: one judge gives you one considered opinion rather than a distribution to sample from, which is a different trade-off, not a shortcut to the same answer. Whichever tool you use, including Rate Cock, the underlying advice does not change - a score you have only seen once is a starting point, not a result.

Read next

Full archive