How it works

Repeated sampling as a stability trick

Running a language-model scorer several times and aggregating tames its randomness at the cost of compute, and hides it from the user.

By 3 min readHow it works

Guides on How it works: Every score is a memory of somebody's opinion, An axis earns its place by being separately observable, What an image becomes before it is scored

Self-consistency means calling the same language-model scorer several times and reporting the median or mean, which tames sampling randomness without changing the model. It is one of the cheaper, more effective ways a tool can make a noisy scorer look stable, and it hides the spread from the user.

How it works

The same photo and the same prompt are sent to the model several times - five and ten runs are typical - and each run returns its own score, which will vary somewhat if sampling is not fully deterministic. The tool then takes the median or mean of the runs rather than showing any single one. A median is more common than a mean here because it is less sensitive to one wildly off run pulling the reported number away from where most of the runs actually landed.

The technique was popularised by Wang and colleagues at Google (2022, published at ICLR 2023), who sampled "a diverse set of reasoning paths instead of only taking the greedy one" and took the most common final answer. On the GSM8K maths benchmark that lifted accuracy by 17.9 percentage points over trusting a single path. Applied to a numeric score rather than a discrete answer, the same instinct holds: several independent draws from the same underlying distribution, combined, are a better estimate of the centre of that distribution than any one draw.

What it costs, and what it hides

The cost is straightforward - five times the model calls means roughly five times the inference cost and, unless the calls are parallelised, five times the wait before a result appears. That tradeoff is why it is not universal; a tool serving results instantly to a large number of free users has a real incentive to skip it.

The less obvious cost is what the technique conceals from the user. A single number presented after five runs looks exactly as certain as a single number presented after one run, even though the underlying model was demonstrably capable of landing on five different figures for the same photo. Averaging is real stabilisation, not a trick, but it also removes the one piece of information - the spread across runs - that would have told a user how much to trust the figure in the first place. Reading the spread rather than the single score is the more honest version of what self-consistency does internally, made visible instead of hidden.

Where this differs from an ensemble

This is not the same as combining several different models, which draws on architectural diversity rather than repeated sampling from one model - a distinct technique with its own tradeoffs, covered separately. Self-consistency uses one model, asked the same question more than once, which only helps with the randomness introduced by sampling, not with any bias the single model shares across every run.

Rate Cock aggregates multiple internal passes before returning a result for exactly this reason, which is part of why repeat submissions of the same photo tend to land close together rather than swinging widely. Measure My Cock's data coverage makes the equivalent point for a physical method: a single careful measurement does not need five repeats to be trustworthy, because there is no sampling step to average away. Penis Rater's scores coverage discusses repeat testing from the user's side. A panel of human reviewers gets a version of the same benefit for a different reason - several independent opinions, not several draws from one model - which is worth keeping distinct even though both end in an averaged figure.

Read next

Full archive