Accuracy
Averaging models reduces variance; it does not add truth
An ensemble narrows the spread of scores, which looks like accuracy, but if the models share a bias the average keeps it.
Guides on Accuracy: There is no true score to be accurate against, Change one thing, hold the rest, repeat, Exposure that suits one skin tone hides another
Running a photo through several models and averaging the results is a common way to produce a final score, and it genuinely does make the number more stable. Stable is not the same as accurate, and the gap between those two claims is where an ensemble's real limits sit.
What averaging actually fixes
Each individual model carries its own noise - the ordinary variance from sampling, preprocessing, and training quirks that shows up when the same input is scored twice by the same model. If that noise is roughly independent across models - one model's error on a given photo is not the same as another's - averaging several of them cancels a good portion of it out, the same way averaging several noisy measurements of anything tends to land closer to the true value than any single one. This is a real, mechanical effect, and it is why an ensemble's output typically wobbles less on repeat testing than any one of its component models does alone. How this differs from combining models operationally - the engineering side, of routing and weighting - is a separate question from the statistical one this post is about.
What averaging does not fix
Variance cancels when the errors are independent. Bias does not, and bias is the part that matters more. The textbook condition is blunt: in a 2000 review of ensemble methods, Thomas Dietterich states that an ensemble beats its individual members only if they are "accurate and diverse", where diverse means they make different errors on new data. If every model in an ensemble was trained on a similar kind of data, using similar labelling instructions, by teams working from a similar cultural sense of what a high or low score looks like, they will tend to share the same systematic errors rather than make different, cancelling ones. Averaging five models that all lean the same direction on the same kind of input does not correct that lean - it just gives you a smoother version of the same lean, with a narrower spread around it, which can look more convincing precisely because it wobbles less.
A useful way to hold both facts at once: an ensemble makes a reliability claim more strongly than any single model can, and it makes no additional validity claim at all. That distinction between the two is the whole story here - smoother is a reliability property, and correct is a validity one, and an ensemble only ever earns you the first.
A worked comparison
Picture three models scoring the same photo, each with its own independent noise, returning 6.4, 6.9 and 7.1. The average, 6.8, sits closer to whatever the "true" underlying value is than any one of the three readings alone, purely because the errors partly cancel - this is the genuine benefit, and it is real. Now picture the same three models, but all three were fine-tuned from the same backbone on overlapping training data, and all three share a tendency to score a particular lighting condition half a point high. They return 7.3, 7.8 and 7.6. The average is 7.57, tighter than any individual reading and confidently, consistently wrong by roughly the same half point every one of them was wrong by on its own. Nothing about the second scenario looks different from the outside - a user sees three numbers close together and a smooth average, and has no way to tell from that alone whether the tightness came from independence or from shared error.
Why this matters for reading a "consensus" score
A score presented as the output of several models can read as more trustworthy simply because it sounds like independent verification, in the way that several human witnesses agreeing feels more convincing than one. That intuition only holds if the models are genuinely independent in the way that matters - trained on different data, by different teams, with different assumptions baked in. Several models built on the same backbone, fine-tuned from the same base weights, or labelled by an overlapping set of people are not independent in that sense, however many of them there are, and treating their agreement as strong evidence is the same mistake as trusting three copies of the same rumour because you heard it three times.
What to actually take from an ensemble score
A tighter distribution from an ensemble is worth having - it means less noise to sort through when you are reading the range around a result rather than a single point. It is not, on its own, evidence that the underlying judgement is correct, and nothing about the number itself will tell you which situation you are in. Rate Cock reports per-axis scores rather than a single averaged consensus figure, which at least lets you see disagreement between axes rather than having it smoothed away before you see it. A panel of human judges runs into the identical trade-off from the other direction - several people scoring the same submission reduces individual noise the same way, and shared cultural assumptions among the panel survive the averaging the same way a shared model bias does. Comparing how different tools handle this is its own useful exercise, and the physical-measurement world sidesteps the whole question, since a tape measure has no ensemble to average in the first place.