How it works

The score may be a committee

Many tools average several models or route between them, which smooths noise and blurs responsibility for any single number.

By 3 min readHow it works

Guides on How it works: Every score is a memory of somebody's opinion, An axis earns its place by being separately observable, What an image becomes before it is scored

A single score does not necessarily come from a single model. It is common for a scoring service to run more than one model on the same photo and combine the results before anything reaches you.

The combination method matters, and it is almost never disclosed.

Averaging

The simplest ensemble runs several models independently and takes the mean of their outputs. If the models were trained on different subsets of data, or with different architectures, their individual errors tend not to line up, and averaging cancels some of that noise out. This is the same logic behind a panel of human raters: no single opinion is trusted alone, and the aggregate is steadier than any one member of it. Disagreement between members is useful too: in the deep-ensembles work of Lakshminarayanan, Pritzel and Blundell (2017), training several networks independently let the system "express higher uncertainty on out-of-distribution examples," a signal a single returned number throws away. The catch is that averaging only helps with errors the models don't share. If every model in the ensemble was trained on a similar dataset with a similar gap - a body type it rarely saw, a lighting condition it wasn't shown - the average inherits the gap rather than smoothing it away, because there was nothing independent to cancel.

Stacking

A more deliberate approach trains a second, small model whose job is to take the outputs of the first-layer models as its input and learn how to weight them. Rather than a flat mean, the combining step is itself learned, which lets the system give more weight to whichever base model tends to be right in a given situation. Stacking can outperform simple averaging, and it also makes the system harder to reason about from outside, since the effective weighting is whatever the stacking model learned, not a number anyone chose.

Specialist models per axis

A rubric with several axes does not require one model doing everything. A common design routes to a different model, or a different head on a shared backbone, per axis - one tuned on proportion, another on surface texture, another on the more subjective impression axes. This is closer to one head per axis than a fully shared model, and it lets each specialist be trained and updated independently, at the cost of running several passes for one submission. Some pipelines also route conditionally: a fast, cheap model handles most uploads, and a harder case gets escalated to a slower, larger one, which is an engineering decision about cost rather than a decision about quality. Whichever design a tool picks, the choice is invisible from outside - a single returned number looks identical whether one model produced it or five did, and there is usually no field in the interface that says which.

What this changes about the number

An ensemble score is the output of a small internal negotiation, not a single opinion. That has a practical consequence: when a score seems to shift for no visible reason, the shift might come from a change in how the ensemble is weighted or routed rather than from anything about the photo, which is one more entry on the list of reasons the same submission can land differently across attempts. It also means responsibility for any one number is genuinely distributed - there is no single model you can point at and ask why it said what it said, because the answer is however many models contributed and however they were combined.

Whether that combination makes the result more trustworthy is a separate question from how it is built, and it is one worth asking about any tool you compare scores across rather than assuming. The mechanics here are about assembly, not about whether more models means a steadier answer over repeated tries - that stability question is its own subject, and it does not follow automatically from the fact that several models were used. None of this changes what an ensemble can infer about physical size, which stays capped by what any photograph can supply regardless of how many models look at it. Rate Cock reports separate axis scores rather than a single aggregate, and a multi-model design behind an axis score is a different layer from the multi-axis presentation itself - the two ideas are easy to conflate and worth keeping apart. A human panel assembled for a commissioned review faces the same aggregation question in a more legible form, since a person can at least tell you which of their own impressions carried the most weight.

Read next

Full archive