How it works

The total is a policy, not a measurement

Combining several axis scores into one requires choosing weights, and every choice encodes an opinion about what matters.

4 min readHow it works

Once a rubric produces several axis scores, something has to turn them into the one number most people actually look at. That step is not a calculation in the way the axis scores themselves are; it is a design decision, made by someone, about how much each axis is allowed to move the total.

The plain mean

The simplest aggregation is an unweighted average of every axis. It treats every axis as equally important by construction, which is a real claim, not a neutral default - it says the rubric's designers believe each axis deserves identical influence, and that claim is worth stating rather than assuming.

A plain mean also has a specific weakness: it lets a very high score on one axis compensate fully for a very low score on another, so two submissions that are, on their component parts, quite different can land at the same total. Whether that compensation is desirable depends entirely on what the total is meant to represent.

Weighted averages

A weighted average assigns each axis a different multiplier before averaging, reflecting a judgement that some axes matter more to the overall impression than others. This is a defensible thing to do, but it moves the actual editorial decision from "the axes" into "the weights," and a tool that will not disclose its weights has hidden the one place where its opinion about what matters actually lives.

Weights are also where correlated axes do the most quiet damage: if two axes are really carrying one signal, giving each its own weight in the total effectively double-weights that shared signal without anyone choosing to.

Geometric mean

A geometric mean - the nth root of the product of n scores, rather than their sum divided by n - behaves differently from an arithmetic mean at the low end of a scale. Because it multiplies rather than adds, a single very low axis score pulls the geometric mean down harder than it would pull an arithmetic mean, which makes compensation much less available: a submission cannot cruise past one badly-scoring axis on the strength of the others the way it can under a plain average.

This property makes a geometric mean attractive whenever the designer wants every axis to function as something closer to a floor, rather than one input among several that can offset each other.

Min-based aggregation

The most extreme version of the same idea takes the lowest axis score as the total, ignoring the rest entirely. This removes compensation completely - nothing can be made up for - which is the right behaviour if the rubric's designer believes the weakest dimension genuinely should cap the whole judgement, and the wrong behaviour if strong performance elsewhere is supposed to count for something, which is most of the time a rubric has more than one axis in the first place.

What each choice rewards

A plain mean rewards balance and lets strengths offset weaknesses, which suits a rubric where the axes are genuinely commensurable and no single one is meant to be a gatekeeper. A weighted mean rewards whatever axes were given the higher multipliers, and is honest only when the weights are disclosed. A geometric mean and min-based aggregation both reward consistency across axes over standout strength on any one, in increasing degrees, which suits a rubric where a serious weakness on one axis should be visible in the headline number rather than buried in an average.

None of these is the correct choice in the abstract. Each is a policy about what the finished total is supposed to communicate, and the choice should follow from what the rubric's designers decided that total is for, made before the arithmetic, not after it.

What this means reading a result

A total without a stated aggregation method is asking to be trusted on a point the tool could simply state. Knowing whether a headline number is a mean, a weighted mean, or something closer to a floor changes what a single strong or weak axis should be expected to do to it, which is worth checking before treating the total as more informative than the axes it was built from - what an axis needs to earn a place on the rubric in the first place is the design question upstream of this one, and worth reading alongside it.

Rate Cock reports a total alongside its six named axes, and the aggregation choice behind any tool's headline figure is exactly the kind of methodological detail measuremycock.com treats as worth stating plainly rather than leaving implicit. penisrater.com's guidance on reading a score covers the user side of this same gap - what a total does and does not tell you - and a human reviewer sidesteps the arithmetic question entirely, since a person forming an overall impression is not running any of these formulas underneath it, for better and for worse.

Read next

Full archive