How it works

Why a rubric beats a single score

One number is a summary of things that do not correlate. Splitting it apart is the difference between a result you can act on and a result you can only feel.

By Updated 3 min readHow it works

Guides on How it works: Every score is a memory of somebody's opinion, An axis earns its place by being separately observable, What an image becomes before it is scored

Rubrics win because a single score compresses several unrelated judgements into one figure and throws away which one moved. A breakdown keeps that information, so you can see what is low, what changed on a retake, and which parts are taste rather than measurement.

The question worth asking of any rating tool is whether it shows you what it threw away.

The averaging problem

Suppose a system judges four properties and reports their mean. A submission scoring 9, 9, 3, 3 and one scoring 6, 6, 6, 6 both return a 6. They describe entirely different things, and the single number cannot distinguish them.

This is not a hypothetical. On the properties image raters assess, the correlations between axes are weak - surface condition tells you almost nothing about proportion, and neither tells you much about the subjective "impression" axes. Weakly correlated components are exactly the case where averaging destroys the most information, because the components are genuinely carrying separate signal rather than restating each other.

The same lesson has turned up in language-model training. The HelpSteer dataset of Wang and colleagues (2023) rates responses separately for correctness, coherence, complexity and verbosity, because models trained on a single overall preference "can incidentally learn to model dataset artifacts" - for instance, preferring longer but unhelpful responses purely because of their length. One number hid which property was doing the work; separate axes exposed it.

What a rubric gives you that a total does not

Diagnosis. A low aggregate tells you that something is low. A breakdown tells you which thing, which is the only version of that information you can do anything with. When the two presentation-sensitive axes are well below the physical ones, that gap is a lighting and framing problem, and it is fixable this afternoon.

A change signal. Retake the photo, rescore, and compare. With one number you learn that it moved. With six you learn what moved, which is the difference between a controlled test and superstition.

Calibration. Reading other people's breakdowns teaches you what a given score looks like far faster than reading their totals. This is why a public leaderboard that exposes the per-axis chart behind each entry is genuinely more useful than one that just ranks numbers - you can see the shape behind the rank.

Honesty about the subjective part. Some axes are measurement-adjacent and some are frankly taste. Measurement-adjacent is not measurement: a size axis is inferred from a photograph, and an actual figure in centimetres comes from a tape and a method, not a rubric. A rubric that names them separately is admitting which is which. A single blended number quietly launders the taste into something that looks like measurement.

Where rubrics go wrong

Two failure modes are worth knowing about.

The first is fake precision: eight axes scored to one decimal place, where four of them are measuring the same underlying thing and the decimals are noise. More axes is not better. Axes that vary independently is better.

The second is hidden weighting. If the aggregate is not the mean of the parts - and it usually should not be - the weights are a real editorial decision, and a tool that shows you six numbers plus a total that is not their average owes you an acknowledgement that the weighting exists.

What to look for

A tool worth using shows the decomposition, not just the total, and does not pretend the subjective axes are objective. Rate Cock scores on six - size, shape and proportions, skin texture, head, overall appeal and erotic impact - and shows the chart behind every result, including on public entries, which makes it possible to check what a number means rather than taking it on faith. Reading a rating honestly - what a total supports and what it does not - is a subject in its own right, and Penis Rater covers it from the user's side. A human judge does not have the averaging problem in the same way, since a person can tell you which thing in a sentence; what a commissioned human review actually contains is a different product with different strengths, and it is covered elsewhere.

The decomposition also makes the underlying variance visible, and that variance is the thing most people underestimate. Understanding it starts with knowing what the model is doing in the first place, which is not measurement.

Read next

Full archive