How it works

Past a point, more axes means more noise

Each added axis needs its own reliable labels; beyond a handful the labels get thin and the axes start repeating each other.

3 min readHow it works

Adding an axis to a rubric looks free. It is a new column in a spreadsheet and a new head on a model. What it actually costs is a fresh set of reliable labels for that specific property, on every training image, and that cost is where the real ceiling on rubric size sits.

The labelling budget is the real constraint

Each axis needs labellers who agree with each other on it, which usually means a written definition, a training pass, and enough labelled examples for a model to learn the pattern rather than memorise the training set. That is a fixed cost per axis, and it does not shrink because you already paid it for the axis next to it. A team that can maintain four well-labelled axes at a given budget can rarely stretch the same budget to nine without letting label quality slide on some of them.

Where the extra axes start repeating

Past a handful, new candidate axes increasingly correlate with ones already in the rubric, because most of the properties a photo can vary along are not independent of each other. An axis added at that point is mostly restating an existing one under a new name, which does not add coverage - it adds a second vote for the same underlying signal, quietly weighting it twice in whatever aggregate the tool reports. That specific failure, two axes moving together and getting counted as if they were separate information, is worth treating as its own subject rather than folding in here.

What this is not about

This is not a claim that correlated axes are a design mistake in themselves - that gets its own treatment. And it is not a checklist for choosing which axes to keep, which is a longer question about starting from what people actually notice and working down to a workable count. This is narrower: the label budget sets a practical ceiling, and it is lower than most first-draft rubrics assume.

What a reasonable stopping point looks like

There is no fixed number that is correct for every rubric, because the ceiling depends on how much labelling budget and reviewer time a team actually has, not on the subject matter. What is consistent is the failure signal: label agreement between reviewers starts dropping on the newest axes first, since those are the ones with the least labelled volume and the least-refined written definition behind them. A rubric that tracks per-axis agreement and stops adding axes when a new one comes in below the others is stopping in the right place, whatever the resulting count turns out to be. Rate Cock settled on six for its own rubric, which is a specific instance of this trade-off worked all the way through, and that walk-through is covered separately. Whether six axes or six thousand ratings, the underlying question multi-axis scoring versus one number asks is the same one this note answers from the labelling side rather than the presentation side.

Reliability at the retake level - whether a single axis holds still on repeat runs - is a related but separate question penisrater.com covers from the user side, and it is worth reading if what you are actually asking is "can I trust this number," not "how was the rubric sized." Human review sidesteps the labelling-budget problem entirely, since ratepenis.com covers a judge writing a paragraph rather than filling in a fixed set of axis boxes, and a size-conscious rater can grow a rubric the way measuremycock.com treats measurement categories, adding one only when the data supports it.

Read next

Full archive