How it works

An axis earns its place by being separately observable

A good rubric axis is something a model can be trained to read on its own, and most proposed axes fail that test.

7 min readHow it works

A rubric is a list of things a model is asked to score separately. Splitting one number into several is worth doing in general, because a total throws away which part of a judgement moved - but that only pays off if each item on the list is a thing the model can actually learn to read on its own. This piece is about building that list: what makes a candidate axis real, and what makes it a decoration.

The test an axis has to pass

An axis earns a place on a rubric if three things are true of it.

It has to be observable in the image - something present in the pixels a vision encoder can extract, not a fact about the world that happens to correlate with the image but is not actually visible in it. Weight, for instance, is not directly observable in a photograph the way surface texture is; a model can infer a correlate of it, but the axis would be scoring the correlate, not the thing its name suggests.

It has to be separable from the other axes - varying independently of them across real photographs, rather than moving in lockstep with something else already on the list. If two candidate axes always go up and down together in your data, they are not two axes; they are one axis with two names, and training both is how a rubric ends up double-counting the same signal.

It has to be stable across incidental changes to the photograph - lighting, angle, crop - that have nothing to do with what the axis claims to measure. An axis that swings wildly because the photographer moved three inches is not reading the subject; it is reading the photograph, which is a different thing wearing the axis's name.

Where the three tests come from

None of these three requirements is specific to body-image scoring; they are the same tests any rubric designer runs when turning a fuzzy human judgement into something a model can be trained to predict, whether the subject is essay grading, product photography, or medical imaging. What changes between domains is which candidate axes tend to fail which test, and on this site's subject the failures cluster in predictable places, which is what makes a worked-through list of them useful rather than a generic checklist.

Observability fails most often for anything that describes an inference about the world rather than a property of the pixels themselves. Separability fails most often for axes drafted by a committee that wanted the rubric to feel comprehensive, and added a fifth or sixth item without checking whether it actually moves independently of the first four. Stability fails most often for axes that sound precise but were never operationally defined, so every labeller quietly applies a slightly different standard, and the resulting noise looks like instability in the axis when the real fault is upstream in the instructions.

Where proposed axes usually fail

Most rubric drafts start longer than they end up, because most proposed axes fail one of the three tests above once someone actually checks.

"Confidence" or "presentation," as often proposed, are not observable in the strict sense - a model reading pixels has no channel to a person's internal state, and what actually gets scored under that name is usually posture, framing and lighting, which are legitimate but different axes wearing a borrowed label. Naming the axis honestly, as something like framing quality rather than confidence, keeps the model's actual target in view instead of implying it read something it structurally cannot.

"Overall aesthetic" proposed alongside several specific axes usually fails the separability test, because it tends to correlate heavily with the specific axes combined - it is closer to a preview of the total than a fifth independent input, and including it both as an axis and inside the total is a soft version of double-counting.

Anything defined by comparison to an ideal rather than by a directly checkable property - "proportionality," undefined - fails stability, because without an anchor, labellers each bring their own implicit ideal, disagree with each other more than they would on a directly observable property, and the resulting noisy labels teach the model an unstable, labeller-dependent target. The fix is not dropping the concept; it is operationalising it, which is its own step, worth doing deliberately rather than assuming a name is enough.

Two kinds of legitimate axis

Axes that pass the three tests still split into two useful categories, worth naming separately because they behave differently.

Subject axes describe the thing in the photo: shape, proportion, surface condition. These are what a rubric is nominally for, and they are also the harder half to keep stable, because incidental photography choices leak into them more easily than a first draft expects.

Presentation axes describe the photograph itself: framing, lighting, sharpness. These are usually easier for a model to read reliably, precisely because they are closer to the pixels and further from inference - a model detects blur far more reliably than it infers proportion - which is why a rubric that quietly lets presentation axes stand in for subject axes tends to score the picture more than the picture's subject. Keeping the two kinds distinct, and being honest about which is which, is a distinction worth treating as its own design question once a rubric has more than a couple of axes on each side.

How many axes, and why not more

Every additional axis needs its own reliable labels, which means its own labelling instructions, its own inter-labeller agreement check, and its own volume of examples across the score range. That cost scales roughly linearly with axis count, while the informational return per additional axis shrinks, because the easy, clearly-separable properties tend to get claimed by the first few axes and what is left over correlates more with what already exists.

There is no universal correct count, but there is a practical ceiling: past a handful of axes, on properties like the ones a body-image rubric covers, new candidates increasingly fail the separability test against axes already on the list, and adding them anyway produces the fake-precision problem of many numbers that do not actually vary independently. Working through a specific rubric size in detail is a useful worked example of where that ceiling tends to land in practice, and the general version of the how-many-is-too-many question is worth reading before finalising a draft list.

Anchoring the scale

An axis definition is not complete until the scale itself has fixed points. Labellers - and, indirectly, the model - need to know what a 3 looks like on that specific axis, not just that 3 is lower than 7.

The standard approach is reference examples: a small set of images pinned to specific scale points, shown to every labeller as a common anchor. What those anchors are and how they get chosen shapes the entire downstream scale, because every later score is implicitly a comparison to those fixed points, whether the person or model doing the scoring is aware of that or not.

Revisiting a rubric over time

A rubric is not a one-time decision. As more labelled data accumulates, correlations that were not visible in a small pilot set can show up clearly in a larger one, and an axis that looked separable at launch can turn out to move in lockstep with another once there is enough data to check properly.

The discipline that keeps a rubric honest over time is the same one that built it: periodically re-running the three tests against the growing label set, rather than treating the original design as settled once the first version ships. A rubric that has never been revisited since its first draft is not necessarily wrong, but it has not been checked, and those are different claims - the tests are cheap to re-run and expensive to have skipped if an axis quietly stopped meaning what its name promised.

Naming and disclosure

An axis's public name is a promise about what it measures, and the design work above is what makes that promise honest or not. A rubric that calls an axis "confidence" when it is really reading posture and lighting is making a claim the pixels cannot support; a rubric that calls the same axis "presentation" is making a claim they can. How a rubric built this deliberately for a machine actually compares to the looser rubric a human judge carries around unwritten is its own worthwhile comparison.

Rate Cock publishes six named axes - size, shape and proportions, skin texture, head, overall appeal and erotic impact - and shows the breakdown behind every result rather than only the total, which is the design choice this whole piece has been arguing toward: an axis that cannot survive being shown on its own, next to its name, has not earned a place on the rubric. Measure My Cock's coverage of method approaches a version of the same discipline from the data-collection side, checking that what gets measured is what the method claims to measure before a single label is ever assigned. Reading the resulting breakdown well, once it exists, is a skill on its own, which is what penisrater.com spends its attention on from the user's side rather than the design side, and a human judge sidesteps most of this design problem entirely, since a person can name which axis moved in a sentence without any of the upfront labelling and anchoring work a machine rubric requires before it can do the same.

A well-designed rubric is not a longer list. It is a shorter list where every entry survived being checked against the three tests, and where the total built on top of it is a separate, disclosed decision rather than an assumption.

Read next

Full archive