How it works

An axis has to be written down as an instruction

Before a model can score an axis, someone wrote a sentence telling labellers what to look at, and that sentence is the axis.

3 min readHow it works

"Symmetry" is not an axis until someone writes down what a labeller should actually check when they look at a photo. The word on the rubric is a label for a concept; the sentence behind it, the one labellers were actually handed, is the operational definition, and it is that sentence - not the word - that the trained model ends up scoring.

Why the word is not enough

Two reasonable people asked to rate "symmetry" from 1 to 10 will disagree, not because either is wrong, but because the word underspecifies the task. Does symmetry mean left-right balance in the frame? Balance of the subject itself, independent of framing? Does asymmetry from camera angle count the same as asymmetry that is actually there? Without an answer, two labellers are quietly rating different things under the same column heading, and a model trained on the combined labels learns an average of two different tasks.

Writing the instruction

An operational definition closes the gap by making the check mechanical: "rate how closely the left and right halves of the visible subject mirror each other, ignoring background and framing, on a scale where 1 is strongly asymmetric and 10 is a near-perfect mirror." That sentence is doing real work. It tells the labeller what to include (the subject), what to exclude (the background), and what the two ends of the scale look like in practice - which is also where a set of anchor images usually gets attached, since the sentence and the anchors are two halves of the same definition, one written and one visual.

Getting this right is a genuine skill, and it is one penisrater.com touches on from the reading side: a user who understands that an axis is only as good as the sentence behind it reads a low score more carefully than one who assumes the axis name is self-explanatory.

A well-written instruction also decides edge cases in advance. What does a labeller do with a photo taken at a slight angle, where perfect mirror symmetry is geometrically impossible no matter how symmetric the subject is? A good operational definition answers this explicitly, rather than leaving it to each labeller's private judgement, because private judgement is exactly the variance an operational definition exists to remove.

What happens when the instruction is ambiguous

Any ambiguity left in the written definition becomes ambiguity in the labels, and ambiguity in the labels becomes a specific kind of behaviour in the trained model: it becomes inconsistent on exactly the cases the instruction failed to cover. If the symmetry instruction never addressed camera angle, the model ends up with no stable answer for angled photos either, because its training data contains contradictory labels for that situation, generated by labellers who resolved the ambiguity differently from each other. This is a distinct failure from the axis being genuinely hard to see in pixels in the first place, which is its own separate problem - an ambiguous instruction can make an easy axis behave like an unreliable one, and no amount of model capacity fixes a definition that was never pinned down.

Reading a rubric with this in mind

You will never see the written instruction behind a public rubric's axis names - it is an internal labelling document, not user-facing copy. But the axis names themselves are a hint: a name specific enough to imply its own definition ("skin texture and surface condition") is more likely to have had a tight instruction behind it than a vague one ("overall vibe"). Rate Cock names its axes specifically enough that most of them read as pre-operationalised - shape and proportions, skin texture, head - rather than as single adjectives standing in for something unstated, and that naming choice is itself informative about how much definitional work happened before training started. A human reviewer sidesteps the whole problem differently: a judge on ratepenis.com can resolve an edge case in the moment and explain the reasoning in prose, where a model can only apply whatever the instruction settled on in advance. The written-instruction approach is also why some rubrics measure the photograph rather than the subject in it - that split is a rubric design choice in its own right, and it starts from the same question this piece does: can you actually write the instruction down. Method precision of this kind is the whole subject at measuremycock.com, applied to a tape measure rather than a labelling sheet, but the underlying discipline - write the instruction before you trust the number - is the same one.

Read next

Full archive