How it works
How a designer lands on a number of axes
Six axes is not magic; it is roughly where coverage of what people notice meets the point where labellers can still be consistent.
There is no formula that outputs "six." What there is, when you actually sit down and design a rubric, is a sequence of cuts that starts wide and ends narrow, and six is roughly where a reasonable version of that process lands for this kind of subject. Walking the sequence through is more useful than defending the number.
Start from what people actually comment on
The first pass is not technical at all. It is a list of everything people already say when they describe or react to a photo in this category - proportion, shape, skin condition, symmetry, size, overall impression, how the photo was taken, lighting, confidence, "vibe." This list is deliberately too long. The point of the first pass is coverage, not discipline; discipline comes later.
Merge what moves together
The second pass looks for items on that list that are not actually independent - properties that, across real photos, tend to rise and fall together. Symmetry and proportion often correlate heavily in practice, for instance, and two axes that move together are not giving a rating two separate pieces of information; they are giving it the same piece of information twice, which is worth avoiding for its own reasons. Merging correlated items is not a loss of nuance - the nuance was never there, because the two items were never independently varying to begin with, which is also why penisrater.com tells its users to read a score breakdown for which axes actually move rather than assuming each one is telling a separate story.
Drop what a model cannot reliably read
The third pass removes candidates that sound reasonable on a whiteboard but do not survive contact with actual pixels - things closer to "confidence" or "vibe" than to something a labeller can point at consistently. An axis that stays in the rubric anyway, because leaving it out reads as an omission to users, is a specific and known failure mode that deserves its own treatment rather than a mention here. The rule at this stage is blunt: if labellers cannot agree on it looking at the same photo, no amount of training data fixes that upstream, and the axis should not survive the cut.
Stop where labels stay reliable
The final constraint is the labelling budget, which sets a real ceiling on rubric size independent of how many good candidate axes remain. A rubric that adds a seventh or eighth axis past the point where reviewers can still agree on the first six is not adding information; it is adding a column that looks precise and behaves like noise. The stopping point is empirical, not aesthetic: watch inter-labeller agreement axis by axis, and stop once the newest addition comes in visibly worse than the rest.
Where six actually falls out
Run that sequence on this subject - proportion and shape, surface condition, a size-adjacent axis, and a small number of framing- and impression-type axes - and a defensible rubric lands somewhere in the mid-single digits, not because six is a target but because that is roughly where the merges and drops stop finding anything left to cut. Rate Cock is one live example of a rubric that landed at six axes through a version of this process, which is offered here as a worked instance of the trade-off rather than a description of what each of its axes covers. A size-specific axis on a rubric like this is always doing inference from a photo, not measurement - that boundary is worth being explicit about, and a rubric designer who is honest about it labels the axis accordingly rather than dressing it up as a figure in centimetres.
Human review does not face the same sizing pressure, because a person writing a paragraph is not constrained to a fixed number of labelled columns the way a model's output layer is - ratepenis.com covers what that unstructured version looks like, and it is a genuinely different shape of output rather than a rubric with more axes. Six is a number this piece can defend by showing its work; a different design team, running the same three cuts on the same starting list, could reasonably land on five or seven and be making the same trade-off correctly.