How it works

What changes when the rater is a model

Rubrics written for human judges rely on context and taste; a model rubric has to be reduced to what shows up in pixels.

4 min readHow it works

Hand a human-review rubric to a model team expecting it to transfer directly, and it will not survive contact with training. A rubric written for a person leans on things a person brings to the task automatically: context, discretion, the ability to notice something the rubric never anticipated. A rubric for a model has none of that available to lean on, and has to be rewritten around what is actually recoverable from pixels.

What a human rubric gets for free

A human reviewer reading an instruction like "assess overall presentation" fills in an enormous amount of unstated context without noticing they are doing it. They know what a typical submission looks like, they can tell a bad photo from a bad subject without being told how, and they can flag something the rubric's authors never thought to mention - an edge case the instruction is silent on gets a sensible judgement anyway, because the person applying it has general understanding to fall back on. None of that transfers to a model. A model applies exactly the pattern it learned from labelled examples and nothing else; there is no general understanding sitting behind it to catch what the rubric failed to specify.

What has to be added for a model

Every piece of context a human rubric assumes has to be made explicit for a model rubric, or the axis will not train reliably. This is the same discipline operational definitions require - a written instruction specific enough that a labeller, and eventually a model, can apply it the same way every time without needing outside judgement. A human rubric can say "assess whether this looks natural" and trust the reviewer to know what that means; a model rubric has to specify what "natural" is checking for, because there is no discretion downstream of the instruction to interpret it sensibly.

Discretion vs granularity

The trade runs in both directions. A human reviewer can catch something genuinely unusual and respond to it in prose, which no fixed rubric axis anticipated - that flexibility is a real strength, not a workaround. A model, in exchange for losing that flexibility, offers granularity and repeatability a human rubric structurally cannot: the same input produces close to the same output every time, which is not true of a person applying even a well-written rubric across a long session. Neither trade is free, and a rubric designer choosing between the two is choosing which failure mode they would rather have - a model missing something it was never told to look for, or a person drifting from their own earlier judgement.

A worked example

Take an axis both kinds of rubric might call "presentation." A human rubric can leave the word alone, trusting the reviewer to weigh lighting, angle and framing together and reach a sensible overall impression, adjusting for whatever is unusual about a given submission without anyone writing that adjustment down in advance. A model rubric cannot leave the word alone. It has to decide, in writing, whether presentation means lighting, or framing, or both, whether they are weighted equally, and what happens when one is good and the other is poor - because whatever the instruction does not specify, the labels will specify inconsistently instead, and the model will learn the inconsistency as if it were signal.

Where the two approaches actually meet

The gap narrows on presentation-type axes - lighting, framing, focus - because those are close enough to raw pixel properties that a model rubric and a human rubric end up checking almost the same thing. It widens on subject-type axes, where a human's context genuinely adds information a model has no channel for, and it widens further on axes that are closer to taste than to anything mechanical, where a model rubric has to either drop the axis or be explicit about how much it is guessing.

What this means in practice

A tool that scores with a model is not a faster version of a tool that scores with a person; it is a structurally different instrument, built around a different rubric even when the axis names on the two look identical. Rate Cock runs the model version, with a rubric written for what pixels can actually support. What a commissioned human review looks like from the inside - the reasoning a judge writes, the discretion it applies, the edge cases it can catch that no fixed axis was told to look for - is ratepenis.com's subject, and it is a genuinely different product rather than the same rubric run by a slower rater. Reading which of the two you actually want for a given question is closer to a methodology decision than a quality one, and the tape-and-method approach measuremycock.com covers sits outside both, since it never asks either kind of rater to infer anything a ruler can settle directly.

Read next

Full archive