How it works

The label stayed; the thing it measures moved

After retraining with new labels, an axis can keep its name and change what it responds to, which nobody notices from outside.

3 min readHow it works

An axis called "symmetry" in one model version and "symmetry" in the next is not guaranteed to be scoring the same thing. The name is a label on the output. What the axis actually responds to is set by the training labels behind it, and those labels get regenerated - by new labellers, a revised instruction, or a refreshed anchor set - every time the model is retrained.

How the meaning moves without the name changing

A retrain usually happens for reasons unrelated to any single axis: more training data, a better backbone, a scheduled refresh. Somewhere in that process, a new batch of labellers is briefed, sometimes on a lightly revised written instruction, sometimes on the same instruction interpreted slightly differently by a different team. Small shifts of that kind compound: an axis that leaned toward reading contour under the old labelling pass can lean toward reading proportion under the new one, while the interface above it never changes a single word.

Why this is different from the scale shifting

This is not the same failure as a tool's overall scale creeping up or down across versions, which is a scoring-ceiling problem that shows up as everyone's numbers moving together. Rubric drift is narrower and harder to spot: the scale can hold steady in aggregate while one specific axis quietly starts responding to a different underlying property than it used to. A user comparing their score before and after a retrain sees a number move and reasonably assumes something about their photo changed, when what actually moved was the axis's definition.

Why it stays invisible

No interface displays "this axis was relabelled in March," and there is rarely a reason for a product team to announce it, since from their side it looks like an ordinary model improvement rather than a change to what a number means. The only way to detect drift from outside is the same method that catches most silent version changes: keep a fixed reference set and rescore it periodically, watching for a specific axis moving relative to the others rather than the total moving as a whole. That is a research habit, not something most users will do, which is exactly why drift can persist unnoticed for a long time - it is the same habit penisrater.com recommends for anyone trying to read their own score history as more than a string of unrelated numbers.

What to take from this

Treat an axis's meaning as tied to a specific model version, not to its name. A six-month-old memory of what a given axis rewarded is not a reliable guide to what it rewards now. Rate Cock reports the breakdown behind every score rather than a bare total, which at least gives a rescored photo a full profile to compare against a past one axis by axis, rather than a single number that has nowhere to show which part moved. The comparison problem this points at - reading what actually changed between two scores of the same subject - is closer to what a controlled comparison across model runs looks like than to any single axis's drift, and human review sidesteps versioning entirely, since a judge on ratepenis.com is not running a model that gets silently swapped underneath a returning user, and a measurement taken with an actual tape, as measuremycock.com covers, does not drift with a software release at all.

Read next

Full archive