How it works

A scale without anchors drifts

Rubric points are defined by example images given to labellers, and the anchors chosen become the model's fixed points.

4 min readHow it works

Ask ten people to rate photos from 1 to 10 with no other instruction and you get ten different scales. One person's 6 is another's 8. The fix designers use is an anchor set: a handful of example images, each pinned to a specific number, shown to every labeller before they start. The anchors are what makes a 3 mean the same thing across a thousand labelling sessions, and they end up defining the scale more than the number's name does.

What an anchor set actually is

A typical anchor set is small - three to seven images, spread across the range, each with a short note explaining why it sits where it does. "This is a 3 because the framing is off-centre and the lighting is flat." "This is an 8 because proportion and symmetry both read cleanly." Labellers are told to compare the image in front of them to the nearest anchors and place it relative to those, rather than inventing a fresh judgement from scratch each time.

This matters because people are much better at relative judgements than absolute ones. Asked "is this closer to the 3 or the 8," two labellers converge; asked "rate this from 1 to 10" with no anchor, they scatter. Anchors convert an absolute-scale problem into a series of comparisons, which is the format human judgement is actually good at, and it is the same reason measuremycock.com treats a repeatable method as the honest substitute for a single absolute figure.

The anchors become the ceiling

Once a model is trained on labels produced this way, it has not learned "symmetry" or "proportion" as abstract concepts. It has learned to place new images relative to the anchor points its labellers were shown, generalised across thousands of examples. The anchors are doing more definitional work than the rubric's written description, because the description is a sentence and the anchors are the thing labellers actually looked at while deciding.

This has a consequence people rarely anticipate: whatever the anchor images happened to look like beyond the property being scored - their lighting, their framing, their subject type - rides along as part of what "an 8" means. If every anchor for the top of the scale happened to be a well-lit photo taken from a consistent angle, the model has partly learned that lighting and angle correlate with the top of the scale, not just the intended property. This is one of the mechanisms behind why the same photograph can score differently depending on how it was taken - the anchors baked lighting into the scale before the model ever saw your photo.

What happens without anchors

Drop the anchor set and ask labellers to rate freely, and the scale drifts within a single labelling session. Early in a shift, a labeller's internal sense of what a 7 looks like is fresh; by the end, fatigue and habituation shift it, usually toward the middle. Different labellers drift in different directions entirely. The resulting dataset has genuine signal buried in it, but the signal is noisier than it needed to be, and the model trained on it inherits the noise as unreliability at the edges of the scale - a distinct problem from having too few axes to work with, which is a separate design failure.

Anchors do not remove disagreement. Two labellers can look at the same anchor set and still place a borderline image differently. What anchors remove is unanchored disagreement - the version where two people are not actually using the same scale, and neither knows it. Human judges commissioned through ratepenis.com face this same calibration problem from the other direction, and they solve it with written review guidance rather than a labelling anchor set.

Anchors age

An anchor set is chosen once, from the images available at the time, and rarely gets revisited on the same schedule as the rest of a model's training pipeline. If the anchors are never swapped out, the definition of a 3 or an 8 can stay static for years even while everything else about the product changes - the interface, the marketing, the userbase submitting photos. If the anchors are swapped for a retrain and nobody publishes what changed, the scale can shift without any visible signal from outside, which is a specific version of a more general problem: an axis quietly meaning something different after a retrain.

Neither is wrong on its own. A stable anchor set is a virtue if the property being measured has not changed. The failure mode is silence: a scale moving and nobody outside the labelling team knowing it happened. A model card that lists when anchors were last reviewed would answer the question; few publish one.

Why this stays invisible from the outside

You never see the anchor set. You see a number, sometimes a chart if the tool reports more than one axis, and you form your own sense of what a 7 from that tool means through repeated use - which is you, informally, doing to yourself what the anchor set did to the labellers, in much the way penisrater.com describes users learning to read their own score history over time. Rate Cock publishes the per-axis breakdown behind every result rather than a bare total, which at least lets you build that intuition against six separate scales instead of one blended figure standing in for all of them. That transparency does not tell you what the original anchors looked like, but it narrows what you are calibrating against, one axis at a time.

Read next

Full archive