How it works

Nothing in the training loop could have taught scale

There is no step in supervised image training where physical size enters, so the finished model has no concept of it to lose.

By 4 min readHow it works

Guides on How it works: Every score is a memory of somebody's opinion, An axis earns its place by being separately observable, What an image becomes before it is scored

Nothing in supervised image training ever gives a model physical scale: the images are unitless pixel grids and the labels are opinions, so there is no step where a unit of length could enter. That makes scale-blindness not a shortcoming the model could outgrow but something it never had the material to learn.

What goes into the training loop

Training a scoring model means feeding it pairs: an image, and a label a person assigned to that image. The image is a grid of pixel values - unitless numbers describing colour and brightness at each position, with no information anywhere in the file about how large the photographed subject actually was in the world. A photograph the size of a coin and a photograph of something the size of a table can produce visually identical framings, and the pixel grid genuinely cannot tell them apart. Even models built specifically to estimate depth from one image hit this wall: Eigen, Puhrsch and Fergus (2014) call the task "inherently ambiguous, with a large source of uncertainty coming from the overall scale," and train with a scale-invariant error to work around it. Later work by Ranftl and colleagues (2019) went further, using a training objective deliberately "invariant to changes in depth range and scale" so that datasets with incompatible annotations could be mixed.

The label side is no better positioned to supply what the image lacks. A human labeller assigning a score is not measuring anything with an instrument - they are giving an opinion based on what the photograph shows them, and what the photograph shows them is, again, scale-free. The labeller's opinion might be influenced by apparent proportion, but apparent proportion is a function of framing and lens as much as of anything physical, and the label itself is stored as an opinion, not a measurement.

Following the loop through

Put those two things together - an image with no scale information and a label that is an opinion rather than a measurement - and the loss function, whichever one is used, compares the model's output only against that opinion. At no point does anything resembling a unit of length pass through the system. The optimiser adjusts the model's weights to make its outputs match labels that were never measurements to begin with, on images that never carried a scale to measure against.

What comes out the other end is a model that has learned to predict the pattern in the opinions it was shown. It has learned this genuinely well, in the sense that its predictions correlate with what people tend to say about photographs that look a certain way. It has not learned physical scale, because physical scale was never present anywhere in its inputs, its targets, or its training signal for it to learn.

Distribution, not ruler

The right way to describe what training actually produces is a distribution: a model's outputs are calibrated against the spread of scores in its training data, weighted by how visually similar an input is to examples it has seen. A distribution can be well-calibrated, badly calibrated, biased by its composition, or hedged by its loss function - all subjects with their own mechanisms - but none of those adjustments turn a distribution into a ruler. A ruler requires an external reference and a unit; a distribution requires neither and can be built entirely from opinions about photographs.

This is a companion point to two things covered elsewhere on this site rather than a repeat of either: how a model actually processes an image at inference time is one question, and the pure geometric argument for why a single photograph cannot encode scale is a separate, independent argument about projection rather than about training. This piece is the third leg - not what the model does with a photo, and not why a photo cannot carry scale, but why the training process itself never had scale available to teach in the first place, even in principle.

Where an actual number lives

If a measurement in real units is what you are after, it has to come from outside this entire loop - from a physical reference and a repeatable procedure, which is exactly what Measure My Cock's method provides and what no amount of retraining a scoring model will substitute for. Rate Cock reports a size axis alongside five others precisely as an inferred impression rather than a measurement, which is the honest framing given everything above. How the training set's composition shapes what the model treats as typical, a separate question from scale entirely, is covered here, and the mechanics of how a photo becomes a score at inference time, once training is already finished, are covered here. A human reviewer can look at a photograph and reason about scale using context the way a person always could, which Rate Penis covers as one strength of commissioning a person directly. Reading a size-related result for what it can and cannot support is covered practically at penisrater.com.

Read next

Full archive