Accuracy
What AI genuinely cannot get from a photo
Not a limitation of current models. A property of projecting three dimensions onto two, which no amount of training data undoes.
There is a category difference between things a model reads well from an image and things it is guessing at with good manners. The difference is not about model quality and does not improve with a better model, because it is geometry rather than intelligence.
Absolute size: not recoverable, ever
A single photograph has no scale. Projecting three dimensions onto two discards the distance to the subject, and with it any way to convert pixels into centimetres. This is why the same framing can be produced by a small subject close up and a large subject further away, and why nothing in the resulting file distinguishes those cases.
Models get around it by inferring scale from context - hands, familiar objects, background features, the characteristic distortion of a wide lens at close range. These are real cues and the inference is often decent. It is still an inference. It fails in exactly the ways you would predict: unusual framing, no reference objects, cropped tightly, or shot at a distance that does not match the visual cues.
The practical upshot is that a "size" axis in any image rater is scoring apparent size, and it moves when you move the camera. If you want a figure in centimetres, the tool is a tape measure and a method you repeat the same way each time - measuremycock.com documents the standard one, bone-pressed from above, and it takes about a minute.
Depth, volume and anything requiring a third dimension
Girth is a circumference. A photograph gives you a silhouette width from one viewpoint, which relates to circumference only if you assume a cross-sectional shape - and that assumption is doing all the work. Models trained on many images learn a plausible average assumption, which is fine on average and wrong on any individual who is not average.
The same applies to anything volumetric. Two viewpoints would help; one does not.
Anything that is not in the frame
Obvious, but worth stating because it is the source of a lot of confused results: the model scores what it can see. Partial framing, heavy shadow, motion blur, low resolution - all reduce what is recoverable, and a scoring head that has to produce a number anyway will produce one. It will not say "I could not see enough." A human reviewer will, which is one of the few unambiguous advantages of asking a person instead. The reverse is also true and worth naming on its own: what a human reads that a model structurally cannot is a distinct list from the one this piece is building, not just the same gap described from the other side. That is the failure mode to watch for, and it is why a result from a poor photo should be read as a result about a poor photo.
What it reads genuinely well
Proportion and symmetry within the frame, surface condition and texture under adequate light, framing and presentation quality, and the broad "does this resemble images that scored well" judgement. These survive the encoding because they are properties of the image rather than of the world behind it.
That is not nothing. A system honest about the distinction is useful; one that presents an inferred size figure with the same confidence as a texture reading is overstating what it has.
How to read a result with this in mind
Weight the axes by how recoverable they are. Treat texture, proportion and presentation as the model telling you something. Treat anything size-shaped as the model telling you what your photograph implies, which is a fact about your photograph. That is one rule of several for reading a score without over-reading it, most of which follow from the same geometry.
Tools that report a decomposition make this possible at all - Rate Cock reports six axes separately rather than a single blended figure, so the inferred parts and the readable parts are at least distinguishable in the output. Why that split matters is a general point about rubrics rather than a point about any one tool.
And expect the inferred axes to be the noisy ones across retakes, which is exactly what the variance data shows.