How it works
Where your photo lands, and what "near" means
The model places every image in a high-dimensional space; a score is a function of where yours lands, not of anything measured on it.
Guides on How it works: Every score is a memory of somebody's opinion, An axis earns its place by being separately observable, What an image becomes before it is scored
Latent space is the space of a few hundred dimensions where a rating model places every photo as a point, and a score is a read-out of that point's position. The score reads from there - not from the photo, and not from anything measured on the subject.
A space you cannot draw
A dimension, in the everyday sense, is something like height or width - one you can point at. A latent space has hundreds of them, each corresponding to some direction the model's training settled on as useful, and none of them individually correspond to a property a person would name. Two dimensions are easy to draw as a map; a few hundred are not, and nobody, including the people who trained the model, has a full picture of what each one represents.
What can be reasoned about is not the individual dimensions but relationships between points - which images sit near each other, which sit far apart, and along what rough directions groups of similar images tend to cluster. That is enough to be useful even without being able to name any single axis, in the same way you can navigate a city by relative position without knowing the name of every street.
What "near" means
Two points being close in this space is usually measured by comparing the angle between them - a method called cosine similarity - rather than by simple straight-line distance, and the reasons that particular measure is standard, plus what it quietly ignores, are covered in cosine similarity: the one formula behind "looks like". For this post, the important fact is just that nearness is well-defined and computable, and that images the model treats as similar end up near each other as a direct consequence of how the encoder was trained. OpenAI's CLIP (Radford and colleagues, 2021), a common starting point for such encoders, learned its space from "400 million (image, text) pairs" by pulling each image toward its own caption, which is why nearness tracks what images tend to be described as rather than any measured quantity. That step is covered from the encoding side in what an image becomes before it is scored.
A score as a read-out of position
Once an image has a position in this space, producing a score is a separate, comparatively simple step: a scoring head - often not much more than a small additional network - takes the position and maps it onto a number. It was trained the same way the encoder was, by adjusting itself until its outputs matched the labels it was shown, and what it has actually learned is a rough correspondence between region of the space and typical label for images that land there.
This has a specific and slightly uncomfortable consequence. Two photographs that are genuinely different in the property you care about, but that happen to land near each other in the space for reasons unrelated to that property - similar lighting, similar composition, similar overall texture statistics - can receive very similar scores, because the scoring head has no way to distinguish them once they are close together as points. The read-out only has access to position; it has no separate channel back to the raw photo to check whether the closeness is meaningful or coincidental.
Clusters, and what they are made of
Training tends to organise the space into rough clusters - regions where similar kinds of image congregate - and a photo that is unusual relative to the training data ends up somewhere sparse, far from any dense cluster, where the scoring head has seen comparatively few examples to learn from. Scores in sparse regions are less reliable than scores in dense ones, not because the model is being deliberately cautious, but because a function fitted mostly from dense-region examples extrapolates worse the further it has to reach. This is the same underlying idea behind distribution shift and behind why unusual features can produce unstable results, without repeating either of those explanations here.
What this rules out
Nothing in this description involves a ruler, a stored reference size, or anything that behaves like measurement. Position in the space is a function of everything the encoder was sensitive to in the photo - framing, lighting, texture, arrangement - compressed together, and a score derived from that position inherits all of it at once, which is one reason the same subject can land at different points, and different scores, across a single afternoon. For an actual physical figure, Measure My Cock's method does not involve a latent space at all, which is precisely the point of using a tape.
Multi-axis systems read several scores off related but distinct regions of structure in this same space rather than a single point-to-number mapping, which is part of why a rubric outperforms one aggregate figure; Rate Cock's six-axis breakdown is one working example of that approach. Understanding a result this way is also useful groundwork for interpreting what Penis Rater's guidance on reading a score actually means in practice, and it is worth knowing that a human reviewer does not operate on anything like a latent space at all - the comparison between the two approaches is covered on the judging side by Rate Penis.