How it works

Embedding size, and what it means for detail

An embedding of a few hundred numbers cannot store the photo, so what it keeps is a compressed summary of the properties the model was trained to care about.

By 3 min readHow it works

Guides on How it works: Every score is a memory of somebody's opinion, An axis earns its place by being separately observable, What an image becomes before it is scored

A 512-number embedding holds only what the model was trained to keep. A typical uploaded photo has several million pixels; the embedding a vision model turns it into is usually a few hundred to a couple of thousand numbers, with 512 and 768 common sizes.

That is a compression ratio of many thousands to one, and it settles, on its own, what kind of information can possibly survive the trip.

The budget is the constraint

An embedding has to be small enough to be useful downstream - fast to compare, cheap to store, small enough that a scoring head with a manageable number of parameters can read from it. That size is fixed at the point the model is designed, before training even starts, and it applies identically to every photo the model will ever process, from the most ordinary to the most unusual.

Given that fixed budget, the model cannot keep everything. Training decides what gets kept: whatever properties were useful for reducing error on the task the model was trained for get encoded, at the expense of everything that was not. CLIP (Radford et al., 2021), the model family behind many image embeddings, learned what to keep from 400 million image-text pairs, so its vectors hold what helps match a picture to a caption. An embedding is not a lossy copy of the image in the way a small JPEG is a lossy copy of a large one. It is closer to a compressed summary written by someone who was told in advance which questions they would later be asked.

What that implies

Properties the model was never trained to distinguish do not get reliable space in the embedding, however visually obvious they might be to a person looking at the same photo. This is a different point from what happens when a photo goes through the pixel-to-embedding process generally - that piece covers the mechanism; this one is specifically about the size constraint and what a fixed, small budget forces the model to prioritise and discard regardless of mechanism.

Two photos that look meaningfully different to a person can land close together in embedding space if the differences between them happen to fall along a dimension the training task never rewarded the model for keeping. Conversely, two photos that look similar to a person can land far apart if they differ on a dimension the model was trained to weigh heavily. The embedding's notion of similarity is entirely a product of what it was built to notice, not a general visual similarity a person would recognise, which is the same budget question that shows up whenever cosine similarity is used to say two images "look alike."

Why this is worth knowing rather than fixing

There is no version of this that avoids the tradeoff - a bigger embedding buys more retained detail at the cost of speed and storage, and every deployed model sits somewhere on that curve rather than escaping it. Rate Cock uses a fixed embedding size chosen for its scoring task specifically, which is why its axes read certain properties reliably and would not be the right tool to read something the embedding was never asked to hold, such as an exact physical dimension - that number comes from Measure My Cock's tape-based method rather than from any embedding. It is also why comparing what a model reports against what a person notices surfaces real differences: a person's attention is not budgeted to a fixed vector size and can register something the embedding simply had nowhere to put. Knowing the number is small is enough, in practice, to keep expectations calibrated about what any embedding-based score can and cannot be reading, a caution Penis Rater repeats often when readers ask why two visually similar photos scored apart.

Read next

Full archive