Privacy
Inversion attacks, and what they recover
An embedding is lossy, but research has shown partial reconstruction from some models, so 'we only keep the vector' is a weaker promise than it sounds.
Guides on Privacy: Who could see it, how, and what each path costs to close, Soft delete, backups, caches and logs, Sometimes, and the policy clause that says so is easy to miss
Partly, yes: research has reconstructed recognisable approximations of photos from the embeddings vision models produce. An embedding is a lossy summary, but lossy is not the same as empty, so "we don't keep the image, just the embedding" is a weaker reassurance than it sounds.
What an embedding actually keeps
A vision model turns a photo into a vector - a list of a few hundred to a few thousand numbers - that encodes the properties the model was trained to care about: shape, texture, arrangement, lighting, and whatever else the training objective rewarded it for noticing. It is not a compressed image file in the way a JPEG is; there is no decoder built to turn it straight back into pixels, because nothing in the training process asked the model to preserve that ability. That absence is exactly why "we only keep the vector" sounds safe.
Why inversion works anyway
Dosovitskiy and Brox (CVPR 2016) showed that a separate network, trained specifically for the purpose, could take the features produced by a standard vision model and reconstruct an approximation of the original input. Not pixel-perfect, but they found that "colors and the rough contours of an image can be reconstructed from activations in higher network layers and even from the predicted class probabilities." The trick is that the embedding was never designed to resist this; it was designed to be useful for a task, and usefulness for a task and resistance to inversion are not the same property, and nothing in ordinary training optimises for the second one. An attacker who can pair enough embedding-image pairs, or who has access to the model itself, can train an inversion network the same way researchers do.
What determines how much comes back
Reconstruction quality is not fixed - it depends on the embedding's dimensionality, on how much of the original signal the training objective preserved, and on how much access an attacker has to the model that produced the embedding in the first place. A larger embedding, holding more of the 512 or so numbers a typical vector carries, preserves more signal and is generally easier to partially invert than a small one deliberately compressed further for storage. Black-box access - querying a public API without seeing the model's internals - makes inversion considerably harder than white-box access, where an attacker has the model weights directly, which is the scenario most published inversion research actually uses. This is also why a hash is a different and generally harder target than an embedding: a hash was explicitly designed to be one-way, an embedding was not designed with that property in mind at all.
What this changes about "we only keep the vector"
It changes the claim from a settled fact to a claim with conditions attached: how large is the vector, is the model or an equivalent one available to an attacker, and how motivated would someone need to be to build an inversion network for it. For most casual scenarios, the effort required is well beyond what makes reconstruction a live day-to-day risk - this is not a reason for alarm about routine use of a scoring tool. It is a reason the sentence "we do not keep the image, only the embedding" should not be read as equivalent to "we keep nothing sensitive," particularly for anything held long after the source photo itself has been deleted.
Reading a policy with this in mind
A policy that addresses embedding retention at all is already ahead of most, since the concept rarely gets a sentence of its own. The stronger version states how long embeddings are kept relative to the photo, whether they are used for anything beyond producing the original score, and whether they are ever shared with a third party for a purpose unrelated to scoring.
Rate Cock publishes what it retains after scoring in its privacy documentation, and reading the embedding-specific line, where one exists, is worth the extra minute over reading only the headline "we delete your photo" claim. The equivalent record for a physical measurement is a number rather than a vector, and Measure My Cock's approach to that data is a useful comparison precisely because a number cannot be inverted into an image the way an embedding sometimes can. Comparing what different tools keep after scoring, feature by feature, is groundwork Penis Rater's tool coverage is set up to make faster than reading ten separate privacy policies from scratch. A human reviewer produces no embedding at all, which is one of several structural differences from a model-scored result that Rate Penis's coverage of commissioned reviews is worth reading against this piece.