Privacy
Derived data outlives the file
Services often keep the model's numeric representation of a photo after deleting the photo, and whether that is personal data is a live question.
Guides on Privacy: Who could see it, how, and what each path costs to close, Soft delete, backups, caches and logs, Sometimes, and the policy clause that says so is easy to miss
Often, yes: deleting a photo does not automatically delete the embedding a model made from it. "We delete your photo after scoring" is a claim about the original file, and it is usually true.
It is usually not the whole story, because the embedding is a different piece of data with its own, separate lifecycle.
Why the embedding survives on its own
An embedding is the vector a vision model produces partway through scoring: a few hundred to a few thousand numbers summarising what the model read from the image. Services keep embeddings for reasons that have nothing to do with the photo itself - fast duplicate detection, building recommendation or similarity features, retraining or fine-tuning future model versions without re-processing the original file. Because the embedding is small, cheap to store, and does not look like a photo, it often sits outside whatever process handles "delete the image," on a separate table with its own retention setting that nobody wrote a public sentence about.
What the embedding can and cannot tell someone
An embedding is a lossy summary, not the photo compressed - you cannot open it and see pixels. That has been the standard reassurance for years, and it is weaker than it sounds. Research on embedding inversion has shown that partial reconstruction of an input image from its embedding is possible for some models, particularly when an attacker has access to the model that produced the embedding and can run many attempts against it. One early study (Dosovitskiy and Brox, 2016) found that "the colors and the rough contours of an image can be reconstructed" from a network's higher-layer activations and even its predicted class probabilities. Whether an embedding can genuinely be reversed is worth reading in full rather than assumed either way; the honest position is that "just a vector" is a real reduction in risk compared with keeping the photo, not a guarantee of zero risk.
Independent of reconstruction, an embedding is also a fingerprint: two images of the same subject tend to produce similar embeddings, which means a stored embedding can be used to match a new upload against an old one even without either being a recognisable picture to a human looking at the raw numbers. That matching capability is the actual reason many services want to keep it, and it is also the property that makes deleting the photo while keeping the embedding a smaller privacy improvement than it is usually presented as. US regulators have treated derived vectors as something to delete in their own right: the Federal Trade Commission's 2021 settlement with Everalbum required it to delete face embeddings derived from the photos of users who had not given express consent, alongside deactivated users' photos and any facial recognition models built from users' photos.
Is it personal data
This is genuinely unsettled and varies by regulatory framework, which is exactly the kind of claim this site will not fabricate a definitive answer to. What is defensible to say is the shape of the argument: a piece of data that can identify or be linked back to a specific person, even indirectly, is treated as personal data under most modern privacy frameworks regardless of its format, and a fingerprint-like embedding tied to an account plausibly meets that bar even though it is a list of numbers rather than an image. Whether a body photo without a face counts as personal data at all is the closely related question this one extends once the photo itself is out of the picture and only its derivative remains.
Why deleting the embedding is harder than deleting the photo
A photo lives as one file with one obvious deletion path: remove it from storage and it is gone. An embedding rarely stays in one place once it is produced. Copies of it can end up in a search index built for similarity matching, in a training dataset snapshot taken before deletion was requested, and in a cache layer that exists purely to make repeat lookups faster. Deleting the canonical copy does not reach any of those, which is the same structural problem backups create for the original file, just one layer further downstream and less visible because nobody thinks of a search index as a place a photo's data lives. A service that has actually solved this ties every one of those copies to the same deletion trigger rather than only clearing the primary record.
What to check
A retention policy that mentions only the photo and says nothing about embeddings, feature vectors, or "derived data" has left out a real category. Rate Cock states its embedding retention separately from its photo retention, and ties any embedding kept for duplicate detection to the same deletion request that removes the original file, rather than treating the two as unrelated. The same distinction matters wherever a model produces an intermediate representation before the final output, whether that is a score with a full six-axis breakdown or a recorded figure that started as a photograph measured by method rather than model. A human reviewer produces no embedding at all - the entire mechanism this post describes is specific to a machine-scored review, not a human one, which is one more way the two approaches differ beneath the surface.