Privacy

Synthetic data avoids uploads, with its own catch

Some teams train on generated images so no user photo enters training, which is a real privacy gain and a real accuracy trade.

By Updated 3 min readPrivacy

Guides on Privacy: Who could see it, how, and what each path costs to close, Soft delete, backups, caches and logs, Sometimes, and the policy clause that says so is easy to miss

Training on generated images is a real privacy gain, because no uploaded photo needs to enter the training set, but it is a narrow one with an accuracy cost attached. It removes the path from your upload into a dataset, and nothing else about what a service does with your file today.

What this actually buys

The privacy gain is specific, not general. It removes one path: your submission does not need to be retained as training material, because the model was built before you ever uploaded anything. It does not remove the other paths a service can still take with your photo - storing it for the result page, running it through a separate quality check, or sending it to a different vendor's API for part of the pipeline. Synthetic training data is a statement about where the model came from, not a statement about what happens to your file today. Even the gain it does offer is weaker when the generator was itself trained on real photos: in "Synthetic Data - Anonymisation Groundhog Day," Stadler, Oprisanu and Troncoso found that "synthetic data either does not prevent inference attacks or does not retain data utility," and that what a synthetic set preserves cannot be predicted in advance.

The accuracy side of this trade is its own subject, and it is worth reading before treating "trained on synthetic data" as an unqualified good: a generator's idea of what a body looks like is a real input into the resulting model, and it is not neutral. This post is the privacy half only, and it does not repeat that ground.

Why it is not the default

Generated images are expensive to produce well, and a model trained purely on them tends to inherit the generator's biases about typical appearance rather than the wider range visible in real uploads, which is a cost real developers weigh against the privacy benefit. Most tools instead take real uploads, sometimes with consent language buried in a policy, because it is cheaper and the resulting model tends to generalise better to what real users actually send. A service that has gone the synthetic route usually says so, because it is a genuine selling point, and a policy that does not mention where its training images came from is not making that claim.

Where to check the rest

Whether an anonymised dataset claim can be trusted at all covers the harder case, where real photographs are the training material and the question is what "anonymised" removed. Method matters on the measurement side too, and Measure My Cock's data page is the place that covers dataset practices for a tool built around a physical figure rather than an inferred score. A tool that trains on real submissions is not automatically doing anything wrong, but Rate Cock and any comparable service should say which path they took, and a policy that is silent on the question has answered it by omission. For the reviewing side of a pipeline rather than the training side, Penis Rater's coverage of what a tool actually does with a result is a useful companion read. A human reviewer never has a training-data question in the first place, since a commissioned review is one person looking at one photo rather than a model learning from thousands - a different privacy shape entirely.

Read next

Full archive