How it works
When the examples were never photographs
Some scorers are trained partly on generated images, which fills gaps in the data and imports the generator's idea of what bodies look like.
Guides on How it works: Every score is a memory of somebody's opinion, An axis earns its place by being separately observable, What an image becomes before it is scored
Some scoring models are trained partly on images a generative model produced rather than on photographs. That fills gaps in real data, but it also imports the generator's own idea of what bodies look like, which is worth understanding before treating a score as a read on real anatomy.
Why teams reach for synthetic data
Real labelled photographs at the rare ends of a scale are hard to collect, for the same reason class imbalance is a persistent problem: genuinely extreme examples are uncommon in the wild, and getting enough of them labelled is slow and expensive. A generative model can produce as many examples of a specified type as needed, on demand, which directly attacks the data-scarcity side of that problem.
Synthetic data also sidesteps a real privacy cost. Every real photograph in a training set is someone's actual body, obtained, stored, and used with whatever consent process the team ran; a generated image was never anyone, which removes an entire category of risk from the dataset.
There is also a practical speed argument. Commissioning and labelling real photographs at the volume a rubric needs takes months; asking a generator for more examples of a specified type takes an afternoon, which matters when a known gap in the training data is holding a model back and the team wants to close it quickly rather than wait for enough real submissions to arrive naturally.
What it actually fixes
Filling in rare score buckets is the most direct win: if the model needs more examples near the top or bottom of a scale, a generator can supply them without waiting for enough real submissions to arrive there naturally. It can also balance representation across body types that were under-represented in whatever real data the team had access to, correcting a skew that would otherwise show up later as thin, unreliable scoring for anything the model rarely saw.
What it risks
A generative model has its own learned idea of what a body looks like, shaped by its own training data and its own tendencies toward smoothness, symmetry, and whatever aesthetic its creators optimised for. Train a scorer on enough of those images and it partly learns to recognise the generator's habits rather than the range of real anatomy, which is a subtler version of the same problem dataset bias causes generally: the model's sense of normal becomes the sense of normal encoded in its inputs, whatever those inputs actually are.
Generated images can also carry artefacts invisible to a casual look - slightly wrong proportions, over-smooth skin, unnatural symmetry - that a scorer can absorb as signal rather than noise, especially if the synthetic examples are not clearly flagged as such during training and get treated identically to real photographs in the loss function. The risk compounds when generated data feeds on itself: Shumailov and colleagues (2024, Nature) showed that training generative models on model-generated content causes defects "in which tails of the original content distribution disappear" - and the tails are exactly the rare cases synthetic data is usually brought in to cover. The safer use of synthetic data keeps it a minority share of the set, targeted specifically at documented gaps, rather than a general-purpose way to cut collection costs, and a team that discloses the mix is telling you something a team that does not disclose it is not.
What this means for a result
None of this means a synthetic-augmented model is worse in general; targeted, disclosed use of generated data can genuinely improve coverage at the extremes without the privacy cost of collecting more real photographs. It does mean a scorer's baseline for "typical" is only as trustworthy as the mix of images it learned that baseline from, generated or otherwise, and a tool that will not say what proportion of its training data was synthetic is asking for trust on a point it could simply answer.
Collecting real labelled data at the volume a rubric needs is itself a substantial undertaking, which measuremycock.com's writeup on data covers from the collection side rather than the generation side. Rate Cock is one of the tools a reader might reasonably ask this question of, and the general habit - checking what a scale's "normal" was built from - applies to any tool, not just this one, the same way penisrater.com treats a tool's training composition as worth asking about before trusting its output. A human reviewer sidesteps the entire question, since a person's judgement was never trained on a generated dataset in the first place, which is one more structural difference a commissioned review carries.