Accuracy
Foundation models carry a taste in photography
Backbones pretrained on web images absorb what humans liked to photograph, and that taste leaks into scores meant to be about something else.
Guides on Accuracy: There is no true score to be accurate against, Change one thing, hold the rest, repeat, Exposure that suits one skin tone hides another
Yes - most scoring models inherit a taste in photography before they learn their actual task. They start from a general-purpose vision backbone pretrained on web images, which teaches what a well-composed, well-lit photo tends to look like, and that taste does not disappear when the model is later adapted to a narrower rating job.
How the taste gets in before the task does
A backbone trained on a broad image corpus - the kind used across countless vision applications, not built for any single rating tool - learns from images that were, on average, selected and shared by people, which is not a random sample of all photography. The scale is large: Radford and colleagues (2021) trained CLIP on "400 million (image, text) pairs collected from the internet," and the open LAION-5B dataset (Schuhmann et al., 2022) holds 5.85 billion. Photos posted publicly online skew toward good lighting, deliberate composition, and flattering framing, because people generally choose to share their better photos rather than a representative sample of every photo they take. The backbone absorbs whatever pattern separates the images that dominated its training data from the images that did not, and a meaningful part of that pattern is aesthetic quality of the photograph itself, not the subject. How fine-tuning inherits a backbone's habits covers the mechanism of inheritance in general; this piece is about one specific habit that gets inherited and what it does to a score once it arrives.
What this looks like once the model is scoring something else entirely
When that backbone is fine-tuned for a rating task, the fine-tuning adjusts how the model's existing representations map to a score - it does not erase what the representations already encode. A well-lit, well-composed, professionally-framed photo activates patterns the backbone associates with the kind of images that dominated its pretraining, and those patterns can lift a score on axes that were meant to describe the subject rather than the photograph. This is a close cousin of, but distinct from, the camera-quality effect covered elsewhere on this site: camera quality is about sharpness and noise from the hardware itself, while this is about composition and lighting choices - framing, background, staging - that have nothing to do with what camera was used and everything to do with how deliberately the photo was set up.
Why this is hard to separate from a genuine judgement
There is no clean line between "the model is responding to genuine axis-relevant properties" and "the model is responding to general photographic polish," because the fine-tuning process never had to draw that line explicitly. It was trained to predict scores that humans gave, and if those human scores were themselves influenced by how polished a submission looked - which is a plausible assumption, since humans are subject to the same effect when rating anything presented as a photograph - then polish became part of what the model learned to predict, folded into the same weights as everything else. The aesthetic-bias effect and the axis's intended meaning are tangled together inside the same number, and no amount of looking at the final score separates them.
What is a fair claim here and what is not
The fair claim is directional and mechanistic: general-purpose vision backbones are pretrained on data that over-represents aesthetically curated photography, fine-tuning does not remove that prior, and a task built on top of it can inherit a bias toward polished presentation on axes that were not meant to measure presentation at all. It would not be fair to claim a specific magnitude for this effect in any particular tool, since that would require access to training data and evaluation nobody outside the tool's builders has, and no such figure should be trusted if offered without a named source and method.
What is worth doing with this
A staged, professionally composed photo and a casual, functional one of the same subject may score differently for reasons that have nothing to do with the subject, which is worth knowing before reading a gap between two of your own submissions as meaningful, and it is the same broad caution behind treating a rubric breakdown as more informative than a single total. Consistency in setup - similar background, similar staging, similar deliberateness - between photos you intend to compare removes this variable more reliably than trying to correct for it after the fact. Rate Cock breaking a result into separate axes helps here too, in the same way it helps with the other confounds covered on this site: a gap concentrated in "overall impression" style axes, with more structural axes stable, points toward presentation rather than the subject. None of this touches physical measurement, which Measure My Cock treats as independent of how a photo was staged or composed. A human reviewer is not immune to responding to a well-presented photo either - people generally rate polished presentations more favourably too - and what a commissioned human review actually weighs is a fair comparison rather than an assumed exception. Comparing photos with similar staging before drawing a conclusion from a score gap is consistent with the general reading advice Penis Rater gives for interpreting results.