Accuracy
An engagement-trained model ranks what spreads
A scorer trained on likes learns the biases of an audience, including which bodies and photos get shown in the first place.
A model trained on likes and views is not learning what people rate well. It is learning what already got shown to people, and what a specific audience chose to reward once it appeared. Those are two different targets, and the gap between them is where popularity bias lives.
What the label is actually tracking
An engagement count is downstream of a lot more than quality. It depends on who saw the post, which depends on a feed algorithm that was itself optimising for engagement on a previous round of posts. It depends on the time of day, the caption, the platform's existing audience, and whether the account had followers already. None of that is the photograph. A model trained to predict likes from pixels is trying to recover a signal that was mostly generated somewhere else in the system, and it can only use what is visible in the image to do it.
Two separate biases, stacked
The first is exposure bias: content from accounts and body types that were already popular gets shown more, so it accumulates more engagement regardless of anything about the photo itself, and a model trained on the outcome learns to reward whatever those accounts had in common. The second is audience bias: engagement reflects the taste of whoever was doing the liking, on that platform, at that time, which is a narrower and less representative group than "people in general." A model does not distinguish between "this is broadly appealing" and "this is what this particular audience rewards," because the training signal cannot tell the two apart.
Why it does not look like an error
This is the part that makes popularity bias hard to catch from outside. The model is still consistent - the same photo gets close to the same score twice, which is the property most people check for. It still correlates with something real, since audience preference is not random. What it will not do is tell you when its sense of "good" has quietly become "what this platform's users tend to reward," which can diverge from a more general judgement in ways that are specific to whoever was doing the labelling. Where a model's training labels come from matters more here than for a scorer trained on judged ratings, because there is no rubric standing between the audience's behaviour and the number.
What to do with a score like this
Treat an engagement-trained score as a measure of "would this get attention on this kind of platform," not as a general judgement, and read aesthetic bias from general photo training alongside it, since the two compound rather than cancel. Rate Cock trains its scoring on judged rubric labels rather than engagement metrics for exactly this reason - the target is a rating, not a prediction of what would spread. Whether a score is even reproducible at all is a separate, prior question that measuremycock.com covers on the data side, and it is worth settling before worrying about which audience a number reflects. Comparing a single score against your own history is a different exercise again, one penisrater.com treats as its main subject, and it does not require solving the audience question first. Human review sidesteps the exposure problem differently again, since a commissioned judge is not scoring based on what a feed already amplified. A popularity-trained number can be a real signal about a real audience. It is a different signal from the one most people assume they are reading.