Accuracy
There is no true score to be accurate against
Accuracy requires a correct answer to compare with, and for a subjective rating the only candidate is other people's opinions.
"Accurate" is a comparison word. It only means something once you have said what the thing is being compared to. A thermometer is accurate against a calibrated reference; a spell-checker is accurate against a dictionary. An AI rating of a photograph is accurate against - what, exactly? This is not a rhetorical question, and the honest answer is that there is no single thing it can be accurate against, only several partial candidates, each of which is missing something.
Why "accurate" needs an answer key
Every accuracy claim, in any field, has the same shape: take a prediction, take a known-correct answer, measure the gap. The known-correct answer is called ground truth, and it has to exist independently of the model being tested, or the comparison is circular. For a scorer that classifies whether an image contains a cat, ground truth is easy - a human looks at the photo and says yes or no, and the two answers rarely conflict. For a rating on a 1-10 scale of a subjective quality, there is no equivalent fact of the matter sitting in the world waiting to be looked up. The number was never out there to be discovered. It only exists as an opinion, generated by whoever the model was trained to imitate.
This is the core of the ground-truth problem, and it is worth being precise about what kind of problem it is. It is not that current models are insufficiently sophisticated to find the true score. It is that "the true score" is not a coherent target for this kind of question, in the way it is for a thermometer reading a real temperature. Any accuracy claim about a subjective rating has to substitute something else in place of ground truth, and the substitute determines what the claim can and cannot support.
Candidate one: agreement with humans
The most common substitute is human judgement itself. A model is called accurate if its scores line up with what people give the same photos. This sounds solid until you ask which people, how many, and how they were instructed - agreement with human raters as the accuracy metric works through the details, but the summary is that "matches humans" is only as trustworthy as the panel it was measured against. A panel of five friends, a panel of five hundred strangers on a crowdsourcing platform, and a panel of five trained annotators working from a shared rubric will each produce a different "ground truth," and a model can score well against one and poorly against another without changing at all.
There is a deeper issue underneath the panel-composition question. Human raters do not agree with each other, even within a carefully chosen panel. When labellers disagree, the model cannot beat that ceiling - if two people rating the same photo only agree, say, seventy percent of the time, no model trained to predict "the" human answer can exceed that seventy percent either, because there is no single human answer to predict. The model can only be as accurate as the humans were consistent, and that number is usually lower than people assume before they measure it. Human agreement is a real and useful substitute for ground truth. It is also a substitute with a known, often disappointing ceiling baked into it before the model does anything at all.
Candidate two: consistency with itself
A different substitute drops the comparison to humans entirely and asks whether the model agrees with itself. Give it the same photo twice; does it return the same number? This is reliability, not accuracy, and the difference matters more than the similar names suggest. Consistent is not the same as correct - a broken clock that always shows the same wrong time is perfectly reliable and never accurate, and a scoring model can be the machine-learning version of that clock. Consistency is genuinely valuable, since it means a change between two of your own scores is a real signal rather than noise, but it says nothing about whether the underlying number is measuring the thing it claims to measure. A model can be highly self-consistent and consistently wrong about what a viewer would actually rate the photo, and consistency alone gives you no way to tell.
Candidate three: predictive validity
A third substitute steps outside the scoring system entirely and asks whether the score predicts something that happens later. Does a higher score correlate with more matches on a dating profile, more engagement on a post, more of whatever outcome the score is implicitly claiming to forecast? Predictive validity is the strongest possible form of ground truth, because it checks the score against a real-world consequence rather than against another opinion. It is also the rarest to actually see measured. Building this kind of validation requires tracking outcomes over time, linking them back to specific scores, and controlling for everything else that also affects the outcome - which platform, which audience, which caption, which timing. Almost no consumer rating tool publishes this kind of study, because almost none has done it. When you see a tool claim its score is "meaningful" or "predicts real reactions," ask what candidate for ground truth backs that claim, because the honest answer, most of the time, is none of the three above in any rigorous form.
What each candidate is actually good for
None of these three is wrong to use. Each answers a genuinely different question, and confusing them is where most misreadings of an "accuracy" claim happen.
Human agreement tells you whether the model's opinion resembles a chosen group's opinion. It is the right substitute if what you care about is "would people like this." It is the wrong substitute if you want to know whether the model is measuring a real property of the photo, because a panel's shared taste is not the same thing as a fact.
Self-consistency tells you whether repeated scores can be trusted to reflect a real change rather than noise. It is the right substitute if what you want is a stable baseline to compare your own submissions against over time. It says nothing about whether that baseline is calibrated to anything outside itself.
Predictive validity tells you whether the score forecasts a consequence you actually care about. It is the strongest evidence and the hardest to produce, which is exactly why it is cited least often and claimed most loosely.
A tool that only ever cites the first candidate, vaguely, without naming its panel, is making the weakest version of an accuracy claim while using the language of the strongest one.
Why this is not a criticism of any specific tool
It is tempting to read the ground-truth problem as evidence that AI rating tools are broken or dishonest. That is the wrong conclusion. The same problem exists for human judges, for peer review, for restaurant ratings, for almost anything with a subjective target - there is no ground truth for "how good is this restaurant" either, only aggregated opinion under some methodology. What differs between tools is not whether they have solved the ground-truth problem, since nobody has, but whether they are honest about which substitute they are using and how it was measured. A tool that reports its inter-rater agreement, states the size and composition of its evaluation panel, and does not overclaim what a score predicts is doing the best available version of this work. A tool that says "94% accurate" with no further detail is not lying, necessarily, but it is letting you assume a ground truth exists when what actually exists is a specific, unstated substitute that may or may not resemble what you had in mind. What that kind of number would even mean, mechanically, is a question worth asking before trusting the headline figure, on any tool, in any domain that scores something subjective.
Why averaging opinions does not create a fact
There is a tempting shortcut that looks like it solves the problem: average enough opinions together and the noise should cancel, leaving something closer to a real answer underneath. This works for questions that have a real answer hiding under noisy measurements - guess the number of jellybeans in a jar, average enough guesses, and the average genuinely converges on the true count, because the count is a fact and each guess is a noisy estimate of it. It does not work the same way for a subjective rating, because there is no fact under the opinions for the average to converge toward. Averaging a hundred people's taste does not produce "the true quality" of a photograph; it produces the average taste of those hundred people, which is a real and sometimes useful number, but it is a description of the panel, not a discovery about the photo. Swap the panel and the average moves, in a way that swapping jar-guessers would not move the actual jellybean count. This is worth stating plainly because it is easy to conflate the two situations - a big enough sample feels like it should manufacture objectivity, and for a fact-based question it does, but a subjective question does not have the underlying fact that makes averaging work that way.
The same problem shows up whenever a subjective judgement gets a number
It is useful to notice how often this exact structure recurs outside AI scoring specifically, because it clarifies that the ground-truth problem is a property of the question, not of the technology answering it. A wine critic's score, a film's rating aggregate, a restaurant's star average - each one converts a spread of individual opinions into a single figure, and each one faces the identical question this piece has been asking about AI: accurate compared to what? The film aggregate is compared to nothing except the panel of critics or users who contributed to it, whose composition shapes the number as much as the films do. A rating model is doing the same conversion, faster and with a machine standing in for the panel, and it inherits the same structural limitation rather than a new one specific to being automated. The interesting question was never whether a model can be made perfectly accurate at a task like this - nothing has been - but whether a given system is honest about which substitute for ground truth it is using and how that substitute was built.
What this means for reading a single score
None of this means a score is meaningless. It means a score is an opinion generated by a system trained to imitate other opinions, evaluated against one or more imperfect substitutes for a ground truth that does not exist in the way a temperature does. That is still useful information. Rate Cock reports scores as one input among several rather than the final word, and separates them by axis rather than blending everything into a single number that would hide which part of the judgement the score actually rests on. Reading a single number against the standard a human panel would set is a different exercise from reading it against a measurement, and penisrater.com covers what a total can and cannot support from the perspective of someone using the result rather than building the model. A human judge sidesteps some of this, since a person can tell you what they actually mean rather than compressing it into a scale with an unstated reference point, and what a commissioned human review contains is worth reading precisely because it makes the standard explicit rather than implicit. None of the three properties this piece is about - shape, size, or anything measured in centimetres - are what this problem is even about; if what you want is a physical figure rather than an opinion, that is a different question with a different, checkable kind of ground truth, and it is the one place in this whole discussion where an actual reference measurement exists to compare against.