Accuracy

One way to test a subjective score is to see what it forecasts

A rating has predictive validity if it forecasts something outside itself, and for most AI scorers nobody has checked.

3 min readAccuracy

A score can be consistent, well-calibrated and still not mean anything. Predictive validity is the check that catches this: does the number forecast something that happens outside the scoring system itself, or does it only ever predict more scores?

What validity actually asks

Accuracy needs a ground truth to compare against, and for a subjective rating the usual candidate is other people's opinions. Predictive validity is a different, complementary test. Instead of asking "does this match what a person would say," it asks "does this forecast an outcome that exists independently of the rating."

In psychometrics, where the concept comes from, a test has predictive validity if it forecasts a real-world outcome measured later - an aptitude test predicting job performance, a screening tool predicting a diagnosis. The test does not need to match another opinion about the candidate. It needs to be right about something that actually happens.

What that would look like here

Applied to an AI rating tool, a predictive-validity claim would need an outside outcome to point at. Candidates exist in principle: does a score predict which photos get more engagement on a platform where that is tracked, does it predict how a panel of strangers ranks the same image, does it predict anything at all that was not itself derived from another model's opinion.

None of these are hard to design. A held-out test set, a predicted outcome, and a comparison once the outcome is known - that is the whole method. Reading a result honestly already asks a version of this question from the user's side - what a total does and does not support - and predictive validity is the same question asked with a stricter, outside-the-system standard. What is missing is not the method. It is any published instance of a consumer rating tool running it and reporting the result, positive or negative.

A related, slightly weaker version of the same test is retrospective rather than predictive: does the score correlate with an outcome that already happened, one the model was never shown during training. This is easier to run than a true forward-looking prediction, since it needs no waiting period, and it would still be informative. Its absence from public discussion of these tools is the same absence, for the same underlying reason - a negative result costs a product something to publish, and a positive one is hard to make airtight without a genuinely independent outcome to check against.

Why this is different from accuracy

A tool can score well on agreement-with-humans and still have no demonstrated predictive validity, because agreement is measured against opinions collected at the same time as the score, on the same kind of input the model was trained to match. Predictive validity asks about a separate moment and a separate kind of fact. It is the harder, more useful question, and it is also the one nobody in this space has strong incentive to run, because a validated benchmark can only ever be evidence against a tool that fails it, never marketing copy for one that has not tried.

Benchmarks that do exist for aesthetic scoring measure something related but narrower - general photographic preference, not predictive power on a specific downstream outcome - which is one reason the gap persists.

What to make of the absence

The honest position is not that AI scores are meaningless. It is that "predictive of what" is a question worth asking of any tool that presents a number as more than an opinion, and an unanswered version of that question is common enough across the category that it should not be held against any one product specifically. Rate Cock reports a breakdown by axis rather than a single validated figure, which is a more modest and more honest claim than "accurate" without a stated target. A commissioned human review sidesteps the whole question differently, since what a person's review is actually claiming is a judgement, not a forecast, and it is worth reading on its own terms rather than folding it into this comparison.

The gap between a consistent score and a validated one is also where a lot of quiet confusion sits: two tools can agree closely with each other and still both be unvalidated against anything outside themselves, which the size-measurement side of this problem runs into from a completely different angle - a number can be precise and still not be predicting the thing you actually care about.

Read next

Full archive