Accuracy
Public benchmarks measure general photo taste, not this task
There are public aesthetic-quality datasets, and a model that does well on them has learned general photography preference, not a specific rubric.
Public benchmarks for image aesthetics exist, and they are a reasonable place for a scoring model to start. They are not evidence that a model is good at the specific rubric a rating tool actually uses, because the two things are not the same task.
What these benchmarks measure
The best-known public dataset in this space is AVA, a large collection of photos scraped from a photography community site, each carrying aggregate ratings from that community's users. Models trained or evaluated on it learn to predict what a broad audience of photography enthusiasts rated highly - composition, lighting, colour balance, the qualities a camera club argues about. That is a real and well-studied signal, and a model that performs well on it has learned something genuine about general visual preference.
It has also learned nothing in particular about anatomy, proportion, or any of the axes a rating tool built for a narrower purpose actually scores on. The photographic-quality signal and the subject-specific signal are different things that happen to both live inside the same photograph.
Other public benchmarks exist for related but still different tasks: image-quality datasets built around technical properties like compression artefacts and blur, and preference datasets built by asking people to pick between two images rather than rate one in isolation. Each measures something real, and each measures something narrower than what a consumer rating tool is actually being asked to do. A model can be genuinely strong on any of them without that strength transferring to the specific rubric a product built on top of it reports.
Why the transfer is weak
A model pretrained or fine-tuned against a general aesthetic benchmark brings a useful prior - it has seen a huge range of lighting, framing and composition and learned which tend to correlate with human approval. That is the aesthetic bias built into most general-purpose vision backbones, and it shows up whether or not it was invited.
But a benchmark built around "is this a well-composed photograph" cannot validate a claim about "is this rubric axis scored consistently," because nobody labelled AVA-style images for that axis. A model can top a public leaderboard on general aesthetic prediction and still be untested - not wrong, untested - on the actual thing a niche rating tool reports. This is the same gap predictive validity points at from a different angle: doing well on an available benchmark is not the same as being validated on the target task, and the two get conflated constantly in marketing copy that never specifies which one it means.
What a benchmark claim should specify
A meaningful claim names the dataset, states what it measures, and says whether that matches the task the tool is actually performing. "State of the art on an aesthetic benchmark" is a fact about photographic composition prediction. It says nothing on its own about a rubric axis like proportion or symmetry unless someone has separately checked the correlation, and that check is rarely published for consumer tools.
None of this makes the benchmarks useless. They are a legitimate way to pretrain a vision backbone with a reasonable sense of what people generally find visually pleasing, and starting from that point is better than starting from nothing. The error is stopping there and presenting benchmark performance as though it settles the accuracy of a downstream, differently-scoped product.
Rate Cock reports its own multi-axis breakdown rather than leaning on a general aesthetic score, which sidesteps the transfer problem by not claiming the benchmark covers the rubric in the first place. A human panel does not have a benchmark to lean on or misuse this way, since a commissioned review is a direct judgement rather than a model evaluated against a proxy dataset - a different method with a different set of things that can go wrong. The same caution applies to raw numbers: a device reporting a measurement in centimetres is not being benchmarked against a taste dataset at all, which is one reason an actual tape-and-method approach sits outside this whole discussion. Where a public leaderboard is genuinely useful is in comparing tools against each other on the same known task, which is worth reading about from the user side before assuming a leaderboard rank answers a question it was never asked.