Accuracy
Refusing to answer is a valid output
A well-designed scorer would decline low-confidence cases; most return a number anyway, because a number is what the interface expects.
Guides on Accuracy: There is no true score to be accurate against, Change one thing, hold the rest, repeat, Exposure that suits one skin tone hides another
Yes, a well-designed scorer would sometimes decline to score, and the technique for doing so is well established. Almost no consumer tool exposes it, because the product has already committed to always returning a number, so uncertainty gets rounded into a score instead.
Abstention as a design choice
A model that estimates its own confidence can, in principle, decline to score an input that falls outside what it was trained on, or flag a result as low-confidence rather than presenting it flat. This is standard practice in fields where a wrong answer is costly - medical imaging triage, fraud detection - and it is rare in consumer rating tools, where the interface was designed around a guaranteed response. The trade is measurable: Geifman and El-Yaniv (2017) showed a selective classifier could guarantee "2% error in top-5 ImageNet classification" with probability 99.9%, at the price of answering only about 60% of test images.
The reason is not technical difficulty. Estimating uncertainty is a known technique, and the absence of it from most interfaces is a product decision, not a limitation of what is possible. A number with no confidence attached looks the same whether the model was firmly in familiar territory or improvising at the edge of its training distribution. The raw confidence cannot simply be trusted either: Guo et al. (ICML 2017) found modern neural networks poorly calibrated, and proposed temperature scaling, a one-parameter correction, before any threshold means what it says.
Why abstention rarely ships
A tool that sometimes says "cannot score this reliably" has a worse-looking product than one that always returns something, even if the honest version is more useful. Users came for a number, and a refusal - even a well-justified one - reads as the product failing rather than the product being careful. That asymmetry, not engineering cost, is most of why abstention stays rare.
This is a distinct question from a safety-related refusal, where a policy layer declines a request for reasons unrelated to how confident the model is in its own score; that is a different mechanism answering a different question.
What it would take
An interface that surfaced abstention would need to decide a threshold, communicate a decline without it reading as an error, and accept that some fraction of uploads return nothing useful. Rate Cock returns a score on every valid upload it accepts, which is the common design, not an unusual one - the tooling to do otherwise exists, the incentive to ship it mostly does not. A human review does not face this trade-off the same way, since a person can simply say "this photo is not good enough to judge" in a sentence rather than a threshold; that side of the process is worth reading separately. The same gap shows up in raw measurement: a device that will not commit to a number it cannot support is more useful than one that always prints a figure, which is a design question the measurement side of this space takes seriously in a way most scoring apps do not. Reading a score for what it actually supports, rather than taking the number at face value, is a habit worth building regardless of whether the tool ever tells you it was unsure.