How it works
Safety layers, and where they sit in the pipeline
A refusal is not the model failing to see; it is a policy layer declining, and it can be triggered by things unrelated to the photo.
Guides on How it works: Every score is a memory of somebody's opinion, An axis earns its place by being separately observable, What an image becomes before it is scored
An AI rater that refuses to score has usually not failed to see the photo. A policy layer - the system prompt, safety tuning or a classifier - has declined to let a response through, and it can trigger on things unrelated to the image. The apology you get instead of a number is that decision, not a visual judgement.
The system prompt as a standing instruction
Every call to a large language model, including a vision-language model doing scoring, is preceded by a system prompt: a block of instructions, invisible to the end user, that sets rules the model is meant to follow regardless of what the visible request asks for. A provider's own system prompt sits underneath whatever a tool builder adds - guidance on tone, on what to refuse, on how to handle ambiguous or sensitive requests. The tool's own instructions, including its scoring rubric, are layered on top of that. A refusal can originate from either layer, and from outside the tool there is no way to tell which one triggered.
Refusal as a policy decision, not a perception failure
When a model declines to score, it has not necessarily failed to process the image. In many cases it has processed it fully and produced an internal judgement that never reaches the user, because a classifier or an instruction upstream of the final answer decided the response should not be delivered as normal output. This is a categorically different event from the model being uncertain about a score - uncertainty still produces a number, however tentative; a refusal produces no number at all, because the policy layer intercepted the response before it could be one.
Refusals can be triggered by things that have nothing to do with image quality or content in the way a user would expect. Ambiguous phrasing in a caption field, an unrelated word in a filename, or content the classifier associates loosely with a sensitive category can each be enough, because these systems are typically tuned to err toward refusing borderline cases rather than risk letting through something that should have been blocked. That asymmetry is deliberate on the provider's part, and it means a refusal rate above zero is closer to expected behaviour than to a bug, even when a particular refusal looks unwarranted to the person who received it. Researchers have a name for the overshoot: Röttger and colleagues' XSTest (NAACL 2024) uses 250 clearly safe prompts to measure "exaggerated safety," where models refuse requests that only superficially resemble unsafe ones.
Why this is a separate layer from content gating on upload
Providers now treat a refusal as its own output type: OpenAI's Structured Outputs documentation notes that a safety refusal "does not necessarily follow the schema" a tool supplied and arrives in a separate refusal field instead, which is why a well-built tool can tell a declined request from a low score.
It is worth distinguishing this from the classifier that gates uploads before they ever reach a scoring model, which serves a different purpose and sits at a different point in the pipeline. The refusal covered here happens after the model has the image and the request, inside the model's own response process, triggered by the system prompt and any safety tuning behind it - not by a separate gate deciding whether to accept the upload at all. Quality gates for blur and exposure are a third, unrelated kind of rejection, based on measurable properties of the file rather than policy.
What a refusal does and does not tell you
A refusal is not feedback about the photo's quality, and reading it that way is a mistake worth naming directly: it tells you a policy layer declined this particular request under this particular phrasing, and nothing more specific than that. Rephrasing a caption or resubmitting sometimes resolves it, which is itself evidence the trigger was incidental to the image rather than a judgement about it - a genuine visual assessment does not change because a word in an unrelated field changed.
Rate Cock is explicit that a refusal carries no verdict about the subject, precisely because users otherwise tend to read a blocked response as a bad result rather than as a declined one, which is a different thing entirely. A human reviewer can decline a job too, but for reasons a person could actually explain if asked - a judge's own boundaries work nothing like a classifier's trigger conditions, and conflating the two misreads both. Measure My Cock's method coverage has no equivalent failure mode, since a tape measure has no policy layer standing between the reading and the number. Penis Rater's tool coverage is a reasonable place to check which scoring tools handle borderline cases gracefully, since refusal behaviour varies considerably between providers even when the underlying models are similar.