How it works
Structured output, and what breaks when it fails
Numbers extracted from free text are unreliable, so tools constrain the model to a schema, and the constraint has side effects on the values.
Guides on How it works: Every score is a memory of somebody's opinion, An axis earns its place by being separately observable, What an image becomes before it is scored
Tools force JSON because a language model asked to score a photo returns text by default, and pulling a number out of "around a 7, maybe leaning 7.5" is unreliable. Structured output removes that parsing step by making the model answer in a fixed, machine-readable shape from the start.
What structured output actually is
Most providers now offer a mode where the model's output is constrained to match a schema - typically JSON with named fields, like {"score": 7, "confidence": "medium"}.
OpenAI's Structured Outputs documentation draws the line clearly: the older JSON mode only guarantees valid JSON, while Structured Outputs guarantees the response adheres to the supplied schema.
Under the hood, this usually works by restricting which tokens the model is allowed to generate at each step, so it becomes structurally impossible for it to produce a string that does not fit the shape.
The model still decides what number to put in the field; it just cannot wander off into a paragraph of hedging instead of answering.
Before this existed, tools had to parse free text with regular expressions or a second, smaller model, hunting for the first digit that looked like a score. That approach breaks in ordinary ways: a model that writes "on a scale where 10 is exceptional" mentions a 10 that is not the score, or one that gives a range instead of a single figure leaves the parser to guess which end to take.
The side effect nobody mentions
Constraining the output format is not free. Forcing a model to answer immediately in a fixed schema removes the space it would otherwise use to reason before answering, unless the schema explicitly includes a reasoning field ahead of the score - and whether that reasoning is trustworthy even when present is a separate question. Research backs this up: Tam and colleagues (2024) reported "a significant decline in LLMs reasoning abilities under format restrictions," with stricter formats doing more damage on reasoning tasks. Their tests were on reasoning benchmarks rather than image scoring, so the size of any shift for a rating tool is unmeasured; the honest claim is that the format is not a neutral wrapper around an unchanged answer.
What it means for a user
None of this is visible from outside a tool. What you can infer is that a service returning results instantly and consistently in the same layout, run after run, is very likely using structured output rather than parsing free text - the alternative shows up as occasional garbled results or missing fields, which structured output is specifically built to prevent.
Rate Cock returns its per-axis breakdown as structured data for this reason, which is also what makes a clean multi-axis chart possible rather than a wall of text a person has to read for the number. Whether the wording asking for that structure is itself doing extra work is a related question about prompts generally. A written measurement guide has no equivalent parsing step to worry about, since a person reading a tape writes the figure down directly. Penis Rater covers the practical side of judging which tools return clean, comparable results. A human reviewer's output has the opposite problem in miniature - there is no schema to enforce on a person's written verdict, only their own consistency.