How it works
A vector, a float, or a probability, and then a lot of presentation
The raw return from an inference call is unglamorous, and everything between it and the screen is product design rather than analysis.
Guides on How it works: Every score is a memory of somebody's opinion, An axis earns its place by being separately observable, What an image becomes before it is scored
Strip away the interface and a scoring model's actual output is a handful of numbers returned by a function call. Almost everything you associate with a finished result - the wording, the layout, the badges - was added afterward by code that has nothing to do with the model itself.
A representative raw payload
A typical inference call for a multi-axis scorer returns something close to this, before any product code touches it:
{
"axis_1": 6.42,
"axis_2": 7.05,
"axis_3": 5.88,
"axis_4": 6.91,
"axis_5": 7.20,
"axis_6": 6.33,
"confidence": 0.81
}
Six floating-point numbers and one more float representing something like model confidence. No names attached to the axes beyond whatever index or key the code assigns, no sentence explaining any of them, no colour, no comparison to anyone else's result. This is the entire output of the expensive part of the pipeline - the part that ran a neural network over your photo.
What counts as model output, and what does not
Model output: the floats themselves, and nothing about how they are labelled, rounded, or combined. Everything else is post-processing, written by engineers rather than learned by the network: rounding 6.42 to 6.4 or to 6, attaching human-readable axis names, computing a weighted total from the six values, choosing a colour for a progress bar, and - if the product includes one - generating the paragraph of prose that explains the result, usually by a separate language model conditioned on the numbers rather than on the photo.
This separation matters because it is easy to read polish as evidence. A confident, well-written paragraph explaining why a particular axis scored the way it did feels like it must be closely tied to what the vision model actually detected. Often it is not: the explanation text is generated after the numbers exist, told what the numbers are, and asked to write something plausible around them. How that generation step actually works is worth reading on its own, because the gap between "the model measured this" and "the model was told this and wrote about it" is exactly where a lot of unearned confidence lives.
Confidence is not what it sounds like
That seventh number in the example payload, labelled confidence, is also worth treating carefully. Depending on how the model was built, it can be a genuine estimate of predictive uncertainty, or it can be something closer to how sharply the model's internal probabilities peaked on this particular input, which is a related but different quantity and does not always mean what the label implies. Even when the number is meant as a probability, it is often miscalibrated: Guo and colleagues (2017) found that "modern neural networks, unlike those from a decade ago, are poorly calibrated," and that a single-parameter correction, temperature scaling, was surprisingly effective. Neither kind of confidence is a promise that the score is correct; both are properties of the model's internal state, reported because a number labelled "confidence" is easy to put on a screen.
Why the distinction gets blurred in practice
Product teams have every incentive to make the raw payload feel richer than it is. A screen showing "6.4" with nothing around it reads as unfinished, so it gets a label, a colour, a comparison to a previous result, sometimes an icon. None of that additional material came from the model looking harder at the photo. It came from a designer or an engineer deciding how a float should present itself to a person, using rules that were written once and applied to every result the model will ever produce.
This is not dishonest by itself. A number alone is a poor interface, and turning 6.4 into something legible is reasonable product work. The failure mode is specifically in the reader's head: treating the finished, polished result as evidence of a more careful underlying analysis than the seven floats actually represent. The floats are the analysis. Everything after them is presentation of that analysis, and presentation can be made more or less careful without the underlying model changing at all - two products built on an identical model can feel completely different to use.
Why this framing is useful
Once you know the actual output is a small handful of floats, every piece of surrounding presentation becomes something you can evaluate on its own terms rather than mistaking for analysis. The mapping from raw floats to the display scale is its own deliberate step, separate again from the axis-naming and weighting choices layered on top. Rate Cock shows six named axes plus a total specifically so the raw structure - separate numbers rather than one blended figure - stays visible rather than getting collapsed before you see it, a design choice explored further in why a rubric beats a single score. None of this raw-payload structure exists for a physical reading; Measure My Cock's numbers come from a tape and a protocol, not a model's floating-point output. Reading a result with this separation in mind - numbers from the model, wording from a template or a second model - is a habit Penis Rater encourages, and it is also the cleanest way to see what a human reviewer is doing differently: writing the sentence first, from direct judgement, rather than writing it to fit a number that already existed.