How it works

Most scorers look at a crop, not your photo

A detector usually finds the subject and crops around it before scoring, so what you framed and what the model rated can differ.

By 4 min readHow it works

Guides on How it works: Every score is a memory of somebody's opinion, An axis earns its place by being separately observable, What an image becomes before it is scored

The photo you upload is rarely the image a scoring model sees. A detection step first finds the subject, draws a box around it, and crops the frame to that box before anything gets scored.

The model that produces your number never saw your background, your framing choice, or most of what was in the original file.

Why detection runs first

Scoring models are trained on images that are already cropped to the subject, because that is how training data gets labelled - a person marks the region that matters, and the model learns from that region alone. Feeding the same model a full, unedited phone photo, with the subject occupying a fraction of the frame, gives it a very different input than the one it was trained on. A lightweight detector - usually a bounding-box model, cheap and fast compared to the scoring model itself - runs first specifically to close that gap: find the subject, crop tight, hand the crop to the model that actually knows what to do with it.

This is a standard architecture in applied computer vision generally, not something specific to this category of tool. Object detection followed by a specialised model on the cropped region is how a huge share of production vision systems work, because training one model to find things and a separate model to judge them is easier than training one model to do both at once. Finding the box is itself a learned task with its own research lineage: the Faster R-CNN detector (Ren, He, Girshick and Sun, 2015) uses a region proposal network to suggest candidate boxes before classifying and refining them, and its authors report 5 frames per second on a GPU with 300 proposals per image.

Padding, and why it matters

A raw bounding box is usually expanded by a fixed margin - padding - before the crop is taken, because a box drawn exactly to the edges of the subject discards context the scoring model may have used in training: a hint of surrounding skin, the transition at the edge, sometimes a hand. How much padding a pipeline adds is a design choice, and it is invisible from outside. Too little padding and edge-of-frame detail vanishes; too much and background starts leaking back into the crop, which is closer to no detection at all. Neither failure produces an error message. The pipeline runs, a score comes back, and nothing in the interface indicates which side of that trade-off the crop landed on.

Failure modes

Detection fails in a small number of predictable ways.

Wrong crop. The detector locks onto the wrong region entirely - a shadow, a fold of fabric, a similarly shaped object in frame - and scores something that is not the intended subject at all. This is rarer than the other failures but produces the strangest results when it happens, because nothing about the returned number looks obviously wrong.

Partial subject. The box is drawn too tight, or the subject extends outside the frame the camera captured, and the crop cuts off part of what should have been scored. A model given half a subject scores what it was given; it has no way to signal that the other half existed and was missing.

No detection at all. Some pipelines fall back to scoring the full uncropped frame when the detector returns nothing, quietly, rather than surfacing a rejection. That fallback path behaves like a different tool wearing the same interface, and its scores are not comparable to the normal path's.

None of these are photography advice - what to do about framing belongs to the product itself, not to a piece explaining the mechanism, and Rate Cock's own guide to taking a photo that scores accurately covers that ground properly.

What it means for reading a score

A score is a judgement of the crop, not of the photo, which is worth holding onto before treating any single number as a verdict on the photo you thought you took - how to read a result without over-trusting it is covered from the user's side elsewhere. Two uploads that look identical to you but produce slightly different bounding boxes - because the detector is not perfectly deterministic near an edge - can carry that difference all the way through to the final number, and it will look like ordinary variance rather than a detection artefact. This sits underneath most of the spread you see across repeated uploads of the same subject, and it is one layer earlier than the encoding step described for what a model sees once cropping is done.

Segmentation is the more aggressive cousin of this step - a mask instead of a box, cutting the subject out of the background entirely rather than just framing it - and how that works, and where masks fail, is its own subject. The equivalent step for a human judge does not exist in the same form: a person reads the whole photo at once and applies judgement about what to ignore, which is one reason what a commissioned human review actually contains differs structurally from an automated score rather than just being slower. Framing also interacts with the size question directly - a crop drawn around the wrong reference points is a bigger problem for an actual measurement taken with a tape and a documented method than for a rubric score, since a tape measurement does not depend on where a detector decided the edges were.

Read next

Full archive