How it works
Pixel-level masks, and what they let a scorer ignore
Some pipelines segment the subject from the background so scoring is not swayed by the room, which works until the mask is wrong.
Guides on How it works: Every score is a memory of somebody's opinion, An axis earns its place by being separately observable, What an image becomes before it is scored
Segmentation cuts the subject out of the background pixel by pixel, so the room stops influencing the score. Where a detector draws a rectangle, a segmentation model decides for every pixel whether it belongs to the subject, producing a mask rather than a box. Everything outside the mask can then be discarded or blacked out before scoring.
Why a box is not enough
A bounding box is a rectangle, and rectangles are a poor fit for an irregular subject against a cluttered background. Anything inside the box that is not the subject - bedding, a wall, a hand, fabric - is still there for the scoring model to read, because a crop only limits where it looks, not what it sees within that region. Background and context leak into scores precisely because most pipelines stop at cropping and never take this extra step. The effect is well measured in general object recognition: Xiao and colleagues (2020) found ImageNet models misclassified images despite correctly classified foregrounds "up to 87.5% of the time with adversarially chosen backgrounds." Segmentation closes that gap directly: once the mask exists, the pixels outside it can be replaced with a neutral fill or dropped before the vector representation is ever computed, and whatever was in the room stops being part of the input.
How it actually works
A segmentation model is trained the same way a detector is - on labelled examples, except the label here is a pixel-accurate outline rather than a box - and it outputs a probability per pixel that the pixel belongs to the subject. Thresholding that probability map produces the final binary mask. General-purpose segmenters have become much stronger recently: Meta's Segment Anything model (Kirillov et al., 2023) was trained on "over 1 billion masks on 11M licensed and privacy respecting images," which is why off-the-shelf masking is now a realistic option for small teams. This is more computationally expensive than detection, since the model has to make a decision at every pixel rather than a handful of coordinates, which is part of why not every pipeline bothers with it: cropping is cheaper and, for many purposes, close enough. The two are often used together rather than as alternatives - a detector finds a rough region first, cheaply, and segmentation then runs only inside that region rather than across the whole original frame, which keeps the more expensive step affordable. A related but distinct approach skips masks entirely and instead predicts a sparse set of specific points on the subject, an approach covered in keypoints and landmark models.
Pipelines that do add segmentation are usually doing it specifically because the scoring task is sensitive to exactly the kind of contextual noise a plain crop leaves in - texture and surface axes are more affected by stray background detail than an overall-impression axis would be, since those axes are reading fine local detail rather than a holistic impression.
Where masks fail
A mask is a model's guess, and guesses are wrong sometimes, in a small number of recognisable ways.
Under-segmentation. The mask misses part of the actual subject - an edge, a section under different lighting that reads as background - and that region gets zeroed out along with the room. The model then scores an incomplete subject without any signal that it was incomplete.
Over-segmentation. The mask includes material that is not the subject - shadow, a similarly toned object nearby - and that gets scored as though it were.
Boundary noise. Even a broadly correct mask has a ragged, uncertain edge at the transition, and that edge is exactly where fine surface detail lives, which is the detail some axes weight most heavily.
None of these failures throw an error. The pipeline proceeds with whatever mask it produced, and the resulting score looks exactly as confident as one produced from a clean mask.
What this changes about a result
A well-masked score isolates the subject the way the axis was designed to be read; a badly masked one silently substitutes a different, damaged input and reports a number anyway. This sits one step past what a detector crops before scoring even starts, and both are upstream of the vector a scoring head ultimately sees. None of it recovers a physical size - masking removes clutter, it does not add scale, and an actual figure in centimetres still requires a tape and a documented method rather than a cleaner crop. It is also a mechanical question rather than a judgement one, which is where this differs most from what a human reviewer naturally filters out without needing a mask - a person ignores the room by default, a model has to be told to. Reading a result with this step in mind is part of treating a score as informative rather than definitive, and Rate Cock is one of the tools that segments before scoring the texture and shape axes specifically, for exactly this reason.