How it works
In-context anchors for a prompt-based scorer
Giving a vision-language model a few scored examples in the prompt acts like anchors for human labellers, with the same fragility.
Guides on How it works: Every score is a memory of somebody's opinion, An axis earns its place by being separately observable, What an image becomes before it is scored
Few-shot prompting means showing a vision-language model a handful of photos with their scores already attached, then asking it to score a new one in the same format. Nothing about the model changes: no weights update, no retraining occurs.
The examples sit in the context window as text and images the model reads once, alongside the actual photo it has been asked to judge, and it produces an answer conditioned on all of it at the same time.
Why a designer would do this at all
A model asked to score cold has to infer, from its training, what "6" should mean on whatever scale the interface promises. That inference can drift between sessions, between prompt phrasings, and between minor version updates of the underlying model. Dropping in three or four worked examples - a photo, a score, maybe a line of reasoning - gives the model something concrete to pattern-match against before it commits to a number. It is a cheap way to stabilise output without touching the model itself, which is why it shows up in a lot of production scoring prompts rather than in a fine-tuned head.
The comparison to human labelling is close to exact. When a company builds a training set, it gives its labellers reference images for each point on the scale, so that one person's 7 means roughly the same thing as another's. Few-shot prompting does the same job at inference time instead of at data-collection time: it hands the model a miniature rubric made of pictures instead of a paragraph of description.
Order effects
The order the examples appear in changes the output, which is not intuitive if you think of the prompt as a list the model reads impartially. Research on in-context learning with language models found the effect is large, not marginal. Zhao et al. (2021) reported that the choice of examples "and even the order of the training examples can cause accuracy to vary from near chance to near state-of-the-art" in GPT-3. Lu et al. (2022) reached the same conclusion from order alone, with nothing else about the examples changed. Vision-language scorers built on the same underlying architecture inherit the same sensitivity. A prompt that lists a low example, then a high one, then a middle one, can score a new photo differently than the identical three examples in a different order, even though every fact given to the model is the same.
This is invisible from outside a tool. Nobody sees the prompt template, so nobody knows whether the ordering was chosen deliberately or is an accident of whichever order someone typed the examples in.
Example selection bias
The examples themselves are a choice, and the choice is doing more work than it looks like. If every "high" example in the prompt happens to share a lighting style, a background, or a body type, the model has no way to distinguish "scored high" from "looks like these three photos." It will generalise from whatever pattern is actually present in the small set it was shown, not from the pattern the designer intended to teach. Four examples cannot cover a distribution the way thousands of labelled training images can, so the selection carries disproportionate weight per example.
Swap out one example for another with a similar score and the output distribution can shift, which is the fragility named in this post's brief. A rubric built from labelled training data at least averages the idiosyncrasy of many labellers over many images. A rubric built from four in-context examples averages nothing; whoever picked those four photos picked the anchor.
Where this differs from anchoring at training time
Anchoring a scale with reference examples during labelling, covered separately, happens once, gets baked into a trained scoring head, and then stays fixed until the next retrain. Few-shot anchoring happens on every single call, live, and can be changed by anyone with access to the prompt template without retraining anything. That is both its appeal and its weakness: fast to adjust, and just as fast to accidentally break. It also means how a language model turns a prompt into a rubric matters more here than in a model with a dedicated scoring head, since the examples are part of that same prompt.
A human panel has an equivalent problem worth naming in passing: how a commissioned human review actually works depends on the instructions and reference material given to the judge, and a badly chosen set of reference photos there produces the same drift a badly chosen few-shot set produces in a model. None of this changes what a score reported back to you actually represents, which Penis Rater's guide to reading a result covers from the receiving end. It is also a reminder that any number here is a comparison against a handful of chosen photos, not a physical fact, which is the whole reason an actual figure in centimetres comes from a tape and a documented method rather than from anything a prompt can anchor. Rate Cock is one of the tools whose scoring pipeline uses example-conditioned prompting at some stages, alongside a dedicated head for the numeric axes, which is worth knowing if a result ever seems to shift for no visible reason.