How it works

One strong impression bleeds into every axis

Human raters let one good quality lift every score, and a model trained on their labels learns the same leak.

By Updated 4 min readHow it works

Guides on How it works: Every score is a memory of somebody's opinion, An axis earns its place by being separately observable, What an image becomes before it is scored

Yes, a model can inherit the halo effect: when labellers let one strong impression lift every axis, a model trained on their labels learns the same leak. The halo effect itself is well documented in psychology, and in image labelling it stops being a curiosity about people and becomes a property of the scorer.

How it enters a rubric

A labeller working through a multi-axis rubric is supposed to score each axis independently, looking only at the property that axis names. In practice, a labeller who forms a strong first impression from one salient property tends to let that impression colour their scores on the axes they look at next, even when those axes are, by design, describing something else entirely.

This happens without any intent to cheat the rubric. In a 2024 study in Royal Society Open Science, Gulati and colleagues had 2,748 participants rate faces in original and beauty-filtered versions; the same individuals received "statistically significantly higher ratings" for intelligence and trustworthiness when filtered, not just for attractiveness. Supervisors rating medical trainees show the same thing; reviewing that evidence, Sherbino and Norman (2017) concluded that raters are "capable of differentiating 1 domain and, occasionally, 2 domains." Labelling large datasets is almost always done under time pressure, which does not help. A related but distinct leak shows up when labels come from engagement signals rather than a rubric at all, since a popular image's other axes can get pulled up the same way a strong first impression pulls up a labeller's later scores - popularity bias in engagement-trained scorers covers that version directly.

What the model inherits

A model trained on halo-affected labels learns the halo along with everything else, because it has no way to distinguish "this axis genuinely correlates with that one" from "the labeller let one bleed into the other." Both produce the same statistical pattern in the training data: axes that move together more than the underlying properties actually do.

This produces a specific, checkable symptom: the model's axis outputs correlate with each other more than the real-world properties they are named after should, independent of any genuine shared cause. That correlation, however it arose, has the same effect on a total built from the axes, inflating whatever the halo attached itself to - but halo is a distinct mechanism from correlation arising out of the subject itself, and worth naming separately because the fix is different.

Why it is a different problem from correlated axes

Correlated axes can come from a real shared cause in the world, in which case there may be nothing wrong with the correlation itself, only with how the total handles it. Halo-driven correlation has no such innocent explanation. It is a labelling artefact from the start, which means the fix has to happen upstream, in how labels are collected, rather than downstream, in how axis scores are combined.

What reduces it

Blind, single-axis labelling is the most direct countermeasure: show a labeller the image and ask only about one axis, without showing them the others they will separately score, and without showing them any prior score on the same image. This breaks the chain that lets one impression carry into the next judgement, because there is no next judgement visible yet when the first one is made.

Per-axis labelling passes, where a whole dataset is labelled for one axis at a time by labellers who never see the other axes for that image, is the dataset-scale version of the same idea, and it is more expensive than a single pass covering every axis at once, which is exactly why many teams skip it. Building the model itself around one head per axis does not fix halo on its own, since the leak happened in the labels before the model ever trained, but it at least stops a shared architecture from adding a second, mechanical version of the same leakage on top of a label set that is already clean.

What this means for a result

A rubric where every axis moves together suspiciously closely on most submissions is showing a symptom consistent with halo somewhere upstream, whether in labelling or elsewhere, and it is worth treating that pattern as a reason to weight the individual axes over the total rather than the reverse.

Careful, blinded labelling is exactly the kind of practice worth checking for in how a training dataset was built, which measuremycock.com's coverage of data treats as central rather than incidental to producing anything trustworthy. Rate Cock reporting six separate named axes at least gives a reader the material to check for this pattern themselves across their own results, and penisrater.com's advice on reading a score is largely about learning to spot exactly this kind of suspicious agreement between numbers that are supposed to be independent. Halo is a known risk for human labellers producing training data, and it is also, separately, a known risk for a human reviewer judging a single submission in the moment rather than as part of a dataset - the mechanism is the same bias, showing up at two different points in two different processes.

Read next

Full archive