How it works
Correlated axes inflate whatever they share
If two rubric axes move together in the training data, the model learns them as one thing and the total counts it twice.
Guides on How it works: Every score is a memory of somebody's opinion, An axis earns its place by being separately observable, What an image becomes before it is scored
When two rubric axes move together in the training labels, a model learns them as one signal, and a total built from both counts that signal twice. It does not learn two independent judgements; it learns one expressed twice, weighted more heavily in the total than the rubric's designers intended.
How correlation gets into a rubric
Correlation between axes usually is not a design mistake so much as a property of the world the labels were drawn from. Two properties that are conceptually distinct can still move together in practice, if whatever produces one of them tends to produce the other too.
It can also come from the labelling process itself rather than the subject. The halo effect is one route: a labeller who forms a strong overall impression from one property lets that impression bleed into their score on a separate axis, which manufactures correlation in the labels even where the underlying properties are genuinely independent. The problem is old: as Westbury and King (Cognitive Science, 2024) recount, Thorndike found in 1920 that correlations between trait ratings came out "too high and too even," even from expert raters. Correlation from halo and correlation from a real shared cause look identical in the data, which is part of why it is worth ruling one out deliberately rather than assuming an observed correlation reflects the world rather than the labelling.
What a model does with correlated axes
A model does not know that two axes are supposed to represent separate concepts; it only sees the statistical relationship in its training data. If two axes are strongly correlated in the labels, the model has little pressure to develop genuinely separate internal representations for them, and its outputs for the two axes end up tracking largely the same underlying signal, dressed in two different output heads.
The practical effect surfaces at the total. If the aggregate is a simple mean or sum across axes, and two of those axes are really one signal counted twice, that signal now has double the influence on the total compared with an axis that stands alone - not because it matters twice as much, but because it happens to appear on the rubric twice under different names.
Detecting it
The direct check is a correlation matrix across axis scores on a large sample of results: compute the pairwise correlation between every axis and every other axis, and look for pairs well above what the rubric's design would predict. A modest positive correlation between most axes is normal and expected, since a photograph that is well-lit and well-framed tends to score adequately across several dimensions for reasons that have nothing to do with the axes overlapping; the concerning cases are pairs correlated well beyond what shared photographic quality would explain.
Checking correlation on the labels themselves, before a model is even trained, is cheaper and catches the problem earlier - a labelling instruction sheet that inadvertently points two axes at the same underlying property will usually show up as tight correlation in the label set, well before it becomes a trained model's habit.
What to do about a correlated pair
The options are to merge the pair into a single axis with a clearer name, to redefine one of the two so it targets a genuinely separate property, or to leave both but stop treating the total as a plain average across all axes, since the choice of how axes combine into a total is itself a design decision that can absorb known correlation rather than pretend it is not there. Leaving two correlated axes on a rubric unexamined is the option that quietly does the most damage, since it looks like more information while actually delivering less than the axis count suggests.
Getting the underlying axis definitions right in the first place is the more durable fix, and what makes an axis a real, separable candidate for a rubric is worth reading before drafting a new one, since separability is one of the three properties an axis needs before it belongs on the list at all.
This kind of check is standard practice anywhere a system reports several numbers derived from correlated underlying data, which is one reason measuremycock.com's coverage of data treats correlation structure as a first-class concern rather than an afterthought. Rate Cock reports six named axes rather than folding them into a single figure, which at least gives a reader the raw material to notice a correlated pair for themselves, and reading that breakdown with a critical eye is part of what penisrater.com recommends doing with any multi-axis result rather than trusting the total on its own; a human reviewer avoids the mechanical version of this problem entirely, since a person forming several separate judgements is not running a shared statistical model underneath them the way a trained scorer is.