Accuracy

How a scale is pinned after training

Calibration adjusts raw outputs so the scale behaves consistently, and it is a step many tools skip or redo without saying.

By 4 min readAccuracy

Guides on Accuracy: There is no true score to be accurate against, Change one thing, hold the rest, repeat, Exposure that suits one skin tone hides another

Calibrating a scoring model means fitting a separate correction, after training, so that a stated number - a 7, specifically - means the same standard across the whole scale. A freshly trained model's raw output might cluster between 5.8 and 7.2 or rarely touch the ends, and calibration maps it back.

What "uncalibrated" actually looks like

An uncalibrated model is not necessarily wrong in its ranking - it can still correctly tell that photo A is better than photo B. What it gets wrong is the mapping from its internal confidence to the number it reports. A common failure is overconfidence: the raw scores cluster near the extremes more than the underlying judgement warrants, so images that should read as solid middle-of-the-road results get pushed toward a 2 or a 9. This is well documented for classifiers: Guo, Pleiss, Sun and Weinberger (2017) found modern deep networks poorly calibrated, with depth, width, weight decay and batch normalisation among the factors. A different, more common failure in scoring systems is the opposite - everything hedges toward the centre, because during training the model was penalised more for confident wrong answers than for cautious average ones, so it learned caution as a strategy rather than as an honest reflection of the input. Both are calibration problems, not ranking problems, and they can exist in a model whose relative ordering of any two photos is basically sound.

How it gets fixed

The general approach is to take the model's raw outputs on a held-out set of examples with known reference labels, and fit a simple correction function that maps raw output to a corrected score - stretching the middle, compressing the extremes, or shifting the whole scale, depending on what the mismatch looks like. This correction function is fit after the main model is frozen, which is what makes calibration a distinct step rather than part of ordinary training. Guo and colleagues found the simplest version, temperature scaling - "a single-parameter variant of Platt Scaling" - to be "surprisingly effective." It does not change what the model perceives. It changes how that perception gets translated into the number you see, which is a narrower and more mechanical job than it sounds, and one that is easy to skip under time pressure since a model without it will still produce plausible-looking scores.

Why it needs redoing, quietly

Calibration is fit against a specific reference set at a specific point in time. When a model is retrained - new data, an architecture tweak, a rebalanced training set - the raw output distribution can shift even if the underlying judgement barely changes, and the old calibration no longer applies cleanly. A tool that retrains and redeploys silently may or may not have recalibrated against the new distribution, and there is usually no way to tell from outside which happened. This is one of the mechanisms behind scores drifting across versions of the same tool even when nothing about the stated methodology changed - the scale itself moved a little, quietly, as a side effect of maintenance work that had nothing to do with rating quality.

Calibration is not the same thing as confidence

It is worth separating this from a related but different concept. A softmax output that looks like a confidence percentage is a per-prediction number describing how sure the model appears to be about one specific input. Calibration is a property of the whole scale, fit once against a reference set, describing whether the numbers as a system behave consistently. A model can be well calibrated on average and still produce an overconfident-looking softmax value on an unusual individual photo, because calibration corrects the aggregate mapping, not every single output. The two get confused often enough that it is worth stating plainly: one is about a scale, the other is about a single answer.

What a calibrated scale buys you

The practical payoff of calibration is comparability. A 7 from a well-calibrated system means roughly the same thing whether it was given to the tenth photo scored that day or the ten-thousandth, and roughly the same thing across the scale's own history, within the limits of whatever reference set the calibration was fit against. Without it, the number is still a number, and it can still be used to compare two photos scored close together in time by the same model version - but any claim that the number means something specific, on an absolute scale, rests entirely on this unglamorous fitting step having been done and kept current. Rate Cock calibrates each of its six axes separately rather than the blended total, on the reasoning that a size-adjacent axis and a taste-adjacent axis do not drift the same way and should not share one correction curve. Reading a single number against your own history, rather than as an absolute claim, is the safer habit regardless, and it is the angle penisrater.com takes on interpreting a total. A human panel has an informal version of the same problem - judges anchor to whatever they scored most recently in a session - and how a commissioned review keeps its own scale steady is a related question with a very different mechanism behind it. None of this touches a measured figure in centimetres, which does not need calibrating against a reference set because the reference is a physical object rather than a fitted curve.

Read next

Full archive