How it works

High accuracy on the training data, strange scores on yours

An overfitted scorer has learned its examples rather than the pattern, and it shows as confident, oddly specific reactions to irrelevant detail.

By 3 min readHow it works

Guides on How it works: Every score is a memory of somebody's opinion, An axis earns its place by being separately observable, What an image becomes before it is scored

Overfitting is what happens when a model gets very good at the exact photos it was trained on and worse at everything else, because it learned the specific examples rather than the underlying pattern those examples were supposed to teach. It is one of the oldest problems in machine learning, and scoring models are not immune to it.

How it is normally caught

Designers hold back a slice of labelled data, called a validation set, that the model never trains on. If the model scores well on the training data but noticeably worse on the validation set it has never seen, that gap is the signature of overfitting - the model has memorised rather than generalised. A well-run training process watches this gap and stops training, or adjusts the setup, before it grows too large. The capacity to memorise is not hypothetical: Zhang and colleagues (2017) showed that standard image networks "easily fit a random labeling of the training data" - labels that meant nothing.

What it looks like from the outside

A user cannot see a validation curve. What surfaces instead is a specific kind of strange behaviour: the score reacting confidently to something that should plainly be irrelevant. A particular background object, a specific piece of furniture, or an incidental detail that happened to correlate with certain scores in the training set can end up mattering to an overfitted model, because it learned that correlation as if it were signal rather than coincidence. Geirhos and colleagues (2020) call these learned coincidences shortcuts: "decision rules that perform well on standard benchmarks but fail to transfer" to real-world conditions. The model is not being random when this happens - it is being specific in a way that has nothing to do with the subject of the photo.

This differs from a model simply being uncertain, which tends to look like a hedged, middling score. Overfitting tends to look the opposite: an unusually confident number attached to a detail that has no business influencing it at all.

Why it is worth knowing about rather than worrying over

Overfitting is a training-time problem that a competently built pipeline catches with a validation set before the model ever ships, so it is not something a user can diagnose from a single result. What it explains, when a score seems to swing on something that plainly should not matter, is that the swing is not necessarily noise or bias in the usual sense - it can be a genuine artifact of memorisation rather than any property of your photo.

This is a distinct issue from a photo simply being unlike anything the model was trained on, which produces a different kind of unreliable score for a different reason and is worth its own separate explanation rather than folding in here.

Rate Cock reports six separate axes rather than one total, and an axis that swings on background detail while the others hold steady is the kind of thing this failure mode would actually look like if you happened to catch it. What the model inherited from its pretrained backbone before task-specific training even began is covered here, and the composition question - which kinds of photo the model saw plenty of versus barely any - is its own subject. A human reviewer does not memorise a training set in this sense at all, which is one of several structural differences Rate Penis covers about commissioning a person directly. None of this touches physical measurement, where Measure My Cock's method has no training set to overfit to in the first place. Reading a result for signs of exactly this kind of oddity, rather than taking any single number at face value, is covered practically at penisrater.com.

Read next

Full archive