How it works

A heatmap is evidence of sensitivity, not of reasoning

Saliency and attention visualisations show which pixels moved the output, which is useful and routinely over-read as an explanation.

By 4 min readHow it works

Guides on How it works: Every score is a memory of somebody's opinion, An axis earns its place by being separately observable, What an image becomes before it is scored

A heatmap laid over your photo looks like the model pointing at what it judged. What it is actually showing is narrower: which pixels, if changed, would have changed the number. That is a real and checkable fact about the model, and it is not the same thing as a reason.

How the heatmap gets made

The common technique, Grad-CAM, traces the gradient of the output score back through the network to a late convolutional layer, then overlays the result on the image as a coloured map. Selvaraju and colleagues described the method in 2017, and it has become the default way vision tools produce a "here's what I looked at" picture. Attention-based models offer a related but distinct signal: the weights a transformer assigns between image patches, which can also be rendered as a map. Both start from something real inside the network - a gradient or a weight - and turn it into a picture, which is where the trouble starts, because a picture reads as an explanation whether or not it is one.

The map is built after the score exists, not alongside a reasoning process the model ran through. Nothing about the pipeline in how a model turns a photo into a number involves a step where it "looks at" a region and decides. The vector that produced the score was computed once, over the whole encoded image, and the heatmap is a retrospective attribution laid on top of that.

What it does establish

A saliency map is a legitimate answer to one question: which pixels, perturbed, move the output. That is useful for debugging. If a model's top region for a body-rating task is consistently the corner of the frame, something is wrong with the model or the training data, and the map caught it. It is also useful for spotting when a model is responding to background or a stray object rather than the subject, which is a real failure mode worth catching before you trust anything else the model says.

Researchers use these maps to audit exactly this kind of thing, and where the highlighted region genuinely and reliably corresponds to the subject, that is mild positive evidence the model is not just keying off some artefact.

What it does not establish

A heatmap does not tell you the model's reasoning, because there isn't one to report - there is a function that maps a vector to a number, and the map estimates local sensitivity around that computation, not a chain of inference. Cynthia Rudin's 2019 argument against relying on post-hoc explanations for high-stakes decisions applies directly: a method built to explain a black box after the fact can produce a plausible-looking answer that does not reflect what the model actually did. Two different internal computations can produce similar-looking heatmaps, and a small change to the network can shift the highlighted region without moving the score at all. Adebayo and colleagues (2018) found something stronger: "some existing saliency methods are independent both of the model and of the data generating process." The map also says nothing about which learned feature - texture, edge density, contrast - is doing the work at that location, only that the location mattered.

This is a different problem from a written explanation under a score, which is text a language model generated after seeing the number, not derived from a heatmap at all. Both are attempts to make an opaque process legible, and both should be read as illustrations rather than proof.

Where this shows up in a rating tool

A tool that shows you a heatmap is giving you a real diagnostic signal, and it is worth reading as exactly that: which region was sensitive, not why it mattered or what it means. Rate Cock reports a breakdown across separate axes rather than a single heatmap, which sidesteps the interpretability problem differently - by naming which axis moved rather than trying to show where, an approach that tells you more without pretending to explain a mechanism that isn't there. The distinction between a checkable region and a checkable reason matters most in how a rating gets read once it leaves the model and reaches a person deciding how much to trust it. A human judge, by contrast, can be asked directly what moved their opinion and give you an actual answer in a sentence, which is a structural advantage covered on Rate Penis that no visualisation closes. None of this is a measurement question - a heatmap over a photo is not gesturing at anything to do with size or method, which is a separate property entirely, sitting closer to the encoding step than to any output.

Read next

Full archive