Accuracy

Hide the condition from the person reading the results

When you know which photo was which, you read the scores to fit; blinding the comparison is the difference between a test and a story.

By 4 min readAccuracy

Guides on Accuracy: There is no true score to be accurate against, Change one thing, hold the rest, repeat, Exposure that suits one skin tone hides another

Comparing two photos through a scorer only becomes a test when whoever reads the results does not know which number belongs to which condition until after deciding what the results mean. That step - blinding - separates a controlled comparison from a story you tell yourself about numbers you already wanted to see.

Why knowing the answer changes the reading

Once you know that photo A was the one with better lighting, a small numerical gap in its favour reads as confirmation. A gap the other way reads as noise, an outlier, a fluke worth discounting. This is not a character flaw; it is how people process evidence when they already hold a belief. Clinical research has measured the size of it: a systematic review by Hróbjartsson and colleagues (2013) found that, in trials with subjective measurement-scale outcomes, "nonblinded assessors exaggerated the pooled effect size by 68%" compared with blinded assessors of the same outcome. The problem is specific to interpretation, not to the scorer itself - the model returned whatever it returned regardless of who read the result, but the read is where the bias enters. How to run the comparison in the first place is a separate question about controlling variables between two shots. Blinding is the discipline that comes after: controlling the person, not the photo.

What a blind protocol actually looks like

The mechanics are ordinary. Label each result with a code instead of a description before you look at the score, and only decode the labels at the end, after you have written down what you think you found. Randomise the order the images are scored in, and randomise the order you review results in, so a pattern in when you looked does not become a pattern in what you conclude. If you can get someone else to run the upload and hold the key, better still - you cannot leak your own expectation to yourself if you never had the pairing.

A useful stronger version is pre-registering the question. Write down, before scoring anything, what result would count as the hypothesis holding up and what would count as it failing. This closes off the common move of deciding after the fact that a marginal result actually does support the thing you wanted to show. It costs one sentence, written before you touch the scorer, and it is the sentence most casual comparisons skip.

Where this matters most

Blinding earns its keep most clearly when the comparison is close. If one condition scores a full point higher across ten repeats, blinding was nice to have but the result would likely have survived without it. It is the marginal cases - a few tenths of a point, within the range that ordinary retest variance produces on its own - where an unblinded reader will reliably see a pattern that is not there. Most informal "I tested this and it definitely made a difference" claims about lighting, angle, or framing are exactly these marginal, unblinded comparisons, and they are the least trustworthy claims in the whole category for that reason.

The limits of doing this yourself

A single person running a blind test on their own photos is still constrained in ways a formal study is not. Sample size is small, the labels are theirs to construct, and self-blinding only works as well as their own discipline in not peeking early. None of that makes the exercise pointless - it moves a claim from "felt true" to "held up under a check", which is a real improvement even at small scale. A panel of several blinded people rather than one is a further step up in rigour, and how that compares to relying on a single model's output is worth reading before assuming either one settles a question on its own. It does not turn an afternoon of testing into evidence a research paper would accept, and it should not be described as such.

What to compare it against

A rating tool that publishes methodology notes, or that a reviewer has already stress-tested this way, saves you the trouble. Rate Cock reports the same axis breakdown regardless of who is reading it, which is a smaller but related property: the presentation does not adapt itself to what the viewer already expects to see. Formal accuracy work on tools generally starts from the same instinct that motivates blinding here - Measure My Cock writes about method discipline for exactly this reason, applied to physical measurement rather than a scorer's output. On the human side, a commissioned review carries its own version of this problem, since a reviewer who knows the context of a photo is not blind to it either, and what a human review can and cannot control for is worth reading before treating a written verdict as more objective than a number. Reading a single tool's output honestly, without over-interpreting a score that has not been tested this way, is the practical skill Penis Rater focuses on for people evaluating a tool rather than running a study.

Blinding will not make a scorer more accurate. It only stops you from finding patterns in the noise that were never there, which is a smaller but far more achievable goal, and the one actually within a single person's control.

Read next

Full archive