How it works
Reported accuracy is a number about a held-out set
An accuracy claim describes performance on a specific test set under a specific definition, and both are usually missing from the claim.
Guides on How it works: Every score is a memory of somebody's opinion, An axis earns its place by being separately observable, What an image becomes before it is scored
An accuracy figure is not a property of a model in the abstract. It is a property of a model measured against a specific set of examples, using a specific definition of what counts as correct, and changing either one changes the number.
The three-way split
Training a scorer normally starts with one large pool of labelled images, which gets divided into three parts before any training happens.
The training set is what the model actually learns from - its weights are adjusted repeatedly against these examples until its predictions stop improving. The validation set is held back during training and used to make decisions about the model itself: which architecture, which settings, when to stop. Because those decisions are made by watching performance on the validation set, the model has effectively been tuned to do well on it, even though it never trained on it directly.
The test set is held back from everything. It is not used to train the model and not used to choose between versions of it. It is used once, at the end, to produce the number that gets reported. If a team keeps going back to the test set to make more decisions, it stops being a clean measurement and starts being a second validation set with a misleading name.
What "accurate" is measured against
Accuracy needs a definition of correct before it means anything. For a classification task, correct usually means the predicted label matches an agreed-upon label. For a scoring task, it more often means the predicted number falls within some tolerance of a target number, and that tolerance is a choice someone made, not a natural fact.
Where does the target number come from. Usually from the same labelling process that produced the training labels, which means an accuracy claim about a scorer is really a claim about how well the model reproduces one particular labelling process, on one particular set of images, under one particular tolerance. Swap the labellers, the images, or the tolerance and the number moves, sometimes by a lot.
Questions worth asking of any accuracy claim
Accuracy against what ground truth. If the ground truth is itself a set of human labels, the claim is bounded by how much those humans agreed with each other in the first place, which is a separate number tools rarely publish alongside the headline figure. Even the reference labels on famous benchmarks are imperfect: Northcutt, Athalye and Mueller (2021) estimated that label errors make up at least 6% of the ImageNet validation set, and an average of at least 3.3% across ten widely used test sets.
Measured on what data. A model can perform very well on a test set that resembles its training data and much worse on the kind of photo a real user actually submits, which is a distribution question rather than an accuracy question, and the two get conflated constantly. The gap can be large even when the new data is collected carefully to match: when Recht and colleagues (2019) built a fresh ImageNet test set by closely following the original process, a broad range of models lost 11% to 14% accuracy on it.
Reported by whom, and checked by whom. A figure a company generated on its own held-out set, using its own definition of correct, is a claim, not an audit. An external benchmark, run the same way across multiple tools, is closer to a comparison, though even those carry their own assumptions about what counts as ground truth.
Why the number alone tells you little
A bare percentage without the test set, the tolerance, and the ground-truth source is a marketing sentence wearing a decimal point. None of that means the figure is false; it means it is incomplete, and incomplete in a way that happens to make every model look better than a fuller accounting would. A tool that shows the breakdown behind a result, the way Rate Cock does on every submission, at least gives you something to check its output against beyond a single trusted digit.
The harder question underneath - accurate compared to what - is its own subject, covered at what a ground-truth comparison actually requires, and a model's ceiling on any of this is set well before evaluation, by how consistently its labellers agreed with each other during the labelling stage.
None of this is specific to rating tools; it is how supervised learning is evaluated everywhere, which is why Measure My Cock's writeup on method and this piece land on the same caution from different directions. A tool that will not name its test set or its tolerance is not unusual, but it is worth reading that silence for what it is, in the same way penisrater.com treats a bare number as a starting point rather than an answer, and it is one more reason a human reviewer's judgement is evaluated on a different basis entirely, since there is no held-out test set for a person's opinion.