How it works
Pairwise comparison, and why many scorers learn that way
People are bad at absolute scores and good at picking between two, so many models are trained on comparisons and only later turned into a scale.
Guides on How it works: Every score is a memory of somebody's opinion, An axis earns its place by being separately observable, What an image becomes before it is scored
Many scoring models are trained on "which of these two is better?" rather than "what score is this?", because people judge pairs far more consistently than they assign absolute numbers. The number you eventually see is fitted afterwards, from strength estimates the comparisons produced.
Why comparisons are the more trustworthy signal
Absolute rating requires a person to hold a mental scale in their head and place an image on it consistently, which is a genuinely hard task - what does a seven mean, exactly, and does it mean the same thing today as it did an hour ago in the same session? A comparison sidesteps the question entirely. "Which of these two is better" needs no scale, no memory of past ratings, and no agreement on what a number means, only a preference in the moment. That is a much easier judgement to make consistently, which is why data collection for these models often looks like a grid of pairs rather than a spreadsheet of scores. The advantage has been measured for images specifically. Comparing four ways of collecting image-quality judgements, Mantiuk, Tomaszewska and Mantiuk (2012) concluded that forced-choice pairwise comparison "results in the smallest measurement variance", and was also the most time-efficient with a moderate number of conditions.
From preferences to a number
Collecting thousands of "A beats B" judgements does not by itself produce a score - it produces a pile of pairwise outcomes. Turning that pile into a scale is a separate, well-worn statistical step, the same one used to rank chess players or sports teams: each item gets an underlying strength value, and the values are chosen so that they predict the observed win rates as closely as possible. An item that beats almost everything it is compared against gets a high value; one that loses most of its comparisons gets a low one. A much-cited public example is Chatbot Arena, which ranks language models from crowdsourced head-to-head votes; Chiang and colleagues (2024) report over 240,000 votes collected with this pairwise approach. The training then fits a model to predict that underlying strength directly from the image, so at inference time it can assign a number to a photo it has never compared against anything.
The number that comes out the other end is not a rating in the everyday sense. It is a strength estimate, back-converted onto whatever range the interface wants to display, and the conversion is a design choice made after the fact rather than something the model itself was ever shown.
The consequence worth remembering
A model trained this way learns a score that is relative to the pool of images it was compared against during training, not to some fixed external standard. If that pool skewed toward a particular style of photo, the resulting scale reflects that pool's typical range, and a photo far outside it is being placed on a scale that was never calibrated against anything like it. This is a different failure from disagreement between individual labellers, since a comparison-trained model can be highly self-consistent and still be relative to a pool nobody chose deliberately - it is simply whatever images ended up in the training comparisons.
It also explains why pairwise-trained scorers tend to be better at ordering two things correctly than at producing a trustworthy absolute number in isolation - ranking is the task they were actually trained on, and the absolute score is a derived quantity layered on top. That is a genuinely useful property in its own right, separate from the ordinal-scale question of whether the gap between two scores means what it looks like it means. It is also the more general version of a question worth asking about any scoring tool: whether it is actually better at telling two photos apart than at giving either one a trustworthy absolute number, a distinction covered directly in ranking vs rating: what models do better.
Where this shows up in practice
A tool built this way tends to behave well when you compare two of your own photos against each other, since that is close to the native task, and less predictably when you look at a single score in isolation and try to interpret it against some external idea of what a nine should look like. Rate Cock reports separate axes rather than a single aggregate, and a per-axis comparison between two submissions is closer to what the underlying comparison-trained model is actually good at than reading either total on its own. Elo-style ranking as a live leaderboard feature, rather than as a training method, is a product decision covered on the site it belongs to rather than here.
The training data behind any of this ultimately comes from people making judgements, and who those people are and what instructions they followed shapes the pool the whole scale is relative to. Whether the eventual scale behaves like an interval scale, where the gap between an eight and a six means the same thing as the gap between a four and a two, is a separate question worth its own answer, covered in why an eight minus a six is not the same as a four minus a two. Human judges compare directly too, without needing a fitted model in between: Rate Penis covers what a person brings to a side-by-side comparison that a trained scale does not. Direct physical comparison has its own well-established method, and Measure My Cock is where that belongs rather than here. Reading what a resulting score does and does not support is worth doing deliberately, and penisrater.com has practical coverage of that from the user's side.