Topic
How it works
What the model is doing between the upload and the number that comes back.
-
Token frequency leaks into the number
A language model picking a number is picking a token, and some numbers are simply more common in text, which shows up in the scores.
-
Reported accuracy is a number about a held-out set
An accuracy claim describes performance on a specific test set under a specific definition, and both are usually missing from the claim.
-
AI penis rating, explained
The number looks like a measurement. It is an inference from one photograph, and knowing the difference tells you how much weight to give it.
-
Most scorers look at a crop, not your photo
A detector usually finds the subject and crops around it before scoring, so what you framed and what the model rated can differ.
-
Pixel-level masks, and what they let a scorer ignore
Some pipelines segment the subject from the background so scoring is not swayed by the room, which works until the mask is wrong.
-
Safety layers, and where they sit in the pipeline
A refusal is not the model failing to see; it is a policy layer declining, and it can be triggered by things unrelated to the photo.
-
Your 12-megapixel photo becomes a few hundred pixels wide
Almost every vision model resizes the input to a fixed small square, which throws away most of the detail you thought you were uploading.
-
One strong impression bleeds into every axis
Human raters let one good quality lift every score, and a model trained on their labels learns the same leak.
-
Where your photo lands, and what "near" means
The model places every image in a high-dimensional space; a score is a function of where yours lands, not of anything measured on it.
-
Spatial detail is averaged away on purpose
Pooling layers trade position for robustness, which is why a model can recognise a shape and still be poor at judging its proportion.
-
A scale without anchors drifts
Rubric points are defined by example images given to labellers, and the anchors chosen become the model's fixed points.
-
What was over-represented in training becomes 'normal
A model's sense of typical is the composition of its training set, and anything under-represented there is scored from a thin sample.
-
The total is a policy, not a measurement
Combining several axis scores into one requires choosing weights, and every choice encodes an opinion about what matters.
-
A 92% is a normalised number, not a belief
When a tool shows a confidence percentage, it is usually a softmax output, which sums to one by construction and says less than it appears to.
-
Some axes exist because users expect them
Rubrics sometimes carry axes the model cannot reliably read, kept because leaving them out reads as incomplete.
-
Features are not the things you would name
The model's features are learned patterns of contrast and texture, not the anatomical parts a person would list, and the mismatch explains a lot.
-
Two architectures, and what differs for a rating task
Convolutional nets and transformers reach a score by different routes, and the difference shows up mostly in what each is sensitive to.
-
Tying a claim to a region of the image
Grounded models tie each statement to a region, which makes a score checkable in a way a bare number never is.
-
Input normalisation, and why it is not colour correction
Every input is shifted and scaled by fixed constants from the training set, which is a technicality with one visible consequence for unusual photos.
-
How a model decides two images are alike
Most "this resembles that" judgements in image models are one dot product, and knowing what it ignores explains some strange results.
-
When the scorer is a language model looking at a picture
A vision-language model produces a score as text, after producing other text, and that ordering changes what the number depends on.
-
When the exact same bytes score differently
A frozen model given identical bytes should return identical output, but GPU maths, batching and preprocessing randomness can each break that.
-
Sampling randomness, and why it should be off for scoring
Language models sample their output, and unless the temperature is zero, the same photo and prompt can return a different number each run.
-
Blur, exposure and resolution checks come first
Tools often run cheap checks for blur, darkness and size before the expensive model, and a rejection there says nothing about the subject.
-
Wrong labels do not cancel out; they smear
Noisy labels do not average away cleanly; they flatten the model's confidence and pull unusual inputs toward the mean.
-
Same photo, seen before, without storing it
A perceptual hash lets a service recognise a near-identical image from a short fingerprint, which serves dedup, caching and moderation.
-
What 8-bit weights do to a number out of ten
Serving a model in lower precision saves cost and shifts outputs by small amounts, which is enough to flip a rounded score.
-
Past a point, more axes means more noise
Each added axis needs its own reliable labels; beyond a handful the labels get thin and the axes start repeating each other.
-
The colour pipeline the model never told you about
Between your camera's file and the model's tensor sit colour-space conversions that can shift tone enough to matter for a skin-tone-sensitive task.
-
The last layer decides what kind of number you get
A model can produce a score by predicting a continuous value or by picking a bucket, and the two behave differently at the extremes.
-
What changes when the rater is a model
Rubrics written for human judges rely on context and taste; a model rubric has to be reduced to what shows up in pixels.
-
The same architecture behaves differently depending on where it lives
On-device inference means smaller models and no upload; server inference means bigger models and a file that leaves your phone.
-
Flips, crops and colour jitter as a list of things that should not matter
Augmentation is how designers tell a model which variations to be blind to, and a rating model's blind spots are exactly that list.
-
Most rating models start from something general
A model fine-tuned from a general vision backbone inherits that backbone's habits, including what it was never shown.
-
A rubric written in English, read by a model
In prompt-based scoring the rubric is a paragraph of instructions, so a wording change can move every score without any retraining.
-
An axis has to be written down as an instruction
Before a model can score an axis, someone wrote a sentence telling labellers what to look at, and that sentence is the axis.
-
The label stayed; the thing it measures moved
After retraining with new labels, an axis can keep its name and change what it responds to, which nobody notices from outside.
-
When the score is 'how well does this match a sentence
Contrastive image-text models can rate a photo by measuring its similarity to a written description, which is elegant and easy to mislead.
-
The photo is cut into tiles, and the tiles vote
A vision transformer splits the image into a grid of patches and lets every patch weigh every other, which is why context matters more than you expect.
-
A second model decides whether the first one runs
Adult-content classifiers gate many pipelines, and their false positives and negatives shape which uploads ever reach the scorer.
-
Few training examples at the ends means timid predictions there
If the training data had few very low or very high examples, the model has little to go on at the extremes and retreats toward the centre.
-
How a designer lands on a number of axes
Six axes is not magic; it is roughly where coverage of what people notice meets the point where labellers can still be consistent.
-
Hallucination in image description, and how it reaches the score
Vision-language models sometimes narrate details absent from the photo, and if the score is derived from that narration it inherits the error.
-
Landmark detection, and what it adds to a rubric
Some pipelines locate landmarks first and derive geometry-like features from them, which is more legible and still not measurement.
-
A score is a rank, not a quantity
Most rating scales are ordinal, and treating the gap between two scores as a measured amount is where interpretation goes wrong.
-
A model trained on engagement learns what got clicks
Some scorers learn from likes and views rather than judged labels, and the result is a model of attention rather than of quality.
-
Retraining changes the ruler you did not have
Tools retrain and redeploy, often silently, and a score from an earlier version is not on the same scale as one from today.
-
Every score is a memory of somebody's opinion
A scoring model learns from labelled examples, and the identity, number and instructions of the labellers set the ceiling on everything after.
-
An axis earns its place by being separately observable
A good rubric axis is something a model can be trained to read on its own, and most proposed axes fail that test.
-
What an image becomes before it is scored
A photo is turned into a list of numbers long before any score exists, and that list is not a picture in any sense you would recognise.
-
Why a rubric beats a single score
One number is a summary of things that do not correlate. Splitting it apart is the difference between a result you can act on and a result you can only feel.
-
What an image model is actually doing when it scores you
There is no ruler anywhere in the pipeline. Understanding what replaces it explains almost every surprising result people get.