Topic
Accuracy
Where AI scoring is reliable, where it drifts, and how to tell which you are looking at.
-
Score compression, and its several causes
Model scores cluster near the centre for reasons that stack: hedging losses, thin data at the extremes and cautious labellers.
-
The top of the scale is structurally hard to reach
A ten needs training examples labelled ten, a head that can output it and a mapping that does not clip, and usually one of the three is missing.
-
The model rewards the camera, not only the subject
Sharper, cleaner images tend to score higher on axes that were meant to be about the subject, which quietly ranks people by phone.
-
Filters change texture, and texture is where the model looks
A filter smooths and reshapes; a model trained on unfiltered photos may score the smoothing as quality or as artefact, and it is hard to predict which.
-
An extreme score is more likely to be followed by a middling one
If your first score was unusually high, the next is likely lower for purely statistical reasons, and people read that as the tool changing.
-
Foundation models carry a taste in photography
Backbones pretrained on web images absorb what humans liked to photograph, and that taste leaks into scores meant to be about something else.
-
How much of the frame the subject fills is a variable
The proportion of frame the subject occupies changes what the resized input contains, and the model reads that as part of the image.
-
Product pressure pushes scales upward
Tools whose users prefer higher scores tend, over versions, to give higher scores, and the mechanism is ordinary incentives rather than better models.
-
Two kinds of soft, two different effects
A model reacts to motion blur and defocus differently because they destroy different frequencies, and quality gates catch one more reliably than the other.
-
Beyond the model's input size, extra pixels do nothing
Resolution above what the model resizes to is discarded, so a higher-megapixel upload changes the score only through the resampling path.
-
Under-represented bodies get thin-sample scores
When a body type was rare in training, the model's score for it comes from few examples and is more a guess than a read.
-
Computational photography is a hidden preprocessing step
Modern phones sharpen, smooth, tone-map and merge frames automatically, so the file the model receives is already an interpretation.
-
Tilt the camera and the geometry the model sees changes
Angle changes silhouette and foreshortening; the model reads the silhouette, so angle is a scored variable whether or not anyone intended it.
-
One way to test a subjective score is to see what it forecasts
A rating has predictive validity if it forecasts something outside itself, and for most AI scorers nobody has checked.
-
Small edits that move the number a lot
Vision models can be pushed by perturbations invisible to people, which says something about what the score was tracking all along.
-
No lens data reaches the score
Focal length, sensor size and distance are stripped or ignored before inference, so any judgement that would need them is being made without them.
-
Where the light comes from matters more than how much
Side light, top light and flash change contours and shadows the model reads as shape, so direction is a bigger variable than brightness.
-
Rare features pull the embedding somewhere odd
Anything unusual in frame lands the image in a sparse region of the model's space, and scores from sparse regions are unstable.
-
How to notice the ruler moving
A fixed reference set scored periodically is the only way to see whether a tool's scale has shifted, and it is cheap to keep.
-
Agreement is high at the ends and low in the middle
Model and human ratings tend to line up on clearly high and clearly low inputs and scatter across the middle, which is where most inputs sit.
-
The container matters less than the re-encoding
The format itself is mostly irrelevant once decoded, but converting between formats re-encodes pixels and that is where small shifts come from.
-
Sharpness, exposure and noise, not composition
When a scorer reports photo quality, it is mostly reading low-level statistics, and those correlate with the phone more than the photographer.
-
The room is in the embedding
Unless the pipeline masks the subject, everything in frame contributes to the embedding, and clutter or bedding can move a number.
-
Optimising for a model stops measuring what the model meant
Once people adjust inputs to raise a score, the score tracks the adjustments rather than the thing it was built to reflect.
-
Fatigue and order affect people, versioning affects models
A human's standards move over an afternoon; a model's move only when it is redeployed. Both are drift, on different clocks.
-
The interval around a number is the honest result
Ten repeats give a range; the range is the finding, and a single number quoted without it is an anecdote dressed as a measurement.
-
Relative judgements are easier than absolute ones
A scorer that struggles to give a stable absolute number can still order two photos correctly, and that is often the more useful output.
-
Colour temperature is an input variable
Auto white balance guesses the light source and recolours the scene; the guess changes apparent skin tone and texture, both of which models read.
-
Feed it noise and see what comes back
A quick sanity check for any scorer is what it does with an irrelevant image; a confident number there tells you how to weigh its other numbers.
-
Models pick up age-correlated cues and score them
Skin texture and tone correlate with age in training data, so a scorer can end up encoding age without anyone asking it to.
-
What you can test without seeing the training set
Paired inputs that differ in one attribute reveal bias in a closed model, and the method needs no access to anything but the upload box.
-
When the scale runs out before the differences do
At the ends of a scale the model can no longer distinguish, so two quite different inputs both get the same top or bottom number.
-
The second upload may not be scored at all
Some services cache by file hash, so re-uploading the same file returns the stored result and tells you nothing about consistency.
-
An engagement-trained model ranks what spreads
A scorer trained on likes learns the biases of an audience, including which bodies and photos get shown in the first place.
-
Public benchmarks measure general photo taste, not this task
There are public aesthetic-quality datasets, and a model that does well on them has learned general photography preference, not a specific rubric.
-
One score is an anecdote; the spread is the data
The number of repeats you need depends on how wide the spread is, and you cannot know the spread until you have taken several.
-
Context, intent and the parts of the frame that are not pixels
A person brings context a model has no channel for, and that difference is structural rather than a matter of model size.
-
Consistency is the model's real advantage, and its limit
A model gives the same answer to the same input, which humans cannot; that is a genuine strength that is routinely oversold as accuracy.
-
A frame grab is a compressed, motion-affected still
Frames pulled from video carry heavier compression, rolling-shutter effects and motion, so scores from frames are not comparable to scores from photos.
-
Depth from one photo is relative, not absolute
Models can estimate depth ordering from a single image, and that still leaves the overall scale unknown, which is the part people want.
-
Reliability and validity are different failures
A model can return the same number every time and still be measuring the wrong thing, and consistency is often mistaken for proof.
-
There is no true score to be accurate against
Accuracy requires a correct answer to compare with, and for a subjective rating the only candidate is other people's opinions.
-
Change one thing, hold the rest, repeat
A test that means something changes a single variable and repeats enough to see past the noise, and almost nobody does that before concluding.
-
Exposure that suits one skin tone hides another
Cameras and datasets are tuned around some skin tones more than others, and a model inherits both the exposure habits and the imbalance.
-
What AI genuinely cannot get from a photo
Not a limitation of current models. A property of projecting three dimensions onto two, which no amount of training data undoes.
-
Why the same subject scores differently every time
The spread across a single afternoon is wider than almost anyone expects. Most of it is yours to control, which is the good news.