How it works

Your 12-megapixel photo becomes a few hundred pixels wide

Almost every vision model resizes the input to a fixed small square, which throws away most of the detail you thought you were uploading.

By Updated 5 min readHow it works

Guides on How it works: Every score is a memory of somebody's opinion, An axis earns its place by being separately observable, What an image becomes before it is scored

Models downscale your photo because a vision network is built around a fixed input budget far smaller than a phone photo. Classic encoders take a small square such as 224 by 224 pixels; newer vision-language models accept more but still cap the long edge, so the extra resolution is discarded before any layer runs.

Why the size is fixed at all

A neural network's layers are built around a specific input shape. The number of weights in the early layers, the way the image gets carved into regions, the size of everything downstream - all of it is set once, when the architecture is designed, and does not flex per photo. Accepting a variable-sized input would mean either retraining the whole thing for every possible size, or engineering something more elaborate to handle variation, and neither buys enough to be worth it for the accuracy gain. So every photo, regardless of whether it started at 300 pixels wide or 12 megapixels, gets resampled down to the one size the network was built for. The Vision Transformer paper by Dosovitskiy and colleagues (2021) is typical: "all training is done on resolution 224," with fine-tuning at 384. Vision-language models have raised that ceiling without removing it: as of September 2026, Anthropic's vision documentation says Claude reads images in 28-by-28-pixel patches and downscales anything with a long edge over 2576 pixels on its newest models, or 1568 pixels on others.

This is not a corner someone cut carelessly. It reflects how these architectures were designed from the start - fixed input sizes go back to the earliest convolutional networks and are still standard in the transformer-based models that followed, for the same practical reasons.

What gets thrown away

A 224-pixel square holds a small fraction of the information in a modern phone photo. Fine texture - skin detail, small marks, subtle shading gradients - is largely gone by the time the resize finishes, because there simply are not enough pixels left to represent it. Small features anywhere in frame are affected worse than large ones, since shrinking treats every region of the photo the same way regardless of what is in it. Exactly how much detail sensitivity a given input resolution buys, and where the returns flatten out, is worth its own look in resolution and small-detail sensitivity.

This matters for a rating model specifically because some of what it is trying to read - surface condition, fine proportion - lives exactly in the resolution that gets discarded first. A higher-megapixel original does not buy you anything once it is past the model's fixed input size, because the extra pixels never reach the layers that would use them; the resize step happens before feature extraction, not after.

Letterboxing versus centre-crop

There are two common ways to get an arbitrary photo down to a fixed square, and they behave differently.

Letterboxing scales the whole photo down until it fits inside the target square on its longer side, then pads the remaining space - usually with black or grey - to fill it out. Nothing in the original frame is lost, but part of the square the model actually analyses is dead padding rather than photo, and the subject occupies a smaller share of the useful pixels than it would otherwise.

Centre-crop instead scales the photo so the shorter side fits the target, then cuts off whatever sticks out on the longer side to make it square. Nothing is padded, but anything near the edges of a tall or wide original - which, for most phone photos taken in portrait orientation, means the top and bottom - can be discarded entirely before the model ever sees it.

Which one a given tool uses is a design choice, not a universal standard, and it changes what the model is actually being handed. A subject framed off-centre survives letterboxing better; a subject framed centrally with a lot of empty space around it survives centre-crop better and gets scored on more of its actual pixels.

A human reviewer does not have this problem in the same way - a person looking at a full photo is not silently discarding the top and bottom before forming an opinion, which is one of several structural differences between the two approaches that Rate Penis lays out from the judging side.

What this is not

None of this is a case for a particular way of photographing anything - that is a practical question that belongs elsewhere, and Rate Cock's guide to getting an accurate result covers it properly rather than as an aside here. It is also a separate question from file compression: a heavily compressed JPEG loses information for different reasons than a resize does, and the compression side of the story is its own post.

What resizing does explain is one specific, common piece of confusion: the belief that a "better camera" or a higher-resolution upload should produce a more accurate reading of fine detail. Past the model's fixed input size, it mostly does not, because the extra resolution is discarded at the door. What survives the resize is closer to what the model's embedding actually encodes - shape, arrangement, gross texture - and that embedding, not the original file, is what everything after this step ever reads from.

Where a system reports several scores rather than one, this resize step affects some axes more than others; the shape and proportion reads survive a small square reasonably well, while the finer surface-condition read is the one that suffers most, which is a detail worth knowing if you are trying to interpret why a breakdown moves the way it does. It is also, mechanically, one reason the same photo can score differently depending on how it was cropped and framed before it ever left your hands - a variable penisrater.com has written up from the tool-selection side, and one Measure My Cock's data notes touch on from a different angle again.

Read next

Full archive