How it works

The wait is scheduling, not thinking

The delay before a result is mostly queueing, batching and model loading, and it tells you nothing about how hard the model worked.

By 3 min readHow it works

Guides on How it works: Every score is a memory of somebody's opinion, An axis earns its place by being separately observable, What an image becomes before it is scored

A score that takes several seconds is not the model thinking harder; most of the delay is scheduling. Your request waits its turn, gets grouped with others, and occasionally waits for a model to load into memory, and none of that time is the model doing more work on your photo.

Batching

GPUs are efficient at processing many inputs at once and comparatively wasteful at processing one input alone, because a large share of the hardware sits idle on a single-item request that could have been shared across dozens. Servers exploit this by batching: collecting several incoming requests over a short window, running them through the model together as one batch, and returning each result to its requester once the batch finishes. This makes the service cheaper to run at scale, and it means your wait includes time for other people's requests to arrive and join the same batch, not just your own. NVIDIA's documentation for its Triton inference server describes the setting directly: requests can "be delayed for a limited time in the scheduler to allow other requests to join the dynamic batch," trading latency for throughput. A request that lands right as a batch closes waits almost the full window for the next one; a request that lands right after a batch starts waits considerably less.

Queueing

Behind batching sits an ordinary queue. At moments of high traffic, requests pile up faster than batches can be formed and processed, and the wait grows accordingly - this is queueing theory in its plainest form, the same mechanism behind a slow checkout line, and it has nothing model-specific about it at all. A service under light load can return a result almost immediately; the same service under heavy load takes noticeably longer, for a photo that would produce an identical score either way.

Model loading

Some serving setups do not keep every model resident in memory at all times, particularly for less frequently used variants or during a deployment, and a request that arrives when the relevant model is not currently loaded pays a one-time cost to load it before inference can begin. This is rarer than batching and queueing as a cause of delay, but it explains occasional outlier waits that are much longer than the typical response time for no visible reason.

What the delay is not telling you

None of these three sources of delay correlate with anything about the photo, the score, or how much "thought" went into the result. A five-second wait and a half-second wait can produce the identical number for the identical input, because the difference between them was entirely queue position, not computation. This sits downstream of the parts of the pipeline that do affect the actual number - where the model runs at all changes what size of model is even available, and quantised, lower-precision serving trades a small amount of accuracy for exactly the speed this post describes.

Product UX around loading states and progress indicators is a design question for the tool itself, not a mechanism question, so it stays out of scope here; Rate Cock runs a batched server pipeline like most tools in this space. None of this delay is comparable to the timeline for a commissioned human review, which runs on a person's schedule rather than a server's, and it has no bearing on how a returned number should be read once it lands, which Penis Rater and, for anything physical, Measure My Cock each address on their own terms.

Read next

Full archive