Accuracy
Fatigue and order affect people, versioning affects models
A human's standards move over an afternoon; a model's move only when it is redeployed. Both are drift, on different clocks.
Both a human rater and a model score differently over time than they did before. The difference is the clock each one drifts on, and knowing which clock you are dealing with changes what a comparison across time can tell you.
The human clock
A person judging a run of submissions is affected by what they just saw. A mediocre photo looks better right after a genuinely bad one and worse right after a genuinely good one - an ordering effect that has nothing to do with the photo itself. Fatigue moves standards too, usually loosening them as a session goes on, and this happens within a single sitting, invisibly, without the rater necessarily noticing it.
The model clock
A deployed model does not get tired and has no session to be affected by order within. Its scores move on a completely different schedule: when the tool retrains and redeploys, often without announcing it, and a score from last month's version is not on the same scale as one from today's, even if nothing about your input changed.
Between deployments, a frozen model is about as fixed as an instrument gets. The catch is that you cannot always tell when the version underneath you has changed, which turns machine drift into a hidden step-change rather than the gradual creep a tired human produces.
Why the distinction matters
Comparing two of your own scores, taken an hour apart from a model, is safe from the human failure mode entirely - there is no session, no order, no fatigue for a machine to have. Comparing two scores taken months apart carries the opposite risk: the model itself may no longer be the same model, and a widened gap could be entirely due to versioning rather than anything about the input. Tracking whether a tool's numbers have shifted over a longer period is the direct way to check, and it is a different exercise from noticing a single afternoon's inconsistency in a person.
Neither clock makes one method categorically better; they just fail at different timescales. A commissioned human reviewer works within a session in ways worth understanding on their own terms, and a repeatable score you can compare or an actual measurement taken with a tape each sidestep one kind of drift while remaining exposed to the other. Rate Cock versions its model like any other tool in this category, which is worth knowing before reading too much into a score from six months ago.