Accuracy
How to notice the ruler moving
A fixed reference set scored periodically is the only way to see whether a tool's scale has shifted, and it is cheap to keep.
A score from six months ago and a score from today look identical - a number between one and ten - and there is no visible marker attached to either one that says whether the scale underneath them is still the same scale. It might not be. Models get retrained, scoring heads get recalibrated, and the only way to notice is to build a check for it deliberately, because nothing in the interface will tell you.
Why drift happens without anyone announcing it
A tool's score is only ever a mapping from a model's internal output to a number a user sees, and that mapping is a design choice, not a law of nature. When a team retrains on fresh data, adjusts a threshold, or swaps a backbone model, the mapping can shift even if the interface, the scale, and the marketing copy stay exactly the same. Nothing about a 1-10 label guarantees that this month's 7 corresponds to what a 7 meant a year ago. This is a different question from why any given version changes in the first place - that post covers the causes; this one covers how you would notice from outside, without access to any changelog.
The method: a held reference set
The only workable check is to keep a small, fixed set of photos aside specifically for this purpose - never used for anything else, never reshot, never edited - and run that same set through the tool on a regular schedule. Because the inputs never change, any shift in the scores over time cannot be explained by anything about the photos. It has to be the tool.
A few points matter for this to actually work. The reference set needs to be genuinely fixed - the exact same files, not "the same photo taken again," since a retake reintroduces the ordinary variance covered elsewhere on this site and muddies the comparison. Run more than one reference image if you can, since a shift on a single image could still be ordinary noise rather than drift, and a handful of stable images moving together is a much stronger signal than one image alone. Keep the schedule consistent - monthly or quarterly is enough for most purposes - and log the date alongside each result, because "the scores felt different lately" is not evidence and a dated table is.
What a real drift shows up as
Genuine drift looks like every image in the reference set moving in a broadly similar direction across a single check-in, not one image jumping around while the others hold steady. A single image moving is more likely ordinary variance; several moving together, on the same schedule, in the same direction, is the signature of the underlying scale having changed rather than the input. If you see that pattern, the honest conclusion is not that your photos got better or worse - it is that the ruler moved, and any comparison you make against an older score from before that point is no longer apples to apples.
A worked example of a drift log
Suppose you keep four reference photos and score them every quarter, logging the result each time. January: 6.8, 7.1, 5.9, 8.0. April: 6.9, 7.0, 6.1, 7.9. July: 7.6, 7.9, 6.8, 8.6. The gap between January and April is small and inconsistent in direction - one photo up, one down, two roughly flat - which reads as ordinary run-to-run noise. The gap between April and July is a different shape entirely: all four numbers moved the same way, by a comparable amount, on the same check-in. That second pattern is what drift actually looks like from outside the tool - not a single number moving, but the whole set shifting together, which a spreadsheet with four rows and a handful of columns is enough to catch. Nothing about spotting this requires statistics beyond eyeballing whether the numbers moved together or independently.
Keeping this cheap
This does not need to be elaborate. Three or four photos, revisited on a calendar reminder every few months, with the results logged in a simple table, is enough to catch the kind of drift that actually matters for a personal record. The same logic applies to any instrument you rely on repeatedly - Measure My Cock's own writing on method covers the physical equivalent, where a tape and a technique are the thing being kept consistent rather than a model. A human reviewer has no version number to drift in the same sense, though an individual judge's standards can shift across their own sessions for entirely different, more personal reasons. Comparing scores across different apps over time runs into a related question, since any list ranking AI tools is itself a snapshot that can go stale as each tool underneath it drifts at its own pace. Whichever tool you track, including Rate Cock, the discipline is the same: keep the input fixed, and let the score be the only variable.