Accuracy
A ruler in frame helps a person, not a scorer
A model trained to score has no step that reads a reference object; unless it was built to, the ruler in the photo is just more background.
Guides on Accuracy: There is no true score to be accurate against, Change one thing, hold the rest, repeat, Exposure that suits one skin tone hides another
A scoring model does not use the coin, card or tape measure you put in frame. A general-purpose scorer has no stage that finds a reference object, measures it in pixels and converts anything else into a real-world unit. The coin is just more pixels, encoded like the background it sits on.
What using a reference object actually requires
Reading a reference object correctly is a small pipeline in its own right, not a side effect of looking at a photo. The system needs to find the object - detect that this particular patch of pixels is a coin and not a button, a ring or a smear of light. It needs to know that object's true size in advance, from a lookup table of coin diameters or card dimensions. It needs to measure the object in the image accurately, which means correcting for the angle it was photographed at, since a coin tilted away from the camera projects as an ellipse rather than a circle and a naive pixel measurement will be wrong. Only after all three steps does a ratio exist that can be applied to anything else in the frame, and even then it only holds for objects at the same distance from the lens as the reference, because perspective means nearer things are bigger in the image regardless of their real size.
None of that happens by accident. It is a deliberate computer-vision pipeline: detection, a size prior, geometric correction, and a distance assumption stated plainly enough to be wrong about. A model trained end to end to output a rating score was never shown that task and never asked to solve it.
Why the standard scoring pipeline skips it
A scoring model of the kind used across most rating tools is trained on pairs of photographs and human-assigned scores. Nowhere in that training does an example say "this coin is 24 millimetres wide, use it." The model's encoder turns the whole image into a vector describing overall appearance, and the scoring head maps that vector onto a number. A coin in frame changes the vector, because it changes what the image contains, but it changes it the way any object changes a photograph's composition and lighting - not the way a calibration reference changes a measurement. The same is true of an unplanned object already in the shot rather than one placed deliberately - tattoos, jewellery and other unexpected features get folded into the vector the same way, with no special handling for what they are. There is no arithmetic step downstream that could use it even if the model "noticed" it clearly.
This is the same gap covered from the size side elsewhere: a photograph has no scale information built in, and a reference object is exactly the kind of context cue that could fix that in principle, if anything in the pipeline were built to read it. For a general scorer, nothing is. Research on depth from a single image names the same wall: Eigen, Puhrsch and Fergus (2014) called the task inherently ambiguous, with a large source of uncertainty coming from the overall scale.
What a purpose-built pipeline would need
A tool that wanted to genuinely use a reference object would need to add the missing stages on purpose: an object detector trained specifically to recognise a small set of known references, a lookup of their true dimensions, a perspective correction step, and a rule for rejecting frames where the reference is unclear, occluded or shot at the wrong angle to trust. That is a meaningfully different system from a rating model, closer to the kind of measurement pipeline that measuremycock.com's method is built around, where the entire point is a repeatable procedure rather than a score. Building that well is possible. Bolting it onto a scoring model that was never designed for it is not, which is why no rating tool worth using claims to read a coin from an arbitrary upload.
Reading a result with this in mind
If a tool's interface encourages placing an object in frame "for scale," treat that as a photography convention rather than a technical one - it may help a person glancing at the result form an impression, the way a human reviewer brings judgement a model cannot, but it is not doing arithmetic on the model's behalf. Rate Cock scores what is in the frame without pretending a coin changes that math, which is the honest version of this: the size axis is inference from the whole image, not a measurement calibrated against an object you placed there. Comparing that kind of inferred axis against what other tools report is its own exercise, and it starts from the same fact - unless a tool says otherwise and shows its pipeline, assume the reference object is decoration. The gap between what a coin could tell a purpose-built system and what it tells a general one is the gap between measurement and inference that runs through every result these tools produce.