Accuracy
Why the same subject scores differently every time
The spread across a single afternoon is wider than almost anyone expects. Most of it is yours to control, which is the good news.
Ask any rating tool to score the same subject twice, photographed slightly differently, and you will get two different numbers. People generally interpret this as the tool being unreliable. It is more accurate to say the tool is reliably scoring two different inputs, and the inputs differ more than they look like they do.
Where the variance comes from
Roughly in order of how much they move the number:
Camera angle is the largest single source and it is not close. Shooting downward foreshortens along the axis of the lens, which changes the proportions the model reads. You are not photographing the same shape from two angles; you are handing the model two different shapes.
Lens distance is second, and it surprises people. Phone cameras are wide-angle, so anything near the lens is enlarged relative to anything further from it - within the same frame. At close range this distortion is severe. Two steps back and a crop afterwards produces geometry that is closer to true, at the cost of resolution nobody was scoring anyway.
Lighting mostly moves the surface and impression axes. A direct flash fires down the lens axis, blowing the surface into flat glare and erasing exactly the detail a texture axis exists to read. Side light at roughly forty-five degrees is the whole fix, and window light on an overcast day is the free version of it.
Crop, framing and background move the subjective axes and leave the physical ones roughly alone. This is worth knowing because it means presentation and proportion are separable in the result - if you can see the axes separately.
Sampling noise - the model itself returning slightly different numbers for a byte-identical input - is real but small, and in practice it is the least of your problems.
How to run a test that means something
Change one variable at a time. Three photos varying only the light, scored in one session, tells you what light does. Three photos varying everything tells you nothing.
Score them close together rather than across weeks. Tools get updated, and a difference across two months might be the model rather than the photo.
Read the breakdown, not the total. The whole reason a per-axis result is worth having is that it shows which judgement moved when you changed something, and a total that stayed the same can easily be hiding two axes that moved in opposite directions. Tools that publish the breakdown make this straightforward - Cock Rate exposes the six-axis chart on every result and on public leaderboard entries, which is enough to run a controlled comparison without any special access.
What the variance means for the number
It means your score is a range rather than a value, and the honest way to report it is as one. If four careful photographs land between 6.8 and 8.1, you do not have an 8.1. You have a range whose top end you reached once with good light.
None of this makes the tool useless. It makes it a measurement of a photograph, which is what it always was, and a photograph is a thing you can get better at taking.