Accuracy
Filters change texture, and texture is where the model looks
A filter smooths and reshapes; a model trained on unfiltered photos may score the smoothing as quality or as artefact, and it is hard to predict which.
Guides on Accuracy: There is no true score to be accurate against, Change one thing, hold the rest, repeat, Exposure that suits one skin tone hides another
A beauty filter can move a model's score, but in no predictable direction - the same smoothing might read as polish or as something synthetic. A filter deliberately smooths skin, evens tone and sometimes reshapes contour on top of the phone's own processing, and where that lands relative to the model's training data decides the result.
What a filter actually does to the pixels
Most filters work in one of two ways. Smoothing filters reduce local contrast and high-frequency detail in skin-toned regions, specifically targeting the kind of fine texture - pores, small blemishes, minor variation - that reads as surface detail to a person and to a model. Warping filters, more aggressive and less common outside dedicated apps, subtly reposition pixels to slim or reshape a region, which is a geometric change rather than a texture change, and closer in kind to the angle and lens distortions covered elsewhere than to anything about lighting or processing. Both categories operate directly on the properties a texture- or shape-sensitive scoring axis was trained to read, which is exactly why the effect on a score is not a side issue.
Why the direction is unpredictable
A model trained mostly on unfiltered photos has learned what unfiltered texture tends to look like across a wide range of quality levels, and it has learned to associate certain texture patterns with certain outcomes in its training labels. Heavy smoothing can push an image toward looking artificially flat in a way some models read as odd or synthetic - a texture pattern that resembles low-resolution or over-processed images more than it resembles anything in the model's normal training distribution, which is the kind of gap covered under distribution shift. Robustness research treats ordinary image processing as a known weak point: Hendrycks and Dietterich (2019) built the ImageNet-C benchmark to test classifiers against common corruptions such as blur and noise, and found "negligible changes in relative corruption robustness from AlexNet classifiers to ResNet classifiers." Alternatively, if smoother, more even texture correlated with higher-rated photos somewhere in the training data - which is plausible, since professional and well-lit photos also tend to have cleaner texture - the model may read the filtered smoothness as a positive signal despite it being artificial. There is no way to know in advance which effect will dominate for a given model, a given filter strength, and a given photo, because the answer depends on where the specific combination lands relative to that model's specific training distribution, which is not something visible from outside the model.
A different case from automatic processing
This is worth separating from what a phone's camera app already does before you touch a filter, which happens by default, invisibly, to every photo regardless of intent. A beauty filter is applied deliberately, after the base photo already exists, and it is usually stronger and more targeted than default camera processing - a difference of degree that becomes a difference in how predictable the result is, since default processing is roughly constant across your photos while filter choice and intensity vary shot to shot. It's also a different question from someone trying to deliberately game a score by exploiting a known weakness, which is a more adversarial framing covered separately - most filter use is aesthetic preference, not an attempt to manipulate a number.
What follows from this
Because the effect can run either direction, filtering a photo specifically to raise a score is not a reliable strategy, and a result obtained with a filter is not comparable to one obtained without it - one more source of the spread behind why the same subject can score differently across attempts. Anyone running a genuinely controlled comparison needs to hold filter use fixed across every shot in the comparison, on or off, rather than mixing the two and attributing the resulting spread to something else. This piece stays at the level of what happens to the pixels and the score; what to do about it as someone deciding whether to use a filter before a real upload is a question Rate Cock's own guidance addresses directly. A human reviewer notices a filter far more reliably than a model does - unnatural smoothing is a familiar visual tell to a person in a way it is not to a scoring function - and what a person catches that a model misses is one more reason the two methods are not interchangeable, a distinction worth keeping in mind when comparing results across either one.