Sure, maybe it isn't the model getting nerved but the harness getting updates that make it better with the new model but substantially worse with the old (at that point still current) model.
The test doesn't differentiate. But neither can the average user, who will also be using the normal auto-updating harness. You still get degrading quality right before each new release
Yes, but then the model wasn't nerfed, the harness/overall product just had a plain old regression.
This is very different from a nefarious inference-side degradation to save cost, promote the new model or anything else frequently proposed as motivation.
The test doesn't differentiate. But neither can the average user, who will also be using the normal auto-updating harness. You still get degrading quality right before each new release