I refuse to believe they "play with their quants" once a model version is labelled and shipped. What does that even mean; could you explain it please? These models aren't just used through claude/codex, they are used through API access and it's quite expensive. Previous regressions were related to harness regression, and platform issues. Not some Nerf conspiracy 99% of the vibe bros believe in.
Note: I know what quantization is so don't hold back.
I would guess that if they do use such methods, it'd be to handle peak loads that go beyond their compute capacity, while they run the models at full capability when there's excess capacity
like before Anthropic signed the Colossus deal, the usage limits were insane and everyone was complaining, I wouldn't be surprised if they'd rather try to make inference faster that way than try to just limit people, at least for those on subscriptions
Trimming parameter size that can be reduced while surviving regression evals. They have so much data they know exactly where to shave the models. Most people will never see it in their work loads. It won’t affect core benches because that is part of the regression evaluation.
During peak hours requests queue and inference slows. During off-peak they can move systems over to training.
Where is the evidence they are "nerfing" the models due to request volume?
Edit: I don't know they do, I mean they could repurpose systems if they are idle. Inference demand is global, and providers like Azure have global routing options that are cheaper. Night time in the USA could be serving inference demand on the other side of the globe.
It's all just conjecture, your hypothesis about moving systems equally so.
But you seem adamant that there's no chance the providers serve slightly quantized models for subscription users during high loads, or otherwise tweak models for requests from those users.
It's tricky to prove either way, but the chance is not zero.
Note: I know what quantization is so don't hold back.