There is a theory that the verbosity and comments help getting better results with the current benchmarks. So the models are theoretically getting better but in practice they are getting worse.
I believe they are genuinely getting better for fully autonomous tasks. Anthropic seems to have gone all in on this, at the expense of more typical usage patterns.