I suspect it's partly because people didn't jump from GPT-3 to GPT-6.1 Sol and partly because SOTA models from the last few(?) months have been able to tackle most of the regular tasks. It means this new model isn't different in that regard from Opus 4.8, if your mental benchmark is that they both are capable of implementing something like a CRUD app.
I’ve seen overwhelmingly that when a model is good people see it, and when it isn’t, they criticize.
I’ve seen nothing but ‘wow this is a huge step up’ from Opus 5.5. I felt this way about Opus 4.5, GPT-5.6 Sol/Luna, and to a lesser extent with Fable and Astra.
But Opus 5 was ass, and the entire gpt 6 line feels like OpenAI’s version of that.
I like the fast releases. 6.1 coming so fast after 6.0 means they found some improvement solid enough for a new rollout.
The only time I remember the newest model making an obvious regression was when the first rolled out MoE. Super speed update but each request had less intelligence at hand. We’re way past that now
Did they find this improvement within the week? If so, given all their whinging about safety, it seems irresponsible to only test the improved model for less than a week.
Did they find the improvement more than a week ago? If so, why bother releasing GPT 6 if they knew they had a better version essentially ready to go?
They probably have 6.2 just about ready, while 7 and 8 are still being cooked up. You can work on multiple releases in parallel.
We're just getting the latest training checkpoints constantly, just to edge out the other lab, while they are trying to come up with something worthy of a new major version number. It's actually worrying, since there's so much pressure to release now.
I'm not convinced. Pipelining doesn't mean they're forced to release a model they know will be replaced in a week.
>We're just getting the latest training checkpoints constantly, just to edge out the other lab, while they are trying to come up with something worthy of a new major version number. It's actually worrying, since there's so much pressure to release now.
This makes me think that this is on purpose to prop up the "we have no control over ourselves, please give us a regulatory moat" narrative. The competition is nowhere near extreme enough to justify weekly releases. The Chinese models are still a handful of months behind the frontier, and the frontier companies in the West haven't been pushing the frontier at this rate.
They've been pretty capable for a very long time. I don't think the models are getting more capable so much as people are getting better at using them and more people are getting the opportunity to be impressed.
Every time a new model comes out, people come out in droves "oh I don't notice anything different".
People have been saying this about <currentModel-1> for 2 years now, and the entire state of AI has changed dramatically.
It cannot be that the next AI model isn't better, but also suddenly what they are capable of is on an entirely different level.