I am starting to wonder how useful these benchmark still are? Aren't all of these models trained to ace these benchmarks?
In any case they give an indication, but I am increasingly looking to real world feedback from real users. I have a solo project I (voice-to-text typing voicewink.app) and I'm using both the Claude Code and Codex coding harnesses and multiple agents doing reviews on the same code.
This has given me real tangible results to compare on real work. Conclusion: GPT-6, Opus 5 / Fable and Grok are "top tier", with Claude good at planning and executing and Codex / Gpt-6 better at finding bugs and fixing them (but tending to overengineer), and Gemini and Kimi 3 clearly behind in capability (more so than the benchmarks suggest in my opinion).
In any case they give an indication, but I am increasingly looking to real world feedback from real users. I have a solo project I (voice-to-text typing voicewink.app) and I'm using both the Claude Code and Codex coding harnesses and multiple agents doing reviews on the same code.
This has given me real tangible results to compare on real work. Conclusion: GPT-6, Opus 5 / Fable and Grok are "top tier", with Claude good at planning and executing and Codex / Gpt-6 better at finding bugs and fixing them (but tending to overengineer), and Gemini and Kimi 3 clearly behind in capability (more so than the benchmarks suggest in my opinion).
Any thoughts?