Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I am starting to wonder how useful these benchmark still are? Aren't all of these models trained to ace these benchmarks?

In any case they give an indication, but I am increasingly looking to real world feedback from real users. I have a solo project I (voice-to-text typing voicewink.app) and I'm using both the Claude Code and Codex coding harnesses and multiple agents doing reviews on the same code.

This has given me real tangible results to compare on real work. Conclusion: GPT-6, Opus 5 / Fable and Grok are "top tier", with Claude good at planning and executing and Codex / Gpt-6 better at finding bugs and fixing them (but tending to overengineer), and Gemini and Kimi 3 clearly behind in capability (more so than the benchmarks suggest in my opinion).

Any thoughts?

 help



I find codex to overengineer for routine tasks compared to claude. I use codex mainly for infra tasks.

Yes, exactly my experience. Do you have tricks / approaches to keep Codex from overengineering?



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: