Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Your benchmarks are really weird though. Opus 5.5 claimed to be trash, Gemini Flash claimed to be the best, Sol 6 claimed to be better than Astra 6? How are you testing? Your findings don't match anyone else's.
 help



I am also surprised by some results, but I also don't favour any model, so all models are tested exactly the same, and some simply fail some tests: incorect answer, requests timing out, not respecting output format requirements, writing code that doesn't generate the correct response, etc.

The top models, when used in a harness, will likely catch those errors or somehow manage to get to the right answer at some point, but those tests test mostly how likely a model is, when used via API, to give the correct response.


The benchmark is a bit sarurated at the top, there it's more relevant to see cost/response times/efficiency.

Opus 5.5 is not trash, still in the top models, but Anthropic models always have struggled with instructions following and refusing to answer questions. That being said, most tests are basically questions or simple tasks, and models could either do them ok or not. Nowadays the models and how good they are in practice is given more by the harness, than the model itself. I do think I should probably find a way to test the models including their harness, and to do so for a complex, long-running task, the so called "agentic" use-case.

Also, Gemini models have the best all-around knowledge, they are way above other models in general knowledge and domain-specific knowledge. No tests have web search enabled, and most modern models are indeed optimized for that use-case nowadays.

tl;dr: the leaderboard simply shows, given any simple question or programming task, which model is most likely to get the answer right.




Consider applying for YC's Winter 2027 batch! Applications are open till November 2.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: