I am also surprised by some results, but I also don't favour any model, so all models are tested exactly the same, and some simply fail some tests: incorect answer, requests timing out, not respecting output format requirements, writing code that doesn't generate the correct response, etc.
The top models, when used in a harness, will likely catch those errors or somehow manage to get to the right answer at some point, but those tests test mostly how likely a model is, when used via API, to give the correct response.
The top models, when used in a harness, will likely catch those errors or somehow manage to get to the right answer at some point, but those tests test mostly how likely a model is, when used via API, to give the correct response.