Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I am also surprised by some results, but I also don't favour any model, so all models are tested exactly the same, and some simply fail some tests: incorect answer, requests timing out, not respecting output format requirements, writing code that doesn't generate the correct response, etc.

The top models, when used in a harness, will likely catch those errors or somehow manage to get to the right answer at some point, but those tests test mostly how likely a model is, when used via API, to give the correct response.

 help



Consider applying for YC's Winter 2027 batch! Applications are open till November 2.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: