Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Big number results, and impressive pricing. That said it really feels like benchmarks have been hyper saturated these days. I’ll wait for hands on before getting too hyped that Google is back. It would be nice having more than just OAI / A\ in the running for SOTA top tier intelligence.
 help



I don't think new benchmarks are saturated. They still give you a clue, they arn't perect but they have value. If model can't even do some easy tasks from benchmark then why would u even consider using it?

With these numbers, I'm holding my breath for the pelicanbench.

~20% for Harvey's Legal Benchmark doesn't seem saturated.

I suppose it could be saturated if we assume we've hit the limit on LLM capabilities.



Consider applying for YC's Winter 2027 batch! Applications are open till November 2.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: