Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

What's the clear best, that you see?


What I’ve seen from LLMs is you have no idea which model does the best at a certain task until you try all the models. Gemini and Chat seem to be the best research models but also hallucinate like crazy. Oddly, Deepseek v4 Flash is the best web research model I’ve used. No idea why! There doesn’t seem to be a “best at everything “ and you don’t know which is the best at the thing you are doing until you try doing it.

Note this is for complicated tasks - simple stuff can be done by whatever pretty well.

Also models will be awesome at doing something at 150k tok of context and terrible at 750k toks so even within the model itself there are capability considerations.

It’s an engineering problem to design around, not something you can escape.


Hilarious to see only different responses


Should have been “clear best and what do you do”


GPT 5.6 Sol


GLM 5.2


Fable 5




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: