Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

It would be very nice if artificialintelligence.ai actually, from a UX perspective, did the models in more than one thinking mode. I use claude, and I wanna build a feeling for what high, medium, etc. actually gives me. So far their comparisons, and having tried several different models for my work, has given me a feel of what 50 intelligence actually is. And I believe it would be be of even greater value to get a feel inside the single model I actually use, as most people do, because not many, I believe, switch heavily between models when working. I understand that the cost here is greater but the model provivders should obviously give you free access, because of the great work you are doing.
 help



> It would be very nice if artificialintelligence.ai actually, from a UX perspective, did the models in more than one thinking mode

But they do. See for example the pareto curve they have, try to locate GPT-6 Luna (max), (xhigh), (high), (medium), (low)


> I wanna build a feeling for what high, medium, etc. actually gives me

Nothing, really. It's like oversampling your data set. You usually get a much better overall baseline performance if you use the default setting.


In my case it actually made a difference. I tested it on my own code review set with Gemini Flash. On low it found fewer bugs than on default around 90% versus 97% but was about three times faster and a lot cheaper. On high it actually found everything in the hardest case but took almost three minutes per call. so I think the difference is real you only see it if you run the same fixed cases a few times and not by feel. And yes I used Gemini rather than the GPT Sol that is mentioned here would be interesting to see the same tests on Sol.

It does!

…but not for all models, which is pretty annoying.




Consider applying for YC's Winter 2027 batch! Applications are open till November 2.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: