It would be very nice if artificialintelligence.ai actually, from a UX perspective, did the models in more than one thinking mode. I use claude, and I wanna build a feeling for what high, medium, etc. actually gives me. So far their comparisons, and having tried several different models for my work, has given me a feel of what 50 intelligence actually is. And I believe it would be
be of even greater value to get a feel inside the single model I actually use, as most people do, because not many, I believe, switch heavily between models when working. I understand that the cost here is greater but the model provivders should obviously give you free access, because of the great work you are doing.
In my case it actually made a difference. I tested it on my own code review set with Gemini Flash. On low it found fewer bugs than on default around 90% versus 97% but was about three times faster and a lot cheaper. On high it actually found everything in the hardest case but took almost three minutes per call. so I think the difference is real you only see it if you run the same fixed cases a few times and not by feel. And yes I used Gemini rather than the GPT Sol that is mentioned here would be interesting to see the same tests on Sol.