You will understand if you run some LLM models with a greedy sampler that does that. The text quality begins to deteriorate rapidly. This is a very counter intuitive result so I don't blame you for not understanding until you actually tried it and experienced it for yourself.
> You will understand if you run some LLM models with a greedy sampler that does that. The text quality begins to deteriorate rapidly.
Right, I've done this, and this makes sense to me, but I'm not following how that falsifies the top probability word being the best choice in any particular instance.
"Picking only the best word at each decision point results in a worse final result" seems like an imminently reasonable hypothesis.
I think you're using different definitions of best. If best = leads to a correct answer overall then by definition anything that leads to a bad outcome can't be best