Has this hypothesis that reasoning tokens help the model engage a wider range of experts been tested? I'm a little skeptical that it's the main driver of extended reasoning traces, particularly because MoE models are generally already trained so that expert activations are as uniformly distributed as possible. But it would be interesting to know to what extent this happens.