The speedup here would be very dependent on the context -- the kind of texts that the models are working with, as it proposes a rather naive n-gram generator (maybe I should say it does not provide any details on this critical component, instead simply refers to Jurafsky textbook). It might not be robust. Instead Apple's work on using the same model to produce n-gram lookahead is robust -- the n-gram generator works as well as the model itself: https://arxiv.org/abs/2402.11131