Do we really have the data on this? I mean, it does happen on a smaller scale, b...

hodgehog11 · 2025-08-08T03:27:44 1754623664

True, we can't say for certain. But there is a lot of theoretical evidence too, as the leading theoretical models for neural scaling laws suggest finer properties of the architecture class play a very limited role in the exponent.

We know that transformers have the smallest constant in the neural scaling laws, so it seems irresponsible to scale another architecture class to extreme parameter sizes without a very good reason.