Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

If I had to summarize this in with sentences then here is my attempt.

Transformer training can be parallelized easily, making it possible to use brute force to train the neural network quickly (bitter lesson rewards compute friendly scalable architectures).

Transformers have perfect retrieval, they re-read the entire context window from scratch for every token.

Explanation over.

If you extrapolate this, then the logical conclusion is that the next model architecture would use even more brute force.

Right now transformers can only append a token at the end. This means they can read any input, but write only one specific output.

If you wanted to extend this, you would want to make the transformer read from any input and write to any output, i.e make it capable of updating the entire KV cache every iteration.

 help



Consider applying for YC's Winter 2027 batch! Applications are open till November 2.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: