I’m talking about attention only transformers. Those don’t have an autodiff but ...

lostmsu · on Feb 9, 2023

> attention only transformers

Can you share any good link on the subject?

adamnemecek · on Feb 10, 2023

https://transformer-circuits.pub/2021/framework/index.html

lostmsu · on Feb 10, 2023

Maybe I am missing something, but I don't see any learning without autodiff.

adamnemecek · on Feb 11, 2023

I thought you were asking about attention only transformers. This paper touches on some of it https://arxiv.org/abs/2212.10559v2.

lostmsu · on Feb 13, 2023

The paper speculates that it is analogous to gradient descent and empirically confirms it is similar in behavior, but it is not a rigorous proof of any kind.

The momentum experiment they made also does not seem related. E.g. it just adds past values to V, which extends the effective context length.

adamnemecek · on Feb 14, 2023

> but it is not a rigorous proof of any kind.

Such is the nature of early theories.