5 ms·
This was done already here as well: https://arxiv.org/abs/2507.04239 https://arxiv.org/abs/2507.04239
by korbip 8mo ago
This was done already here as well: https://arxiv.org/abs/2507.04239 https://arxiv.org/abs/2507.04239
- cubefox 8mo agoSounds interesting, but... > these models dominate both exponential attention and linear attention at long-context training There is no exponential attention; standard attention is quadratic. Strange mistake.