6 ms·
Only for the reference implementation, flash attention is just based on optimization of the memory bandwidth between caches. The ThunderKittens[0] library has w
by f_devd 2y ago
Only for the reference implementation, flash attention is just based on optimization of the memory bandwidth between caches. The ThunderKittens[0] library has what is effectively flash attention v3, and they are working on supporting AMD arch.
[0]: https://hazyresearch.stanford.edu/blog/2024-05-12-tk https://hazyresearch.stanford.edu/blog/2024-05-12-tk