8 ms·FlashAttention – optimizing GPU memory for more scalable transformers1 points by mpaepper 2y ago