5 ms·7x speed improvement for LLaMA in less than 10 lines of code2 points by hack_ml 3y agobrucethemoose2 3y agoIs that 5s per token?