6 ms·
Also relevant: LLM in a flash: Efficient Large Language Model Inference with Limited Memory Apple seems to be gearing up for significant advances in on-device
by amitprasad 3y ago
Also relevant: LLM in a flash: Efficient Large Language Model Inference with Limited Memory
Apple seems to be gearing up for significant advances in on-device inference using this LLMs
https://arxiv.org/abs/2312.11514 https://arxiv.org/abs/2312.11514