8 ms·
This project needs webgpu -- I did it on cpu about a year ago. My demo uses 2 bit quantization to run llama3 models on any device with enough ram. https://gal
by om8 2mo ago
This project needs webgpu -- I did it on cpu about a year ago.
My demo uses 2 bit quantization to run llama3 models on any device with enough ram.
https://galqiwi.github.io/aqlm-rs/ https://galqiwi.github.io/aqlm-rs/