5 ms·
> emerging practice of using Q4 quants and Q8 KV cache for local inference That's not an emerging practice, it's a tested strategy that is these days only used
by zargon 2mo ago
> emerging practice of using Q4 quants and Q8 KV cache for local inference
That's not an emerging practice, it's a tested strategy that is these days only used as a last resort by those desperate to fit a model in memory. Some models do better than others, but generally the model quality suffers greatly under those conditions.
- walrus01 2mo agoI have never seen anyone report "this produced really great results" from intentionally quantizing their context vs. leaving it at full precision which is the ordinary default.
- trollbridge 2mo agoGemma's QAT is surprisingly good (although Gemma isn't that great to begin with).
- KerrAvon 2mo agoIME: Gemma is not great for programming, but it is fantastic at following directions compared to anything else in its size class.