7 ms·
So this could universally decrease the memory requirements by un-quantitized LLMs by 30%? Seems big if true.
by wills_forward 1y ago
So this could universally decrease the memory requirements by un-quantitized LLMs by 30%? Seems big if true.
- moffkalast 1y agoNot as big when Q8 quantization is already considered overkill and cuts it down to 50% (and a flat 2x speed boost without any additional compute overhead mind you) and the more common Q4KM is more like 30%. Definitely interesting if it can be added to existing quantization, but K quants do already use different precision levels for different layers depending on general perplexity impact which is similar to this entropy metric they use, e.g. Q6 using a mix of 4 bits and 8 bits. And that's not even considering calibrated imatrix which does something conceptually similar to FFT to compress even higher.
- janalsncm 1y agoQuantization is not lossless.
- danielmarkbruce 1y agoNobody really cares if it meets a strict definition of lossless.
- moffkalast 1y agoAnd when you consider that the usual final step in the pipeline is that a sampler goes ham on the probabilities and just picks some random nonsense, the tolerance for lossy compression is fairly high. In fact, there's this funny occurrence where Q4 models on occasion perform better than their fp16 counterparts on benchmarks ran with top_k=1 since the outputs are slightly more random and they can less deterministically blunder past the local maximum into a more correct solution.
- Der_Einzige 1y agoWe got an oral at ICLR for calling out how shit samplers like top_p and top_k are. Use min_p!
- moffkalast 1y agoTrue yep, I wish more people benchmarked models with more representative sampler settings and then took the average of 5 or 10 responses.
- kridsdale3 1y agoThat's not true. If there are measurable performance differences.
- kadushka 1y agoIf you get any accuracy degradation with full 8 bits of precision you're doing it wrong.
- omneity 1y agoOr your model wasn't trained so well (weights are too spiky)
- danielmarkbruce 1y ago"strict" means something. People, including yourself, only care if there is a practical difference in performance. "this is lossless and that isn't lossless" is a completely useless statement in this realm. In many domains lossy compression is either not tolerated, not legal or not practical.
- throwaway314155 1y agoSeems reductive.
- BoorishBears 1y agoI do? I spend a ton of time post-training models for creative tasks. The effects of model quantization are usually qualified in terms of performance on benchmaxxed tasks with strong logit probabilities, temp 0, and a "right" answer the model has to pick. Or even worse they'll be measured on metrics that don't map to anything except themselves like perplexity (https://arxiv.org/pdf/2407.09141 https://arxiv.org/pdf/2407.09141) I agree Q8 is strong but I also think the effects of quantization are constantly being underappreciated. People are often talking about how these models perform while fundamentally using 10+ variants of a single model with distinct performance profiles. Even knowing the bits per weight used isn't enough to know how exactly a given quant method is affecting the model: https://docs.unsloth.ai/basics/unsloth-dynamic-v2.0-ggufs https://docs.unsloth.ai/basics/unsloth-dynamic-v2.0-ggufs
- danielmarkbruce 1y ago"Nobody really cares if it meets a strict definition of lossless" != "quantization can be done haphazardly."
- BoorishBears 1y agoIf you're trying to really snarkily refer to the article on Dynamic Quants 2.0 and how carefully developed they were, they're comparing their quants to the methodology 99.99% quants out there use. The problem is not that people are making quants "haphazardly", it's that people keep parroting that various quants are "practically lossless" when they actually have absolutely no clue how lossy they are given how application specific the concept is for something as multidimensional as an LLM. The moment anyone tries a little harder to quantify how lossy they are, we repeatedly find that the answer is "not any reasonably definition of lossless". Even in their example where Q4 is <1% away in MMLU 5-shot is probably massively helped by a calibration dataset that maps to MMLU-style tasks really well, just like constantly using WikiText massively helps models that were trained on... tons of text from Wikipedia. So unless you're doing your own calibrated quantization with your own dataset (which is not impossible, but also not near common), even their "non-haphazard" method could have a noticeable impact on performance.