5 ms·
To compare a 1 bit quant to the full fat model is misleading. Honestly this model people at home can tinker with, if you have a big enough Mac. Maybe 4 Strix
by ilc 1mo ago
To compare a 1 bit quant to the full fat model is misleading.
Honestly this model people at home can tinker with, if you have a big enough Mac. Maybe 4 Strix Halo/DGX Spark, and then at 1 bit quant? Nah.
Use the right sized model, for your hardware. You'll get better results.
- guardiangod 1mo agoExtremely large 1 bit models are usually within 50-60% of KV divergence to lossless models. In this case I think the comparison to Opus 4.5 is a fair assessment. Extremely large models don't suffer as much from quantization due to its weight topology also contains encoded information, so the loss of info from any one weight is somewhat mitigated.
- ilc 1mo agoAny one weight, but all of them. And also crushing the architecture itself? I wouldn't pick up 400gb of hardware to run in that mode. I might try it for fun, but even then you are looking at handling a 95GB active parameter set. This is NOT a model for most home labs. I'm sure some can and will use it. But most, should steer clear.
- kadoban 1mo agoI wouldn't just rush out and buy hardware, but there will be benchmarks after a while to make an informed decision. 95GB active is not _too_ bad, would require some creativity and $$, but I bet I could do that at home for less than a cheap car.
- dist-epoch 1mo agoKL divergence (you misspelled it) doesn't tell you anything about capability drop - how much did this particular benchmark (thus ranking among models) change after 10% or 50% KL divergence?
- frotaur 1mo agoNot only that, but KL divergence is not a '%'. It's just a number ranging from 0 to infinity that tells you the 'distance' between two probability distributions.