5 ms·
Have you benchmarked against full precision models for accuracy/ performance?
by htrp 22d ago
Have you benchmarked against full precision models for accuracy/ performance?
- cpburns2009 22d agoNot full precision. I've only benchmarked 27B across Q3-6 quants using lm-eval. I lack the hardware to bench 27B at BF16 but I might be able to do Q8_0. I haven't gotten around to doing 35B. I really should upload my collection of results to Github or somewhere. Here's a summary of what I have for 27B. I used unsloth's UD-Q{3-6}_K_XL quants across 11 evals. The values are pretty linear between Q3 and Q6. Qwen3.6-27B Q3 Q6 ARC-Challenge 97.0 97.0 BIG-Bench Hard 57.9 59.3 GPQA Diamond 77.8 83.3 GSM8K 92.4 92.6 Hendrycks Math 35.5 38.9 HumanEval 80.5 85.4 HumanEval+ 75.0 79.3 IFEval 87.3 88.0 MBPP 75.2 77.2 MBPP+ 88.4 88.9 MMLU-Pro 83.1 83.5