7 ms·
256+: GLM-5.2 (swap with Kimi K3 when it goes open-weight on Monday) <256: Actually, surprisingly, not a Chinese model but probably Laguna S2.1. The best Chine
by reissbaker 2mo ago
256+: GLM-5.2 (swap with Kimi K3 when it goes open-weight on Monday)
<256: Actually, surprisingly, not a Chinese model but probably Laguna S2.1. The best Chinese model at this size is DeepSeek V4 Flash though
<96: Qwen 3.6 27B
<32: Still Qwen 3.6 27B (NVFP4)
<16: Oof, not sure. Nothing will feel great at this size TBQH without finetuning on a specific task. Pick your poison of tiny Qwen or tiny Gemma (although again Gemma is not Chinese)
- 3eb7988a1663 2mo agoExactly the kind of breakdown I was hoping to see. Thanks.
- pbgcp2026 2mo agoAnd we can quickly see that the real problem is not the models, but HW to run them. You can build whole Enterprise on Gemma 4 31B full precision without significant problems. If you can afford not to lobotomise it by quantisation.
- greenavocado 2mo ago[dead]