6 ms·
I don't know what black magic you're up to but I see more like 30-35t/s on a 16" M5 Max using 3.8:27b Q4, regardless of whether it's mlx or gguf. qwen3.5:122b-
by Casteil 22d ago
I don't know what black magic you're up to but I see more like 30-35t/s on a 16" M5 Max using 3.8:27b Q4, regardless of whether it's mlx or gguf.
qwen3.5:122b-a10b is significantly faster at around 60-65.
- syntaxing 22d agoWith MTP? I get 25-30 TPS on a strix halo. 50+ on a M5 max should very doable. Dflash (2) will push your TG even further
- Casteil 22d agoIt's a bit deceptive to state inference speeds without mentioning the additional things you're doing to achieve them
- smcleod 22d agoNo magic, just oMLX with MTP. You can look through the speed the community is getting here: https://omlx.ai/benchmarks/performance?model=qwen3.8&chip=&chip_full=M5%7CMax%7C40&quantization=&context=&pp_min=&tg_min=&sort=tg_tps&order=desc https://omlx.ai/benchmarks/performance?model=qwen3.8&chip=&c...