14 ms·
I was already rolling around the idea of a 128GB M5 Max MBP. Now this! A 4-bit MLX quant with 128k window should fit perfectly, in the 50-70 tok/s range.
by pwython 22d ago
I was already rolling around the idea of a 128GB M5 Max MBP. Now this!
A 4-bit MLX quant with 128k window should fit perfectly, in the 50-70 tok/s range.
- sscaryterry 22d agoI have a 128GB M5 Max, and it sucks at this stage. 50-70 tok/s might be something...
- smcleod 22d ago50-70tk/s is what I get on my m5 max on a 5-6bit Qwen 3.8 27B?
- Casteil 22d agoI don't know what black magic you're up to but I see more like 30-35t/s on a 16" M5 Max using 3.8:27b Q4, regardless of whether it's mlx or gguf. qwen3.5:122b-a10b is significantly faster at around 60-65.
- syntaxing 22d agoWith MTP? I get 25-30 TPS on a strix halo. 50+ on a M5 max should very doable. Dflash (2) will push your TG even further
- Casteil 22d agoIt's a bit deceptive to state inference speeds without mentioning the additional things you're doing to achieve them
- smcleod 22d agoNo magic, just oMLX with MTP. You can look through the speed the community is getting here: https://omlx.ai/benchmarks/performance?model=qwen3.8&chip=&chip_full=M5%7CMax%7C40&quantization=&context=&pp_min=&tg_min=&sort=tg_tps&order=desc https://omlx.ai/benchmarks/performance?model=qwen3.8&chip=&c...
- sscaryterry 22d agoI tried 8-bit, perhaps I should try 6-bit.
- irthomasthomas 22d agoIDK, prefill speed is a bigger concern for most wokflows, like agent coding, and I heard that this is quite low on macs?
- smcleod 22d agoThat was mainly before the M4 generation when they didn't have matmul instructions.
- jasonjmcghee 22d agoM5 prefill is much faster than M4. I've seen benchmarks that show 4-5x faster of M5 Max vs. M4 Max. For local models you're likely using M5 Max, prefill is low thousands of tokens per second, as opposed to, say high hundreds with M4 Max. For larger dense models, some fraction of that, but similar multiple.
- smcleod 22d agoYes, I have the M5 Max. But there was no matmul acceleration before the M4 which made things a lot slower.
- deleted 22d ago[deleted]
- Eric_WVGG 22d agoJust out of curiosity, why run "local-local" when you could just set up a Mini or Studio at home and query it over http? [edit] whole conversation about this in another thread https://news.ycombinator.com/item?id=49433413 https://news.ycombinator.com/item?id=49433413 I’m personally considering retiring my MBP for a Studio + 15" Air whenever this MBP ages out.
- LeBit 22d agoThis is the way. I’m doing that. Mac Mini M4 Pro with 48G RAM as a headless llama.cpp server. I much prefer using " thin clients " as the interface to the big VMs running in my homelab
- kamranjon 22d agoI actually do this with my MBP - it's a LLM server when I'm working - and then when I'm not it's just a really great machine for video editing and other media work.
- rdsubhas 22d agoHow do you folks code at 40-50 tps? With an extremely lightweight harness (pi) and just 8k system and tools context, and ~40tps on qwen 3.8 27B 4-bit on low thinking mode, it still takes me nearly 30-45 mins for a basic coding session... Does it work? yeah... But I'd pick a subscription anyday...
- hgoel 22d agoDo you find subscriptions to be meaningfully faster? I didn't really feel too much of a speed difference compared to Opus.
- latentsea 22d agoAs someone who uses Opus daily for professional work and Qwen3.8-27B for all my private stuff, yes, Opus sub is faster for me, but I'm only rocking an R9700. If you're lucky enough to have sold a kidney on the blackmarket and purchased a 5090 and you're running ninfer, then actually... I think you'd be seeing fairly comparable performance!
- julianlam 21d agoWhen you hear stories like "Opus 5 thought for 20 minutes and then denied my request" it really puts wind in this sails of Local LMs
- bicepjai 21d agoThere is a finite amount of time left for these companies to become next Facebook/Google, hoarding our interaction and privacy will be a point of contention pretty soon. At that moment, Qwen will be the knight in shining armor.