5 ms·
I get ~5 tokens/s on an M4 with 32G of RAM, using: llama-server \ -hf unsloth/Qwen3.6-27B-GGUF:Q4_K_M \ --no-mmproj \ --fit on \ -np 1 \ -c 65
by benob 5mo ago
I get ~5 tokens/s on an M4 with 32G of RAM, using:
llama-server \
-hf unsloth/Qwen3.6-27B-GGUF:Q4_K_M \
--no-mmproj \
--fit on \
-np 1 \
-c 65536 \
--cache-ram 4096 -ctxcp 2 \
--jinja \
--temp 0.6 \
--top-p 0.95 \
--top-k 20 \
--min-p 0.0 \
--presence-penalty 0.0 \
--repeat-penalty 1.0 \
--reasoning on \
--chat-template-kwargs '{"preserve_thinking": true}'
35B-A3B model is at ~25 t/s. For comparison, on an A100 (~RTX 3090 with more memory) they fare respectively at 41 t/s and 97 t/s.
I haven't tested the 27B model yet, but 35B-A3B often gets off rails after 15k-20k tokens of context. You can have it to do basic things reliably, but certainly not at the level of "frontier" models.
- dunb 5mo agoWhy use --fit on on an M4? My understanding was that given the unified memory, you should push all layers to the GPU with --n-gpu-layers all. Setting --flash-attn on and --no-mmap may also get you better results.
- halJordan 5mo agoMeaningless question, fit will put everything on the gpu if it fits. Fa is default on. No-mmap is not an inference tradeoff and if you do turn it off you need to turn on direct io via -dio What he should actually do is enable speculative decoding
- danielhanchen 5mo agoWe also made some dynamic MLX ones if they help - it might be faster for Macs, but llama-server definitely is improving at a fast pace. https://huggingface.co/unsloth/Qwen3.6-27B-UD-MLX-4bit https://huggingface.co/unsloth/Qwen3.6-27B-UD-MLX-4bit
- DarmokJalad1701 5mo agoWhat exactly does the .sh file install? How does it compare to running the same model in, say, omlx?
- danielhanchen 5mo agoSorry on the delay - so it installs https://github.com/Blaizzy/mlx-vlm https://github.com/Blaizzy/mlx-vlm and other components and sets up the commands - you don't need to use it but we thought it might be easier for folks
- kpw94 5mo agoWhen you say tok/s here are you describing the prefill (prompt eval) token/s or the output generation tok/s? (Btw I believe the "--jinja" flag is by default true since sometime late 2025, so not needed anymore)
- zargon 5mo agoIf someone doesn't specifically say prefill then they always mean decode speed. I have never seen an exception. Most people just ignore prefill.
- kpw94 5mo agoBut isn't the prefill speed the bottleneck in some systems* ? Sure it's order of magnitude faster (10x on Apple Metal?) but there's also order of magnitude more tokens to process, especially for tasks involving summarization of some sort. But point taken that the parent numbers are probably decode * Specifically, Mac metal, which is what parent numbers are about
- zargon 5mo agoYes, definitely it's the bottleneck for most use cases besides "chatting". It's the reason I have never bought a Mac for LLM purposes. It's frustrating when trying to find benchmarks because almost everyone gives decode speed without mentioning prefill speed.
- mercutio2 5mo agooMLX makes prefill effectively instantaneous on a Mac. Storing an LRU KV Cache of all your conversations both in memory, and on (plenty fast enough) SSD, especially including the fixed agent context every conversation starts with, means we go from "painfully slow" to "faster than using Claude" most of the time. It's kind of shocking this much perf was lying on the ground waiting to be picked up. Open models are still dumber than leading closed models, especially for editing existing code. But I use it as essentially free "analyze this code, look for problem <x|y|z>" which Claude is happy to do for an enormous amount of consumed tokens. But speed is no longer a problem. It's pretty awesome over here in unified memory Mac land :)
- wuschel 5mo agoHow is the quality of model answers to your queries? Are they stable over time? I am wondering how to measure that anyway.
- cyanydeez 5mo agoUsing opencode and Qwen-Coder-Next I get it reliably up to about 85k before it takes too long to respond. I tried the other qwen models and the reasoning stuff seems to do more harm than good.
- fuomag9 5mo agoI confirm with the GGUF version at q4, 35B-A3B starts going in thinking loops at 60k basically