5 ms·
Simon Willison's post about this gives a good context on why exactly this is happening. While it doesn't mention this in the Artificial Analysis page, this is l
by phsource 1mo ago
Simon Willison's post about this gives a good context on why exactly this is happening. While it doesn't mention this in the Artificial Analysis page, this is likely with Max reasoning, which has extremely long reasoning traces:
https://simonwillison.net/2026/Aug/16/qwen-38-27b/ https://simonwillison.net/2026/Aug/16/qwen-38-27b/
It seems like the token usage is 2.3x GPT Luna Max and almost 2x Kimi K3!
https://imgur.com/a/dDSyhr2 https://imgur.com/a/dDSyhr2
I'm curious if they can make up for this with insanely high tokens-per-second especially when served from hosted providers, though, given how tiny it is (37B!)
- skohan 1mo agoI'm running 3.8 27B locally, and the results from the past few days have been excellent. I find raw speed is less of an issue when you can trust the model more to reach the right result.
- ArvidSu 1mo agoA ThinkingCap variant of Qwen 3.8 27b would be extremely interesting. https://huggingface.co/bottlecapai/ThinkingCap-Qwen3.6-27B https://huggingface.co/bottlecapai/ThinkingCap-Qwen3.6-27B And then a Bonsai ternary on top of that model.
- kees99 1mo agoRe: bonsai - unsloth's quants have Q2 (UD-IQ2) variants, which are more or less same in size. ...or did Prism do something special with their "bonsai" releases? I didn't notice anything like QAT being mentioned.
- drob518 1mo agoThere is some special sauce that they have. It’s not just a simple quant of another release. Or so they imply. I don’t have any insight into how it works or what the Bonsai special sauce is.
- DoctorOetker 1mo agoI suggest nobody repeats claims of "special sauce" if such "special sauce" is not understood.
- kees99 1mo agoQwen models are slower in tokens/s, compared to similarly sized gemma4 and others, and they use more tokens per task, in part thanks to that xhigh default. On the other hand, there are some of us who are stuck with hardware that has plenty of compute, but limited (V)RAM. The new 27B is just perfect for that.
- petu 1mo ago> Qwen models are slower in tokens/s, compared to similarly sized gemma4 and others No? Gemma 31B and Qwen 27B are about the same speed. Gemma 26B-A4B and Qwen 35B-A3B are about the same speed.
- stymaar 1mo agoThere's no Qwen3.8-35B-A3B though.
- hadlock 1mo agoI benched Qwen 3.6 35B-A3B against Qwen 3.8 27B with the same parameters, thinking set to low. Despite 35B having 9x fewer active parameters, it benched only 2.34x slower. The 35B got only 50% more agentic tasks done per hour.
- trouve_search 1mo agoWhat configuration are you using? On both vllm and llama-cpp, I get significantly higher speeds from gemma4 than qwen3.6 (with their respective speculative decoding methods). Output TPS in vllm for instance: - Gemma4 26B-A4B: 200-300TPS - Qwen3.6 35B-A3B: 120-180TPS - Gemma4 31B: 80-120TPS - Qwen3.6 27B: 60-80TPS This is for a first request on a dual 5090 setup, with their respective speculative decoding methods.
- petu 1mo agoSingle 3090 under llama.cpp: | model | size | test | t/s | | ------------------- | ------- | ------ | ---- | | gemma4 31B Q4_0 | 16.1 GB | pp2048 | 1248 | | gemma4 31B Q4_0 | 16.1 GB | tg512 | 40 | | qwen35 27B Q4_K | 15.9 GB | pp2048 | 1248 | | qwen35 27B Q4_K | 15.9 GB | tg512 | 39 | | gemma4 26B.A4B Q4_0 | 13.3 GB | pp2048 | 4304 | | gemma4 26B.A4B Q4_0 | 13.3 GB | tg512 | 160 | | qwen35 35B.A3B Q3_K | 15.7 GB | pp2048 | 3329 | | qwen35 35B.A3B Q3_K | 15.7 GB | tg512 | 144 | > with their respective speculative decoding methods You're benchmarking drafter acceptance rate, then. Which is real life values, yes, but attributing worse drafter performance to the other 95% of the model being inherently slower.
- drob518 1mo agoIt’s still going to chew up context quickly. Surely, some of the added tokens are helping the model, but does it require as many as it generates? What happens on long, multi step tasks as it pushes old tokens out of context? I’m not sure we know the answers to those.
- stymaar 1mo ago> insanely high tokens-per-second especially when served from hosted providers, though, given how tiny it is (37B!) It's a dense model so it will use all of its parameters per token. 37B active parameters isn't tiny at all, it's almost what Deepseek R1 had, and it's 2/3 of what Kimi k3 uses, so it's not going to be “insanely high” tps: it's going to be three times slower than Deepseek Flash (Prefil speed is going to be quite high though, but not token generation).
- Azantys 1mo agoIts 27B not 37B and having just 27B in total and 3T and like 30B active of those is still totally different. A 120B with 5B active is still much slower than a proper 5B. Just like the new Ling 3.0 Tiny with 8B and 1B active only gets around 120tk/s compared to 250tk/s which a real 1B one gets on my hardware.
- jakswa 1mo agoThanks for mentioning Ling 3 Tiny. This model has completely bypassed me and seems promising for how small it is.
- 2001zhaozhao 1mo agoOn the other hand, it used about the same tokens as GLM 5.2 and got 1 point lower score. The fact that we have a GLM 5.2-class model that can run on two 3090's comfortably at Q8 is absolutely insane. It wasn't long ago that GLM 5.2 was considered amazing for open weight models.
- thousand_nights 1mo agoi feel like Simon omitted an important part of how Qwen's "reasoning" levels work. they are just one sentence additions/omissions to the system prompt xhigh -> "Reasoning effort is set to xhigh. Please think carefully through the task, validate key assumptions, consider plausible alternatives, and prioritize correctness, consistency, and clarity in the final answer." medium -> no mention of effort (sentence omitted) low -> "Reasoning effort is set to low. Keep your thinking brief and focused, moving directly to the conclusion without unnecessary elaboration." in my testing this doesn't seem to produce exactly deterministic thinking levels, because it's just a system prompt nudge. i had instances where medium thought longer than xhigh
- anon373839 1mo agoThe models are post-trained on these prompt additions so they’re more structural than thinking of them as “system prompts” suggests. (All LLMs ever see is tokens going in, so even the concept of a system prompt is just formatting they’ve seen in post-training.) You can also apply fixed token budgets for the reasoning blocks, though it will decrease quality in some cases.
- thousand_nights 1mo agoyes of course, I understand that. but I feel like it would've been nice to include in the article because the main point of it is the effort and overthinking
- nixon_why69 1mo agoWhy not invent a few magic token values for reasoning level instead? It would be like 4 out of a vocabulary of 200k and save like 30 tokens in every prompt
- Bombthecat 1mo agoDamn, beating Kimi k3 is crazy. It produces already a ton of tokens. But I think the next step will be even more thinking on smaller models. Maybe fine-tuned and we get really crazy stuff
- segmondy 1mo agoIt doesn't beat k3, it doesn't even beat deepseekv4Flash-0731 let alone glm5.2