7 ms·
Artificial Analysis hasn't posted their independent analysis of Qwen3.6 35B A3B yet, but Alibaba's benchmarks paint it as being on par with Qwen3.5 27B (or bett
by coder543 5mo ago
Artificial Analysis hasn't posted their independent analysis of Qwen3.6 35B A3B yet, but Alibaba's benchmarks paint it as being on par with Qwen3.5 27B (or better in some cases).
Even Qwen3.5 35B A3B benchmarks roughly on par with Haiku 4.5, so Qwen3.6 should be a noticeable step up.
https://artificialanalysis.ai/models?models=gpt-oss-120b%2Cgpt-5-4%2Cgemini-3-1-pro-preview%2Cgemma-4-31b%2Cclaude-sonnet-4-6-adaptive%2Cclaude-opus-4-6-adaptive%2Cclaude-4-5-haiku-reasoning%2Cglm-5-1%2Cqwen3-5-27b%2Cqwen3-5-35b-a3b https://artificialanalysis.ai/models?models=gpt-oss-120b%2Cg...
No, these benchmarks are not perfect, but short of trying it yourself, this is the best we've got.
Compared to the frontier coding models like Opus 4.7 and GPT 5.4, Qwen3.6 35B A3B is not going to feel smart at all, but for something that can run quickly at home... it is impressive how far this stuff has come.
- naasking 5mo agoQwen models commonly get accused of benchmaxxing though. Just something to keep in mind when weighing the standard benchmarks.
- coder543 5mo agoEvery model release gets accused of that, including the flagship models.
- naasking 5mo agoLess so for Gemma-4 because it falls behind Qwen on benchmarks. Tests for benchmaxxing are also strongly suggestive: https://x.com/bnjmn_marie/status/2041540879165403527 https://x.com/bnjmn_marie/status/2041540879165403527
- coder543 5mo agoNo… seriously. Every model release is accused. Including Opus, GPT-5.4, whatever. And yes, including smaller models that are not the top in every benchmark. My own experiences with Gemma 4 have been quite mediocre: https://www.reddit.com/r/LocalLLaMA/comments/1sn3izh/comment/ogjsfjv/?context=3 https://www.reddit.com/r/LocalLLaMA/comments/1sn3izh/comment... I would almost be tempted to call it benchmaxed if that term weren’t such a joke at this point. It is a deeply unserious term these days. Gemma 4 is worse than its benchmarks show in terms of agentic workflows. The Qwen3.x models are much better; not benchmaxed. I have tested this extensively for my own workflows. Google really needs to release Gemma 4.1 ASAP. I really hope they’re not planning to just wait another calendar year like they did for Gemma 3 -> 4 with no intermediate updates. And the lead author on the paper replied to that tweet to say that the scores would need to be greater than 80 to show actual contamination: https://x.com/MiZawalski/status/2043990236317851944?s=20 https://x.com/MiZawalski/status/2043990236317851944?s=20