8 ms·
Given that it apparently defaults to 'xhigh', this is probably the answer. Granted, it's still much lower tokens/s than you'll get out of many MoE models. Edi
by Casteil 1mo ago
Given that it apparently defaults to 'xhigh', this is probably the answer.
Granted, it's still much lower tokens/s than you'll get out of many MoE models.
Edit: Even set to medium or low there's still a lot of second guessing, less consistency, lower 'acceptable response' rate, and slower/more token churn vs gemma4:26b-a3b. I think gemma4 is just a better 'general purpose' model.
- eek2121 1mo agoI haven't tried lowering thinking, however, I actually asked a solid question earlier regarding a real world scenario I encountered and all that excessive thinking made it give me an amazing answer. The thinking actually all made sense, and honestly I found it thought of similar stuff to what I thought when I drew my own conclusion. I don't usually rely on AI for much (I'm actually kind of anti-AI, although I follow stuff like this enthusiastically because of the rapid advancements, "average joe" access, and openness), however for what I asked? It was spot on. The subject was a bit personal, so I won't share. I was just curious what AI would say about the situation and it definitely surprised me, especially since Qwen, while usually great on development/coding stuff, has shown weaknesses in other areas. Definitely a solid release, and this one runs on my 4090 with minimal loss of quality!