7 ms·
you'd be surprised how good small models have gotten. Size of the model isnt all that matters.
by wesammikhail 5mo ago
you'd be surprised how good small models have gotten. Size of the model isnt all that matters.
- verdverm 5mo agoMy experience with qwen-3.6:35B-A3B reinforces this, gonna give this a spin when unsloth has quants available Gemini flash was just as good as pro for most tasks with good prompts, tools, and context. Gemma 4 was nearly as good as flash and Qwen 3.6 appears to be even better.
- cassianoleal 5mo ago> when unsloth has quants available https://huggingface.co/unsloth/Qwen3.6-27B-GGUF https://huggingface.co/unsloth/Qwen3.6-27B-GGUF
- verdverm 5mo agoThat was quick (compared to the 1T Kimi-2.6, not surprising)
- danielhanchen 5mo agoHaha :) We had some issues with Kimi-2.6 since it was int4 and we were investigating how to handle it :)
- verdverm 5mo agoAppreciate what y'all do! We were slacking about how many HGX-B300 it would take to run Kimi and it looks like we could actually fit 2-3 Kimis on a single HGX.
- danielhanchen 5mo agoSorry on the delay - oh haha that would be cool :) We did release 2bit dynamic ones, but unsure if they'll be helpful
- freedomben 5mo agoPlus you can control thinking time a lot more, so when Anthropic lobotomizes Opus on you...
- dudefeliciano 5mo ago> Size of the model isnt all that matters. What matters is the motion in the tokens