5 ms·
By their own benchmarks it is about 10% lower scoring than Qwen 3.6 35b-a3b, but I've added it to my list. Always looking for MoE to compare to it so we can squ
by hadlock 14d ago
By their own benchmarks it is about 10% lower scoring than Qwen 3.6 35b-a3b, but I've added it to my list. Always looking for MoE to compare to it so we can squeeze more out of our local LLM system.
- Bluestein 13d agoI found it has some "tail" errors, wherein it would make up important details (ie. "happypath.exp" vs "happypaws.exp" and then claim your "DNS is having issues" - where the second domain does not exist), things like that.- ... but correctly supervised it does get some things done.-
- aftbit 13d agoI've found that's generally true of smaller / weaker models. They're quite capable, but you need to distrust them a lot and give them very detailed instructions. Even the free Gemini in Google Search is like this - it lies a lot, clips off important info, and generally goes off the rails if you do too many turns, but it's still very useful if you keep all that in mind.
- Bluestein 13d agoSpot on. Guardrails is the game.-