7 ms·
And to read the tea leaves a little: 3.8 actually performs slightly worse than 3.6 on AA-Omniscience Accuracy, which could imply that they traded out world kno
by anana_ 1mo ago
And to read the tea leaves a little:
3.8 actually performs slightly worse than 3.6 on AA-Omniscience Accuracy, which could imply that they traded out world knowledge for capability in other areas.
It also produces nearly twice as many tokens per task as 3.6 (and by extension, time), which may be a tradeoff required to achieve correctness at this parameter size.
- skohan 1mo agoImo it makes sense for things to move in the direction of small, focused models that excel in one area. I use LLMs for technical work 99% of the time, I could care less about general world knowledge, or if the model is good at creative writing. With good orchestration and delegation you can get surprisingly far with small models running on consumer hardware.
- anana_ 1mo agoAgreed. Luckily, this model also scores high in AA non-hallucination, so it knows what it doesn't know -- perfect for situations where it can just tool call a web search.
- tancop 1mo agoThe biggest untapped market is pure agentic models that are built for tool calling and non hallucination instead of memorizing facts. You need some world knowledge (as in common sense) to build a useful model, but I don't think perfect recall on general QA is a good use of space when you have web search and structured knowledge in Wikidata or Wolfram Alpha. Training should focus on tasks that require real intelligence instead of memory. Creative writing is actually good for this if you score it on coherence instead of getting random real life details right. Basic level of coding (simple prompt to code, don't need to one shot complex projects) is also great because writing a small script is more efficient than 20 separate tool calls.
- CamperBob2 1mo agoWith respect to Wolfram Alpha, it's worth noting that VibeThinker-3B is basically a match for the larger frontier models -- hundreds of times larger -- in the narrow domain of mathematical and logical reasoning problems. It doesn't seem necessary to resort to external models or tools for that, at least in principle.
- drob518 1mo agoYep, exactly. I keep saying that I want the “coding expert” extracted from these multi-T parameter models to run locally on reasonable hardware (large laptops, not servers). Yea, I know there’s no single “coding expert” that you can actually extract in these models, but you get what I mean. Like you said, when I’m coding, I don’t care about world knowledge, and I’m fine with consulting another model when I need that.