5 ms·
The past DeepSeek models and now these new checkpoints score very badly on the ArtificialAnalysis AA-Omniscience and hallucination rate benchmarks. I wonder whe
by schmorptron 1mo ago
The past DeepSeek models and now these new checkpoints score very badly on the ArtificialAnalysis AA-Omniscience and hallucination rate benchmarks. I wonder where that's from? Maybe they're overindexing on coding even more than others? I can't say I've noticed it in my (coding) usage so far, has anyone seen it make up potential root causes or other speculative stuff more than other models?
- gunalx 1mo agoYeah. I stopped ising deepseek v4 flashbecause it is awful (even worse than my local qwen3.6 35B model) at multilingual prose.
- isqueiros 1mo agoWorking on a language related app makes me realize that all these supposed language models don't have many good language benchmarks
- mordae 1mo ago0731 is definitely tuned for coding. I mean: https://gertlabs.com/rankings?mode=agentic_coding https://gertlabs.com/rankings?mode=agentic_coding But it is also a decent translator from English to Czech in my experience.