6 ms·
Those 27.3% are still in the ballpark of modern models: - Sonnet 5 - 12.4% - Luna - 17.3% - Grok 4.6 - 20.3% - Sol - 37.3% - GLM 5.3 - 41.8% - Opus 5 - 51
by fallingbananna 6d ago
Those 27.3% are still in the ballpark of modern models:
- Sonnet 5 - 12.4%
- Luna - 17.3%
- Grok 4.6 - 20.3%
- Sol - 37.3%
- GLM 5.3 - 41.8%
- Opus 5 - 51.8%
- Readerium 6d agoDeepSeek v4.1 Flash 31.2% Source: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash#comparison-with-frontier-models-max-reasoning-effort https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash#compa...
- p1esk 6d agoAstra is 58%. The current title says it's "rivaling Astra"
- thereitgoes456 6d agoIt is rivaling Astra, on their own benchmark that they made (FrontierCode), that they ran themselves in their own closed-source ecosystem that isn’t reproducible by anyone.
- nijave 6d agoThis explains a lot about Sonnet 5.
- mokre 6d agoGLM 5.3 looks strange, because of this Chinese labs benchmaxx moto. So rather they have emergent abilities or... Also a lot of questions to benchmark because opus 5 is completely useless model right now. I think that the main problem with opus that they try to solve context size optimization problem, and that is the main reason why it speaks like alien with only one technical dictionary at hand. So why it is so good?
- gpt5 6d agoTerminal Bench 4.0 did not introduce new questions. All tasks were public for a while. If you look at GLM 5.2, which is using the same base model as 5.3, but was released prior to most tasks, it does extremely horribly on terminal Bench 3.0 (4 to 8 times worse than every other model) - source: https://benchlm.ai/benchmarks/terminal-bench-3 https://benchlm.ai/benchmarks/terminal-bench-3
- didibus 6d agoOpus 5 generates really good code and terminal commands though. It's just bad at the accompanying text it tells you. These benchmarks don't grade the text generation of the response I don't think, only the task outcome.
- simondotau 6d agoOpus 5 keeps making howlingly stupid errors, like one recently where a regexp would catch an invalid date because -\d{2} won’t match -00 Seriously.
- didibus 4d agoThat's just all models though. They're not 100% accurate.
- simondotau 4d agoI regularly have issues like that with Opus. I rarely do with the supposedly sub-frontier Grok 4.6, or Kimi K3, or even with Sol.
- airstrike 6d agoSo, better than Sonnet and Luna? lol