5 ms·
I think you're being overly cynical here. First, I don't see any claim that is the world's best model for agentic coding. Second, it is absolutely the best mode
by ActivePattern 2mo ago
I think you're being overly cynical here. First, I don't see any claim that is the world's best model for agentic coding. Second, it is absolutely the best model in terms of coding performance vs. dollar, and it's raw performance seems very close to the frontier.
- adam_arthur 2mo agoGPT 5.6 is far more token efficient at most tasks with similar performance. Especially so for Opus 4.8, still to be seen with Opus 5. Where are you getting cheaper per dollar?
- ActivePattern 2mo agoHow are you supporting the claim that GPT 5.6 is "far more token efficient" than Opus 5? Tokens equal, output is cheaper for Opus 5 ($25/1M) than GPT-5.6-Sol ($30/1M), and it seems to outperform slightly on agentic coding benchmarks.
- adam_arthur 2mo agoThe first chart in the blog post shows a similar $/performance curve to GPT 5.6. Where 5.6 has optionality to run much cheaper along the same performance curve at lower thinking levels. There's a later chart that shows Opus 5 ahead, but seems like an esoteric benchmark rather than for common use. (Novel problem solving) If they had a more efficient model at coding they would lead with that chart.
- km144 2mo agoHere is one data point for cost: https://artificialanalysis.ai/models?cost=intelligence-vs-cost-per-task&model-filters=large-models%2Cproprietary%2Creasoning-models#cost-tabs https://artificialanalysis.ai/models?cost=intelligence-vs-co... Here is another data point for output token efficiency: https://artificialanalysis.ai/models?cost=intelligence-vs-cost-per-task&model-filters=large-models%2Cproprietary%2Creasoning-models&intelligence-index-token-use=intelligence-vs-output-tokens-per-task#intelligence-index-token-use-tabs https://artificialanalysis.ai/models?cost=intelligence-vs-co...
- HarHarVeryFunny 2mo agoToken cost and token efficiency are two unrelated metrics, and anyways what really matters is neither in isolation - it's cost to complete a task.
- conradkay 2mo agohttps://www.anthropic.com/_next/image?url=https%3A%2F%2Fwww-cdn.anthropic.com%2Fimages%2F4zrzovbb%2Fwebsite%2F08499ed7c3c2b6416700fa47c70d36dff5eb8461-3840x2160.png&w=3840&q=75 https://www.anthropic.com/_next/image?url=https%3A%2F%2Fwww-... It seems roughly equal according to Anthropic's benchmarks
- shwaj 2mo agoIt would still be the best model per dollar if the score was 2% lower instead of 0.1% lower. Would it be ok to still give it the highlight color then? How big of a lie is too big? Especially when no lie needed to be told at all: many including myself would have noticed the tiny 0.1% deficit and been suitably impressed by the Opus 5 result. I’ll admit this is a small deception by today’s standards. I’m one of those who believes in truth for truth’s sake. Edit: typo