6 ms·
I like how they highlighted Opus 5 as the best for “Agentic Coding” even though the number is slightly lower than Fable. Close enough for marketing, I guess!
by shwaj 2mo ago
I like how they highlighted Opus 5 as the best for “Agentic Coding” even though the number is slightly lower than Fable. Close enough for marketing, I guess!
- ceejayoz 2mo agoBest can describe multiple things. Almost as good for half the cost is something I'm very comfortable describing that way.
- ProofHouse 2mo agoBest marketing
- lelanthran 2mo ago> Almost as good for half the cost is something I'm very comfortable describing that way. It's also not unusual in this context - many people describe the Chinese models as "best", because it's 80% as good for 20% of the price (or similar).
- moffkalast 2mo agoHopefully it's not like old Opus, where it was actually more expensive than Fable cause it thought for half an hour, got it wrong, and then thought until you ran out of credits trying to come up with a correction, while Fable just went for it and did it in one go, getting it right the first time without thinking more than a few seconds. Got an endless list of stuff done with Fable, Opus 4.8 was like a flailing braindead idiot in comparison. Maybe this one is a bit better if it's distilled.
- deleted 2mo ago[deleted]
- entropicdrifter 2mo agoI mean that certainly makes it best-in-class
- dd8601fn 2mo agoAt half the price and less likely to auto-downgrade, it sounds like a reasonable claim.
- binsquare 2mo agogiven that i couldn't even use fable without it downgrading to Opus, this is just a straight upgrade for me
- Freedumbs 2mo agoOpus 5 also downgrades. it's now Fable -> Opus 5 ; Opus 5 -> Opus 4.8. Unclear why they want to nerf their own products with sometimes right classifiers. I guess the government ban might've been real and not coordinated marketing?
- kettleballroll 2mo agoI don't understand why people believe this conspiracy theory of "oh the government ban was just marketing". That claim feels so incredibly ridiculous to me. It cost Anthropic a ton of money and reputation, and worst of all: it absolutely killed their competitive advantage. They were 1-2 months ahead of OpenAI, but trump conveniently gave OpenAI the time they needed to catch up and push 5.6 out the door without having to lose their subscriber base to the competitor.
- benjiro29 2mo ago> At half the price and less likely to auto-downgrade, it sounds like a reasonable claim Two benchmarks (artificial analysis and vals) show a increase in cost (a insane increase for vals compared to Opus 4.8). Already posted this before, so here is the link. https://news.ycombinator.com/item?id=49041158 https://news.ycombinator.com/item?id=49041158
- MagnumOpus 2mo agoBut better cost for the same performance. According to AA, Opus 5 _medium_ is as smart as Opus 4.8 _max_, at 1/3 the cost and twice the speed. And if you need a better response, you can turn it up to 11.
- Aurornis 2mo agoUsing the most expensive model for all of your agentic coding work hasn’t been good practice for a long time. Not unless you have infinite money to spend. Fable is typically used for key planning, architecting, and review tasks. I think this is a case where you don’t understand the use case, not that the marketing department is making mistakes.
- airstrike 2mo agoThey cost the same if you're already at $200/mo
- akmarinov 2mo agoEh, not really. Fable does a lot better on coding than Opus 4.8. Just this past week Fable was able to figure out a couple of small issues for me where Opus was failing to. Also both are still somewhat bad at UI implementation. Opus more so
- tshaddox 2mo agoThe blog posts figure cites Frontier-Bench for its agentic coding score, and shows Opus 5 beating Fable 5 43.3% to 33.7%.
- ActivePattern 2mo agoI think you're being overly cynical here. First, I don't see any claim that is the world's best model for agentic coding. Second, it is absolutely the best model in terms of coding performance vs. dollar, and it's raw performance seems very close to the frontier.
- adam_arthur 2mo agoGPT 5.6 is far more token efficient at most tasks with similar performance. Especially so for Opus 4.8, still to be seen with Opus 5. Where are you getting cheaper per dollar?
- ActivePattern 2mo agoHow are you supporting the claim that GPT 5.6 is "far more token efficient" than Opus 5? Tokens equal, output is cheaper for Opus 5 ($25/1M) than GPT-5.6-Sol ($30/1M), and it seems to outperform slightly on agentic coding benchmarks.
- adam_arthur 2mo agoThe first chart in the blog post shows a similar $/performance curve to GPT 5.6. Where 5.6 has optionality to run much cheaper along the same performance curve at lower thinking levels. There's a later chart that shows Opus 5 ahead, but seems like an esoteric benchmark rather than for common use. (Novel problem solving) If they had a more efficient model at coding they would lead with that chart.
- km144 2mo agoHere is one data point for cost: https://artificialanalysis.ai/models?cost=intelligence-vs-cost-per-task&model-filters=large-models%2Cproprietary%2Creasoning-models#cost-tabs https://artificialanalysis.ai/models?cost=intelligence-vs-co... Here is another data point for output token efficiency: https://artificialanalysis.ai/models?cost=intelligence-vs-cost-per-task&model-filters=large-models%2Cproprietary%2Creasoning-models&intelligence-index-token-use=intelligence-vs-output-tokens-per-task#intelligence-index-token-use-tabs https://artificialanalysis.ai/models?cost=intelligence-vs-co...
- manojlds 2mo agoWhich numbers are you seeing? It does show that it's better than Fable 5 in most things related to coding?
- jsLavaGoat 2mo agoIn my opinion, the frontier is passed what is really needed for coding. Fable is good as a supervisor.
- dbbk 2mo agoYeah I spotted this immediately too. I'm sorry. You're supposed to be a multi billion dollar company and you can't even highlight your chart honestly?
- unclebucknasty 2mo agoRecent releases have said something to the effect (paraphrasing here): "Use <less expensive or older model> for everyday tasks and <other non-critical stuff>. Use <more expensive or recent model> for complex coding tasks, refactoring large code bases, etc.". Then, the next model/release emerges and the previous "best for complex" gets demoted to "everyday". Obviously, it's all relative. But, it does beg the question: was the previous model really good for complex coding tasks or no? I mean, how is it now suddenly only good for the "easy" stuff?
- hvb2 2mo ago> I mean, how is it now suddenly only good for the "easy" stuff? Because your expectations have changed.
- unclebucknasty 2mo agoI'm sure the marketeers would love for the public's assessment of complex versus easy to conveniently shift per their release cycles; or for the public to simply forget their prior marketing.
- toephu2 2mo agoAlso it scored worse on DeepSWE than chatgpt 5.6 sol
- throw03172019 2mo agoAnd no data retention for 30 days.