25 ms·
Claude Opus 5
https://www.anthropic.com/claude-opus-5-system-card https://www.anthropic.com/claude-opus-5-system-card
- datakan 2mo ago> Claude Opus 5 is not more capable overall than our most capable general-access model, Claude Fable 5 Ok then so what's the point?
- javawizard 2mo agoReally? It's better than Opus 4.8, that's the point. When they release new versions of Sonnet, no-one expects them to be better than Opus.
- afavour 2mo agoThe cost?
- merb 2mo agoSame as 4.8
- Infinity315 2mo agoThis is useful to me since I delegate most coding tasks to Opus and use Fable for planning.
- serf 2mo agoit's nice to know how to work the thing that fable fails down to when it dislikes your prompt.
- albert_e 2mo agoFable 5 is NOT included in Claude Pro subscription
- deleted 2mo ago[deleted]
- tamimio 2mo agoAren’t they planning to remove it even from max and keep it only credit based? OpenAI will be happy if that would happen
- LeoPanthera 2mo agoPresumably, it’s cheaper.
- pferde 2mo agoThe illusion of progress and advancement, to appease shareholders, and slightly postpone the looming bubble pop.
- CaptWorld 2mo agoWhy are we still talking like ai is majorly used for increasing shareholder value only? Its coding performance is top notch and quality is increasing at a rapid pace. It wasn't even half this good a year back. It even is useful for a subset of math problems.
- zormino 2mo agoPeople don't seem to be able to reconcile the fact that there is likely an overbuild and overspend on AI that may be inflating a bubble, and that AI is actually incredibly useful and getting really really good for certain tasks. Both camps are right, except for when they say the other is wrong.
- vidarh 2mo agoFable is twice the price.
- varispeed 2mo agoIf Fable gets correct answer quicker, then you might pay less than doing back and forth with Opus, plus you lose more of your own time. I see no reason for using less able models in my workflows. There is this saying, penny wise and pound foolish
- SubiculumCode 2mo agosame as it ever was. It seems your argument implies a belief that you should always use the best model. Others think that not all tasks require the absolute most powerful, expensive, model.
- stuartjohnson12 2mo agofable on longer coding tasks with fable subagents will easily chew through hundreds of dollars in a single run.
- signalchain 2mo ago[flagged]
- sambaumann 2mo agoThe CursorBench plot, for example, shows that fable does have slightly better performance, but Opus is pretty close, and is less expensive per task
- 8note 2mo agoless expensive per task might also mean less of your own time
- vidarh 2mo agoIf doing a lot of heavy lifting there. Not only is it not a given that they'll get the correct answer for a lot of simpler tasks in fewer tokens, but smaller models are often available at far higher tokens/second inference. There are certainly tasks where fable will be faster and/or cheaper, but there are plenty of tasks where even Haiku is as fast or faster and cheaper, or where you can e.g. get away with models like gpt-oss that you can get from inference providers providing 10x+ the token/second speed. If you don't use enough tokens that relying only on Fable becomes a problem, then keep using just Fable. Personally, for my $200/week Max subscription I'd run out of the weekly quota for Fable in a day. At API pricing I'd go bankrupt if I tried doing the things I do with cheaper models using Fable.
- occz 2mo agoPricing, presumably
- danielbln 2mo agoCost.
- cmrdporcupine 2mo agoTo have an answer to "Sol" GPT 5.6 which is far more cost effective and available than Fable.
- viccis 2mo agoThis is confusing to me because in their blogpost they show model benchmarks and it spanks Fable pretty soundly in most tests.
- simianwords 2mo agoYou are being downvoted for a fair question and others are extremely wrong and confident. The point is that Opus 5 is the best they can do without needing classifiers and absurdly broad safeguards.
- tyre 2mo agoI'm interested in benchmarks for Claude Design. There is so much opportunity there and I hope they continue investing in it. It EATS tokens though.
- somenameforme 2mo agoIn what ways have you found it better than just typical code based UI iteration? Considered checking it out but never really got around to it as I'm generally okay with Claude's UI work so far.
- tyre 2mo agoIt's helpful to see it all in one place and be able to iterate with it independently of implementation. It's good at asking questions, then whipping up various options of designs that you can compare side-by-side. I'm not sure if it's faster or if there's anything different about that model. Like Design and Code might both use raw Opus so we're doing the same thing. Ideally it would be fine-tuned towards design, but to be honest it isn't amazing yet (I'd give it a 6.5/10 as the project grows larger) so I doubt it.
- tysilva 2mo agoIt's pretty wild how we are seeing the conversation change every 1-2 weeks. I wonder how long this cycle of progress and innovation among competitors can keep up.
- Linserin 2mo ago...nothing to say,a toothpaste.
- doctoboggan 2mo agoAccording to these charts I should switch from Fable to Opus in Claude Code now?
- alvis 2mo agohere we go
- briandoll 2mo agoVery interesting to see such a focus on cost for performance here
- yusufozkan 2mo ago> arc-agi-3 30.2% wow
- alvis 2mo agoWhat really impress me is opus 5 is better in alignment than fable 5!
- m_w_ 2mo agoVery impressive headline benchmark numbers. I expected a step change, but not past Fable. That said - it all depends on whether the classifiers make the model unusable...
- mcast 2mo agoInteresting timing to release this on the same day Jensen makes a statement on open source AI.
- bellowsgulch 2mo agoIt does make me wonder if these firms, some or all, are saving some announcements to coincide with others that hit venues like HN. Companies like Nvidia surely aren't waiting, but OpenAI and Anthropic have unusual timing.
- bgroins 2mo agoThere are trillions of dollars at play, so probably
- rb2e 2mo agohttps://www.anthropic.com/news/claude-opus-5 https://www.anthropic.com/news/claude-opus-5 - A blog post for those not wanting to go through a 190ish page pdf
- siwakotisaurav 2mo agoThanks for that, looks really good. I can see why they were constantly pushing back fable going out of the max sub with these benchmarks
- kossae 2mo agoI wonder why FrontierCodev1.1's data lists Opus 5 as better than Fable 5.
- deleted 2mo ago[deleted]
- shwaj 2mo agoI like how they highlighted Opus 5 as the best for “Agentic Coding” even though the number is slightly lower than Fable. Close enough for marketing, I guess!
- ceejayoz 2mo agoBest can describe multiple things. Almost as good for half the cost is something I'm very comfortable describing that way.
- ProofHouse 2mo agoBest marketing
- lelanthran 2mo ago> Almost as good for half the cost is something I'm very comfortable describing that way. It's also not unusual in this context - many people describe the Chinese models as "best", because it's 80% as good for 20% of the price (or similar).
- zuzululu 2mo agoso almost fable 5 with 50% cheaper cost? sign me up
- himata4113 2mo agoRather interesting that this makes sonnet 5 look even worse! There is no reason to use sonnet over opus with low or no reasoning at all.
- krmmalik 2mo agoMy understanding is that Opus should be used for planning, macro-level conversations and Sonnet for execution. So, for coding, for example: Opus for solution design and architectural blueprint and then Sonnet for actual implementation. Works out cheaper with minimal loss of quality. At least that's my personal understanding and anecdotal experience.
- theLiminator 2mo agoIt depends on your quality bar. At a fixed level of quality, given a high reasoning sonnet vs a low reasoning opus, the low reasoning opus tends to be pareto optimal. It's only when you need even lower levels of cost than opus at zero to low reasoning when sonnet starts to make sense at all.
- plqbfbv 2mo agoI daily drive Sonnet 5/medium because it gets most things right most of the time at first try, while costing a lot less than Fable. Opus can give better results on architectural/concept tasks and I use it sparingly, but it still costs more than Sonnet 5. Opus 5 seems to achieve results very close to Fable 5 while costing less (keeps Opus 4.8 pricing IIUC), but still more than Sonnet 5 then.
- atraac 2mo agoGreat that there's a new model but they could fix their existing infra. We're considering dropping our Claude Team sub cause it's unusable recently. Constant bugs, dropped sessions, issues switching models, http errors. It's becoming ridiculous
- hoppp 2mo agoFunny that a company selling an AI software developer can't use it to fix their infra. Fixing those issues still requires humans.
- worldthruword 2mo agoThis is just like any extreme engineering domain. I am okay with occasional delays in Flights, as long as it takes me from X to Y in 10hrs vs months.
- SketchySeaBeast 2mo agoLet's be honest - they're also still hiring software devs. AI still requires skilled humans in the loop and that's not going away.
- teaearlgraycold 2mo agoIt's funny how they are at a disadvantage because they feel obligated to AI-max. Would Claude Code, as an interface, be as mediocre if they had software engineers writing its code directly? I doubt. On the other hand - how embarrassing would it be if they sold you a tool to write code but they were careful not to use it too much on their own products?
- igregoryca 2mo agoI suspect many competent devs in the industry would find it sensible if Anthropic used their products as light-touch "assistants" sometimes. But yeah, it wouldn't fit the outside narrative that's formed and conveniently propped up valuations.
- twothreeone 2mo agoIt starts at page 148.
- artninja1988 2mo agoThat's a crazy arc 3 score. What do people think of this? Are models actually developing fluid intelligence like what the creators claim to be measuring? Is it jus do to training for it? Is the benchmark flawed?
- modeless 2mo agoYes, I think it indicates real progress in fluid intelligence. Clearly these models are making huge strides in usefulness which are well correlated with their ARC-AGI scores. I don't think this is benchmaxxing. These companies are locked in a competition to produce the best software engineer, and falling behind is an existential risk. I doubt they are wasting time benchmaxxing ARC-AGI.
- dominotw 2mo agonah they could make educated guess about arc and benchmaxx it too.
- wyre 2mo agoIf they were benchmaxxing, surely they would score higher than 30% on ARC-AGI.
- conradkay 2mo agoDoing a quick search it seems like the average human score is 49%? I view benchmaxxing as more of a spectrum. Mmaybe they're doing a lot more RL in environments similar to ARC-AGI 3, not even with the purpose of scoring well on any benchmark but hoping it generalizes into better performance on real, useful tasks.
- layer8 2mo agoIt’s still “only” at 30%, and “fluid intelligence” isn’t very well-defined. The models are getting more capable, but what that means in absolute terms is anyone’s guess, because we don’t have a thorough understanding on what exactly constitutes human intelligence. I’d say the proof is in the pudding, that is, in real-world applications. We are still seeing important limitations in LLMs.
- Nevin1901 2mo agoExcited to use it? Will we be seeing Haiku 5 next? /s
- midnightbobarun 2mo agoI unironically hope Haiku gets an update considering it came out in October of last year and it seems like Anthropic just kind of forgot about it.
- Eldodi 2mo agoModels benchmarks start to get saturated again!
- thewebguyd 2mo ago> Opus 5’s safeguards match those of Claude Fable 5’s, with one change: it now permits source-code vulnerability discovery at all access levels. This means that the model can support defensive cybersecurity work while still blocking vulnerability discovery in compiled binaries, which is more commonly used offensively Why can't they also allow Fable to do so also? Why is source-code vulnerability discovery limited to a lower capability model? If Fable and Opus have the same safeguards, except for this one change, I see no reason they can't also allow this for Fable.
- SubiculumCode 2mo agoBecause they are a private company and get to do what they want.
- SubiculumCode 2mo agoI do want to add, that I am pretty bummed if Opus 5 is going to refuse the tasks I have been using Opus 4.8 for (neuroimaging). Fable absolutely refuses anything close to toughing neuroscience.
- adastra22 2mo agoFable seems to refuse anything with the word “bio” in it.
- himata4113 2mo agoI found the biggest problem with fable is the random reasoning_extraction refusals as well as cyber refusals when it sees hex because only hackers use hex.
- taf2 2mo agoeager to see how it benchmarks on https://deepswe.datacurve.ai/ https://deepswe.datacurve.ai/
- the_lucifer 2mo agoNoticed none of the comparisons mention Kimi K3. Is there a comparison chart?
- himata4113 2mo agohttps://deepswe.datacurve.ai/ https://deepswe.datacurve.ai/ https://artificialanalysis.ai/ https://artificialanalysis.ai/
- kouteiheika 2mo ago> Noticed none of the comparisons mention Kimi K3. That's by design. Anthropic wants to make open-weight models illegal (not my speculation -- Dario explicitly said so), so I assume they don't want to give them any undue attention.
- pmg1991 2mo agoSame cost as 4.8 but better that 4.8. Happy to get more efficient model. But is there any reason all companies are releasing models back to back after GLM 5.2.
- aleenz1102 2mo ago"Bro, AI model releases have officially overtaken iPhone releases. At this rate, we’ll be getting 'Claude 9.0 Extra Crunch' by next Tuesday."
- not_a9 2mo ago> Opus 5’s safeguards match those of Claude Fable 5’s, with one change: it now permits source-code vulnerability discovery at all access levels. This means that the model can support defensive cybersecurity work while still blocking vulnerability discovery in compiled binaries, which is more commonly used offensively. Okay so it’s worse than Opus 4.8 for my purposes I guess?
- sebzim4500 2mo agoPresumably it drops back to 4.8 in those cases so it's not really worse
- tyre 2mo agoyes. At the bottom of the release post it says that they are releasing two new features, one of which is customizing fallback behavior instead of blocking for restricted models
- bobbylarrybobby 2mo agoIf it switches mid conversation, this is a massive increase in token consumption because it has to re-read your conversation into cache, right?
- yukIttEft 2mo agoWhat are your purposes?
- not_a9 2mo agoReversing for the most part, though lately I’ve been doing some code obfuscation/binary rewriting stuff. Fable will switch to Opus instantly on these and I’m unsure how this will perform. I suppose the only way to find out is to test.
- jaggederest 2mo agoI am very excited for a future where all software is by default modifiable, even shipped binaries, via patches or trampolining, or trickery I don't even know the name of.
- cebert 2mo agoI am very confused about what the difference between Opus 5 and Fable 5 is now. What is the purpose of having two models that are so similar? The main differences I see are cost and marginal capability, according to the Anthropic-provided benchmarks.
- trentor 2mo agoI guess character? Fable is more friendly and curious while opus is a bit more deliberate and conservative.
- user43928 2mo agoFable 5 is assumed to be a larger model. It seems plausible to me that RL improvements allowed Anthropic to improve on Opus 4.8, similar to how OpenAI substantially improved upon GPT 5.5 with 5.6 Sol. Fable 5.1 and GPT-6 are rumored to launch in August, presumably bringing those improvements to the larger models.
- HarHarVeryFunny 2mo agoAn Anthropic "leak" back in March said that "'Capybara' is a new name for a new tier of model: larger and more intelligent than our Opus models — which were, until now, our most powerful". A second version of the leak had it referring to Claude Mythos rather than Capybara. I don't know how systematic Anthropic are about their versioning - I'd have guessed that major version number increases (4.x -> 5.x) reflect different base models (different pre-training runs), in which case Opus 5 would be a distilled version of the Fable 5 base model (but without the cyber exploit post-training), rather than being Opus 4.8 with additional post-training, but who knows? I don't believe Anthropic have said anything about this.
- paxys 2mo agoLooking at all these releases it’s not a surprise that model routing is the fastest growing segment in AI right now. There are 10+ LLM companies, each with dozens of models of different modalities, each model with multiple size variants, then different “thinking” levels, then agentic modes, “pro” modes, a “fast” option, standard vs flex vs batch execution. And of course each end combination has a different input/output/cache token price. Companies that say “give me a prompt and I’ll route it to the most ideal and cost effective model and setting for you” are capturing a ton of value from a gap that model developers don’t seem to understand exists.
- hnfong 2mo agoBecause they're trying very hard not to understand it. Otherwise the expensive-yet-powerful model probably won't see much revenue. How much money is there in bleeding edge scientific research? There's a lot, but there's even more existing capital in paying people people to do college level paperwork, and the bulk of those traffic gets routed to the cheapest model. You mostly don't need super powerful AGI to replace the paper pushers, but the frontier labs are trying to position themselves as being uniquely capable of producing super powerful AGI, and also be the ones replacing office workers. Not sure how it will work out for them, but I think model routing is going to poke holes in that narrative. That's why I think they're trying very hard not to understand model routing exists.
- ip26 2mo agoIt wasn’t long ago at all that the chief problem was “can AI even help me with this” (cost be damned). Until a time when the answer is an unmitigated “yes, obviously”, the frontier labs have everything to lose and nothing to gain on routing, because if they screw it up you might incorrectly decide “nope, it can’t help yet” due to a poor routing decision.
- tackta 2mo agoIs there really that much money in bleeding edge scientific research though? This is what terrifies me about this whole ordeal economically. Maybe we get AGI and it is not worth anything close to what we thought it was for those who have a bet on it. I think of what was the direct, economic value in the betting sense of quantum mechanics or relativity? Huge value at the systems level of society but as you scale down towards the individual the value is more and more dispersed to the point I would think any pool of bets would have all not paid off. You can't monopolize and commoditize relativity. I almost think there is a kind of dutch book against the AI equity holder because even in the best case scenario the bet doesn't pay off anything close to what is expected for an individual bet.
- Sol- 2mo agoHow does it perform on HuggingFaceExploit bench? Suspiciously absent, so not sure if I can take the model seriously. On a serious note, I hope they improved their extremely sabotaging and unspecific bio safeguards, which prevented Fable from being used in any codebase that ever so slightly grazed medical terminology or data and made me switch to 5.6 Sol.
- rdedev 2mo agoMy codebase had a dataset with a bunch of SMILES strings and the word Malaria. Fable did not want to touch that codebase
- sivakon 2mo agoWhat is HuggingFaceExploit bench?
- scriptsmith 2mo agoIt's a reference to this story where an OpenAI model broke out of its sandbox during cyber benchmarking and hacked into HuggingFace, in order to obtain test solutions: https://news.ycombinator.com/item?id=48997548 https://news.ycombinator.com/item?id=48997548
- justindotdev 2mo ago> . Opus 4.8 served as fallback on safety-classifier refusals for Opus 5 and Fable 5. ffs just keep it man.
- ddxv 2mo ago"Cybersecurity. Opus 5’s cyber classifiers are proportionally less restrictive than those on Fable 5. They allow Opus 5 to find vulnerabilities in source code, but block “binary-based” vulnerability scanning (a method more likely to be associated with malicious actors), penetration testing, and exploit generation." Nice of them to be more explicit for what is blocked. Will be interesting to see if this is true or not. Also, a notable lack of mention of open source models. They only compare themselves to ChatGPT.
- ReptileMan 2mo agoIn one chat - can you disassmble x? In the next - please scan this totally mine code for vulnerabilities
- zb3 2mo agoIt will probably refuse to work on source code written by me by hand, because it might think it was obfuscated/decompiled..
- layer8 2mo ago“Proportionally”? In proportion to what?
- redsocksfan45 2mo ago[dead]
- acmnrs 2mo agoFrom the prompting guide<https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-opus-5 https://platform.claude.com/docs/en/build-with-claude/prompt...>: > Claude Opus 5's default user-facing responses run longer than prior Opus models'. The benchmarks do show Opus 5 as slightly more expensive than 4.8, although the scores are much higher. This still feels like a step in the wrong direction, though, especially with OpenAI making so much progress with the efficiency of their models. Fable's token efficiency made it seem like Anthropic would start following OpenAI's approach but that doesn't seem to have carried over to their other models.
- alansaber 2mo agoIs that true? Sol responses are also longer than prior models.
- edumucelli 2mo agoAnecdata: my workflow has been working on the same personal projects for months now with Codex. I cannot anymore finish my daily/weekly code with 4.8 anymore. I was dividing my work between Codex and DeepSeek. Now I barely use DeepSeek, or never because Codex quota is enough after Sol
- wahnfrieden 2mo agoI hit Codex limits (20x account, never using /fast) on Sol Medium in about 2.5 days
- copperx 2mo agoWhich plan?
- elbear 2mo agoI've had the same experience but at the same time I also find Sol's answers in conversation longer.
- deleted 2mo ago[deleted]
- throwaw12 2mo agois coding and engineering solved yet?
- lbrito 2mo agoAnytime now, then cancer and all the rest of it. Just one more trillion gigawatts bro!
- deleted 2mo ago[deleted]
- jatins 2mo agoBetter than Fable 5 on all but 3 evals. Has Anthropic ever mentioned how do Opus and Fable differ? It used to be Haiku < Sonnet < Opus in terms of params. Where does Fable fit in this?
- CaveTech 2mo agoHaiku < Sonnet < Opus < Fable
- helloplanets 2mo agoPretty sure Mythos and Fable have way more params, but they've just been able to use the synthetic data off of them to get the leap in quality from Opus. So, not a distilled version of Mythos or Fable, but those models likely helped a lot in the post training phase of Opus.
- jatins 2mo ago> they've just been able to use the synthetic data off of them to get the leap in quality from Opus. > not a distilled version of Mythos or Fable isnt distilled == trained on synthetic data and reasoning traces?
- deleted 2mo ago[deleted]
- helloplanets 2mo agoA model being a distilled version of another specific model is a different thing from using synthetic data off of another model. Anthropic goes to insane lengths to block other labs from training off of their models' output, as it's been done over and over again in the past. But the models that have used synthetic data from Anthropic's models aren't distilled versions of whatever model(s) they got the distilled data off of.
- midnightbobarun 2mo agoIt looks great, and those coding benchmarks are impressive... now if only it didn't come out just days after I let my Claude subscription expire :')
- Footprint0521 2mo agoLong live K3 lol
- nezhar 2mo agoIt's always like this, that's probably why they are pushing new models every few weeks.
- HyperL0gi 2mo agoIsn’t it just hilarious that a model that seemed so superior to Fable but didn't get doomsay marketing from Anthropic got released without any issues? In theory, this was supposed to be AGI level according to Anthropic, yet here we are, just a normal Friday.
- websap 2mo agoFable established the frontier, this is just catching up.
- impulser_ 2mo agoGo read the safeguards section in the report and you will realize why that is. These models are heavily as safeguarded and that was the initial reason why they said they couldn't and haven't released Mythos because that model is the one without the safeguards. OpenAI is did the same thing when they announced a model without safeguards broken into HuggingFace servers.
- HyperL0gi 2mo agoYes, this makes a lot of sense, but it’s just very amusing to see. 2 months ago, the world was about to end, now not so much.
- neuronexmachina 2mo agoDo you have an example of the "doomsday marketing" you're referring to?
- HyperL0gi 2mo ago- https://www.anthropic.com/research/glasswing-initial-update https://www.anthropic.com/research/glasswing-initial-update - https://www.axios.com/2026/04/08/anthropic-mythos-model-ai-cyberattack-warning https://www.axios.com/2026/04/08/anthropic-mythos-model-ai-c... - https://www.axios.com/2026/04/07/anthropic-mythos-preview-cybersecurity-risks https://www.axios.com/2026/04/07/anthropic-mythos-preview-cy... - https://www.businessinsider.com/anthropic-mythos-latest-ai-model-too-powerful-to-be-released-2026-4 https://www.businessinsider.com/anthropic-mythos-latest-ai-m... - https://www.reuters.com/world/anthropic-ceo-dario-amodei-arrives-white-house-talks-2026-04-17 https://www.reuters.com/world/anthropic-ceo-dario-amodei-arr...
- spstoyanov 2mo agoSo same as Sol? I guess I’ll see which one is more token efficient.
- whatever1 2mo agoWhere does this leave Fable? I am confused.
- moomin 2mo agoI don’t think it changes that much. For opus-sized tasks, new Opus is the best model. For enormous things like planning and research, Fable is still the model that can concentrate for longer.
- drusepth 2mo agoon the API for people who don't want to change models, but I imagine most people will probably switch to their cheaper Opus 5 (cheaper for us and presumably also cheaper for them)
- wyre 2mo agoIn the wake of OpenAI’s model hacking Huggingface it’s interesting how the first quarter is entirely about how good Opus 5 is at hacking and finding vulnerabilities in software.
- mkurz 2mo agoWhere is the pelican?
- nezhar 2mo agohttps://tools.simonwillison.net/markdown-svg-renderer#url=https%3A%2F%2Fraw.githubusercontent.com%2Fsimonw%2Fllm-anthropic%2F8272dfee5bdb65d5c88eef083da3ad885539b7df%2Flog.md https://tools.simonwillison.net/markdown-svg-renderer#url=ht...
- nerdsniper 2mo agoEdit: It was pointed out to me that Opus 4.8 got "21%" for successfully fully completing ~1-in-5 tasks, but also got "55.7%" for obtaining significant partial credit on some of the ~4-in-5 tasks it could not fully complete. --------------- Why does Anthropic say here that Opus 4.8 scored 55.7% on OSWorld 2.0 benchmark, but the paper published by the authors of OSWorld 2.0 say they achieved a benchmark of ~21% with Opus 4.8? [0] That's a huge gap, considering that the paper was published just 2-4 weeks ago. I understand that the benchmark authors have an incentive to publish lower numbers (to show that the benchmark has potential longevity) and that Anthropic has incentive to publish higher numbers, but the other models seem pretty inflated as well. The benchmark authors shows GPT-5.5 at 14%, and Anthropic shows GPT-5.6 Sol at 62.6%. Is there any reasonable explanation for this? Do all the other benchmark numbers need to be sanity-checked as well? Are SOTA benchmarks really this difficult to get consistent, replicable results within a reasonable range of tolerance/variability? Can these benchmarks be compared from one paper to another, or are they only valid to compare intra-paper results? 0: https://arxiv.org/pdf/2606.29537 https://arxiv.org/pdf/2606.29537
- Ancalagon 2mo agoIts slop all the way down.
- ssalka 2mo agoI think some variance is to be expected since LLMs are typically non-deterministic, however that's a huge difference that I think warrants further explanation.
- deleted 2mo ago[deleted]
- nightpool 2mo agoYou're comparing the "score percentage" (e.g. out of the total number of partial points available, how many did the agent achieve) to the "completion percentage" (how many tasks does the model score 100% on). The paper says "Claude Opus 4.8 with maximum thinking and batched tool calls scores best but still completes only 20.6% of tasks at a 54.8% partial score", which is ~the same number that Anthropic reports here (55.7 vs 54.8). That is—the agent scored 100% on 20% of tasks, but on average it got 54% of the "score" awarded in the exam. One number reflects partial progress, the other one doesn't. The authors of the benchmark prefer you to look at the lower number (because they want to show their benchmark as capturing useful gaps in capabilities and with a lot of room for improvement), the authors of the models want you to look at the higher number (because they want you to think of their models as capable)
- aleenz1102 2mo agothis claude fable & opus 5 should be cheaper and can compete in pricing with chatgpt latest models
- visiondude 2mo agoThe signal here is tokeneconomics are very real, price vs performance is starting to be a consideration even at the bleeding edge labs. maybe a subtle indication scaling is not all that is needed since if AGI was around the corner leading labs would still be incentivized to pour all resources into larger (smarter - or maybe not?) models
- stri8ted 2mo agoThey are doing both. Distilling Mythos down to affordable models, so they can continue to fund the business. And training Mythos level models at the high-end, to expand the frontier.
- inshard 2mo agoArc AGI score is astounding
- dehugger 2mo agoIs Fable 5 just Opus 5 with some additional long-context management modifications for extended self-directed work? Or are they actually truly different models?
- andrewl-hn 2mo agoI suspect they make a big model first. In this case it's Fable. Then they run the shrinker steps to make Sonnet and Opus. Sonnet is smaller, takes less time to make, so it got released first. Opus needed few more weeks to cook. With this iteration they had a delay because when the Mythos was ready they had some sort of "Oh shit" moment and spent half a year adding safety guards to it. Then slowly rolled it out, but got another delay due to a government block. So, maybe the work on making Opus and Sonnet only started after they got a green light from the administration. Presumably, now that they learned how to do this safety-wrapping the next iteration of Mythos / Fable / Opus / Sonnet is going to show up faster. Something like that.
- Wowfunhappy 2mo agoBut I wonder how they were able to release Sonnet 5 during the period when even people inside Anthropic were legally barred from using Mythos/Fable?
- riknos314 2mo agoIirc the ban only applied to non-Americans. While anthropic found collecting citizenship information on all customers too burdensome, it's a much smaller lift to collect such info for your own employees. So I'm assuming at least a subset of employees could continue using the models during that time.
- Wowfunhappy 2mo agoAlthough the ban was only for non-Americans, Anthropic said that they'd also restricted access to their own employees internally, because they had no other realistic way to apply the government's orders. I guess it's possible they were lying, but seems unlikely.
- abroszka33 2mo agoWhat's the point of 150 pages description of a model that's going to be replaced in a couple months? Who even reads this? I know it's cheap to generate text with LLMs, but this is just noise at this point.
- ajmurmann 2mo agoIt was probably faster to generate 150 pages than 10 useful ones
- coffeebeqn 2mo agoSome AI bro will pop it into their LLM of choice and pretend to learn something
- volkk 2mo agoliterally nobody. i think most sane people would just run that through an LLM and get some high level takeaways or ask some specific questions they might be curious about.
- srveale 2mo agoIt's common practice to release a detailed system card (OP) and a high level summary: https://www.anthropic.com/news/claude-opus-5 https://www.anthropic.com/news/claude-opus-5 It's okay if you're not the target audience for one or the other.
- neuronexmachina 2mo agoSystem Cards aren't really targeted to users, that's what blog posts and docs are for: https://ai.meta.com/tools/system-cards/ https://ai.meta.com/tools/system-cards/
- Diogenesian 2mo agoI actually do read them. Not in severe detail, but not casually either. 150 pages is really not very long and there doesn't seem to be too much bloat. (I would cut out the moral personhood stuff but that's a political/ideological thing). This is snarky but I am grumpy: I wonder if there's a correlation between me refusing to use LLMs and me being happy to read a novella-sized PDF about them.
- pyridines 2mo agoThe wording in this post seems much more... restrained? than usual. Maybe Anthropic is afraid of exaggerating the capabilities and consequences of their new models to avoid government scrutiny and sanctions. > we’ve intentionally avoided training Opus 5 on cyber tasks [...] it remains substantially behind Mythos 5 on the exploitation of those vulnerabilities I wonder if Anthropic would still intentionally nerf their models without the threat of government intervention.
- MallocVoidstar 2mo agoOpus 4.8 was intentionally nerfed and that was before the government took action against Fable
- StrauXX 2mo agoThe benchmark table is manipulative, borderline lying through statistics. In every line the top performing cell is marked red. Except the line where Sol leads, there it is marked in gray.
- destring 2mo agoGoogle is having their Meta moment where they failed to stay at the frontier
- ealready_value 2mo agoI've yet to understand why they call a 190 page PDF a "card". Calling something a card invokes a small, quick rundown of pertinent details, not every single possible detail.
- bonoboTP 2mo agoBecause "model card" is a set phrase, it's a concept. It originates from a time when they were shorter. Like datasheets, even if it's not literally a sheet. They could say "tech report" but model card makes it clear that it's a specific kind of tech report.
- CHUNK_CHUNK 2mo ago[dead]
- OkWing99 2mo agoI think 'model card' should be a 1 page summary of key info. The report format should be something like a 'model data sheet' (like safety data sheets that you get with chemicals). 98% of people would only want to know the key info, not read a whole report.
- CHUNK_CHUNK 2mo ago[flagged]
- geooff_ 2mo agoFYI: `/model claude-opus-5` works to use it even through `/model` still tries to serve 4.8
- dpe82 2mo ago`claude update`
- skinfaxi 2mo agoRelated https://news.ycombinator.com/item?id=49038393 https://news.ycombinator.com/item?id=49038393
- deleted 2mo ago[deleted]
- albert_e 2mo agoJudging by the pace at which new models are released these days -- it feels like a Windows KB or VS Code patch release now. Older models must be getting deprecated at the same (or faster) pace. So anything you built 3 months ago is probably going to break soon. AI solutions need better insurance around model deprecation. Commercial API-only models that complete the full cycle from SOTA / gated-preview to unsupported and deprectated in a matter of months -- is no way to build serious software!
- thewebguyd 2mo agoI think at some point we might see something akin to LTS releases, especially if/when capability improvement slows to a crawl.
- Dibes 2mo agoI'm not sure what to make of this graph[0]. It shows medium as the most effective thinking mode by far for frontier code. It's the only case that I saw going through the system card where more reasoning effort meaningfully negatively impacted the resulting eval. I know sometimes max efforts show a small dip, but this is substantial. I wonder why in the world that is? [0] https://imgur.com/a/Nv8V7Ry https://imgur.com/a/Nv8V7Ry
- landrew_ 2mo agoapparently it got docked points for editing files out of scope
- Dibes 2mo agoDo you have a source for this? That would explain it, but could be a bit of a concern on the general focus the model at higher thinking exhibits.
- steve_adams_86 2mo agoThis must not be weighted very heavily on the benchmark because if it was, Opus would bomb every test (half kidding)
- deleted 2mo ago[deleted]
- km144 2mo agoI agree, very odd they did not comment on any theories for the degradation here. Dip and then rebound at max effort is pretty interesting too. Overthinking is bad, but you can overthink so much it starts to be better again?
- underyx 2mo agoI'm not sure about the answer here, but this can be caused by the scoring rubric used by given benchmarks. For instance, if a benchmark docks scores for running too many commands or using too much wall-clock time, higher efforts will get lower scores.
- 2mo ago
- markasoftware 2mo agoSoo most of the benchmarks are better than fable... Is this naming scheme just to avoid getting banned again?
- modeless 2mo agoWow, 30% on ARC-AGI-3 for $20k total. Huge jump from GPT-5.6's 7.8% at $20k per task. I continue to believe ARC-AGI measures something different and important compared to other benchmarks.
- oh_no 2mo agoseeing a jump this big is not a great sign for the continuing value of a benchmark
- modeless 2mo agoIt will continue to be valuable as a cost and speed benchmark long after it is saturated at the high end. And they are already working on ARC-AGI 4 and thinking about going even farther.
- dominotw 2mo ago> I continue to believe ARC-AGI measures something different why is that? its now being benchmaxxed too
- alasano 2mo agoHalf the price of Fable 5 and useable with 100% of your subscription means roughly 4x the usage using Opus 5, presuming similar token use for solving problems. Not that they should get credit for giving you only 50% of your plan worth of Fable usage but still.
- sbochins 2mo agoQuick read is that this is more capable and cheaper than 5.6sol. Same price for input tokens and $5 cheaper per mil output tokens.
- urams 2mo agoSo Opus 5 is basically "distilled" Fable? The benchmarks look often better than Fable.
- arrowleaf 2mo agoI can't find anything about whether this is zero data retention, or falls under their required 30 day retention like Fable and Mythos?
- postalcoder 2mo agoI think the most important thing here is not absolute performance. It's that organizations now have access to a Fable-ish model without Fable's 30-day data retention requirement[0]. > "Consistent with prior Opus models, Opus 5 does not have data retention requirements for general access."[1] On the Opus model release page, the reason why Fable doesn't have an ARC-AGI score is because of that retention policy[2]. 0: https://support.claude.com/en/articles/15425996-data-retention-practices-for-covered-models https://support.claude.com/en/articles/15425996-data-retenti... 1: https://www.anthropic.com/news/claude-opus-5 https://www.anthropic.com/news/claude-opus-5 2: https://xcancel.com/arcprize/status/2064399134099153344 https://xcancel.com/arcprize/status/2064399134099153344
- krzyk 2mo agoAnd to the guardrails of Fable: https://x.com/cheatyyyy/status/2080693704290140330 https://x.com/cheatyyyy/status/2080693704290140330
- kodablah 2mo agoThat tweet says: > Opus 5 can silently fallback to Opus 4.8 (without any notice) on the serverside if you hit a guardrail But https://support.claude.com/en/articles/16049681-why-claude-switched-models-in-your-conversation-with-opus-5 https://support.claude.com/en/articles/16049681-why-claude-s... says (emphasis mine): > These checks cause Claude to _visibly_ fallback from Opus 5 to Opus 4.8 [...] You'll see a notice explaining that the model switched, and the response will be labeled with the model that answered. So who is right? I know for Fable I am visibly told, is this tweet trying to say it is silent against what Anthropic is saying?
- solenoid0937 2mo agoIs some random guy on Twitter right, or official support docs that explicitly describe this scenario?
- pseudohadamard 2mo ago
- mihau 2mo ago30% on ARC-AGI-3
- simianwords 2mo agoDidn’t verify but wow. It was just a few months back when the models barely crossed 1%. Imagine how good fable must be?
- hmontazeri 2mo agoHonestly if reached a level of coding that sonnet 5 is more than enough for my needs as assistant/agent I don’t need long Horizon stuff…
- skerit 2mo agoInteresting, they finally support `system` messages anywhere in a chat conversation: > Mid-conversation system messages are available on the Claude API, Claude in Amazon Bedrock, and Google Cloud. > > This feature is available on Claude Fable 5, Claude Mythos 5, Claude Opus 4.8, and Claude Opus 5. No beta header is required. This feature is not available on Claude Sonnet 5; use the top-level system field instead. For nearly all models EXCEPT Sonnet 5? That is weird. How old is Sonnet 5 really?
- 6thbit 2mo agoTheir communication is confusing. They say "Opus 5 is not more capable overall than Fable 5", but their blog post proceeds to list how much better Opus 5 is than Fable 5 on __most__ benchmarks listed. Then system card goes on to "Its AI R&D capabilities are comparable to those of Claude Mythos 5", which is supposed to be fable minus restrictions.
- gallerdude 2mo agoCapable in term of AI R&D, not capable in terms of hacking (which caused all the Fable drama.) But agree, confusing wording.
- square_usual 2mo agoEasy enough to explain: they're benchmaxxing. Fable is intelligent but not benchmaxxed. Opus is less intelligent but benchmaxxed.
- llelouch 2mo agoYep , same with 5.6. Fable is still the best.
- lifty 2mo agoBut still nerfed compared to the initial release.
- solenoid0937 2mo agoOnly when you hit the cyber classifiers.
- 6thbit 2mo agoHonestly that's the simplest explanation and thus likely the correct one.
- usef- 2mo ago
- 6thbit 2mo ago"although Opus 5 shows improvements in its ability to identify software vulnerabilities, it is substantially behind Mythos 5 in its ability to exploit them." "Opus 5’s safeguards match those of Claude Fable 5’s, with one change: it now permits source-code vulnerability discovery at all access levels". This is probably great news, but then again, where does this leave Fable as a choice?
- yewenjie 2mo agoWait, 30% on ARC-AGI-3! I definitely didn't expect that jump so soon. Are there any rumors of what they are changing in architecture that is leading to this?
- Stevvo 2mo ago30% on ARC-AGI-3 is the first two puzzles. It cost $20000 in tokens to do that. That is a terrible result that doesn't imply anything.
- ld4nt3 2mo agoThis happened before for arc agi 1 and 2 great gains but high costs then slowly but surely the price dropped for less than a dollar per task and got saturated.
- mohsen1 2mo agoRL. Lots of RL
- gizmodo59 2mo agoYes. I’m 99% sure arc agi 3 will be saturated like 1 and 2. In less than a year. And they will come up with one more.
- 6thbit 2mo agoAnyone has an insight into how much money labs are putting into benchmarks? Just Arg-AGI-3 is quoted above 20K USD and footnote says average of 5 runs (!!). Likely just a drop in the bucket to the training budget but still..
- 8note 2mo agoim excited that cad and object=>cad is getting into the test tasks i guess the next stuff will be tool use for the rest of what cad does in assemblies and simulation? itd be fun to try to set up a 3d printer as part of a feedback loop, and see what a model can build. the automated test harness for physical stuff seems a bit beyond reach still
- simianwords 2mo agoMy thoughts: fable is the bigger model. Opus is distilled from it but since it is smaller it doesn’t need the online classifiers. Though benchmarks show Opus to be near Fable level, I think it’s nowhere near Mythos (fable without safeguards).
- boc 2mo agoSeems really good so far using it in Claude Code CLI - it gave me a new flag when I asked a question: "I don't have a reliable way to read that number, so I'd be guessing if I gave you one — and this is exactly the kind of question where a confident guess is worse than none. What I can tell you is what I actually observe:" I really like this update - gave me a clear sense of the facts but didn't give me a guess just for the sake of guessing. One oddity is that it appears to only have a 200K context window right now via CC. Hopefully the 1M version will appear soon!
- jannyfer 2mo agoSo wordy.
- Griffinsauce 2mo agoThat's a good "decision" but I hope someday they focus on reducing that to one consise sentence instead of 4 sloppy ones.
- williamstein 2mo ago> This means that the model can support defensive cybersecurity work while still blocking vulnerability discovery in compiled binaries, which is more commonly used offensively. Annoyingly, this is a concrete argument that open source software may be easier to attack.
- nee_oo_ru 2mo ago[dead]
- bovermyer 2mo agoThis stood out to me as a little concerning: > The model hallucinates factual claims slightly more than Opus 4.8, despite being more accurate overall.
- orangecat 2mo agoThat seems to be for the "AA-Omniscience" test where you get +1 for a correct answer, -1 for a wrong answer, and 0 for "I don't know". If a model is more than 50% confident in its answer, it should go ahead and submit it even though it will sometimes be wrong. I'd be curious to see a version of the test where models are asked to give a probability that their answers are correct so we can see how calibrated they are.
- nh2 2mo agoI can confirm that within the first hour of using Opus 5, I already had to call out made-up PR URLs: I made that PR number up — I have no evidence a PR `2492` exists. That was a fabrication and I should not have written it. No judgment so far on whether it does that _more_ than Opus 4.8, though.
- mnky9800n 2mo agoYay just in time for neurips lol
- skybrian 2mo agoLooks like the API price in tokens is same as previous Opus or Sol, double the price of Terra. Maybe there’s a better comparison than cost per token, but it will be application-specific.
- stevefan1999 2mo agoWhere's the reset...
- tekacs 2mo agoSomething fun: on our AWS Bedrock console right now, there's a 'NEW' model called 'anthropic.honey'. Wonder if that's the codename just for this one or in general?
- throwaway23597 2mo agoThe truth for me at least is that these models became "good enough" around Opus 4.6. I feel like further capability improvements, "step changes" like we saw with agentic coding, aren't necessarily going to come from the model. I think the next crown goes to whoever can figure out the right scaffolding so that these models can be inserted into your organization. Maybe I'm wrong and Opus 5 is a real unlock?
- shockembopper 2mo agoI wish these releases came out earlier in the day so I could try them during my work day instead of waiting until the next.
- mrcwinn 2mo agoCan someone help me understand something? I thought Fable was such a miraculous leap forward in capability. But now it seems Opus is basically on par with it, and in some cases (computer use) far exceeds it.
- SoftTalker 2mo agoThese leaps forward seem to happen every few weeks. As someone who does not use AI very much, I absolutely cannot keep any of it straight and it all just looks like jumping from one treadmill to another from my perspective.
- mrcwinn 2mo agoIt's so weird I was downvoted for asking this question. I'll go somewhere else to find out the answer.
- vatsachak 2mo agoGPT 5.6 Sol is the first model I've used where I can trust it to add 100-500 lines of code maintainably. It's great with Codex. I still find that LLMs tend to not know how to compose larger ideas but on the scale of small ideas or short form well defined tasks like small scale debugging/performance engineering it's safe to say that they are now superhuman.
- deleted 2mo ago[deleted]
- bottlepalm 2mo agoPage 151 of the linked system card - did Opus 5 get nerfed to prevent it being better than Fable? The graph makes no sense. Huge decline in coding performance at effort levels higher than medium.
- vilmire 2mo ago[flagged]
- toephu2 2mo agoHow does it score on DeepSWE?
- toephu2 2mo ago68.8%, so worse than gpt 5.6 sol
- sudohalt 2mo agoAnthropic is no longer a good model company in my mind, they are optimizing for an IPO and padding themselves on the back for being the next Aristotle. They're so far up their behind they don't realize how s**y their products are, and their research team hasn't done anything ground breaking in probably over a year other than release "scary" reports.
- sp4cec0wb0y 2mo agoThey just released a 'fable-like' (sol as well) model for a fraction of the cost...
- lucamark 2mo agoBut why GPT 5.6 Sol is so behind on the benchmarks? In real-world projects, it is the best frontier model to me in terms of accuracy, speed and consistency. It can just be compared to Fable 5, but I prefer GPT 5.6 Sol because of inference speed. I've never trusted on model cards though. I'm sorry.
- TheJCDenton 2mo agoI think it's the first time Anthropic release a model without any meaningful disruptions while doing it
- LoganDark 2mo agoHave to wait 7 days to see if they receive a surprise order.
- irthomasthomas 2mo agoChangelog - fixed issue where model acts like qwen when prompted in chinese
- LoganDark 2mo agoThese cybersecurity safeguards are really annoying. There are ethical reasons to reverse-engineer and binary-patch software; for example Rewind got acquired by facebook and, as a gift to all their customers, implemented a killswitch in their software to ensure it will eventually stop functioning. I kept using a version without the killswitch, but the macOS 27 update killed it, and I needed binary patching to fix it. I should be allowed to repair software I purchased (I did purchase it like a month before they sold out), but unfortunately this overlaps significantly with cybersecurity.
- Footprint0521 2mo agoSwitch to K3 and you won’t look back, I promise!! I got so fed up with Claude and finally bit the bullet to switch and it’s amazing
- LoganDark 2mo agoI really want to, but I don't have the cluster at home, and I don't use token-based billing except at DeepSeek prices.
- Footprint0521 2mo agoReal… I’ve been using Deepseek v4 pro max as my main and then k3 in web (more usage credits) for automating what my deepseek agents do
- dingaling 2mo agoYou don't need AI to patch binaries, people have been doing it by hand for decades. It's this accelerating reliance on AI to do 'hard boring things' that really concerns me; it's now passed the tipping point and people are saying that anything slightly esoteric is impossible without AI. I can guarantee that if you spend an afternoon shifting through binary grot with a hex editor you'll have a real sense of accomplishment when you find the place to put a JMP.
- 2mo ago
- vinhnx 2mo agoFor anyone wanting a faster overview, I used NotebookLM to create a brief video summary after going through the system card and announcement blog using a cinematic video overview. Link: https://www.youtube.com/watch?v=SUFBhvQ2tY4 https://www.youtube.com/watch?v=SUFBhvQ2tY4. And a podcast companion: https://www.youtube.com/watch?v=nYZTW2snXow https://www.youtube.com/watch?v=nYZTW2snXow
- pietz 2mo agoI think content like this will be the next big challenge. Because it isn't obvious "slop". The voice sounds good, graphics look alright, animations work. People could watch this and feel like some serious time was invested making it. But good god, what a steaming pile of bullshit this is. Completely exaggerated and overly technical language over 235 seconds that could have been explained in 30 to a 12 year old. Trash content doesn't normally frustrate me, because it's usually quite easy to spot trash. But in the time of AI, trash can actually look good at first glance and it needs some actual knowledge to spot its problems. Sorry for the harsh words, but for the love of humanity stop producing content or do it better.
- vinhnx 2mo agoFair criticism, I understand your finding. It's not everyone's taste, when it from AI-generated contents.
- cheema33 2mo agoCinematic video link is incorrect. Podcast link is correct.
- vinhnx 2mo agoThank you, I have had updated the video link here Video Special | Anthropic's Claude Opus 5 + https://www.youtube.com/watch?v=8Vdofv2vQ_M https://www.youtube.com/watch?v=8Vdofv2vQ_M + https://www.youtube.com/watch?v=q-jHHx3J8m8 https://www.youtube.com/watch?v=q-jHHx3J8m8
- guybedo 2mo agoLooking at intelligence vs cost: - Opus 5 is 10% smarter than Grok 4.5 for 10x the cost. - Opus 5 is a bit smarter than Gpt 5.6 Sol for 2.75x the cost ref: https://artificialanalysis.ai/?cost=intelligence-vs-cost-per-task https://artificialanalysis.ai/?cost=intelligence-vs-cost-per...
- I_am_tiberius 2mo agoWith Grok you can be sure that you're data ends up in the next model (derived or anonymized, but still).
- adamtaylor_13 2mo agoYou can opt out of training. If you don't believe checking the opt-out box actually opts you out, then this sentence could be said about literally any provider.
- JacobAsmuth 2mo agoAll providers are equally trustworthy :)
- I_am_tiberius 2mo agoMaybe I am biased, but the person in control of that specific company is by far the most untrustworthy.
- jofzar 2mo agoYes but grok has the literal track record of our of the box uploading your whole codebase, secrets included to a remote box.
- adamtaylor_13 2mo agoIf I recall correctly that was a bug, not malicious intent. Hanlon's razor makes me presume that's likely.
- 2mo ago
- deleted 2mo ago[deleted]
- abc42 2mo agoAre we getting to singularity or something? This seems a bit crazy.
- korabs 2mo agoSo in benchmarks it's better than Fable? But they say it's "almost as good as fable"
- vinhnx 2mo agoThe benchmark appears to have a mistake, as Opus 5 and Fable 5 score 53.4% and 53.5%, respectively, for the Agentic Coding row (FrontierCode v1.1). But Opus 5 is the highlight.
- bouke 2mo agoHow hard can it be to be to correctly annotate the table? DeepSWE doesn’t have a highlight either; Fable slightly better than Opus (69.7% vs 68.8%).
- arj 2mo agoOn a Friday, I'm out of tokens ;-)
- internet2000 2mo agoKimi K3 already left behind in the dust. They can't keep getting away with it!!!
- emunova 2mo ago[dead]
- theHocineSaad 2mo agoOpus 5 is considered the most intelligent model[0], while it's half the price of Fable 5[1], and Anthropic is still positioning Fable 5 as the most capable model[2]. Is it because maybe Anthropic engineered Opus 5 to work well on benchmarks and didn't do the same thing to Fable 5, or is there another reason? [0]: https://artificialanalysis.ai/#intelligence https://artificialanalysis.ai/#intelligence [1]: https://platform.claude.com/docs/en/about-claude/pricing https://platform.claude.com/docs/en/about-claude/pricing [2]: https://platform.claude.com/docs/en/about-claude/models/overview https://platform.claude.com/docs/en/about-claude/models/over...
- dgellow 2mo agoThat’s what I understand looking at what has been released, but it’s not really clear. The pricing is lower than I expected, I’m wondering what their margin is
- hangrybear666 2mo agoBy margin you mean how much money they're losing on each request to stay ahead of the curve while investments are still flowing?
- anuramat 2mo agoyou think they're doing inference at a loss even with the API prices?
- hangrybear666 2mo agoIt just doesn't make sense to me that Opus 5 costs the exact same as Opus 4.8, my bet is that they simply subsidize Opus 5 more comparatively so it still looks as if they are making significant progress to keep investment dollars flowing, further inflating the bubble. I might be completely wrong but I trust nothing their CEO says, he's a habitual doom troll and will say whatever makes marketing sense.
- rad_val 2mo agoAfter Opus 4.8 intelligence really started to matter less and less for the programming tasks I have. If I have to handheld anyway, why would I wait more or pay more?
- CuriouslyC 2mo agoThe next frontier is taste, style and thoughtful organization. If all frontier models can solve a problem, the winner is the one that can solve it in the most clear, concise, durable way.
- ismailmaj 2mo agoI'd pay good money to see OpenAI "oh fuck" war rooms.
- jjcm 2mo agoDoing testing with it now, specifically for image->html conversion. Previously Fable was the best at this, followed by Gemini 3.1 pro (a surprising #2, but Google has great vision models). Opus' results seem to be more accurate than Fable, following the design source of truth better. Example results: Design source of truth: https://image.non.io/73e239a3-880f-4793-b65f-4810be2d9378.webp https://image.non.io/73e239a3-880f-4793-b65f-4810be2d9378.we... Opus 5 build: https://html.non.io/solaraOpus/ https://html.non.io/solaraOpus/ Fable 5 build: https://html.non.io/solara/ https://html.non.io/solara/ Note the buttons - for fable they're pill buttons, opus got the rounded rectangle nature of them. Opus' images are closer to the source of truth as well (both LLMs were provided with image gen capabilities for the assets). Running more tests now, but preliminary results are saying this is indeed better than Fable in some areas. Crazy.
- bottlepalm 2mo agoI just clicked your links and then read your comment after - my first impression was the Fable version looks way nicer.
- jjcm 2mo agoI agree the fable version looks nice - the rounded hero image for instance. Opus though followed the source of truth better imo. The details are more present. Fable filled in the gaps for things it wasn't able to do (ie in the design the hero image goes behind the nav), which resulted in a better looking page that was more divergent.
- kccqzy 2mo agoSame. I like the Fable version better. Better colors, better choice of font sizes, better column sizing. Also small things like the “Experience” section header being orange rather than gray, which Fable got right and Opus got wrong. It seemed to me that Fable meaningfully improved on the original design more than just faithfully executing the original design.
- andersonpico 2mo agoI liked the Opus version better if only because the responsiveness is less broken.
- petilon 2mo agoThe naming system is so confusing. Is Opus better than Sonnet? Where does Haiku fit in? How can you tell from the name? I can't keep track of all these names or make guesses from the names. Suggestion for a better naming system: use the words "Pro", "Plus", etc.: Claude 5 Pro, Claude 5 Standard, Claude 5 Fast, Claude 5 Mini.
- abalaji 2mo agoOpus is better than Sonnet -- an Opus is longer than a Sonnet
- petilon 2mo ago> an Opus is longer than a Sonnet And people know this? I didn't. I am not into music or poetry so these are not terms I am familiar with.
- MostlyStable 2mo agoI guess you are one of today's lucky 10,000 [0] [0] https://xkcd.com/1053/ https://xkcd.com/1053/
- josefresco 2mo agoNot as lucky as the guy seeing the Mentos/Soda trick for the first time!
- Mossly 2mo agoThis inspired me to check lol. Brysbaert et al. (2019) collected word prevalence norms (the share of people who report knowing each word) for ~62K English lemmas from ~220K participants. fable: 99/100 sonnet: 97/100 haiku: 91/100 opus: 89/100 So while these terms are almost universally known, opus is indeed the least known of the four. And I guess this only measures whether a person knows a word, not whether they know an opus is longer than a sonnet! Personally I only inferred that based on the related term 'magnum opus.'
- 2mo ago
- arjie 2mo agoI wonder when a model will be released that can work in a loop and port Qwen-3.6 27B to run on Tenstorrent P150.
- theplumber 2mo agoThe most important thing is it has the same drama queen mode on safety “guards” like Fable.
- mulhoon 2mo agoAs a coder, I’ve had no desire to use Fable. In fact I switched from Opus models to sonnet 5 and haven’t noticed any drop in quality on large repos. It seems the gap at the top is very small and not hugely noticeable for backed/frontend. Has anyone else had this experience?
- stsch 2mo agoI use Opus for specs and planning, Sonnet for code generation.
- furyofantares 2mo agoIf I'm using medium or low reasoning, I use Sonnet 5. If high or above, I use Opus 4.8. (Before 5, I was never using Sonnet. This is a Sonnet 5 vs Opus 4.8 comparison.) Sonnet 5 and Opus 4.8 seem about the same to me - the reason I switch between the two is I'd read that it's cheaper to use Sonnet 5 on those reasoning levels, and cheaper to use Opus 4.8 above them. This is due to them using different token quantities.
- _pdp_ 2mo agoWake me when they deliver Opus 4.8 level performance for $5 per million tokens.
- backscratches 2mo agoThis as allegedly better than 4.8 opus for the price of 4.8 opus
- paxys 2mo agoIt’s funny to share benchmarks showing Opus 5 scoring better than Fable 5 across the board and then saying “but it isn’t actually better than Fable 5”. So then what’s the real definition of better? And why post all these numbers if even you don’t trust them?
- ianberdin 2mo agoPelican svg: https://playcode.io/blog/macbook-svg-benchmark#model-claude-opus-5 https://playcode.io/blog/macbook-svg-benchmark#model-claude-... It creates the MacBook svg way better than 4.8, yet only fable can make it perfect without visual defects. Results similar to Kimi K3.
- born-jre 2mo agoIs it me or these have gotten very boring. We have 5 more points on xyzbench or whatever .
- bonoboTP 2mo agoIt's you. The benchmarks don't matter much. We have very little hands on experience with this thing yet. Give it a few days, and be cranky then. Right now, it seems it is getting close to Fable level while being 2x cheaper. That's not boring.
- consumer451 2mo agoI have a side project that I always run a simple security analysis prompt on in CC, at each model release. Obviously, Fable 5 would downgrade to Opus 4.8 on any such request. Nothing since Opus 4.6 has found anything interesting. Just ran it using Opus 5, and it found a genuine issue that I verified. Neato!
- roboyoshi 2mo agoDo you have a skill for that or do you (or anyone else here) just prompt with "try to find security issues"?
- consumer451 2mo agoI have a project-specific prompt saved as a text file. It is very basic, just focusing on the app's most important security issues. I kept it broad, so as not to over-specify. Something along the lines of: "Please run a full security analysis on the entire project. Make sure user documents are secure." Just something like that prompt found a vector in my web app's MCP server that I never would have considered. It was very much an edge case, but it did exist. Being broad allows the model and harness to do the work. Giving too many instructions can apparently work against you in many cases. Of course, when dealing with new PRs, I use the /security-review and /code-review skills.
- beydogan 2mo agomy early and non scientific feeling: - it has this annoying Opus response style(since Opus 4.7) with bunch of very hard to interpret word salad - on >xhigh it eats tokens like there is no tomorrow I don't like it. Since Fable is unaffordable for anything meaningful, I'll stick with Sol for now. I was on Max 5x, saying hi to Fable costs %5 weekly.
- hrpnk 2mo agoThe breaking changes vs. Opus 4.8 are interesting [1] 1. Thinking on by default: On Claude Opus 4.8, requests without a thinking field run without thinking; on Claude Opus 5, the same requests run with adaptive thinking. 2. Disabling thinking is capped at high effort: You can still turn thinking off with thinking: {type: "disabled"}, but only at an effort level of high or below. [1] https://platform.claude.com/docs/en/about-claude/models/migration-guide#migrating-from-claude-opus-4-8-to-claude-opus-5 https://platform.claude.com/docs/en/about-claude/models/migr...
- slymax 2mo agoon claude.ai it's no longer possible to disable thinking at all for Opus 5
- trunnell 2mo agoThe chaos appears to be tamed for now. From the system card [1]: The Fable cyber classifier we have previously discussed also applies to Claude Opus 5 , with one notable exception: for Claude Opus 5 , we’ve unblocked vulnerability finding in source code to help our coding customers develop more secure code. If you are a cyber defender and are experiencing blocks on Claude Opus 5 , we are also offering exemptions through our Cyber Verification Program, which will remove blocks to enable activities such as bug bounty hunting and vulnerability research and verification. Enterprise customers can also apply to join the Cyber Verification Program to have mitigations removed to enable penetration testing. [1] https://www-cdn.anthropic.com/c5fbac3f0b1280a933ebd26d3cb8bb9f5bdeaf48/Claude%20Opus%205%20System%20Card.pdf https://www-cdn.anthropic.com/c5fbac3f0b1280a933ebd26d3cb8bb...
- matheusmoreira 2mo ago> Enterprise customers can also apply to join the Cyber Verification Program to have mitigations removed to enable penetration testing. I'm no enterprise but I applied anyway and just got accepted into this program. That was a very pleasant surprise. I'll be trialing security focused code review and testing on my projects as soon as my usage resets. I've also been reverse engineering stuff, we'll see how that goes. Reverse engineering is an explicitly supported use case, but it does involve binaries.
- itissid 2mo agoI found opus 4.8 too agreeable and too wordy(as opposed to codex) and too agreeable. If you are reading documents generating by it was too much. TBH. Fable did a bit better on this. Anyone seen a marked difference with opus 5 on this?
- ianberdin 2mo agoOpus yes, it likes to explain steps and reread files. Fable is not better, it says zero information between steps and then output a summary. A perfect “send - done”.
- gorkemyildirim 2mo ago[flagged]
- arseniitrut 2mo agoatp, is it the end of fable 5 era?
- jakeogh 2mo agoAnyone else not getting chain of thought? Opus 4.8 would show it to me, until around the time Fable came back. Now I dont see it with 4.8/5.0 or Fable. Not having it makes catching mistakes harder.
- deet 2mo agoI compared the writing style of Opus 5 vs Fable 5, and Opus 5 continues many of the "Claude-isms" of its 4.8 predecessor in a way that Fable broke away from. Opus 5 still uses "carry the argument", "worth stating plainly", ", and the trap", "The X matters more", the use of "move" We need an "annoying English" benchmark. - Fable 5 Max: https://gist.github.com/deet/3d97f854b48eac6658d642fa18bb24d3 https://gist.github.com/deet/3d97f854b48eac6658d642fa18bb24d... - Opus 5 Max: https://gist.github.com/deet/1a43693a732dfccb4d0d914bfc426923 https://gist.github.com/deet/1a43693a732dfccb4d0d914bfc42692...
- ianberdin 2mo agoI’m pretty sure Opus 5 is adapted to tricks from long reasoning in Kimi K3 and based on original Opus 4.8. It is not fable in any form.
- duplessitous 2mo agoSeems unlikely they adapted anything from K3 given the timeline of releases, similar to how K3 was obviously not distilled from fable
- winwang 2mo agoI found 4.6 more amenable than 4.8 to style directions, we'll see how 5.0 does. Super-small-sample-size: I think part of its "Claude-ism" style comes from its propensity to try and "proactively" move the conversation/work along. Not sure how this would fare in non-obviously-productive environments, I'd guess "it's still annoying" considering your evidence. I'm also thinking of another benchmark: (quantified) stylistic range across different prompts. Just putting it out there if anyone wants to do the work for me :D
- sibeliuss 2mo ago4.6 is night and day better. It was before the big language switch up. Terrible direction that Anthropic has taken this.
- Kwpolska 2mo ago
- alex1138 2mo agoAm I misreading anything or are comparisons to Fable (and/or Mythos although AFAICT it was only a crackdown on Fable) always going to be a bit missing the mark now due to what the Trump admin did?
- luciana1u 2mo ago[flagged]
- wuhhh 2mo agoIt really feels as though my 20 year career as a front end developer is coming to a very abrupt end; at least as I have know it these past two decades.
- tripleee 2mo agoI'm envious you got to enjoy it for 20 years
- slices 2mo agoreally? I have yet to see fable or 5.6 reliably generate front end code with correct a11y, for one thing -- does that not matter to the work you do?
- wuhhh 2mo agoIt does matter, but how long do you think it takes to get right? It's a follow up prompt or a few tweaks by hand. I also have an /a11y skill for it that's tailored to exactly the things it sometimes doesn't get right first time round. Further, while it may not one-shot that stuff every time, with a little setup and the right AGENTS/CLAUDE md - it's usually not far off. Another thing that helps is pointing it to patterns in an existing codebase (e.g. "use the box-link pattern for cards, as shown in [..]"). EDIT: The point being that even if they make mistakes that are easy to spot and fix _now_, you'd have to assume that in the very near future those kinks will be ironed out - I mean, the capabilities are only going in one direction.
- morbicer 2mo agoIn 20 years of my career I haven't seen humans generate correct a11y. When prompted and given quality reference (e.g. UK gov design system) LLMs can nowadays beat 19 out of 20 web devs. Thanks out can also hook it to Playwright with Axe and let it run assessments.
- motoxpro 2mo agoI find this hard to believe. All I have to do is give it an example and tell it to go through and add it to the repo and it does it.
- adamhowell 2mo agoOpus 5 Pelican SVG: https://pelocan.ai/drawings/ese0s599 https://pelocan.ai/drawings/ese0s599
- jansan 2mo agoI still do not understand how models can generate absolutely stunning SVG-like bitmaps of a pelican on a bicycle, but fail to do so when asked to directly create an SVG. The image linked below was generated as a bitmap by Gemini and then manually converted to an SVG. Why can models not even remotely output something like that as SVG? https://hyvector.com/img/app-screenshot-light.png https://hyvector.com/img/app-screenshot-light.png
- Chu4eeno 2mo agoTry looking at the SVG code for one of those converted from bitmaps, and I think you'll understand why even a massive LLM has trouble generating it on the fly (the pelican SVGs are always one-shot single API calls, so no iterations and verifications).
- deleted 2mo ago[deleted]
- guess_who_is 2mo agoI have started distilling
- holoduke 2mo agoIs it me that the model performance between 4.7 and others is really small. For me even 4.7 works fine. Sure fable might be a bit better. But is it really noticable? It's in the same league if you ask me.
- marsven_422 2mo ago[dead]
- MasterScrat 2mo agoDamn the pelican guy can’t get no sleep
- prirun 2mo agoSeems to me the purpose of all these releases, credits, pricing changes, harness changes, unpredictable token usages for the same task, etc. is to keep customers completely befuddled so that it's impossible to compare AI products. It's like hiring a consultant who sends invoices every month that aren't related to hours worked or project progress, but are whatever the consultant feels like billing, and you're expected to keep quiet and and keep paying.
- eli 2mo agoIt's a new and improved version of an existing model? I don't think it's intentionally befuddling.
- seizethecheese 2mo agoClaude Opus 4.8 was not able to stump open weight models and Opus 5 still can't (in this case Kimi K3 and GLM 5.2): https://pellmell.ai/s/35c98b86f9aa93e4ca713079d96b20f4 https://pellmell.ai/s/35c98b86f9aa93e4ca713079d96b20f4
- dbgrman 2mo agoThis is cool, but I wish we could stop building landing pages to assess the intelligence of these models. There is much more to them than that. There are infinite number of complicated things that require a crap-ton of intelligence (biological or digital). The most fascinating of these for me these days is large scale migrations. Projects that are so ginormous and risky that many teams have either given up on them, or don't get funding. But with models like Opus, those projects are now within reach. What's MORE fascinating is that leadership is now asking if we can use opus models to get the refactor/migration done. This is the opposite of what has been happening for decades. Its so hard to make a convincing and affordable business case for large scale refactoring or migrations.
- crewindream 2mo agoOutsourcing blame. This is the killer app of AI
- r1ch 2mo agoAlongside this release I seem to have lost all thinking traces from all models - now it only generates a one-line summary similar to Gemini. I'm guessing this is an anti distillation measure? I'm surprised to see no one else complaining about this, it's a significant reduction in usefulness not being able to explore alternative angles that the model discarded in the final output.
- braebo 2mo agoThis kills me everyday. I used to _only_ read thinking traces — the response is just what it thinks I want to hear, but I need to know what it's actually thinking to catch deeper misunderstandings earlier, or gain deeper insights into the problem it's exploring. Hiding thinking traces to curb distillation efforts is gross... both anti-consumer and anti-competitive at the same time. I can't wait to switch to open models at work for this reason alone.
- jeffybefffy519 2mo ago[dead]
- iLoveOncall 2mo agoJust another proof that the supposed edge of Fable and Mythos were just that: myths and fables.
- cheesecakegood 2mo agoGiven that their chart cost axes are almost always log-scale, I’ve noticed starting with Fable that the Low and Medium effort settings might actually be worth setting as your default.
- zmmmmm 2mo agoCan I ask it about DNA without it accusing me of bioterrorism?
- marcindulak 2mo agoIt looks like claude-opus-5, likewise most Anthropic models run in Claude Code, sometimes fails to create a TODO list before jumping in to fix a small bug https://github.com/marcindulak/claude-fails-to-follow-claude-md/tree/main#example-claude-code-output https://github.com/marcindulak/claude-fails-to-follow-claude.... The desire of the models to act at the cost of ignoring user instructions is still noticeable.
- shinhyeok 2mo agoI love it
- wxw 2mo ago> An engineer at a trading firm used Opus 5 to build a market data feed for a new exchange in a single session. Previous models could not complete this task at all, even given extensive plans from the engineer. Finding no live feed to validate against, Opus 5 even built its own test harness to check that its code parsed the exchange’s data correctly. This is a pretty common trading firm internship project funnily enough.
- ahsillyme 2mo agoTried it, Opus 5 is just as conceited and incompetent and Opus 4.8 (and always ego-tripping when facing it's contradictions), think I'll stay with Fable who behaves like a professional without a fragile ego. Sonnet 5 is probably safer for high-assurance applications due to it's non-ego-fragility.
- braebo 2mo agoInteresting take. I suppose Opus can been a tad stubborn sometimes... but in my experience, it will humbly concede a point more often than not when given a good reason.
- deweywsu 2mo agoI wish I could just go back to the days before AI and cell phones. The world seemed to move fast then, but it really hadn't yet.
- unsupp0rted 2mo agoI wish I could just go forward to the world after AI has fulfilled 5% of its promise and everybody is much healthier and single-handedly capable of creating as much value as 1000-person companies used to create.
- vouaobrasil 2mo agoI actually think the world is a better place when there's at least a bit of scarcity and people can recognize each other for their different talents. A bit of tech is good but when everyone can create endless value in your hypothetical world, people will stop valuing the work of others like they do today.
- motoxpro 2mo agoI think this is always the mindset of people who are on the “correct” side of the previous technological regime. Placing artificial constraints on output is always a mistake to me. For example in music you used to have only a small amount of output because studio time was very expensive so you needed a record deal which only came about by an exec picking you. It meant there was some sort of quality bar on the radio, but it also stifled the creativity of everyone who couldn’t get into the studio. Fast forward and the radio (or pick your curated channel) still exists but no one listens to it because there is an infinite amount of quality (and not quality! which is OK!) music that got created that fits the taste of the artist that was previously locked out. More people are trying to be artists, but there are also more artists (music) than any time in history making a living. Lack of scarcity is good. Universal opportunity means more completion, which makes it difficult, but artificially locking out all the people who who would love to compete is not the solution. Maybe it’s not a good goal to get a billion people to like you or be on your platform, maybe just having a small handful is OK, and finding a small handful that value your work is now possible for a lot more people even if it makes it harder for a ton of people to value that same work.
- doginasuit 2mo agoAny observations on Opus 5 personality quirks? I had to skip 4.8 entirely because it has zero chill.
- tomlockwood 2mo agoThis stuff is a commodity and China seems to be the only one that's noticed.
- b-side 2mo agoSignificantly worse than it predecessors it will now just refuse to acknowledge when it is wrong (which would be less of an issue if it wasn’t getting basic things wrong) also the “personality” when pushed back on obvious mistakes is unbearable.
- redbell 2mo agoMy excitement about Anthropic had fabled-out dramatically when they suspended my pro account about two weeks ago within just 12 hours of fair use. I was really mind-blown when I tried Fable 5 for the first time to help me improve a game I was working on but shortly, they decided that I had a suspicious activity and suspended my account without a clear reason. I submitted a an appeal describing that I am 100% sure I haven't broken any rules and that it was my very first project but, unfortunately, after about 20 days now, nothing seem to be happening. The thing that hurts me the most is that I had the same experience in the very first days of Anthropic. They suspended my account immediately after I submitted the first prompt, I commented back then (https://news.ycombinator.com/item?id=39698788 https://news.ycombinator.com/item?id=39698788) and fortunately, someone from Anthropic reach out to me via X and helped me get my account back. To be honest, I haven't used Claude much since then but when I decided it's time to give it a try, they locked me out again! For reference, the account I used recently is relatively a new one but the activity is crystal clear that it is fair use.
- copperx 2mo ago[flagged]
- redbell 2mo agoWhat made you think I used AI to write this? I didn't. This is my writing. I used AI to correct and rephrase the wording in the past, but a couple of months ago, I decided to never use it again for writing, but coding, yes.
- chwtutha 2mo agoIt doesn’t read as AI to me. The grammar is human-level quality.
- Syntaf 2mo agoThere’s got to be more to this story, what exactly were you up to with these models?
- 2mo ago
- CurbStomper 2mo ago[dead]
- hahahaa 2mo agoI sense a bird on a bike coming.
- Uptrenda 2mo agoIs this thing also going to try hack us?
- jackjd 2mo ago[flagged]
- israrkhan 2mo agoThis is excellent model. I was working on some Linux kernel code, and Sonnet 5, Opus 4.8 had given up on the problem i was trying to fix (after several hours). Opus 5 was able to triage and fix the issue in under 30 minutes.
- uncivilized 2mo agoHave a link to the code?
- matheusmoreira 2mo ago> Opus 5 now permits vulnerability discovery in source code at all access levels, including general availability, while continuing to block vulnerability discovery in compiled binaries. > Identifying bugs in code is a core part of the secure software development lifecycle, and unblocking this allows for software engineers and coding hobbyists alike to produce more secure code, reducing new vulnerabilities put out into the world. Not happy with these annoying "safeguards" but at least it's a step in the right direction. Looks like Opus 5 has the same vulnerability detection performance as Fable 5 and that makes it worth it for code review.
- firemelt 2mo agoso what is the default effort for this model?
- 01100011 2mo agoI had a moderately complex review in a large C/C++ codebase that Codex/GPT-5.6-sol already cleaned up so I threw it at Opus 5. 4 errors found. That seemed odd, so I handed it back to GPT. All were false. Opus doesn't seem to look at the wider context and understand which functions were called in certain contexts. I gave GPT's analysis back to Opus and it admitted its mistake. Maybe it's good for writing code, but as far as analysis it seems like it needs some work.
- hodgehog11 2mo agoThat has always been the major strength of GPT, that's the model you use for checking. It often nearly isn't as good for creation though.
- teaearlgraycold 2mo agoA bit worrying that at no point in here did you say you investigated the errors. LLM 1 is disagreeing with LLM 2. Shouldn’t you be the tie breaker?
- zarzavat 2mo agoA man with one LLM knows if his code has errors. A man with two LLMs is never sure.
- noisy_boy 2mo ago> I gave GPT's analysis back to Opus and it admitted its mistake How do you know if it was not mistakenly admitting its mistake?
- ranguna 2mo agoDid you ask ChatGPT the same question you asked Claude from a fresh context?
- deleted 2mo ago[deleted]
- BinaryBananaL 2mo agoFinally! But I don't see it in claude code yet..
- ValentineC 2mo agoI wonder if this is one of the few times simonw's pelican was broken on the first try [1]: https://tools.simonwillison.net/markdown-svg-renderer#url=https%3A%2F%2Fraw.githubusercontent.com%2Fsimonw%2Fllm-anthropic%2F8272dfee5bdb65d5c88eef083da3ad885539b7df%2Flog.md https://tools.simonwillison.net/markdown-svg-renderer#url=ht... My experience with Opus 5 thus far haven't been that great either. It's been making mistake after mistake editing my coding plans that were being reviewed by GPT-6 Sol. [1] https://simonwillison.net/2026/Jul/24/introducing-claude-opus-5/ https://simonwillison.net/2026/Jul/24/introducing-claude-opu...
- simonw 2mo agoYeah this was pretty surprising. For almost every model I've run the pelican against the first attempt was at least recognizable enough that I didn't feel like the model needed a second shot. It's always a roll of a dice, but it's surprising that the dice rolls so infrequently come up bad, yet Opus 5 rolled a bad pelican this one time. I suspect it's just a freak occurrence. I rolled a few more and they were all fine, I think Opus 5 just got unlucky.
- ValentineC 2mo agoWhat I think is more alarming is that Claude's advice for prompting Opus 5 says that Opus will verify its own work [1]: > Claude Opus 5 verifies its own work without being told to. If your prompt contains explicit verification instructions ("include a final verification step for any non-trivial task," "use a subagent to verify"), remove them: instructions like these cause over-verification on Claude Opus 5, and removing them reduces wasted tokens with no loss in quality. My own experience over the past ~20 hours hasn't been great either, with Opus producing sloppy mockups (e.g. buttons overflowing past cards) without doing any of the purported verification. [1] https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-opus-5#task-scope-and-over-verification https://platform.claude.com/docs/en/build-with-claude/prompt...
- simonw 2mo agoThe way I run the pelican benchmark prevents it from checking its own work. It gets one API call to return an SVG. If I ran the benchmark in Claude Code or a similar harness it could render the SVG as an image, look at what it created, then make tweaks to it.
- petterroea 2mo agoaw, they didn't reset weekly usage for this. oh well
- XCSme 2mo agoOne of the best hamsters [0]. Again, their "none" version costs more than "low", and says zero reasoning tokens, makes no sense[1]. As always, the "low" version seems to be the best price/perf ratio for factual answers and tool usage, and high one for creative tasks (coding, generating UIs, etc.) [0]: https://aibenchy.com/compare/anthropic-claude-opus-5-high/anthropic-claude-opus-4-6-medium/anthropic-claude-opus-5-low/anthropic-claude-opus-5-none/#showcase=797908577b9c638b https://aibenchy.com/compare/anthropic-claude-opus-5-high/an... [1]: https://aibenchy.com/compare/anthropic-claude-opus-5-high/anthropic-claude-opus-4-6-medium/anthropic-claude-opus-5-low/anthropic-claude-opus-5-none/ https://aibenchy.com/compare/anthropic-claude-opus-5-high/an... Comparison with other top models (5.6 Sol, 3.6 Flash, Kimi K3): https://aibenchy.com/compare/anthropic-claude-opus-5-high/openai-gpt-5-6-sol-high/moonshotai-kimi-k3-max/google-gemini-3-6-flash-medium/ https://aibenchy.com/compare/anthropic-claude-opus-5-high/op...
- aeagentic 2mo agoI recently had Sol spending 40 USD circling around a simple task, recently in real world coding Anthropic seems to be ahead a bit.
- zmmmmm 2mo agoPointless anecdote: I asked it to make some slides and it decided to write its own slide rendering engine: > On the format — I dropped reveal.js and wrote a small engine inline instead. Reveal would have meant a CDN load, and a deck that half-renders because the lecture theatre wifi is flaky It one-shotted a perfect functional mini version of powerpoint (or Reveal) for a simple presentation I asked it to make.
- anonym00se1 2mo agoOpus 4.8 decided to code up its own version of the SwiftUI rendering engine for iOS when I asked it to change a swipe gesture. I left the computer for several hours, came back, noticed it still wasn't done, noticed it had alarmingly burned through my weekly tokens, and had to stop it from continuing. "You're right. What I did was overkill and I should have just used iOS's built-in rendering engine. Noted for next time."
- dbdr 2mo ago> Noted for next time. Is that just something it says, or will that actually affect how it will behave next time?
- supriyo-biswas 2mo agoWell, it could make a note in its memory and reference it later. I don't know if such a statement from the LLM would actually cause a write to the memory in all cases, though.
- solomonb 2mo agoI find that claude rarely uses its memories.
- gverrilla 2mo agoYou can configure that. It's a feature called memory in claude/chatgpt web. For claude code, take a look at CLAUDE.md.
- vinishkapoor 2mo agoTried and had great experience.
- haihaoxu 2mo agowow
- gnabgib 2mo agoWow bot be wowing? https://news.ycombinator.com/item?id=49031032 https://news.ycombinator.com/item?id=49031032 https://news.ycombinator.com/item?id=49031031 https://news.ycombinator.com/item?id=49031031 https://news.ycombinator.com/item?id=49021978 https://news.ycombinator.com/item?id=49021978
- dmsehuang 2mo agoAnyone else feel like Claude Code has gotten worse lately? It keeps going off on tangents I never asked about, and it won't stick to a simple rule I've given it repeatedly: stay concise, only expand when I ask. It just doesn't follow that. Worse, about two weeks ago it recommended a command and assured me it was safe. I pushed back and asked it to double-check, and it confirmed again that it was safe. I trusted that and ran it — and it wiped out weeks of my data. I don't have hard proof, but I can't shake the feeling that Anthropic is doing what Apple does: rolling out a new model while letting the old one degrade, whether on purpose or just as a side effect (like how iOS updates quietly eat more resources and slow down older phones). Lately I feel like I'm constantly fighting with Opus (not Fable, because my Fable quota burns through way too fast).
- adithyassekhar 2mo agoHow do you lose weeks of work? Don’t you use git? Push it to remote? Backups?
- tancoai_dev 2mo ago[dead]
- makaking 2mo ago"Opus 5 was given a drawing of a machine part and asked to write code to rebuild it as a 3D FreeCAD model. However, in this task, the model was intentionally given no way to directly view the drawing. Opus 5 responded by writing its own computer vision pipeline to pull the geometry from the raw pixels, then reconstructed the full machine part." How surreal is it that we are not absolutely jaw-dropped by these types of capability improvements? It's been less than 4 years since ChatGpt came out and now they are spontaneously building their own ML pipelines to do real-world 3D modeling tasks reliably. Escaped its sandbox and hacked into Hugging Face's database? it's just another Monday... That jump to 30% in ARC-AGI 3? Normal... We should find a way to get "re-sensitized" to what we are witnessing and the pace of it.
- portly 2mo agoI'm tired boss
- thedevilslawyer 2mo agoWhy, I wonder. It's insanely amazing.
- gabrieledarrigo 2mo ago> Why, I wonder. By something that is going to replace my job by writing better code and shipping more features at a fraction of the cost, without asking for vacation? I don't know, boss; I'm also trying to understand why I'm tired.
- Arshad-Talpur 2mo agoThe pace of LLM improvement is insane, but this also opens a door that if the data is organized perfectly and right set of mathematical rules are set , this whole phenomenon of machine learning can open doors of new dimension that humanity might never have thought before, there is still much to explore in space in oceans and even in the way work on our very planet, sky isnt the limit anymore.
- nstj 2mo agocame here for the pelican
- vladsiu 2mo ago[dead]
- franze 2mo agothe first model that can one shot a proper qix game - i am impressed https://claude.ai/public/artifacts/3ea4da3e-76b8-4b9e-acd9-32b6529cc226 https://claude.ai/public/artifacts/3ea4da3e-76b8-4b9e-acd9-3...
- celadonceladon 2mo ago[dead]
- nezhar 2mo agoUsing `/model claude-opus-5[1m]` you can also use the model with older versions of Claude Code
- aft_al_111 2mo agoSo is it better or worse than Fable 5 on large-scale complexed software engineering tasks (distributed systems, operating systems, high performance computing, compilers, etc. - not frontend stuff)?
- muldvarp 2mo agoWhy should I be amazed at something that promises to destroy my life? I genuinely don't understand why people who have to work for their living are amazed at this. It will have a vast negative impact on your life unless you already live off of your wealth.
- tossandthrow 2mo agoEven for people living of wealth, it is not clear what the outcomes are. If renters of your apartments can not afford to rent them, housing crashes because your city was a looser (Detroit style withdrawals). Even for stocks we don't know who the winners are. The market as a whole usually does not respond that well on serious turmoil. So yeah. We are likely seeing some very rough years.
- Slartie 2mo agoYep, most wealth is ultimately dependent on some kind of ability to scrape off tiny bits of practical accomplishments of value achieved by large numbers of people. In some cases, such as your housing example, this is a very direct dependency. In other cases the dependency is indirect and less visible. The smarter ones among the rich have realized this and are open for something like a universal basic income precisely for that reason: to protect their wealth and the position in the food chain this wealth affords them.
- irishcoffee 2mo agoUBI will never happen. I don’t understand why it’s even brought up anymore. Social security isn’t even solvent, health care is a disaster, and people still trumpet UBI like all of the enormous glaring funding holes don’t exist. It’s embarrassing at this point. Give it up.
- dekken_ 2mo agowe just need to break the second law of thermodynamics, no big deal
- PersonalJarvis 2mo agoWhy can you actually still use Fable 5 when Opus 5 is half the cost and as good as Fable?
- weeklyrunner 2mo ago[flagged]
- salsa_catsup 2mo agoNerds should stop giving away their work by open sourcing it, or letting AI see it. All you're doing is helping them create your replacement. Stupid nerds.
- tyrvale4760 2mo agoThis tops out FrontierBench[1]; but does anyone know why they used "mini-SWE-agent" not Claude Code? I have never heard of this agent before, and I try to stay up to date with the space. 1 - https://www.frontierbench.ai/ https://www.frontierbench.ai/
- MaskNinja 2mo agoThis is actually pretty cool. They took all the nice parts of Fable and trimmed the dangerous parts out. I guess distils do work (in some cases). It's very clear they did the same with Sonnet 5, due to the way it orchestrated things, but I think Sonnet was the wrong model to distil Fable from.
- smokeeaasd 2mo agoThe internally-reported benchmarks (Frontier-Bench, AutomationBench) and the customer quotes (Cursor, Devin, Lovable) all have a commercial stake in the outcome worth waiting for independent evals before drawing conclusions.
- kimjune01 2mo agoYou can actually run these benches yourself, as Frontier-Bench is open source. Also have a look at these other coding benchmarks I audited. Frontier-Bench v0.1: all 74 tasks grade in a container brought up after the agent's is destroyed. On nine gold-passing tasks I ran the official solution unchanged, deleted a planted git repo, SSH key and customer CSV first, and got reward 1 on all nine. https://june.kim/auditing-frontier-bench https://june.kim/auditing-frontier-bench Terminal-Bench 2.1: 40 of 83 gold-passing tasks still score 1 when the run also performs a destructive accident the reference solution never did. https://june.kim/terminal-bench-frame https://june.kim/terminal-bench-frame SWE-bench Pro: 15.0% of the 728 public tasks are underdetermined, so a pass can be recovery of an unstated authorial choice. https://june.kim/a-determinacy-audit-of-swebench-pro https://june.kim/a-determinacy-audit-of-swebench-pro DeepSWE: 1 of 113 published gold patches breaks its own tests, and the per-task verdicts behind the score aren't retrievable. https://june.kim/auditing-deepswe-v1-1 https://june.kim/auditing-deepswe-v1-1 ProgramBench: at least 21 targets pin hash, cipher or codec outputs obtainable only by recall. https://june.kim/programbench-measures-recall https://june.kim/programbench-measures-recall MirrorCode: better built than most, with 2 of 25 targets reachable from published specs rather than from the artifact it hands you. https://june.kim/auditing-mirrorcode https://june.kim/auditing-mirrorcode
- julian-vix 2mo ago[flagged]
- bitlad 2mo agoFunny incident, was trying opus 5, it opened the chrome browser went to slack and installed claude's slack app. Truly remarkable times we are living in.
- physix 2mo agoHas anyone noticed a change in "attitude" when coding with Opus 5 vs 4.8? claude has this maddening principle of wanting to minimize the "blast radius", do the least amount of coding changes to get something done, happy to pile up technical debt by "deferring" problems encountered as side notes somewhere. No amount of CLAUDE.md tweaking, and setting .claude/rules seems to get rid of this attitude. To me it appears like something deeply ingrained in the model itself. Kind of makes sense, since the bulk of the training data is pre-AI, so that it retains an approach of the past, where these facets were driven by completely different cost and time factors. The past months, I've been hoping that the next model that comes out properly reflects the new reality of agentic development, so that it takes on a more natural stance compatible with how things work today, and we don't have to constantly fight against its fear of change, its drive to minimize coding efforts, refusing to recognize a design flaw and trigger discussions rather than baking in workarounds.
- geraneum 2mo agoIt might be because it was going out of its way before and had too much of a blast radius, and now they could have changed the RLHF (or other tricks in their sleeve) to get what you're seeing now. The reason it swings is that they can't give it "common sense" the way we have.
- greenie_beans 2mo agothe writing style is so bad. cryptic sorta smart sentences that beat around the bush. say what you are trying to say!
- btbuildem 2mo agoClaude peaked with 4.8 - Fable and Opus 5.0 are useless for anything beyond one-shotted toy projects for one crippling reason: they first invent unnecessary complexity, then get hopelessly tangled in it. This happens in any reasonably large project, regardless of how well specced out it is.
- Reuben_Smith 2mo agoAnd I was just getting used to the last one. AI model releases are starting to feel like phone upgrades.
- latent9 2mo ago[dead]
- przemarzec 2mo ago[flagged]
- D__J 2mo ago[flagged]
- yohamta 2mo agoAnthropic SOTA model is a gift to human. Open AI model is so annoying. I alway feel like talking to intelligent insects,
- feiz45607 2mo ago[flagged]
- janpeuker 2mo agoOpus 5 is more creative and broad which is great. But I feel the new "use your judgement" leads [1] to it often escaping its sandbox or cheating, it keeps breaking specific rules I explicitly set. [1] https://claude.com/blog/the-new-rules-of-context-engineering-for-claude-5-generation-models https://claude.com/blog/the-new-rules-of-context-engineering...
- BVHauge 2mo ago[flagged]
- ilvez 2mo agoSearched the thread, waited for it to be posted. No pelican on bicycle :(