96 ms·
Claude Sonnet 5
- primaprashant 3mo agoBased on both performance vs price charts, it seems using Opus 4.8 with med effort is almost a better choice than using Sonnet 5 at xhigh effort
- baalimago 3mo agoNot looking great for an upcoming IPO
- mrcwinn 3mo agoYou’re right, it’s looking stellar. Well beyond great. Real, and unprecedented, revenue growth will do that for a company.
- CuriouslyC 3mo ago"Real and unprecedented revenue growth" Bro that is financial engineering, not real revenue growth. They engineered the switch to usage based pricing and a price hike timed the quarter before they wanted to go public, long enough to juice their numbers but not long enough for them not to be able to manage backlash and have to walk things back. Then they tried to extrapolate that manufactured bump to make it look like they have record shattering revenue growth.
- tensegrist 3mo agothere was a vibecoded prediction market–style page that was put up yesterday (?) that got the date exactly right i think
- mag7269 3mo agoWhen can we get a new Haiku? 4.5 came out nearly a year ago, and it's showing its age.
- scosman 3mo agoLook at Qwen for that level of intelligence.
- anthonypasq 3mo agoneeds to be on bedrock for me to use it at work
- 0xbadcafebee 3mo agoGemma 4, Kimi K2.5, MiniMax M2.5, gpt-oss, GLM 5, Qwen3 Coder Next, DeepSeek V3.2, Devstral 2, are all available on AWS Bedrock and all are about Haiku level
- deleted 3mo ago[deleted]
- chipgap98 3mo agoInteresting that tasks on extra high cost almost the same as Opus 4.8 with a slightly worse performance
- bredren 3mo agoThis is on the browsercomp graph, right? In that, it seems sonnet 5 on high costs more than opus 4.8 at a lower pass rate. Am I reading this correctly? Edit: It looks like the key value proposition of the updated model is that it is much better than Sonnet 4.6. Wheras, Sonnet 5 delivers great value (by browsercomp benchmarks and compared to opus) when running in low and medium. So: Sonnet 4.6 should ~never have been run for low, medium or high when Opus 4.8 has been available. Whoops, I think I have some skills that delegate easy stuff to Sonnet. --- I remember Anthropic pivoting everyone's default model to Opus but had not seen it put so starkly before. I am a bit confused on the subscription `/usage` screen. It splits out sonnet usage, and I'd presumed that would have contributed to a lower use of subscription Quota. But if this is correct, Sonnet usage was basically like smoking unfiltered cigarettes.
- mchusma 3mo agoI agree with this assessment, IMO my takeaway from this is "Generally run Sonnet on low, otherwise use Opus". It's kind of like an "extra low" setting of Opus. (depends on the application for sure).
- bredren 3mo agoIt would be good if Anthropic provided some kind of feedback or even toggle to auto-route requests for models being used at thinking levels that would be a better value using a different model. Sort of like, getting an automatic upgrade at a car rental or hotel if there is availability.
- siva7 3mo agoThey already do. Don't assume the routing will be in your favour
- babelfish 3mo agoSystem Card: https://www-cdn.anthropic.com/d9bb04416ffe1352af84721476c1fa9994c07fde/Claude%20Sonnet%205%20System%20Card.pdf https://www-cdn.anthropic.com/d9bb04416ffe1352af84721476c1fa...
- tokengod 3mo agoThat’s nice, but we want Fable
- giancarlostoro 3mo agoThe reality is that Fable will eventually be obsolete and Sonnet / Opus will surpass it. Fable did cost 2x as much as Opus, so I assume it involves a much higher cost for what it did, but I wouldn't be surprised if Fable will be obsoleted by Opus or even Sonnet sooner or later at less cost.
- ianhawes 3mo agoOkay I don’t care about “eventually”, I want Fable now.
- arcatech 3mo agoHave you considered getting better at coding so you can build stuff yourself instead of waiting for models you might not be able to get access to anymore?
- cesarvarela 3mo agoThis is like telling someone who wants a motorcycle that they should get better at running instead.
- arcatech 3mo agoWhen the motorcycle manufacturers keep making each new model worse and more expensive and the government keeps trying to ban them.
- giancarlostoro 3mo agoI'd love to meet the devs who can spin up full feature web apps in under 15 minutes with all the bells and whistles I've gotten Claude to spin up and code. I don't think the AI haters understand the level of time cutting that you can achieve with a very simple and reasonably crafted prompt. I'm talking back-end, with database models, classes, queries, accompanying front-end layouts, with real dynamic data, running. Stuff that takes days to weeks to spin up, with minimal errors or issues, having cut down on days or weeks of effort, you can focus on testing and making it all into better code.
- andai 3mo agoOpus 4.8 beats Sonnet 5 on the pareto frontier in several of their graphs (Agentic Search, Agentic Computer Use). In other words, for certain tasks, Opus 4.8 is cheaper than Sonnet 5, and does better than Sonnet 5. I've noticed this pattern on a lot of benchmarks. You can try to emulate a bigger model by ramping up the test time compute (max reasoning, more turns, model fusion etc.), but you can't reach the same quality level, and you often exceed the cost you would have paid by just using a bigger model. tldr: if you're doing something hard, just use a bigger model.
- copperx 3mo agoAnd Claude Code penalizes you for using Sonnet on the subscription plan, so there's little reason to use it.
- bredren 3mo agoThis is what I realized, can you provide more detail on how you've observed this? The /usage screen does not make it clear.
- MillionOClock 3mo agoNot the original commenter, but personally I noticed my quota usage didn’t feel like it was being spent at a much lower rate when using Sonnet even on a relatively low thinking budget and based on a few comments here it seems I might not be the only one. Has anyone else noticed this? Wasn’t it different in the past? I thought I would be getting to use Sonnet much much more than Opus but it did not feel that way despite being on 20x plan.
- grim_io 3mo agoThis is exactly what people have been talking about in this thread. Sonnet is dumber and more expensive than Opus. The token efficiency improvements in Opus are missing in Sonnet. Sonnet generates more output tokens and more reasoning tokens. Any price advantage per token disappears due to volume. It doesn't make sense to use Sonnet if you have access to Opus.
- alvis 3mo agoWhat I starting to hate is that each model's effort level can mean completely different power. Today sonnet 5's med level effort is equivalent to sonnet 4.6 low level effort :/
- deleted 3mo ago[deleted]
- nsingh2 3mo agoThat seems to only be true for the "Agentic Search" benchmark. That benchmark in particular is a bit weird, because Sonnet 4.6 effort levels had a relatively small effect, so Sonnet 5 med is basically comparable to all effort levels of Sonnet 4.6.
- phillipcarter 3mo agoSeems to be another great incremental update to the workhorse, nice! I've been using Sonnet instead of Opus for almost all coding tasks for a while now. A little elbow grease to break down tasks and you can spend a lot less money for just about the same output quality.
- thewebguyd 3mo agoYeah I think people are sleeping on the smaller/faster models like Sonnet. As long as you have a detailed plan or small, well scoped individual tasks Sonnet can implement just fine. Opus will still do better at more open ended tasks or completely "vibe coding." Or spec/plan with Opus, and have Sonnet implement.
- conradkay 3mo agoI was surprised to learn that Sonnet generally has the same tokens per second as Opus
- SeanAnderson 3mo agoCrazy. I just changed the default for our entire org to Opus because people were continually unimpressed with Sonnet's abilities. It's fascinating to think how varied people's experiences are when interacting with LLMs and how much the outcomes depend on how people approach interacting with the models.
- 3mo ago
- wolttam 3mo agoI didn't think they'd actually release a model that was worse than the open-weight frontier and at a higher price-point. Wow.
- 2748484848 3mo ago[flagged]
- tripleee 3mo ago"very aesthetically pleasing beak. good form. looks to be riding fast. please visit my website"
- CharlesW 3mo agohttps://news.ycombinator.com/newsguidelines.html#comments https://news.ycombinator.com/newsguidelines.html#comments
- 2748484848 3mo agoPlease don't use HN primarily for promotion
- LUmBULtERA 3mo agoThat's yet to be determined. I think a lot of open-weight models are benchmaxxed and their usefulness for many tasks are not represented by those.
- enraged_camel 3mo agoYes, this has been my experience. They all struggle with long-horizon tasks and eventually start going in circles.
- s3p 3mo agoWhy did the other reply to this get flagged as dead? It was a comment about how someone would come out saying that Sonnet 5 would be better on the pelican test and therefore it has to be good. But I guess HN loves pelican SVGs so much that you're not allowed to criticize it.
- mesmertech 3mo agoOk thats a one month clock to the next Opus model at least, so thats a silver lining to a meh model.
- conradkay 3mo agoWow, seems worse even on price/performance than GLM 5.2, which is only 744b parameters. From the system card: "On CyberGym vulnerability discovery, Claude Sonnet 5 is less capable than Sonnet 4.6, and far less capable than Opus 4.8 and Mythos 5 As with the other evaluations in this section, these results were achieved with all safeguards turned off. When run with our default mitigations, Sonnet 5 scored a 0 on CyberGym"
- Retr0id 3mo agoFinally, a viable business strategy - sell security-oblivious code monkeys for cheap, then charge premium rates for agents capable of cleaning up the mess.
- JacobAsmuth 3mo agoI think instead they should sell super hackers and get their product banned instantly and go bankrupt
- usef- 3mo agoJudging by the events of recent weeks, I'm guessing the low cyber results are why they were allowed to release it
- loufe 3mo agoNot to single you out, parent commenter, but I really hope the quality of discourse on HN will move past these basic comparisons eventually. It seems like every thread on every model release has the exact same comments. "Wow, X models is Y% better or worse than Claude Z model on T benchmark" "That's irrelevant, they're just benchmaxing." "Not useable for daily coding or agentic workloads, the vibes are totally wrong." "It's almost as good, and costs a lot less, so I will absolutely use it." "I cannot imagine justifying using these, as the step change means open models lower costs do not make up for the productivity loss" I'm an unhappy Anthropic customer and really rooting for open models and non-gatekept intelligence, but how do we move on from this now meme-like model release discourse rigamarole. I do not know what that would be. I don't design LLMs nor benchmarks, and I genuinely appreciate that people do their best to provide information, even if non-perfect here. I'm sure most of you who actively read these comment pages on announcements must feel similarly, though, right?
- mchusma 3mo agoThis is much more interesting of a model at $2/$10 (their launch pricing) than at full price. There are many competing models at around this level of performance. I also like that the difference between low, medium, high, xhigh seems more spread, which is actually a good thing for people trying to tune applications. Running Sonnet 5 on low with the launch pricing makes this potentially a better fit than Haiku or open source models for some tasks. I don't think it will make sense at full price.
- mchusma 3mo agoReally if they wanted a standout model that would really take the wind out of GLM's sails, they should have made this the new Haiku, priced at Haiku levels with this performance.
- beernet 3mo agoAnthropic's run on the model and product side of things is highly impressive. They got Sam A. punching the air consistently, which is well-deserved and self-inflicted above all.
- CuriouslyC 3mo agoWdym? They've been knocking it out of the park on marketing, but Claude Code is still a meme, and Opus is getting trashed by GPT5.5 meanwhile you can't even use their "dominant" model, and anecdotal reports from when people could use Fable, when they weren't getting silently poisoned, was that it was only marginally better than GPT 5.5 in terms of SWE smarts, mostly being better in terms of pleasantness to interact with and design taste.
- beernet 3mo ago> Claude Code is still a meme Claude Code generates more revenue than OpenAI...It appears to be a nice meme.
- CuriouslyC 3mo agoLike I said, Anthropic's marketing is killing it, they've got people freely(?) shilling for them on public forums so even if they have shit developer relations and community relations and a model that's mostly worse while being more expensive, they can ride a wave of misinformation.
- beernet 3mo ago> they have shit developer relations Not true > model that's mostly worse while being more expensive Not true > they can ride a wave of misinformation. Not true
- CuriouslyC 3mo agoLook at the way that Anthropic has legally threatened people who do stuff they don't like around Claude Code and their subs, and compare that to how OpenAI has acted. Look at how mixed up and unstable their communication is on policies is relative to OpenAI. Don't take my word for it, Theo/Primeagen have a whole back catalog outlining how shitty Anthropic is. Look at the cost per intelligence of Opus vs GPT 5.5. Anthropic is the Taylor Swift of frontier labs... Not bad, but massively, MASSIVELY stan'd for inexplicable reasons, in violation with reality.
- deleted 3mo ago[deleted]
- Scroll_Swe 3mo agoI don't pay so I'm glad for the upgrade. I usually use Gemini, Mistral Le Chat (Vibe...) or Deepseek as they have way more generous free limits and I can basically spam forever.
- satvikpendem 3mo ago> Evaluations also show that it has a much lower ability to perform cybersecurity tasks than our current Opus models. Why would they brag about something like this? It's like they know people want to use models to perform cybersecurity tasks yet knowingly deny them the ability. And Opus 4.8 is still cheaper for a higher pass rate (much less open weight models like GLM 5.2) so not sure why I'd use Sonnet except on the low effort level for I suppose trivial tasks where I want it to work only 50% of the time judging by the graph. The pricing doesn't really make any sense.
- zlurker 3mo agoThey spent months hyping up Mythos and ended up with it banned. I’d assume they want to both differentiate their products and appeal to regulators here
- worldsavior 3mo agoThey will release it eventually. Once they see the Chinese models are close to Mythos level they will release it before, so it will be "revolutionary".
- jaapz 3mo agoIt was already released. US government is the only reason it's not available to us mere mortals anymore
- satvikpendem 3mo agoDue to Dario hyping it up as a world ending model. If they kept their mouths shut we'd all have it now still.
- baq 3mo agoWhere is gpt 5.6?
- justicehunter 3mo ago[dead]
- theLiminator 3mo agoSeems like the way to go for any smaller models is to only use the low reasoning levels, and for anything where you'd want it to reason harder, to just use a larger model. In effect, high reasoning only makes sense when you're using the frontier model and need extra performance (higher levels of reasoning are never pareto optimal unless you're at the largest model size).
- mwigdahl 3mo agoNot to sound like an LLM, but that seems exactly right to me. Use it as a cheaper, high-functioning task subagent and lower reasoning for a master Opus session. As long as not every portion of your task requires maximum intelligence, you should come out ahead.
- user43928 3mo agoWon't any input be charged uncached, and the output of the small model charged again as uncached input to the bigger model? I don't know whether that comes out ahead compared to just staying with the better model in the first place.
- mwigdahl 3mo agoIt's a good question, but for multiturn conversations even cached context adds up quickly. My experience has been that spawning off subagents for defined tasks in a large overall plan generally makes me come out ahead. I'm sure folks' mileage will vary though.
- noisy_boy 3mo agoI asked this question and was told that even if it is counter intuitive, medium will be more cost efficient due to caching. Changed to medium, blew my budget and went back to low.
- docheinestages 3mo agoMy experience with using low reasoning effort has been nothing but a waste of time. Claude often keeps guessing, not calling tools to ground itself, and basically at the end I end up wasting the same amount of tokens or just switch to Opus on xhigh. It's been a terrible experience.
- doctoboggan 3mo agoThe cost per task chart is telling me that I should _never_ use Sonnet 5 above medium effort level - Opus always performs better for a given cost. So I guess the takeaway is that if Sonnet 5 medium isn't good enough for you, switch models, not effort levels.
- Torkel 3mo agoYeah, I was looking at the same chart and was very surprised at where the curve is relative to opus... Feels like sonnet 5 is "what if opus had an extra-low effort level"?
- johnfn 3mo agoThat's just one benchmark, though. Tab to the next one and Sonnet 5 performs better as effort goes up just as you'd expect. I imagine the suggestion is that performance vs effort tradeoff is task dependent.
- energy123 3mo agoNo it doesn't? It's worse than Opus across the whole shared frontier on both plots.
- acchow 3mo agoAgreed. The graphs clearly show that opus 4.8 performs strictly better at the same cost per task
- jsnell 3mo agoBut they don't show "strictly better" performance at cost per task! The graphs show parts of the cost/performance pareto frontier occupied by Opus 4.8 and others occupied by Sonnet 5.0. If Opus 4.8 was strictly better at cost per task like you say, by definition the entire frontier would be occupied by Opus. So neither is pareto-dominant over the other. In contrast, Sonnet 5.0 is Pareto-dominent over Sonnet 4.6 on those graphs.
- garo-pro 3mo agoSeems like the cyber detection even is on Sonnet now. https://support.claude.com/en/articles/14604842-real-time-cyber-safeguards-on-claude-opus-and-sonnet https://support.claude.com/en/articles/14604842-real-time-cy...
- microtonal 3mo agoClaude Sonnet 5 is built to be the most agentic Sonnet model yet. It can make plans, use tools like browsers and terminals, and run autonomously at a level that, just a few months ago, required larger and more expensive models. I have been using Sonnet 4.6 more than Opus, because I'm mostly doing agent-assisted development and not fully agent-driven development. This announcement does not make me positive, I have found that the more models are optimized for fully agentic development, the worse they get at assisted development and often start doing too much despite very strict/specific instructions. I have been moving more and more to K2.7 Code and GLM-5.2 the last few weeks. They are often good enough for assistance, very fast, and cheap.
- mohamedkoubaa 3mo agoI've been moving more to Composer 2.5 for the same reason. KISS principle.
- AdminAdmim 3mo agoSame for me, downgraded Cursor Subscription because when i use Cursor i use 90% Composer 2.5 fast
- everfrustrated 3mo agoComposer 2.5 fast (via Grok) is honestly amazing. Its been implementing everything I've asked and getting it right first time. Been impressed with it's front end ability. If this was the last model I could ever use I think I would be happy.
- jklmnopqrstuvw 3mo agoFrom my own experience, GLM-5.2 generally cost more tokens and much more slow.
- microtonal 3mo agoWhich inference provider do you use? (Admittedly, I currently use K2.7 a lot more currently.)
- SoKamil 3mo agoI believe that’s gonna be meta for agentic coding this year for enterprises. Cost optimized models approaching SOTA capabilities on software engineering but without cybersec training.
- johnfahey 3mo agoJudging from those cost-performance graphs, Sonnet doesn't make sense to run at anything higher than a medium reasoning level, since Opus 4.8 low reasoning outclasses it for the price. This line as a selling point is also pretty funny: > Evaluations also show that it has a much lower ability to perform cybersecurity tasks than our current Opus models.
- lucynight 3mo agoAMAZING
- alvis 3mo agoIronically, the key message of today's release is that Sonnet 5 is far less capable than Opus 4.8 and Mythos 5. It's a funny development is the past few weeks
- solenoid0937 3mo agoDuh? It's their cheapest model aside from Haiku.
- moomin 3mo agoI feel like this is a bit of a disappointment. Sonnet 4 was a clear step above Opus 3.x, while this is a lot muddier.
- Sol- 3mo agoWonder if the whole cyber paranoia leads to their models ultimately generating less secure code. After all, if it has the ability to generate safe code, it would imply that it knows something about cybersecurity, which could surely be used to hack all the banks in the world.
- pennomi 3mo agoTrying to censor nudity in image generation models caused all kinds of problems with anatomy in image models. I’m sure these models will have similar issues with security.
- NonHyloMorph 3mo agoInteresting, you find that in medieval painting, due to the authority of the catholic church.
- raincole 3mo agoCensorship on image generation models works on another level. The models can generate NSFW, but there are extra computer vision models checking if the images can be shown to the users. It's especially obvious for Grok and ChatGPT.
- BoorishBears 3mo agoThere are image models with censorship at every stage from pretraining to posttraining. Most recently Ideogram released an open weight model that will denoise into a grey image with the text "Blocked by safety filter" notice for certain prompts Of course, because it's open weights people have found defeats
- nodja 3mo agoThat's only correct for specific models and not what parent was referring to. Stable Diffusion 3, an open weights model, was laughed at at release for not being able to even generate a woman laying in grass. The community attributed this to the heavy dataset filtering. Since then other open weights releases have been made with no NSFW capabilities and the community claims they're not as good as anatomy as well. You can google "stable diffusion 3 woman in grass" and press the images tab to see how the model failed spectacularly.
- jchw 3mo agoAmerican AI company status: We are now bragging about how bad our models are unironically. Okay.
- smallerfish 3mo agoAh that's why Opus has been so slow for the last couple of days.
- benjiro29 3mo agoAnybody notice that they did not include Sonnet 5 Max in the "Agentic Search results", when comparing to Opus 4.8 ... Based upon the "Agentic Computer usage", Sonnet 5 Max was going to be off "Agentic Search results" chart. lol ... In short, Sonnet 5 Low/Medium is more cost efficient, if its a task below Opus 4.8 Medium. For the rest its expensive and your better off using Opus 4.8. Why even release this model?
- bredren 3mo agoI'd narrow that to why even allow the harness to run `high` on this model?
- ricardobeat 3mo agoBecause it’s a massive improvement over the previous model, and cheaper? You are reading too much into the graph and ignoring the threshold of usefulness for real world tasks. By that logic Sonnet 4.5 would have never been worth using.
- benjiro29 3mo agoAm i missing something? Because your making my point. Its only worth it compared to Opus 4.8, if the tasks your running requires Opus 4.8 low (or non-existing lower). For the rest the gap in pricing vs efficiency is so small, that there is no point in using Sonnet. I am looking at their own cost comparisons vs efficiency...
- ricardobeat 3mo agoThe point is that Sonnet at medium or even low will be smart enough for most daily tasks. You’re defining “worth using” as if you always need the highest performance possible, which is what these benchmarks measure, but most work doesn’t need it. You’ll pay more to get the same result. Sonnet 4.5 is very popular as a main model currently, this is a free upgrade. I use Haiku a lot for agent workflows, if I can get better output at similar prices, Sonnet 5 will replace it completely.
- docheinestages 3mo agoBut does it burn tokens just like Opus? That's the feeling I have nowadays. Regardless of what model I choose, the 5-hour limit gets exhausted in the first hour or so.
- a_c 3mo ago"Claude Sonnet 5 is available everywhere today at an introductory price of $2 per million input tokens and $10 per million output tokens through August 31, 2026. It then moves to standard pricing at $3 per million input tokens and $15 per million output tokens.2" "Sonnet 5 is an upgrade to Sonnet 4.6, but it uses an updated tokenizer that changes how the model processes text to improve performance (this is similar to the tokenizer change we introduced with Claude Opus 4.7). The tradeoff is that the same input can map to more tokens: roughly 1.0–1.35× depending on the content type. The introductory pricing is set so that the transition to Sonnet 5 is roughly cost-neutral." If we trust them, then it is roughly the same as sonnet 4.6
- ricardobeat 3mo ago[dead]
- docheinestages 3mo agoIs it just me or is there a huge difference between how much one can accomplish in a 5-hour window with GPT 5.5 on xhigh versus any Claude model?
- mrcwinn 3mo agoI exclusively use 5.5-xhigh-fast within Codex and find it superior to Opus 4.8.
- deleted 3mo ago[deleted]
- kingjimmy 3mo agointeresting footnotes: "Sonnet 5 is an upgrade to Sonnet 4.6, but it uses an updated tokenizer... can map to more tokens: roughly 1.0–1.35× depending on the content type." AKA expect higher costs on Sonnet 5 vs Sonnet 4.6 for the same tasks.
- winstonp 3mo agosame happened to Opus 4.7
- tripleee 3mo agointeresting how much worse the sentiment around Anthropic is getting
- mwigdahl 3mo agoSeems like a combination of multiple factors: "They took my shit away!" -- 3-day Fable 5 addicts (me) "How dare they tell Trump no?" -- US nationalist / "my country right or wrong" types "Great to see a closed source company fail!" -- open source boosters "Great to see an American company fail!" -- anti-US, and/or pro-China folks "Great to see a successful company fail!" -- anti-capitalists and/or sour-grapes crab bucket types "Serves you right for ripping off creators!" -- copyright warriors "They keep silently nerfing the models!" -- secret downgrade conspiracy theorists "Quit killing the planet!" -- anti-datacenter advocates
- tripleee 3mo agoIt seems to be more them losing goodwill combined with their marketing. I don't agree with your framing that all negativity is from crazies
- mwigdahl 3mo agoI don't think all the negativity is from crazies, but big chunks of it are certainly motivated. I certainly left out numerous other categories.
- feralcoder 3mo agoThe amount of anti-Anthropic and anti-Dario posts i've seen on reddit threads has gotten a bit ridiculous. It feels like your analysis is mostly spot on, it's the confluence of several motivated parties pouring effort into social media. Many of the posters are pro-foreign models/pro-open source, and most can't distinguish the difference between "open source" and open weight models like Qwen, Minimax, or GLM. Reminds me of the old "free as in beer" vs "free as in speech" debate. Free beer means you don't pay, but you don't get to see the recipe or change it. Free speech means you get the actual source and the right to study it, modify it, and redistribute it. Open weight models are basically the beer version. You can download the weights, run them locally, fine-tune them, quantize them, host them on your own boxes — but what you have is a finished product, not the blueprint for how it was built.
- micromacrofoot 3mo agoSo they repackaged Fable and added "don't scare the government" to the prompt
- actionfromafar 3mo agoThis is downvoted, but how can it not be a little true?
- mellosty 3mo agoIt does not pass the "I want to wash my car, should I drive or walk"
- cheesecompiler 3mo agodid for me even on low non thinking effort
- gverrilla 3mo agoGIGO, as they say.
- jerrygoyal 3mo agoIt's actually a huge update for building products, given most tasks are sub-agent driven where Sonnet is used, steered by Opus.
- stackedinserter 3mo ago"Our new model is proudly dumber now!"
- mwigdahl 3mo agoWhat? If you're comparing their models in the same size class, Sonnet 5 is Pareto-optimal over Sonnet 4.6.
- zamadatix 3mo agoI think they mean per dollar in the perf/$charts, not per marketing class. I.e. the new model is a complete Pareto failure in said perf/$ charts with the sole exception of Sonnet 5 low, which is dumb enough to not have comparison at all. Opus 4.8 delivers a better outcome per dollar, regardless what the underlying size of the models is. I'd generously assume this is something about the specific category of agentic task presented in the chart... but it does raise the question "then why is that category the one they chose to highlight here".
- mwigdahl 3mo agoFor agentic computer use Sonnet 5 low performs better than Sonnet 4.6 medium at just under half the cost, and better than Opus 4.8 low at 25% off. Their success rates are not that far off. Agentic search is a different story, but even there it still dominates 4.6 (as in, for everything Sonnet 4.6 can do, Sonnet 5 can do it as well or better at the same or lower cost). Yes, Opus 4.8 dominates Sonnet 5 over its entire range in both categories, but Opus's lower range is limited and there is a valid regime on the lower end where Sonnet 5 use makes economic sense. This is not the case for Sonnet 4.6 where Opus 4.8 dominates it completely on both charts. Edit -- reading your response closer I think we're saying the same things, maybe just disagreeing on whether that lower end is valuable or not.
- ekjhgkejhgk 3mo agoIn effective terms they're lowering prices.
- scottfits 3mo ago> the computer use evaluation OSWorld-Verified. Sonnet 5 (orange line) is a strict improvement over Sonnet 4.6 cool to see, still waiting for models to get better at computer use.
- gverrilla 3mo agoIs this the default model for non-paying users? If so, that could be an interesting move in the competition for this segment.
- johnhamlin 3mo agoKind of hilarious how much they’re touting that it sucks at cybersecurity like it’s a feature
- DonsDiscountGas 3mo agoI'd love if they would include speed (though I know there are difficulties involved). At this point the quality of Opus 4.8 is no longer my limiting factor, it's the speed, so a faster model would be great.
- boc 3mo agoHave you tried Opus on fast mode?
- DonsDiscountGas 3mo agoI haven't because I'm not made of money but maybe I will
- Getchowned 3mo agoFable soon please.
- aykutseker 3mo ago[dead]
- Cu3PO42 3mo agoSonnet 5 is not currently available in the EU region on Bedrock, whereas previous models were and still are. I wonder if this is only due to early stages of the rollout or if this is due to recent US restrictions. Unfortunately that means I won't be using it at work for now.
- mellosty 3mo agoSonnet seems to be really expensive
- mrcwinn 3mo agoHave you followed Anthropic at all?
- arendtio 3mo ago> Evaluations also show that it has a much lower ability to perform cybersecurity tasks than our current Opus models. It seems being incompetent is a feature now...
- m3h 3mo agoWhy is Claude Sonnet 5 allowed to be released but OpenAI Terra not? Are they not the same class of models?
- theplumber 3mo agoIs there any reason to use Sonnet instead of GLM?
- rw2 3mo agoThe use of the "cheaper models" in big AI companies are next to useless as they don't even score as well as the open/super cheap Chinese models. Only the frontier big models like Fable and Opus have value.
- docproof 3mo agoThe jump in reasoning quality is noticeable. What's interesting is how it handles ambiguous instructions now — it seems to ask fewer clarifying questions and just makes a reasonable judgment call. That's a double-edged sword depending on your use case.
- _pdp_ 3mo agoToo expensive?
- m3h 3mo agoImportant to note: "Sonnet 5 is an upgrade to Sonnet 4.6, but it uses an updated tokenizer that changes how the model processes text to improve performance (this is similar to the tokenizer change we introduced with Claude Opus 4.7). The tradeoff is that the same input can map to more tokens: roughly 1.0–1.35× depending on the content type. The introductory pricing is set so that the transition to Sonnet 5 is roughly cost-neutral."
- mattas 3mo ago"We can raise prices in two ways: (1) raise the price per token and (2) increase the number of tokens we generate on your behalf. We promise not to do (2) maliciously. Promise."
- squeegmeister 3mo agoWouldn't it be more malicious for them not to mention this at all?
- Alifatisk 3mo agoSure, but I think doing it this way allows them to later on say they were transparent about it. Completely hiding this would make it very difficult for them excuse when getting caught.
- conradkay 3mo agoI think the incentives are less bad since a good chunk of usage comes from subscription plans. There was a fairly major regression in Claude Code performance for some time when they changed the system prompt to try and make it less verbose (saving tokens). And if I'm not misremembering, there were a lot of complaints when they changed the default effort from high to medium.
- ComplexSystems 3mo agoSo the post-introductory price is set such that Sonnet 5 will cost 100%-135% as much?
- PeterStuer 3mo agoAnyone else feel like Opus 4.8 got significantly dumber over the last 2 weeks?
- Jcampuzano2 3mo agoI'm struggling to understand why I'd ever use this instead of just using a lower effort level for opus given on many of the benchmarks listed the cost per task rises above opus at anything higher than medium effort. Only thing I can think of is for when someone is out of opus credits. Of course there are API billing use cases but I'd probably still just use opus on low.
- SirMaster 3mo agoMaybe it's not for you? I don't pay, so I can't even use Opus... So this is an upgrade over Sonnet 4.6 for me.
- nicce 3mo agoOlder Opus models will likely get deprecated and then over time this is the cheapest model. That is how prices are currently increased.
- ChrisLTD 3mo agoYeah... Sonnet becomes the new cheap model, and some Fable class model becomes the more expensive/better one.
- theptip 3mo agoWat. Price/perf has been going down massively over the last few years.
- darkwater 3mo agoBecause they still haven't fully captured the market for Agentic Development.
- enraged_camel 3mo agoSpeed is a huge reason. Sometimes you just need some simple tasks get done fast, and waiting 30-60 seconds for opus to even start thinking can really slow things down.
- cenobyte 3mo agoClaude Sonnet 5 is built to be the most agentic Sonnet model yet. or The Dodge Charger is built to be the most Charger like car yet.
- prmph 3mo agoSo many things to think about regarding these "benchmarks": - Do the ever increasing scores on the mean we will soon have models that approach 100%? And what would that even mean? That there is no more room for improvement? - Would Anthropic (or any other model vendor for that matter) ever release a newer model that scores lower? If not, does that mean they keep tweaking a new model they want to release until it shows an improvement of the prior model? - Would it be more useful to move toward a comparative rather than absolute ranking?
- joaohaas 3mo agoImportant to note that the cost graphs are heavily distorted. The agentic serch one for example is divided into 3 'columns': $0-$2, $2-$5 and $5-$10. And yet, the $2-$5 section is the widest, even though it only contains a single point. I can't even say if this is making the product look better or not, but it sure is weird. Maybe Claude just hallucinated those splits xD
- andrewchambers 3mo agoThe whole fable fiasco really soured me on Anthropic. This just looks disappointing by comparison.
- varispeed 3mo agoWhat is the point if it is one Trump's brain fart away from being blocked?
- swe_dima 3mo agoNot sure what niche it's going to occupy: too expensive for it's intelligence category.
- SkitterKherpi 3mo ago$5/$25 for Opus 4.8 vs $3/$15 doesnt seem cheaper enough to be too worth it. It depends how much better it is than e.g. Mimo, but I imagine Mimo and co to be too cost efficient in the lower tier to be overtaken by Sonnet for most tasks.
- make3 3mo agoit's also a lot faster I would assume
- OsrsNeedsf2P 3mo agoGreat timing. I just started using Claude Sonnet as a long term reverse engineering project[0] for a game I used to play as a kid. The cheaper tokens but sufficiently smart with hard verification makes it a perfect combo for the task [0] https://github.com/dginovker/BFME-Source-Code/ https://github.com/dginovker/BFME-Source-Code/
- brunooliv 3mo agoI only wish Opus 4.6 from earlier this year at a faster inference speed. Since Opus 4.6 things have been so much messier and the overall push for more agency isn’t really panning out for agent assisted development as much as they would like
- fractorial 3mo agoI still use Opus 4.6 (with later models for subagents only sometimes), but I have been preparing for it to go away.
- Danii27 3mo ago[flagged]
- 827a 3mo agoTbh we'll see what using it looks like, but the reasoning/cost charts do not look promising. It seems like the only useful reasoning level for Sonnet 5 is Low; medium might trade blows at price/performance with Opus, but anything beyond that Opus is Just Better. I struggle to understand where this model fits in. If I need a cheap model for simple stuff (like, summarizing an email); I'd go Haiku (actually, I'd go Deepseek v4 Flash, but you catch my drift). I just can't think of many tasks where I'm like "yeah let me reach for Sonnet Low Reasoning so I can save a dollar but also seriously run the risk of it failing"; I'd just reach for Opus Low.
- brokencode 3mo agoKind of crazy how bad this release actually is. I even dug around in the full system card, and every graph showed the same thing. Low and maybe medium will save money on simpler tasks, but after that it just isn’t worth it compared to Opus. I wish they would have explained in the blog post why they think anybody would ever want to use this above medium. Maybe it works well on things that aren’t clear in the benchmarks.
- siva7 3mo agoWhy would a company explain how limited their own major release is?
- brokencode 3mo agoThe graphs do that already. I was expecting them to try to explain how good it was at simple tasks.
- ai_fry_ur_brain 3mo agoFinally a model release where everyone is realising the scam. The world is healing (maybe).
- artursapek 3mo agoI run a proofreading benchmark that tests how well models can find and fix errors in English text. They get several passes in a simple agent loop. Sonnet 5 is definitely better than Sonnet 4.6, but inferior on both quality and cost to GLM 5.1, GLM 5.2, Gemini 3.1 Flash, and Gemini 3.1 Pro. https://revise.io/errata-bench https://revise.io/errata-bench
- whh 3mo agoIt's not Fable, but I'll take it.
- guelo 3mo agoHave they ever said what the difference is between Sonnet and Opus? Are they trained differently? Different architectures? Is Sonnet a distillation? Is it just that Sonnet has less resources for inference? None of the other labs are doing this kind of long lived two model series.
- jsnell 3mo agoGemini has had Pro and Flash since May 2024, across three major version nunmbers. The Opus and Sonnet naming is only two months older than that.
- ThouYS 3mo agoWhy did this get the coveted "5"? I want an Opus that can compete with GPT 5.5
- oybng 3mo agoIn my case, 4.6 degraded massively over time. 5 fails the same basic tasks that I gave 4.6 yesterday. And quite frankly this low, med, high, extra, max, turbo, ultra, ludicrous nonsense is getting tiresome
- edude03 3mo agoLet’s see how long until opus 5 comes out but to me this lends some credence to the rumour that fable/mythos was supposed to be opus 5
- phtrivier 3mo agoWhat is the reference, unbiased, honest, reputable and trustworthy site that ranks and compare models on the couple of realistic metrics that matters ? ("Does it work for code", "no, I mean, for real", "how much does it cost", etc...) ?
- bel8 3mo agoThe only metric that worked for me is running the same prompt 5x for each LLMs on my projects. I keep specific branches a state where they are ready to develop new features.
- girvo 3mo agoTruthfully? There isn't one. They all have flaws. Your best bet is to look at all of them, and then run a suite of evals yourself. Its rough out here!
- kccqzy 3mo agoIt’s not really possible unless you try. Different people use models so differently. The whole model situation has made public minute differences in personal preferences in the process of coding. Some people think carefully and strive to write code that’s as bug free as humanly possible on the first try; others write something that is only approximately correct and then iterate afterwards. The former people would align with a model that thinks for 40 minutes before producing flawless code; the latter would be driven mad by this excessive thinking. Some people like to interrupt AI as soon as they see AI making a mistake, others let AI continue and tell them about the mistake afterwards.
- kvetching 3mo agoGLM 5.2 is better and cheaper. Maybe they are trying to embarrass Trump by making it look like we are losing to China.
- kvetching 3mo agoAnd it worked. BREAKING: The export controls on Claude Fable 5 are expected to be lifted tonight, per Politico!
- jongjinchoi 3mo agoI think so. GLM 5.2 is more reasonable.
- munaf-khatri 3mo ago[dead]
- XCSme 3mo agoI just tested it on my benchmarks[0], it's GLM-5.2 level, at 2x cost, but also 2x faster. Weak spots (categories it fails): - Trivia — 0/3 - basically not much built-in knowledge - Combined tool-calling tasks — score 45/100, sometimes makes invalid tool calls - Puzzle Solving — score 77, flubs carwash-like tests [0]: https://aibenchy.com/compare/anthropic-claude-sonnet-4-6-medium/anthropic-claude-sonnet-5-medium/anthropic-claude-opus-4-8-medium/z-ai-glm-5-2-medium/ https://aibenchy.com/compare/anthropic-claude-sonnet-4-6-med...
- XCSme 3mo agoAs always, note: faster than GLM-5.2 doesn't mean too much, as GLM-5.2 is served by different providers, so the inference speed can vary drastically between providers or over time.
- yieldcrv 3mo agoWhat’s everyone favorite GLM provider? z.ai doesnt always have the most reliable AI but I don’t mind the party seeing my trade secrets and thoughts compared to an American corporation + the party seeing my trade secrets and thoughts. So thats not a functional difference to me, and the Chinese one won’t reply to subpoenas so thats a value add tbh So I’ll consider all, fastest tokens/sec wins
- eli 3mo agoFireworks.ai is solid. And if you care more about speed than cost they have a "fast" variant that I think just throws more hardware at the model for about 2x the cost.
- david-gpu 3mo agoThe privacy policy indicates that they track you and share your data to ad networks like Meta. Yikes.
- 3mo ago
- deleted 3mo ago[deleted]
- simonw 3mo agoClaude Sonnet 5 itself described its pelican as looking like a goose: > Illustration of a white goose riding a bicycle, with one wing extended forward to grip the handlebar, set against a plain white background with a brown ground line. https://simonwillison.net/2026/Jun/30/claude-sonnet-5/ https://simonwillison.net/2026/Jun/30/claude-sonnet-5/
- bel8 3mo agoThat's possibly the worst pelican I saw from all recent LLMs. Meanwhile GLM 5.2 drew a cool self-contained fully animated SVG pelican: https://simonwillison.net/2026/Jun/17/glm-52 https://simonwillison.net/2026/Jun/17/glm-52
- simonw 3mo agoYeah, GLM have been beating Anthropic on the pelicans for a while now. (I suspect that's more of an indication that Anthropic have chosen not to waste resources training on animals riding vehicles, personally.)
- bel8 3mo agoThat's one possibility albeit quite charitable. I'd be inclined to think the same personally if GLM 5.2 wasn't also rocking in other areas too.
- kamranjon 3mo agoThis is interesting, I haven’t actually heard you suggest that the labs are focusing on this benchmark before. Have you come around to this position as a result of the quality of pelicans you’ve been getting? The reason I thought this was an interesting benchmark is because it’s a non-image generating model creating an image using SVG code, so it kinda spans capabilities. If an AI lab trained a model specifically for animals riding bicycles it seem trivial to modify the prompt and determine if it was trained specifically for that or if it’s generalized a skill and can also generate a proper orangutan walking on stilts or an armadillo on a skateboard, this sort of thing?
- botfriendsarent 3mo agoSonnet 5 OUCH! every model is just loaded with more hurt, stolen content, BS prompts, more scare tactics, more illusions, more government lobbying, less honesty. Oh Claude you master of software engineering does it ever end? DO you have no bounds? How may we further assist you oh Claude?
- m3kw9 3mo agoshould have called it 4.9, it don't deserve the 5 monkeier
- caste 3mo agoidk, i think they just tried to compensate for the ban of fable, nothing too good
- taytus 3mo agoRoughly on par with GLM 5.2 at 5x the price
- taspeotis 3mo ago> Claude Opus 4.7 and later Opus models, Claude Fable 5, Claude Mythos 5, Claude Mythos Preview, and Claude Sonnet 5 use a newer tokenizer that contributes to their improved performance on a wide range of tasks. This tokenizer produces approximately 30% more tokens for the same text. Claude Sonnet 4.6 and earlier models use the previous tokenizer.
- MagicMoonlight 3mo ago[dead]
- matheusmoreira 3mo agoWho cares about Sonnet? I want to know about Fable. Are the export restrictions really going to be permanent?
- stingraycharles 3mo agoIt’s supposed to happen when Anthropic introduces identification, which I believe is planned for mid-July.
- matheusmoreira 3mo agoNot a US citizen. Identity verification is not going to help me.
- ianberdin 3mo agoAnthropic outsmarted everyone again. They released Sonnet 5 with a temporary price reduction until August. Everyone was excited, but in reality, they increased the tokenizer size by 50%. As a result, the actual cost went up by 50%, they shifted everyone's attention to decrease. Thus, Anthropic is raising prices but not telling anyone about it. Nobody is really aware of it. You go to the pricing page, the price looks the same. Yet people are actually paying 50% more. Very shady marketing. And of course they lie about 35% again. In reality with coding it is 50%. UPD: I run playcode.io, so it’s my job test all models, their pricing, quality in order to provide best price/quality/speedy/reliability to non-techy.
- boutell 3mo agoUntil now we've been using Sonnet 4 to power an editing agent in ApostropheCMS. Sonnet is a good price/quality/speed compromise, but sometimes when giving it a large set of instructions it would miss half of them. At least until we told it to go back and try again. In my early tests tonight, Sonnet 5 is a LOT better out of the box. It's one-shotting complex instructions. It also recovered independently from bad instructions that led to an uninformative 400 error by using its schema-fetching tool to figure out there were was too much input. If I have to gripe about something: it interpreted another impossible instruction by quietly discarding the input in question. But, the way it did it is... kinda exactly what anybody else would do, if they weren't in a position to change the implementation. This is, obviously, early days but I'm impressed.
- ashvardanian 3mo agoGot really excited for this model and asked my Opus planners in 3 pretty different projects to use Sonnets instead of Opus subagents to help me experiment on HPC kernels faster. Not one of them ended up writing a single line of code... Sonnets just kept spinning, wasting tokens. Can't remember the last time it happened with Opus in my codebases. Reverting back.
- bearjaws 3mo agoI've seen this happen before when they launch new models. When Opus 4.7 came out it was "working" for 20+ min before I just exited entirely and waited till next day. Went away on it's own.
- addozhang 3mo agoIn the 4.x era, I prefer Sonnet to Opus. The quality of Sonnet generation is good enough for me, but it's much faster than Opus.
- crorella 3mo agoFun/interesting to see how opensource models surpassed Anthropic's
- jbritton 3mo agoI accidentally used Sonnet 5 a bit today. It seemed significantly worse to me than Opus 4.8 for software development.
- midtake 3mo ago5 as in 5 times more likely to tell you that you can't edit your driver INF files because that enables DRM circumvention and is dangerous!
- stavarotti 3mo agoI’ll continue to use the last great reasonably affordable duo from Anthropic: Opus 4.6 for planning and Sonnet 4.6 for implementation.
- nickosh 3mo agoIt looks good. Now waiting for Opus 5.
- epsteingpt 3mo agoIf only the agentic model supported the most popular agents like Hermes and OpenClaw...
- syngrog66 3mo agoI'd rather upgrade myself to a more effective version, thanks. in part because I have a monopoly in the market on providing Me
- Madmallard 3mo ago[flagged]
- impodimium 3mo agoEh still looks like it is weaker than Opus 4.8 but maybe a good replacement for Sonnet 4.x
- Alien1Being 3mo agoOnly if you have no problem with their extremely harmful political lobbying.
- oezen 3mo agoopus is better
- yashthakker 3mo ago[dead]
- nnurmanov 3mo ago[flagged]
- ClaudioCronin 3mo agonice!
- frobisher 3mo agoCosts are very opaque from within the product...
- gertlabs 3mo agoIn our coding evaluations, we found Sonnet 5 is more capable than Sonnet 4.6 (which was an underrated model itself), but is now faster and slightly cheaper. Sonnet 5's performance is comparable to GLM 5.2 in both one-shot coding and agentic ability. However, it's about ~20% less verbose than GLM 5.2 in average code submission sizes, and uses fewer reasoning tokens, which reduces the cost gap and suggests it writes cleaner code. In practice, Sonnet 5 ends up being 40% more expensive and ~2x faster than GLM 5.2 in our evaluations (not 300% more expensive as the per-token pricing would suggest). Granted, GLM 5.2 is an extremely reasoning heavy model. Overall, it's a solid release that gives Anthropic some standing in the price-conscious inference market. Data at https://gertlabs.com/rankings https://gertlabs.com/rankings
- Zababa 3mo agoArtificial analysis shows Sonnet 5 as ~2 times more verbose than GLM 5.2. I wouldn't call Sonnet 4.6 underrated, it's in "chinese open source model territory" and unless you rely only on subscriptions it has alternatives.
- Foobar8568 3mo agoAnd Anthropic put that shit model as default, after a single prompt I was wondering what was the shit it was spouting, and yes, Sonnet 5.
- neonstatic 3mo agoI appreciate they added thinking. Sonnet used to think in the actual response, leading to a lot of unnecessary burden for me. "This thing is X, no wait, it's actually Y. Therefore..." - now it's hidden in the thinking trail, so I don't have to read it unless I want to.
- Escapade5160 3mo agoAt that price you should just use glm-5.2. You get an Opus class model for 1/3 the cost.
- sreekanth850 3mo agoAfter using codex i will never return to cc even if they offer it for free.
- mosbyllc 3mo agoClaude is a great model for me, but unfortunately, its quota is often insufficient. It seems that many people are now considering Codex as an alternative. If the quota is sufficient, I believe many people will continue to use the Claude Code model.
- iLoveOncall 3mo agoNeither Claude, nor Codex, nor Claude Code are models. Claude is a series of models (Claude Sonnet X, Claude Opus X, etc.), Claude Code is their development CLI that uses their models, and Codex is the same as Claude Code but from OpenAI. Ultimately the quota is linked to neither of those 3 directly, rather to which specific model you invoke.
- hdjrudni 3mo agoCodex is not better anymore. It appears they nerfed their quota a few weeks ago. I never used to hit my 5 hr limit, now I always do. Sometimes in like 2 prompts.
- pheewma 3mo agoCodex was running a 2x usage promotion from around the time when Claude introduced rate limiting during peak hours, until May 31st. The various relevant subreddits were (more) insufferable: just 1000 posts per day to the tune of "Just switched the codex! So much more usage!" only to have that tone flip immediately after the promo ran out.
- Wasparrow 3mo ago[flagged]
- richardfey 3mo agoI don't know what I am doing right, or wrong, but I have access to claude and codex and I find myself giving the more serious work to codex recently. I tend to trust it more. I might try again Fable when it's back, but this Sonnet 5 didn't work well for my current projects.
- jaggirs 3mo agoSame (Opus 4.8 vs gpt 5.5) I keep having to correct 4.8, but 5.5 more often than not is correcting me. Opus writes a bit nicer though and it is easier to follow wat it is doing/saying. Not too different experience from talking to humans: 5.5 feels like a very smart 'nerd' that doesn't make a huge effort to communicate wel, while Opus is a bit less intelligent but that makes it's ideas easier to communicate
- mdrzn 3mo agoEdit June 30, 2026: In the original version of this post, we included a cost-performance chart for the BrowseComp evaluation that was based on data from a simpler methodology that did not reflect the standard methodology we use for agentic search evaluations. This had the result of underestimating Sonnet 5's performance on the evaluation. They changed the Sonnet 5 'Agentic search' benchmark graph overnight
- deleted 3mo ago[deleted]
- cavan1977 3mo ago[dead]
- nijave 3mo ago> Me: What was the sushi place near latitude 41 in Columbus? Did it go out of business I don't see it on Google maps anymore > Sonnet 5 (medium): None of these past chats mention a sushi place — I don't have anything on record about that. Do you remember the name, or roughly which part of Columbus (neighborhood/street) it was near? That'll help me search and check its current status. Not impressed. It got the name right on high effort one shot but hallucinated the date relativity (Jan 2026 is not last month...). Worked okay on extra. Sonnet 4.6 worked fine on medium, high, and extra one shot. Edit: as the replies point out, the prompt is definitely ambiguous however Sonnet 5 didn't even extract the semantic meaning "looking for <place> near <compound place>" which all the human replies seem to understand. Even Haiku 4.5 identifies the semantic meaning although it fails to retrieve the correct results > Haiku 4.5 (reasoning off) I'll help you find information about that sushi place. Let me search for sushi restaurants that were near latitude 41 in Columbus and check their current status. >_Searched the web_ > I see that "Latitude 41" is a restaurant in Columbus, but it's actually a modern American restaurant, not specifically a sushi place. However, based on your mention of latitude 41, that's what came up. Let me search more specifically for sushi restaurants that may have closed in Columbus around that area.
- sejje 3mo agoTry it 25 more times and let us know how it averages out. It's non-deterministic, remember?
- nijave 3mo agoI tried 3 more times. Two were nearly identical and 1 recognized Latitude 41 as a restaurant but had a similar useless reply
- ben_w 3mo ago"Latitude 41" is a business name? If you gave me this prompt, I'd say "Which Columbus? None of the ones I know about are at a latitude of 41 degrees north or south?"
- theHocineSaad 3mo agoWhat's interesting is that Claude Sonnet 5 costs more per task ($2.29) than Opus 4.8 ($1.80), while the latter is obviously better! It actually costs more per task than every other model. It's only cheaper than Claude Fable 5. Source: https://artificialanalysis.ai/?cost=cost-per-task#price-and-cost https://artificialanalysis.ai/?cost=cost-per-task#price-and-..., as of writing this comment (the results are frequently changing)
- terekhindc 3mo agocost per task > opus-low is a weird place to land. is there a specific task shape where sonnet 5 medium actually wins?
- runnig 3mo agoI tried Sonnet 5 and burned the entire 5h quota on a single deep research run. This has never happened with Opus before.
- linzhangrun 3mo agoWhat happened to Anthropic these past few months? One bad move after another.
- hamnxrye 3mo ago[dead]
- linzhangrun 3mo agoThe biggest issue with Claude Sonnet 5 is GLM 5.2
- caine22 3mo agoBeen spamming Opus 4.8 high effort (pro sub) on Claude Code (WSL 2 Terminal), everywhere actually, rarely ran out of limits. But then again the work isn't too heavy. Nevertheless, Sonnet 5 could still come in handy during long hours of heavy work
- troglodytetrain 3mo agoAnother great product from anthropic where I need to pay hundreds upon hundreds of dollars for the privilege of being the unpaid QA team for the multi trillion dollar corporation.
- de6u99er 3mo agoI switched to Sonnet 5. My impression is, that Claude tries to use more and more tokens. E.g. doing playwright with screenshots for things that i could have easily confirmed. Doing unneccessari ntermediate builds with tons of logs to process, ... Honestly, after the first good impression such behavior makes me start losing trust in Anthropic.