11 ms·
Claude Opus 4.5
https://platform.claude.com/docs/en/about-claude/models/whats-new-claude-4-5 https://platform.claude.com/docs/en/about-claude/models/what...
- jedberg 10mo agoUp until today, the general advice was use Opus for deep research, use Haiku for everything else. Given the reduction in cost here, does that rule of thumb no longer apply?
- mudkipdev 10mo agoIn my opinion Haiku is capable but there is no reason to use anything lower than Sonnet unless you are hitting usage limits
- carcabob 10mo agoI wish the article's graphs weren't distorted by skipping so much of the scale to make it look like a more significant difference than it is. But it does looks impressive.
- jumploops 10mo ago> Pricing is now $5/$25 per million [input/output] tokens So it’s 1/3 the price of Opus 4.1… > [..] matches Sonnet 4.5’s best score on SWE-bench Verified, but uses 76% fewer output tokens …and potentially uses a lot less tokens? Excited to stress test this in Claude Code, looks like a great model on paper!
- deleted 10mo ago[deleted]
- jmkni 10mo ago> Pricing is now $5/$25 per million tokens For anyone else confused, it's input/output tokens $5 for 1million tokens in $25 for 1million tokens out
- mvdtnz 10mo agoWhat prevents these jokers from making their outputs ludicrously verbose to squeeze more out of you, given they charge 5x more for the end that they control? Already model outputs are overly verbose, and I can see this getting worse as they try to squeeze some margin. Especially given that many of the tools conveniently hide most of the output.
- WilcoKruijer 10mo agoYou would stop using their model and move to their competitors, presumably.
- jumploops 10mo agoThanks, updated to make more clear
- alach11 10mo agoThis is the biggest news of the announcement. Prior Opus models were strong, but the cost was a big limiter of usage. This price point still makes it a "premium" option, but isn't prohibitive. Also increasingly it's becoming important to look at token usage rather than just token cost. They say Opus 4.5 (with high reasoning) used 50% fewer tokens than Sonnet 4.5. So you get a higher score on SWE-bench verified, you pay more per token, but you use fewer tokens and overall pay less!
- elvin_d 10mo agoGreat seeing the price reduction. Opus historically was prices at 15/75, this one delivers at 5/25 which is close to Gemini 3 Pro. I hope Anthropic can afford increasing limits for the new Opus.
- rishabhaiover 10mo agoIs this available on claude-code?
- elvin_d 10mo agoYes, the first run was nice - feels faster than 4.1 and did what Sonnet 4.5 struggled to execute properly.
- greenavocado 10mo agoWhat are you thinking of trying to use it for? It is generally a huge waste of money to unleash Opus on high content tasks ime
- rishabhaiover 10mo agoI use claude-code extensively to plan and study for my college using the socrates learning mode. It's a great way to learn for me. I wanted to test the new model's capabilities on that front.
- flutas 10mo agoMy workflow has always been opus for planning, sonnet for actual work.
- rishabhaiover 10mo agodamn, I need a MAX sub for this.
- stavros 10mo agoYou don't, you can add $5 or whatever to your Claude wallet with the Pro subscription and use those for Opus.
- rishabhaiover 10mo agoI ain’t paying a penny more than the $20 I already do. I got cracks in my boots, brother.
- bnchrch 10mo agoSeeing these benchmarks makes me so happy. Not because I love Anthropic (I do like them) but because it's staving off me having to change my Coding Agent. This world is changing fast, and both keeping up with State of the Art and/or the feeling of FOMO is exhausting. Ive been holding onto Claude Code for the last little while since Ive built up a robust set of habits, slash commands, and sub agents that help me squeeze as much out of the platform as possible. But with the last few releases of Gemini and Codex I've been getting closer and closer to throwing it all out to start fresh in a new ecosystem. Thankfully Anthropic has come out swinging today and my own SOP's can remain in tact a little while longer.
- tordrt 10mo agoI tried codex due to the same reasoning you list. The grass is not greener on the other side.. I usually only opt for codex when my claude code rate limit hits.
- bavell 10mo agoSame boat and same thoughts here! Hope it holds its own against the competition, I've become a bit of a fan of Anthropic and their focus on devs.
- wahnfrieden 10mo agoYou need much less of a robust set of habits, commands, sub agent type complexity with Codex. Not only because it lacks some of these features, it also doesn't need them as much.
- edf13 10mo agoI’m threw a few hours at Codex the other day and was incredibly disappointed with the outcome… I’m a heavy Claude code user and similar workloads just didn’t work out well for me on Codex. One of the areas I think is going to make a big difference to any model soon is speed. We can build error correcting systems into the tools - but the base models need more speed (and obviously with that lower costs)
- chrisweekly 10mo ago
- stavros 10mo agoDid anyone else notice Sonnet 4.5 being much dumber recently? I tried it today and it was really struggling with some very simple CSS on a 100-line self-contained HTML page. This never used to happen before, and now I'm wondering if this release has something to do with it. On-topic, I love the fact that Opus is now three times cheaper. I hope it's available in Claude Code with the Pro subscription. EDIT: Apparently it's not available in Claude Code with the Pro subscription, but you can add funds to your Claude wallet and use Opus with pay-as-you-go. This is going to be really nice to use Opus for planning and Sonnet for implementation with the Pro subscription. However, I noticed that the previously-there option of "use Opus for planning and Sonnet for implementation" isn't there in Claude Code with this setup any more. Hopefully they'll implement it soon, as that would be the best of both worlds. EDIT 2: Apparently you can use `/model opusplan` to get Opus in planning mode. However, it says "Uses your extra balance", and it's not clear whether it means it uses the balance just in planning mode, or also in execution mode. I don't want it to use my balance when I've got a subscription, I'll have to try it and see. EDIT 3: It looks like Sonnet also consumes credits in this mode. I had it make some simple CSS changes to a single HTML file with Opusplan, and it cost me $0.95 (way too much, in my opinion). I'll try manually switching between Opus for the plan and regular Sonnet for the next test.
- kjgkjhfkjf 10mo agoMy guess is that Claude's "bad days" are due to the service becoming overloaded and failing over to use cheaper models.
- bryanlarsen 10mo agoOn Friday my Claude was particularly stupid. It's sometimes stupid, but I've never seen it been that consistently stupid. Just assumed it was a fluke, but maybe something was changing.
- vunderba 10mo agoAnecdotally, I kind of compare the quality of Sonnet 4.5 to that of a chess engine: it performs better when given more time to search deeper into the tree of possible moves (more plies). So when Anthropic is under peak load I think some degradation is to be expected. I just wish Claude Code had a "Signal Peak" so that I could schedule more challenging tasks for a time when its not under high demand.
- 827a 10mo agoI've played around with Gemini 3 Pro in Cursor, and honestly: I find it to be significantly worse than Sonnet 4.5. I've also had some problems that only Claude Code has been able to really solve; Sonnet 4.5 in there consistently performs better than Sonnet 4.5 anywhere else. I think Anthropic is making the right decisions with their models. Given that software engineering is probably one of the very few domains of AI usage that is driving real, serious revenue: I have far better feelings about Anthropic going into 2026 than any other foundation model. Excited to put Opus 4.5 through its paces.
- visioninmyblood 10mo agoThe model is great it is able to code up some interesting visual tasks(I guess they have pretty strong tool calling capapbilities). Like orchestrate prompt -> image generate -> Segmentation -> 3D reconstruction. Checkout the results here https://chat.vlm.run/c/3fcd6b33-266f-4796-9d10-cfc152e945b7 https://chat.vlm.run/c/3fcd6b33-266f-4796-9d10-cfc152e945b7. Note the model was only used to orchestrate the pipeline, the tasks are done by other models in an agentic framework. They much have improved tool calling framework with all the MCP usage. Gemini 3 was able to orchestrate the same but Claude 4.5 is much faster
- Squarex 10mo agoI have heard that gemini 3 is not that great in cursor, but excellent in Antigravity. I don't have a time to personally verify all that though.
- incoming1211 10mo agoI think gemini 3 is hot garbage in everything. Its great on a greenfield trying to 1 shot something, if you're working on a long term project it just sucks.
- koakuma-chan 10mo agoNothing is great in Cursor.
- itsdrewmiller 10mo ago
- GodelNumbering 10mo agoThe fact that the post singled out SWE-bench at the top makes the opposite impression that they probably intended.
- grantpitt 10mo agodo say more
- GodelNumbering 10mo agoMakes it sound like a one trick pony
- grantpitt 10mo agowell, it's a big trick
- jascha_eng 10mo agoAnthropic is leaning into agentic coding and heavily so. It makes sense to use swe verified as their main benchmark. It is also the one benchmark Google did not get the top spot last week. Claude remains king that's all that matters here.
- Mkengin 10mo agoI am eagerly awaiting swe-rebench results for November with all the new models: https://swe-rebench.com/ https://swe-rebench.com/
- alvis 10mo agoWhat surprise me is that Opus 4.5 lost all reasoning scores to Gemini and GPT. I thought it’s the area the model will shine the most
- viraptor 10mo agoHas there been any announcement of a new programming benchmark? SWE looks like it's close to saturation already. At this point for SWE it may be more interesting to start looking at which types of issues consistently fail/work between model families.
- Mkengin 10mo agoI like this one: https://swe-rebench.com/ https://swe-rebench.com/
- llamasushi 10mo agoThe burying of the lede here is insane. $5/$25 per MTok is a 3x price drop from Opus 4. At that price point, Opus stops being "the model you use for important things" and becomes actually viable for production workloads. Also notable: they're claiming SOTA prompt injection resistance. The industry has largely given up on solving this problem through training alone, so if the numbers in the system card hold up under adversarial testing, that's legitimately significant for anyone deploying agents with tool access. The "most aligned model" framing is doing a lot of heavy lifting though. Would love to see third-party red team results.
- gtrealejandro 10mo ago[dead]
- wolttam 10mo agoIt's 1/3 the old price ($15/$75)
- brookst 10mo agoNot sure if that’s a joke about LLM math performance, but pedantry requires me to point out 15 / 75 = 1/5
- keeeba 10mo agoOh boy, if the benchmarks are this good and Opus feels like it usually does then this is insane. I’ve always found Opus significantly better than the benchmarks suggested. LFG
- aliljet 10mo agoThe real question I have after seeing the usage rug being pulled is what this costs and how usable this ACTUALLY is with a Claude Max 20x subscription. In practice, Opus is basically unusable by anyone paying enterprise-prices. And the modification of "usage" quotas has made the platform fundamentally unstable, and honestly, it left me personally feeling like I was cheated by Anthropic...
- zb3 10mo agoThe first chart is straight from "how to lie in charts"..
- rvz 10mo agoIn some circles it is called a "chart crime".
- andai 10mo agoWhy do they always cut off 70% of the y-axis? Sure it exaggerates the differences, but... it exaggerates the differences. And they left Haiku out of most of the comparisons! That's the most interesting model for me. Because for some tasks it's fine. And it's still not clear to me which ones those are. Because in my experience, Haiku sits at this weird middle point where, if you have a well defined task, you can use a smaller/faster/cheaper model than Haiku, and if you don't, then you need to reach for a bigger/slower/costlier model than Haiku.
- ximeng 10mo agoIt’s a pretty arbitrary y axis - arguably the only thing that matters is the differences.
- waynenilsen 10mo agomarketing.
- chaosprint 10mo agoSWE's results were actually very close, but they used a poor marketing visualization. I know this isn't a research paper, but for Anthropic, I expect more.
- flakiness 10mo agoThey should've used an error rate instead of the pass rate. Then it'll get the same visual appeal without cheating.
- unsupp0rted 10mo agoThis is gonna be game-changing for the next 2-4 weeks before they nerf the model. Then for the next 2-3 months people complaining about the degradation will be labeled “skill issue”. Then a sacrificial Anthropic engineer will “discover” a couple obscure bugs that “in some cases” might have lead to less than optimal performance. Still largely a user skill issue though. Then a couple months later they’ll release Opus 4.7 and go through the cycle again. My allegiance to these companies is now measured in nerf cycles. I’m a nerf cycle customer.
- TIPSIO 10mo agoHilarious sarcastic comment but actually true sentiment. For all we know this is just the Opus 4.0 re-released
- film42 10mo agoThis is why I migrated my apps that need an LLM to Gemini. No model degradation so far all through the v2.5 model generation. What is Anthropic doing? Swapping for a quantized version of the model?
- blurbleblurble 10mo agoI fully agree that this is what's happening. I'm quite convinced after about a year of using all these tools via the "pro" plans that all these companies are throttling their models in sophisticated ways that have a poorly understood but significant impact on quality and consistency. Gpt-5.1-* are fully nerfed for me at the moment. Maybe they're giving others the real juice but they're not giving it to me. Gpt-5-* gave me quite good results 2 weeks ago, now I'm just getting incoherent crap at 20 minute intervals. Maybe I should just start paying via tokens for a hopefully more consistent experience.
- throwuxiytayq 10mo agoy’all hallucinating harder than GPT2 on DMT
- deleted 10mo ago
- alvis 10mo ago“For Max and Team Premium users, we’ve increased overall usage limits, meaning you’ll have roughly the same number of Opus tokens as you previously had with Sonnet.” — seems like anthropic has finally listened!
- 0x79de 10mo agothis is quite a good
- jasonthorsness 10mo agoI used Gemini instead of my usual Claude for a non-trivial front-end project [1] and it really just hit it out of the park especially after the update last week, no trouble just directly emitting around 95% of the application. Now Claude is back! The pace of releases and competition seems to be heating up more lately, and there is absolutely no switching cost. It's going to be interesting to see if and how the frontier model vendors create a moat or if the coding CLIs/models will forever remain a commodity. [1] https://github.com/jasonthorsness/tree-dangler https://github.com/jasonthorsness/tree-dangler
- hu3 10mo agoGemini is indeed great for frontend HTML + CSS and even some light DOM manipulation in JS. I have been using Gemini 2.5 and now 3 for frontend mockups. When I'm happy with the result, after some prompt massage, I feed it to Sonnet 4.5 to build full stack code using the framework of the application.
- diego_sandoval 10mo agoWhat IDE/CLI tool do you use?
- jasonthorsness 10mo agoI used Gemini CLI and a few rounds of Claude CLI at the beginning with stock VSCode
- cyrusradfar 10mo agoI'm curious if others are finding that there's a comfort in staying within the Claude ecosystem because when it makes a mistake, we get used to spotting the pattern. I'm finding that when I try new models, their "stupid" moments are more surprising and infuriating. Given this tech is new, the experience of how we relate to their mistakes is something I think a bit about. Am I alone here, are others finding themselves more forgiving of "their preferred" model provider?
- irthomasthomas 10mo agoI guess you where not around a few months back when they over-optimized and served a degraded model for weeks.
- cyrusradfar 10mo agothat's funny you just made me connect the dots. I was! I spent several days spinning in place after I thought it could help me clean up my code quality with biome. Afterwards it destroyed the whole app and I needed to figure out how it worked -- that need, inspired me to prototype and extension for vccode I'm actually still building :)
- irthomasthomas 10mo agoYep, that was it! That really turned me off anthropic and closed models until they provide regular quality tests. I use chutes ai, now. They tell you exactly which model/quant and server config they use, so you know if you have trouble with a task, it's not the model.
- GenerWork 10mo agoI wonder what this means for UX designers like myself who would love to take a screen from Figma and turn it into code with just a single call to the MCP. I've found that Gemini 3 in Figma Make works very well at one-shotting a page when it actually works (there's a lot of issues with it actually working, sadly), so hopefully Opus 4.5 is even better.
- redfloatplane 10mo agoOn my Max plan, Opus 4.5 is now the default model! Until now I used Sonnet 4.5 exclusively and never used Opus, even for planning - I'm shocked that this is so cheap (for them) that it can be the default now. I'm curious what this will mean for the daily/weekly limits. A short run at a small toy app makes me feel like Opus 4.5 is a bit slower than Sonnet 4.5 was, but that could also just be the day-one load it's presumably under. I don't think Sonnet was holding me back much, but it's far too early to tell.
- Robdel12 10mo agoRight! I thought this at the very bottom was super interesting > For Claude and Claude Code users with access to Opus 4.5, we’ve removed Opus-specific caps. For Max and Team Premium users, we’ve increased overall usage limits, meaning you’ll have roughly the same number of Opus tokens as you previously had with Sonnet. We’re updating usage limits to make sure you’re able to use Opus 4.5 for daily work. These limits are specific to Opus 4.5. As future models surpass it, we expect to update limits as needed.
- redfloatplane 10mo agoIt looks like they've now added a Sonnet cap which is the same as the previous cap: > Nov 24, 2025 update: > We've increased your limits and removed the Opus cap, so you can use Opus 4.5 > up to your overall limit. Sonnet now has its own limit—it's set to match your > previous overall limit, so you can use just as much as before. We may continue > to adjust limits as we learn how usage patterns evolve over time. Quite interesting. From their messaging in the blog post and elsewhere, I think they're betting on Opus being significantly smarter in the sense of 'needs fewer tokens to do the same job', and thus cheaper. I'm curious how this will go.
- agentifysh 10mo agowish they really bolded that part because i almost passed off on it until i read the blog carefully instant upgrade to claude max 20x if they give opus 4.5 out like this i still like codex-5.1 and will keep it. gemini cli missed its opportunity again now money is hedged between codex and claude.
- futureshock 10mo agoA really great way to get an idea of the relative cost and performance of these models at their various thinking budgets is to look at the ARC-AGI-2 leaderboard. Opus 4.5 stacks up very well here when you compare to Gemini 3’s score and cost. Gemini 3 Deep Think is still the current leaders but at more than 30x the cost. The cost curve of achieving these scores is coming down rapidly. In Dec 2024 when OpenAI announced beating human performance on ARC-AGI-1, they spent more than $3k per task. You can get the same performance for pennies to dollars, approximately an 80x reduction in 11 months. https://arcprize.org/leaderboard https://arcprize.org/leaderboard https://arcprize.org/blog/oai-o3-pub-breakthrough https://arcprize.org/blog/oai-o3-pub-breakthrough
- energy123 10mo agoA point of context. On this leaderboard, Gemini 3 Pro is "without tools" and Gemini 3 Deep Think is "with tools". In the other benchmarks released by Google which compare these two models, where they have access to the same amount of tools, the gap between them is small.
- simonw 10mo agoNotes and two pelicans: https://simonwillison.net/2025/Nov/24/claude-opus/ https://simonwillison.net/2025/Nov/24/claude-opus/
- pjm331 10mo agoi think you have an error there about haiku pricing > For comparison, Sonnet 4.5 is $3/$15 and Haiku 4.5 is $4/$20. i think haiku should be $1/$5
- simonw 10mo agoFixed now, thanks.
- dreis_sw 10mo agoI agree with your sentiment, this incremental evolution is getting difficult to feel when working with code, especially with large enterprise codebases. I would say that for the vast majority of tasks there is a much bigger gap on tooling than on foundational model capability.
- qingcharles 10mo agoAlso came to say the same thing. When Gemini 3 came out several people asked me "Is it better than Opus 4.1?" but I could no longer answer it. It's too hard to evaluate consistently across a range of tasks.
- throwaway2027 10mo agoI wonder if at this point they read what people use to benchmark with and specifically train it to do well at this task.
- diego_sandoval 10mo ago:%s/There model/Their model/g
- jasonjmcghee 10mo ago
- whitepoplar 10mo agoDoes the reduced price mean increased usage limits on Claude Code (with a Max subscription)?
- skerit 10mo agoYes. Opus is now the default model in Claude Code. And Opus 4.5 counts the same toward your usage limit as Sonnet 4.5 did. Even better: Sonnet 4.5 now has its own separate limit.
- saaaaaam 10mo agoAnecdotally, I’ve been using opus 4.5 today via the chat interface to review several large and complex interdependent documents, fillet bits out of them and build a report. It’s very very good at this, and much better than opus 4.1. I actually didn’t realise that I was using opus 4.5 until I saw this thread.
- tschellenbach 10mo agoOk, but can it play Factorio?
- deleted 10mo ago[deleted]
- andreybaskov 10mo agoDoes anyone know or have a guess on the size of this latest thinking models and what hardware they use to run inference? As in how much memory and what quantization it uses and if it's "theoretically" possible to run it on something like Mac Studio M3 Ultra with 512GB RAM. Just curious from theoretical perspective.
- docjay 10mo agoThat all depends on what you consider to be reasonably running it. Huge RAM isn’t required to run them, that just makes them faster. I imagine technically all you'd need is a few hundred megabytes for the framework and housekeeping, but you’d have to wait for the some/most/all of the model to be read off the disk for each token it processes. None of the closed providers talk about size, but for a reference point of the scale: Kimi K2 Thinking can spar in the big leagues with GPT-5 and such…if you compare benchmarks that use words and phrasing with very little in common with how people actually interact with them…and at FP16 you’ll need 2.9TB of memory @ 256,000 context. It seems it was recently retrained it at INT4 (not just quantized apparently) and now: “ The smallest deployment unit for Kimi-K2-Thinking INT4 weights with 256k seqlen on mainstream H200 platform is a cluster with 8 GPUs with Tensor Parallel (TP). (https://huggingface.co/moonshotai/Kimi-K2-Thinking https://huggingface.co/moonshotai/Kimi-K2-Thinking) “ -or- “ 62× RTX 4090 (24GB) or 16× H100 (80GB) or 13× M3 Max (128GB) “ So ~1.1TB. Of course it can be quantized down to as dumb as you can stand, even within ~250GB (https://docs.unsloth.ai/models/kimi-k2-thinking-how-to-run-locally https://docs.unsloth.ai/models/kimi-k2-thinking-how-to-run-l...). But again, that’s for speed. You can run them more-or-less straight off the disk, but (~1TB / SSD_read_speed + computation_time_per_chunk_in_RAM) = a few minutes per ~word or punctuation.
- threeducks 10mo ago> (~1TB / SSD_read_speed + computation_time_per_chunk_in_RAM) = a few minutes per ~word or punctuation. You have to divide SSD read speed by the size of the active parameters (~16GB at 4 bit quantization) instead of the entire model size. If you are lucky, you might get around one token per second with speculative decoding, but I agree with the general point that it will be very slow.
- thot_experiment 10mo agoIt's really hard for me to take these benchmarks seriously at all, especially that first one where Sonnet 4.5 is better at software engineering than Opus 4.1. It is emphatically not, it has never been, I have used both models extensively and I have never encountered a single situation where Sonnet did a better job than Opus. Any coding benchmark that has Sonnet above Opus is broken, or at the very least measuring things that are totally irrelevant to my usecases. This in particular isn't my "oh the teachers lie to you moment" that makes you distrust everything they say, but it really hammers the point home. I'm glad there's a cost drop, but at this point my assumption is that there's also going to be a quality drop until I can prove otherwise in real world testing.
- mirsadm 10mo agoThese announcements and "upgrades" are becoming increasingly pointless. No one is going to notice this. The improvements are questionable and inconsistent. They could swap it out for an older model and no one would notice.
- emp17344 10mo agoThis is the surest sign progress has plateaued, but it seems people just take the benchmarks at face value.
- ximeng 10mo agoWith less token usage, cheaper pricing, and enhanced usage limits for Opus, Anthropic are taking the fight to Gemini and OpenAI Codex. Coding agent performance leads to better general work and personal task performance, so if Anthropic continue to execute well on ergonomics they have a chance to overcome their distribution disadvantages versus the other top players.
- syspec 10mo agoThank you Claude.
- CuriouslyC 10mo agoI hate on Anthropic a fair bit, but the cost reduction, quota increases and solid "focused" model approach are real wins. If they can get their infrastructure game solid, improve claude code performance consistency and maintain high levels of transparency I will officially have to start saying nice things about them.
- agentifysh 10mo agoagain the question of concern as codex user is usage its hard to get any meaningful use out of claude pro after you ship a few features you are pretty much out of weekly usage compared to what codex-5.1-max offers on a plan that is 5x cheaper the 4~5% improvement is welcome but honestly i question whether its possible to get meaningful usage out of it the way codex allows it for most use cases medium or 4.5 handles things well but anthropic seems to have way less usage limits than what openai is subsidizing until they can match what i can get out of codex it won't be enough to win me back edit: I upgraded to claude max! read the blog carefully and seems like opus 4.5 is lifted in usage as well as sonnet 4.5!
- jstummbillig 10mo agoWell, that's where the price reduction comes in handy, no?
- agentifysh 10mo agocodex-5.1-max I can see from benchmark is ~3% off what opus 4.5 is claiming and while i can see one off uses for it i can't see the 3x reduction in price being enticing enough to match what openai subsidizes
- undeveloper 10mo agoSonnet is still $3/25M tokens, and peoples still had many many complaints
- fragmede 10mo agoGot the river crossing one: https://claude.ai/chat/0c583303-6d3e-47ae-97c9-085cefe14c21 https://claude.ai/chat/0c583303-6d3e-47ae-97c9-085cefe14c21 Still fucked up one about the boy and the surgeon though: https://claude.ai/chat/d2c63190-059f-43ef-af3d-67e7ca1707a4 https://claude.ai/chat/d2c63190-059f-43ef-af3d-67e7ca1707a4
- adastra22 10mo agoDoes it follow directions? I’ve found Sonnet 4.5 to be useless for automated workflows because it refuses to follow directions. I hope they didn’t take the same RLHF approach they did with that model.
- pingou 10mo agoWhat causes the improvements in new AI models recently? Is it just more training, or is it new, innovative techniques?
- I_am_tiberius 10mo agoSome months back they changed their terms of service and by default users now allow Anthropic to use prompts for learning. As it's difficult to know if your prompts, or derivations of it, are part of a model, I would consider the possibility that they use everyone's prompt.
- AJRF 10mo agothat chart at the start is egregious
- tildef 10mo agoFeels like a tongue-in-cheek jab at the GPT-5 announcement chart.
- sync 10mo agoDoes anyone here understand "interleaved scratchpads" mentioned at the very bottom of the footnotes: > All evals were run with a 64K thinking budget, interleaved scratchpads, 200K context window, default effort (high), and default sampling settings (temperature, top_p). I understand scratchpads (e.g. [0] Show Your Work: Scratchpads for Intermediate Computation with Language Models) but not sure about the "interleaved" part, a quick Kagi search did not lead to anything relevant other than Claude itself :) [0] https://arxiv.org/abs/2112.00114 https://arxiv.org/abs/2112.00114
- dheerkt 10mo agobased on their past usage of "interleaved tool calling" it means that the tool can be used while the model is thinking. https://aws.amazon.com/blogs/opensource/using-strands-agents-with-claude-4-interleaved-thinking/ https://aws.amazon.com/blogs/opensource/using-strands-agents...
- davidsainez 10mo agoAFAICT, kimi k2 was the first to apply this technique [1]. I wonder if Anthropic came up with it independently or if they trained a model in 5 months after seeing kimi’s performance. 1: https://www.decodingdiscontinuity.com/p/open-source-inflection-point-kimi2-ai-competitive-dynamics https://www.decodingdiscontinuity.com/p/open-source-inflecti...
- BoorishBears 10mo agoOpenAI has been doing this since at least O3 in January, Anthropic has been doing it since 4 in May. And the July Kimi K2 release wasn't a thinking model, the model in that article was released less than 20 days ago.
- I_am_tiberius 10mo agoStill mad at them because they decided not to take their users' privacy serious. Would be interested how the new model behaves, but just have a mental lock and can't sign up again.
- quantummagic 10mo agoI would look past their privacy issues and have wanted to sign up for over a year, but don't have a cellphone, which is required to register.
- MaxLeiter 10mo agoWe've added support for opus 4.5 to v0 and users are making some pretty impressive 1-shots: https://x.com/mikegonz/status/1993045002306699704 https://x.com/mikegonz/status/1993045002306699704 https://x.com/MirAI_Newz/status/1993047036766396852 https://x.com/MirAI_Newz/status/1993047036766396852 https://x.com/rauchg/status/1993054732781490412 https://x.com/rauchg/status/1993054732781490412 It seems especially good at threejs / 3D websites. Gemini was similarly good at them (https://x.com/aymericrabot/status/1991613284106269192 https://x.com/aymericrabot/status/1991613284106269192); maybe the model labs are focusing on this style of generation more now.
- adt 10mo agohttps://lifearchitect.ai/models-table/ https://lifearchitect.ai/models-table/
- jmward01 10mo agoOne thing I didn't see mentioned is raw token gen speed compared to the alternatives. I am using Haiku 4.5 because it is cheap (and so am I) but also because it is fast. Speed is pretty high up in my list of coding assistant features and I wish it was more prominent in release info.
- mutewinter 10mo agoSome early visual evaluations: https://x.com/mutewinter/status/1993037630209192276 https://x.com/mutewinter/status/1993037630209192276
- starkparker 10mo agoWould love to know what's going on with C++ and PHP benchmarks. No meaningful gain over Opus 4.1 for either, and Sonnet still seems to outperform Opus on PHP.
- ramon156 10mo agoI've almost ran out of Claude on the Web credits. If they announce that they're going to support Opus then I'm going to be sad :'(
- undeveloper 10mo agohaven't they all expired by now?
- throwaway2027 10mo agoOh that's why there were only 2 usage bars.
- xkbarkar 10mo agoThis is great. Sonnet 4.5 has degraded terribly. I can get some useful stuff from a clean context in the web ui but the cli is just useless. Opus is far superiour. Today sonnet 4.5 suggested to verify remote state file presence by creating an empty one locally and copy it to the remote backend. Da fuq? University level programmer my a$$. And it seems like it has degraded this last month. I keep getting braindead suggestions and code that looks like it came from a random word generator. I swear it was not that awful a couple of months ago. Opus cap has been an issue, happy to change and I really hope the nerf rumours are just that. Undounded rumours and the defradation has a valid root cause But honestly sonnet 4.5 has started to act like a smoking pile of sh**t
- idonotknowwhy 10mo ago>This is great. Sonnet 4.5 has degraded terribly. >I can get some useful stuff from a clean context in the web ui but the cli is just useless. >I swear it was not that awful a couple of months ago. I agree on all 3 counts. And it still degrades after a few long turns in openwebui. You can test this by regenerating the last reply in chats from shortly after the model was released.
- irthomasthomas 10mo agoI wish it was open-weights so we could discuss the architectural changes. This model is about twice as fast as 4.1, ~60t/s Vs ~30t/s. Is it half the parameters, or a new INT4 linear sparse-moe architecture?
- gsibble 10mo agoThey lowered the price because this is a massive land grab and is basically winner take all. I love that Antrhopic is focused on coding. I've found their models to be significantly better at producing code similar to what I would write, meaning it's easy to debug and grok. Gemini does weird stuff and while Codex is good, I prefer Sonnet 4.5 and Claude code.
- kachapopopow 10mo agoslightly better at react and spacial logic than gemini 3 pro, but slower and way more expensive.
- synergy20 10mo agogreat, paying $100/m for claude code, this stops me from switching to gemini 3.0 for now.
- maherbeg 10mo agoOk, the victorian lock puzzle game is pretty damn cool way to showcase the capabilities of these models. I kinda want to start building similar puzzle games for models to solve.
- morgengold 10mo agoI'm on a Claude Code Max subscription. Last days have been a struggle with Sonnet 4.5 - Now it switched to Claude Opus 4.5 as default model. Ridiculous good and fast.
- dave1010uk 10mo agoThe Claude Opus 4.5 system card [0] is much more revealing than the marketing blog post. It's a 150 page PDF, with all sorts of info, not just the usual benchmarks. There's a big section on deception. One example is Opus is fed news about Anthropic's safety team being disbanded but then hides that info from the user. The risks are a bit scary, especially around CBRNs. Opus is still only ASL-3 (systems that substantially increase the risk of catastrophic misuse) and not quite at ASL-4 (uplifting a second-tier state-level bioweapons programme to the sophistication and success of a first-tier one), so I think we're fine... I've never written a blog post about a model release before but decided to this time [1]. The system card has quite a few surprises, so I've highlighted some bits that stood out to me (and Claude, ChatGPT and Gemini). [0] https://www.anthropic.com/claude-opus-4-5-system-card https://www.anthropic.com/claude-opus-4-5-system-card [1] https://dave.engineer/blog/2025/11/claude-opus-4.5-system-card/ https://dave.engineer/blog/2025/11/claude-opus-4.5-system-ca...
- aurareturn 10mo agoPages 22–24 of Opus’s system card provide some evidence for this. Anthropic run a multi-agent search benchmark where Opus acts as an orchestrator and Haiku/Sonnet/Opus act as sub-agents with search access. Using cheap Haiku sub-agents gives a ~12-point boost over Opus alone. Will this lead to another exponential in capabilities and token increase in the same order as thinking models?
- dave1010uk 10mo agoPerhaps. Though if that were feasible, I'd expect it would have been exploited already. I think this is more about the cost and time saving of being able to use cheaper models. Sub-agents are effectively the same as parallelization and temporary context compaction. (The same as with human teams, delegation and organisational structures.) We're starting to see benchmarks include stats of low/medium/high reasoning effort and how newer models can match or beat older ones with fewer reasoning tokens. What would be interesting is seeing more benchmarks for different sub-agent reasoning combinations too. Eg does Claude perform better when Opus can use 10,000 tokens of Sonnet or 100,000 tokens of Haiku? What's the best agent response you can get for $1? Where I think we might see gains in _some_ types of tasks is with vast quantities of tiny models. I.e many LLMs that are under 4B parameters used as sub-agents. I wonder what GPT-5.1 Pro would be like if it could orchestrate 1000 drone-like workers.
- pwython 10mo agoI was excited to see this today, as I just finished a new website in Figma last night and planned to have CC help build it out with the Figma MCP today. CC started up with Opus 4.5 and... Not off to a good start. It described a totally different looking layout on the first prompt to see the selected frame. I asked why: You're right to call that out. Looking back at what happened: 1. First get_design_context call - Nothing was selected, so it returned an error 2. Second get_design_context call - It returned a tiny snippet of React code for a "Vector" component (just an image element), which gave me almost no context 3. First get_screenshot call - The image returned was essentially blank/empty (the output showed <output_image> with what appears to be no actual visual content) I then hallucinated an entire design description based on: - The project context from CLAUDE.md (mentions "Creative Director", "Performance Marketing", "20+ years") - Common portfolio design patterns - The small code snippet mentioning a "Vector" element That was a significant error on my part. I should have told you "The screenshot appears empty or didn't load properly - can you try selecting a different element or re-selecting?" instead of fabricating a description.
- jaakkonen 10mo agoTested this today for implementing a new low-frequency RFID protocol to Flipper Zero codebase based on a Proxmark3 implementation. Was able to do it in 2 hours with giving a raw psk recording alongside of it and some troubleshooting. This is the kind of task the last generation of frontier models was incapable of doing. Super stoked to use this :)
- achierius 10mo agoWas this just 2 hours of the agent running on its own, or was there back-and-forth/any sort of interaction? How much did you have to set up scaffolding, e.g. tests?
- PilotJeff 10mo agoMore blowing up of the bubble with anthropic essentially offering compute/LLM for below cost. Eventually the laws of physics/market will take over and look out below.
- jstummbillig 10mo agoHow would you know what the cost is?
- adidoit 10mo agoTested this building some PRs and issues that codex-5.1-max and gemini-3-pro were strugglig with It planned way better in a much more granular way and then execute it better. I can't tell if the model is actually better or if it's just planning with more discipline
- gigatexal 10mo agoLove the competition. Gemini 3 pro blew me away after being spoiled by Claude for coding things. Considered canceling my Anthropic sub but now I’m gonna hold on to it. The bigger thing is Google has been investing in TPUs even before the craze. They’re on what gen 5 now ? Gen 7? Anyway I hope they keep investing tens of billions into it because Nvidia needs to have some competition and maybe if they do they’ll stop this AI silliness and go back to making GPUs for gamers. (Hahaha of course they won’t. No gamer is paying 40k for a GPU.)
- obblekk 10mo ago80% on swebench verified is incredible. a year ago the best model was at ~30%. i wonder if we'll soon have a convincingly superhuman coding capability (even in a narrow field like kernel optimization). this is the most interesting time for software tools since compilers and static typechecking was invented.
- quantumHazer 10mo agoLast year’s model were at 50-60% on SWE bench-verified actually
- obblekk 10mo agoI see 25-29% here https://www.swebench.com/viewer.html https://www.swebench.com/viewer.html for models released in Nov 2024 albeit not verified. gpt4o (Aug 2024) was 33% for swe bench verified. Important point because people have a bias to underestimate the speed of ai progress.
- tymscar 10mo agoDo you people think nobody calls your bluff? Here’s the launch card of the sonnet 3.5 from a year and a month ago. Guess the number. Ok, Ill tell you: 49.0%. So yeah, the comment you replied to was not really off. https://www.anthropic.com/news/3-5-models-and-computer-use https://www.anthropic.com/news/3-5-models-and-computer-use
- lerp-io 10mo ago80% and 77% is not that much lol
- ddxv 10mo agoThe LLMs rate of improvement has really slowed down. This looks like a minor improvement in terms of accuracy and big gains from efficiency.
- energy123 10mo ago14 months ago we had GPT-4 and now we have models that can get a gold medal at the IMO. But sure, if you curve fit to the last 3 months you could say things are slowing down, but that's hyper fixating on a very small amount of information.
- ddxv 10mo agoYes, that is what I'm saying, that 14 months ago the rate of change was noticeably faster. Lately the new models are much less groundbreaking and increasing in the volume of output and decreasing in cost.
- energy123 10mo agoThe private model that got gold at IMO was 4 months ago. 14 months ago we had o1-preview, we didn't have that gold medal winning approach yet. You could only say that things have slowed down since 4 months ago, but in my view that's reading the tea leaves too much. It's just not enough time and too little visibility into the private research.
- riku_iki 10mo agoit could be results of corps focusing resources on IMO in PR wars, and results is not as generalizable outside this niche.
- nickandbro 10mo ago"Create me a SVG of a PS4 controller" Gemini 3.0 Pro: https://www.svgviewer.dev/s/CxLSTx2X https://www.svgviewer.dev/s/CxLSTx2X Opus 4.5: https://www.svgviewer.dev/s/dOSPSHC5 https://www.svgviewer.dev/s/dOSPSHC5 I think Opus 4.5 did a bit better overall, but I do think eventually frontier models will eventually converge to a point where the quality will be so good it will be hard to tell the winner.
- sbinnee 10mo agoAs much as I am excited by the price, the tools they called "the advanced tool"[1] look so useful to me; Tool search, programmatic tool calling (smolagents.CodeAgent by HF), and tool use examples (in-context learning). They said that they have seen 134K tokens for tool definition alone. That is insane. I also really liked the puzzle game video. [1] https://www.anthropic.com/engineering/advanced-tool-use https://www.anthropic.com/engineering/advanced-tool-use
- andai 10mo agoThis one is different. IYKYK...
- robertwt7 10mo agothis is very impressive! as much as I love Claude though, is it just me or their limit is much lower compared to others (Gemini and GPT)? At the moment I'm subscribed to Google One AI ($20) which gives me the most value with the 2tb google drive and Cursor ($20). I've subscribed to GPT and Claude as well in the past, I find that I was hitting the limit much faster in Claude compared to all the others, it made me reluctant to subscribe again. from the blog post it seems like they've been prioritising the Max users most of the time?
- ofermend 10mo agoCan't wait to try Opus 4.5 We just evaluated it for Vectara's grounded hallucination leaderboard: it scores at 10.9% hallucination rate, better than Gemini-3, GPT-5.1-high or Grok-4. https://github.com/vectara/hallucination-leaderboard https://github.com/vectara/hallucination-leaderboard
- nickandbro 10mo agoI use the following models like so nowadays: Gemini is great, when you have gitingested the code of pypi package and want to use it as context. This comes in handy for tasks and repos outside the model's training data. 5.1 Codex I use for a narrowly defined task where I can just fire and forget it. For example, codex will troubleshoot why a websocket is not working, by running its own curl requests within cursor or exec'ing into the docker container to debug at a level that would take me much longer. Claude 4.5 Opus is a model that I feels trustworthy for heavy refactors of code bases or modularizing sections of code to become more manageable. Often it seems like the model doesn't leave any details out and the functionality is not lost or degraded.
- OhioMan2943 10mo agoSo are we in agreement that claude is the thinking persons model and openai is for the masses
- rw2 10mo agoGemini 3 in antigravity is significantly better than Claude code with either Opus or Sonnet that I struggle to see how they can compete. And I'm someone with the 100 dollar/month plan. I can't even use Opus for a day before it runs out before. This will make it better but Antigravity has way better UI and also bug solving.
- clbrmbr 10mo agoDoes anyone have a benchmark that clearly distinguishes the larger models? I would think that the high parameter count models would have capabilities distinct from the smaller ones, that would easily be read out. For example, Opus 4 has apparently memorized many books. If you ask it just right (to get around the infuriating copyright controls), it will complete a paragraph from The Wealth of Nations or Aristotle’s Nicomachean Ethics in Ancient Greek. That cannot be possible on a smaller model that needs to compress more.
- rutagandasalim 10mo agoclaude opus 4.5 is an incredible model i just one-shoted https://aithings.dev https://aithings.dev with it
- jstummbillig 10mo agoWhat was the prompt?
- Havoc 10mo agoInteresting that the number of hn comments on big model announcements seems to be dropping. I recall previous ones easily surpassing 1k Maybe models are starting to get good enough/ levelling off?
- sd9 10mo agoIt's fatigue. This is the third major model announcement in the last week. On the other hand, this is the one I'm most excited by. I wouldn't have commented at all if it wasn't for your comment. But I'm excited to start using this.
- jstummbillig 10mo agoIt's not fatigue. It's just our new normal that we have a tool that gets % better every few month. Which is fairly insane but we don't have to sweat it.
- lm28469 10mo ago"we gained 2.7% in these artificial benchmarks and here is a picture of a pelican on a bicycle, get excited and give us $7 trillion please"
- swapnilt 10mo agoOpus 4.5's scaling is impressive on benchmarks, but the usual caveats apply: benchmark saturation is real, and we're seeing diminishing returns on evals that test pattern-matching vs. genuine reasoning. The more relevant question: has anyone stress-tested this on novel problems or complex multi-step reasoning outside training data distributions? Marketing often showcases 'advanced math' and 'code generation' where the solutions exist in training data. The claim of 'reasoning improvement' needs validation on genuinely unfamiliar problem classes.
- sd9 10mo agoAfter experimenting with Gemini 3, I still felt like Sonnet 4.5 had the edge. So I'm very excited to start playing with this in the wild.
- AbstractH24 10mo agoAmazing how every company's newest model performs best in the benchmarks they share in the announcment....
- dent9 10mo agoAll the users in the comments here complaining about API limits and usage limits have missed the boat. You're not the target audience. This AI is not for you. It's not for consumers and end users. This AI is for the multi-billion and trillion-dollar businesses who are signing massive contracts to get these models enabled for their entire company. I've been using Sonnet 4.5 for months and never had a usage limit ever. And I used every model before that, all day and all night, and never once saw any mention of usage limits. Never saw a bill either. If "price per token" is a concern to you then you already lost.
- bigmadshoe 10mo agoHow could price per token not be a concern for any “multi-billion” or “multi-trillion dollar” business? Do they just burn money to remain profitable?
- ranyume 10mo agoYou'd be surprised.
- winrid 10mo agoSo far this seems like a huge downgrade from Opus 4.1. Please add back 4.1 as an option...
- johnnycombin 10mo agothe most overhyped model ever, not even close to Gemini3 or GPT5.1 after 8h of complex tasks.