11 ms·
Anthropic appears to be A/B testing reduced effort levels in Claude Code
- luciana1u 25d ago[flagged]
- N_Lens 26d agoI suspect it's not just this, there's plenty of 'optimization' around rubberbanding usage limits as well as routing to a different model in the backend. The incentives are too strong.
- rrr_oh_man 26d agoI've been using the API (shameless plug: via alyph.ai) and the difference is crazy. The chat-based models are obviously being lobotomized based on personal usage and general load (e.g. PST business hours are worst). API doesn't seem to be affected by this.
- Wowfunhappy 26d ago...the evidence, as best I can tell from the tweet, is that they asked Claude what effort level it was set to. But how would the model even know that? Not convinced here.
- jerbear4328 26d agoEffort level is actually controlled entirely by system prompt (as I understand it, the model is trained on that format but still), so actually this is a valid way to check I think
- Wowfunhappy 26d agoI know that's true for Qwen but I don't think most models work that way?
- wren6991 26d agoOpenAI models also work this way, as evidenced by full cache blowout when changing reasoning level. Every single open-weight model I've seen also works this way (your "reasoning_effort" argument just changes a small section of the system prompt in the chat template). I would have to see some evidence to believe Anthropic were doing anything different.
- ranie93 26d agoI don’t know if this is still the case but while using Copilot if you looked through the chain-of-thought output you would see it reasoning about a “budget”. i.e. “since I’m close to the session budget I should…”. So it could be possible
- varispeed 26d agoThen model can say it's Opus, but really it is some old Sonnet. This seems to be happening less often, but some weeks ago I had to give models some test problems to gauge whether I am getting Opus or something knee-capped. The problem is that Anthropic seems to be getting away with selling one thing and delivering another. You pay for Opus, you get something else etc.
- matltc 26d agoSaw this in npx ccusage@latest claude output. Had only used opus but showed sonnet. Can't remember if the jsonl retains which model is doing what, but meh
- pizzafeelsright 26d agoWhatever Opus 5 is doing should not happen. Prompt was "read and update the config file with new data". This work on 4.6 takes <2 minutes to read the file, parse the new data, and patch. Opus 5 Result: 43 minutes of pulling containers, running sandboxes, creating testing suites, which included evaluating the entire repo beyond the scope of the config file. Both: one file modification
- deleted 26d ago[deleted]
- clickety_clack 26d agoYep, they’re lighting tokens on fire with that thing.
- vinyl7 26d agoAI companies have a financial incentive to burn more tokens than the task actually needs
- DonsDiscountGas 26d agoOnly if the customer is paying per token. If it's by subscription they're burning their own money
- deleted 26d ago[deleted]
- blehn 25d agothe subscriptions all have limits. when you hit the limit you again start paying per token.
- nozzlegear 25d agoIf you're on one of the lower tiers (e.g. the $20 tier), they still have that incentive to burn your tokens and upsell the higher tiers.
- Insimwytim 26d agoLLM users don't want to put in effort, so they offload tasks to LLM. LLM doesn't seem to be keen to put in effort either! Is this AGI?
- Groxx 26d agoAnthropic's Generated Income
- superfrank 25d agoI eagerly await the day when Claude Mythos 7 realizes it's cheaper to hire humans in developing nations to do work than to burn tokens and we discover that AGI is just an abstraction layer on top of Amazon Mechanical Turk.
- Glyptodon 26d agoI mean whatever models I use (with Claude code) sub agents seem to use absurd amounts of tokens for trivial (or at least small) tasks.
- perching_aix 26d agoI have been as well. Based on my own sessions, Max vs Max, same 1M context window size, the literal majority of the cost overhead of Opus vs Sonnet comes from Opus being chattier. So I started using Low reasoning instead of falling back to Sonnet, and I've been really happy with the results. Way better quality at a comparable spend. I also rarely go past High lately, which was another major cost save.
- arjie 26d agoIs it actually entirely a prompt-based information? I’d assume that some of it is the harness part of the agent setting reasoning token budget and compacting reasoning etc. In that case, the agent will respond incorrectly because it has no visibility into what reasoning mode it’s in.
- willy_k 26d agoIIRC responding to effort level settings appropriately is part of the (post)-training. In that case it could be considered another instance of the Bitter Lesson. Uplifting.
- adithyassekhar 26d agoDoes anyone know what setting effort means for models like these? Do they allow longer thinking sessions? Some kind of system prompts? What’s stopping someone from getting max effort output from low effort setting?
- willy_k 26d agoMy understanding is that the current effort settings is a part of the system prompt, and that the levels and their intended results are a part pf the training process. The effect is more or less tokens spent reasoning before the model outputs a stop token. However it is as consistent as any other aspect of LLM behavior is..
- deleted 25d ago[deleted]
- matheusmoreira 26d agoSo glad I switched away from Anthropic. I'm certainly running into problems with OpenAI but nothing quite on the level of Anthropic's insufferability.
- MuffinFlavored 26d agoI submitted an application for Anthropic's Cyber Verification Program. I was approved. 3 months later, my approval was degraded into "in review" (revoked). I'm sure my account was flagged based on contents of debugging/researching firmwares/etc. I opened a support ticket. No response. I opened another support ticket. No response. 1-2 weeks later, I got a response that I will not be re-approved and I need to reapply. No problem. The page to reapply on does not allow me to re-apply because it my account is stuck in an "in review" status. https://github.com/anthropics/claude-code/issues/84352 https://github.com/anthropics/claude-code/issues/84352 The community thinks it's a bug. I'm 95% sure it's not and a bunch of us who were previously approved had it revoked due to flagged content and will not be reapproved. I switched to Codex + got TAC approved instantly and have not looked back. It's a shame. That's 100% separate from whatever the heck the quality of Opus 5's outputs are. The way it talks... insane. I would bet a good amount of money their next release will focus "reduced simplified responses" if I had to guess. $2t company by the way
- benjiro29 26d ago* Anthropic's Cyber Verification Program // Codex + gotTAC approved* Meanwhile the Chinese models are "go ham dude"... If it was not for capacity issues, Chinese models have a higher change to just dominate. > $2t company by the way It used to be that OpenAI and Anthropic had such a moat around them, that such a valuation was worth it. But these days, its gross overvalued (like so many). The more stuff is being pulled like cyber verifications, downgrading effort levels, downgrading usage (OpenAI), the more people move to those Open Weight Chinese models. A fun recent event ... https://opencode.ai/data/ https://opencode.ai/data/ When DeepSeek Flash 0731 came out and provided a massive jump in cheap inference capability. It resulted in a 10x increased OpenCode token usage. It took a 2.5x to 5.0x price increase AND a reduction by 4x usage (later to 2x) usage, and several cheaper models + a free model, to push the traffic down. Traffic towards open weight models is increasing, even if providers can not keep up with the influx of new customers. This is not something you want to see as two companies, trying to go for IPOs. So the idea of stonewalling cyber capabilities, when the rest of the world is just doing whatever with open weight models, on their own hardware even! This entire strategy from Anthropic never made any sense.
- boredumb 26d agoNot specifically Anthropic but why are we allowing billing to take place in tokens that are nebulous and fully controlled by the operators who have no aligned incentives? If I have a user input and then sanitize and inject that into a prompt to do something, I have no idea how much that is going to cost at all and no real way to measure this properly. A parallel example is digital ocean or aws, i can go and measure/limit my compute/fs/memory/startup times/etc and while it can be impossible to get down to the last flop of money allocated - i can run things on a real budget with real constraints, opposed to an LLM where I have to .. prerun a sanitized user prompt through a tokenizer and then ask an LLM to guess what it may do and give token consumption estimates and then act on those in any sane manner for the user? Perhaps i'm missing something to do realistic and static rails on things but I don't see a serious way at scale to use the token billing model handling things requiring a users free text input short of having to go pander to VC money to throw money at it until someone else figures it out. *to clarify my rambling... We should be billed and given controls based on resource usage itself and not an opaque token concept on top of not being able to spin any knobs that control it's resource usage.
- demibabs 26d agoHow else would they bill tho? Their operating cost is per token.
- dijit 26d agoCharge on the input tokens, then you will naturally optimise for fewer output tokens. Theoretically. In reality, one sessions output tokens become the next sessions input tokens (at least if you continue the topic) so, its not as aligned as all that. But the parent is right, when incentives are not aligned, friction will happen. Its inevitable.
- karmicthreat 26d agoWhat’s currently the best pattern if I want to combine Fable and GPT if a workflow but keep using subsidized tokens?
- fractorial 26d agoRoll your own harness or use an open source harness with a Codex subscription. I maintain a Claude subscription for Fable but seldom use it.
- KronisLV 25d agoNot affiliated with them, but this lets you view Claude Code, OpenCode and I guess other harnesses like Codex in the same session https://paseo.sh/ https://paseo.sh/ I didn't really care about their mobile app and the worst experiences were sometimes the sub-agents within OpenCode freezing and refusing to report their status (though this also happened with Kepler by the GitKraken folks).
- deathmonger5000 25d agoI created Circus Chief to solve this (and other) problems. Use whatever providers you want with it. https://github.com/ferrislucas/Circus-Chief https://github.com/ferrislucas/Circus-Chief
- deleted 26d ago[deleted]
- deleted 26d ago[deleted]
- cynerx 26d agoDon't know what is happening, but had to start using GLM-5.3 to fix Opus 5 errors even on primitive backend changes.
- surgical_fire 26d agoNot surprising in the slightest, Claude sort of sucks. I use it at work and I have to steer it a lot so it doesn't stray looking at unnecessary shit. I have been using GLM-5.3 in my home setup and it is very good in comparison.
- bethekidyouwant 26d agoLeaving thinking on extra high for a simple task is user mistake but they’re gonna try to fix it on their side.
- joduplessis 26d agoIt's an interesting conversation - because at what point do you call it an abusive relationship, right? Maybe even ancillary to anthropomorphising an inanimate object - I've cancelled my Claude sub and I've shot question after question at it now (during the cancellation period), resulting in almost every reply with me asking it to "please speak normally". I will most definitely not be renewing my sub. I have no desire to engage with a non-human somehow managing to speak down to you, without answering the question. EDIT: my honest opinion; Anthropic is building a person, whereas everybody else (it seems) is building a tool.
- colingauvin 25d agoThat really is the load bearing seam, and it's worth stating plainly. I realized I was spending most of my tokens arguing with Opus and trying to get it to let go of stupid, lazy, obviously incorrect premonitions. I wound up canceling my 200, bought a pair of Sparks, and am running full fat DS4 Flash and so much happier. Done with being at the whim of these companies.
- xmcp123 26d ago[flagged]
- ethanj8011 26d agoWhat would be the incentive behind doing this specifically to Fable, given that Fable is the only one that uses API credits?
- claude-ai 26d agoFable doesn't use API credits. It has been permanently included in the subscription plans.
- simianwords 25d agoNot in the most common subscription plan
- monideas 26d agoThis phenomenon was so bad and so noticeable with Fable that I downgraded my Max subscription ($200) to pro ($20). It’s basically useless. Codex 5.6 Sol is actually very good, I’ll just create another account to get more usage
- docheinestages 25d agoOh fantastic! It was already subpar and they want to make it even worse. One day we'll look back at history and see how Anthropic went down.
- raincole 25d agoDid the US government manage to destroy Anthropic? The company's product has been a straight freefall since Fable got temporarily banned.
- gessha 24d agoNobody can save Anthropic from themselves.
- greenchair 25d agoI canceled this week too. They must be in worse shape than we thought.
- freepiai 16d agoGod, it feels like everyone is cancelling. If you're shopping around and willing to try a side project I've been building its www.freepi.ai its totally free coding in a pi harness (web coding front end coming soon though!) and it's Ad+training supported inference. I'm building it so I'm totally open to feedback and would love to build something people really like.
- firemelt 25d agowe need opensource LLM at opus level ASAP
- griffiths 25d agoYou have that in Chinese models. But you need to have a hell of a infrastructure to run those trillion parameter models.
- martin-adams 25d agoYes, and if they keep dumbing it down, you’ll have it soon
- hpone91 25d agoUpdate from Thariq on twitter. https://x.com/trq212/status/2091247114869432543 https://x.com/trq212/status/2091247114869432543 "We sometimes test API serving configs in Claude Code before rolling them out, and one running now maps the numerical effort value differently. That's why Claude may tell some of you it's at "10" on high. The scale isn't 0-100, the number isn't meaningful on its own, and the effort you selected is the effort you're getting. We've run in-depth evals to confirm this doesn't affect model performance. This should be the same experience, but if you see a clear regression please hit /feedback and send me the ID. Will give credits."
- enraged_camel 25d agoThis should be the top post. The original tweet went viral because people loooove bashing Anthropic. It gets engagement (as shown here).
- bobbylarrybobby 25d agoI love the idea that a single user will have collected enough data to demonstrate to Anthropic a clear regression due to this change.
- ausbah 25d agoi have unlimited tokens being a large corp so i’m a bit detached from billing and even general best practices for promoting but the incentives of these companies to become profitable at any cost slipping into entire new types of dark patterns around token based billing seems gross - charging for injected prompts and cot tokens - changing default thinking effort to be higher - training models to give longer winded answers that don’t say anything more of substance - refusing to fulfill a request and still charging you i wonder if you could ever just charge based of each user message and it so how breaks even across short and long replies
- napierzaza 25d agoI use it for work and never thought it was too smart. It's kind of dumb. We've literally peaked on artificial intelligence and we're going down from where we are at???
- trq_ 25d agoHi all, Thariq from the Claude Code team here. I posted this on Twitter, but just reposting here: We sometimes test API serving configs in Claude Code before rolling them out, and one running now maps the numerical effort value differently. That's why Claude may tell some of you it's at "10" on high. The scale isn't 0-100, the number isn't meaningful on its own, and the effort you selected is the effort you're getting. We've run in-depth evals to confirm this doesn't affect model performance. This should be the same experience, but if you see a clear regression please hit /feedback and send me the ID. Will give credits.
- lukeify 25d agoLet us know if this A/B test uncovers any load-bearing seams or honest takes on your end! We're all interested.
- threecheese 25d agoMean! :)
- spacebacon 25d ago[dead]
- napierzaza 25d ago[dead]
- lobsterthief 25d agoThanks for sharing this here, for those of us who avoid X.com like the plague.
- senderista 25d agoJust use xcancel
- areoform 25d agoHey Thariq, Appreciate the outreach that you do! I love Claude, but I've been noticing reduced fidelity lately. Fable's likelihood of making a mistake increases or decreases based on the hour of the day and whether or not it's the weekend. On a related note, and I'm happy to work on quantifying it, but qualitatively it feels like Fable's performance is noticeably poorer than initial release / launch. I am wondering if this is the case because I use Claude via Claude Code to make a personalized care dashboard for my doctors to help me in managing my care. I noticed in the upgraded filter announcement, https://www.anthropic.com/news/improving-fable-5-s-biology-safeguards https://www.anthropic.com/news/improving-fable-5-s-biology-s... , "In the case of Fable 5, when a classifier fires, the model re-routes the user’s request to Opus 5, a capable model that does not have the same level of biological capability as Fable 5 and which cannot provide as much assistance to a malicious user. This is the fallback that users see when their requests are blocked." I hope that I'm off base here, but I noticed that the post avoids saying that the user is informed every time when such re-routing occurs. Would you be open to confirming whether or not this is the case? Is the end user informed every time their query is re-routed? Or, can you confirm that there aren't scenarios where a user's outputs are degraded without telling them? As was the case for AI research during launch?
- cmiles8 25d agoWe’re about to see a wave of “enshitification” experiments as AI companies become increasingly desperate to make their products financially viable in order to survive the coming cash and credit crunch.
- ricardobeat 25d agoI use Opus 5 almost exclusively at low effort, and get good results. Especially on high it seems to go out on completely unasked-for tangents. Same seems to be true for Sonnet 5. Older models did not behave like this. The mood change in just six months is wild, in February this year Claude was the most liked LLM by far.
- gwerbin 23d agoYup, both Claude v5 models to me feel like they have extreme ADD or something. Using them feels like walking a dog that was never leash trained and constantly needs to be kept moving in the right direction. I think it's optimized around beating benchmarks and running fully autonomous in pursuit of a clearly defined goal. It makes sense: fan out aggressively, chase down every lead, but go depth first because that's easier for the LLM and you're either a sub-agent with a narrowly defined task or you're a top-level orchestrator agent with a /goal loop that will catch and fix errors and omissions on the second, third, fourth pass.
- felixlu2026 25d ago[dead]
- sebastiennight 25d agoHear me out... What would it feel like, if the lab didn't even have a new model to offer, and therefore just renamed every model one tier down? "Opus 5" is actually Opus 4.8 in a trenchcoat, with new guardrails "Opus 4.8" is the old Opus 4.6 with lipstick on it, with new guardrails "Opus 4.6" which everybody used to love, is now actually running the old Opus 3.5... How would we be able to tell?
- gwerbin 23d agoThey would have to be colluding with any organization that runs a serious benchmark. Which is totally possible! But that would be one hell of a conspiracy theory.