6 ms·
Anthropic Claude and API service outages
- vikrantrathore 24d agoAnthropic has constant failure since morning.
- beredon 24d ago[dead]
- cromka 24d agoAnyone can explain why do they fail so often and so seriously? Is this inherent to managing LLMs or is their Ops really just that bad?
- vinni2 24d agoBecause they can’t keep up with the demand?
- tvbusy 24d agoStruggling with demands is a performance problem, not a reliability problem.
- swe_dima 24d agooverly high demand can often put a system in an unstable state
- therein 24d agoOnly if you're incapable of throttling to assure quality of service. Are you saying they have so much demand that even their load balancers were overloaded?
- slopinthebag 24d agoThere is no point having this discussion. Some people have had their brains genuinely broken by Anthropic and will defend them to the last breath, despite having no clue what they are talking about.
- lelanthran 24d agoThrottling helps to prevent a 30% failure rate (with a quick recovery time) turning into a 100% failure rate (with practically infinite recovery time), but some users are still going to see failures, and the news is still going to be "Anthropic is down again" because some users are seeing an outage. I'm not really sure how they can avoid getting in the news with "Anthropic is down", TBH. Although, at least by throttling they might get "Anthropic is down for some requests" and not "Anthropic is down for all requests".
- ACCount37 24d agoIt is a reliability problem - because if the API ends up rejecting 30% of the incoming queries, no one cares if it's 500 Internal Server Error or 529 Overloaded. Having your infrastructure at the knife's edge of load to capacity also means you have no redundancy when something fails.
- cherioo 24d agoIf this incident had happened during the day I could buy that argument. But it is happening in the dead of the night. Are there really that much people scheduling claude runs over-night, taking more capacity than day-time? (assume most claude users are west coast programmers) My other guess is they are scheduling training runs and their capacity isolation is bad.
- deleted 24d ago[deleted]
- Recursing 24d agoI think it's because customers don't seem to mind enough to switch. I really want to support Anthropic because of their commitments to effective altruism, but it's getting harder and harder to justify on technical merits. Sol is just faster, cheaper, and more reliable, and as good as Opus for my use cases
- therein 24d ago> commitments to effective altruism That's a downside if anything.
- eru 24d agoDifferent folks have different preferences.
- KingMob 24d agoYes, some people think effective altruism means helping people here and now as efficiently as possible. Others think the sheer number of potential sentients in the future justify any amount of suffering tolerated in the here and now. Notably, people in the latter group never assume THEY will be the ones suffering for the future.
- eru 24d ago> Notably, people in the latter group never assume THEY will be the ones suffering for the future. What makes you think so? What evidence do you have? > Yes, some people think effective altruism means helping people here and now as efficiently as possible. Most effective altruists seem pretty relaxed about the 'here' part, but you are right that the 'now' part receives different answers from different subgroups. > Others think the sheer number of potential sentients in the future justify any amount of suffering tolerated in the here and now. There's also the people who care a lot about shrimp. They are less crazy than it seems at first.
- Smaug123 24d ago
- slopinthebag 24d agoEverything is vibe coded, would you expect anything different?
- alansaber 24d agoQuick, release a new claude product then depricate it after 90 days
- ifwinterco 24d agoBecause coding is solved and there’s no need for software engineers any more. Being charitable to Anthropic though: coding is sort of solved for a very limited definition of “solved”. Software engineering of huge distributed systems at massive scale is very much not solved, and it seems Anthropic don’t have the human capital to do it particularly well
- nixon_why69 24d agoIf we want to be even more charitable to Anthropic, they're going from nothing to hyperscaler in like 3 years. I'm as happy as anyone to dunk on the general bubble situation and people claiming their technology is too dangerous to share, then something like this happens, it's hilarious. But it's still actually hard to build a service with that much usage and every org has had some outages on their way to figure it out.
- ifwinterco 24d agoYes I agree, I don't doubt they are dealing with some serious technical challenges. It's not just the enormous rate of growth, LLMs and workflows are also evolving rapidly. Whatever they planned for 1-2 years ago in terms of how they expected their stack to be used is probably already well out of date, so they'll be rearchitecting bits of a production system while it's being used
- chrisjj 24d agoAnthropic's product is inherently unreliable even when operating at its best. Why would they bother with reliable delivery when their customers are the type clearly not interested in reliability?
- vasco 24d agoClearly they just want to be down all the time. They have the ability to throttle down traffic. I have no clue why they decide to throttle it down just enough to have the worst of both worlds, throttling and incidents. Throttle more and have zero incidents, what an incredible revelation.
- Barbing 24d agoThey didn’t want to be down so bad they rented capacity from Elon and uptime improved dramatically. Super flaky before that, they were even playing games with response quality to lighten their load nontransparently. Not sure what the issue is today. But on average, past 90 days, good availability and quality right?
- gagan2020 24d agoMind you, They use latest Anthropic Claude models for programming. /s
- tvbusy 24d ago"Coding is solved", how ironic.
- chrisjj 24d agoBad coding is solved.
- champagnepapi 24d ago“Coding is solved, bugs are not yet solved. Fix incoming” This was said recently lolol https://x.com/bcherny/status/2090649326032945591 https://x.com/bcherny/status/2090649326032945591
- indrex 24d agoHow they go down globally? All regions use the same infra?
- nullsanity 24d ago[dead]
- Iolaum 24d agoIt was a configuration error ...
- singularity2001 24d agoAre they really though? The graphic seems to be highly misleading. And marks a whole day as red when there was only a partial outage for one hour.
- deleted 24d ago[deleted]
- hyperionultra 24d agoLately Dario has to many bad news from customer side. Of course there is no bad or good news, just marketing, but still…
- cloveyxD 24d ago[dead]
- rvz 24d agoOne of the Anthropic salesmen told everyone to use loops all over the place, auto mode and use Fable 5 and Opus 5 with multiple sub agents. Then it seems they forgot about their own infrastructure which goes down once every two weeks. Looks like most of all the engineering knowledge that was needed at Google was lost after the whole org at Anthropic is now vibing their work. Rather than screaming “coding is solved!”, “AGI with a month” and spooking everyone with the bogus mysticism of “Mythos” which now everyone has an equivalent strong cyber model, maybe keep the lights on first before plotting the next doomsday story.
- eru 24d ago> Looks like most of all the engineering knowledge that was needed at Google was lost after the whole org at Anthropic is now vibing their work. Sorry, what does Google have to do with Anthropic's outage?
- tcp_handshaker 24d agoThey also serve inference from GCloud
- eru 24d agoOh, ok, thanks. It seems reasonably likely that maybe Google Cloud might be to blame.
- alansaber 24d agoYes yes it is key that we vibe code critical infrastructure. I can't wait to see a claude-ism in my next BIOS update.
- tcp_handshaker 24d agoThey cant solve the infrastructure problem until Claude is back up....
- dannyw 24d agoGreat reminder to be using multiple providers (i.e. Anthropic 1P, Bedrock, Vertex AI, Azure) with auto-fallback. Bedrock has been fine throughout today; just pick two providers and you have much more stable Claude :) Also, sometimes older models work fine, Opus 4.8/4.6/4.5 are worth trying.
- wartywhoa23 24d agoNever put all AI taxes in one basket!
- bob1029 24d agoIt's fun to offer this advice but I think actually practicing it is nearly impossible in many settings. If you are a solo developer who has the patience to tinker endlessly, this advice is probably fine. If you are responsible for provisioning AI services in a team setting, this advice starts to fall apart rapidly. OAI and Anthropic might as well be oil and water when it comes to what tools and descriptions are most ideal. Swapping the inference provider like it's some interchangeable module is a total fantasy in most real world settings. > Also, sometimes older models work fine > sometimes My users are hoping for slightly more definitive results. "Usually" or even "often" would be much preferred.
- ageitgey 24d agoYou don't neccesarily have to fallback across models. Both Anthropic and AWS Bedrock provide the same models with the same API at different endpoints and different infra. So if you have a proxy endpoint to point users at, that specific fallback case is pretty easy. I'm not saying that works in every business or billing scenario, though.
- renezander030 24d agoFailover saves the call, not the run. If the provider fails mid-agent run, the second provider starts without knowing which tool calls have already been written, and the retry on a non-idempotent write simply creates the record twice. Without an idempotency key for each tool call and a checkpoint after each write, two providers make you available, but not correctly.
- roschdal 24d agoIt's a reminder to never depend of something as flaky as AI on the Internet for your important business processes.
- UnfitFootprint 24d agomy senior suggested quoting prices based off an online LLM fed a rules table and what the customer did :thinking:
- 36799753298 24d ago[flagged]
- chrisjj 24d agoLoad is not demand.
- deleted 24d ago[deleted]
- BoredomIsFun 24d agoJust buy 2x5060ti's and run local Qwen 3.8. Not Claude of course, but a good backup anyway.
- virajk_31 24d agoI really love using Claude, but running its API in production was a constant headache (they hv most horrible status graph), switched to Gemini and didn't hit any major downtime since...
- ex1fm3ta 24d agocheckout my claude-code plugin to display Claude Code status directly in the status line (below prompt textfield) https://github.com/moumine9/claude-status https://github.com/moumine9/claude-status