5 ms·
Hello, I'm one of the engineers who worked on the incident. We have mitigated the incident as of 14:43 PT / 22:43 UTC. Sorry for the trouble.
by palcu 9mo ago
Hello, I'm one of the engineers who worked on the incident. We have mitigated the incident as of 14:43 PT / 22:43 UTC. Sorry for the trouble.
- dan_wood 9mo agoCan you divulge more on the issue? Only curious as a developer and dev op. It's all quite interesting where and how things go wrong especially with large deployments like Anthropic.
- Chance-Device 9mo agoThank you for your service.
- nickpeterson 9mo agoThe one time you desperately need to ask Claude and it isn’t working…
- dgellow 9mo agoHope you have a good rest of your weekend
- g-mork 9mo agoit's still down get back to work
- l1n 9mo agoAlso an engineer on this incident. This was a network routing misconfiguration - an overlapping route advertisement caused traffic to some of our inference backends to be blackholed. Detection took longer than we’d like (about 75 minutes from impact to identification), and some of our normal mitigation paths didn’t work as expected during the incident. The bad route has been removed and service is restored. We’re doing a full review internally with a focus on synthetic monitoring and better visibility into high-impact infrastructure changes to catch these faster in the future.
- 999900000999 9mo agoWas this a typo situation or a bad process thing ? Back when I did website QA Automation I'd manually check the website at the end of my day. Nothing extensive, just looking at the homepage for piece of mind. Once a senior engineer decided to bypass all of our QA, deploy and took down prod. Fun times.
- weird-eye-issue 9mo ago[flagged]
- spike021 9mo agoDepending on how long someone's been in the industry it's more a question of if, not when, an outage will occur due to someone deciding to push code haphazardly. At my first job one of my more senior team members would throw caution to the wind and deploy at 3pm or later on Fridays because he believed in shipping ASAP. There were a couple times that those changes caused weekend incidents.
- MobiusHorizons 9mo agoI think you meant to write “when, not if” instead of “if, not when”
- spike021 9mo ago
- giancarlostoro 9mo agoAny chance you guys could do write ups on these incidents similar to how CloudFlare does? For all the heat some people give them, I trust CloudFlare more with my websites than a lot of other companies because of their dedication to transparency.
- l1n 9mo agoWe're considering this!
- giancarlostoro 9mo agoI already love the product, and I think it would be great to see. Even if its not as "quickly" as CloudFlares (they post ASAP its insane) I would still be happy to see postmortem threads. We all learn industry wide from them.