8 ms·
Is Northern Virginia still the least reliable AWS region?
- kankerlijer 9mo agoThere are only two kinds of cloud regions: the ones people complain about and the ones nobody uses
- Manouchehri 9mo agoYeah, I was often the single source of reporting Claude outages (or even missing support completely) on less commonly used Amazon Bedrock regions.
- htrp 9mo agoWhich regions were you using ? ( Thought claude had global inference support that routed to all regions)
- Manouchehri 9mo agoI believe I was using us-east-2. In the early days of cross-region inference, less people were using it, and there was basically no monitoring (and/or alerting) on Amazon's side. The cross-region and global inference routing is... odd at times.
- kachapopopow 9mo agoI like this a lot, this is a great comparison for hetzner american offerings since it's not big enough for them to even bother investing much into it so there's not that many complains about it. People just dumping it (me included) after discovering the amount of random issues it has probably also doesn't help. if you are using hetzner: avoid everything other than fra region, ideally pray that you are placed in the newer part of the datacenter since it has the upgraded switching spine I haven't seen the old one in a bit so they might have deprecated it entirely.
- jeltz 9mo agoHetzner does not have any "fra region". They have Helsinki, Falkstien and Nuremberg in Europe. None of them which has any issues as far as I know. They used to have some issues with the very old stuff in Falkstien.
- kachapopopow 9mo agosorry, fsn* I have them typod fra internally and keep messing it up since it's stuck in my head.
- joe_the_user 9mo agoA sound banker, alas, is not one who foresees danger and avoids it, but one who, when he is ruined, is ruined in a conventional and orthodox way along with his fellows, so that no one can really blame him. JM Keynes
- paulddraper 9mo agoThat is incredibly appropriate.
- vasco 9mo agoEu-west-1 is miles better and is huge
- david_shaw 9mo agoYes, it's the least reliable. Thanks for summarizing the data here to illustrate the issue. It's often seen as the "standard" or "default" region to use when spinning up new US-based AWS services, is the oldest AWS center, has the most interconnected systems, and likely has the highest average load. It makes sense that us-east-1 has reliability problems, but I wish Amazon was a little more upfront about some of the risks when choosing that zone.
- Forgeties79 9mo agoNobody ever got fired for connecting to us-east-1
- yibers 9mo agoAss covering-wise, you are probably better off going down with everyone else on us-east-1. The not so fun alternative: being targeted during an RCA explaining why you chose some random zone no one ever heard of.
- riffic 9mo agohow about following the well-architected framework and building something with a suitable level of 9s where you can justify your decisions during a blameless postmortem (please stamp your buzzword bingo card for a prize.)
- paradox460 9mo agoWe vibe code everything in flavor of the month node frameworks, tyvm, because elixir is too hard to hire for (or some equally inane excuse)
- DANmode 9mo agoI agree with your post conceptually. However: Don’t underestimate community support (in the areas you’re likely to want it) when comparing development stacks.
- paradox460 9mo agoConversely, a community means nothing if they flit from one "best practice" to another
- transcriptase 9mo agoI look forward to the eventual launch of a new and improved version of your app using electron. What’s the point in having 64 Gb of DDR5 and 16 cores @ 4.2 GHz if not to be able to have a couple electron apps sitting at idle yet somehow still using the equivalent computational resources of the most powerful supercomputer on earth in the mid 1990s.
- theturtle 9mo agoI searched for it, and did not find, the word "backhoe." Big fail. I have said for years, never ascribe to terrorism what can be attributed to some backhoe operator in Ashburn, Virginia. We got a lotta backhoes in northern Virginia.
- therobots927 9mo agoOf course it is, all of the NSA men in the middle add a lot of overhead that can interfere with regular operations.
- arusahni 9mo agoThe sorting for the "Duration" column appears to be lexicographical, not numeric.
- davidfstr 9mo agoI intentionally avoid using us-east-1 for anything, since I’ve seen so many outages.
- temp0826 9mo agous-east-1 is often a lynchpin for services worldwide. Something hinky happening to dns or dynamodb in us-east-1 will probably wreck your day regardless of where you set up shop.
- secondcoming 9mo agoWe get constant resource issues in GCP’s us-east4 region
- nadis 9mo agoCackling while reading this visiting my family in Northern Virginia for the holidays. Despite it being a prominent place in the history of the web, it's still the least reliable AWS region (for now).
- rayiner 9mo agoIts nice to know that where I grew up is Too Big to Fail lol.
- the__alchemist 9mo agoI don't know if this is still true, or related, but that area used to be (Circa 10-30 years ago) very highly prone to power outages. The reason was lots of old trees near the lines that would inevitably fall; blackouts in local areas were common due to this.
- Fhch6HQ 9mo agoThat's an interesting data point, but I don't think it's relevant. The datacenters themselves are designed with a high level of power reliability and can island themselves if needed. We've started to see some rather interesting consequences for grid reliability: https://blog.gridstatus.io/byte-blackouts-large-data-center-loads-new-issues-pjm/ https://blog.gridstatus.io/byte-blackouts-large-data-center-...
- noosphr 9mo agoAt 34 hours of downtime that's two nines of uptime At this point my garage is tied for reliability with us-east-1 largely because it got flooded 8 month ago.
- emersonrsantos 9mo agoGlad to use us-west-2 for reasons.
- alexjurkiewicz 9mo agoI think part of this is that Status Page updates require AWS engineers to post them. In the smaller Tokyo (ap-northeast-1) region, we've had several outages which didn't appear on the status page.
- calmbonsai 9mo agoAnswer these questions: - Is X region and its services covered by a suitable SLA? https://aws.amazon.com/legal/service-level-agreements/ https://aws.amazon.com/legal/service-level-agreements/ - Does X region have all the explicit services you need? (note things like certs and iam are "global" so often implicitly US-East-1) - What are your PoP latency requirements? - Do you have concerns about sovereign data: hosting, ingress, and egress? https://pages.awscloud.com/rs/112-TZM-766/images/AWS_Public_Sector_Day_Navigating_Digital_Sovereignty_and_Data_Protection_with_AWS.pdf https://pages.awscloud.com/rs/112-TZM-766/images/AWS_Public_...
- JojoFatsani 9mo agoYes
- bzGoRust 9mo agoThe test environment is deployed on us-east-1, whereas the production environment deployed on us-west-2 on our side.
- yearolinuxdsktp 9mo agoUs-east-1 is far far from least reliable. It’s one of the more reliable ones. Smaller regions tend to have more reliability issues affecting the entire AZ. This analysis is skewed due to the major incident in 2025. What was the data for 2024 and over the last, say, 5 years? So the proclamation of least reliable of us-east-1 is based on 1 year of data, and it’s probably fair to say that at least last 3 years if not 5 are a better predictor of reliability. us-east-1 also hosts some special things, so it will have more services to lose.
- mlhpdx 9mo agoI stopped deploying to a single region for production years ago, so I don’t really have a horse in this region comparison race. That said, I’ve seen network level issues in every region I use — nothing like the big outage, but issues that may disrupt a service. Designing for how the world is rather than how I wish it was makes a lot of sense to me.
- bob1029 9mo agoI think if you need something more reliable than us-east-1 that you should be hosting on prem in facilities you own and operate. There aren't that many businesses that truly can't handle the worst case (so far) AWS outage. Payment processing is the strongest example I can come up with that is incompatible with the SLA that a typical cloud provider can offer. Visa going down globally for even a few minutes might be worse than a small town losing its power grid for an entire week. It's a hell of a lot easier to just go down with everyone else, apologize on Twitter, and enjoy a forced snow day. Don't let it frustrate you. Stay focused on the business and customer experience. It's not ideal to be down, but there are usually much bigger problems to solve. Chasing an extra x% of uptime per year is usually not worth a multicloud/region clusterfuck. These tend to be even less resilient on average.
- jl6 9mo ago> worst case (so far) It’s kind of amazing that after nearly 20 years of “cloud”, the worst case so far still hasn’t been all that bad. Outages are the mildest type of incident. A true cloud disaster would be something like a major S3 data loss event, or a compromise of the IAM control plane. That’s what it would take for people to take multi-region/multi-cloud seriously.
- svelle 9mo ago> A true cloud disaster would be something like a major S3 data loss event So like the OVH data center fire back in 2021?
- jl6 9mo agoNo, a major one. (No shade on OVH, but they are ~1% market share player)
- dijit 9mo agoI mean, EBS went offline and people were ok to continue using AWS… https://arstechnica.com/information-technology/2011/04/amazons-lengthy-cloud-outage-shows-the-danger-of-complexity/ https://arstechnica.com/information-technology/2011/04/amazo...
- bmitch3020 9mo agoThis story missed a glaring detail. There are simply more data centers in northern VA [0]. More than the rest of the US by a wide margin, or the entire EU+Asia. Things break here because it's where most things are. [0]: https://www.datacenters.com/providers/amazon-aws/data-center-locations?view=Map https://www.datacenters.com/providers/amazon-aws/data-center...
- vivzkestrel 9mo agosort by duration on that page is broken