14 ms·
Tell HN: AWS appears to be down again
Console is flickering between "website is unavailable" and being up for my team. This is happening very frequently just now, reliability seems to have taken a hit.
- schnebbau 5y agoSo, how many execs are going to push to move to self-managed hosting in the new year? Packaging a way to migrate off AWS could be a unicorn idea.
- qwertyuiop_ 5y agoNone. Amazon hired all ex VPS, CTOs, Directors of small, medium large companies with Rolodexes.
- dehrmann 5y agoDepends on how many customers are ready to move to a different vendor. I suspect most customers are forgiving because either they were also down or half the services they use were down. You don't get fired for hosting in AWS.
- mikece 5y agoWould need one hell of a compressional algorithm to keep the data exfiltration costs down.
- pm90 5y agoPied Piper
- adamm255 5y agoAnyone using VMware Cloud services is probably laughing. Just chuck it at Azure or GCP or back on prem.
- wallacoloo 5y agoAWS has its Outpost product for on-prem hosting. not 100% self-managed, but maybe enough to satisfy the execs and make your market a bit smaller.
- Nextgrid 5y agoDoes it come with its own locally-hosted console or does it still rely on the main AWS control plane? If the latter then it could be affected too.
- dolibasija 5y agoOne of our EC2 instances in us-east-1c is unavailable and stuck in "stopping" state after a force stop. Interestingly enough, EC2 instances in us-east-1b don't seem to be affected. The console is throwing errors from time to time. As usual no information on AWS status page.
- chrishynes 5y agoI had the same issue with unavailable, but on an instance in us-east-1b. Finally just got the force stop to go through a minute ago and it's now running and available again.
- mike-cardwell 5y agoYour us-east-1b may be the parents us-east-1c. The letters are randomised per AWS account so that instances are spread evenly and biases to certain letters don't lead to biases to certain zones.
- chrishynes 5y agoHuh, that's interesting. Didn't know that, but makes sense.
- thrtythreeforty 5y agoIt's pretty cool. If I recall, they call it "shuffle sharding."
- ciceryadam 5y agoYou can check which availability zone is with: aws ec2 describe-availability-zones --region us-east-1
- throwaway984393 5y agoI'm not sure if we should say "AWS is down" if only us-east-1 is down. That region is more unstable than Marjorie Taylor Greene on a one-legged stool.
- lukeqsee 5y agoI can't get to the console either, receiving a "Temporarily unavailable" notice without branding.
- pawelduda 5y agoBitbucket is affected, pages randomly take forever to load or return 500
- Pandabob 5y agoYep, just botched a merge likely because of this.
- el_duderino 5y agoBitbucket just completed their migration to AWS too. Rough start.
- sprite 5y agoMy Elastic Beanstalk instances are completely unreachable. Seems at the very least ELB is down. Looking @ down detector it looks like this is taking a bunch of sites down with it. As usual AWS status page shows all green.
- jakub_g 5y agoWhere are you located? "X is down" without location is only moderately useful. I'm having issues with Slack from central EU (Poland) -- can't upload images, or send emoji reactions to post; curiously, text works fine). Wondering if linked
- riknox 5y agoAWS Console runs in us-east-1 so that points to at least that region having issues IIRC. I am also having Slack issues in EU.
- hdjjhhvvhga 5y agoYou should complain to Slack then. It's their problem to choose a reliable provider, and AWS seems to have trouble with keeping this status.
- RobertKerans 5y agoAssuming crates.io is AWS-backed? Getting fun situation where direct dependencies of an application are downloading but then the sub-dependencies aren't.
- mwcampbell 5y agoYeah, and I can't publish a crate.
- lukeqsee 5y agocrates.io is directly hosted on GitHub, but I'm sure some dependencies use S3 or other AWS services for things.
- RobertKerans 5y agoYep, S3 possibly the villain here
- RobertKerans 5y agoAh, back to normal now. Getting intermittent flickers on some of our apps but all seems solid-ish again
- withinboredom 5y agoI wonder if there's an s3 compatible service with similar pricing that can be used as a fallback? Are digital ocean s3 compatible storage accounts's backed by real s3?
- Ancapistani 5y agoWould Wasabi.com meet your requirements? I’m not affiliated with them, and haven’t even really used them other than to explore a bit. They come highly recommended by my acquaintances, though.
- RobertKerans 5y agoafaik there's nothing tying it specifically to GH (where the metatada is), and then the actual code is just in an S3 bucket, so in theory should be reasonably easy [ha!] to just host anywhere. In theory, I mean that's a massive lump of stuff, and surely wherever it gets hosted is going to face exactly the same issues (though if it does become very widely used, then you'd think every major provider that controls infra could easily have a mirror)
- darkwater 5y agoFields of green here https://status.aws.amazon.com/ https://status.aws.amazon.com/ Anyway I can access the web console with no issue (eu-west)
- lordnacho 5y agoThe elite DevOps teams are always assigned to the status page
- temp0826 5y agoChanges to this page require very high level management approvals (source: used to work at aws)
- hnarn 5y agoI think it's pretty widely accepted that AWS' own status pages are utterly useless.
- darkwater 5y agoYeah, it was just to confirm that this time was no different :)
- hdjjhhvvhga 5y agoIn Russia they have a specific name for it: https://en.wikipedia.org/wiki/Potemkin_village https://en.wikipedia.org/wiki/Potemkin_village
- s_dev 5y agoYou would think that but there always a few contrarian AWS evangelists in the comments going on about the "difficulty" in operating a status page as though it were trying to conjure a N=NP proof. Like how come down detector can do a superb job of detecting when AWS goes down and AWS can't? Because AWS doesn't want account managers of SLAs asking for credits for the uptime they're paying for but not getting. https://downdetector.co.uk/status/aws-amazon-web-services/ https://downdetector.co.uk/status/aws-amazon-web-services/
- captn3m0 5y ago4:35 AM PST We are investigating increased EC2 launched failures and networking connectivity issues for some instances in a single Availability Zone (USE1-AZ4) in the US-EAST-1 Region. Other Availability Zones within the US-EAST-1 Region are not affected by this issue. via https://stop.lying.cloud/ https://stop.lying.cloud/
- snth 5y agoWhat is this website? Is there an "about" or something? What is it doing differently from the official AWS status page?
- junon 5y agoCan anyone explain the affiliation of stop.lying.cloud to Amazon? All of the legalese in the header/footer seem to indicate it's actually owned and run by Amazon. If so... why? Why not just... use the real status page? I mean I'm glad it exists, don't get me wrong. Just weird that they'd have two status pages, one seemingly existing only to sort of 'mock' themselves...
- deadbunny 5y agoFWIW `lying.cloud` is registered with Namecheap. `amazon.com`/`aws.com`/`amazon.ca` are all registered with Mark Monitor. And I know that AWS uses ghandi behind the scenes for domain reg. Given that, I'd hazard a guess that it's not owned by Amazon. Definitely not a guarantee though.
- bithavoc 5y agoI think it was built[0] by @quinnypig [0] https://twitter.com/quinnypig/status/1468331194471178241?s=21 https://twitter.com/quinnypig/status/1468331194471178241?s=2...
- andyjih_ 5y agoIt's not official. The people making the page probably just copied everything, including the legalese.
- taspeotis 5y ago
- sswaner 5y agoNot down as of 7:40 EST. US-EAST-1 hosted site (athene.com). Cognito, API Gateway, Lambda, S3, DynamoDB, RDS, S3, Cloudfront.
- omosubi 5y agoI do wonder if the great resignation has anything to do with this. My team (no affiliation with Amazon) was cut in half from last year and we are struggling to keep up with all the work
- clavicat 5y agoHow much more frequent do these outages need to become before it starts triggering SLA limits?
- IceWreck 5y agoHonestly my server at home has more uptime than US-East-1
- BossingAround 5y agoDoes your server at home handle similar traffic to that of US-East-1 since you're comparing uptime? Simiarly, my laptop, if I keep it plugged in the wall, and enable httpd on localhost, will surely have better uptime than any of the top clouds. I'd bet that it'd have 100% uptime if I plugged in a UPS and cared for traffic on my local network only.
- IceWreck 5y agoNo but I access my home-server remotely from my university all the time and it hasn't gone down once. Better uptime than paying for EC2 on AWS US-East-1. Obviously this approach isn't scalable but it serves me well.
- amelius 5y ago> Obviously this approach isn't scalable but it serves me well. It's perfectly scalable. Just give everybody their own home server.
- Sammi 5y ago> Does your server at home handle similar traffic to that of US-East-1 since you're comparing uptime? Of course it doesn't. Why are you asking antagonistic questions?
- streamofdigits 5y agoSomebody call the IT department
- 300bps 5y agoCan we please stop saying, “AWS is down”? AWS consists of over 200 services offered in 86 availability zones in 26 regions each with their own availability. If one service in one availability zone being impaired equals a post about “AWS is down” we might as well auto-post that every day.
- omh2 5y agoAWS doesn't follow their own advice about hosting multi-regional so every time us-east-1 has significant issues pretty much every AZ and region is affected. Specifically large parts of the management API, and IAM service are seemingly centrally hosted in us-east-1. If your infrastructure is static you'll largely avoid the fallout, but if you rely on API calls or dynamically created resources you can get caught in the blast regardless of region
- sawmurai 5y agoIt's like my grandma saying "Honey, the internet is broken again." xD
- KptMarchewa 5y agoWould be cool if this wasn't the region where AWS hosts their internals, making other regions unusable, right?
- satya71 5y agoSeems enough services in us-east-1 are down to cause most apps to fail. My simple app uses 10s of AWS services, at least some of which are out.
- 300bps 5y agoI may have seen more of these posts than you. The last one I saw where “AWS is down” was us-west-1.
- mule1 5y agoFeel for devops peeps who are just trying to chill for Christmas
- izietto 5y agoI guess that's why I'm experiencing weird issues with Heroku: remote: Compressing source files... done. remote: Building source: remote: remote: ! Heroku Git error, please try again shortly. remote: ! See http://status.heroku.com for current Heroku platform status. remote: ! If the problem persists, please open a ticket remote: ! on https://help.heroku.com/tickets/new
- dijit 5y agoYes. Another thread: https://news.ycombinator.com/item?id=29648325 https://news.ycombinator.com/item?id=29648325
- sascha_sl 5y agoquay.io is also dead, as well as giphy, some parts of slack just the weekly internet apocalypse, happy holdidays fellow SREs
- anshumankmr 5y agoIf AWS, GCP and Azure go down, we will be back in the stone ages, right?
- dijit 5y agoThe only stuff that will work will probably depend on things in AWS in some form. That, or people never took the “if AWS goes down then lots of people will have a problem, so we’ll be fine” line seriously; there are few such cases.
- deleted 5y ago[deleted]
- rsp1984 5y agoBitbucket having issues too: https://bitbucket.status.atlassian.com/ https://bitbucket.status.atlassian.com/
- debarshri 5y agoHubspot seems to be down too [0]. [0] https://status.hubspot.com/ https://status.hubspot.com/
- Demcox 5y agoImgur is suffering from this too, I think.
- temptemptemp111 5y ago
- ItsBob 5y agoI've built out many 42U racks in DC's in my time and there were a couple of rules that we never skipped: 1. Dual power in each server/device - One PSU was powered by one outlet, the other PSU by a different one with a different source meaning that we can lose a single power supply/circuit and nothing happens 2. Dual network (at minimum) - For the same reasons as above since the switches didn't always have dual power in them. I've only had a DC fail once when the engineer was performing work on the power circuitry for the DC and thought he was taking down one, but was in fact the wrong one and took both power circuits down at the same time. However, a power cut (in the traditional sense where the supplier has a failure so nothing comes in over the wire) should have literally zero effect! What am I missing? I've never worked anywhere with Amazon's budget so why are they not handling this? Is it more than just the imcoming supply being down?
- lordnacho 5y agoWhat about a UPS/battery thingy? That's saved me a few times, though it normally just gives enough time for a short outage. Is it uncommon in cloud infra?
- vel0city 5y agoFor even regular datacenters they'll often have UPS systems the size of a small car, usually several of these, to power the entire datacenters for a few minutes to get the diesel generator started.
- uluyol 5y agoWhy spend the cost on dual X and Y when you can failover to another cluster? For big DC workloads, it is usually, though not always, better to take the higher failure rate than add redundancy.
- ItsBob 5y agoReally? You'd think at Amazon's scale an additional PSU in a 1U custom-built server (I assume they're custom) would be a few tens of $ at most. Actually, now that I type that it makes sense. Scaling a few tens of dollars to a bajillion servers on the off-chance that you get an inbound power failure (quite rare I'd reckon) might cost more than what they'd lose if it does actually fail. So yeah, they're potentially just balancing the risk here and minimising cost on the hardware. Edit: changed grammar a bit.
- camdenreslink 5y agoWho needs chaos monkey? Just host on AWS for a similar effect.
- exabrial 5y agoStat That.
- exabrial 5y agoAs an industry, can we please stop making products like vacuums that can't operate unless someone else's computer is working in a field in Virgina? There's literally no reason for it.
- ChrisMarshallNY 5y agoI can't play Borderlands 3 this morning (Epic). Wonder if it's connected?
- whoomp12342 5y agothe cloud is great they said...
- amai 5y agoA problem with log4j/logshell?
- antihero 5y agoI wonder how many 9s AWS is going for. Can't be a lot of 9s anymore.
- iso1631 5y agoAhh, the cloud https://imgflip.com/i/5yrt24 https://imgflip.com/i/5yrt24
- sprite 5y agoMy app running on AWS is currently down. Having intermittent problems with console as well.
- sh4un 5y agoDamn you all eggs in one basket.
- potas 5y agoSlack seems to have some issues because of that - I'm not sure if anyone is receiving messages, as it became completely silent for the last 15 minutes or so.
- oneeyedpigeon 5y agoNew messages seem to be ok for me, but editing old ones and uploading images both seem to be broken right now.
- jenoer 5y agoSending and receiving messages works here, but editing them does not, it throws an error. Statuses such as "calling" also do not seem to be updated any longer. Edit: Restarting Slack does update the edited messages. Edit 15:24 CET: Slack is back up.
- jakub_g 5y agoSame: only normal text seems kinda working - edits failing or working with big lag; - "Threads" view slow; - can't emoji-react; - can't upload images; - people also say they can't join new channels.
- deleted 5y ago[deleted]
- jakub_g 5y agohttps://status.slack.com/2021-12/a17eae991fdc437d https://status.slack.com/2021-12/a17eae991fdc437d > We are experiencing issues with file uploads, message editing, and other services. We're currently investigating the issue and will provide a status update once we have more information. > Dec 22, 1:58 PM GMT+1
- Pandabob 5y agoUploading images doesn't work for me.
- aden1ne 5y ago
- hnarn 5y agoIs there a history of AWS downtimes available somewhere? This makes what, three times in as many months? edit: The question isn't necessarily AWS specific, just any data on amount of downtime per cloud provider on a timeline would be nice.
- LuciusVerus 5y agoI'd say three times in as many weeks, give it or take
- MatteoFrigo 5y agoI don't know about AWS, but both Google Cloud and Oracle Cloud maintain at least a high level history of past outages. See https://status.cloud.google.com/summary https://status.cloud.google.com/summary and https://ocistatus.oraclecloud.com/history https://ocistatus.oraclecloud.com/history
- dijit 5y agoGiven the hilariously awful reputation of the AWS status page I would hazard a guess that such a page would also be incredibly inaccurate. If you can’t even admit you’re having an issue how can you keep an accurate record?
- cassianoleal 5y agoSimilar with GCP. We had a pretty bad outage once where the status page was showing all green. Google informed us that because the actual issue was further down the stack and didn't trigger any internal SLOs the status didn't get an update. It took them hours to acknowledge and fix it.
- dijit 5y agoAssuming you have a support contract the rep should send out a post-mortem page. This is what happens when we've been affected by outages (even without involving support).
- allocate 5y agoAlso running a big production app in east-1 and we're experiencing issues.
- sprite 5y agoI'm also in east-1 and completely down.
- fipar 5y agohttps://downdetector.com/status/aws-amazon-web-services/ https://downdetector.com/status/aws-amazon-web-services/
- throwaway81523 5y agoOk, enough AWS outages to say I'm tired of hearing about low end stuff being flaky.
- henriquez 5y agoHeroku isn’t “low end,” it’s a PaaS built on top of AWS. So you’re really just hearing about another AWS outage lol
- christophilus 5y agoThey're not saying Heroku is low end. They're saying, "I'm tired of hearing that it's irresponsible to run your own servers." At least, that's what I understood.
- deleted 5y ago[deleted]
- ryanbrunner 5y agoAny place I've worked at that managed their own servers (to be fair, the last time I worked at a place like that was 2010) definitely had more protracted downtimes than AWS - it just felt not as bad because we were in control of the situation, but at the end of the day that didn't get us up any faster. Another side benefit of being with AWS is when you do have an outage, a lot of other people have outages, and so you sort of blend in with the noise. It's not great to be down, but if you're down and also "big service X" who's also an AWS customer is down, it makes your downtime look less like a lack of competence and more like an unavoidable force of nature.
- dijit 5y agoI guess it's extremely dependent on an org to org basis. I worked at a company that's bread and butter was online services (e-commerce SaaS platform, similar to Netsuite) and we had significantly fewer outages than AWS had. But we had redundancies built in to most things, I'm not saying it was perfect but it worked. The major difference might be that almost nobody is willing to spend 20% of what they spend on AWS/GCP to have a self-hosted solution. The reason "cloud is so expensive" is because they're essentially telling you what the price will be and even if they only spend 40% of that on actual hardware and operations: it's more than most companies would invest in themselves. This is absurd, of course, but it's absolutely true.
- sreitshamer 5y agoConsole is sluggish for me, but S3 (us-east-1) seems to work fine.
- loudtieblahblah 5y agoYay! Adult snowday!
- RobertKerans 5y agoApropos of nothing, but a few Christmasses ago the place I worked had a dedicated fibre line that some workmen doing gas line repairs sawed straight through, took out everything; I was just drone worker at the time & it was a beautiful thing
- throwaway875487 5y agoOur RDS instances have completely packed up. Hell knows what's going on. Here come the customer support tickets.
- bognition 5y agoWhat a way to start my day
- biznickman 5y agoWhy isn't Heroku showing a status error despite being offline?
- mikece 5y agoBecause it's built on AWS and uses the AWS status page for it's status info?
- vegai_ 5y ago5ish years ago it was common knowledge that us-east-1 is generally the worst place to put anything that needs to be reliable. I guess this is still true?
- thow-58d4e8b 5y agoUnfortunately, the fact that us-east-1 is roughly 10% cheaper than other regions usually overrides any other concerns
- beermonster 5y agous-east-1 seems to be AWS’s not so well kept little dark secret! In all seriousness though - even non-regional AWS services seem to have ties to us-east-1 as evidenced by the recent outages. So you might be impacted even if it looks like (on paper at least) you’re not using any services tied to that region.
- taf2 5y agoI don't know about that. It was more like common knowledge that one availability zone in us-east-1 was a problem - you would have to figure out which one it was usually by spinning up instances in all 4 zones (now 6)... and that it was the largest of all regions making it ideal place to put your service if you wanted to be close to other vendors/partners in AWS...
- rswail 5y agoSo why are people not migrating out of us-east-1? Operating in ap-southeast, we weren't that affected by the us-east-1 down time, although our system is reasonably static and doesn't make lots of IAM calls (which seems to be a large SPOF from us-east-1).
- taf2 5y agolatency. us-east-1 is positioned very nicely relative to many large businesses in North America and Europe. This gives you pretty good access to a very large percentage of the economies of the world with good latency... while not requiring you to architect your application around multiple regions...
- dijit 5y agoSome “global” systems run in us-east1 even if you’re not hosted there a service you depend on might be. Notably: cognito, r53 and the default web UI. (You can work around the webui one I’m told, by passing a different domain instead of just console.aws.amazon.com)
- watermelon0 5y agoDon't forget about CloudFront, which can only be configured via us-east-1.
- andyjih_ 5y agoThe most hilarious irony of not being able to acknowledge a 4AM page in the PagerDuty mobile app because AWS is down.
- exikyut 5y ago(Which was about AWS being down?)
- deleted 5y ago[deleted]
- stunt 5y agoIt seems that it's due to powerloss. [05:01 AM PST] We can confirm a loss of power within a single data center within a single Availability Zone (USE1-AZ4) in the US-EAST-1 Region. This is affecting availability and connectivity to EC2 instances that are part of the affected data center within the affected Availability Zone. We are also experiencing elevated RunInstance API error rates for launches within the affected Availability Zone. Connectivity and power to other data centers within the affected Availability Zone, or other Availability Zones within the US-EAST-1 Region are not affected by this issue, but we would recommend failing away from the affected Availability Zone (USE1-AZ4) if you are able to do so. We continue to work to address the issue and restore power within the affected data center.
- aledalgrande 5y agoIf you haven't seen yet, news is it was a power loss: > 5:01 AM PST We can confirm a loss of power within a single data center within a single Availability Zone (USE1-AZ4) in the US-EAST-1 Region. This is affecting availability and connectivity to EC2 instances that are part of the affected data center within the affected Availability Zone. We are also experiencing elevated RunInstance API error rates for launches within the affected Availability Zone. Connectivity and power to other data centers within the affected Availability Zone, or other Availability Zones within the US-EAST-1 Region are not affected by this issue, but we would recommend failing away from the affected Availability Zone (USE1-AZ4) if you are able to do so. We continue to work to address the issue and restore power within the affected data center.
- notyourday 5y agoWe are being told that the are still issues in the USE1-AZ4 and some of the instances are stuck in the wrong state as of 16:15 PM EST. There's no ET for resolution.
- codeduck 5y agoanother example of a single dc in a single AZ rendering an entire region almost unusable. This has shades of eu-central-1 all over again.
- nightpool 5y agoAmazon is claiming the failure is limited to a single AZ. Are you seeing failures for instances outside of that AZ? If not, how has this rendered "the entire region almost unusable"?
- londons_explore 5y agoA lot of people will automatically fail over jobs to other AZ's. That often involves spinning up lots more EC2 instances and moving PB's of data. The end result is all capacity on other AZ's gets used up, and networks get full to capacity, and even if those other zones are technically working, practically they aren't really usable.
- networkisfine 5y agoIsn't the point of the design of an availability zone having multiple data centers so that if a single data center in the availability zone fails, services aren't affected?
- temptemptemp111 5y ago
- RONROC 5y agoThe prevailing wisdom throughout the last couple of years was: “ditch your on-prem infrastructure and migrate to a major cloud provider” And its starting to seem like it could be something like: “ditch your on-prem infrastructure and spin up your own managed cloud” This is probably untenable for larger orgs where convenience gets the blank check treatment, but for smaller operations that can’t realize that value at scale and are spooked by these outages, what are the alternatives?
- Victerius 5y agoI'm tempted to found a startup to help businesses migrate from cloud providers to on-prem infrastructure.
- datavirtue 5y agoSlinging some of that sweet Tanzu or Ranger?
- f6v 5y agoSelf-managed infrastructure doesn’t fail now?
- iso1631 5y agoNot at this rate. I remember we had a power outage in 2006, it actually took one of my services off air. Since then of course that has been rectified, and the loss of a building wouldn't impact on any of the critical, essential or important services I provide.
- ctvo 5y ago> Not at this rate. And what rate is this? It gets attention because it impacts more people, but AWS / GCP / Azure uptime is still better than what I've seen for small / mid size businesses trying to manage their own infrastructure.
- 5y ago
- CaptRon 5y agoAt least HN works.
- sctgrhm 5y agoInvision image uploads are down too because of this : https://status.invisionapp.com/ https://status.invisionapp.com/
- bobviolier 5y agoSeems unlogical that this is just a single region in a single US region We are having issues pulling images from public.ecr.aws from an EU region.
- saxonww 5y agoI don't know what's still true, but at one point us-east-1 seemed more critical than other regions because there were some things that had to be there. One thing that comes to mind is ACM certificates used with things like API Gateway (probably Cloudfront), they had to be in us-east-1 no matter where the rest of your infrastructure was. So it's not shocking to me that something going down in us-east-1 could have impact on other regions.
- reactive55 5y agoBitbucket is down as well
- reactive55 5y agoBitbucket is down as well because of this. https://bitbucket.status.atlassian.com/incidents/r8kyb5w606g5 https://bitbucket.status.atlassian.com/incidents/r8kyb5w606g...
- Hippocrates 5y agoEvery time a major cloud provider has an outage, Infra people and execs cry foul and say we need to move to <the other one>. But does anyone really have an objective measure of how clouds stack up reliability-wise? I doubt it, since outages and their effects are nuanced. The other move is that they want to go multi-cloud... But I’ve been involved in enough multi-cloud initiatives to know how much time and effort those soak up, not to mention the overhead costs of maintaining two sets of infra sub-optimally. I would say that for most businesses, these costs far exceed that occasional six-hour-long outage.
- sdevonoes 5y agoPerhaps is us, the customers (and our customers, and the customers of our customers, ...), the ones who should get used to the status of "things can go wrong"? Except for some specific scenarios (medical-related stuff, for instance), if my favourite online shopping place is down, well, it's down, I'll buy later.
- indigomm 5y ago> I doubt it, since outages and their effects are nuanced. Your point here deserves highlighting. A failure such as a zone failing is nowadays a relatively simple problem to have. But cloud services do have bugs, internal limits or partial failures that are much more complex. They often require support assistance, which is where the expertise of their staff comes into play. Having a single provider that you know well and trust is better than having multiple providers where you need to keep track of disparate issues.
- mongrelion 5y agoI agree with you. I think that having multi-AZ is the first thing to figure out before wanting to do multi-cloud, which is just another buzzword taken out of management's bullshit bucket :)
- jtc331 5y agoI’ve seen at least half a dozen full region AWS issues in the past 8 months. You really need multi-region and also not be relying on any AWS service that’s located only in us-east-1 (including everything from creating new S3 buckets to IAM’s STS).
- devoutsalsa 5y agoWe'll never really know the answer, but I have to wonder what percentage of comments on this thread are from Amazon downplaying the severity & other cloud providers hyping it up.
- mongrelion 5y agoYou give HN too much credit.
- sydthrowaway 5y agoSwitch to Azure
- anonu 5y agoBetter polish off your BCP docs. People will be asking for them quite a bit more in the new year.
- gtsop 5y agoQuestion to the sysadmins here: Is it really that outrageous of amazon to have such issues or are people way to spoiled to appreciate the effort that goes into maintaining such a service? Edit: Not supporting amazon, i generally dislike the company. I just don't understand the extend to which the criticism is justified
- dsr_ 5y agoThe issue is in three parts: 1. Did AMZN build an appropriate architecture? 2. Did AMZN properly represent that architecture in both documentation and sales efforts? 3. What the heck is going on with AMZN? Let's say that they build an environment in which power is not fully redundant and tested at the rack level, but is fully redundant and tested across multiple availability zones. Did they then issue statements of reliability to their prospective and existing customers saying that a single availability zone does not have redundant power, and customers must duplicate functionality in at least 2 AZs to survive a SPOF?
- quantumfissure 5y agoMe: Hesitation at last job moving absolutely everything (including backups) to AWS because if it goes down it's a problem I'm a firm believer in some kind of physical/easily accessible backup. Coworkers: "You're an f'n idiot. Amazon and Facebook don't go down, you're holding us back!" <-Quite literally their words. Me: leaves cause that treatment was the final straw Amazon and Facebook both go down within a month of each other, and supposedly they needed backups Them: shocked pikachu face
- rafale 5y agoDid u file a complaint on the use of swear words?
- dookahku 5y agoSend Your former colleagues a group email asking how it is
- lmilcin 5y agoThink about it this way: 1) Can you make your on prem infrastructure go down less than Amazon's? 2) Is it worth it? In my experience most people grossly underestimate how expensive it is to create reliable infrastructure and at the same time overestimate how important it is for their services to run uninterrupted. -- EDIT: I am not arguing you shouldn't build your more reliable infrastructure. AWS is just a point on a spectrum of possible compromises between cost and reliability. It might not be right for you. If it is too expensive -- go for cheaper options with less reliability. If it is too unreliable -- go build your own yourself, but make sure you are not making huge mistake because you may not understand what it actually costs to build to AWSs level. For example, personally, not having to focus on infra reliability makes it possible for me to focus on other things that are more important to my company. Do I care about outages? Of course I do, but I understand doing this better than AWS has would cost me huge amount of focus on something that is not core goal of what we are doing. I would rather spend that time thinking how to hire/retain better people and how to make my product better. And adding all that complexity of running this infra to my company would cause entire organisation be less flexible, which is also a cost. So you can't look at cost of running the infra like a bill of materials for parts and services. And if there is an outage it is good to know there is huge organisation there trying to fix it while my small organisation can focus preparing for what to do when it comes back up.
- ClumsyPilot 5y agoNow that everyone and their dog is on AWS, it is not just 'a website stops working', half the world, from telephones to security doors and Iot equipment, stops working? I am not sure if the movement the cloud has reduced amount of failures, but it definitely has made these failures more catastrophic. Our profession is busy makin the world less reliable and more fragile, we will have our reconning just like the shipping industry did.
- dehrmann 5y agoIt's more like it's making downtimes correlated rather than random. For everything other than urgent communication, I'm not sure if this is a big deal.
- madeofpalk 5y agoall I've noticed is slack was a bit unreliable for a little bit, but i just carried on and otherwise ignored it. my world did not stop working.
- ClumsyPilot 5y agoMy apartment block has a dialing system, that, instead if using a cale that goes to your apartment, relies on IP telephony and calls your mobile phone. It stos working if there is no internet, or your phone is out of battery, or you are not home but your wife is.
- deleted 5y ago[deleted]
- KronisLV 5y agoSame, maybe that was a related issue. Today, on Slack i could not edit messages, could not edit statuses and could not post attachments. Pretty annoying!
- kingsloi 5y agoOf all the AWS outage, my team and I have dodged them all, except this one. 3 instances down and unavailable > Due to this degradation your instance could already be unreachable >:(
- electroly 5y agoFWIW I don't think that message has anything to do with this outage. I think it's just a coincidence that you got some degraded hosts. They didn't send out emails like that for this AZ outage (nor would I expect them to -- that email is for when host machines die).
- exogenousdata 5y agoLooks like the SEC's Edgar website is affected. This is the site the SEC uses to post the filings of public companies. Normally there are a hundred or more company filings in the morning starting at 6am ET. This morning there are two. https://www.sec.gov/cgi-bin/browse-edgar?action=getcurrent https://www.sec.gov/cgi-bin/browse-edgar?action=getcurrent
- JCM9 5y agoAWS didn’t “go down”. They had an outage in one AZ, which is why there are multiple AZs in each region. If your app went down then you should be blaming your developers on this one, not AWS. Those having issues are discovering gaps in their HA designs. Obviously it’s not good for an AZ to go down but it does happen and why any production workload should be architected to have seamless failover and recover to other AZs, typically by just dropping nodes in the down AZ. People commenting that servers shouldn’t go down ect don’t understand how true HA architectures work. You should expect and build for stuff to fail like this. Otherwise it’s like complaining that you lost data because a disk failed. Disks fail… build architecture where that won’t take you down.
- matharmin 5y agoAWS is under-reporting the severity of the issue though. The primary outage may be in a single AZ, but there are parts of the AWS stack that affected all AZs in us-east-1, and potentially other regions as well. For example, even now I'm unable to create a new ElastiCache cluster in different AZs of us-east-1.
- zymhan 5y ago> I'm unable to create a new ElastiCache cluster in different AZs of us-east-1 Isn't that because Elasticache will distribute the cluster across AZs automatically? https://docs.aws.amazon.com/AmazonElastiCache/latest/red-ug/AutoFailover.html https://docs.aws.amazon.com/AmazonElastiCache/latest/red-ug/...
- matharmin 5y agoIn this case, this was specifically with a single-AZ setup, using an AZ that was supposed to be unaffected.
- dkryptr 5y ago100% agree. I'm actually surprised AWS hasn't built in a Chaos Monkey into their APIs/console so people can test their resiliency regularly if an AZ goes down. edit: of course, AWS does have this: AWS Fault Injection Simulator
- bob1029 5y ago2 of our servers are fucked right now. VOIP services down. Only with AWS and Github do I seem get panicked text messages on my phone first thing in the morning... Our workloads on Azure typically only have faults when everyone is in bed.
- 13daug 5y agoThis S3 how you gonna get you investment back from it
- pkulak 5y agoI used to think it was silly to have your own hardware (like a NAS) in your house. What makes you think you can do it better than AWS? Santa is bringing me a Synology in three days.
- darkstar999 5y agoWhy not both? I just got a Synology NAS and it makes cloud sync dead simple. Now the most important things are on my PC, mirrored on 2 drives in my NAS, and on AWS S3 (or any other cloud storage).
- pkulak 5y agoOh yeah. My plan is to migrate everything to the NAS, then have that back up to Glacier and/or Rsync.net. By S3, do you mean Glacier?
- darkstar999 5y agoI have some in glacier, some in Infrequent Access.
- jorgeudajer 5y ago
- richardfey 5y agoAs far as I understood a whole availability zone went down; today is also the day a lot of people understand why "multi-AZ" matters, so I don't think it's fair to say that services are down because the whole AWS is down.
- joelbondurant 5y ago
- l0b0 5y agoMeta: I posted a "PyPI is down" link a few days ago, and the post got insta-flagged. Is there some rule about this sort of thing?
- amai 5y agoThank goodness we host all IT services in the same cloud. Imagine the chaos we had if everything would not fail at the same time.
- kemals 5y agoHere is The Internet Report episode on the topic of recent AWS outages that covers outage and root causes: https://youtu.be/N68pQy8r1DI https://youtu.be/N68pQy8r1DI
- joe_chip 5y ago
- tomerbd 5y agoRumble was up all this time.
- j10c 5y agoI also had problem with loading youtube at the same time(for 10-15 minutes) . It looks like a coincidence, but who knows if google uses some of the infrastructure from aws.