7 ms·
Cascading errors caused AWS to go down
- eli 14y agoIn case your browser doesn't speak RSS: Service is operating normally: Root cause for June 14 Service Event June 16, 2012 3:15 AM We would like to share some detail about the Amazon Elastic Compute Cloud (EC2) service event last night when power was lost to some EC2 instances and Amazon Elastic Block Store (EBS) volumes in a single Availability Zone in the US East Region. At approximately 8:44PM PDT, there was a cable fault in the high voltage Utility power distribution system. Two Utility substations that feed the impacted Availability Zone went offline, causing the entire Availability Zone to fail over to generator power. All EC2 instances and EBS volumes successfully transferred to back-up generator power. At 8:53PM PDT, one of the generators overheated and powered off because of a defective cooling fan. At this point, the EC2 instances and EBS volumes supported by this generator failed over to their secondary back-up power (which is provided by a completely separate power distribution circuit complete with additional generator capacity). Unfortunately, one of the breakers on this particular back-up power distribution circuit was incorrectly configured to open at too low a power threshold and opened when the load transferred to this circuit. After this circuit breaker opened at 8:57PM PDT, the affected instances and volumes were left without primary, back-up, or secondary back-up power. Those customers with affected instances or volumes that were running in multi-Availability Zone configurations avoided meaningful disruption to their applications; however, those affected who were only running in this Availability Zone, had to wait until the power was restored to be fully functional. The generator fan was fixed and the generator was restarted at 10:19PM PDT. Once power was restored, affected instances and volumes began to recover, with the majority of instances recovering by 10:50PM PDT. For EBS volumes (including boot volumes) that had inflight writes at the time of the power loss, those volumes had the potential to be in an inconsistent state. Rather than return those volumes in a potentially inconsistent state, EBS brings them back online in an impaired state where all I/O on the volume is paused. Customers can then verify the volume is consistent and resume using it. By 1:05AM PDT, over 99% of affected volumes had been returned to customers with a state 'impaired' and paused I/O to the instance. Separate from the impact to the instances and volumes, the EBS-related EC2 API calls were impaired from 8:57PM PDT until 10:40PM PDT. Specifically, during this time period, mutable EBS calls (e.g. create, delete) were failing. This also affected the ability for customers to launch new EBS-backed EC2 instances. The EC2 and EBS APIs are implemented on multi-Availability Zone replicated datastores. The EBS datastore is used to store metadata for resources such as volumes and snapshots. One of the primary EBS datastores lost power because of the event. The datastore that lost power did not fail cleanly, leaving the system unable to flip the datastore to its replicas in another Availability Zone. To protect against datastore corruption, the system automatically flipped to read-only mode until power was restored to the affected Availability Zone. Once power was restored, we were able to get back into a consistent state and returned the datastore to read-write mode, which enabled the mutable EBS calls to succeed. We will be implementing changes to our replication to ensure that our datastores are not able to get into the state that prevented rapid failover. Utility power has since been restored and all instances and volumes are now running with full power redundancy. We have also completed an audit of all our back-up power distribution circuits. We found one additional breaker that needed corrective action. We've now validated that all breakers worldwide are properly configured, and are incorporating these configuration checks into our regular testing and audit processes. We sincerely apologize for the inconvenience to those who were impacted by the event.
- namidark 14y agoSounds like Amazon is doing something wrong, shouldn't it fail over to Battery then Generator?
- WestCoastJustin 14y agoYou're correct. It goes battery then generator. If they didn't use battery first, then when the power initial failed all systems would be off-line as the generator takes about 15-30 seconds to kick in.
- ericabiz 14y agoHe's correct or incorrect, and that entirely depends on the facility. Many newer datacenters don't use battery banks--they are expensive to maintain and often cause more failures than they prevent.
- rdoherty 14y agoFrom what I can remember of a datacenter tour, most generators supply power to the batteries, which then supply power to servers. The batteries can only supply a few minutes (I think) of power, so the generators need to turn on immediately.
- mrkurt 14y agoBattery backup isn't usually considered a "failover" step, just like your desktop battery backup doesn't actually do much other than stop charging when the power goes out. Datacenters only really have battery backup to let the emergency generators come up.
- ericabiz 14y agoThe batteries at colocation facilities are only designed to hold power long enough to transfer to the generator. They're also a huge single point of failure. A better design is a flywheel that generates enough power. But datacenters are often hit with these generator failures (in my experience, once every year or so.) Amazon had a correct setup--but not great testing. By the way, these are great questions to ask of your datacenter provider: Are there two completely redundant power systems up to and including the PDUs and generators? How often are those tested? How do I set up my servers properly so that if one circuit/PDU/generator fails, I don't lose power? There is a "right way" to do this--multiple power supplies in every server connected to 2 PDUs connected to 2 different generators--but it's expensive, and many/most low-end hosting providers won't set this up due to the cost. (I ran a colocation/dedicated server company from 2001-2007.)
- heretohelp 14y ago1 in a (million/billion/trillion) I guess. That'll make for a great horror story to tell though.
- kbutler 14y agoIsn't the moral of the story, "Check your backups"? There was a defective fan in one generator (sounds like it was findable via a test run?) and a misconfigured circuit breaker (sounds like it was findable by a test run). Redundancy is only helpful if the redundant systems are actually functional.
- olefoo 14y agoHaving been affected several times by colocation facilities bouncing the power during a test of the failover system, I can tell you that such tests are not without risk. Yes, you should test redundant systems, but how often, at what cost, and what risks are you willing to run while doing so. It's a fact of life that when dealing with complex, tightly coupled systems with multiple interactions between subsystems that you will routinely see accidents caused by improbable combinations of failures.
- spartango 14y agoI wonder if it's better to create an accidental outage during a scheduled test, or to have an outage completely out of the blue. Obviously mitigation is tricky even during a scheduled test, but perhaps its plausible?
- dkulchenko 14y agoWith a scheduled test, you have the benefit of having the main power actually working if the backup being tested comes crashing down; seems to me that mitigation would be much quicker in a scheduled test than in a real outage.
- excuse-me 14y ago
- jluxenberg 14y ago"Those customers with affected instances or volumes that were running in multi-Availability Zone configurations avoided meaningful disruption to their applications" "Meaningful disruption" is a bit of a weasel word; Amazon's own EBS API was down for almost two hours[1] despite being designed to use multiple AZs [1] "the EBS-related EC2 API calls were impaired from 8:57PM PDT until 10:40PM PDT ... The EC2 and EBS APIs are implemented on multi-Availability Zone replicated datastores" Guess the moral of the story is, if you require high availability then you must test your system in the face of an availability zone outage.
- jrockway 14y agoThe RSS link was quite amusing. My Chrome instance downloaded the RSS file without displaying it. Then I clicked it to open, and it opened Firefox. Firefox showed its file download box, suggesting I open the RSS with Google Chrome. Deadlock detected.
- nowarninglabel 14y agoDoes anyone know a solution for this? When I got upgraded to this version of Chrome (Version 20.0.1132.34 beta), it started downloading RSS feeds instead of displaying them. Much sadness has ensued :(
- anaheim 14y agoAFAIK Chrome has never detected RSS correctly for me. I assumed this was special to Mozilla, with their "we should probably provide a basic version of everything - newsreader, FTP client, etc., even if it makes the browser a bit more bloated." Bit like emacs. Chrome is a bit like vi, if you want more stuff, there's probably some sort of extension. Sorry for the emacs/vi analogies, I'm not trying to flame :)
- rapind 14y agoYes there is an extension for it, by Google too. https://chrome.google.com/webstore/detail/nlbjncdgjeocebhnmkbbbdekmmmcbfjd https://chrome.google.com/webstore/detail/nlbjncdgjeocebhnmk...
- nowarninglabel 14y agoWell, no, the extension only allows you to subscribe to an RSS feed, whereas the aforementioned feature in Chrome was that it would open the RSS feed inline.
- tedunangst 14y agoOn Linux, right? Firefox, or whatever gnome/dbus/opendesktop/gtk fuckery it uses has all sorts of strange notions about file types. When I download a tar.gz file, it saves a copy to /tmp, then launches a new instance of firefox with a file:// url, which opens a save file dialog.
- anaheim 14y agoTL; DR: Shit happens. Don't use AWS as your only platform, you will get burned sometime. Guaranteed, you will also get burned if you try to host and run your own stuff. How competent you are determines which way you get burned less.
- mhartl 14y agoOr, if you use AWS as your only platform, accept that shit will happen from time to time. Unless your application is a matter of life and death, or unless billions of dollars are at stake, a little downtime now and then probably isn't that big a deal. (All my sites went down when Heroku did (including railstutorial.org, which pays my bills), but the losses are acceptable given the convenience of not having to run my own servers.)
- rdl 14y agoI think it's reasonable to escalate criticism of Heroku for remaining in a single AZ. They have had plenty of time and resources to fix this, and haven't, despite being quite competent. I don't know if it is that they don't think it's necessary (due to the profile of their current customers) or what, but I wouldn't use Heroku for anything as long as they remain in a single AZ, and would be really reluctant to advise other people to do so. I obviously really like the Heroku team and product and would love to use them otherwise. It wouldn't even need to be true seamless failover across AZs right away -- just offering a us-west and us-east Heroku would be enough for me, with shared nothing (maybe billing, or not even that), and then figure out redundancy yourself inside your app. Multiple regions is WAY better than multiple AZs within a region, too -- both for reliability and for locality. Obviously a real seamless multi AZ/multi region solution would be much more technically impressive, useful to users, and Heroku-like, but they shouldn't let the perfect be the enemy of the good here.
- spartango 14y agoWhile I'd agree with the general premise that diversification is a good thing in platform use if high-availability is a requirement, given that this outage was single-AZ, this particular outage should really highlight the point that your application should be multi-AZ scaled if it needs to be up.
- jtchang 14y agoData Center Operator: We've lost our main power. No problem though we have a backup generator so we are good! ... 5 minutes later ... Uhh boss, our backup generator's fan crapped out. But no worries we have a secondary generator just for this kind of scenarion! ...10 minutes later and lights go out... "Well damn...looks like we configured the breaker wrong. This is not a good day."
- lutorm 14y agoThings could be have been worse -- it could have been a nuclear power plant. oh wait...
- forgotusername 14y agoIn my time at larger companies, DC power seems to be one of the weakest links in the reliability chain. Even planned maintenance often goes wrong ("well we started the generator test and the lights went out, that wasn't supposed to happen. Sorry your racks are dead"). Usually the root cause appears simple - a dead fan, breaker set to the wrong threshold, alarm that didn't trigger, incorrect component picked during design phase, or whatever else that gets the blame - things it would seem to a software guy that good processes could mitigate. Can any electrical engineers elaborate on why power networks fail (in my experience at least) so frequently? I guess failure modes (e.g. lightning strike) are hard to test, but surely an industry this old has techniques. Is it perhaps a cost issue?
- mrkurt 14y agoIt's really incredibly complicated, and difficult to test fully. The bits of Amazon's DC that failed seem like stuff normal testing should catch, but the DC power failures I've dealt with in the past always had some really precise sequence of events that caused some strange failure no one expected. As an example, Equinix in Chicago failed back in like 2005. Everything went really well, except there was some kind of crossover cable between generators that helped them balance load that failed because of a nick in its insulation. This caused some wonky failure cycle between generators that seemed straight out of Murphy's playbook. They started doing IR scans of those cables regularly as part of their disaster prep. It's crazy how much power is moving around in these data centers, in a lot of way they're in thoroughly uncharted territory.
- rdl 14y agoThe even crazier thing is big industrial plants where they are using tens or hundreds of MW and have much lower margins than datacenter companies, so they run with dual grid (HV, sometimes like 132kV) feeds and no onsite redundancy. As in, when the grids flicker, they lose $20mm of in-progress work.
- bigiain 14y agoI'd guess that's because "tens or hundreds of MW" of on-site backup power would be _ludicrously_ expensive to own/maintain, and the tradeoff against the risk of both ends of their dual grid flickering at once and trashing the current batch is less expensive. (or maybe the power supply glitches are insurable against, or have contract penalty clauses with the power companies?)
- damian2000 14y agoI love this sentence: Those customers with affected instances or volumes that were running in multi-Availability Zone configurations avoided meaningful disruption to their applications; however, those affected who were only running in this Availability Zone, had to wait until the power was restored to be fully functional. Translation: If you have a redundant (multiple-AZ) installation, then you were ok, if not then your server died.
- drags 14y agoDid anyone else run into issues with ELB during the outage? We're multi-AZ and could access unaffected instances directly without a problem, but the load balancer kept claiming they were unhealthy.
- ferringham 14y agoExcuse me being blunt: don't they routinely test power going off? This all sound like they have never tested. At Intel we tested power going off from time to time...
- tzury 14y agoSeems like deploying on two _physical_ regions (or more) is the best and only proven approach. That could be within the global AWS, or even say, one cluster at AWS and the other at RackSpace/Linode, etc.
- robryan 14y agoThen you just need to worry about your application consistency with the replication lag. No silver bullet I guess.
- mleonhard 14y agoI'm running https://www.rootredirect.com/ https://www.rootredirect.com/ and http://www.restbackup.com/ http://www.restbackup.com/ in us-east-1, in multiple availability zones. Both sites remained up with no problems.
- aparadja 14y agoDoes your rootredirect service actually attract paying customers? I'm genuinely interested to know.
- joelcollinsdc 14y agoAlso, how do you do this without having to handle the customers DNS lookups as well?
- mleonhard 14y agoThe title is incorrect. It should say something more like "Cascading failures cause part of AWS to go down."
- tysont 14y agoOn the plus side, the level of transparency that AWS displays and the detail that they provide seems above and beyond the call of duty. I find it refreshing and I hope that other companies follow suit so that customers can understand the details of operational issues, calibrate expectations appropriately, and make informed decisions.
- rdl 14y agoThey're less transparent and responsive than most datacenter or network providers -- it's just that most of those providers hide their outage information behind an NDA, so only customer contacts get it, vs. making it public.
- smackfu 14y agoYeah, a good datacenter will have SLAs around the root-cause analysis document for any failures. Like a preliminary report within a day and a final report within 7 days.
- acdha 14y agoI've also had a few cases where providers either outright lied or only gave details if you persisted in requesting them. Having to play that game gets old…
- moe 14y agoCan someone translate that to control rods and manifolds?
- gosub 14y agoCould it be possible to have power management the same way Erlang manages processes? Instead of 2 or 3 enormous backup power unit, hundreads of small ones to come in and out of use "fluently".