6 ms·
"Currently, Netflix uses a service called "Chaos Monkey" to simulate service failure. Basically, Chaos Monkey is a service that kills other services. We run thi
by mattew 15y ago
"Currently, Netflix uses a service called "Chaos Monkey" to simulate service failure. Basically, Chaos Monkey is a service that kills other services. We run this service because we want engineering teams to be used to a constant level of failure in the cloud. Services should automatically recover without any manual intervention. We don't however, simulate what happens when an entire AZ goes down and therefore we haven't engineered our systems to automatically deal with those sorts of failures. Internally we are having discussions about doing that and people are already starting to call this service "Chaos Gorilla"."
I am wondering how they could simulate the loss of an AZ. Any ideas?
- efsavage 15y ago> I am wondering how they could simulate the loss of an AZ. Any ideas? Nelson: How many chaos monkeys will there be? Bart Simpson: One at first, but he'll train others.
- ceejayoz 15y ago> I am wondering how they could simulate the loss of an AZ. Any ideas? Kill all instances in an AZ?
- RyanKearney 15y agoPerhaps they have groups set up in their "Chaos Monkey" tool? Like a sort of take down ALL services in GROUP B type of command?
- joenorton 15y ago"One of the first systems our engineers built in AWS is called the Chaos Monkey. The Chaos Monkey’s job is to randomly kill instances and services within our architecture. If we aren’t constantly testing our ability to succeed despite failure, then it isn’t likely to work when it matters most – in the event of an unexpected outage." Sounds like Chaos Monkey distributes his chaos randomly.
- gbelote 15y ago> I am wondering how they could simulate the loss of an AZ. Any ideas? They could instrument whatever library they use to interact with AWS and make it report failures or fail to respond to "create new instance"-like commands.
- SpikeGronim 15y agoThere are several ways to do it. Kill all the instances. Use a firewall to blackhole all the instances. Use traffic shaping to degrade the latency or packet loss of all the instances.