7 ms·
Joyent us-east-1 rebooted due to operator error
Due to an operator error, all compute nodes in us-east-1 were simultaneously rebooted. Some compute nodes are already back up, but due to very high load on the control plane, this is taking some time. We are dedicating all operational and engineering resources to getting this issue resolved, and will be providing a full postmortem on this failure once every compute node and customer VM is online and operational. We will be providing frequent updates until the issue is resolved.
- alrs 12y agoJoyent's messaging about "we're cloud, but with perfect uptime" was always broken. It's mildly gross that the current messaging sounds like they're throwing a sysadmin under the bus. If fat fingers can down a data center, that's an engineering problem. I care about an object store that never loses data and an API that always has an answer for me, even if it's saying things that I don't want to hear. 99.999 sounds stuck-in-the-90s.
- evan_ 12y ago> sounds like they're throwing a sysadmin under the bus at least they didn't name the operator in question...
- elijahwright 12y agoOur internal culture is such that everyone on the team would rather be blamed for something than accuse someone else of doing it. That's shitty, and not something you do to someone. You fix the problem and then you move on. If it makes you happy, blame me - I don't mind.
- hack_edu 12y agoBTW, this is the right way to do it. :)
- niels_olson 12y ago"elijahwright" shall henceforth be used in place of "scapegoat"
- elijahwright 12y agoAwesome! It's what I've always wanted!!!
- arakkisu 12y agoit was that way at tech, no reason for it to change now
- elijahwright 12y agoNow I have to figure out who you are. :-)
- sokoloff 12y agoAt my $DAYJOB, we are always careful to figure out exactly what happened, including by whom. It's not to assign personal blame, but I believe it's critical that everyone agrees on the facts (who, what, when, where, and [if possible] why). Response and conversation is always focused on "how do we prevent this in the future?", not on punishing whoever was involved in the past. IOW, I agree with I believe is your intent, but differ on the implementation. Blameless transparency is the term we use (and we probably stole that from somewhere else). It's a very powerful signal to the whole team when you first see individuals "admitting" to exactly what they did, how it caused or contributed to the outage, and to hear them thanked for their contribution of understanding in the post-mortem. Senior leadership (including myself, who originally instituted the entire process a decade ago) is very clear that we want to know the facts and that in seeking and using those facts, we're only focused on the future, no matter how boneheaded the individual actions appear with the benefit of hindsight and knowledge that they'd lead (in)directly to an outage. I run operations and also participate in the promotion discussions for all technologists, and in 11 years, I've never heard a negative shadow cast onto a sysadmin/sysengineer from their actions during or leading to a production outage. And we've (collectively) made our fair share of mistakes over the years. That doesn't stop good employees from feeling bad about it, but that's a personal feeling they have, not from the fear of it being a professional black mark.
- M2Ys4U 12y agoI think there's a difference in how you approach this with an internal-facing view and an external one. Internally, You're right. But externally the company fucked up, not the individual.
- sokoloff 12y ago100% agree, and it is my oversight to not draw that distinction more clearly. We have the luxury (so far) of only reporting internally.
- deleted 12y ago[deleted]
- knodi 12y agoSure blame on the engineers. You give power, people use it badly blame the engineer for giving too much power. You don't give enough power sysadmins/users bitch and yell why don't we have enough power, we're not children. Its always the engineer fault. :(
- CHY872 12y agoIt's a combined fault. Clearly the operator made a mistake, but the system shouldn't have let such a calamitous operation take place without at least three levels of "Are you sure" (or something smarter like "Confirm how many servers would you like to reboot:") before it lets you take down thousands of servers.
- alrs 12y agoSystems engineers, software engineers, architects, whatever. We're all in the same gang. My point is that the problem in this case is likely the system's design, not one engineer's typing abilities.
- jsmthrowaway 12y agoThis comes down to operational philosophy, in the end. The point you're dancing around is whether the system should permit grave actions that don't make any sense when you're designing the system. By the rules, every single system on a commercial aircraft has a circuit breaker. Pilots make the "what if X catches on fire?" case, which is actually pretty compelling. However, that also means there are several switches overhead that will ostensibly crash the airplane if pulled. Pilots lobby very strongly for the aircraft not to fight them in any way because they are the only ones with the data, in the moment, now. They have final command over the aircraft in every way. I use this to point out that as you're designing systems for operations people -- something we're increasingly doing ourselves as devops/SRE takes hold -- you might think you can anticipate every scenario and design suitable safeguards into the system. However, sometimes, when Halley's Comet refracts some moonlight into swamp gas and takes your fleet down, you as an operator have to do some really crazy shit. It's in that moment, when all hell has broken loose, I'm at the helm, and based on the data available to me I have made a decision to shoot the system in the head: if the system fights me and prolongs an outage because we argued about whether we'd ever need to reboot a fleet all at once, I'm replacing the system as the first item in my postmortem. If you make me walk row to row flipping PDUs, we're going to have words. That's just my philosophy. Give the operators the knives and let them cut themselves, trusting that you've hired smart people and understanding mistakes will happen. Your philosophy may vary. By all means, ask me to confirm. Ask me for a physical key, even. But if you ever prevent me from doing what I know must be done, you are in my way. I have yet to meet a system that is smarter than an operator when the shit hits the fan (especially when the shit hits the fan). There's probably a broader term for operational philosophy like this.
- berns 12y agoJoyent's marketing is not the most transparent. They haven't updated AWS prices in their pricing page since AWS lowered their prices two months ago.
- 4ad 12y agoWhat? Joyent doesn't use AWS.
- SEJeff 12y agoAs I've always said, "You can never protect a system from a stupid person with root". You can limit carnage and mitigate this type of thing, but you can't fully protect against sysadmins doing dumb things (unless you just hire great sysadmins)
- llamataboot 12y agoI don't think "just hiring great sysadmins" is possible. People have off-days or are tired or sick, new people get on-boarded, even great people make mistakes, etc.
- protomyth 12y ago...or accidentally switch which of the 25 term sessions they had open I tend to color my production terms in a red background / yellow font scheme. It tends to inspire the tired brain to understand you are in production.
- wmf 12y agoSo don't give anyone root on an entire data center.
- akerl_ 12y agoIs this like Captain Planet? It's a bit exceptional to divide access servers of similar type between administrators such that individuals have full access to a portion of the fleet. Do they meet up and put their rings together to roll out updates? What if one of them goes on vacation?
- bcantrill 12y agoIt should go without saying that we're mortified by this. While the immediate cause was operator error, there are broader systemic issues that allowed a fat finger to take down a datacenter. As soon as we reasonably can, we will be providing a full postmortem of this: how this was architecturally possible, what exactly happened, how the system recovered, and what improvements we are/will be making to both the software and to operational procedures to assure that this doesn't happen in the future (and that the recovery is smoother for failure modes of similar scope).
- mixologic 12y agoI feel bad for the person who made the mistake. Even though its obviously a systemic problem, and highly unlikely to be an act of negligence, Im sure he/she doesnt feel too hot right now.
- socceroos 12y agoMy thoughts exactly. Poor fella. I've seen worse though. A newish officer spilled his morning coffee into the circuitry of a device worth over 10 zeros. Immediately short circuited.
- tedsanders 12y agoThat sounds so awful. I can't imagine living the rest of my life knowing that I had been a net negative in the world. All of my life's earnings would just be a partial restitution of that one second of destruction.
- JetSpiegel 12y agoIf you have an EXTREMELY reductive point of view, that equates revenue with human worth.
- ddorian43 12y agoreally? what can cost that much ?
- dharbin 12y agosalt '*' system.reboot
- quickdry21 12y ago> hubot restart all on prod oh shit i meant stag fuckfuckfuckfuck
- qbrass 12y agoSo have hubot second guess any changes to production unless you specifically told it you were messing with prod beforehand. Have it wait a few seconds before doing something important and listen for sounds of regret. >hubot restart all on prod hubot: > say "Hubot isn't responsible for hosing production because I actually meant staging" >Hubot isn't responsible for hosing production because I actually meant staging hubot: okay, don't say I didn't warn you. >oh shit i meant stag fuckfuckfuckfuck hubot: I hadn't started yet, but I'm doing it anyway just to teach you a lesson.
- angersock 12y agoWhy in the name of all that is holy do you have Hubot getting access to your production boxen? Why does that seem like a good idea, ever?
- akoumjian 12y agoMy thought, exactly. Time to setup some good ACL :-) http://docs.saltstack.com/en/latest/ref/clientacl.html http://docs.saltstack.com/en/latest/ref/clientacl.html
- shiftpgdn 12y agoLet this be a lesson to linux admins. Re-alias shutdown -r now into something else on production servers. I once took down access to about 6000 servers because I ran the script to decommission servers on our jump box when I got the SSH windows confused.
- cordite 12y agoI once put `shutdown -h now` (halt) instead of `shutdown -r now` (reboot) Once I realized what had happened on the production server I ended up calling OVH (and they were helpful but not immediately acting). It's not a good feeling.
- icebraining 12y agoI tend to use /sbin/reboot instead, it amounts to the same (calls shutdown), but it's harder to get it mixed up.
- smtddr 12y agoThis happened to me once; I don't know if this works on all linux distros but if you quickly follow a halt/shutdown with a "sudo init 6"(reboot) before your ssh-session gets SIGTERMed/KILLed, the box comes back up. This at least worked on some Ubuntu version a few years back. Give it a try on some system that's not critically important :)
- cordite 12y agoYeah, but the problem is when you honestly didn't realize calling a halting shutdown until the server doesn't come back 5 minutes later and then you review the terminal
- baconhigh 12y agoMight I suggest molly-guard: https://packages.debian.org/unstable/admin/molly-guard https://packages.debian.org/unstable/admin/molly-guard
- 12y ago
- deleted 12y ago[deleted]
- jameshart 12y agoDevOps means being able to take out an entire datacenter with a single keysstroke...
- akurilin 12y agoDevOps Borat is going to have a field day today.
- stephengillie 12y agoAs a Devops, I can't justify building any automated way to down or restart all of my systems at once. We've only had to do that to resolve router reconvergence storms when changing out (relatively) major infrastructure pieces, such as our Juniper router.
- tommu 12y agoSorry - are you telling us you had to reboot all nodes because you swapped a router out? Sounds like you need a network engineer.
- tommu 12y agoAnd I'm being downvoted for that? Seriously? In 13 years of networking I have never once had to reload machine to help with OSPF or BGP convergence. Good networking architecture and planning should mitigate anything other than a couple of minute outage. No routing change should ever require a reload of a server or end node.
- cpayne 12y agoI believe you were down voted not for what you said, but the way you have said it. I've been down voted several times for (what I see) as relatively minor remarks. The HN readers are a sensitive bunch...
- stephengillie 12y ago
- jordanthoms 12y agoLooks like the janitor needed somewhere to plug in the vacuum cleaner again...
- saganus 12y agoI assume bash.org?
- gknoy 12y agoHe might be referring to The daily WTF (worse than failure): http://thedailywtf.com/Articles/I-Didnt-Do-Anything.aspx http://thedailywtf.com/Articles/I-Didnt-Do-Anything.aspx Unintentional Mishap while Contractor Unplugs X to fix/maintain Y is a relatively common theme on their list of horror stories. edit: I think he might actually have meant this one: http://thedailywtf.com/Articles/I-Told-You-So.aspx http://thedailywtf.com/Articles/I-Told-You-So.aspx
- wiml 12y agoIt's a truly ancient anecdote; it probably predates the Internet. The first example in RISKS is in 1994: http://catless.ncl.ac.uk/Risks/15.59.html#subj3.1 http://catless.ncl.ac.uk/Risks/15.59.html#subj3.1 but the canonical version of the story is in a Cape Town hospital in 1996: http://web.archive.org/web/20040624065333/http://www.legends.org.za/arthur/cleanfaq.htm http://web.archive.org/web/20040624065333/http://www.legends...
- saganus 12y agoAnd I get 2 downvotes for this? really? downvoters care to explain why, just for asking if it was a reference from bash? Wow... Edit: Thanks to the other 2 posters who provided alternative sources. You learn by asking, no? or at least some of us do..
- hack_edu 12y agoNot even just the plug. I've had outages from bits flipped simply by the static electricity generated when vacuuming near servers.
- lukasm 12y agoMandatory DevOps Borat "To make error is human. To propagate error to all server in automatic way is #devops" and my fav "Law of Murphy for devops: if thing can able go wrong, is mean is already wrong but you not have Nagios alert of it yet."
- shanselman 12y ago"What's this button do?"
- devinegan 12y agoJoyent has been having some serious issues over the past month or two. I am not sure if it is growing pains, bad luck or what is happening, but we had already lost faith and trust in their Cloud prior to today. This is the nail in the coffin from our perspective. Moving on...
- rincebrain 12y agoHowso?
- devinegan 12y agoThanks for asking rather than just down-voting. I wanted others to know that this isn't isolated. We have been having issues with their service for a few months now. They never know when there is a problem with hardware, for instance. Joyent support will gladly tell you everything is fine. After you insist, and insist they will actually have someone look at the underlying infrastructure. Eventually they will acknowledge the problem and fix it (maybe). I believe the monitoring and reporting for their team is flawed or incomplete which leads to more downtime of affected systems. Just one observation, but we have had three incidents over the past month and a half. Two within a week of each other.
- bcantrill 12y agoI'm sorry to hear about your experience; we pride ourselves on being able to root-cause problems regardless of where they might be in the stack, but it sounds like your problem didn't get properly escalated. If you want to reach out to me privately (my HN login at acm.org), we can try to figure out what happened here -- with my apologies again for the subpar experience.
- Diederich 12y agoThe 'devops' automation I made at my last company (and am building at my current company) had monitoring fully integrated into the system automation. That is, 'write' style automation changes (as opposed to most 'remediation' style changes) would only proceed, on a box by box basis, if the affected cluster didn't have any critical alerts coming in. So, if I issued a parallel, rolling 'shutdown the system' command to all boxes, it would only take down a portion of all of the boxes before automatically aborting because of critical monitoring alerts. Parallel was calculated based on historical but manually approved load levels for each cluster, compared to current load levels. So parallel runs faster if there's very low load on a cluster, or very slowly if there's a high load on a cluster. One way or another, most automation should automatically stop 'doing things' if there's critical alerts coming in. Or, put another way, most automation should not be able to move forward unless it can verify that it has current alert data, and that none of that data indicates critical problems.