7 ms·
As DigitalOcean's CTO, I'm very sorry for this situation and how it was handled. The account is now fully restored and we are doing an investigation of the inci
by bcooks 7y ago
As DigitalOcean's CTO, I'm very sorry for this situation and how it was handled. The account is now fully restored and we are doing an investigation of the incident. We are planning to post a public postmortem to provide full transparency for our customers and the community.
This situation occurred due to false positives triggered by our internal fraud and abuse systems. While these situations are rare, they do happen, and we take every effort to get customers back online as quickly as possible. In this particular scenario, we were slow to respond and had missteps in handling the false positive. This led the user to be locked out for an extended period of time. We apologize for our mistake and will share more details in our public postmortem.
- harshalizee 7y agoSure, but the email he received basically said "your account is locked. No other info. Thank You". That to me is a much scarier thing than anything else in the thread. How can anyone trust in your infrastructure if your standard protocol is literally just shutting down their entire operation without any form of review or communication?
- throwaway34241 7y agoYou can't, obviously. Even though I've used them before I really doubt I'll ever use DigitalOcean again. I can almost understand terminating customers (with notice) via automated heuristics for suspicious behavior, especially on the low end of the hosting market, but locking out a legitimate paying customer from backups with no notice or recourse is terrifying.
- jasonjayr 7y agoWe have a relatively large spend ($5k+) @ DO, for a unique client (most of our other clients can be served by our colocated facility), and I'm going to second this. Or with any other provider. They should always explain exactly which rule was broken. If the customer is legit + genuine, they will promptly fix the issue and won't be a further problem. Being vague makes it super troublesome to rely on any service that takes that tactic. (Like Google, for example) If they continue to re-offend, and find other ways to skirt the rules, that's when you move on to account termination.
- nathan-io 7y agoMistakes happen, and algorithms are sometimes a necessary part of scale/efficiency. Everyone understands that. That said, what's highly troubling as a DO customer (and someone who is planning to deploy startup infrastructure of my own with DO) is: 1) The discrepancy between this customer's experience and clear assurances made on this very forum by high-level DO employees that: a. warnings are ALWAYS issued before suspensions. b. even in the event of a suspension, services remain accessible (though dashboard access and/or the ability to spin up NEW services may be impacted), ie. the affected customer could still retrieve data or SSH in to droplets. 2) The relatively trivial nature of the customer's offending usage (temporarily spinning up 10 droplets). What happens if, for example, a startup gets a press mention somewhere that leads to a massive traffic spike, necessitating a sudden and significant spin-up of new droplets (especially if this is done programmatically versus by hand in the dashboard)? 3) The apparent lack of consideration of the customer's history, or investigation into their usage. It seems the threshold for suspending services of longstanding customers who are verifiably engaging in commerce (taking a moment to look at their website and general online presence for indicators of legitimacy), should be SUBSTANTIALLY higher than for, say, an account who signed up a week ago. Context matters.
- tnolet 7y agoNot sure why you’re being downvoted. Point 2 is very relevant. Scaling instances due to sudden peaks should be totally safe. Even when automated. Guess AWS is still lonely at the top.
- jdcro 7y ago"In this particular scenario, we were slow to respond and had missteps in handling the false positive. This led the user to be locked out for an extended period of time." This didn't seem like a case of being "too slow" - the customer in question went through your review process (which was slow, yes), and the only response he got was "We have decided not to reactivate your account, have a nice day". That just seems like a lack of interest in supporting your customers that are falsely flagged.
- lapnitnelav 7y agoPublic post mortem? Brilliant. Hope you can share what you learnt from this incident and hopefully you'll take a hard look at your processes. I'd hate to be caught in the same issue, especially that we are already customers, and I'm not sure I'll have as much clout as Nicolas here to get your attention.
- kbenson 7y ago> and I'm not sure I'll have as much clout as Nicolas here to get your attention. It's occurring to me now that while I've successfully ignored twitter for years, I should probably rectify that just so I have somewhere to type my hopes and prayers when this eventually happens to me, and hope for a miracle. It sure seems like the only place they're listened to.
- sampo 7y ago> I'm not sure I'll have as much clout as Nicolas here to get your attention. Maybe keeping a twitter (and other social media) account with at least a certain number of follower should be considered a part of a company's security strategy? You'd also need to post something interesting periodically, to keep your follower, so that you have their attention when you need it.
- rdiddly 7y agoYou've got an additional problem though, which is that this tells us you have two support channels: one that doesn't work (i.e. yours, the one you built), and one that does (Twitter-shaming). The first channel represents how you act when no one's watching; the second, how you act when they are. Most people prefer to deal with people for whom those two are the same.
- buzzerbetrayed 7y agoAs someone who has been blown off by DO support, you hit the nail on the head.
- deleted 7y ago[deleted]
- dkersten 7y agoAs a DO user who was planning on ramping up usage in the coming weeks and months, this is what scares me and what is making me seriously reconsider.
- xvector 7y agoDo not use DO. The very fact that their default response to suspected spam is to cause prod downtime is so bizarre and unacceptable that it does not make any sense whatsoever for a business to rely on them.
- dkersten 7y agoThanks, I’ll stick with AWS then.
- sneak 7y agohttps://github.com/fog/fog/issues/2525 https://github.com/fog/fog/issues/2525 https://news.ycombinator.com/item?id=6983097 https://news.ycombinator.com/item?id=6983097 Running anything business or privacy critical on DO is madness.
- mirimir 7y agoWill the customer be compensated for business losses?
- system2 7y agoAre you dreaming?
- mirimir 7y agoSadly enough, yes. I'm sure that it's covered in DO's ToS. But the DO CTO did basically admit fault in a public forum.
- kartickv 7y agoIt would probably be worth it to restore trust: Refund all the money they've taken from this company for the last year, and apply a credit to their account for 3x that amount, say.
- bufferoverflow 7y agoI hope they sue and win. This bull###t needs to be fought.
- mirimir 7y agoIANAL, but DO's ToS is loaded with weasel words.[0] So if they can sue in some jurisdiction where the binding arbitration and liability limitations don't apply, maybe they could at least get a fair settlement. 0) https://www.digitalocean.com/legal/terms-of-service-agreement/ https://www.digitalocean.com/legal/terms-of-service-agreemen...
- bcooks 7y agoThanks for the replies. Let me try to address a few of the things I have seen here. We haven't completed our investigation yet which will include details on the timeline, decisions made by our systems, our people, and our plans to address where we fell short. That said, I want to provide some information now rather than waiting for our full post-mortem analysis. A combination of factors, not just the usage patterns, led to the initial flag. We recognize and embrace our customers ability to spin up highly variable workloads, which would normally not lead to any issues. Clearly we messed up in this case. Additionally, the steps taken in our response to the false positive did not follow our typical process. As part of our investigation, we are looking into our process and how we responded so we can improve upon this moving forward.
- good_guy 7y agofuck you cunt.
- lioeters 7y agoThank you for jumping in personally to clarify what happened. As a business owner with much of our infrastructure depending on DigitalOcean, the incident is concerning. It affects the reputation of DO as well as its customers. The demographics on Twitter and especially here on HN represents a sizable crowd with decision-making influence on DO's bottom line. I hope to see some effort being made to prevent situations like this in the future, and to regain the trust. As a (so far) satisfied customer, it's great to hear that: > A combination of factors, not just the usage patterns, led to the initial flag. > We recognize and embrace our customers ability to spin up highly variable workloads, which would normally not lead to any issues. > we are looking into our process and how we responded so we can improve upon this
- stevenjohns 7y agoI’ll be awaiting the post-mortem and, depending on that and the procedures proposed to stop this from happening again, will hold off moving everything I have from DO. The real “mess up” here was the bit where you blocked the account with no reason given and no further communication - other than the one-liner your intern wrote for the email. I’m expecting you to sit down with your legal team and rewrite your TOS to be more customer-focused and less robotic.
- good_guy 7y agoyou can go fuck yourself.
- chris_wot 7y agoIt's not the false positive that is the issue here. The issue is that a. it took way too long to get the business back up and running, and b. the second response gave no explanation and no recourse for the business to become operational again. The very fact that this can happen from an automated script with no oversight should give every one of your customers pause as to whether they continue with your service.
- Cakez0r 7y agoI'd say the issue is that DO is shutting down servers for any reason at all (legal issues aside). If DO sells a product with a particular capacity, why should they intervene at all if a user is using all of the capacity they're paying for?
- OBLIQUE_PILLAR 7y agoGcp bans mining. Most providers frown on running tor exit nodes.
- ryanackley 7y agoI'm genuinely curious. What type of fraud or abuse are you trying to prevent? Maybe cover that in the postmortem.
- LeonM 7y agoIf your DO (or other cloud provider) credentials are compromised, it's usually a matter of seconds before someone fires up the largest possible number of instances to start crypto mining.
- bcooks 7y agoYup. LeonM, you are correct. In this case that was the cryptocurrency mining detector that was triggered. More details in the postmortem.
- deleted 7y ago[deleted]
- jameshilliard 7y agoAre you aware that Viasat has blacklisted a huge amount of Digitalocean /24 subnets? I can't access many of my servers when I'm on a satellite connection in addition to other websites hosted on Digitalocean. I've talked with the Viasat NOC and they told me they were blocking Digitalocean subnets due malware.
- zhte415 7y agoThis is probably worth it's own post, it would be very interesting to see more detail. I'm also probably certain that this is also not exclusive to DO.
- jameshilliard 7y agoPosted here: https://news.ycombinator.com/item?id=20067036 https://news.ycombinator.com/item?id=20067036
- sneak 7y agoDO is famously bad at dealing with abuse reports; in a lot of cases they simply do nothing. I block their netblocks for a lot of things, too.
- MagicPropmaker 7y agoSo unless a person is popular enough to get enough people talking about it on twitter or hacker news, someone whose account is flagged by your bad script is going to lose his business. That doesn't sound good to me.
- amrx431 7y agoA year of Data backup lost. Do you realize how that alone may cause the clients to dump a company and do you realize that startups may never recover from fiasco like these? I understand that it was false positive triggered by internal systems. But how do you explain the delay in restoring the services and reflagging again within hours after the services were restored?
- rat9988 7y agoAt the end it was not even a matter of delay. It was more like, we locked your account, we don't want to hear you anymore until the end of times.
- system2 7y agoShould we be concern about our 40+ droplets with DO now? We built our business on DO, we really can go bankrupt as well as our 30+ clients if anything like this happens to us. Please change your support system ASAP otherwise we will be switching to another platform. We are expecting a very serious response from you.
- bufferoverflow 7y agoClearly you should be doing regular backups of everything, and not on DO. And make sure to test your backups. And make sure you have a fast migration plan into another cloud. Ideally you should be cloud-agnostic, but that's quite hard to achieve.
- xvector 7y agoDo not use DO. The very fact that their default automated response to spam is prod downtime is unacceptable. It requires so many failures in understanding the service being provided across the company for this decision making process to have ever actualized that there is no reasonable expectation of safety or trust from DO at this point.
- bpicolo 7y agoEvery cloud company has anti abuse systems that will limit your access to their APIs / take down your machines if abuse is suspected - for example if it looks like you're mining bitcoin. Your prod isn't any different from your staging for them
- locusm 7y agoCurious how you will compensate them?
- justforyou 7y ago>> and we take every effort to get customers back online as quickly as possible. In this particular scenario, we were slow to respond and had missteps in handling the false positive. You clearly don't make every effort, and did not -- so why waste the extra verbiage and switch from active to passive voice? Based on your cliche response I have zero confidence that DO will do anything substantial to address the root causes of the issue.
- mercer 7y agoI've found DO's public posts to be particularly grating in the "we are listening to YOU, our customer. we take feedback extremely seriously" department.
- scarface74 7y agoThat’s all well and good. But how do you plan to reimburse your customer for this gross negligence? I have never heard of such incompetence or lack of communications from anyone on AWS’s business support plan. Why should anyone trust DO over AWS or Azure?
- exabrial 7y agoNo offense, as I'm sure this has been hard, but a screwup like this publicly demonstrates DO is not ready for prime time competition against AWS, Azure, GCP and the like. I'd gladly do whatever it takes to KYC, send you my business license, tax returns, EIN, invoice billing, etc so you know there is someone behind my account. We spend thousands of hours eliminating single points of failure. If an automated system can undermine that work, DO is not an option for us to host anymore.
- justforyou 7y agoDo you realize that by abusing this thread to make a single PR focused comment with no intention of participating in the conversation -- you've disrespected the community here and the few remaining DO customers within said community.
- bcooks 7y agoLast week ended on a real low note for many of us at DO. We took a perfectly good customer and gave them an experience no one should have to go through (all while he was trying to leave on vacation no less). We can and must do better. To do better we need to learn from our mistakes. To that end, we also think sharing the information about this incident openly is the best way to help all our customers understand what happened and what we are doing to prevent it in the future. Yesterday we completed our postmortem analysis of the incident involving Nicolas (@w3Nicolas) and his company Raisup (@raisupcom). With their permission we are sharing the full report on our blog here: https://blog.digitalocean.com/an-update-on-last-weeks-customer-shutdown-incident/ https://blog.digitalocean.com/an-update-on-last-weeks-custom...
- bofadeez 7y agoUnacceptable. I just instructed my team to begin a transition to AWS.
- bofadeez 7y agoYou ruined your brand