16 ms·
Google Cloud outage brings down Layer
- nwrk 10y agoDon't like the attitude. Pointing fingers doesn't help paying customers trapped by Layer poor design choices.
- smt88 10y agoEspecially when they seem to be referencing only a single region. Multi-region deployments is the most basic protection against outages when using IaaS.
- theDoug 10y agoYup. Just as too few companies realize that cloud doesn't mean one remote box replacing one local box, all eggs in one regional basket (or even single cloud provider) is unwise.
- andyfleming 10y agoThey should at least be in multiple availability zones. Multiple regions often comes with a lot of challenges, but there isn't much reason not to be redundant in multiple AZs.
- Artemis2 10y agoWere they or were they not in multiple AZs? Developing for multiple availability zones is trivial when creating cloud-first software (and it's irresponsible not to use AZs!), multi-region comes with its own set of problems.
- inlined 10y agoIn the report they say they're looking at moving to a new region, but Google apparently told them that us-central1-a was down. The "-a" makes it an AZ. It sounds like they're only on one AZ and may not fully understand the difference. [correction: they accidentally called usc1-a a region, but everything mentioned in their outage was a zone. They specifically called it a "deployment zone" not an availability zone, so it sounds like an issue of inexperience with best practices.] [obligatory disclaimer: I'm a Google employee. I don't have a relationship with Layer]
- smt88 10y ago> Multiple regions often comes with a lot of challenges As an almost-customer of Layer (before their massive price increase), they led me to believe that this was one of the problems they would be solving for me. Nowhere on their website does it say, "We save money by not following best practices, so plan accordingly for occasional outages!"
- bogomipz 10y agoThat was my first thought as well. Why did they need to start migrating customers to another AZ? I hope their customer's started asking that as well. The title should be "Poor Design Choices Brings Down Layer"
- blakewatters 10y agoBlake from Layer here: I've reviewed the updates from last night and I don't feel like the tone was out of line. We were simply trying to provide our customers with complete transparency about where the issue was and where we were in restoring service. With that said, we do feel that Google came up in short in their responses to us over the course of the issue. We pay handsomely on a support contract to get off-hours responses and issue escalations. The responses we received were hand-wavy and vague, leaving us without sufficient data to make decisions. We have raised these concerns with our Google representative and will be working with them to tighten our partnership going forward. We take full responsibility for this event and are working to cover the exposure. Building a system and business with resource constraints and complex distributed technologies is a long game of managing risk and trade-offs. We're human and we make bad calls along the way. We are very sorry and violated our commitments to our customers and their users. The entire Layer engineering team is head down right now working to make it right.
- mbesto 10y ago> As we are now several hours into this outage and do not have satisfactory timeline for resolution, we have begun the process of migrating our hosts into another deployment zone within GCE. We will have a baseline set of services migrated within the hour and evaluate our ability to operate in a split deployment. Should we need to pursue a complete migration of hosts across zones then we would expect another 4-5 hours to return to full operational capacity. Wait, their service isn't setup to operate in a split environment out of the box? I think it's time SaaS companies start documenting their IaaS setup so purchasers can do a high level audit before they decide to use it for potentially a core part of their own product/service.
- neom 10y agoI agree. If you're going to use abstracted infrastructure but you don't understand basic distributed architecture you shouldn't really be blaming your cloud provider.
- niftich 10y agoI imagine if one were a customer of this SaaS, it's on the customer to ask what availability to expect. Clearly this vendor thought that their savings on their IaaS bill outweighed any operational or reputational risk they'd suffer from an outage at a lower layer (pun unintended).
- blakewatters 10y agoBlake from Layer here: We are forthright with all our customers about our current deployment configuration and the roadmap timelines for evolving into a deployment with higher availability characteristics. There is real complexity in operating a system such as ours in a widely distributed configuration and like any other company at our stage we regularly assess risks and make trade-offs. Sometimes we get things wrong. We are very sorry to all our customers for the downstream impacts their businesses. We came up short and are doing everything we can to make it right.
- tschellenbach 10y agoI like how Algolia does that, https://www.algolia.com/infra https://www.algolia.com/infra (their blogposts and presentations go into much more detail) Currently thinking of creating a similar page for getstream.io, at the moment we always explain it during sales/onboarding calls. (we replicate our data to 3 different instances across multiple AZs)
- wiradikusuma 10y agoSlightly OOT: Anyone know good alternative to Layer?
- CometChat 10y agoCometChat works seamlessly on web, mobile & desktop! Your users can be on any platform and communicate with each other. Check out the demo here: https://www.cometchat.com/demo https://www.cometchat.com/demo
- DoubleMalt 10y agoUse https://matrix.org https://matrix.org if you don't want to depend on the infrastructure deccisions of a third party.
- herman5 10y agoFirebase is a great platform for building chat functionality
- alecsmart1 10y agoYou can try cometchat.com which is a self hosted. So no such issue.
- joshmarinacci 10y agoPubNub. We have fantastic uptime and do over a trillion transactions a month flawlessly. We also do some cool blogging too. (my job :) http://www.pubnub.com/ http://www.pubnub.com/
- billychia 10y agoTwilio's chat SDKs https://www.twilio.com/ip-messaging https://www.twilio.com/ip-messaging
- ben_jones 10y agoMost concise summary of Layer I could find on the internet quickly. > Layer is an amazingly elegant and light-weight solution for video communication. Layer is currently in a private beta primarily focused on Video, Voice and Chat on Android and iPhone. [1] Comment was in 2014. [1]: https://www.quora.com/What-is-the-difference-between-PubNub-and-Layer https://www.quora.com/What-is-the-difference-between-PubNub-...
- JasonSage 10y agoIf you just go to layer.com, the first text you see on the page does a pretty good job of spelling out what it is. At least, it did for me. It's also more up-to-date than that comment, it would seem.
- ben_jones 10y ago> Everything you need, from UI to infrastructure, to boost retention, engagement or drive transactions with the power of rich messaging. Wasn't enough for me. And if you click "Learn more" it's more marketing drivel. Granted my quora quote isn't much better.
- blakewatters 10y agoBlake from Layer here. Have you taken a look at our developer documentation on developer.layer.com? I felt like we did a pretty good job of presenting the product capabilities. Our homepage and the developer documentation speak to different audiences. Let us know how the developer side matches up to your expectations.
- tschellenbach 10y agoLayer is just a building block for adding chat to your app. Similar to how you would use Elastic for search or Sendgrid for email.
- jameskegel 10y agoThat reminds me, I wonder what ever came of the Adria Richards v. Sendgrid issue.
- flyt 10y agoIt doesn't do any good to point the finger at your vendors when your service goes down; that data isn't useful for your customers. Never forget the lesson of http://www.whoownsmyavailability.com/ http://www.whoownsmyavailability.com/
- ngrilly 10y agoI'm not sure I agree. Customers like to know why it doesn't work. If it was a physical machine, they would have said something like "the disks are broken and we are replacing them". But it is cloud and they said "Google persistent disks are currently unavailable and they are fixing it".
- knorker 10y agoBut the real reason is "we didn't set up our system properly". This is like saying "Hitachi Storage hard drives broke" when you actually mean "we didn't run RAID".
- ngrilly 10y agoYou can't compare persistent disks failing in a whole zone, with a RAID array failing in a single machine. There is a reason why Amazon and Google takes EBS/Persistent Disk failures very seriously: there are not supposed to be unavailable during several hours, except if the whole datacenter is unable to operate (flood, fire, etc.), but it's not the case here. If your RAID fails, and you have a support contract which guarantees restoration within 1 hour, and it's not restored within 1 hour, then I think you can legitimately say something was wrong at your provider. It's not pointing fingers. Everyone does mistakes. It's taking responsibility. That said, I agree they should have run in multiple zones, as recommended by Google, if they need/want to avoid that kind of downtime. But I maintain Google Compute Engine Persistent Disk are not supposed to fail in such a way, and I'm quite sure Google will do whatever they can to avoid this in the future, instead of saying "don't point finger at us, it's supposed to happen".
- jpatokal 10y ago
- johnm1019 10y agoI am so turned off when I click a "Pricing" link and get a contact form. Even more so when I read that, "our pricing team" [will get back to you]. So, you have an entire team of people who will try and maximize how much I pay? Sounds like a great experience doing business with you. /heavysarcasm
- homero 10y agoI automatically skip any product where I have to speak to a human at any time
- theprotocol 10y agoIs it me, or are a lot of web-based service providers very chatty lately? I won't name and shame any particular ones, but I will say I've found myself regretting signing up for trials of certain services because of the almost sycophantic attention I'd receive from the oh-so-personable and friendly CEOs who make it a point to personally message all customers. I usually respond, initially, but then it quickly becomes pushy and intrusive, e.g. "Hi, I've noticed you haven't used [x] feature yet." "Hello? Are you getting my emails?" "Hello?" I don't mean to be rude, but I didn't sign up for the "omg you're so friendly and amazingly helpful" show. I just wanted to try the service out. Kindly stop breathing down my neck! :/
- j_s 10y agoThis happens because it works, though not necessarily so much for the HN crowd.
- theprotocol 10y agoI'd wager the initial greeting works, but I question whether the person's lack of self awareness (when it's clear that the customer doesn't have time to do small-talk and has been evading you for 3 weeks straight) wouldn't be grating to most people who are trying to evaluate several products and get some work done.
- knorker 10y ago> As we are now several hours into this outage and do not have satisfactory timeline for resolution, we have begun the process of migrating our hosts into another deployment zone within GCE Wait, what? Isn't running in multiple zones something like rule #1 or #3 in "how to run in the cloud"? So why did they not already do this?