11 ms·
Google Cloud networking issues in us-east1
- geogram 7y agoLonger than 4 hours. We have stackdriver setup to monitor uptime/latency and its been acting up since 2am PST.
- tlynchpin 7y agoObPedant: notice in google's status page "...as of Tuesday, 2019-07-02 09:11 US/Pacific." This notation is useful because it's stable year round. I don't recommend 'PDT', instead colloquially 'out here on the left coast' or specifically US/Pacific.
- geogram 7y agoThanks. Good point. Regardless, gcloud has been having issues for nearly 12hours. (timezone agnostic)
- bob33212 7y agoDown or just high latency? For some folks that is the same thing.
- crankylinuxuser 7y agois a 4 hr latency, "latency"? You make a good point though. Downtime seems to be awfully overloaded.
- geogram 7y agoOn our tests the latency is surprisingly low (20-40ms) but it has an error rate of 10-30%.
- thsowers 7y agoWhy so many problems at Google lately? Calendar down two weeks ago[0], and Google Cloud had a larger outage a month ago[1] [0]: https://news.ycombinator.com/item?id=20213092 https://news.ycombinator.com/item?id=20213092 [1]: https://news.ycombinator.com/item?id=20077421 https://news.ycombinator.com/item?id=20077421
- zbowling 7y agounrelated. very big company with thousands of products that don't suffer outages. two incidents doesn't make a pattern.
- thsowers 7y agoI would argue that two direct Google Cloud outages within a month is pretty concerning for GCP customers, and that it's possible that the calendar outage could also be related in someway since it is likely hosted on GCP, although that is speculation
- tdhoot 7y agoDoubt Calendar is hosted on GCP. Generally Google does not run first-party systems on GCP, instead putting them on Borg (internal cloud).
- judge2020 7y agoRelated reading: https://news.ycombinator.com/item?id=9428043 https://news.ycombinator.com/item?id=9428043 https://news.ycombinator.com/item?id=17576720 https://news.ycombinator.com/item?id=17576720
- samcday 7y agoWhich, IMO, is actually a big problem. AFAIK Amazon are running a lot of actual production loads on AWS. Dogfooding can be extremely valuable, especially if a massive portion of your staff have the same profession as your target market. I've been using Google Cloud in a new role I started recently. There's definitely some parts of GCP I like, but whenever I use the Web Console I get the distinct impression nobody at Google actually uses it. If they did, I'm fairly sure all the annoying little warts I encounter would not exist.
- fastest963 7y agoHere's the original issue: https://status.cloud.google.com/incident/cloud-networking/19015 https://status.cloud.google.com/incident/cloud-networking/19... Not sure why they closed that one at 9:12 just to open a new one at 10:25. We didn't see any traffic coming to us-east1 during that time period so I would assume the original issue is still the root cause.
- joshuamorton 7y agoHopefully the thread title can be updated. (If it were actually down, this thread would have been posted 3 hours ago and have 400+ comments).
- boulos 7y agoYeah, that happens sometimes based on which team notices, thinks it might be different and then opens an outage. Sorry for the confusion, and yes, the fiber link issue is the root cause. Draining the Google.com traffic presumably resolved the issue for you, though you may still be seeing elevated latency as the updates suggest.
- fastest963 7y agoSince we use GCP Global LBs I presume that "draining the Google.com traffic" also meant that you're diverting all global LB traffic, which is what we see. The second incident (the OP's link) indicates that but at first it was very confusing to a customer when the first issue was marked as resolved but we still saw no traffic being sent to us-east1 via our global LBs. If that makes sense.
- boulos 7y agoThis part was somewhat nuanced, so I wasn’t sure to post it: yes, if you are using GCLB, and have more than 1 healthy Region, we will also rebalance to avoid us-east- for now (though not so statically as that sounds, mumble mumble). Edit: added this to the top level comment so more folks see it.
- inlined 7y agoHoly crap. It’s an outage in all zones? What’s the point of AZs if you lose whole DCs at a time.
- klodolph 7y agoAvailability is hierarchical.
- neonate 7y agoCan you explain that more?
- klodolph 7y agoThere is no service with 100% availability. You put multiple AZs in one region but nobody was ever pretending that regional failures were impossible, just that single-AZ failures are more common than regional failures. You want high availability, you want multi-regional. Above that you want multi-provider. The same decisions that make regions fail also makes infra-region traffic cheaper. This is true for all large cloud providers. If you are okay paying more for internal network traffic you can get multiregional. But multi-AZ is still better than single-AZ. Up to you to decide if it’s worth it. For that you need good SLAs and (IMO) support contracts.
- neonate 7y agoThanks, I understand what you meant now.
- dekhn 7y agoregions are the point. this is known as a "meteor outage".
- MrStonedOne 7y agoOperational Consistency creates a hidden single point of failure
- 7y ago
- estsauver 7y agoIn a moment that's likely to be very, very frustrating for a large number of you that have businesses and customers that depend on G cloud, let's try to remember that somewhere there's an engineer or an SRE having a really hard day just trying to fix things. Please, be kind and decent to each other, especially when things are hard.
- danaur 7y agoI don't follow comments like these, should people refrain from criticising giant companies because there are people working at them? I don't understand the purpose of this comment
- vidar 7y agoHe is asking people to be constructive in their criticism
- jpitz 7y agoThe purpose of the comment, to me, is to remind folks to refrain from taking your frustration with a product or a company out on a person.
- noob_slayer 7y agoAccording to some US Code, person means company [0]; so, we should avoid taking frustrations out on companies altogether? [0] https://www.law.cornell.edu/uscode/text/26/7701 https://www.law.cornell.edu/uscode/text/26/7701
- highesttide 7y agoComplaining about the communication and response time of a company is different from yelling in the direction of some stressed engineer that they are useless and incompetent at everything they do. Sadly you get too much of the latter around the Internet.
- StreamBright 7y ago
- partiallypro 7y agoIt's been down for 4 hours and it's just now being posted on HN? Is it intermittent?
- larkeith 7y agoFrom another comment, original issue [1] was closed at 9:12, so looks like they got it back up for a bit over an hour before it went down again. Post-mortem will be interesting. [1] https://status.cloud.google.com/incident/cloud-networking/19015 https://status.cloud.google.com/incident/cloud-networking/19...
- boulos 7y agoDisclosure: I work on Google Cloud. There were (and continue to be) connectivity issues due to a subset of the fiber links having trouble. But that’s different from being “down”, it’s “just” an outage. We won’t declare the outage over until the impact is minimal.
- noncoml 7y agoBad config push again?
- deleted 7y ago[deleted]
- deleted 7y ago[deleted]
- gaogao 7y agoRunning a betting pool on cloud service outage root causes would be fairly fun. I'm going to guess load balancer cascading failures.
- notriddle 7y agoNope. Physical destruction of fiber-optic cables is to blame, according to the GC status page. https://status.cloud.google.com/incident/cloud-networking/19016 https://status.cloud.google.com/incident/cloud-networking/19...
- hnaccy 7y agoWhat's the actual number of 9s for the major cloud services these days? My impression from their PR seems to mismatch the number of outages and issues lately.
- zzzcpan 7y agoI don't think there is any mismatch, it was always three nines. I guess the only mismatch is in claims that three nines is enough to not be noticeable or annoying to people.
- Johnny555 7y agoAWS EC2 promises 4 9's (4.3 minutes of downtime/month) before their SLA kicks in, but they only give a 10% discount until availability dips below 99% (7.5 hours of downtime/month) when they give a 30% discount. If availability is below 95% (36 hours) in a month, they give a full refund. For an individual instance, they only promise 90% availability.
- user5994461 7y agoAvailability of what? I've noticed entire afternoon where it wasn't possible to provision instances of some types, when I was working with AWS daily.
- Johnny555 7y agoAvailability of running instances, I don't think they make any guarantees for availability unreserved on-demand instances (I don't see how they could).
- sudosteph 7y agoAvailability of network access to existing instances. What you're talking about with provisioning capacity is a totally different matter. Provisioning availability is not guaranteed (unless you purchase reserved instances) and there are frequently periods where certain instance types are not available in certain AZs, though they do try to resolve that as fast as practicality allows them to. It really stinks sometimes though - especially if you get into a situation where something fails in your autoscaling group and there is no capacity available for a replacement instance. Usually you can get around that though by making sure your ASG is set up for multiple AZs, or worst case changing instance types (though that can be problematic in it's own way). source: I used to work for AWS Support.
- pkaye 7y agoLooks like all that high end engineering talent and processes still has its limits.
- pastor_elm 7y agolaughs in aws
- frostyj 7y agoleetcode can't buy you stability I guess
- Thaxll 7y agoLooks like an external issue. "The Cloud Networking service (Standard Tier) has lost multiple independent fiber links within us-east1 zone. Vendor has been notified and are currently investigating the issue."
- foobiekr 7y agoIt’s surprisingly hard to avoid shared fate links and it’s one of the things I would have thought google would be expert at.
- vinay_ys 7y agoIt's not that hard. In India because of so much construction related digging cuts OFCs, we do the path planning quite well and our redundancies get tested quite regularly whether you want to or not.
- tyingq 7y agoIt can be hard. Getting redundant separated paths under/over railroad tracks, for example, might require political power that not everyone has. Google, of course, has plenty.
- dragonwriter 7y ago> Getting redundant separated paths under/over railroad tracks, for example, might require political power that not everyone has. Google, of course, has plenty. But Google's vendors might have less. One would hope that Google is auditing claims of independence from vendors at least somewhat, but at some level they have to rely on vendor representation and SLAs if they aren't going to do it all themselves.
- vinay_ys 7y agoThe companies who operate the cross-country backbone fibers have independently verified fibre maps and you can also audit them with their cooperation. And those who operate last-mile metro networks are usually highly reputed ISPs (at least in India where there is decent competition in this space) who have a lot to lose if their reputation is damaged. Also, the community of their customers is small and they all talk to each other. So it is hard to make fake claims and get away with it. Usually, when cable cuts happen, it is more a question of whose traffic is rerouted on the available paths and whose traffic is dropped. If you are high-paying customer with strong SLAs then your traffic is usually safe and will displace a lower SLA customer's traffic. You will notice latency spikes due to rerouting and maybe temporary glitches w.r.t link stabilization. Since you see this so often, your BGP timers etc are all tuned to be patient and avoid cascading failures.
- deleted 7y ago[deleted]
- mrmattyboy 7y agoTo whomever commented something like 'laughs in AWS' (comment was removed before I submitted the comment)... please don't... glass house and all that... but I also share the same glass house as you.. I don't want bad luck ... and it's only a fluke that this happened to google in eu-east1 and not AWS in X region and then you (and I) would be having a time of hell! :/
- dymk 7y agoAnd did we forget about the insane AWS east outages of two years ago?
- fjp 7y agoDuring that AWS outage I was training people on [enterprise software] as part of the certification portion of [enterprise software company annual conference]. Nobody really wanted to be [enterprise software]-certified, but it was a way to get their employers to pay for them to go to the conference with cool talks and perks and such. We delayed the training most of the day, and couldn't say it was AWS' fault because they were sitting in the audience, waiting to get certified. People were about to riot, that was not a fun day.
- deleted 7y ago[deleted]
- outworlder 7y agoGoogle seems to be more forthcoming with their issues. We have seen incidents in AWS where the status never got updated, but support confirmed issues.
- deleted 7y ago[deleted]
- deanCommie 7y agoShow me a GCP post-mortem that's as detailed and proactive about future improvement as https://status.aws.amazon.com/s3-20080720.html https://status.aws.amazon.com/s3-20080720.html Their last one was laughable in it's lack of self-awareness.
- saltminer 7y agoThe title says "almost 4 hours" (was posted at around 3 PM EST), but the incident was created at 10:25 AM PST, which is 1:25 PM EST. Has it been more like 2 hours or is there more to this incident?
- garyb2 7y agoNotice they did not get around posting the next status update on time.
- deleted 7y ago[deleted]
- deleted 7y ago[deleted]
- username444 7y agoCloudflare was returning a 502 this morning, wonder if they're related. Lots and lots of sites down for about an hour, including all of Shopify.
- benbristow 7y agoI highly doubt Google are using CloudFlare networks. Must be just a coincidence.
- totaldude87 7y agoor CloudFlare using GCP :)
- itslennysfault 7y agoNah, CloudFlare runs on bare metal. They run their own data centers.
- jgrahamc 7y agoNot related
- mirceal 7y agonope. cloudflare had a bad push / deployment.
- naniwaduni 7y ago"bad push / deployment" seems like it covers 108% of breakage.
- mirceal 7y agosure. i believe you are 110% right on the 108% number :)
- boulos 7y agoAs jgrahamc (Cloudflare CTO) noted below, these aren't related. They had a push that they rolled back, we lost some fiber links.
- codingslave 7y agoGoogle engineers ran into a coding problem that wasnt on leetcode
- sieabahlpark 7y agoNow isn't that the truth
- dang 7y agoPlease don't post unsubstantive comments here.
- dx87 7y agoKind of related to this, but these types of outages are why I moved from Google Play to Spotify for streaming music. Their infrastructure seems so large that things that should be a standalone service, like streaming music, are bound to be collateral damage when they mess something up on another service. Having everything provided by one company is convenient until it all goes down at the same time and you can't access your email, videos, or music because they all run on the same infrastructure.
- mav3rick 7y agoSpotify is on Google Cloud.
- tomschlick 7y agoSpotify is hosted on google cloud: https://www.wired.com/2016/02/spotify-moves-itself-onto-googles-cloud-lucky-for-google/ https://www.wired.com/2016/02/spotify-moves-itself-onto-goog...
- mav3rick 7y agoLol.
- crusader76 7y agoI think the point OP was trying to make was relating to google services and their dependencies on each other.
- lern_too_spel 7y agoI have the opposite conclusion as OP. Google doesn't use Google Cloud for anything critical, so I wouldn't use Google Cloud for anything critical or services that run on Google Cloud for that matter.
- cbhl 7y agoIn my personal opinion, you should move off of Google Play Music, but not because of the dependency on Google infrastructure. https://9to5google.com/2018/05/23/google-play-youtube-music-2019/ https://9to5google.com/2018/05/23/google-play-youtube-music-... https://www.digitaltrends.com/music/what-happens-to-google-play-music-youtube-music/ https://www.digitaltrends.com/music/what-happens-to-google-p...
- wwwpppddd 7y agoApp Engine and Cloud functions were apparently returning error rates of > 30 percent overall between 11 a.m. and 3 p.m., with some projects experiencing a 100 percent error rate. GCS was also experiencing issues for the first half, which was attributed to the networking issues. Google said the networking issues were resolved initially but then stated they were investigating the GAE issues. Those issues were resolved, and the networking issue has been reopened as of 2:35 eastern: https://status.cloud.google.com/incident/cloud-networking/19016 https://status.cloud.google.com/incident/cloud-networking/19.... GAE and all other services still show green here, of course: https://status.cloud.google.com/ https://status.cloud.google.com/
- z3t4 7y agoWhen choosing a big cloud provider people forget that it's many orders of magnitude more complicated to run something at Google scale then to maintain one single server. For example the whole Stack overflow website runs on one or two servers. World of Warcraft also used to run on one single (blade) server. Chances are one server will be good enough for most use cases. And if you don't want to have it in your closet there are plenty of dedicated hosting and colocations.
- avocado4 7y agoHow can Stack Overflow run on a single server? Do you mean single cluster?
- davedunkin 7y agoAs of 2016, Stack Overflow ran on dozens of servers in two data centers. https://nickcraver.com/blog/2016/03/29/stack-overflow-the-hardware-2016-edition/ https://nickcraver.com/blog/2016/03/29/stack-overflow-the-ha...
- mehrdadn 7y agoAlso interesting what their minimum requirements were in 2014 :https://nickcraver.com/blog/2013/11/22/what-it-takes-to-run-stack-overflow/#core-hardware https://nickcraver.com/blog/2013/11/22/what-it-takes-to-run-...
- edwintorok 7y agoWith a cloud it also means that when there is an outage there are potentially many sites/services affected all at once, and there is potentially nothing customers can do to fix it other than wait (or plan in advance, and use/pay for multi-AZ/multi-region/multi-provider redundancy). Such outages are also possible with traditional hosting providers, and when an outage does happen I'm not convinced whether a large public cloud would recover more quickly (due to better resourcing/expertise available to fix the problem), or a small hosting provider (which may have a smaller team, but the problems they deal with are at a smaller scale and more easily fixable). Either way you probably want some kind of CDN independent of your cloud/hosting provider that can help survive some of these glitches.
- boulos 7y agoDisclosure: I work on Google Cloud (but I'm not in SRE, oncall, etc.). As the updates to [1] say, we're working to resolve a networking issue. The Region isn't (and wasn't) "down", but obviously network latency spiking up for external connectivity is bad. We are currently experiencing an issue with a subset of the fiber paths that supply the region. We're working on getting that restored. In the meantime, we've removed almost all Google.com traffic out of the Region to prefer GCP customers. That's why the latency increase is subsiding, as we're freeing up the fiber paths by shedding our traffic. Edit: (since it came up) that also means that if you’re using GCLB and have other healthy Regions, it will rebalance to avoid this congestion/slowdown automatically. That seemed the better trade off given the reduced network capacity during this outage. [1] https://status.cloud.google.com/incident/cloud-networking/19016 https://status.cloud.google.com/incident/cloud-networking/19...
- deleted 7y ago[deleted]
- ricardobeat 7y agoTangential question: does Google allow employees, not directly tasked with it, to represent the company online as they wish? Most companies I know of have a strict ‘do not speak for the company’ policy.
- jauer 7y agoIt's probably less "as they wish" and more "here's an approved statement" or "your role involves engaging with external parties, here are some guidelines"
- kyrra 7y agoIt's a fine line. We are not allowed to represent Google in any kind of public discussion. But we can talk about some things we do, as long as we state it's our own opinion and we don't represent Google's views.
- IX-103 7y ago
- harshreality 7y agoHacker News: The real status page and help desk for the internet. Do companies realize how absurd this is? ETA: It seems someone at Google had a change of heart, and most of what boulos posted in this thread has been added as updates to the official google status page. Better late than never, I guess, especially if this is the start of a trend in outage reporting.
- notatoad 7y agoseriously, they've got a text field on the official status page, why not put the text boulos posted here in that instead of the meaningless text they've got there?
- boulos 7y agoCan you expand on why you find it “meaningless”? As my other comment says, I’m not in SRE and the real people fixing it are trying their best to remediate the problem. I agree that the text I posted (with blessing from SRE!) gives you some more detail, but you can’t do anything differently with it, right? What about the new text do you prefer? (We’re happy to improve!)
- toufka 7y agoYour, even brief, description is interpretable by your clients and some customers - and is actually really informative. It helps estimate the magnitude of the issue, and the types of downstream problems to expect or avoid. Knowing an astroid took out the entire continent tells you something about the repairability, resources required to fix the problem, and generally provides context for later updates, as opposed to other contexts like a cut fiber line, a burning datacenter or a bad power supply.
- notatoad 7y agoI think the difference between your comment here and the info on the status page is that after reading your comment i feel like i know what's happening. you're right, there's no additional actionable information there, the status page contains everything i actually need to know. but a bit more information makes me feel better. I guess the difference is your comment reassures me that you actually know what's going on. the status page text (prior to the 14:31 update) could equally mean "we've got this under control" or "shit's broken and we don't know why"
- verdverm 7y agoI've been working out of us-east1 all day and haven't noticed
- lgats 7y agoPretty sure I've read before that us-east1 is one of the older Google data centers presumably with older equipment
- boulos 7y agoDisclosure: I work on Google Cloud. I think you’re thinking of AWS’s us-east-1 in Virginia. I don’t recall when us-east1 for us was constructed, but this wasn’t any sort of “old equipment” issue. Even there, while your experience may vary, AWS certainly has both old and new equipment.
- pupdogg 7y ago> The disruptions with Google Cloud Networking and Load Balancing have been root caused to physical damage to multiple concurrent fiber bundles serving network paths in us-east1. I am assuming some sort of construction zone at or nearby the facility and the backhoe operator dug in and accidently cut the cables?
- mehrdadn 7y agoDoes anybody else feel like there have been a lot of outages in recent months? And I don't mean Google -- I mean lots of others too (I seem to recall CloudFlare, Facebook, etc.)... are they really increasing or are we just hearing more about them? Seems a bit odd.
- jimmaswell 7y agoI came here to say this - it's like the cloud as a whole is imploding lately.
- rossdavidh 7y agoIt's almost as if we had made an overly complicated system with too much "efficiency" and thus not enough redundancy, centralizing on too few pieces of what used to be a quite widely dispersed system. The more "the cloud" replaces many, many servers at lots of different places, the more the outages (which once happened all the time, but to many different organizations at different times) will become big enough to notice. So, yeah, not just your imagination.
- mehrdadn 7y ago> It's almost as if we had made an overly complicated system with too much "efficiency" and thus not enough redundancy, centralizing on too few pieces of what used to be a quite widely dispersed system. The more "the cloud" replaces many, many servers at lots of different places, the more the outages (which once happened all the time, but to many different organizations at different times) will become big enough to notice. This is just for the last few months...?
- m0zg 7y agoThat's more or less inevitable. As complexity increases (which it does naturally, if there's no effort to decrease it) at some point it begins to outstrip the limits of human understanding. I've been saying this repeatedly (and downvoted for it repeatedly): if you want truly reliable systems, use simple, boring technology, and don't fuck with it after it's set up, and run it yourself. 99.99% of all these outages are due to screwing up something that already works, something that if it was in your own rack you could just leave alone and not touch at all.
- rco8786 7y ago2019 has been a really rough year for GCP
- awinter-py 7y agodo they not have extra hands on staff to dedup the messages? what's with the identical messages at 14:31, :44, :48? This happened last time too.
- deleted 7y ago[deleted]
- digitalsanctum 7y agoI routinely see notices of outages like this posted on HN while HN itself never seems to be impacted. This begs the question: Where and how is HN hosted in a way that avoids being impacted by widespread network and provider outages?
- hunter2_ 7y ago$ host news.ycombinator.com news.ycombinator.com has address 209.216.230.240 https://whois.arin.net/rest/net/NET-209-216-230-0-1/pft?s=209.216.230.240 https://whois.arin.net/rest/net/NET-209-216-230-0-1/pft?s=20... M5 Computer Security https://www.m5hosting.com https://www.m5hosting.com Unrelated: https://begthequestion.info/ https://begthequestion.info/
- MisterPea 7y agoThe begs the question site is one of my pet peeves. Language is not moderated by a select few who want to claim it, this isn't France. This is why Ebonics is still a valid form of English - as long as it is used consistently. If everyone uses "begs the question" and everyone else understands it as "raises the question" then it is perfectly valid.
- joemag 7y agoAny phrase that succeeds at transferring an idea out of your head into mine is good enough for me.
- hunter2_ 7y agoVernacular creates validity, for sure. The site's author acknowledges this, but nonetheless maintains that preserving this particular phrase is useful because there's not really another synonymous and popular phrase that means this particular fallacy, just the Latin petitio principii and the modern translation "laying claim to the principle" which is pretty clumsy if you ask me. Not that "begging the question" is crystal clear either, but at least it's googleable.
- jsjohnst 7y ago
- mountainofdeath 7y agoAnother day, another Google outage. It feels like it's once a month this year
- dragonwriter 7y ago> The disruptions with Google Cloud Networking and Load Balancing have been root caused to physical damage to multiple concurrent fiber bundles Is this concurrent damage to separated bundles or damage to colocated bundles?
- deleted 7y ago[deleted]
- spullara 7y agoWow. GCP is always a networking issue. Their QA on networking changes needs work. Maybe they should spend 20% on it.