14 ms·
Level 3 Global Outage
- RedShift1 6y agoIronically this page doesn't load for me
- mathieubordere 6y agostackoverflow seems to be unreachable
- vld 6y agoI'm having issues reaching IP addresses unrelated to Cloudflare. Based on some traceroutes, it seems AS174 (Cogent) and AS3356 (Level 3) are experiencing major outages.
- jbotz 6y agoIs there any one place that would be a good first place to go to check on outages like this? It would be really cool and useful to have an "public Internet health monitoring center"... this could be a foundation that gets some financing from industry that maintains a global internet health monitoring infrastructure and a central site at which all the major players announce outages. It would be pretty cheap and have a high return on investment for everybody involved.
- thejosh 6y agoUntil that site also goes down.
- lioeters 6y agoIndeed, if we're to have a public Internet health meter, it must be distributed and hosted/served from "outside" somehow, to be resilient to all or parts of the network being down.
- johnisgood 6y agoHere's a thought: we should all be outside. :D
- efreak 6y agoSomething something anycast.
- mnadkvlb 6y agoSounds like a good idea. The closest i know is the one from pingdom which i use the most. Its not detailed enough though. https://livemap.pingdom.com/ https://livemap.pingdom.com/
- guerby 6y agoIn the network world there's the outages mailing list: https://puck.nether.net/mailman/listinfo/outages https://puck.nether.net/mailman/listinfo/outages Public archives: https://puck.nether.net/pipermail/outages/ https://puck.nether.net/pipermail/outages/ Latest issue reported: https://puck.nether.net/pipermail/outages/2020-August/013187.html https://puck.nether.net/pipermail/outages/2020-August/013187... "Level3 (globally?) impacted (IPv4 only)"
- svdr 6y agoI go here :-)
- cultureulterior 6y agohttps://www.thousandeyes.com/outages https://www.thousandeyes.com/outages
- tuukkah 6y agoBased on that map, Telia seems to be one of the most affected which might explain why Scandinavia is so badly affected.
- dexterdog 6y agoYou just imagined the first target in an attack. Might as well just call it honeypotnumber1.
- exikyut 6y agoThis is an excellent idea and simple but moderately expensive for anyone to set up. Just have a site fetch resources from every single hosting provider everywhere. A 1x1 image would be enough, but 1K/100K/1M sized files might also be useful (they could also be crafted images) The first step would be making the HTML page itself redundant. Strict round robin DNS might work well for that. But yeah, moderately expensive - and... thinking about it... it'll honestly come in handy once every ten years? :/
- aeyes 6y agoCan confirm for a number of sites, even Hacker News was unreachable for me.
- rewtraw 6y agoReddit, HN, etc. are inaccessible to me over my Spectrum fiber connection, but working on AT&T 4G. It’s not DNS, so a tier 1 ISP routing issue seems to be the most likely cause.
- kossTKR 6y agoLots of local sites not working in Scandinavia either. So seems more global than a single Tier 1?
- phoe-krk 6y agoProbably relevant Fastly update: > Fastly is observing increased errors and latency across multiple regions due to a common IP transit provider experiencing a widespread event. Fastly is actively working on re-routing traffic in affected regions.
- dbetteridge 6y agoHN and reddit out on my talktalk link in London, 3 mobile 4g working normally.
- one2know 6y agoBased on twitter, the outage was on multiple continents. What would cause that? Subsea cable broken?
- tambre 6y agoFastly is also seeing problems. [0] However, they report that they've identified the issue and are fixing it. [0]: https://status.fastly.com/ https://status.fastly.com/
- haunter 6y agoI'm in Hungary EU. My fiber works fine but 4G gone except for domestic addresses can't connect to anything
- stordoff 6y agoIt wasn't a total outage for the site I was trying to reach. It took about 20 minutes to make an order, but after multiple retries (errors were reported as a 522 with the problem being somewhere between Manchester, UK and the host), it did go through.
- vbsteven 6y agoI'm having lots of issues with Hetzner machines not being available (and even the hetzner.com website). Don't know if this is related.
- zepearl 6y agoFyi I'm not having any problems right now with hetzner.com nor hetzner.de - my own dedicated server hosted at Hetzner datacenter in Germany seems to be reachable/working as well. Connecting from Switzerland.
- deleted 6y ago[deleted]
- system2 6y agoI am having trouble with Hulu right now. I bet it is related.
- suby 6y agoI was doing development work which uses a server I've got hosted on digital ocean. I started getting intermittent responses which I thought weird as I hadn't changed anything on the server. I spent a good ten minutes trying to debug the issue before searching for something on duckduckgo, which also didn't respond. Cloudfare shouldn't be involved at all with my little site, so I don't think it's limited to just them.
- one2know 6y agoYeah, something happened to ipv4 traffic worldwide. Don't see how that could happen.
- pps43 6y agoLet me guess: somebody misconfigured BGP again?
- ra 6y agolikely
- johnisgood 6y agohttps://puck.nether.net/pipermail/outages/2020-August/013198.html https://puck.nether.net/pipermail/outages/2020-August/013198...
- Sebb767 6y agoThat's definitely going to be an interesting postmortem.
- opan 6y agoSeconding this. Had some ssh connections timing out repeatedly just a bit ago. Also got disconnected on IRC.
- mikegioia 6y agoMe too. I can only connect to one of my DO servers. The rest are all unreachable.
- ffpip 6y agoDDG, down detector are all very slow. Both are on cloudflare. Fastly, HN, Reddit too. Only Google domains are loading here.
- deleted 6y ago[deleted]
- thejteam 6y agoFrom where I am (mid-altantic US) Google site are completely down (google.com, youtube)
- rantanplan 6y agoIncidentally I can't connect to HN directly from Greece, but only if I use my VPN through New York. Probably somehow related?
- janmo 6y agoThere is a major internet outage going on. I am using Scaleway they are also affected. According to Twitter, Vodafone, CityLink and many more are also affected.
- mikiem 6y agoM5 Hosting here, where this site is hosted. We just shut down 2 sessions with Level3/CenturyLink because the sessions were flapping and we were not getting complete full route table from either session. There are definitely other issues going on on the Internet right now.
- exikyut 6y agoOooh, maybe that's why HN wasn't working for me a little while ago (from AU)...
- chkaloon 6y agoWonder if that's that why Feedly is down
- pinkano 6y agoYes
- blantonl 6y agoEverything to Oracle Cloud's Ashburn US-East location is down. Their console isn't responding at all and all my servers are unreachable. Their status console reports all normal though.
- system2 6y agoStatus pages of the companies are just PR disasters for them. Most of the time they don't report what's up.
- _fool 6y agoBut when they do, it can be amazing. https://www.atlassian.com/blog/statuspage/how-to-write-a-good-status-update https://www.atlassian.com/blog/statuspage/how-to-write-a-goo...
- eric_khun 6y agoShameless plug: I spent too much time losing precious time when github/npm/cloudflare are going down, until I figure out it was them. So currently working on a project[1] to monitor all the 3rd party stack you use for your services. Hit me up if you want, access I'll give free access for a year+ to some folks to get feedbacks. [1] https://monitory.io https://monitory.io
- reimertz 6y agoFYI: Your site is down because of GitHub pages maintenance. Edit: it’s up again! Just want to let you know about the spelling error ”Save titme” :)
- naavis 6y agoMaybe fix this typo? "Save titme on issues investigation"
- kzrdude 6y agoAnd > Monitor all 3rd parties services *3rd party services or possibly 3rd parties' services
- vladvasiliu 6y agoAnother typo: > Know when services you depend on goes down "Services go down", not "goes".
- 6y ago
- bovermyer 6y agoSalesForce/Office365 is also having trouble.
- Yetanfou 6y agoOdd, I'm trying to reach a host in Germany (AS34432) from Sweden but get rerouted Stockholm-Hamburg-Amsterdam-London-Paris-London-Atlanta-São Paulo after which the packets disappear down a black hole. All routing problems occur within Cogentco. 3 sth-cr2.link.netatonce.net (85.195.62.158) 4 te0-2-1-8.rcr51.b038034-0.sto03.atlas.cogentco.com 5 be3530.ccr21.sto03.atlas.cogentco.com (130.117.2.93) 6 be2282.ccr42.ham01.atlas.cogentco.com (154.54.72.105) 7 be2815.ccr41.ams03.atlas.cogentco.com (154.54.38.205) 8 be12194.ccr41.lon13.atlas.cogentco.com (154.54.56.93) 9 be12497.ccr41.par01.atlas.cogentco.com (154.54.56.130) 10 be2315.ccr31.bio02.atlas.cogentco.com (154.54.61.113) 11 be2113.ccr42.atl01.atlas.cogentco.com (154.54.24.222) 12 be2112.ccr41.atl01.atlas.cogentco.com (154.54.7.158) 13 be2027.ccr22.mia03.atlas.cogentco.com (154.54.86.206) 14 be2025.ccr22.mia03.atlas.cogentco.com (154.54.47.230) 15 * level3.mia03.atlas.cogentco.com (154.54.10.58) 16 * * * 17 * * *
- cotillion 6y agoWhat seems to have happened is that Centurylinks internal routing has collapsed in some way. But they're still announcing all routes and they don't stop announcing routes when other ISPs tag their routes not to be exported by Centurylink. So as other providers shut down their links to Centurylink to save themselves the outgoing packets towards centurylink travel to some part of the world where links are not shut down yet.
- Cyphase 6y agoI just experienced HN down for several minutes before it loaded and I saw this story at the top. I'm doing something with the HN API as I type this, so for a moment I was trying to decide if I'd been IP blocked, even though the API is hosted by Firebase. I haven't noticed any obvious issues elsewhere yet. (Just got a delay while trying to submit this comment.)
- Darmody 6y agoHalf of the internet is down. Crazy... I can't even access the private WoW server I play.
- tc313 6y agoFWIW, I can’t connect to Madden NFL online servers.
- johnchristopher 6y agoSo, that's why HN is unreachable from Belgium at the moment (right when I was trying to figure a dns cache problem in Firefox,of course). An ssh tunnel through OVH/gravelines is working so far. edit: Proximus. edit2: also, Orange Mobile
- iso1210 6y agoHN working for me from the UK on BT, but traceroute showing lots of different bouncing around and a lot of different hops in the US 7 166-49-209-132.gia.bt.net (166.49.209.132) 9.877 ms 8.929 ms 166-49-209-131.gia.bt.net (166.49.209.131) 8.975 ms 8 166-49-209-131.gia.bt.net (166.49.209.131) 8.645 ms 10.323 ms 10.434 ms 9 be12497.ccr41.par01.atlas.cogentco.com (154.54.56.130) 95.018 ms be3487.ccr41.lon13.atlas.cogentco.com (154.54.60.5) 7.627 ms be12497.ccr41.par01.atlas.cogentco.com (154.54.56.130) 102.570 ms 10 be3627.ccr41.jfk02.atlas.cogentco.com (66.28.4.197) 89.867 ms be12497.ccr41.par01.atlas.cogentco.com (154.54.56.130) 101.469 ms 101.655 ms 11 be2806.ccr41.dca01.atlas.cogentco.com (154.54.40.106) 103.990 ms 93.885 ms be3627.ccr41.jfk02.atlas.cogentco.com (66.28.4.197) 97.525 ms 12 be2112.ccr41.atl01.atlas.cogentco.com (154.54.7.158) 106.027 ms be2806.ccr41.dca01.atlas.cogentco.com (154.54.40.106) 98.149 ms 97.866 ms 13 be2687.ccr41.iah01.atlas.cogentco.com (154.54.28.70) 120.558 ms 122.330 ms 120.071 ms 14 be2687.ccr41.iah01.atlas.cogentco.com (154.54.28.70) 123.662 ms be2927.ccr21.elp01.atlas.cogentco.com (154.54.29.222) 128.351 ms be2687.ccr41.iah01.atlas.cogentco.com (154.54.28.70) 120.746 ms 15 be2929.ccr31.phx01.atlas.cogentco.com (154.54.42.65) 145.939 ms 137.652 ms be2927.ccr21.elp01.atlas.cogentco.com (154.54.29.222) 128.043 ms 16 be2930.ccr32.phx01.atlas.cogentco.com (154.54.42.77) 150.015 ms be2940.rcr51.san01.atlas.cogentco.com (154.54.6.121) 152.793 ms 152.720 ms 17 be2941.rcr52.san01.atlas.cogentco.com (154.54.41.33) 152.881 ms te0-0-2-0.rcr11.san03.atlas.cogentco.com (154.54.82.66) 153.452 ms be2941.rcr52.san01.atlas.cogentco.com (154.54.41.33) 152.054 ms 18 te0-0-2-0.rcr12.san03.atlas.cogentco.com (154.54.82.70) 162.835 ms te0-0-2-0.nr11.b006590-1.san03.atlas.cogentco.com (154.24.18.190) 146.643 ms te0-0-2-0.rcr12.san03.atlas.cogentco.com (154.54.82.70) 153.714 ms 19 te0-0-2-0.nr11.b006590-1.san03.atlas.cogentco.com (154.24.18.190) 151.212 ms 145.735 ms 38.96.10.250 (38.96.10.250) 147.092 ms 20 38.96.10.250 (38.96.10.250) 149.413 ms * *
- karpolan 6y agoDeployment to Netlify fails on installing of any version of Node :)
- _fool 6y agomore specifically, npmjs.com and nodejs.org are not available from Netlify's datacenter due to this outage.
- iso1210 6y agoLooks like Centurylink/Level3 (as3356) might not be withdrawing routes after people close their peering?
- josephb 6y agoThat's what various networks have reported. It kind of makes it hard to route around an upstream, if they keep announcing your routes even when there isn't a path to you!
- swinglock 6y agoQuick hack; split all your announcements in two, making the new announcement route around their old stale announcement by being more specific.
- regolithori 6y agoWhat could cause this? I wonder what the technical problem is.
- crizzlenizzle 6y agoWe had this once with one of our former ISPs configuring static routes towards us and announcing them to a couple of IXPs. I have no idea why they did it, but it caused a major downtime once for us and basically signed the termination.
- jcims 6y agoI would love to hear the inside scoop from folks working at CenturyLink. I’ve used their DSL for years and the network is a mess. I don’t know if it them here or legacy Level3 but i have a guess. Edit: Looks like i would have guessed wrong :P. Still want that inside scoop!
- iso1210 6y agoUsed level3 IP for a long time professionally with limited issues, ceratainly not on the list of worst ISPs. Also used a company that over the years has gone from Genesis, GlobalCrossing, Vyvx, Level3 and now of course Level 3 is CenturyLink, which has been fine.
- nottorp 6y agoI have two pipes from two different (consumer ISPs) at home. One can reach HN, the other can't. Incidentally, uBlock Origin seems to be completely broken. It doesn't have any local blacklists to work when their ?servers? are unavailable?
- TreeInBuxton 6y agoLooks like an issue with AS3356, they are advertising stale routes - lots of unrelated services impacted
- tiernano 6y ago1.1.1.1 warp is having issues too...
- xyst 6y agoInternet infrastructure is broken. Why do a few companies control the backbone of the internet? Shouldn’t there be a fallback or disaster recovery plan if one or more of these companies become unavailable?
- kzrdude 6y agoWhy doesn't stuff just route around this automatically, if one provider has problems?
- q3k 6y agoThings mostly routed around the problem. Issues arose because a) some people are single-homed to Level3/CenturyLink b) apparently Level3/CenturyLink continued announcing unreachable prefixes, which breaks the Internet BGP trust model.
- johncolanduoni 6y agoThe problem is the provider having problems is still sending misconfigured routes after the other providers have tried to pull them in response to the outage. So it’s as if CenturyLink was doing a massive BGP attack against their peers, pointing at a black hole.
- ezconnect 6y agoNamecheap is also having network connection issues.
- corford 6y agoNo impact here in Lisbon, PT (using MEO). I can access: HN, twitter, cloudflare, AWS, DO, Hetzner, DDG, Scaleway etc.
- tictok4 6y agoThis (thread) explains why we've been having internet problems this morning.... lots of sites not working.
- wiremine 6y agoIn a situation like this, what are the best "status" sites to be watching?
- ojagodzinski 6y agohttps://downdetector.com/ https://downdetector.com/ client perspective is best perspective ;) Problem in this outage is that site X works ok but transit provider for clients in US works badly and generates "false positives"
- dkdk8283 6y agoNanog is also pretty helpful for this specific type of issue
- lucb1e 6y agoYou mean nanog.org? I don't see a stats page linked in their menu.
- toomuchtodo 6y agoIt’s a mailing list for network operations/engineering folks. The emails are the status updates. You’ll have to look to each network’s own site if you want connectivity, peering, and IXP red/green/ up/down status.
- zowanet 6y agoHere's a direct link to this month's messages: https://mailman.nanog.org/pipermail/nanog/2020-August/thread.html https://mailman.nanog.org/pipermail/nanog/2020-August/thread...
- tn890 6y agohttps://www.internetweathermap.com/map https://www.internetweathermap.com/map
- tillinghast 6y ago
- vinni2 6y agoI had to use a VPN With US location to post this comment. I am in Europe.
- lucb1e 6y agoHN works fine from Germany with Telefonica (O2) and also from the Netherlands with XS4ALL. Edit: Somewhere between 14:00 and 14:46Z it also went down from O2; XS4ALL still works, and O2 can reach XS4ALL.
- minxomat 6y agoNo luck on T-Mobile
- omnibrain 6y agoYes, I had to switch to my Vodafone eSIM for data to connect to Hacker News.
- crizzlenizzle 6y agoYup. ``` Prefix 209.216.230.0/24 BGP as_path 3356 21581 21581 ``` As seen from AS3320.
- minxomat 6y agoEven NordVPN to the nearest German hub is screwed. Have to vpn to the US to access HN.
- lucb1e 6y agoI see a lot of ads for NordVPN, but you should know they're not necessarily reliable. Just look for NordVPN on hacker news search: https://hn.algolia.com/?dateRange=all&page=0&prefix=true&query=nordvpn&sort=byPopularity&type=story https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que... (see e.g. the second hit: https://news.ycombinator.com/item?id=21664692 https://news.ycombinator.com/item?id=21664692 covering up security issues, using your connection to proxy other people's traffic, a related company does data mining...). The only VPN that seemed to fit the bill when I looked for one about a year ago was ProtonVPN, but I certainly didn't manage to look at every VPN on the planet and I'm just a random internet stranger so... take that with a grain of salt.
- iso1210 6y agoNo peering problems from my network with Level3 in London Telehouse West, maybe a minute or so of increased latency at 10:09 GMT Routing to a level3 ISP I have an office in in the states peers with London15.Level3.net No problem to my Cogent ISP in the states, although we don't peer directly with Cogent, that bounces via Telia Going east from London, a 10 second outage at 12:28:42 GMT on a route that runs from me, level3, tata in India, but no rerouting.
- tpmx 6y agoFrom the other (Cloudflare) thread (post: https://news.ycombinator.com/item?id=24322603 https://news.ycombinator.com/item?id=24322603), the outages list (https://puck.nether.net/mailman/listinfo/outages https://puck.nether.net/mailman/listinfo/outages). https://puck.nether.net/pipermail/outages/2020-August/thread.html https://puck.nether.net/pipermail/outages/2020-August/thread... Not a network engineer, but based on the comments there it looks like it's a BGP blackhole incident. Edit: removed details about the similarity to a 1997 incident based in input from commenters.
- jsjohnst 6y ago> Not a network engineer, but based on the comments there it looks like it's a BGP blackhole incident, possibly reminiscent of the https://en.wikipedia.org/wiki/AS_7007_incident https://en.wikipedia.org/wiki/AS_7007_incident in 1997. As you aren’t a network engineer, I can understand making that leap based on the context, but no, this is nothing like the AS7007 event. The “black hole” in this case is due to networks pulling their routes via AS3356 to try and avoid their outage, but when they do, CenturyLink is still announcing those routes and as such those networks blackhole.
- Benjamin_Dobell 6y agoThis explains a lot. Initially thought my mobile phone Internet connectivity was flakey because I couldn't access HN here in Australia, whilst it's fine over wi-fi (wired Internet).
- abhishekjha 6y agoIts reverse for me. The broadband fails to connect to HN but my mobile ISP is able to reach it fine.
- josephb 6y agoBecause networks are connected to others via different paths, it's not unusual that one method of connectivity would work and one doesn't. Also the Internet has lots of asymmetric traffic, just because a forward path towards a destination may look the same from different networks, it doesn't mean the reverse path will be similar.
- willis936 6y agoSame for me in midwest US. I first thought I had broken my DNS filter again through regular maintenance updates, then I suspected my ISP/modem because it regularly goes out. I have never seen the behavior I saw this morning: some sites failing to resolve.
- bmlzootown 6y agoI thought Cloudflare was having issues again, since I use their DNS servers, so I started by changing that. Then I tried restarting everything, modem/router/computer. Wasn't until I connected to a VM that a friend hosts that I was finally able to access HN, and thus saw this thread. Hopefully this will get fixed within a reasonable timespan.
- every 6y agoycombinator.com pinged just fine but news.ycombinator.com dropped 100% packets. But all better now...
- 6y ago
- nurettin 6y agoalso related https://www.cloudflarestatus.com/incidents/hptvkprkvp23 https://www.cloudflarestatus.com/incidents/hptvkprkvp23
- ramshanker 6y agoI hope these kind of “ipv4” only outages encourages more and more websites to upgrade to ipv6. #OutageBenefit ;)
- eskaytwo 6y agoIt would appear from the limited info so far, to be an issue in the v4 routing configuration - I haven’t seen anything that says this couldn’t have been the other way around.
- cuu508 6y agoSadly, in my experience, ipv4 is generally more reliable than ipv6 still. Set up two hosts, host A and host B in two different data centers. Make them send HTTP requests to each other over ipv4 and over ipv6. You'll see that latency spikes, packet loss is more frequent over ipv6.
- bigdict 6y agoWhy is that?
- chaboud 6y agoWe’ve observed this in end-user devices, especially on some ISPs. It makes sense if the overall adoption and resource allocation are comparatively smaller, making individual or small-group coincident spikes more impactful against the amortized whole. It’s a lot like a market with low volume/liquidity. Someone wanders in with a big transaction and blows everything up.
- iso1210 6y agoFew people care if ipv6 breaks so it doesn't make headlines
- itguy4321365 6y agoThis doesn't have anything to do with IPv4 vs IPv6. It is a routing issue with BGP. To give an analogy, if every website were a house, and every house has a house number (IP address-- either IPv4 or IPv6), and a group of houses form cities and towns that can be identified by a number (AS/ Atonomous System number), the highways between cities are similar to BGP routes, and if half of the world's internet traffic goes through the city of Centurylink (AS3356), If the city of CenturyLink (AS3356) shut down traffic, either on purpose or on accident. ...then it doesn't matter if your house number / IP address is a 32bit number or a 128bit number because traffic needs to take a different route. This is what everyone is worried about BGP routes, not IP addresses.
- mikro2nd 6y agoHad to laugh: "I'm seeing complaints from all over the planet on Twitter" The one site I can't see is Twitter. (Not a heart-wrenching loss, mind you...)
- quickthrower2 6y agoI could not get on HN as a logged in person (logged out was OK) during this. I wondered how big the cloudflare thread would be if people could get on to comment on it :-)
- blooalien 6y agoI experienced this issue while reading docs at "Read the Docs" (and ironically had connection issues while trying to read this very exact page right here, too.)
- EE84M3i 6y agoI'm confused about why Cloudflare had problems but other CDN providers/sites with private CDNs like Google did not. Is there something different about how Cloudflare operates?
- danecek099 6y agoEven https://downdetector.com/ https://downdetector.com/ has problems loading for me. Middle Europe *internetweathermap is down
- neuronic 6y agoWho watches the Watchmen...
- danecek099 6y agoBroadband here just fell down for few minutes, mobile ISP's are ok
- MLij 6y agoIndia just lost to Russia in the final of the firstever online chess olympiad, probably due to connection issues of two of its players. I wonder if it's related to this incident and if the organizers are aware. Edit: the organizers are aware, and Russia and India have now been declared joint winner.
- redwood 6y agoInteresting. How would connection issues cause them to lose? Was it a timed round?
- MLij 6y agoYes, two players lost on time.
- cyphar 6y agoAll professional chess games have a time limit for each player (if you've ever heard of "chess clocks" -- that's what they're used for). In "slow chess" each player has a 2-hour limit and all of the other time control schemes (such as rapid and blitz) are much shorter.
- hinkley 6y agoThere’s an interesting protocol for splitting a Go or chess game over multiple days so that neither party has the entire time to think about their response to the last move: at the end of the day the final move is made by one player but is sealed, not to be revealed until the start of the next session. For this to work on an internet competition, the judges would need a backup, possibly very low bandwidth communication mechanism that survives a network outage. This wouldn’t save any real-time esports, but would be serviceable for turn based systems.
- FR10 6y agoYes, this is call Adjournment[0] and they used to do it until 20 or so years ago when computer analysis became too good/mainstream. [0] https://en.wikipedia.org/wiki/Adjournment_(games) https://en.wikipedia.org/wiki/Adjournment_(games)
- redwood 6y agoCould this be a Russia move vis a vis today's expected Belarus protests? (I hope this doesn't mean a violent crackdown is imminent) Oy https://mobile.twitter.com/HannaLiubakova/status/1300064535697555456 https://mobile.twitter.com/HannaLiubakova/status/13000645356...
- badrabbit 6y agoI don't see any bgpmon alerts, that's unlikely.
- 2fast4you 6y agoCenturylink is my isp, it looks like traffic drops out after 2 hops. It’s been this way for a few hours
- 2fast4you 6y agoYoutube is still trucking though, not sure how that works
- t0mas88 6y agoThey have servers inside a lot of ISPs. Same for Netflix.
- adamcharnock 6y agoThey probably peer into Google at the local IX/Data centre. Google traffic will therefore take a different path which isn’t suffering the current outage.
- badrabbit 6y agoYoutube colocates at most major ISPs on the planet, that might help.
- _eigenfoo 6y agoCan somebody please clarify - what exactly is this an outage of, and how serious is it?
- jsjohnst 6y agotl;dr One of the large Internet backbone providers (formerly known as Level3, but now known as CenturyLink usually) that many ISPs use is down. Expect issues connecting to portions of the Internet. Usually the Internet is a bit more resilient to these kinds of things, but there are complicating factors with this outage making it worse. Expect it to mostly be resolved today. These things have happened a bit more frequently, but generally average up to a couple times a year historically.
- rmrfstar 6y agoHere is a fantastic, though somewhat outdated overview [1]. Section 5 is most relevant to your question. The network topology today is a little different. Think of Level3 as an NSP, which is now called a "Tier 1 network" [2]. The diagram should show links among the Tier 1 networks ("peering"), but does not. [1] https://web.stanford.edu/class/msande91si/www-spr04/readings/week1/InternetWhitepaper.htm https://web.stanford.edu/class/msande91si/www-spr04/readings... [2] https://en.wikipedia.org/wiki/Tier_1_network https://en.wikipedia.org/wiki/Tier_1_network
- g105b 6y agoIs this affecting all geographic regions?
- dredmorbius 6y agoUS, Europe, and Asia that I'm aware of (NANOG mailing list).
- kitteh 6y agoMassive reconvergence event in their network, causing edge router bgp sessions to bounce (due to cpu). Right now all their big peers are shutting down sessions with them to give level3s network the ability to reconverge. Prefixes announced to 3356 are frozen on their route reflectors and not getting withdrawn. Edit: if you are a Level3 customer shut your sessions down to them.
- deleted 6y ago[deleted]
- kitteh 6y agoMost of level3s settlement free peers aka "tier 1s" have shutdown or depreffed their sessions with them. Example: https://mobile.twitter.com/TeliaCarrier/status/1300074378378518528 https://mobile.twitter.com/TeliaCarrier/status/1300074378378...
- beagle3 6y agoHistory doesn't repeat, but it rhymes .... There was a huge AT&T outage in 1990 that cut off most US long distance telephony (which was, at the time, mostly "everything not within the same area code"). It was a bug. It wasn't a reconvergence event, but it was a distant cousin: Something would cause a crash; exchanges would offload that something to other exchanges, causing them to crash -- but with enough time for the original exchange to come back up, receive the crashy event back, and crash again. The whole network was full of nodes crashing, causing their peers to crash, ad infinitum. In order to bring the network back up, they needed to either take everything down at the same time (and make sure all the queues are emptied), but even that wouldn't have made it stable, because a similar "patient 0" event would have brought the whole network down. Once the problem was understood, they reverted to an earlier version which didn't have the bug, and the network re-stabilized. The lore I grew up on is that this specific event was very significant in pushing and funding research into robust distributed systems, of which the best known result is Erlang and its ecosystem - originally built, and still mostly used, to make sure that phone exchanges don't break. [0] https://users.csc.calpoly.edu/~jdalbey/SWE/Papers/att_collapse https://users.csc.calpoly.edu/~jdalbey/SWE/Papers/att_collap...
- pgoodjohn 6y agoPressing F for everyone else who was on call today
- jetru 6y agoOh lord. I'm oncall and we were like "WHATS HAPPENING"
- b3lvedere 6y agoSame here :) Couple of companies started complaining. Told them it's a worldwide issue. It seems going better at the moment.
- tyfon 6y agoSeems like "the internet" works again here in Norway. I've been limited to local sites all day. Hacker news has been off for several hours for me. Whatever it was it must have been nasty.
- deleted 6y ago[deleted]
- naringas 6y agoseems like the internet in 2020 has a diminished ability to route around damage
- Frost1x 6y agoMy opinion is that this, like many issues were seeing today, is largely an issue with ongoing consolidation trends. The less diversity of systems/solutions we have for a given problem or set if problems, the less chance we're protected from unknown unknowns that creep up. The more diversity you have in systems, the more likely you have some option that is hardened against unknown unknowns when they arrive and the quicker we can work around them. Modern society is all about consolidating systems into a few efficient solutions typically dictated by market forces which I argue, don't concern themselves much with these sorts of problems. As a result, when we run into problems, we're left with fewer options to resort to and instead have to identify problems and develop new solutions on-the-fly. Consolidation leads to complacency and stagnation. Sometimes this is reasonable (and even desirable) for certain non-critical systems, it just doesn't make financial sense to pour resources into system diversity for certain systems we could do without--find the one that works best/most efficiently and use it. If it breaks, it's not critical and the work around can wait. On the other hand, if a system is critical, then I think it behooves us to continue looking at improvements of existing systems and alloting resources to investigating new approaches.
- sp332 6y agoBGP has always had this issue. It depends on trustworthy information being available. Any trusted source who starts lying (or just screws up) is going to cause routing problems.
- salawat 6y agoNote, trustworthyness jumps off of being a technical problem, and becoming a human/people problem. Level 8 as someone mentioned, or GIGO (Garbage-In-Garbage-Out) as others may know it. To safely use a system, your operator needs to be 10% smarter than the system being operated. It is clear that we have problems in that department with certain AS's. This is about, what the third major outage attributed to CenturyLink in the last handful of years? I have no idea what exactly their process must look like, but good heavens, a better look need be taken, as this is becoming a bit regular for my tastes.
- deleted 6y ago[deleted]
- hkc 6y agoChess.com was down due to the outage and some of the Indian players got disconnected and lost on time, so FIDE declared India-Russia joint winner of the Online Chess Olympiad 2020.
- eatmyshorts 6y agoI was doing a big release over the evening. I was working fine up until about 6 hours ago, when I signed off. Our network monitors show an outage started about half an hour later (at about 4:05am CST). Service restored a few minutes ago, at about 9:44am CST. I don't know if our problem is the same as this problem, but we are on CenturyLink.
- emilstahl 6y agoCloudflare status page: Update - Major transit providers are taking action to work around the network that is experiencing issues and affecting global traffic. We are applying corrective action in our data centers as the situation changes in order to improve reachability Aug 30, 14:26 UTC https://www.cloudflarestatus.com https://www.cloudflarestatus.com
- dathinab 6y agoI guess that why HN was temporary unreachable from my home?
- protomyth 6y agoand why Cloudflare was having so many issues https://www.cloudflarestatus.com/ https://www.cloudflarestatus.com/
- aosaigh 6y agoHas anyone any good resources for learning more about the "internet-level" infrastructure affected today and how global networks are connected?
- marmshallow 6y ago+1 Many of the comments here presume knowledge about this stuff, and I can’t follow.
- albertTJames 6y agohttps://mobile.twitter.com/Level3 https://mobile.twitter.com/Level3 (not an internet level, just a company :)
- Vinnl 6y agoWere their tweets protected (i.e. only visible to approved followers) when you posted that link, or is that in response to this event?
- iso947 6y agoLevel3 was qquired/merged/changed to century link a year or so back, I think they closed their old twitter account then When someone says level3, read century link. L3 have been a major player for decades though (including providing the infamous 4.2.2.2 dns server), so people still refer to them as level3. The account to follow for them now is https://mobile.twitter.com/CenturyLink https://mobile.twitter.com/CenturyLink but it won’t tell you much.
- Godel_unicode 6y agoNote that L3 is a separate company from Level 3 Communications, which was the ISP that was acquired by CenturyLink. L3 is an American aerospace and C4ISR contractor. CenturyLink's current CEO, Jeff Storey, was actually the pre-acquisition Level 3 CEO.
- vesh 6y ago
- dredmorbius 6y agoNANOG are talking about a CenturyLink outage and BGP flapping (AS 3356) as of 03:00 US/Pacific, AS209 possibly also affected. AS3356 is Level 3, AS209 is CenturyLink. https://mailman.nanog.org/pipermail/nanog/2020-August/209359.html https://mailman.nanog.org/pipermail/nanog/2020-August/209359...
- gailees 6y agoThe beginning of WWIII probably looks something like this.
- based2 6y agohttps://status.ctl.io/history/f19a0555-abbd-4038-91cb-b55a7645c1f5 https://status.ctl.io/history/f19a0555-abbd-4038-91cb-b55a76... https://twitter.com/g_bonfiglio/status/1300022993251446785?s=19 https://twitter.com/g_bonfiglio/status/1300022993251446785?s... https://old.reddit.com/r/networking/comments/ijb8tn/global_as3356_level3_outages/ https://old.reddit.com/r/networking/comments/ijb8tn/global_a...
- emilstahl 6y agoCNN just blames Cloudflare.. :facepalm: https://edition.cnn.com/2020/08/30/tech/internet-outage-cloudflare/index.html https://edition.cnn.com/2020/08/30/tech/internet-outage-clou...
- ihatecloudflare 6y agoCNN is absolutely right. Every day I read news that something goes down at CloudFlare. CloudFlare do much more harm than they "fix" with their services.
- emilstahl 6y agoCenturyLink/Level3 on Twitter: "We are able to confirm that all services impacted by today’s IP outage have been restored. We understand how important these services are to our customers, and we sincerely apologize for the impact this outage caused." https://twitter.com/CenturyLink/status/1300089110858797063 https://twitter.com/CenturyLink/status/1300089110858797063
- ystad 6y agoI hope they provide a root cause analysis
- colde 6y agoBased on experience it will probably not public, or at least very limited. But customers are likely to get one, at least if they request it.
- rootsudo 6y agoBeing it was pretty big, they'll probably make it public.
- deleted 6y ago[deleted]
- jlgaddis 6y agohttps://news.ycombinator.com/item?id=24324280 https://news.ycombinator.com/item?id=24324280
- gnyman 6y agoThis had me really confused until I saw it was a global outage. I have been getting delayed iOS push notifications (from prowl) now for the last few hours, from a device I was fairly sure I had disconnected 3 hours ago (a pump) Got questioning if I really disconnected it before I left. I'm wondering if we're at the point where internet outages should have some kind of (emergency) notification/sms sent to _everyone_.
- tpmx 6y agoHow the xxxx did it take CenturyLink/Level3 like 3-4 hours to fix this problem? Again (https://news.ycombinator.com/item?id=24322988 https://news.ycombinator.com/item?id=24322988) not a network engineer, but it seemed like their routers actively stopped other networks from working around the problem since L3 would still keep pushing other networks' old routes, even after those networks tried to stop that. Also: BGP probably needs to redesigned from the ground up by software engineers with experience from designing systems that can remain working with hostile actors.
- q3k 6y ago> Also: BGP probably needs to redesigned from the ground up by software engineers with experience from designing systems that can remain working with hostile actors. This has been attempted a number of times, but this is a political problem, not a technical problem: there's no single agreed source of truth for routing policy. A lot of US Internet providers won't even sign up for ARIN IRR, or even move their legacy space to a RIR - so there isn't even any technical way of figuring out address space ownership and cryptographic trust (ie. via RPKI). Hell, some non-RIR IRRs (like irr.net) are pretty much the fanfiction.net equivalent of IRRs, with anyone being able to write any record about ownership, without any practical verification (just have to pay a fee for write access). And for some address space, these IRRs are the only information about ownership and policy that exists. Without even knowing for sure who a given block belongs to, or who's allowed to announce it, or where, how do you want to fix any issues with a new dynamic routing protocol?
- kitteh 6y agoRPKI is a totally diff problem here, though. If people refuse to sign ROAs, then they don't get protection. The ARIN TAL thing is real and people have to keep fighting that. As it is right now you can xfer v4 out of ARIN but not v6. So even if you wanted to you can't.
- tpmx 6y agoBuild an industry coalition. Put pressure on those who don't join. Randomly throw away 1 out of 10000 packets from the providers that fail to get with the times. Increase that frequency according to some published time function.
- lemiffe 6y agoI had this earlier! A bunch of sites were down for me, I couldn't even connect to this site. The problem is I don't know where to find what was going on (tried looking up live DDOS-tracking websites, "is it down or is it just me" websites, etc. I couldn't find a single place talking about this. Is there a source where you can get instant information on Level3 / global DNS / major outages?
- kitteh 6y agoDdos tracking sites are eye candy and garbage. Stop using them. Outages and nanog lists are your best bet, short of being on the right IRC channels.
- bregma 6y agoMisread the headline as "Level 3 Global Outrage" and thought "someone had defined outrage levels?" and "it doesn't matter, he'll just attribute it to the Deep State". In some ways I'm a little bit disappointed it's only a glitch in the internet.
- deleted 6y ago[deleted]
- osipovas 6y agoA service I run on Digital Ocean was affected by this early this morning. Looks like it was mitigated by DO - so I'm very grateful for that. Although, the service I run is time sensitive so failures like this are pretty unfortunate for me. Where would I get started with building in redundancy against these sort of outages?
- jlgaddis 6y ago> "Root Cause: An offending flowspec announcement prevented BGP from establishing correctly, impacting client services." -- That doesn't really explain the "stuck" routes in their RRs... maybe it'll make sense once we've gotten some more details...
- quickthrower2 6y agoThis might be a silly question but is there such a thing as CI/CD for this sort of thing that may have caught the problem?
- dsr_ 6y agoThere are two aspects to this: 1. Is there syntax correctness checking available, so you don't push a config that breaks machines? Yes. 2. Is there a DWIM check available, so you can see the effect of the change before committing? No. That would require a complete model of, at a minimum, your entire network plus all directly connected networks -- that still wouldn't be complete, but it could catch some errors.
- jsumrall 6y agoThe iDeal payment network used by most online stores of the Netherlands was down/flaky all afternoon.
- rglover 6y agoThis knocked out the Starbucks app and some of their systems this morning. A bunch of people in line couldn't log in and they were saying parts of their whole internal system were down, too.
- skee0083 6y agoGood. It's about time ISP switched to ipv6.
- tpmx 6y agoBased on what I've seen: They essentially "shut down the Internet" for probably a quarter of the global population for about 3-4 hours. That response time is atrocious. It wasn't that they needed to fix broken hardware, rather they needed to stop running hardware from actively sabotaging the global routing via the inherently insecure BGP protocol. That took 3-4 hours to happen. As an example: Being in Sweden with an ISP that uses Telia Carrier for connectivity things started working around the time of https://twitter.com/TeliaCarrier/status/1300074378378518528 https://twitter.com/TeliaCarrier/status/1300074378378518528
- swinglock 6y agoSeems they didn't even get around to doing so, rather asking other carriers to stop peering with them. https://twitter.com/TeliaCarrier/status/1300074378378518528?s=20 https://twitter.com/TeliaCarrier/status/1300074378378518528?...
- deleted 6y ago[deleted]
- matsur 6y agoCenturyLink requested depeering to give them some breathing room and stop the bleeding. Hug ops.
- deleted 6y ago[deleted]
- tpmx 6y agoThat is a fantastic euphemism. Personally I'm disappointed Telia didn't de-peer two hours earlier, after diagnosing the issue for 30 minutes, since that whole lack of functioning routning to very large parts of the internet forced me to use VPN in north america to access many web services, including HN. I realize I'm going to get insanely downvoted by the elite internetworking crowd again but I think this needs to be said. From an outsider's POV: There seems to be a very strange and almost incestual relationship between the networking companies. Or maybe it's just their hangaround supporters? I dunno.
- dz0ny 6y agoSummary: On August 30, 2020 10:04 GMT, CenturyLink identified an issue to be affecting users across multiple markets. The IP Network Operations Center (NOC) was engaged, and initial research identified that an offending flowspec announcement prevented Border Gateway Protocol (BGP) from establishing across multiple elements throughout the CenturyLink Network. The IP NOC deployed a global configuration change to block the offending flowspec announcement, which allowed BGP to begin to correctly establish. As the change propagated through the network, the IP NOC observed all associated service affecting alarms clearing and services returning to a stable state. Source https://puck.nether.net/pipermail/outages/2020-August/013229.html https://puck.nether.net/pipermail/outages/2020-August/013229...
- kitteh 6y agoFlowspec strikes again. Its a super useful tool if you want to blast out an ACL across your network in seconds (using BGP) but it has a number of sharp edges. Several networks, including Cloudflare have learned what it can do. I've seen a few networks basically blackhole traffic or even lock themselves out of routers due to a poorly made Flowspec rules or a bug in the implementation.
- parliament32 6y agoIs "doing what you ask" considered a sharp edge? Network-related tools don't really have safeties, ever (your linux host will happily "ip rule add 0 blackhole" without confirmation). Every case of flowspec shenanigans in the news has been operator error.
- mrguyorama 6y agoIt's possible that if a tool allows you to destroy everything with a single click, that tool (or maybe process) is bad
- tpmx 6y agoObservation: You're not really adjusting your language to the target audience. (Edited away potentially hurtful language.) Learn from Feynman. Explain things using concepts the target audience can be expected to understand. Real mastery of a concept is when you can explain it using simple terms to any other reasonably intelligent person.
- kitteh 6y agoBGP is a routing protocol that is mostly used for propagating routing/reachability information that also includes additional data that can be used (communities as tags, etc). A few years ago folks wanted to bake in additional functionality. For example, packet filters (aka ACLs) normally are deployed to router configuration files using each operators own tooling. To deploy this against hundreds or thousands of routers rapidly was a challenge for them (not good at swdev, etc.). So the idea was we already have a protocol that propagates state to every router rapidly in the network, let's find a way to bake ACLs into the BGP updates. The result wasnt that good for a few reasons: 1) bgp state isn't sticky. If a router goes offline or bgp sessions reset, acls go away. That means if you are using flowspec for a critical need like always on packet filters you've got the wrong tool. 2) the implementation had various bugs. 3) most importantly it gave people a really easy way to hurt themselves globally. There was no phased deployment with pre and post checks. What you deployed led to packet filters being installed across the network in seconds. In most cases (depends on your config) the only way to remove it is remove the specific flowspec route or have bgp reset to it. I've seen bad flowspec routes core dump the daemon on a router responsible for programming ACLs that led to them being unable to withdraw the programmed entry. I've seen as bugs on tcp/UDP port matches go wrong and eat lot more than intended. I've seen so many flowspec rules installed on a network where it exhausted routers ability to inspect and process packets and you'd see flat lining of packets being dropped. In my opinion, it's a hack around not having a good ACL deployment tool that has led to many outages in its wake. Edit: another flowspec gotcha. Some folks like to integrate ddos tooling systems into flowspec. An example of this is if I run a network and some IP address behind me gets lit up, deploy a rule for that specific IP and rate limit traffic to it. Unfortunately, sometimes folks don't put a lot of care into making sure it can't mess with internal IPs that should be off limits. Like route reflectors, router loopback IPs, etc. I've seen situations where some networks have had a bad day due to a ddos or traffic mis classified as ddos by auto installing rules to protect something but actually impair legitimate communications to network infrastructure which then causes the outage. Also, flowspec doesn't work like regular ACLs where you have input and output on a per interface basis - it applies to all traffic traversing a router, which makes it difficult to say which interfaces should be exempt (think internal vs external).
- ausjke 6y agoI wasted two hours for this, diagnosis, reboots,etc.
- eastdakota 6y agoAnalysis of what we saw at Cloudflare, how our systems automatically mitigated the worst of the impact to our customers, and some speculation on what may have gone wrong: https://blog.cloudflare.com/analysis-of-todays-centurylink-level-3-outage/ https://blog.cloudflare.com/analysis-of-todays-centurylink-l...
- ngold 6y agoGreat write up. It is embarrassing that most of America has no competition in the market. >To use the old Internet as a “superhighway” analogy, that’s like only having a single offramp to a town. If the offramp is blocked, then there’s no way to reach the town. This was exacerbated in some cases because CenturyLink/Level(3)’s network was not honoring route withdrawals and continued to advertise routes to networks like Cloudflare’s even after they’d been withdrawn. In the case of customers whose only connectivity to the Internet is via CenturyLink/Level(3), or if CenturyLink/Level(3) continued to announce bad routes after they'd been withdrawn, there was no way for us to reach their applications and they continued to see 522 errors until CenturyLink/Level(3) resolved their issue around 14:30 UTC. The same was a problem on the other (“eyeball”) side of the network. Individuals need to have an onramp onto the Internet’s superhighway. An onramp to the Internet is essentially what your ISP provides. CenturyLink is one of the largest ISPs in the United. Because this outage appeared to take all of the CenturyLink/Level(3) network offline, individuals who are CenturyLink customers would not have been able to reach Cloudflare or any other Internet provider until the issue was resolved. Globally, we saw a 3.5% drop in global traffic during the outage, nearly all of which was due to a nearly complete outage of CenturyLink’s ISP service across the United States.
- throwaway344655 6y agoI remember working the support queue _before_ this automatic re-routing mitigation system went in and it was a lifesaver. Having to run over to SRE and yell "look! look at grafana showing this big jump in 522s across the board for everything originating in ORD-XX where the next hop is ASYYYY! WHY ARE WE STILL SENDING TRAFFIC OVER THAT ARRRGHH please re-route and make the 522 tickets stop" it's cool to see something large enough that the auto-healing mechanisms weren't able to handle it on their own, though shoutout to whoever was on the weekend support/SRE shift; that stuff was never fun to deal with when you were one of a few reduced staff on the weekend shifts
- ihatecloudflare 6y agoIt's probably just another daily outage at CloudFlare, they are famous for their the most unreliable infrastructure on the entire planet.
- ihatecloudflare 6y agoI don't use CloudFlare, so my website has ZERO issues. I'm so glad CloudFlare and their customers are affected so badly. They deserved it.
- gnicholas 6y agoCan anyone help me understand why I can't access HN from my iPhone, but I can from my computer? both are on the same network. I'm getting "Safari cannot open the page because the server cannot be found", and many apps won't work at all either.
- lmm 6y agoOne might be using IPv6 and the other v4. Or you might have different DNS settings.
- deleted 6y ago[deleted]
- person_of_color 6y agoImagine a ransomware attack against these jokers.
- dancemethis 6y agoProbably due to the incredibly ugly name this company has. No one in their right mind should shake hands with a thing called Level 3.
- CarCooler 6y agoYep, internet has been horrible out here, I had to use Cloudflare DNS to reach websites!