7 ms·
“An ISP in Asia is leaking routes to a Tier 1 transit provider”
- _jomo 11y agoApparently this is AS4788 (Telekom Malaysia) leaking routes. https://twitter.com/rrbone_net/status/609282081420652544 https://twitter.com/rrbone_net/status/609282081420652544
- bbrazil 11y agoNanog discussion: http://mailman.nanog.org/pipermail/nanog/2015-June/076187.html http://mailman.nanog.org/pipermail/nanog/2015-June/076187.ht...
- sbarre 11y agoWhat is the impact of this problem? Any TL;DR from an expert on here?
- buro9 11y agoI am no expert. But a route leak forces traffic to take a different path, and this potentially results in a pipe being inundated with a significant volume of traffic that cannot be handled by ISPs along that route (or simply a dead-end). i.e. imagine an ISP in Rio suddenly declaring that they are the best route to reach the networks that contain facebook.com and google.com ... that ISP will DDoS itself or one of their downstream partners. An ISP may not even notice this, but their peers and partners probably will. A couple of years ago a BGP route leak took out the entire internet for a few hours for people in Australia. They're fairly common unfortunately, but tend to have limited impact and are resolved quite quickly. More info: https://labs.apnic.net/?p=139 https://labs.apnic.net/?p=139
- tluyben2 11y agoThanks for asking that; this is the first time I open HN and think ; WUT???
- bbrazil 11y agoPeople cannot access some internet sites, depending on who they have internet connectivity with. https://blog.cloudflare.com/route-leak-incident-on-october-2-2014/ https://blog.cloudflare.com/route-leak-incident-on-october-2... has more information from a previous time this happened. A key idea is that internet routing depends a lot on trust, and it's possible for a single misconfigured site to cause serious issues across the internet.
- hhw 11y agoAny networks that both accept the prefixes and see these advertisements as the best path will send their traffic towards the route leaker, who will lack the capacity to handle this traffic, effectively blackholing these routes. In theory, a large portion of the Internet would have still seen the legitimate routes as better paths, as they would have a shorter AS path (fewer networks in between) than the leaked ones. However, many networks often implement other BGP metrics for traffic engineering, depending on whether they are seeing the routes from a transit/peer/downstream customer, which may override the shortest AS path.
- nissehulth 11y agoRumor has it that a lot of World of Warcraft players left their homes and went outside.
- yebyen 11y agoMy connection through a Sprint hotspot is only able to reach www.google.com and plus.google.com (but not news.google.com, reddit, hn, my own website on the other side of the city I'm in, or practically any other website I can think of.) I got on with tech support this morning and they had no indication that anything was wrong regionally, had me do a factory reset on the device, and verified it was connecting to their network successfully. So, maybe my issue is resolved, but Sprint is still having peering issues. That is the kind of behavior I would expect to see during an issue like this. Of course I go into the store and they tell me there have been regional issues for the last two weeks and it's not my device (we can't help you), but everything works fine for me while I'm inside of the store and on the way back. Then 2 miles before back at the office, broken again, call again and phone support guy hasn't heard anything about this Malaysia Telekom debacle or any regional outages in my area. At this point I notice that I can still reach some sites, but most sites are down. Not completely sure this is related, since I've been having issues since Wednesday afternoon (in fact that would seem to indicate it's unrelated). Also possible I have been having one issue that was fixed by the factory reset, so now I can freely experience the other issue that everyone else around the globe is apparently having today. I am interested to know from The Expert as well, when Telekom Malaysia fixed their issue (it is actually fixed at the source now, right?) would there be some propagation delay or aftershocks, or should everything just return to normal relatively immediately.
- adamlj 11y agoI can't access a couple of US sites (like github, linkedin etc.) from my Swedish ISP Bahnhof.
- radiospiel 11y agoHad the same issue from Berlin. (and probably unrelated to the original post anyways.) But now we are back online!
- deleted 11y ago[deleted]
- tomjen3 11y agoProbably why I couldn't access Stackoverflow from Denmark earlier.
- peterwaller 11y agoThis is rather comic: https://twitter.com/TMCorp/status/609167065300271104 https://twitter.com/TMCorp/status/609167065300271104 (That's the twitter account for the currently implicated provider which messed things up, and for the record it has a minion in a Hawaiian grass dress saying "Happy Friday!", posted about 10 hours ago.)
- deleted 11y ago[deleted]
- mfoy_ 11y agoHaha, that's pretty great. Whoever runs the Twitter account must feel pretty bad about the timing on that...
- nathas 11y agohttp://devopsreactions.tumblr.com/post/87284390953/friday-deployments-and-leaving-afterwards http://devopsreactions.tumblr.com/post/87284390953/friday-de...
- brazzledazzle 11y agoSome of the replies are so unnecessary and trite.
- jameswyse 11y agoLooks like it's party day at Telecom Malaysia. https://twitter.com/TMCorp/status/609167065300271104 https://twitter.com/TMCorp/status/609167065300271104
- deleted 11y ago[deleted]
- johansch 11y agoThey added the route at 16:44 local KL time on a friday afternoon. The network team started partying early? :)
- jessaustin 11y agoThe bumiputra probably shouldn't schedule any deployment for Friday. b^) Then again GLBX probably should know that too.
- ars 11y agoMalaysia is an Islamic country, so Friday is the equivalent of the American Saturday to them, i.e. Friday is the first day of the Weekend, and Sunday is a normal workday.
- atomengine 11y agoI'm an American living in KL. Malaysia follows the Monday through Friday work week. The Twitter message was posted in the AM and this event didn't occur until the evening.
- deleted 11y ago[deleted]
- piyushpr134 11y agodid this happen last night too (about 15 hours from when this post was made)? I am in India and was not able to reach my servers in Singapore or do a git push. All other sites worked okay. I scratched my head and changed my dns to 8.8.8.8. It still did not work.
- davidgerard 11y agoIt's not DNS - it's how the actual packets are supposed to get between you and the server, that got messed up.
- hhw 11y agoTelecom Malaysia (AS4788) was leaking a full routing table to a Tier 1 network, Global Crossing (AS3549), who in turn was advertising the prefixes to its peers like Level3 (AS3356). Large portions of the Internet would have been affected. This is a double fail, both for Telecom Malaysia for leaking a full routing table, and for GBLX who apparently isn't filtering prefixes from their downstream customers or even restricting to a max number of prefixes.
- namecast 11y agoI just remembered, GBLX was bought by L3 a few years back - I'm guessing someone has a "tighten up route maps for GBLX ASN" ticket open in their backlog at the Level 3 NOC. The people managing peering for AS 3356 and AS 3549 should be the same group, no? ("That DB we don't share the URL of" seems to imply as much.)
- deleted 11y ago[deleted]
- peterwaller 11y agoCould someone elaborate on the scope of the problem? Is it a sensible question to ask "How many routes were leaked?" how big were the prefixes of the networks leaked?
- elktea 11y agoFull table was leaked
- lucb1e 11y agoI see only this: > Your IP address 5.79.68.161 has been flagged as a scanner. Scanners are not permitted. If you are seeing this message in error, please contact security@statuspage.io. Guess I better send an e-mail to see this status page..?
- decasteve 11y agoAfter I switched Tor circuits it worked fine for me.
- ryanlol 11y agoFuckups like this should result in criminal charges (and immediate depeering). DoS attacks are illegal in most countries and this is definitely gross negligence.
- alice715 11y agoShit happens. Not that it was intentional. Anyway, who cares about software stuff, nobody dies. Unlike in engineering sector where a negligence could result in injury or death; those could be charged.
- therealmarv 11y agoThis reminds me of arguments some people use in the EU against net neutrality... they want fast lanes for Industry 4.0, online emergency and self driving cars. Incredible thoughts... I know.
- laumars 11y agoThere's a big difference between accidentally breaking something and doing it intentionally. I think some people get too caught up on blaming individuals for accidents. Which I think stems from the culture of litigation and offsetting unforeseen costs. While there obviously should be processes in place to prevent accidents from happening, gross negligence aside I do also think we should stop hunting down individuals to compensate our own greed.
- ryanlol 11y agoPeople are usually still held responsible when they accidentally break things, especially if it's because of their own gross negligence. I don't think anyone is blaming individuals here, in most countries businesses can face criminal charges.
- laumars 11y agoThere hasn't actually been any crime committed though. Unless you want to turn miscoding config files into an international crime - in which case every single person who works in IT would be a criminal since we all error at some point in our careers (usually frequently, to be honest). Granted the scale here is greater than your usual sysadmin would have access to, but the nature of the error isn't any different.
- deleted 11y ago[deleted]
- deepnet 11y agoCan confirm UK ISP was unable to reach arxiv imgur mit or nytimes. Ycombinator and reddit were A OK though.
- pyvpx 11y agothat is because cloudflare has direct peering with many access networks. most providers see cloudflare routes directly and not through a transit such as Level3.
- amalcon 11y agonanog thread: http://seclists.org/nanog/2015/Jun/586 http://seclists.org/nanog/2015/Jun/586 Basically what's happened here is that Telecom Malaysia told one of Level3's networks (AS3549 - Global Crossing) that it was capable of delivering traffic to... anywhere on the Internet. Global Crossing apparently didn't have the usual sanity checks in place. Because of how BGP works, once GBLX decided the route seemed reasonable, it immediately proceeded to dump huge amounts of traffic on the tiny Telecom Malaysia. This isn't incorrect behavior on GLBX's part (accepting the routes was incorrect, but sending the traffic was correct once they were accepted). This is way more traffic than Telecom Malaysia was prepared to handle. Every tier 1 network gets rid of traffic as soon as possible (because it reduces their costs), completely ignoring performance or whether a route seems sensible. Telecom Malaysia therefore claimed to the most attractive route from anywhere near Malaysia to most of the world. Level3 is one of the biggest telecoms in the world -- and in fact the biggest by far in that region. This means that most of the Internet probably stopped working for anyone anywhere near Malaysia.
- mikedb 11y agoIs it actually normal to have some checks for this now? When this happened a few years ago in Australia the apnic writeup (someone has linked below) essentially said there was no scalable solution yet.
- pyvpx 11y agowhen the number of originated or announced routes are "reasonable", prefix lists, as they are called, are generated and applied to BGP sessions. when the originated or announced routes is substantial and/or fluid, there are little to no prefix lists; allowing anything to be announced and accepted. in an ideal world that has not existed for over a decade, IRR data could be used to automagically generate accurate prefix lists but for all the usual reasons (inaccurate, stale data; obtuse interface; etc) that just isn't done on a wide scale. You can put a prefix limit on a session (that's the "little" in little to no) but that doesn't stop someone from announcing 10,000 more specific routes to popular (YouTube, Facebook, Twitter, Akamai) destinations.
- 11y ago
- sumedh 11y agoI was able to open reddit but I was not able to open wikipedia from Australia. So I did a tracert. 6 29 ms syd-sot-ken-int2-be-20.tpgi.com.au [203.219.35.68] 7 197 ms 10ge3-7.core1.sjc1.he.net [72.52.66.21] 8 199 ms 10ge1-4.core1.pao1.he.net [72.52.92.113] 9 * Request timed out.
- ryan-c 11y agoLooks like it made it to California (San Jose and Palo Alto) with what appears to be a reasonable round trip time. Weird. Edit: It seems like that must have come through the Southern Cross Cable[0] which is nowhere near Malaysia. 0. https://en.wikipedia.org/wiki/Southern_Cross_Cable https://en.wikipedia.org/wiki/Southern_Cross_Cable
- h43k3r 11y agoI am very much interested in learning about such issues. Any links or blogs that put light on such issues.
- ethbro 11y agoGoogle has a couple good starting points: http://www.enterprisenetworkingplanet.com/netsp/article.php/3615896/Networking-101-Understanding-BGP-Routing.htm http://www.enterprisenetworkingplanet.com/netsp/article.php/... http://blog.pluralsight.com/bgp-border-gateway-protocol http://blog.pluralsight.com/bgp-border-gateway-protocol
- georgerobinson 11y agoThere is a great chapter on BGP by Hari Balakrishnan. I believe the chapter is available on MIT Open Courseware. It should be under the Computer Networks class. Edit: found it. See Lecture 4: http://ocw.mit.edu/courses/electrical-engineering-and-computer-science/6-829-computer-networks-fall-2002/lecture-notes/ http://ocw.mit.edu/courses/electrical-engineering-and-comput...
- gordon_gecco 11y agohttps://www.tm.com.my/OnlineHelp/Announcement/Pages/INTERNET-SERVICES-DISRUPTION-12-June-2015.aspx https://www.tm.com.my/OnlineHelp/Announcement/Pages/INTERNET...
- ckvamme 11y agoHere is an interactive snapshot of this outage: Capital One uses Level 3 GLBX as a primary ISP: https://oxqyi.share.thousandeyes.com https://oxqyi.share.thousandeyes.com (Jump to BGP route visualization and you can see AS 3549) On the routing plane, you can see that the London GLBX monitor had issues to many services and that AS4788 Malaysia Telecom was advertised in the route. Here is an example from a LinkedIn snapshot: https://wbpkq.share.thousandeyes.com https://wbpkq.share.thousandeyes.com
- prusswan 11y agojust want to ask if there's anything a home user could do to detect unusual routing behavior/phenomenon? I'm envisioning something like a browser plugin that logs/monitors outgoing connections and traceroute data With that, instead of just looking at 404s we can make more informative observations like this request got stuck at this node, or that request is routed through a new node which has never been seen before recently..