9 ms·
The 5-Hour CDN
- legrande 5y agoI like to blog from the raw origin and not use CDNs because if a blogpost is changed I have to manually purge the CDN cache, which can happen a lot. Also CDNs have the caveat in that if they're down, it can make a page load very slow since it tries to load the asset.
- tshaddox 5y agoIf you’re okay with every request having the latency all the way to your origin, you can have the CDN revalidate its cache on every request. Your origin can just check date_updated (or similar) on the blog post to know if the cache is still valid without needing to do any work to look up and render the whole post. To further reduce load and latency to your origin, you can use stale-while-revalidate to allow the CDN to serve stale cache entries for some specified amount of time before requiring a trip to your origin to revalidate.
- cj 5y ago> If you’re okay with every request having the latency all the way to your origin, you can have the CDN revalidate its cache on every request. It's also worth mentioning that even when revalidating on every request (or not caching at all), routing through a CDN can still improve overall latency because the TLS can be terminated at a local origin server, significantly shortening the TLS handshake.
- spondyl 5y agoAh, the TLS shortening aspect of a CDN is something that seems obvious in hindsight but I'd never really thought about it. Thanks!
- dilyevsky 5y agoNot just tls but generally tcp will slowstart faster on lower rtt connection (and edge can keep origin connection always open so it stays “warm”)
- hinkley 5y agoSome of the HTTP/2 and HTTP3 design choices are seen as trying to solve this problem another way. If a round trip to New York is too long, then twenty of them is way worse. So I can either do 20 round trips to Nevada, which does <20 round trips to Chicago, which does <<20 round trips to New York. Or, I can do some more cleverness with transport and session bootstrapping and end up with 14 round trips to New York.
- champtar 5y agoAlso CDN providers will hopefully have good pearing. My company uses OpenVPN TCP on port 443 for maximum compatibility. When around the globe the VPN is pretty slow, so I proxy the tcp connection via a cheap VPS, and speed goes from maybe 500kbit/s to 10Mbit/s, just because the VPS provider pearing is way better than my company "business internet". (The VPS is in the same country as the VPN server).
- mrkurt 5y agoWe've seen people use background revalidation to great effect, particularly in front of S3. You can get pretty close to one stale request per cache entry this way. And if-modified-since requests are really cheap.
- raro11 5y agoI set an s-maxage of at least a minute. Keeps my servers from being hugged to death while not having to invalidate manually.
- cortesoft 5y agoYou can fix this with proper cache headers
- babelfish 5y agofly.io has a fantastic engineering blog. Has anyone used them as a customer (enterprise or otherwise) and have any thoughts?
- alopes 5y agoI've used them in the past. All I can say is that the support was (and probably still is) fantastic.
- joshuakelly 5y agoYes, I'm using it. I deploy a TypeScript project that runs in a pretty straightforward node Dockerfile. The build just works - and it's smart too. If I don't have a Docker daemon locally, it creates a remote one and does some WireGuard magic. We don't have customers on this yet, but I'm actively sending demos and rely on it. Hopefully I'll get to keep working on projects that can make use of it because it feels like a polished 2021 version of Heroku era dev experience to me. Also, full disclosure, Kurt tried to get me to use it in YC W20 - but I didn't listen really until over a year later.
- cgarvis 5y agojust started to use them for an elixir/phoenix project. multi region with distributed nodes just works. feels almost magically after all the aws work I've done the past few years.
- tiffanyh 5y agoWhat’s magically? I was under the impression that fly.io today (though they are working on it) doesn’t do anything unique to make hosting elixir/Phoenix app easier. See this comment by the fly.io team. https://news.ycombinator.com/item?id=27704852 https://news.ycombinator.com/item?id=27704852
- mcintyre1994 5y agoThey're not doing anything special to make Elixir specifically better yet, but their private networking is already amazing for it - you can cluster across arbitrary regions completely trivially. It's a really good fit for Elixir clustering as-is even without anything specially built for it. I have no idea how you'd do multi-region clustering in AWS but I'm certain it'd be a lot harder.
- amirhirsch 5y agoThis is cool and informative and Kurt's writing is great: The briny deeps are filled with undersea cables, crying out constantly to nearby ships: "drive through me"! Land isn't much better, as the old networkers shanty goes: "backhoe, backhoe, digging deep — make the backbone go to sleep".
- tptacek 5y agoWe can't take credit for the backhoe thing; that really is an old networking shanty.
- deleted 5y ago[deleted]
- chrisweekly 5y agoThis is so great. See also https://fly.io/blog/ssh-and-user-mode-ip-wireguard/ https://fly.io/blog/ssh-and-user-mode-ip-wireguard/
- simonw 5y agoThis article touches on "Request Coalescing" which is a super important concept - I've also seen this called "dog-pile prevention" in the past. Varnish has this built in - good to see it's easy to configure with NGINX too. One of my favourite caching proxy tricks is to run a cache with a very short timeout, but with dog-pile prevention baked in. This can be amazing for protecting against sudden unexpected traffic spikes. Even a cache timeout of 5 seconds will provide robust protection against tens of thousands of hits per second, because request coalescing/dog-pile prevention will ensure that your CDN host only sends a request to the origin a maximum of once ever five seconds. I've used this on high traffic sites and seen it robustly absorb any amount of unauthenticated (hence no variety on a per-cookie basis) traffic.
- anonymoushn 5y agoDo you know if varnish's request coalescing allows it to send partial responses to every client? For example, if an origin server sends headers immediately then takes 10 minutes to send the response body at a constant rate, will every client have half of the response body after 5 minutes? Thanks!
- simonw 5y agoI don't know for certain, but my hunch is that it streams the output to multiple waiting clients as it receives it from the origin. Would have to do some testing to confirm that though.
- elithrar 5y agoI don’t know about Varnish, but having worked on other implementations, you would usually have a timeout on the initial lock (semaphore) to prevent a slow connection from impacting all clients. But this is much, much harder to do once you are already streaming the response - if the time to first byte (TTFB) is quick, but the connection is low-throughout, you can’t do much at this point. But nearly all modern implementations stream the bytes to all clients immediately; they don’t try to fill the cache first (they do it simultaneously). Some implementations might avoid fanning in too much - maintaining a smaller pool of connections rather than trying get to ”1”, but that’s ultimately a trade-off at each layer of the onion, as they can still add up. (I worked at both Cloudflare and Google, and it was a common topic: request coalescing is a big deal for large customers)
- youngtaff 5y agoSome of the things they miss in the post are Cloudflare uses a customised version or Nginx, same with Fastly for Varnish (don't know about Netlify and ATS) Out of the box nginx doesn't support HTTP/2 prioritisation so building a CDN with nginx doesn’t mean you're going ti be delivering as good service as Cloudflare Another major challenge with CDNs is peering and private backhaul, if you're not pushing major traffic then your customers aren't going to get the best peering with other carriers / ISPs…
- mike_d 5y agoHTTP/2 prioritization is a lot of hype for a theoretical feature that yields little real world performance. When a client is rendering a page, it knows what it needs in what order to minimize blocking. The server doesn't.
- youngtaff 5y agoYes, which is why the browser send priorities with the requests but many servers ignore these and just server responses in what ever order suits them. If a low priority response is served before a high priority one the page is likely to be slower to render etc.
- vmception 5y ago>The term "CDN" ("content delivery network") conjures Google-scale companies managing huge racks of hardware, wrangling hundreds of gigabits per second. But CDNs are just web applications. That's not how we tend to think of them, but that's all they are. You can build a functional CDN on an 8-year-old laptop while you're sitting at a coffee shop. huh yeah never thought about it I blame how CDNs are advertised for the visual disconnect
- lupire 5y agoIt's misleading. CDN software might be simple in the basic happy case, but you still need a Network of nodes to Deliver the Content.
- mrkurt 5y agoWell it's a self serving article! It's easy to turn up a network of nodes on Fly.io. It's a little harder, but not impossible, to do the same elsewhere.
- Rd6n6 5y agoSounds like a fun weekend project
- jabo 5y agoLove the level of detail that Fly's articles usually go into. We have a distributed CDN-like feature in the hosted version of our open source search engine [1] - we call it our "Search Delivery Network". It works on the same principles, with the added nuance of also needing to replicate data over high-latency networks between data centers as far apart as Sao Paulo and Mumbai for eg. Brings with it another fun set of challenges to deal with! Hoping to write about it when bandwidth allows. [1] https://cloud.typesense.org https://cloud.typesense.org
- mrkurt 5y agoI'd love to read about it.
- ksec 5y agoIt is strange that you put a Time duration in front of CDN ( content delivery network ), because given all the recent incident with Fastly, Akamai and Bunny, I read it as 5 hours Centralised Downtime Network.
- parentheses 5y agoAuthor has a great sense of humor. I love it!
- cortesoft 5y agoThe hard part of building a CDN is not setting up an HTTP cache, it is setting up an HTTP cache that can serve thousands of different customers.
- mrkurt 5y agoMaking a service multitenant is more complex, yes. But many companies roll their own CDNs. There are lots of good reasons to do that, and it's a problem that can be reduced to a single developer for understanding.
- intricatedetail 5y agoDoes Nginx still not support cache invalidation? If you setup long TTL, is there a way to remove some files from cache without nuking entire cache and restarting an instance?
- daniel_iversen 5y agoYears ago I was involved with some high performance delivery of a bunch of newspapers, and we used Squid[1] quite well. One nice thing you could do as well (but it's probably a bit hacky and old school these days) was to "open up" only parts of the web page to be dynamic while the rest was cached (or have different cache rules for different page components)[2]. With some legacy apps (like some CMS') this can hugely improve performance while not sacrificing the dynamic and "fresh looking" parts of the website. [1] http://www.squid-cache.org/ http://www.squid-cache.org/ [2] https://en.wikipedia.org/wiki/Edge_Side_Includes https://en.wikipedia.org/wiki/Edge_Side_Includes
- 3np 5y agoAs someone who’s mostly clueless about BGP but have a fair grasp of all the other layers mentioned, I’d love to see posts like this going more in depth on it for folks like myself.
- mbStavola 5y agoFly is great and I love reading their blog posts. Just hoping they come back around on CockroachDB-- I feel like it's a match made in heaven for what they're providing.
- tptacek 5y agoWe love CockroachDB. There are people tinkering with it on Fly.io. I think anything formal would involve our companies talking to each other, which we're happy to do, but everybody is busy all the time. :)
- mrkurt 5y agoWe're getting there: https://github.com/fly-apps/cockroachdb https://github.com/fly-apps/cockroachdb
- awoods187 5y agoPM at CRL here--we love Fly too! Definitely can see our two products working together!
- jusssi 5y ago> 3. Be like a game server: Ping a bunch of servers and use the best. Downside: gotta own the client. Upside: doesn't matter, because you don't own the client. "If you can run code on it, you can own it". Your front page could just be a tiny loader js that fires off a fetch() for a zero byte resource to all your mirrors, and then proceeds to load the content from the first responder.
- marcosdumay 5y agoNow you just have the bad latency of the non-cached content, plus the ok latency of your CDN.
- amelius 5y agoWaiting for IPFS to shake this all up.
- deleted 5y ago[deleted]
- cpascal 5y ago> DNS: Run trick DNS servers that return specific server addresses based on IP geolocation. Downside: the Internet is moving away from geolocatable DNS source addresses. Upside: you can deploy it anywhere without help. Can anyone expand on how/why "the Internet is moving away from geolocatable DNS source addresses"?
- mritzmann 5y agoSome public/recursive DNS Servers like Cloudflare (1.1.1.1) do not tell the authoritative dns server the ip address or subnet of the requestor. Your ISP's DNS server usually does this. This makes CDN via DNS more difficult, as it is not always entirely clear from where the request comes (Cloudflare itself does not need this, they do everything with Anycast).