5 ms·
TL;DR meta point is: "The reason that we haven’t been doing any significant work on this HAProxy stack is that we’re moving towards Envoy Proxy for all of our i
by debt 6y ago
TL;DR meta point is: "The reason that we haven’t been doing any significant work on this HAProxy stack is that we’re moving towards Envoy Proxy for all of our ingress load-balancing"
The legacy system supporting Slack in production was heavily resource-constrained as they were moving to a new fancy system. Slack admits here that the legacy system likely wasn't getting the attention it needed and lo-and-behold it started failing in mysterious ways.
Organizational failure by not properly calculating all the risks caused by rotating out part of their load-balancing system. They probably should've asked for more budget here to keep their existing system functional as they slowly transitioned to their new system.
They admit that COVID caused all their systems to become stressed, they probably had appropriately budgeted for the transition to Envoy whenever they asked management(probably pre-covid). The team likely was never meant to support both the load they're now seeing during COVID while transitioning to a new system.
Either way during any transition, there's a period where you must support both systems at full capacity until the legacy system can be gracefully decommissioned.
- sailfast 6y agoI really feel this comment. Especially during transition it’s hard to see the “to-be legacy” system as worthy of the effort because the new stuff will be here so soon! Until it isn’t because life happens and you’re left with a really shaky platform. These are tough investment decisions especially when resource constrained. As a rule, perhaps ensuring that current state platforms are secure, as you say, before attempting a migration is the best way to go, but one person’s critical work is another person’s “too much insurance.”