4 ms·
I can't comment on the current incident, as I've been gone for more than a decade. But my recollection is that most SRE teams would have had exactly such thing
by clavoie 6y ago
I can't comment on the current incident, as I've been gone for more than a decade.
But my recollection is that most SRE teams would have had exactly such things -- either as really crazy borgcfg (sorry, I mean kubectl) code, or as a borgmon (sorry, prometheus) alert.
The former were a real PITA to work with (borgcfg hadn't been designed with that in mind), and the latter were only reactive (and thus unable to warn about things until they were fast becoming a real problem).
So most likely those safeguards were somehow disabled or made ineffective by the exact circumstances of the problem, possibly the speed at which things developed.