6 ms·
Ask HN: Connecting Kubernetes dependencies to application telemetry
I’m building Cardinal (cardinalhq.io). One thing that drives me nuts about Kubernetes debugging is the number of steps it takes to go from an application issue to its underlying kubernetes root cause.
Like for example, if I am seeing an error rate spike, and I group by deployment, I want to immediately see the deployment's k8s neighborhood, which would then show me the configmap that was updated recently and incorrectly.
Similarly, say a node is low on disk and that is causing a ripple effect on a statefulset trying to write new files, which then affects request latency.
The list goes on.
To solve this, one idea we had was: build a smart chart legend that recognizes a k8s tag (based on OTEL semantic conventions), and then for every entry pops an "infra map" icon. Click and you see the full k8s neighborhood of that entity with health annotations. So now it's easier to go from app telemetry to k8s infra.
I am curious if this UX makes sense, and if not how do you solve this correlation problem today?
https://cardinalhq.io/use-cases/kubernetes
- andythomasc83 9d ago[flagged]
- jesol 11d agoPersonally I think security is far ahead here compared to normal observability tools. I decided to work on a side-project to try and add SIEM like functionality to a clickhouse backed otel platform; by the end of it the thing I came to believe the tooling blue teams use should be used in observability generally, not just security. Incident/event management is a very powerful concept, and provides a clean framework to hang all of this information on. Then your solution for mechanically finding the neighborhood in k8s is one way to add observations to an event. Some SIEMs have started having agents recommend stuff to be added to an event, which is a nice middle-ground of having agents help but not completely control the discovery and diagnostic effort (as well as providing a clean feedback loop for training data synthesis). That's all to say, have you considered that framing, and if so, have any opinions on why more general observability tools haven't gone that direction?
- chipfixer 9d agoAre you saying that if everything became an event, its easier to correlate because the event is more free form and it's easier to add custom metadata? In theory, I agree with you. But you can't eliminate 'joins' all together, IMO, because it's not practical to know beforehand what context will be useful.
- deleted 10d ago[deleted]
- jeansilga 10d agoconfigmap that was updated recently and incorrectly: how would you know the cm update is the cause? How would you sort that out if you made 5/10 updates? say a node is low on disk: what about setting up monitoring alerts for those king of things. In general, alerts are of great help. An alert fired after an update is a big smell about that update causing the issue.
- chipfixer 7d agoSo if there is a spike on say errors, you do a group by pod, click on the pod in the legend, see the k8s neighborhood which means discover the pod --> config map association. You see the config map and all the recent updates to it, and can even diff them like code. An alert on the other hand, say on a node that is low on disk, doesn't automatically tell you all the affected workloads. Not saying the alert doesn't help, but this experience helps identify the affected pods and apps much faster.
- vancekai 10d ago[dead]