Why Prometheus couldn’t see Cilium metrics at 2 a.m.

TL;DR AI
2 min readKey summary
A production Kubernetes team found blank Grafana panels because Prometheus was not scraping Cilium agent and operator pods.
The issue was not broken software, but a missing ServiceMonitor integration between otherwise healthy CNCF projects.
The example highlights a broader cloud-native risk: outages and blind spots often come from poor platform wiring, not component failure.
It underscores the operational cost of day-2 integration work across tools like Prometheus, Cilium, Grafana, and Hubble.
