The Monitoring Stack We Actually Use in Production

TL;DR AI
2 min readKey summary
A production monitoring stack using Prometheus, Grafana, PagerDuty, and Sentry was tuned for better operations.
The team cut alerts from 200 to 30, added runbook links, and reviewed dashboards weekly to reduce noise and rot.
Incident response improved sharply: mean time to detect fell from 45 minutes to 8 minutes, and mean time to resolve from 3 hours to 45 minutes.
The changes helped reduce missed incidents, ease on-call stress, and speed up troubleshooting in production.
