An alert fires, but it only shows a symptom. Small incidents keep recurring, but nobody connects them. Internal dashboards look healthy while customers can’t access the service.
Sarah Kaplan from Apple, Heather Osborn from Cars Commerce, and Barnadeep Bhowmik from SLB share how they chased down problems like these, what they missed, and what they changed afterward.
Hear how Kubernetes 503 errors led to a DNS bug, recurring PostgreSQL failures went unconnected for months, and an Active Directory issue triggered downstream alerts and three separate war rooms.
The discussion covers using SLOs, error budgets, burn rates, synthetic monitoring, and user impact to prioritize incident response, plus where AI can help correlate alerts and why engineers still need to verify its conclusions.
Watch the on-demand panel for real production lessons on alert fatigue and root cause analysis, plus an OpenObserve demo of SLO creation and burn rate alerting.