Alert fatigue, AI triage, and incidents: Lessons from observability experts at Cyera, PlayHQ & NAB
Observability looks perfect in a slide deck – in practice, it's messier. In this panel, engineering leaders from Cyara, PlayHQ, and National Australia Bank share what really happened when they scaled observability: unexpected cloud bills, alert fatigue, a weekend database outage caught by an AI-assisted triage agent, and a vendor dispute settled by a single chart. They also cover moving beyond legacy tooling, using AI to close the PromQL skills gap, and what's next – from agentic SDLC integration to continuous profiling. Real stories, real numbers, real lessons.
Chapters
- 00:00 — Welcome and introductions
- 01:01 — Cyara: the disk cost discovery
- 02:11 — PlayHQ: scaling community sports at 50,000 games a week
- 03:44 — National Australia Bank: central observability at scale
- 04:53 — Driving cultural adoption
- 07:04 — Closing the PromQL gap with AI
- 08:20 — NAB's shift to OpenTelemetry
- 09:54 — Managing telemetry costs and alert fatigue
- 10:39 — Adaptive Telemetry and quarterly alert reviews
- 12:22 — Meta-monitoring: alerting on the alerts
- 14:43 — Cutting noise at the collection layer
- 16:29 — Real incidents, real wins
- 17:16 — Tracing cuts investigation time (NAB)
- 18:42 — A weekend outage, caught fast (PlayHQ)
- 19:40 — Settling a vendor dispute with data (Cyara)
- 22:09 — What's next: AI root cause, continuous profiling
- 24:21 — Final advice for engineering leaders