Operations | Monitoring | ITSM | DevOps | Cloud

Dashboards aren't (quite) dead

Historically, non-technical stakeholders would’ve had most of their data questions answered either through pre-built dashboards or by asking their Data team (or equivalent). Self-serve analytics tools went a step further by offering safe, governed datasets built by Data teams which let non-technical users dig into data without having to worry about how it joins together, how metrics like “revenue” are defined, and so on.

Don't add a read replica until you've read this

As the size and complexity of their relational database workload grows, every company eventually goes through the process of off-loading work on a read replica. It comes with lots of benefits, but at a cost of increased complexity. This article is about how we dealt with that, a lot of learnings, and some useful techniques. incident.io is an incident management product relied on by thousands of customers to be the thing that supports them through anything from a minor blip to a full outage.

How Zendesk ditched 15 years of patchwork tooling, in 10 weeks

Zendesk replaced 15 years of homegrown incident tooling and PagerDuty by migrating 1,200 engineers across 150 teams onto incident.io in just 10 weeks, cutting mean time to triage by 32%, saving $500k+ in year one, and eliminating 800+ hours of annual toil, with zero incidents on go-live day. Tom Monaghan (VP of Engineering Productivity & Product Reliability) and Anna Roussanova (Engineering Manager) share how they pulled it off and what's next as Zendesk helps build Investigations, our AI agent that starts digging into incidents the moment an alert fires.