Operations | Monitoring | ITSM | DevOps | Cloud

The latest News and Information on Incident Management, On-Call, Incident Response and related technologies.

The Failure Mode Your Runbook Probably Does Not Cover

Operations teams rehearse plenty of scenarios. Failed deployments, database corruption, certificate expiry, a region going dark, the on-call engineer who cannot be reached. What gets rehearsed far less often is the building losing power for eleven hours, because that feels like somebody else's problem, filed under facilities alongside the air conditioning and the parking barrier. It stops being somebody else's problem at the moment the UPS batteries drain and everything still running on premises goes down at once.

Ensuring Business Continuity in Adverse Conditions

Businesses will always face disruptions. Whether it's a big storm, a broken supply chain, or a power outage, unexpected problems can bring operations to a halt, hurting your income, your reputation, and how much customers trust you. The companies that make it through these tough times, and those that don't, often come down to one thing: resilience. Being a resilient organization isn't about building an unshakeable fortress. It's about being flexible, thinking ahead, and having the right systems to bounce back when disruptions occur.

3 Things IT Leaders Are Learning About AI-First Operations: Key Takeaways From PagerDuty on Tour 2026

In December 2025, an AI coding agent at AWS suddenly decided to delete and rebuild an entire production environment, causing a 13-hour service disruption and a PR headache for Amazon. As rapid adoption of AI leads to more high-profile, revenue-impacting incidents, resilience has moved from a technical concern to a board-level financial risk.

Incident Review with Factory / Sentry

Learn how to user agents to turn Slack alerts into autonomous RCA sessions, build incident memory, and help on-call engineers move from signal to fix faster. ​Join us for the live stream demo of Incident Response. We will break things on stream and let agents fix them. We will watch Droid run a real incident from alert to fix, and we show exactly what it read to get there.

The 13 Questions CEOs Ask After an Incident (And What IT Leaders Must Be Ready to Answer)

It’s 2:47 p.m. Your checkout service has been down for 11 minutes. Customers are screenshotting errors and calling in. Your CEO walks into your office and starts asking questions. In this moment, there are two kinds of IT leaders: Whether you walk out with more budget authority (and executive trust) or less depends on your answers, and the infrastructure that supports them. But preparation isn’t just about surviving the incident. It’s actually a revenue opportunity.

IT problem management VS. IT incident management, and how agentic ITOps improves both

Picture a familiar scene: a critical application goes down during peak business hours, and your on-call engineers scramble to restore service. Two weeks later, the same application fails again, frustrating your teams with the same symptoms, the same scramble, and the same customer frustration. If this pattern feels familiar, your organization may be strong at IT incident management, but underinvested in IT problem management.

Why Regular IT Health Checks Help Prevent Downtime and Improve Business Resilience

Most IT problems do not announce themselves. A backup job quietly fails for three weeks before anyone notices. A firewall rule left open "temporarily" during a project stays open for a year. A former employee's account still has admin rights nobody remembered to remove. None of these cause trouble on the day they happen. They cause trouble later, usually at the worst possible time. This very difference between when these problems start and when they finally become a source of trouble is precisely what a routine IT health check is supposed to bridge.

Your AI agents are lost: give them a graph

The biggest limitation facing enterprise AI agents may not be the model. It may be the context surrounding it. Anthony Alcaraz, Senior AI/ML Portfolio Growth Manager at AWS and co-author of O'Reilly's *Agentic GraphRAG*, joins Humans of Reliability to explain why reliable agents need more than a vector database and a large context window. They need structured knowledge they can navigate, memory they can prune, constraints they can follow, and feedback loops that help them improve.