Operations | Monitoring | ITSM | DevOps | Cloud

The latest News and Information on IT Operations Management and related technologies.

Get the Context Your Alerts Are Missing with Event Enrichment

This blog post is part of PagerDuty’s ongoing series on how we’re helping customers navigate their journey towards autonomous operations. Read on to learn about how Event Enrichment builds towards this vision. Every on-call engineer knows the drill. An alert fires. It tells you something is wrong, but not what it means. Is this asset in maintenance? Which team owns it? Is it customer-facing?

SRE Agent Enhancements: Faster Triage, Greater Access Controls, Deeper System Connectivity

This blog post is part of PagerDuty’s ongoing series on how we’re helping customers navigate their journey towards autonomous operations. Read on to learn about how recent SRE Agent Enhancements build towards this vision. During an incident, everything is competing for attention at once. Responders lose time swiveling between tools, insights gathered by AI stay siloed instead of feeding into the next decision, and the pressure to move fast means learnings rarely stick.

Bring Your Backstage Context Into Every PagerDuty Incident

This blog post is part of PagerDuty’s ongoing series on how we’re helping customers navigate their journey towards autonomous operations. Read on to learn about how Custom Field Mapping for PagerDuty’s plugin for both Spotify for Backstage and Spotify Portal for Backstage now generally available builds towards this vision. It’s 2am. A Sev-1 fires, and your on-call responder opens the incident in PagerDuty. What’s waiting for them? A service name, and not much else. No tier.

10 Best IT Operations Management Tools (ITOM) Compared for 2026

No IT team sets out to run ten monitoring tools. It happens one purchase at a time. You add a network monitor after one outage, a server monitoring tool after the next, then a log analyzer, then an APM product when the app team gets tired of guessing. Then something breaks, and every one of those tools has an opinion. Each fires its own alerts. None of them agrees on the cause, and the first hour of the incident goes to deciding which screen to believe.

3 Things IT Leaders Are Learning About AI-First Operations: Key Takeaways From PagerDuty on Tour 2026

In December 2025, an AI coding agent at AWS suddenly decided to delete and rebuild an entire production environment, causing a 13-hour service disruption and a PR headache for Amazon. As rapid adoption of AI leads to more high-profile, revenue-impacting incidents, resilience has moved from a technical concern to a board-level financial risk.

The 13 Questions CEOs Ask After an Incident (And What IT Leaders Must Be Ready to Answer)

It’s 2:47 p.m. Your checkout service has been down for 11 minutes. Customers are screenshotting errors and calling in. Your CEO walks into your office and starts asking questions. In this moment, there are two kinds of IT leaders: Whether you walk out with more budget authority (and executive trust) or less depends on your answers, and the infrastructure that supports them. But preparation isn’t just about surviving the incident. It’s actually a revenue opportunity.

Custom shifts for one-off requirements or complex schedules

While most on-call schedules are built to represent regular rotations, often on a weekly basis, not all of your on-call needs require the same coverage every week. We’ve added Custom Shifts to the Shift-Based Schedules for maximum flexibility. Custom shifts are a feature of our new Shift-Based Schedules. With Custom Shifts, your team can cover ad hoc needs for special events, major deploys, Failure Fridays, gamedays, or whatever comes up that needs some extra coverage.

Make the most of shift-based schedules

We recently updated our Schedules to better reflect how teams are currently managing their on-call responsibilities. Not everyone is working on weekly shifts or providing 24×7 coverage for all of their services, and that should be easy to schedule in our new tooling. To give you some examples, I’ve gone back through some of the questions we’ve gotten on the PagerDuty Commons over the past couple of years for questions about custom schedules that we weren’t really thinking about.

ServiceNow Runs Your IT. PagerDuty Makes Sure It Never Stops.

For most enterprises, ServiceNow has become the backbone of IT operations, the platform where workflows are governed, compliance is maintained, and every incident, change, and request is tracked from start to finish. If you’re running ServiceNow, you’ve made a serious investment in how your IT operates. PagerDuty is built to make that investment work even harder.

Why Faster Recovery Beats Faster Shipping in the AI Era

A year ago, AI coding tools worked alongside developers—suggesting the next line, completing a function, accelerating work that a human was already doing. Today, they’re writing entire modules and services independently, producing code that no human has reviewed line by line, built from components that no single person has fully mapped. And adoption is only accelerating: According to our recent AI Resilience Survey, 84% of organizations are now using AI to write, review, or suggest code.