Operations | Monitoring | ITSM | DevOps | Cloud

Cloud Outage Preparedness: On-Call Lessons for 2026

Cloud outage preparedness stopped being a nice-to-have this month. In a span of roughly 48 hours, Microsoft Azure lost a big chunk of its West US footprint and Amazon Web Services dropped connectivity between its us-west-2 region in Oregon and the Seattle metro. The AWS event alone rippled outward and knocked DoorDash, Reddit, Hulu, Apple Pay, Snapchat, Fortnite, and the PlayStation Network offline for millions of users, according to incident trackers. Neither outage was caused by anything exotic.

Cloud Outage Incident Response: Lessons From 2026

Cloud outage incident response stopped being a hypothetical exercise this summer. In a single stretch of July 2026, three of the biggest cloud providers stumbled in quick succession, and the ripple effects reached apps that millions of people use every day. If your team runs anything on a hyperscaler, the events of the last few weeks are a direct message: the question is no longer whether your provider will have a bad day, but whether your on-call rotation is ready when it does.

When Status Pages Lie: The Incident Detection Gap

On July 28, 2026, roughly 30,000 people flooded Downdetector with reports that Reddit was broken. Feeds would not load, logins failed, and the mobile app hung. Reddit's own status page, meanwhile, showed a calm wall of green: all systems operational. That contradiction is the whole story, and it is not unique to Reddit. It is one of the most common and most damaging failure modes in modern on-call, and it has a name: the incident detection gap.

T-Mobile SOS Outage: Incident Response Lessons

When more than 140,000 people reach for their phones at once and see nothing but the letters SOS, the topic of incident response stops being an abstract engineering concern and becomes something everyone feels. That is exactly what happened on the evening of July 27 into the morning of July 28, 2026, when a nationwide T-Mobile outage knocked huge numbers of devices into SOS only mode, cutting people off from regular calls, texts, and data.

Dashboards aren't (quite) dead

Historically, non-technical stakeholders would’ve had most of their data questions answered either through pre-built dashboards or by asking their Data team (or equivalent). Self-serve analytics tools went a step further by offering safe, governed datasets built by Data teams which let non-technical users dig into data without having to worry about how it joins together, how metrics like “revenue” are defined, and so on.

Cloud Outage Response: Lessons From July's Bad Week

In a single week, two of the largest cloud providers on earth failed at almost the same time, and a good chunk of the internet went with them. Effective cloud outage response stopped being a theoretical exercise and became the difference between a calm 30 minutes and a chaotic afternoon for thousands of on-call engineers. On July 23, 2026, a maintenance bug inside Microsoft Azure pulled IP routes off more devices than intended in the West US region, cutting Microsoft 365 access for millions.

Cloud Outage Response: AWS us-west-2 Lessons

Cloud outage response got another live fire drill on July 24, 2026, when AWS lost network connectivity between its us-west-2 region in Oregon and the Seattle metro. For most customers the pain lasted about 20 minutes, and a small set on AWS Direct Connect saw errors for roughly an hour and seventeen minutes. That is short as major cloud incidents go. What makes it worth your attention is not the duration.

How Centralized Knowledge Cuts MTTR During Major IT Incidents

Centralized knowledge cuts MTTR by attacking the phase of an incident where most of the clock actually burns: diagnosis. When responders can pull the right runbook, past incident records, and system documentation from one searchable place, they skip the twenty minutes of paging people and digging through wikis that normally precede any real troubleshooting. The fix itself is often quick. Finding out what to fix is what takes an hour.

AWS Outage Incident Response: What July 24 Taught Us

On the morning of July 24, 2026, a large slice of the internet blinked out at once. An AWS outage centered on the US-West-2 region in Oregon rippled outward and took DoorDash, Reddit, Hulu, Apple Pay, Snapchat, Fortnite, and the PlayStation Network offline for millions of users. If your team runs anything on Amazon Web Services, this is the incident to study, because the hard part was never fixing AWS. The hard part was AWS outage incident response.

The July 24, 2026 AWS us-west-2 Outage: Network Routing and a Long Recovery Tail

On July 24, 2026, AWS lost connectivity between the us-west-2 (Oregon) region and the Seattle Metro. The initial impact window was 20 minutes for most and 1 hour 17 minutes for a few customers using AWS Direct Connect through EqSe2, Westin Building Exchange, Seattle. Any traffic that both started and ended inside the region kept working, whereas anything crossing the region boundary saw timeouts and errors. This included the AWS Management Console for some customers.

Azure outage on July 23, 2026: StatusGator detected it 1 hour before Microsoft acknowledged it

On July 23, 2026, Azure users around the world began hitting gateway timeouts, DNS failures, and unreachable virtual machines well before Microsoft posted anything on its status page. The first reports reached StatusGator at 15:06 UTC. By 15:28 UTC, StatusGator had sent an Early Warning Signal to subscribers. Microsoft did not acknowledge the incident until 16:29 UTC.

The July 23 2026 Azure West US Outage: IP Route Removal and Downstream Impact

On July 23, 2026, Microsoft Azure experienced a connectivity outage in the West US region that blocked traffic entering or leaving the region for nearly five hours. Workloads that stayed entirely inside West US were not affected. Microsoft's preliminary Post Incident Review (PIR) attributes the failure to a bug in maintenance request conversion software that removed IP routes from more devices than intended during routine device maintenance.

Incident Response Communication: Why Ops Teams Own the Narrative

Your monitoring stack flagged the outage in 90 seconds. A customer posted about it in 40. That gap is now the defining challenge of incident response communication. Ops teams have spent years driving down recovery times, yet very few track how quickly a public explanation takes shape. This article looks at how teams can monitor both timelines - and respond before speculation hardens into accepted fact.

The Failure Mode Your Runbook Probably Does Not Cover

Operations teams rehearse plenty of scenarios. Failed deployments, database corruption, certificate expiry, a region going dark, the on-call engineer who cannot be reached. What gets rehearsed far less often is the building losing power for eleven hours, because that feels like somebody else's problem, filed under facilities alongside the air conditioning and the parking barrier. It stops being somebody else's problem at the moment the UPS batteries drain and everything still running on premises goes down at once.

Ensuring Business Continuity in Adverse Conditions

Businesses will always face disruptions. Whether it's a big storm, a broken supply chain, or a power outage, unexpected problems can bring operations to a halt, hurting your income, your reputation, and how much customers trust you. The companies that make it through these tough times, and those that don't, often come down to one thing: resilience. Being a resilient organization isn't about building an unshakeable fortress. It's about being flexible, thinking ahead, and having the right systems to bounce back when disruptions occur.

3 Things IT Leaders Are Learning About AI-First Operations: Key Takeaways From PagerDuty on Tour 2026

In December 2025, an AI coding agent at AWS suddenly decided to delete and rebuild an entire production environment, causing a 13-hour service disruption and a PR headache for Amazon. As rapid adoption of AI leads to more high-profile, revenue-impacting incidents, resilience has moved from a technical concern to a board-level financial risk.

Incident Review with Factory / Sentry

Learn how to user agents to turn Slack alerts into autonomous RCA sessions, build incident memory, and help on-call engineers move from signal to fix faster. ​Join us for the live stream demo of Incident Response. We will break things on stream and let agents fix them. We will watch Droid run a real incident from alert to fix, and we show exactly what it read to get there.

The 13 Questions CEOs Ask After an Incident (And What IT Leaders Must Be Ready to Answer)

It’s 2:47 p.m. Your checkout service has been down for 11 minutes. Customers are screenshotting errors and calling in. Your CEO walks into your office and starts asking questions. In this moment, there are two kinds of IT leaders: Whether you walk out with more budget authority (and executive trust) or less depends on your answers, and the infrastructure that supports them. But preparation isn’t just about surviving the incident. It’s actually a revenue opportunity.

IT problem management VS. IT incident management, and how agentic ITOps improves both

Picture a familiar scene: a critical application goes down during peak business hours, and your on-call engineers scramble to restore service. Two weeks later, the same application fails again, frustrating your teams with the same symptoms, the same scramble, and the same customer frustration. If this pattern feels familiar, your organization may be strong at IT incident management, but underinvested in IT problem management.

Why Regular IT Health Checks Help Prevent Downtime and Improve Business Resilience

Most IT problems do not announce themselves. A backup job quietly fails for three weeks before anyone notices. A firewall rule left open "temporarily" during a project stays open for a year. A former employee's account still has admin rights nobody remembered to remove. None of these cause trouble on the day they happen. They cause trouble later, usually at the worst possible time. This very difference between when these problems start and when they finally become a source of trouble is precisely what a routine IT health check is supposed to bridge.

Your AI agents are lost: give them a graph

The biggest limitation facing enterprise AI agents may not be the model. It may be the context surrounding it. Anthony Alcaraz, Senior AI/ML Portfolio Growth Manager at AWS and co-author of O'Reilly's *Agentic GraphRAG*, joins Humans of Reliability to explain why reliable agents need more than a vector database and a large context window. They need structured knowledge they can navigate, memory they can prune, constraints they can follow, and feedback loops that help them improve.

Don't add a read replica until you've read this

As the size and complexity of their relational database workload grows, every company eventually goes through the process of off-loading work on a read replica. It comes with lots of benefits, but at a cost of increased complexity. This article is about how we dealt with that, a lot of learnings, and some useful techniques. incident.io is an incident management product relied on by thousands of customers to be the thing that supports them through anything from a minor blip to a full outage.

Custom shifts for one-off requirements or complex schedules

While most on-call schedules are built to represent regular rotations, often on a weekly basis, not all of your on-call needs require the same coverage every week. We’ve added Custom Shifts to the Shift-Based Schedules for maximum flexibility. Custom shifts are a feature of our new Shift-Based Schedules. With Custom Shifts, your team can cover ad hoc needs for special events, major deploys, Failure Fridays, gamedays, or whatever comes up that needs some extra coverage.

From AIOps to agentic ITOps: Why AI for IT operations has entered a new era

Enterprise IT has reached an inflection point. Your teams are responsible for hybrid cloud infrastructure, microservices, third-party dependencies, and shipping AI-generated code at unprecedented velocity. IT environments are becoming more complex faster than traditional tools and processes can keep pace. Alert volumes keep climbing. Institutional knowledge keeps walking out the door. And the pressure to do more with flat or shrinking budgets isn’t letting up.

H1 2026 Cloud and SaaS Reliability Report

The first half of 2026 reinforced a key idea about Cloud and SaaS reliability - dependency risk. IncidentHub tracked 30,246 outages across 1,082 providers between January and June 2026. May was the busiest month, with 6,070 incidents. Cloud providers led in the total number of outages (4,723), followed closely by developer tools (4,589).

The July 2026 AWS CloudFront Outage: VPC Origins, Cascade Impact, and What Broke

On July 16, 2026, AWS experienced a disruption in its CloudFront service, which affected a large number of websites and applications. The outage was caused by a configuration loading failure in CloudFront's VPC Origins feature. This was AWS's most widely-felt outage after last year's outage on October 20th, which caused widespread damage.

Trust, Resilience & AI: A Customer Panel with TD Bank & New York Life

What does it really take to be "the calm in the storm" during a major incident? In this candid panel from PagerDuty on Tour, Chris Conklin (Technology Executive AIOPs, TD Bank) and Sam Brinley (CVP Enterprise Cloud Solution Architect & Engineer at New York Life) sit down with PagerDuty to talk through two decades of evolution in IT operations – from the "Wild West" of early network management to today's push into AI and agentic operations.

What is MTTR, and how can agentic ITOps reduce it?

Mean time to resolution (MTTR) measures the average duration to restore regular operation for an application, service, or infrastructure component. It’s a key performance indicator (KPI) for IT incident management. To tie MTTR directly to customer satisfaction, you first need to understand how it affects service and application reliability and availability. From there, you can make informed decisions, operate efficiently, and provide a seamless customer experience.

On call? Don't miss the next World Cup match.

Plans change - and your on-call schedule should be able to change with them. With SIGNL4, you can quickly arrange shift coverage from your smartphone, so your team stays fully staffed while everyone knows exactly who's on duty. Whether it's a World Cup match, a family event, or any other last-minute plan, SIGNL4 helps you manage stand-ins and shift handovers without phone calls, spreadsheets, or confusion.

AI vs. AI: from alert fatigue to agentic cybersecurity

AI is transforming cybersecurity on both sides of the battlefield. Attackers can now launch highly personalized phishing campaigns at scale and build malware capable of making autonomous decisions. At the same time, security teams are using AI agents to investigate alerts, reduce noise, and respond to threats faster. In this episode of Humans of Reliability, we speak with Nir Soudry, Head of R&D at 7AI, about the shift from alert fatigue to agentic cybersecurity.

PagerDuty Announces Arnaud Lagarde, Vice President of EMEA

PagerDuty, Inc. announces the appointment of Arnaud Lagarde as vice president of EMEA. Lagarde will lead PagerDuty's next phase of growth in the EMEA region, bringing the entire incident management lifecycle to customers across EMEA to solve their biggest digital challenges.

How to lay the data foundation to support agentic ITOps

Agentic IT operations have arrived. It’s no longer a question of if enterprise IT departments will adopt agentic ITOps, but how quickly. Every year, IT environments grow more distributed, complex, and difficult to monitor with legacy tools and processes. At the same time, the pace of AI development is accelerating the volume of changes and incidents, straining teams that are still trying to manage them manually, reactively, and one alert at a time.

Stop Triaging in the Dark: Full Visibility Across Every IT Domain

Alert correlation solved the noise problem. But noise was never the whole problem. Today’s most disruptive incidents cascade across networks, infrastructure, applications, and services simultaneously, without clear visibility into the true root cause. As a result, L1 teams are left manually piecing together context from multiple dashboards and tools to find the primary root cause while SLA clocks keep ticking and end user tickets add up.

The Value of Preventive Maintenance in Modern Business Operations

Preventive maintenance helps businesses reduce downtime, avoid costly breakdowns, extend equipment life, and maintain safer, more efficient operations. By addressing small issues early, companies can keep workflows running smoothly and protect productivity in a competitive business environment.

Where Status Pages Fit in a Modern Incident-Response Workflow

An incident-response process has two audiences from the moment a service begins to fail. Engineers need evidence detailed enough to isolate the fault. Customers need a clear account of what is affected, what still works, and when they should expect another update. Trying to serve both groups from the same dashboard usually leaves each with the wrong information.

From BigQuery to ClickHouse: How we made our analytics 5× faster

‍For years, ilert has given our customers extensive analytics across their alerts, notifications, and on-call activity, a comprehensive overview of how their teams and services respond to incidents. These capabilities were backed by a separate analytical database running on Google BigQuery. It held the numbers behind every reporting dashboard in ilert, and for a long stretch it was perfectly fine. Then three problems grew too big to ignore.

We rebuilt Spike app for Slack

The new Spike app for Slack brings incident response into the channel your team already works in. This walkthrough covers the @Spike AI assistant, the redesigned incident alert template, Statuspage syncing, and on-call overrides. To get started, head to Slack settings inside Spike and reconnect the app. Chapters Statuspage syncing is available on all plans. Spike is an incident response and on-call management platform. Alert routing, escalation policies, on-call schedules, and incident management, built for engineering teams.

Slack overview video

The new Spike app for Slack brings incident response into the channel your team already works in. This walkthrough covers the @Spike AI assistant, the redesigned incident alert template, Statuspage syncing, and on-call overrides. To get started, head to Slack settings inside Spike and reconnect the app. Chapters Statuspage syncing is available on all plans. Spike is an incident response and on-call management platform. Alert routing, escalation policies, on-call schedules, and incident management, built for engineering teams.

AT&T Email-to-Text Replacement: Best Alternatives for Critical Alerts

AT&T is permanently shutting down its email-to-text and text-to-email gateway, which means alerts sent to @txt.att.net or @mms.att.net no longer reach phones. For IT teams, MSPs, facilities teams, building management, utilities and incident response teams in general, this creates a serious gap. Critical alerts from monitoring tools, ITSM platforms, building systems, IoT devices, and other operational systems still need to reach the right person quickly, especially after hours.

5 AT&T Email-to-Text Alternatives to Improve MTTR in 2026

On June 17, 2025, AT&T permanently shut down its email-to-text and text-to-email gateway. Emails sent to @txt.att.net and @mms.att.net stopped reaching phones, and any automated workflow that relied on that address went dark overnight (AT&T support) . For IT Ops, MSPs, facilities and energy ops and incident response teams, this was not a minor inconvenience.

How Zendesk ditched 15 years of patchwork tooling, in 10 weeks

Zendesk replaced 15 years of homegrown incident tooling and PagerDuty by migrating 1,200 engineers across 150 teams onto incident.io in just 10 weeks, cutting mean time to triage by 32%, saving $500k+ in year one, and eliminating 800+ hours of annual toil, with zero incidents on go-live day. Tom Monaghan (VP of Engineering Productivity & Product Reliability) and Anna Roussanova (Engineering Manager) share how they pulled it off and what's next as Zendesk helps build Investigations, our AI agent that starts digging into incidents the moment an alert fires.

Enhanced Slack Experience

PagerDuty’s slack experience is evolving to help your teams organize better and resolve incidents faster. Use Triage Channels to collect telemetry and updates from your systems. Create dedicated Incident Channels for coordination and resolution. Give stakeholders the updates they need in Announcements Channels. Everyone in your organization can get the information they need easily.

ilert introduces dedicated incident management

Not all alerts are created equal. Some are resolved quickly by the on-call engineer. Others signal something serious enough to affect your business and require your whole team to coordinate. That is why we redesigned incidents as a dedicated coordination workspace for the alerts that have the most business impact.‍ Until now, incidents in ilert were used to communicate status updates to customers and stakeholders. Creating one meant publishing to your status page. We have separated the two.

StepbyStep Guide to Automating Alert Management for IT Ops

Your monitoring stack never sleeps. Datadog fires a spike, ServiceNow spins up a ticket, your RMM flags a failed backup, and every one of those signals competes for attention across email, dashboards, and chat channels. For IT Ops teams running on-call rotations, the volume itself becomes the problem. Alert fatigue sets in, critical notifications blend into the noise, and the one incident that matters at 3 a.m. gets buried under a hundred that don’t. The cost is real.

Introducing the BigPanda AI Incident Assistant

AI incident assistant from BigPanda gives L2, L3, and SRE teams instant answers to resolve incidents faster without manual triage or tool-switching. IT teams lose critical minutes during incidents because context is scattered across Slack threads, bridge calls, monitoring tools, and historical tickets. The BigPanda AI Incident Assistant fixes that by surfacing relevant knowledge exactly when and where responders need it. It gives responders evidence-based resolution paths drawn from historical incidents and live system data, without leaving your workflows.

Introducing AI Incident Prevention from BigPanda

AI Incident Prevention from BigPanda stops change-related outages before they occur by leveraging risk scores, trend analysis, and guided remediation steps. Manual IT changes are still a leading cause of IT outages and disruptions. BigPanda AI Incident Prevention addresses this by automatically scoring change requests against historical data, flagging high-risk changes before they go live, and surfacing the recurring problems that cause service degradation.

They stopped shipping features for half a year, now they're thriving

When incidents pile up fast enough, every part of the company bleeds: support is fielding angry customers, AEs are on apology calls, and engineering is burning cycles on retrospectives instead of shipping. For Eran Kampf (VP of Engineering at Twingate, Co-founder Monday.com) where the product is the network, that was the moment he made a call most engineering leaders won't: stop all feature work for a quarter and fix reliability.

Duty Scheduling 101: Building Reliable On-Call Coverage

Many teams start with a simple approach to on-call coverage. One person carries the phone this week. Someone else covers next week. Vacation requests are handled through emails, chat messages, or spreadsheets. When someone is unavailable, everyone is expected to remember who is covering. This works for a small team until the first missed alert. Duty scheduling is the foundation of reliable alerting.

Grafana & PagerDuty: Automate incident management with ServiceDesk Plus Cloud

������ �������� ������������ ���� ��������! This month, we're bringing you two new ManageEngine Marketplace extensions for ServiceDesk Plus Cloud that help bridge the gap between your monitoring tools and your service desk. With the Grafana extension, alerts automatically create and resolve tickets in ServiceDesk Plus Cloud—eliminating manual ticket creation and ensuring incidents are tracked the moment they're detected.

Office Relocation as an Ops Project: Runbooks, Rollback Plans, and Zero-Downtime Moves

Engineering teams that would never push a config change to production without review will happily move their entire company to a new building on the strength of a shared spreadsheet and a group chat. Then the first Monday in the new office arrives: the ISP install slipped two weeks, the badge system doesn't talk to the identity provider, on-call is paging someone whose desk is in a moving box, and the conference room where the incident bridge usually happens no longer exists.

ServiceNow Runs Your IT. PagerDuty Makes Sure It Never Stops.

For most enterprises, ServiceNow has become the backbone of IT operations, the platform where workflows are governed, compliance is maintained, and every incident, change, and request is tracked from start to finish. If you’re running ServiceNow, you’ve made a serious investment in how your IT operates. PagerDuty is built to make that investment work even harder.

Make the most of shift-based schedules

We recently updated our Schedules to better reflect how teams are currently managing their on-call responsibilities. Not everyone is working on weekly shifts or providing 24×7 coverage for all of their services, and that should be easy to schedule in our new tooling. To give you some examples, I’ve gone back through some of the questions we’ve gotten on the PagerDuty Commons over the past couple of years for questions about custom schedules that we weren’t really thinking about.

AI on AI Challenges

Building AI agents is easy until they launch into production and start behaving unpredictably. In this presentation, João Freitas, Chief AI Officer at PagerDuty, dives into the messy reality of scaling non-deterministic systems and shares how PagerDuty manages multi-agent complexities. Speaker: João Freitas, Chief AI Officer, PagerDuty Recorded during GenAI Community x Google Developer Group Lisbon at PagerDuty Portugal offices, July 2026.

Why Faster Recovery Beats Faster Shipping in the AI Era

A year ago, AI coding tools worked alongside developers—suggesting the next line, completing a function, accelerating work that a human was already doing. Today, they’re writing entire modules and services independently, producing code that no human has reviewed line by line, built from components that no single person has fully mapped. And adoption is only accelerating: According to our recent AI Resilience Survey, 84% of organizations are now using AI to write, review, or suggest code.

Why Modern IT Incident Response Needs Social Sentiment Analysis

IT operations teams face an ongoing battle against alert fatigue. Despite running sophisticated telemetry and baseline Application Performance Monitoring, engineers are often bombarded with notifications that lead nowhere. Relying purely on internal dashboards creates a massive visibility gap, and when critical incidents slip through the cracks, the financial damage is swift and severe. To close this gap, DevOps professionals are increasingly looking beyond traditional server metrics and turning to a surprising source for early warning signals: public social sentiment.

PagerDuty agent app in GitHub

PagerDuty's agent app shows live incident state, incident history and change correlations inside GitHub so you can get context right within your PR without interrupting your flow. Automatically correlate incident data with recent commits and deployments to identify root causes, then generate fix PRs with proper incident linking.#IncidentResponse.

PagerDuty agent app in GitHub: incident context where you already work

This blog post is part of PagerDuty’s ongoing series on how we’re helping customers navigate their journey toward autonomous operations. Read on to learn about the PagerDuty agent app in GitHub (Early Access) and how it builds toward this vision. How many tabs do you have open right now? And how many more do you open the moment an incident hits? Context switching during incident response is one of the most persistent sources of toil in engineering.

AI Orchestrations: Your easy button for proactive operations

This blog post is part of PagerDuty’s ongoing series on how we’re helping customers navigate their journey towards autonomous operations. Read on to learn about how AI Orchestrations builds towards this vision. “We should automate this.” Sound familiar? For many operations teams, that sentence never becomes action. Building event orchestration rules demands deep platform expertise, time no one has, and the ability to spot which patterns in your data actually matter.