Operations | Monitoring | ITSM | DevOps | Cloud

What's new from BigPanda: September 2026 Product Updates

Most teams we talk to are fighting the same battle. The knowledge needed to make a decision already exists somewhere. However, it’s locked in a tool your team isn’t looking at, or in the head of the one engineer who’s seen this same type of incident before. This month’s updates all chip away at that same problem. Here’s what’s new.

The enterprise changed. ITOps didn't.

The modern enterprise runs on a technology stack that changes faster than the operating model responsible for keeping it available. Applications that once moved through scheduled releases now change continuously. Infrastructure is distributed and dynamic. Services depend on other services, teams depend on other teams, and operational data arrives from more places than any individual can reasonably inspect. The business asked for agility, flexibility, and velocity, and the tech delivered.

Introducing Swarm Investigation from BigPanda: Autonomous, multi-agent IT incident investigation

When a major incident opens, the opening minutes often become a race across disconnected tools and competing theories. One engineer checks a monitoring tool. Another scrolls through change records, looking for the one line that explains everything. A third pings Slack, asking if anyone has seen this before. While these are reasonable steps, taken one at a time, in sequence, they are far too slow.

Extending autonomous L1 ops with new suppression and runbook capabilities

Earlier this year, I had the chance to meet with one of our airline customers. During the meeting, we discussed how to use agentic technology to automate L1 workflows. As one of the largest global airlines, they have many applications and service teams focused on flight-critical, tier 1 environments. Any downtime can cause costly delays and unhappy customers.

The high cost of low-quality L1 NOC outsourcing

New research from BigPanda reveals what enterprises spend on outsourced IT operations support, what they get in return, and why leaders are ready to rethink the model. Enterprises spend an average of $5.4 million a year on outsourced IT operations support. That’s a substantial investment. It should produce reliable frontline operations: incidents detected, understood, routed, and resolved with the speed and accuracy the business expects. The research shows a different picture.

Introducing the next generation of the BigPanda AI Incident Assistant

Effective incident response depends on having all of the context surrounding what’s happening. You have to understand your systems, services, architecture, and teams deeply enough to correctly interpret whatever alert just fired. Too often, that context doesn’t arrive packaged neatly in one place. Gathering and interpreting context correctly under time pressure is one of the most difficult parts of the job.

What data sources does agentic ITOps use

Agentic IT operations have arrived. It’s no longer a question of if enterprise IT departments will adopt agentic ITOps, but how quickly. The question we hear most often at BigPanda isn’t “what are agentic ITOps,” it’s “what data do we actually need to get started?” That’s the right question to ask. Agentic AI is only as good as the data and context that feeds it. Real-time observability and telemetry data from machines. Structured ITSM and workflow records.

IT problem management VS. IT incident management, and how agentic ITOps improves both

Picture a familiar scene: a critical application goes down during peak business hours, and your on-call engineers scramble to restore service. Two weeks later, the same application fails again, frustrating your teams with the same symptoms, the same scramble, and the same customer frustration. If this pattern feels familiar, your organization may be strong at IT incident management, but underinvested in IT problem management.

From AIOps to agentic ITOps: Why AI for IT operations has entered a new era

Enterprise IT has reached an inflection point. Your teams are responsible for hybrid cloud infrastructure, microservices, third-party dependencies, and shipping AI-generated code at unprecedented velocity. IT environments are becoming more complex faster than traditional tools and processes can keep pace. Alert volumes keep climbing. Institutional knowledge keeps walking out the door. And the pressure to do more with flat or shrinking budgets isn’t letting up.

What is MTTR, and how can agentic ITOps reduce it?

Mean time to resolution (MTTR) measures the average duration to restore regular operation for an application, service, or infrastructure component. It’s a key performance indicator (KPI) for IT incident management. To tie MTTR directly to customer satisfaction, you first need to understand how it affects service and application reliability and availability. From there, you can make informed decisions, operate efficiently, and provide a seamless customer experience.