Operations | Monitoring | ITSM | DevOps | Cloud

The latest News and Information on Monitoring for Websites, Applications, APIs, Infrastructure, and other technologies.

Cloud Incident Management: Process, Tools, and Practices

How do you resolve an outage your organization has no authority to fix? A managed database drops into read-only mode and stops accepting writes. There's no host to reach, no configuration file to edit, and no restart command available to your engineers. Cloud incident management begins at that boundary, where the response depends on a support channel and a provider status page. Plenty of what you already know still applies here.

Build and Launch AI Agents from Your Splunk Workflows

Introducing the Splunk Agent Launchpad! Let’s face it—your team is busy. Between managing alerts, digging through investigations, and constant context-switching, it’s hard to stay ahead of the noise. What if you could turn your existing operational knowledge into custom AI agents that do the heavy lifting for you? And the best part? No coding required. Watch this exclusive look at the Splunk Agent Launchpad. We’re showing you how to build, deploy, and manage AI agents that help you investigate, enrich, summarize, and act—all without leaving the Splunk environment you already know and trust.

August 2026 product update: hosted MCP and more

Your MCP client doesn’t need your whole API key just to look up an error anymore. Honeybadger's hosted MCP server now supports OAuth. You can approve it through your browser, scope your permissions, revoke your permissions, and rest easy knowing that our tokens auto-refresh and don’t sit around in a config. Keep reading to see how it works and get a quick recap of everything else that shipped this cycle.

The Most Important Improvements Are Often the Ones You Never See

When organizations evaluate software platforms, attention naturally gravitates toward visible outcomes. New capabilities, expanded functionality, improved user experiences, and innovative technologies often dominate conversations about platform value. These improvements are important because they directly influence how teams interact with technology and how organizations achieve business objectives.

Observability's Sixth Sense: Grounding Anomaly Detection in Reality

Summary: Machine learning-based anomaly detection improves observability by learning normal system behavior instead of relying only on static thresholds. This article explains how vmanomaly, its MCP server, purpose-built skills, and an LLM-powered UI copilot help engineers explore telemetry, investigate anomalies, build MetricsQL queries, select suitable models, apply business constraints, and validate configurations through natural language.

Monitors now support the HTTP QUERY method.

If your API answers on QUERY, you can now point a monitor straight at it. QUERY joins HEAD, GET, POST, PUT, PATCH, DELETE, and OPTIONS in the HTTP method list for HTTP, keyword, and API monitors. You can find it under the “Advanced settings” when editing or adding your monitor. QUERY became a standard in June 2026 as RFC 10008. It’s safe, idempotent, and it carries a request body.

Icinga Director: Targeted Health Checks and Sync Rule/Import Source Deletes from CLI

Icinga Director’s CLI commands have supported import sources and sync rules for a long time, which includes listing, checking and running them. Before Icinga Director v1.11.6 two things you couldn’t do from the CLI, though, were narrowing a health check down to a single object, or deleting an import source or sync rule without opening the web UI. I’ll walk through both, using examples.

Platform engineering is not just a developer trend, but a practice ITOps should be paying attention to

Riya has managed IT operations at a mid-sized FinTech company for six years. She knows the infrastructure inside out: Every server, monitoring alert, and compliance requirement is owned by her team. So when Riya heard the engineering lead mention their new internal developer platform in a quarterly review, she assumed her team would be looped in eventually. This did not happen. Three months later, Riya's team was called in to investigate an outage.