Operations | Monitoring | ITSM | DevOps | Cloud

The latest News and Information on Monitoring for Websites, Applications, APIs, Infrastructure, and other technologies.

Grafana Campfire - Assistant powered Dynamic dashboards - (Grafana Community Call - August 2026)

Many times, it feels like you're maintaining multiple versions of the same dashboard (with a slight modification), *OR* simply spending more time writing queries rather than actually looking at the actual data? In this Campfire community call, we're taking a deep dive into two things that are reshaping how Grafana dashboards get built: Dynamic Dashboards and the Grafana AI Assistant and showing you how to combine them to go from a blank canvas to a reusable, production-ready dashboard in minutes.

Trace an AI SRE Agent: AURA Docker Quickstart with Phoenix and OTel

You get an answer from the agent and no way to check how it got there. The route it took is recorded, and so is the reason it gave for taking it. AURA emits OpenTelemetry spans, and the Docker quickstart wires them straight into Phoenix. Four services come up together: AURA Web Server as the persistent agent harness, LibreChat as a browser interface for chatting with the agent, Phoenix to receive the spans, and MongoDB to store stateful data for LibreChat. The Compose file arrives pre-configured to point AURA at Phoenix and to enable content recording for the local demo.

Monitor dependencies now available in the v3 API

Monitor dependencies are now available through the StatusGator v3 API. The new endpoint lets you programmatically retrieve the relationships and dependencies associated with a monitor, giving your integrations and internal tools more context about the services each monitor relies on. Dependencies are already available in the StatusGator UI for Website, Ping, and Custom monitors, while StatusGator automatically identifies relationships for many Service monitors.

Migrating from Patch Manager Plus On-Premises to Cloud: A practical evaluation guide

Ever heard of Murphy's Law of IT Administration? It states that the server hosting your security tools will go down at the exact moment a critical vulnerability is making headlines everywhere. Whether it's a sudden power shutdown, a database hiccup, or a local network failure, losing access to your central management tools right when you need them most is every IT team's worst nightmare.

The role of AI in website monitoring : How AI is rewriting the rules of website monitoring

A peak sale season, missed transaction or availability issues, spiking customer tickets, and unhappy customers. Well, you know the trope. A few years ago, this was just part of doing business online. Today, it’s a problem you can avoid, thanks to artificial intelligence. We’ve quietly reached an important turning point in website monitoring. For most of the internet’s history, monitoring meant setting thresholds: set a number, wait for it to be crossed, get an alert, and fix the issue.

Monzo's Stand-In Held Up on Wednesday. Some Customers Still Had a Bad Day.

On Wednesday 19 August, Monzo had an outage. DownDetector logged more than 3,000 reports by midday. Monzo’s own statement was direct about what it did next: it activated Monzo Stand-in, its fully independent backup bank, while it investigated an issue affecting customers. By the end of the day, Monzo said the issue was resolved and all services were back. Stand-in did roughly what it was built to do, keeping essential banking functions running while the primary platform had a problem.

Why AI Agent Orchestration Needs Runtime Context Between Agents

Every multi-agent system depends on one agent handing its output to the next, and nothing in the architecture confirms that the handoff carried what it should have. Orchestration adds a failure surface that single-agent architecture doesn’t have: a point between every two agents where one has to trust that the other passed along everything it needed, unverified.

Automate Your Entire Incident Response with Skylar Automation

See how Skylar Automation transforms incident response by orchestrating workflows across the tools your teams already use. In this demo, watch Skylar Automation respond to a critical service degradation by automatically creating a ServiceNow incident, paging the on-call engineer in PagerDuty, notifying the Microsoft Teams operations channel, and keeping updates synchronized across platforms. With Skylar Automation, teams can.