Operations | Monitoring | ITSM | DevOps | Cloud

Monitor dependencies now available in the v3 API

Monitor dependencies are now available through the StatusGator v3 API. The new endpoint lets you programmatically retrieve the relationships and dependencies associated with a monitor, giving your integrations and internal tools more context about the services each monitor relies on. Dependencies are already available in the StatusGator UI for Website, Ping, and Custom monitors, while StatusGator automatically identifies relationships for many Service monitors.

Migrating from Patch Manager Plus On-Premises to Cloud: A practical evaluation guide

Ever heard of Murphy's Law of IT Administration? It states that the server hosting your security tools will go down at the exact moment a critical vulnerability is making headlines everywhere. Whether it's a sudden power shutdown, a database hiccup, or a local network failure, losing access to your central management tools right when you need them most is every IT team's worst nightmare.

The role of AI in website monitoring : How AI is rewriting the rules of website monitoring

A peak sale season, missed transaction or availability issues, spiking customer tickets, and unhappy customers. Well, you know the trope. A few years ago, this was just part of doing business online. Today, it’s a problem you can avoid, thanks to artificial intelligence. We’ve quietly reached an important turning point in website monitoring. For most of the internet’s history, monitoring meant setting thresholds: set a number, wait for it to be crossed, get an alert, and fix the issue.

Why AI Agent Orchestration Needs Runtime Context Between Agents

Every multi-agent system depends on one agent handing its output to the next, and nothing in the architecture confirms that the handoff carried what it should have. Orchestration adds a failure surface that single-agent architecture doesn’t have: a point between every two agents where one has to trust that the other passed along everything it needed, unverified.

Automate Your Entire Incident Response with Skylar Automation

See how Skylar Automation transforms incident response by orchestrating workflows across the tools your teams already use. In this demo, watch Skylar Automation respond to a critical service degradation by automatically creating a ServiceNow incident, paging the on-call engineer in PagerDuty, notifying the Microsoft Teams operations channel, and keeping updates synchronized across platforms. With Skylar Automation, teams can.

When to Use Grafana Assistant vs. MCP vs. gcx: Part 3

When should you use gcx? If Grafana Assistant is the brain and Grafana MCP is the easy hand, gcx is the power hand. Built for AI agents working in the terminal, gcx gives them deep access across Grafana Cloud—so they can pull telemetry, verify code, automate workflows, and access places MCP doesn’t. Coding agents? gcx. Need the full Grafana Cloud surface? gcx. Automating in CI/CD? gcx. Here’s where it fits, and when to use it — explained by Nicole van der Hoeven.

Internet Performance Monitoring: From Visibility to Control with LogicMonitor

Internet performance monitoring (IPM) gives IT leaders, operations teams, and network engineers visibility into the ISPs, carriers, and SaaS services their business depends on but doesn't control. In this LogicMonitor and Catchpoint webinar, Callum Brown (presales, EMEA, LogicMonitor) and Brandon Dunlap (solution engineering, Catchpoint, a LogicMonitor company) show how to operationalize IPM, moving from visibility to control.

MSP Observability: Proactive Monitoring to Autonomous IT with SCC Digital

SCC replaced fragmented tooling, including Nagios, with unified observability the whole team can use. The session covers proactive monitoring, SLA protection, and serving more customers without adding headcount per account. It's made for MSP leaders exploring AIOps for MSPs and observability for MSPs.

Storage Monitoring Tools and the KPIs Behind Each Failure Domain

When an application slows down, how long does it take to confirm whether storage caused it? The answer depends entirely on whether anything is collecting from the array itself. The server dashboard reports healthy CPU and memory, the network graphs look clean, and the array holding the data says nothing at all. Storage failures announce themselves late.

How Network Documentation Software Keeps Network Diagrams Current

When did anyone last open your network diagram and trust what it showed? A diagram drawn in a static drawing tool is accurate on the day it is saved. One quarter, two circuit upgrades and a hardware refresh later, it describes a network that no longer exists. Nothing warns you that this has happened. The file still opens, still prints, and still gets attached to change requests, which is what makes it risky during an incident.

Recurring Office Hours with the AI SRE Agent Team Behind AURA

Building an agent and not sure how to approach something? Bring it. AURA office hours are recurring working sessions with the people who build it. The team has been talking to people trying out AURA and hearing the same good questions come up more than once. Office hours are the answer to that: a standing slot on a schedule, rather than one conversation at a time. The format is deliberately loose. Nobody is arriving with thirty slides to spend an hour talking at you. The session goes wherever the questions go.

How to scale Alloy as a central telemetry gateway: capacity planning, load testing, and production lessons

Running Alloy as a single-instance sidecar is simple. Running it as a centralized gateway that absorbs the full telemetry stream of an enterprise platform—tens of millions of active series, terabytes of logs per day, and tens of thousands of trace spans per second—is a different challenge altogether. To get it right, you need deliberate capacity planning, honest load testing, and a monitoring setup that doesn't rely on the very thing you're testing.