Operations | Monitoring | ITSM | DevOps | Cloud

The latest News and Information on Monitoring for Websites, Applications, APIs, Infrastructure, and other technologies.

Cut AI agent cost and improve accuracy with Code Execution in the Datadog MCP Server

Observability investigations rarely follow a straight line. A latency question might cause an AI agent to start with a metric, pivot into traces, compare a deployment window, and finish by reducing thousands of logs to a few patterns. Each individual query is easy, but propagating context throughout an entire investigation can be tricky and expensive.

What Is a Network Interface (and Why Your Monitoring Tool Should Care)

A user opens a ticket: "The Internet is slow." IT checks the network. Bandwidth looks fine, no outages, no alerts firing. They check with the ISP, but nothing on their end either. Everything upstream checks out, and yet the user is still stuck watching a spinning wheel. What rarely gets checked is the one component sitting closest to the problem: the machine's own network interface.

Distributed Tracing Is Now in Beta for Ruby, PHP, and Python

A request comes in, enqueues a job, and returns. Twenty seconds later the job runs, and it’s slow. You have a trace of the request and a trace of the job, and nothing joining them. Time Detective has always helped you reconstruct what happened. Now we join it up for you, across applications, services, background jobs and infrastructure, even when they’re built in different languages.

What Is a Network Topology Diagram? Types, Examples and How to Build One That Stays Current

Most network diagrams are accurate exactly once: the day they are finished. The network keeps changing, the drawing does not, and the gap shows up during the next outage. According to the Uptime Institute Annual Outage Analysis 2026, failure to follow established procedures remains the leading driver of human-error outages. A wrong diagram is how a right procedure hits the wrong port. The fix is a network topology diagram that matches the live network topology.

When AI Agents Attacked Their Own Evaluators, the Industry's Own Leaders Started Asking for Guardrails

When AI agents attacked their own evaluators in July 2026, it exposed a gap no policy commitment can close. The OpenAI Hugging Face incident revealed that enterprise agent governance requires in-flow runtime controls, not retrospective auditing or industry safety agreements.

Bleemeo and ilert: two European companies, one alerting chain

Some alerts only need to reach a Slack channel. Some need to reach one specific person, at 3am, and keep trying until they answer. For the second kind, we are partnering with ilert — an incident response platform covering the full lifecycle: from the moment an alert arrives, through paging the right responder, coordinating the response, telling customers what’s happening, and learning from it afterwards. The integration is live today, on both sides.

Why FIPS Mode Is Not Enough: What Federal Teams Should Expect Their Vendors to Prove

Federal teams need more than a system setting. They need defensible evidence that the cryptography protecting federal information is validated, correctly configured, and actually used. For platforms such as ScienceLogic, that product-level evidence should come from the vendor, not be reconstructed by the customer.

Enterprise AI governance framework: A practical guide to governing AI

AI adoption is accelerating across enterprises, but governance isn't necessarily keeping pace. As AI becomes part of everyday business workflows and applications, organizations need to understand where it is being used, what data it can access, and who is responsible for managing the risks. ManageEngine's shadow AI researchhighlights this challenge.