Operations | Monitoring | ITSM | DevOps | Cloud

Cut AI agent cost and improve accuracy with Code Execution in the Datadog MCP Server

Observability investigations rarely follow a straight line. A latency question might cause an AI agent to start with a metric, pivot into traces, compare a deployment window, and finish by reducing thousands of logs to a few patterns. Each individual query is easy, but propagating context throughout an entire investigation can be tricky and expensive.

What Is a Network Interface (and Why Your Monitoring Tool Should Care)

A user opens a ticket: "The Internet is slow." IT checks the network. Bandwidth looks fine, no outages, no alerts firing. They check with the ISP, but nothing on their end either. Everything upstream checks out, and yet the user is still stuck watching a spinning wheel. What rarely gets checked is the one component sitting closest to the problem: the machine's own network interface.

17: There's No Life Without AI: Agents, MCP, and the Future of Automation With Viktor Farcic

On this episode of Kubex Talks, technology critic Viktor Farcic returns to talk with Andrew Hillier about the rapidly changing landscape in tech. Viktor has gone from AI skeptic to believer, claiming that there really is no life without AI anymore, from a professional standpoint.

How to Send Critical Alerts to the OnPage App | 3 Ways

What are the different ways to send a critical alert to the OnPage app? This video shows three ways people and external systems can trigger a high-priority OnPage mobile alert that's also HIPAA compliant (secure or healthcare use cases): These options are particularly useful for users with OnPage mobile licenses who do not have a Silver or Gold plan.

Best LLM inference providers 2026: 16+ on cost per outcome

An LLM inference provider hosts open-weight models like Llama, DeepSeek, and Qwen behind a pay-per-token API, handling GPUs, scaling, and serving for you. The same Llama 3.3 70B model ranges from $0.10 to $1.04 per million input tokens depending on who serves it, so provider choice is a pricing decision. Top picks as of September 2026: Groq and Cerebras for speed, DeepInfra for price, Together and Fireworks for breadth, Baseten for custom models.

Distributed Tracing Is Now in Beta for Ruby, PHP, and Python

A request comes in, enqueues a job, and returns. Twenty seconds later the job runs, and it’s slow. You have a trace of the request and a trace of the job, and nothing joining them. Time Detective has always helped you reconstruct what happened. Now we join it up for you, across applications, services, background jobs and infrastructure, even when they’re built in different languages.

Self-Improving Agents: A Practical Guide to Continuous Learning

We build agents to take work off engineers’ plates. Then we give those engineers a new manual job: reading failed runs and babysitting prompts. Agents will improve themselves automatically. We’re not there yet, but this is the future I’m betting on. We’ve been working on this ourselves at Komodor over the past year. We know how hard it is to turn a failure into an improvement that holds up beyond a few examples.

What Is a Network Topology Diagram? Types, Examples and How to Build One That Stays Current

Most network diagrams are accurate exactly once: the day they are finished. The network keeps changing, the drawing does not, and the gap shows up during the next outage. According to the Uptime Institute Annual Outage Analysis 2026, failure to follow established procedures remains the leading driver of human-error outages. A wrong diagram is how a right procedure hits the wrong port. The fix is a network topology diagram that matches the live network topology.

Incident Management System: What It Is and How to Choose One

An alert fires. A ticket opens. Someone gets paged. Then the real work begins: gathering context, finding the affected service, deciding who owns the issue, running diagnostics, applying a fix, validating recovery, and documenting the result. Many IT teams assume that an incident management system is simply the application that opens and tracks the ticket. That is part of the job, but it’s not the whole operating model.