Operations | Monitoring | ITSM | DevOps | Cloud

Building an AI Observability Agent: Lessons from the Trenches - Stripe at O11yCon 2026

Stripe shares lessons from building an incident investigation agent, from context-window blowups to why the final 5% still needs a human. In this O11yCon 2026 talk, they dig into what it takes to go from 'it works' to 'it works reliably,' including how pointing agents at like Honeycomb's speeds up on-call investigations.

Signal vs. Spend: Building Cost-Aware Observability at Slack - O11yCon 2026

It started with a single log line taking up a massive amount of volume: 500 million emissions per hour. Pulling that thread led Emma and Steven into Slack's broader logging pipeline: 311 billion logs per day at 4.4M/sec peak, with no volume limits, no per-service attribution, and no feedback to the teams generating the noise.

Your AI agents don't need better models. They need the same context.

Everyone on the team has good coding agents. That is not what made the team fast. The shared repo, the shared docs, and a glossary nobody is allowed to drift from did more than any model choice. In this Product Highlights conversation, Patrick Dawkins, a principal engineer at Upsun who is building Upsun Dispatch under a hard deadline, explains what his team changed to sustain that pace. His take: "As well as the team all having access to the same things and the same vocabulary, all the agents also have access to those docs.".

Netdata Network Topology: Live SNMP Discovery & Container Connection Maps

Netdata now builds live network topology maps directly in the agent: no scheduled scans, no stale picture the next morning. In this walkthrough, we cover both sides of the new Topology view: Network device discovery (SNMP): Container & process network connections.

How to install Kubernetes using OpenShift's CLI | Site24x7

Running Kubernetes on Red Hat OpenShift adds powerful enterprise capabilities—but also introduces operator-driven workloads, stricter RBAC and SCC policies, and platform-specific complexity. In this video, learn how Site24x7 enables platform-aware monitoring for OpenShift environments, helping DevOps and platform teams gain complete visibility without blind spots.

Site24x7 Free Training Day 2: Infrastructure monitoring, custom plugins, cloud cost management

This session covers everything you need to know about infrastructure monitoring, starting from agent-based server monitoring to database monitoring, container monitoring (Docker & Kubernetes), multi-cloud monitoring (AWS, Azure, GCP, OCI), IT automation, custom plugins, and ManageEngine CloudSpend for cloud cost management. If you're looking to master full-stack infrastructure observability, this hands-on walkthrough shows you exactly how to set up and use each Site24x7 module inside the live console.

Product leaders talk safer, faster releases and deeper analysis with Bits | This Month in Datadog

In July’s This Month in Datadog, Jeremy is joined by Datadog product leaders for in-depth conversations about how Bits enables you to confidently evaluate and release features containing AI-generated code, and use natural language to ask, understand, and act across Datadog.