Operations | Monitoring | ITSM | DevOps | Cloud

Making agentic token costs visible in production

In some organizations, high token counts have become a proxy for productivity. Some engineering teams are being pushed to max out context windows and wire in sprawling tool sets. More tokens can mean better agent reasoning and richer context during development, but token costs compound in production. Tokens accumulate across sessions, users, and tool calls in ways that are easy to overlook. Datadog’s 2026 State of AI Engineering report quantifies the scale of this problem.

OpenSearch 3.6: Agentic Applications Meet Long-Term Support

TL;DR OpenSearch 3.6 makes agentic search production-ready, with the AI-powered Launchpad provisioning full search apps in minutes and faster default vector search, and it's the first LTS release, bringing 18+ months of guaranteed support, SBOMs, and an upstream-first commitment (every fix goes back to the main project) so teams get fast-moving open source and a stable, supported platform at once.

How to Use Your Knowledge Base to Increase AI Chatbot Deflection

Ticket deflection is the metric IT leaders point to when they talk about AI chatbot ROI, and the knowledge base is the part of the equation that determines whether that number moves. A chatbot can run natural language processing well and still deflect almost nothing if the content behind it is thin, outdated, or scattered across articles that don't match how people actually ask questions.

Why Cash Flow Still Matters in an AI-Driven Economy

Artificial intelligence is changing how businesses operate. Companies are using AI tools to automate customer service, generate content, analyze data, improve forecasting, and streamline everyday tasks. For many business owners, the promise is simple: work faster, reduce costs, and improve efficiency.

Building AI SRE Agents, Part 1: Start Local, Break Things, Learn Fast

The first stage of AI SRE maturity is a laptop, a throwaway cluster, and zero production access. Here’s how to set it up, and what to watch for. AI SRE (Site Reliability Engineering) agents are AI-powered systems that automate the most time-consuming parts of incident response: triaging alerts, correlating logs and metrics, generating root-cause hypotheses, and proposing remediation steps.

Claude Code Monitoring at Scale: Gateways and Routing With OpenTelemetry

Chelsea and I recently wrote a guide on how we monitor Claude Code usage internally with Bindplane. TLDR; We remotely manage a Bindplane Distribution of the OpenTelemetry Collector (BDOT) that runs on every engineer's laptop. This setup is great, but it has one downside. Sending to Google Cloud Monitoring, Swarmia, and any other destination directly from an engineer’s laptop is limited to local processing. You can’t get the benefit of centralized routing and processing on a gateway.

Rethinking Sprite Creation Costs for Indie Developers

The gap between game design ambition and art production has never been more visible. Over the past twelve months, a growing number of indie teams have discovered that the bottleneck isn't always code, mechanics, or level design-it's the sheer volume of sprite frames required to bring a single character to life. A walking cycle alone can consume an entire weekend. A complete character with idle, attack, and jump animations often stretches into weeks of pixel-by-pixel work.

The Future of Governing AI Agents in Enterprise Order Processing

Governing AI agents are rapidly reshaping how enterprises manage complex, high-volume order workflows, from automated validation to exception resolution and fulfillment routing. As AI models become more reliable and context-aware, organizations are shifting from human-heavy processes to autonomous systems that can handle end-to-end order lifecycle management with minimal manual intervention.