Operations | Monitoring | ITSM | DevOps | Cloud

The latest News and Information on DevOps, CI/CD, Automation and related technologies.

AI didn't kill tech debt, it just changed the currency you pay it in.

Ganesh Datta on why every team still has a finite budget, now it's tokens instead of headcount. $500 to spend: ship the feature or fix the P2? The orgs building a framework for that call now will have it a lot easier when the CFO puts a cap on spend. From Braintrust by Cortex. Full episode out next Thursday.

API-first DCIM: Reduce Integration Friction and Keep Control Across Tools

Disconnected tools slow your operations down more than missing data ever could. Every day, teams waste hours stitching together systems that don’t naturally talk to each other. API-first DCIM cuts through that drag, turning scattered signals into one clear operational view. With Hyperview DCIM, you connect faster, keep control, and make decisions without extra manual work.

How we create a Canonical Academy exam

Open source provides the world with access to cutting-edge software, and the learning that comes with it. But how do you validate someone’s skills in an open ecosystem? Canonical Academy is a highly rigorous, job-focused qualification platform designed to empower individuals and enterprises with industry-recognized credentials. The platform addresses a critical gap in tech: validating real-world, hands-on capability rather than rote memorization.

Cloud Cost Management for Observability: A Practical Guide

Observability spend is outgrowing infrastructure budgets. What drives the cost up, how pricing models work, and a practical framework to manage it. Sejal Pandey works on content and growth at Last9, writing about observability, reliability, and SRE practices.

Why you should (not) build your own observability stack

If you are able to build it better than your vendor, then change your vendor. Not build it. Rishi builds large-scale observability systems at Last9, focusing on reliable and cost-efficient telemetry infrastructure, and writes about the practical lessons learned while operating ClickHouse, VictoriaMetrics, and OpenTelemetry in production.

Why compliance keeps slowing your releases (and what to change first)

A team ships at a steady pace for most of the year. Then an audit approaches, and delivery slows. Engineers get pulled off feature work to support the audit, producing the configuration exports, logs, and environment checks that the evidence depends on. The slowdown lasts as long as the audit does. It is tempting to read this as a team that needs to move faster or be bigger. It is usually neither.

How Is AI Changing IT Operations? Building Production-Ready AI Agents with Alex Zinovy

How is AI changing IT operations, and what does it take to move AI agents from impressive demos to production-ready systems? In this episode of Agents of IT, Resolve’s Zack Austin sits down with Alex Cinovoj, Founder and CTO of TechTide AI, to explore what enterprise AI looks like when it has to work in the real world. Alex brings years of hands-on IT, infrastructure, DevOps, and AI engineering experience to a conversation about the shift from experimenting with AI to building trustworthy systems that deliver measurable outcomes.

LLM token cost: pricing per token explained

LLM token cost is the price a provider charges per token a model reads or writes, quoted in dollars per million tokens. Input and output bill at separate rates, with output priced at roughly 5x input. As of September 2026, published rates range from under $0.10 to more than $180 per million tokens on top-end reasoning tiers. In late 2025, Hardik Sonetta of Thomson Reuters Labs published a warning about the most common prompt caching mistake in production.

Shipped: Rightsize Kubernetes workloads without leaving your MCP client

Changing a Kubernetes resource request takes two numbers: what the workload requests, and what it uses. The CloudZero MCP server now returns both, by cluster, namespace, or workload. This gives you a number you can defend. Usage comes back as P95 over the date range you query, 30 days by default. When an engineering lead asks whether a service runs on a smaller request, that is the figure that settles it. Over-provisioning and under-provisioning show up on the same query.

Our Customer Success AI bill tripled. Here's why we're spending more.

Pop quiz: If you spend $40,000 per month on Anthropic, and you’ve got two customers, what’s your cost per customer? If you bypassed the easy answer of $20,000 and said, “Scott, you old trickster, that’s not enough information to answer that question,” you’ve won today’s prize: a lesson in the perils of average costs. Let’s flesh out the situation: You put an AI feature in your product, a document assistant powered by Claude.