Operations | Monitoring | ITSM | DevOps | Cloud

Shipped: Project this month's AI cost before the invoice closes

The question comes up on the 10th, the 15th, and again on the 25th. Where is AI spend going to end up this month? The invoice won’t tell you until it’s closes, and by then there’s nothing left to forecast. “If I’m looking at this on the 15th, I want to know where we’re going to land.” That’s how a finance lead put it during a persona session in September, and it’sthe whole job. You have half a month of real usage behind you.

Claude Opus pricing in 2026: every model, every rate, and whether it's worth it

Claude Opus pricing is $4 per million input tokens and $20 per million output tokens on Claude Opus 5.5, the current model, with cache reads at $0.20 and batch jobs at $2/$10. Opus 5 and the legacy 4-series bill at $5/$25. The 1M context window carries no surcharge. Every Opus model Anthropic shipped in 2026 held the same line: $5 in, $25 out, per million tokens. Opus 4.6 in February, 4.7 in April, 4.8 in May, Opus 5 in July. Four releases, one price. On September 22, 2026, the line broke.

10 Best AI Agent Infrastructure Platforms in 2026

AI agent infrastructure is the set of platforms that run agents and the code they write. It has three layers: sandboxes that isolate untrusted, model-generated code, runtimes that run the agent and its services in production, and orchestration layers that save an agent’s progress so a long run can resume after a failure. Most production agents need more than one layer. This guide compares 10 platforms across all three, with isolation, state, deployment, and compliance for each.

Shipped: Views now work for every access level

Most people who open CloudZero care about one slice of the spend, like their team, their product, or their region. A View gives them that slice in one click, with the grouping and filters already set, so nobody has to rebuild the same Explorer query every week. Views now work for everyone in your organization, including people with scoped access.

That 2am incident cost you more than downtime

It's 2:14am. A pager alert wakes an on-call engineer for a checkout failure hitting a slice of customers. By the time they've pulled the right logs, cross-checked the deploy history, and finally gotten the bug to happen again in front of them, the sun's coming up. The incident report will list two hours of downtime. It won't say anything about the day that engineer just lost, or the fact that nobody on your team could have told you in advance how long that reproduction step was going to take.

10 Best Multi-Cloud Management Platforms in 2026

A multi-cloud management platform is software that lets you run, provision, govern, or optimize workloads across two or more environments, such as AWS, Google Cloud, Azure, and your own data centers, from one place. The products in this category do different jobs: some run your applications across clouds, some provision infrastructure as code, some govern hybrid estates, and some optimize cost. Most teams combine two or three.

18: Building an Agentic Future: AI and Optimization with Sachin Gharge

On today's episode, Andrew Hillier chats with Sachin Gharge, Head of Cloud Platform at Scandinavian Airlines (SAS). They discuss AI, agents, Kubernetes, MCP, and optimization. Sachin shares how he and his team are optimizing cloud costs, leveraging automation, and experimenting with agentic AI, including bots and Slack integrations, to make operations easier and more effective for developers and the business.

Run your first workflow in minutes, no sales call

The regression nobody catches passes a busy review and ships. An off-by-one, a change that reads as sensible and quietly breaks something, gets a nod from a tired reviewer and lands in production, where it erodes trust one small defect at a time. You can have an AI code reviewer running on your own repository in the time it takes to read this page. Get started without having to book a demo or contact sales.

Shipped: See what your AI spend is actually paying for

Most AI spend comes in with no tags and no owner attached. Your provider console shows total spend, maybe broken out by API key or model. It won’t tell you that the sales team spent $1,700 on Claude this week, let alone what the work was. And the problem is growing. McKinsey found that 56% of organizations now use AI in three or more business functions. More teams means more spend, and most companies respond with a spending cap. Set it too low and you slow down the work you wanted AI to help with.

How to Guarantee a Website or Service Never Goes Down (And What You Can Actually Promise)

No one can guarantee that a website or service never goes down. What you can promise is a measured availability target, and with a multi-location, active-active design you can reach 99.999% (five nines), about 5 minutes 15 seconds of downtime a year. That takes redundancy at every layer, automatic health-based failover across regions and ideally providers, safe deployments, failure testing, and outside-in monitoring. Control Plane is built for that tier.

Azure in Bleemeo: your subscription next to your servers, with one read-only role

Most teams that run on Azure do not run only on Azure. There is a database on a VM nobody wants to move, a Kubernetes cluster somewhere else, a few servers in a rack, and a monitoring setup that grew around all of it. Azure Monitor sees the Azure part very well and nothing else, so the picture of an incident ends up split across two consoles, two alerting configurations and two sets of dashboards. Bleemeo now connects to Azure the same way it already connects to AWS.

Shipped: Start every session where your work lives

Most people who use CloudZero spend their time in one or two places. For some it’s AI Signals, and for others it’s Optimize or Anomalies. If Explorer isn’t one of those places, every sign-in starts with a click to get where you need to be. Dates and numbers are another friction. A date like 04/07 means April 7 in the US and July 4 in much of Europe. When the platform shows a format your team doesn’t use, you end up having to convert each value before you can work with it.

Moving Mainframe Data to Snowflake and AWS Through Apache Kafka: Treehouse Software and meshIQ

Treehouse Dataflow Toolkit moves mainframe data from Db2, VSAM, and IMS through Apache Kafka pipelines into Snowflake and AWS targets. meshIQ keeps the streaming layer visible and stable. Together they give data science teams continuously updated enterprise data for AI and ML.

Get your agents off laptops and onto shared infrastructure

There's a specific, recognizable point where a team's use of AI agents changes shape. Not when they adopt agents; most teams already have. It's when agents stop running on someone's laptop and start running on infrastructure that the whole team can see. This is a real technical shift, not a policy change or a maturity score. Here's specifically what's different on each side of it.

How Honeycomb Private Cloud Drinks From the Fire Hose

Back in July, I wrote about how the Tenant team (the team behind Honeycomb Private Cloud (HPC)) has embraced the code review bottleneck to focus more of its work. One of the other challenges we have is that we're downstream of almost all the other teams at Honeycomb, meaning that we have to package up everyone's code and services, and how it gets provisioned! This is something impossible to handle through code review since there are so many engineers on other teams, and so few of us.

Why engineers ignore cloud cost governance (and fixes)

Discover why engineers ignore cloud cost governance and how to build developer cost accountability. Learn how Harness helps empower engineering teams. Engineers often overlook cloud costs due to friction in traditional FinOps tools and a lack of real-time visibility. By embedding automated guardrails and shift-left cost insights into developer workflows, organizations can drive accountability without slowing velocity.

Where do you see organizations hitting their limits?

In this clip, Virtana Chief Product Officer, Amit Rathi explains why more data does not automatically lead to better operations. As system complexity grows, organizations are collecting more telemetry than ever while struggling to turn it into actionable insights. At the same time, rising observability costs are forcing some teams to monitor only part of their environments. Watch the video to learn why intelligence, not just visibility, is becoming essential for modern IT operations.

Azure Monitor pricing: What you pay vs. what you get

Azure Monitor doesn't have a price. It has a bill and those are very different. There's no plan tier to pick, monthly set fee, or simple number to sanity-check against your budget. Instead, you're charged across a handful of separate meters: log ingestion, log queries, retention, alert rules, web tests, and custom metrics. Each of these metrics scale independently. Most teams don't see the full cost until the bill arrives.

Cycle's Hosted MCP Release Announcement

We are stoked to announce that we have released our very own hosted MCP that allows you to use natural language to interact with and control anything you have running on Cycle.io. Alexander Mattoni, CTO and co-founder of Cycle, shares a few scenarios where the MCP can be useful. With the MCP, Cycle users can communicate with the Cycle platform directly from their preferred AI tool. In minutes, users can provision a new server, deploy an application, or troubleshoot an issue directly from their AI assistant like Claude, Claude Code, ChatGPT or Codex.

Cloud Shell, New Integrations, and More

VirtualMetric DataStream now includes a PowerShell console in the browser, more than a dozen new integrations, and a new way to collect data from servers and virtualization hosts, where teams choose exactly what they collect. This update also brings guided learning for new users, built-in monitoring rules, and new controls for organizations managing branding and sign-in. Here’s what’s new.

Shipped: Get alerted when AI spend spikes, with the cause attached

AI spend now comes from every department, and it can double in a week without anyone deciding it should. The invoice arrives after the month closes. By then the usual response is a spend cap, which slows every team, including the ones getting real work done with AI. You need to know about spend that breaks its normal pattern while there’s still time to act. The alert should reach the person who can act on it, with proper context and detail.