Operations | Monitoring | ITSM | DevOps | Cloud

Shipped: One CloudZero for everyone, starting October 1

On June 3, we made the new CloudZero experience the default for every customer. Since then, we’ve shipped around 30 improvements a week: side-by-side period comparisons in Explorer, budgets you can create and edit right in the app, threshold alerts on dashboard tiles, and Monitors, which flags AI and cloud spend that moves outside its normal pattern and shows you what changed. Pages load 28 to 61% faster. JavaScript execution is 85% faster.

AI cost allocation: how to attribute AI spend by team, product, and customer

AI cost allocation is the practice of attributing every dollar of AI spend to the team, product, feature, or customer that generated it. That spend includes API tokens, GPU compute, per-seat tools, and shared infrastructure. It's harder than cloud allocation because AI spend arrives untagged, spans vendors, and pools in shared resources. Four methods cover most cases: tag-based, key-based attribution, proportional split, and usage-telemetry.

Why Engineers Ignore Cloud Cost Optimization & Fixes

Learn why engineers ignore cloud cost optimization and how to build a culture of FinOps governance. See how Harness helps. Engineers often overlook cloud costs due to lack of visibility, fragmented tooling, and competing delivery priorities. Organizations can fix this by embedding FinOps guardrails into developer workflows and providing real-time cost feedback during build cycles.

Shipped: Every AI provider, one cost story

If you were anywhere near LinkedIn last week, you probably saw us launch AI Signals. We weren’t exactly quiet about it. (Press release, a couple of blog posts, and more social posts than we’d like to admit. Sorry about your feed.) We covered the why behind AI Signals already, but I wanted to actually walk you through what you’re seeing on the screen. Sooner or later someone asks what the company spent on AI last month.

Shipped: Don't ask an AI agent what its work will cost

If you set the budget for your team’s AI agent work, or answer to someone who does, you need a rough idea of what a job will cost before it starts. That’s hard to get. Stanford researchers found the same agent, given the same task, can use up to 30 times more tokens from one run to the next, and you usually find out afterward. Most developers just run the job.

Cost per AI outcome: tying AI spend to results

Cost per AI outcome is your total attributed AI spend divided by the business results it produced: resolved tickets, converted leads, merged pull requests. It includes the cost of failed attempts, sits at the top of the AI unit-cost ladder, and it's the number that makes vendor outcome pricing, ROI claims, and build-versus-buy decisions comparable.

Shipped: A customer support experience that starts with an answer

When you have a question about your cloud or AI spend, you want an answer quickly, not a ticket that disappears into a queue. Support should not mean waiting for business hours, repeating your account details to multiple people, or wondering whether anyone picked up your message. That changed this week for every CloudZero customer. You get answers to most product and account questions immediately, at any hour, and when your question needs a person, they already have context.

We stopped asking an LLM how much its own work would cost

There’s a specific kind of measurement problem worth naming precisely rather than dramatizing: this month we found that our model-routing agent was assigning a token budget to every unit of work, and that budget was noise in the strict sense. Fixing it meant improving a system that’s mostly right, not tearing one down.

AI finally plans like every other line in my budget

September is associated with football, foliage, flannel and, for some, the Financial Plan. As we put pen to paper (or agents to harnesses), there’s a few core elements that have always driven the P&L outlook for the following year: rep productivity and new product releases driving new sales, expansion and contraction against the install base, employee roster changes, and discretionary spend.

Stop capping your best people.

Somewhere in your company, a team is three weeks into the AI project that’s going to matter. Somewhere else, a support pilot from the spring is still summarizing every ticket with a frontier model, and nobody has looked at it since it started working. On the invoice they’re identical, and the company has two moves: leave everything open, which funds the waste, or cap everyone, which kills the bet.

Why Cloud Cost Optimization for Engineers Fails

Learn why cloud cost optimization for engineers fails and how to fix it with developer-centric FinOps practices. See how Harness helps. Engineers often ignore cloud costs due to a lack of visibility, context, and ownership in their daily workflows. By shifting cost governance left and integrating real-time cost insights into CI/CD pipelines, teams build lasting cost accountability.

Best LLM inference providers 2026: 16+ on cost per outcome

An LLM inference provider hosts open-weight models like Llama, DeepSeek, and Qwen behind a pay-per-token API, handling GPUs, scaling, and serving for you. The same Llama 3.3 70B model ranges from $0.10 to $1.04 per million input tokens depending on who serves it, so provider choice is a pricing decision. Top picks as of September 2026: Groq and Cerebras for speed, DeepInfra for price, Together and Fireworks for breadth, Baseten for custom models.

Build vs. buy: should you build your own AI cost management tooling?

Build when the problem is narrow (one provider, one team, simple attribution) and the tooling is strategically yours to own. Buy when AI spend spans providers, arrives untagged, and needs unit costs finance will trust, because that build is a multi-quarter platform project with a permanent maintenance tail. Price both paths in engineer-years before deciding. CloudZero sells the “buy” side.

Perplexity pricing in 2026: Free vs. Pro vs. Max, and who should pay

Perplexity pricing runs $0 for Free, $20 a month for Pro, and $200 a month for Max, with enterprise seats listed from $40 per user. Pro fits most people who search for work daily. Max exists for heavy automation. The prices are verified against Perplexity's live plans page as of September 2026. When Perplexity published a customer quote on its own pricing page, it chose an unusual one.

Kling AI pricing in 2026: plans, credit costs, API packages, and the spend no invoice shows

Kling AI pricing runs from a free tier of 66 daily credits to a reported $180 per month, with annual billing about 34% cheaper. Kling 3.0 bills per second: 6 to 12 credits for standard resolutions and 30 for native 4K. The API sells separate prepaid packages from $9.80 to $7,560.

ElevenLabs pricing in 2026: plans, credits, and agent costs

ElevenLabs pricing runs from a free plan to $990 per month across six published tiers, with custom Enterprise above that. Plans meter usage in credits, where one credit roughly equals one character of speech. Voice agents cost $0.08 per minute on every tier, plus separate LLM and telephony charges. The AI agency PxlPeak published its own ElevenLabs invoice: $303 for January 2026, covering voice content for six clients, IVR systems for two, and live voice agents for three.

Token-based pricing: how AI usage billing works (2026)

Token-based pricing charges for AI by the volume of text a model processes, metered separately for input tokens (what you send) and output tokens (what the model returns). As of 2026, OpenAI, Anthropic, and Google all bill their APIs this way, and the model is spreading into enterprise chat products. Bills scale with usage rather than seats.

Manage Cursor costs with Datadog Cloud Cost Management

AI coding tools such as Cursor are becoming a significant source of engineering spend. But Cursor costs can be difficult for FinOps teams to manage. Cursor’s usage data alone doesn’t tell you how costs break down across users and models, and fixed-threshold alerts may not catch an unusual cost spike if spend remains below the threshold. Datadog Cloud Cost Management (CCM) brings Cursor costs into the same place where you monitor cloud, SaaS, and other AI spend.

150+ AI statistics for 2026: spend, cost, and AI ROI

Worldwide AI spending will reach $2.59 trillion in 2026, up 47% from 2025, according to Gartner. Yet only 37% of organizations report any earnings impact from AI, McKinsey finds. These AI statistics cover what companies spend, what AI costs to run, and the ROI they're actually getting. That gap between the two headline numbers is the story of AI in 2026.

From a $60K invoice to a $200B earnings call, few can explain the AI bill

CloudZero’s own AI Economics Pulse for September found the 75th percentile of its 430-company customer panel crossed 10% of its cloud bill on AI for the first time in August. Gartner’s latest survey found only 22% of organizations have scaled AI successfully and 11% don’t know what their own function spent on it last year. CJ Gustafson showed what that gap looks like on an actual invoice this week.

Shipped: Know what you actually pay per token on OpenAI

Picture two teams running the same million input tokens through the same model. One team’s tokens are cache hits, queued through the batch API. The other team’s are fresh, sent live. On a current-generation OpenAI model, cached input runs about a tenth the price of a fresh token, and batch processing cuts whatever’s left in half. Stack the two: at a list rate of $2 per million tokens, one team’s bill comes to 10 cents, the other’s to two dollars.

Why engineers ignore cloud costs, and how AI Cost Management Agents fix it

Engineers ignore cloud costs because of broken feedback loops, not apathy. Learn what AI cost management is, why AEO matters more than ever, and how a cost management agent embeds accountability directly into engineering workflows. Engineers ignore cloud costs because cost data arrives too late and too disconnected from their workflow to act on.

Best LLM gateways in 2026: 30+ AI gateways compared on cost control

An LLM gateway is a proxy that sits between your applications and model providers, handling routing, failover, caching, and cost controls through one API. The strongest picks in 2026: LiteLLM for self-hosted control, OpenRouter for instant multi-model access, Portkey for managed governance, and Bifrost for production-scale throughput. Enterprises spent $37 billion on generative AI in 2025, a 3.2x jump in one year, per Menlo Ventures.

How to right-size the handoff between two agents

model-right-sizer-schema is a Claude Code skill that designs the typed contract between one agent and the controller that dispatches it. Point it at an agent plus its controller and it returns a JSON prescription with typed in/out fields, an exclusion list that keeps raw logs out of the reply, a before/after size delta, then writes the contract into the agent's own file. It picks from nine portable output-shape families, or your repo's own.

n8n pricing in 2026: every plan, the execution math, and what AI agents change

n8n pricing runs €24 per month for 2,500 workflow executions (Starter), €60 for 10,000 (Pro), and €800 for 40,000 (Business), with 17 percent off on annual billing and custom Enterprise pricing above that. Every plan includes unlimited users and unlimited workflows. The self-hosted Community Edition is free with unlimited executions; you pay only for your server.

CoreWeave pricing in 2026: every GPU rate and what a node really costs

CoreWeave, a GPU cloud provider, prices start at $6.16 per GPU hour for an Nvidia H100 and reaches $8.60 for a B200, sold as fixed multi-GPU nodes: an 8x H100 node lists at $49.24 per hour on demand. Spot rates run up to 60 percent below on demand, reserved contracts discount up to 60 percent, and egress is free.

Shipped: Find your saved Explorer queries faster

Most people rebuild the same handful of Explorer queries: the monthly close view, spend by team for the staff meeting, the filter set that isolates a service you’ve been watching for two months. When we shipped query history and favorites earlier this year, it gave you a way to save up to 12 Explorer configurations.

Shipped: Turn on the ServiceNow integration yourself in Labs

The ServiceNow integration is in Labs now, which means any CloudZero admin can switch it on and start routing cost work into their incident queue the same day. What that gives you is a full loop between the money and the work. A monthly Optimize pass surfaces recommendations with a dollar figure on each one. Pick the ones worth acting on, open incidents for all of them at once, and each ticket shows up carrying the resource, the finding, and the context an engineer needs.

Your AI Economics Pulse for September 2026

Across a same-store panel of 430 CloudZero customer organizations, AI reached 2.66% of the median company's cloud bill in August 2026, up from 2.61% in July and roughly four times its level a year ago. The 75th percentile crossed 11%. The share of organizations with at least 10% of cloud spend attributed to AI jumped to 28.2% from 23.9%, the largest one-month move that tier has posted. Two-thirds of the panel now spends at least $1,000 a month on AI. The typical bill barely shifted.

GPT-6 Astra pricing: What OpenAI's new flagship costs in 2026

GPT-6 Astra is OpenAI's flagship reasoning model, released September 3, 2026. It costs $10 per million input tokens and $50 per million output tokens on the standard API tier, with cached input at $1 and cache writes at $12.50. That is 2.5 times GPT-5.6 Sol's promotional rate and matches Anthropic's Fable 5.1 on both headline numbers. Batch and Flex halve those rates, Fast mode doubles them, and any prompt past 272K input tokens reprices the entire request.

AI cost calculator: estimate your total spend

An AI cost calculator for the whole wallet adds four lanes: seats and subscriptions, API and token usage, cloud AI services, and GPU infrastructure. Average 2026 totals run $25 per employee per month at light adoption, $100 to $150 at active adoption, and $300 or more at AI-heavy companies. Getting to your number takes four lane subtotals and three corrections.

Shipped: Self-serve your MCP server credentials

Enterprise agent platforms need a client ID and client secret in hand before they will connect to anything. An admin with the Modify MCP Settings permission can now issue that pair directly in Settings, connect the platform, and manage the credential lifecycle on whatever schedule your security policy requires. No support request, no wait.

AI Spend Is a Capacity Problem, Not a Billing Problem

Every organisation running models in production eventually reaches the same point: the AI portion of the cloud bill grows faster than expected, and the immediate response is to invest in visibility. Calls are tagged, spending is attributed, dashboards are created, and the results are shown to the teams responsible.

Our Customer Success AI bill tripled. Here's why we're spending more.

Pop quiz: If you spend $40,000 per month on Anthropic, and you’ve got two customers, what’s your cost per customer? If you bypassed the easy answer of $20,000 and said, “Scott, you old trickster, that’s not enough information to answer that question,” you’ve won today’s prize: a lesson in the perils of average costs. Let’s flesh out the situation: You put an AI feature in your product, a document assistant powered by Claude.

Shipped: Rightsize Kubernetes workloads without leaving your MCP client

Changing a Kubernetes resource request takes two numbers: what the workload requests, and what it uses. The CloudZero MCP server now returns both, by cluster, namespace, or workload. This gives you a number you can defend. Usage comes back as P95 over the date range you query, 30 days by default. When an engineering lead asks whether a service runs on a smaller request, that is the figure that settles it. Over-provisioning and under-provisioning show up on the same query.

LLM token cost: pricing per token explained

LLM token cost is the price a provider charges per token a model reads or writes, quoted in dollars per million tokens. Input and output bill at separate rates, with output priced at roughly 5x input. As of September 2026, published rates range from under $0.10 to more than $180 per million tokens on top-end reasoning tiers. In late 2025, Hardik Sonetta of Thomson Reuters Labs published a warning about the most common prompt caching mistake in production.

Cloud Cost Management for Observability: A Practical Guide

Observability spend is outgrowing infrastructure budgets. What drives the cost up, how pricing models work, and a practical framework to manage it. Sejal Pandey works on content and growth at Last9, writing about observability, reliability, and SRE practices.

Shipped: Find the S3 buckets paying early delete fees

S3 lifecycle rules move data to Standard-IA or Glacier to cut storage cost. CloudZero now flags the buckets where that move backfires: an early delete fee is charged when an object leaves its tier before the tier’s minimum storage duration. The cause isn’t always a misconfigured lifecycle rule. A manual delete, an overwrite, or an object written straight into the tier by a replication or backup job produce the identical charge.

AI usage tracking: Monitor spend by team, feature & model

AI usage tracking means measuring who and what consumes AI across your company, by team, feature, and model, then converting the usage into spend and cost per unit of work. Provider consoles stop at totals per API key. Tracking puts names on those totals: which team, which product, which model, and whether any of it was worth the money. In May 2026, CNBC reported that “almost every Fortune 500 is tracking overall AI usage,” quoting ModelOp CTO Jim Olsen. The same reporting carried his warning.

Repo rightsizing: audit every model call in a repo you already shipped

Repo rightsizing is a single-pass audit of every real model call in a codebase you already shipped: SDK invocations, sub-agent dispatch sites, and agent frontmatter pins. Each call site is scored on the job it actually does, and the result commits as one blueprint file you can diff next quarter. It replaces one-skill-at-a-time reviews, which miss files where a single model key covers two different jobs.