Operations | Monitoring | ITSM | DevOps | Cloud

How task containers give AI agents real infrastructure without idle cost

Infrastructure for AI agents usually forces a choice between two bad options. A sandbox is safe but blind, cut off from the data and services that would make the agent's output useful. Full access means paying to keep a container idle between runs, waiting on a prompt that might not arrive for hours. Task containers, which Upsun released on August 12, 2026, are built to avoid that choice. A task container is a single-purpose container defined in a project's.upsun/config.yaml file.

Shipped: Explorer refresh: show more, scroll less

CloudZero Explorer answers a cost question in two parts. The chart shows what your spend did and the table underneath shows which service, account, or team did it. Until now, the chart pushed the table below the fold, and actions like creating a View or checking Anomalies were buried multiple clicks deep. Now the chart and table share the screen and a new right rail puts Favorites, Views, Anomalies and Insights one click away without covering your data.

Turn every branch into a production-like environment, automatically

You push a branch. If your team is like most, that branch now waits: for the shared staging server to free up, for someone to remember to refresh the seed data, for whoever broke staging last to fix it. By the time you actually test your change, you're testing it in an environment that's drifted from production in ways nobody fully tracked. The alternative isn't a better staging server. It doesn't need one.

Private cloud vs. Public cloud: Which delivers greater control and flexibility?

As businesses evolve in today’s digital landscape, the need for efficient and scalable computing resources has become paramount. In the early days of the Internet, large corporations would build or rent out large data centers to run their applications and serve customers. This was great as they could use dedicated hardware and expand as they pleased.

Trust you can verify: security assurance for the AI era

When you choose a cloud platform, you're entrusting a provider with sensitive business information, customer data, critical applications, and a growing share of your operational resilience. Increasingly, you are also entrusting it with AI. And that changes the questions you should be asking. Marketing claims cannot answer these questions. Independent evidence can. Here is what that evidence looks like at Upsun and why it matters to your next supplier review.

The infrastructure work you should not have to touch just to ship a feature

You wrote the feature. It works locally. Then you spend the next two hours on things that have nothing to do with the feature: a Terraform plan that wants to replace a database you didn't touch, a Kubernetes manifest that needs a new ingress rule, an IAM policy that's one permission short of what the deploy needs. None of this is the job. All of it is the job today. Here's what that list actually looks like, and why none of it should be sitting on your plate.

You Vibe Coded an App...Now What?

"Hey, I built this over the weekend. I want to get it in front of customers." And it always hits architecture, security, and infrastructure. Ross Hendrickson, CTO at Inspectiv, calls that gap the chasm. His team crosses it on Control Plane: AI-written code secured, reviewed, and released in a day. Control Plane combines AWS, GCP, Azure and your own hardware into one virtual cloud shaped to your workloads.

Token budgets: capping AI agent and LLM spend

AI costs are changing. As noted by research from EY, outputs that cost just $0.04 in 2023 now cost $1.20, a 30x increase over just three years. It’s worth noting that task operations and complexity have also changed. In 2023, the process was simple. Users input a question, retrieval engines found relevant data, and AI models returned a response. Today, many tasks are handled by orchestrated AI agents capable of much more complex reasoning and analysis.

Shipped: A changelog that keeps up with how fast we ship

When the changelog doesn’t keep pace with the product, two things can happen. One, you keep working around something that was already fixed weeks ago. Or two, a behavior changes, you assume it’s a bug, and you spend an afternoon on triage and a support ticket before learning it was an intentional improvement. CloudZero now ships around 30 improvements a week, a pace driven by the Next Gen Platform and the AI-first approach we’re building for our customers.

The Waiting Game for Data Centre Capacity (And How UK Businesses Can Beat It)

UK data centre occupancy hit 91% in 2024, according to Arizton market data, and new capacity is not arriving fast enough to close the gap. Grid connection wait times for new projects now run between five and 15 years, reports Data Center Dynamics, and Savills has attributed the 11% year-on-year drop in new capacity delivery to power constraints rather than a lack of demand or investment. Rising wholesale energy costs are addingpressure to an already tight market.

What build-versus-buy actually looks like in agentic engineering

Most build-versus-buy debates assume you're choosing once, at the start, and living with it. Agentic engineering doesn't work that way. The decision shows up at every layer of the stack, and the teams getting it right aren't the ones who picked "build" or "buy" as a philosophy. They're the ones who know which layer is which.

Trace AWS Lambda durable functions with Datadog

AWS Lambda durable functions let you build long-running, multi-step workflows for use cases such as payment processing, order fulfillment, and AI workflows with human approval. A single durable execution can pause for a wait or callback, retry failed work, and resume in a fresh Lambda invocation without losing its state. The strong resilience provided by durable executions, however, creates an observability challenge because each invocation produces its own telemetry data.

Shipped: Personalized cost access, powered by SSO

Instead of building a separate role for every team, region, or department, admins can create a single role that automatically personalizes access for each user based on their SSO attributes. Someone moves teams or a new group gets created, and the new access takes effect at their next login with no CloudZero configuration. As AI spend grows, more companies are looking to give teams visibility into their own AI costs without exposing every individual’s usage across the org.

HOA Tech Trends: Managing Communities Faster

Community association leaders face growing administrative demands as modern neighborhood operations become more intricate. Modern technology helps volunteers and board members manage daily tasks with far greater speed and precision. Adopting tailored software reduces manual paperwork and improves operational clarity across the entire neighborhood. Leaders can allocate time toward long range planning instead of spending late evenings chasing routine documents and payment receipts.

Best Azure monitoring tools: Compare the leading solutions

Microsoft Azure has become one of the most widely adopted cloud platforms for running business applications, databases, containers, analytics workloads, and enterprise services. Modern Azure environments now extend far beyond virtual machines, encompassing services such as Azure Kubernetes Service (AKS), Azure SQL Database, Azure Functions, storage accounts, networking services, and serverless applications.

Migration feasibility checklist for IT leaders

Feasibility is a prioritization question that comes before strategy. The five-question check produces a "now, later, or fix blockers first", before anyone touches a target architecture. Strategy earns its place once feasibility returns "now." Feasibility comes before strategy. Before anyone designs a target architecture, builds a runbook, or commits to a multicloud operating model, the question is whether the migration is the right move now, and what would make it fail.

Cloud cost management: how repatriation improves control for UK enterprises

Hyperscale providers are nothing if not consistent in their temptation of enterprise IT buyers. They bombard leaders with a simple message: migrate to the public cloud, shut down data centres, and enjoy both financial savings and operational agility. However, as UK enterprises have scaled their digital footprints, a more nuanced reality has bitten. Public cloud costs have swollen.

The finance dashboard I actually use, built from CloudZero and Campfire in an afternoon

Every finance person I know lives in the same loop approaching the end of the month, quarter, or fiscal year. Leadership wants to know where the financials will land (most times before the close has occurred). CS wants customer margins. Someone on the People team needs each department’s AI spend for an OKR review, and they need it quickly to make business decisions. Each answer sits in a different tool or a different spreadsheet, and I bounce across all of them several times a day.

Shipped: In-app help, right beside your work

You are mid-investigation, chasing a spike or pulling a number for finance, and you hit a term or a workflow you need to look up. You should not have to lose your place to find an answer. Guide lives in a fixed spot in the left sidebar, always one click away. It opens a panel on the right side that sits beside your page instead of covering it. Your chart, filters, and time range stay exactly where they were. Nothing gets rebuilt and you keep the thread of what you were investigating.

How to ensure compliance with private cloud providers in regulated sectors

The compliance question isn't "are we using a private cloud?" Rather, it’s "does our private cloud actually do what compliance requires?" Private cloud has a reputation for solving compliance problems that it doesn't always deserve. The logic seems straightforward: keep data off shared public infrastructure, maintain more direct control, and satisfy the auditors.

Run an AI SRE Agent Entirely Inside AWS with Bedrock and S3: AURA

An on-call question returns the threshold and the escalation owner from your own runbooks, and the answer comes back without a call to anyone outside. AURA runs against Bedrock as its model provider, using Claude Sonnet 5 served by AWS in the same region. Authentication is the normal AWS credential chain: a profile on a laptop, an IAM role in EKS.

Your FY27 plan deserves a real AI number, not a hedge

Budget season is starting and most finance teams are finding the AI line is the most evasive line on the page. You lived through the year. AI spend came in higher than planned and moved in ways nobody could foresee or forecast. And when the board asked what it produced, the honest answer probably was “we’re working on it.”

Shipped: Codex spend tied to the work behind it

People run Codex on their own laptops. When Codex is signed in with a ChatGPT subscription, OpenAI’s own admin console shows who used it and how much: messages and credits. What it doesn’t show is what any of that usage was for, or how it compares to what your team spent on other AI tools. The CloudZero desktop agent for macOS installs on a Mac, sees the traffic from AI coding tools, and prices what those tools use.

Why is AI so expensive? The real cost drivers of AI

AI is expensive because the model bill is only part of the cost. Three components set the floor: model subscriptions, per-token API pricing, and infrastructure. Three more make it move: adapting models to your business, catching and fixing errors, and rising energy and datacenter costs. Efficiency doesn't fix it, because cheaper AI gets used more, not less. Businesses are willing to spend on AI. Research from Deloitte found that in 2025, 85% of organizations increased their AI investments.

Shared context for AI coding agents beats better tooling

The instinct when adopting AI coding agents is to optimize the agent. Compare models, tune prompts, argue about which editor has the better completion, and treat the agent as the thing that determines how fast the team moves. Then the commits go up and the product does not. The team building Upsun Dispatch took a different route, and the result is worth copying. They did not find a better agent.

Shipped: Monthly cost comparison in Explorer gets a glow up

Months have different numbers of days, and a monthly cost chart built on raw totals mixes that calendar difference into the trend. A 28-day February next to a 31-day March shows a 10.7% increase even when daily spend never moved. The same math works in reverse: real growth in a short month can look flat, hiding an increase worth investigating. That costs you time in two places. The first is triage.

AI budgeting: how to plan and forecast AI spend

AI budgeting is the process of planning, allocating, and forecasting an organization's AI spend: model and API costs, AI infrastructure, tooling, and the people running it all. It differs from traditional budgeting because AI spend is usage-based, scales with product success rather than headcount, and often spans multiple providers.

Building an End-to-End Drone Ecosystem: The Technologies That Need to Work Together

Commercial drone technology is rarely a single application running alongside an aircraft. A complete solution may include flight software, onboard sensors, telemetry, cloud infrastructure, web and mobile interfaces, data processing pipelines, analytics tools, and integrations with existing business systems.

The Architecture Question That Never Dies: From BPMN and M&A to MCP

Twenty years ago at RMIT, I became preoccupied with a question that sounded technical but was really about corporate value: could you predict how difficult a company would be to acquire by looking at the shape of its APIs? It was 2006. I was completing Honours in a Bachelor of Applied Science in Software Engineering, and the brief for my research project was unusually open: find an impactful software research hypothesis that hasn’t been done before.

Cloud Incident Management: Process, Tools, and Practices

How do you resolve an outage your organization has no authority to fix? A managed database drops into read-only mode and stops accepting writes. There's no host to reach, no configuration file to edit, and no restart command available to your engineers. Cloud incident management begins at that boundary, where the response depends on a support channel and a provider status page. Plenty of what you already know still applies here.

Shipped: Cost anomalies and savings recommendations, delivered into ServiceNow

If your engineering teams run on ServiceNow, incidents are where they get work done. Putting cost work into an incident gives it the same path to resolution as any other work item your team handles. When a cost anomaly arrives as an incident, your teams route it, assign it, and resolve it on their usual SLAs. When a savings recommendation arrives as an incident, an engineer owns it and acts on it. Now you can send either straight into ServiceNow.

How to build the business case for AI

A strong AI business case ties a specific goal to a measured outcome and a fully-loaded cost. Most fail because they skip one of the three: no clear mandate, an over-broad "AI fixes everything" scope, or a cost estimate that ignores adaptation and error-correction. Build it in six steps: define goals, identify uses, break work into tasks, evaluate models, assess total cost, then launch and refine. Most companies are now spending on AI. Far fewer can show what they got back.

How to right-size your existing Claude skills

You shipped a skill. It worked. You closed the tab. That’s the whole problem. Model choice is a decision you make once, at the moment you’re least equipped to make it: before the skill is even authored. Then you never revisit it, because the skill stopped being interesting the day you got it working. So go back and check. Here’s how.

JFrog Artifactory Now Integrates Natively with Artifact Registry in Google Cloud

Teams running containerized workloads on Google Cloud have long relied on JFrog as their single source of truth for container images. The missing piece has been getting Google Cloud’s own runtime services — like Cloud Run and Google Kubernetes Engine (GKE) — to pull directly from JFrog for every container image pull. I’m happy to say that the gap is now closed. Artifact Registry in Google Cloud has introduced a new repository mode called Connector that addresses this requirement.

Peak Cloud: Decentralising for resilience

For more than a decade, the prevailing wisdom in enterprise IT was simple: move everything to the public cloud. Hyperscale platforms promised unlimited scalability, lower costs, agility and freedom from the burdens of managing infrastructure. Cloud-first has been rapidly gaining momentum as the de facto path to a modern digital footprint. Until now.

Private cloud disaster recovery: How to design for business continuity without public cloud dependency

Disaster recovery (DR) is one area where organizations often assume public cloud has the answer already. Multi-region deployments, managed backup services, automated failover - the hyperscaler catalog is full of DR-flavored offerings, and the marketing suggests that resilience is a solved problem once you're on cloud infrastructure. For many workloads, this is roughly true.

Shipped: Cut the notification noise so real cost anomalies stand out

A view is scoped to the costs your team cares about, and now its notifications are too. Weekly and monthly trend summaries, and global anomaly alerts, only reach a channel when your team wants them there. That keeps a shared channel signal, not static, so the alerts that need action don’t get lost next to irrelevant updates. Your team decides, per view, which notifications reach its channel.

LLM cost management: a practical guide for teams that own the budget

LLM cost management is the practice of tracking, allocating, budgeting, and governing large language model spend so every dollar maps to a feature, team, and business outcome. It has five levels: provider visibility, business allocation, unit economics, model governance, and a continuous optimization loop. It matters because 68% of companies say AI initiatives ran over budget last year, and per CloudZero's 2026 survey, 30% of finance leaders still reconcile AI spend manually.

Pentagon-shaped org charts are coming. Intellectually curious leaders will get a head start.

If you spend even fifteen minutes reading about AI’s impact on the future of work, you’ll take in a lot of fear-based analysis. The fears are real — 40% of workers fear losing their jobs (Metaintro), 60% believe AI will eliminate more jobs than it creates (Yardi Kube), and 52% generally worry about the impact of AI in the workplace (Pew Research) — but the analysis is all wrong.

Data localization for Indian Fintech: RBI rules and your cloud choice

Indian fintech operates under one of the most specific data localization regimes in the world. The Reserve Bank of India has published progressive guidance since 2018 requiring payment system data to be stored in India, with subsequent extensions to other categories of financial data. The rules aren't optional. For fintechs operating in India - whether payment providers, lending platforms, wealth managers, or neo-banks - the localization requirements shape fundamental infrastructure choices.

Don't build the autonomous AI factory first

Here's a scene playing out in engineering teams right now. An engineer spends the weekend running four or five coding agents in parallel. Monday morning, a teammate opens their laptop to 53 changed files with 2000+ diffs and a message that says, more or less, "should be good to merge." Nobody asked for this much output. Nobody has time to review it properly. The team doesn't feel faster. It feels ambushed.

Shipped: Stop guessing why that billing connection exists

Every team with more than a few data connections has had this moment: someone opens the connections list, points at one, and asks “what is this for?” The answer lives in a former teammate’s head or in a Slack thread. And cleaning up the wrong connection can break cost ingestion. Now each connection can carry a note that explains why it exists, and anyone who opens the connection sees it.

AI agent cost: what agents really cost to run

AI agent cost in 2026 is mostly a consumption bill, not a subscription. Running an agent costs anywhere from fractions of a cent for a simple routed task to $5 or more for a complex multi-step job, because one request can trigger 3 to 10 model calls behind the scenes. Average production deployments land between $3,200 and $13,000 per month in operational spend. Here is where that money actually goes.

Upsun recognized for third consecutive year in the Gartner Magic Quadrant for Cloud-Native Application Platforms

Upsun acknowledged for its Ability to Execute and Completeness of Vision. Upsun is proud to be recognized for a third year in the 2026 Gartner Magic Quadrant for Cloud-Native Application Platforms alongside other evaluated CNAP platforms. Upsun empowers development teams to ship better software, faster, not just by simplifying infrastructure management, but by rethinking how the entire software development lifecycle works in an era of AI-powered development.

Cloud Outage Resilience: On-Call Lessons for 2026

Cloud outage resilience has quietly become the most important reliability topic of the year. Analysts now treat large scale cloud downtime as a matter of when, not if. Forrester has predicted at least two major multi day hyperscaler outages in 2026, and the reasoning is hard to argue with. AWS, Azure, and Google Cloud together account for well over half of enterprise cloud spending, so when any one of them stumbles, a huge slice of the digital economy stumbles with it.

SaaS Tools: An Essential Guide for Navigating Insurance Compliance for Franchise Businesses

For franchise businesses, managing insurance compliance can be a complex and time-consuming task. The intricacies of ensuring every franchisee adheres to corporate policies, while also meeting local regulatory requirements, can overwhelm even the most organized teams. Moreover, the need to regularly update and verify insurance certificates adds another layer of complexity. This article will delve into the key challenges faced by franchise businesses in maintaining insurance compliance and explore how SaaS tools specifically designed for this purpose can streamline these processes.

Shipped: Get anywhere in CloudZero with a keystroke

You know exactly where you want to go in CloudZero. Getting there sometimes takes a moment as you click into the nav, open a menu, scroll a dropdown, find the thing, click again. Every trip back to a familiar spot can take a few steps. Shortcuts remove that friction. Press command+K on Mac or ctrl-K on Windows anywhere in CloudZero, type where you want to go, and hit Enter. That means there’s no clicking through the nav and no scrolling to find what you already know the name of.

What is AI ROI? Definition and why it matters

In 2025, 85% of organizations increased AI investment, and 91% plan to do the same this year, according to Deloitte. Despite continued spending, however, ROI lags behind, with just 6% seeing payback within one year. While AI use cases tend to have a longer payback period, often in the 2-4 year range, companies can’t afford to keep spending money without some measure of its practical impact both immediately and over time.

What are AI tokens? The unit your AI bill is written in

AI tokens are the small chunks of text, roughly four characters or three quarters of a word each, that language models read and generate. Every prompt and every response is measured in tokens, and AI providers bill per million of them. That makes the token the base unit of AI spend: 1,000 tokens is about 750 words, and every AI feature you ship is a token meter running.

Ai4 2026: Measuring AI spend is solved. Now it's time to prove its worth.

CloudZero had a full team on the ground at Ai4 in Las Vegas during the first week of August 2026. The team included CTO Erik Peterson, who spoke on a panel about AI cost economics. The same problem surfaced everywhere we went: teams can see what they’re spending, but not whether it’s working. DIY cost tooling that fails time and time again, agent sprawl, and a widening gap between finance and engineering kept coming up throughout the week.

Inference Optimization Techniques. Ray vs. vLLM vs. KubeRay

Serving large language models at scale is fundamentally a distributed systems problem. A single GPU, or even a single node, is rarely enough once you need multiple models, multiple replicas, tensor-parallel sharding across GPUs, or high-availability rollouts. Kubernetes solves general container orchestration well, but it has no native concept of a GPU-aware, actor-based compute cluster.

How to Reduce Photo File Size on iPhone: Easy-to-Follow Guide

Our iPhones now have the capability to take amazing, high-quality photos and videos, but they can quickly eat up your storage if you don’t know how to reduce photo size on iPhone, or how to compress a photo on iPhone. If your cloud storage runs out, you can use the methods in this article to free up the space on your phone or back up your photos on another private alternative to iCloud Photos, like Internxt Drive and Photos.

Shipped: Catch the S3 object-tag charge before it scales with you

There’s an S3 charge that stays invisible in a normal storage cost review. AWS bills S3 object tags per tag, per hour, so the cost scales with how many objects you have, not how much data you store. It gets its own line item, which is easy to miss when you’re scanning storage spend. It can sneak up on you. Tags get added in a dev environment to drive lifecycle rules, where object counts are small and the cost is nothing.

Generative AI ROI: benchmarks and how to prove it

Generative AI ROI measures the financial return on generative AI investments relative to their total cost. Benchmarks diverge sharply: Google Cloud's 2025 study found 74% of enterprises see ROI within the first year, while MIT's NANDA initiative found 95% of pilots deliver no measurable P&L impact. The difference is not the AI. It is whether the organization can actually measure cost and outcome at the use case level.

AI isn't a black box. It's Pandora's Box.

When CFOs talk about AI budgets, they tend to describe it the same way: it’s a black box, offering little or no transparency. The bill arrives at the end of the month, it’s bigger than last month, and nobody can really explain why. Meanwhile, engineering keeps asking to raise the token budget. I think that framing undersells what’s actually happening out there. If the black box is the bill, the Pandora’s box is what you opened when you brought AI into the company.

Kubernetes GPU Scheduling for MLOps and GPU Sharing

The default Kubernetes scheduler was built for stateless services: web servers, APIs, databases. It schedules a pod, checks that a node has enough of whatever resources were requested, and binds it. For CPU and memory, that model works fine. For GPUs, it falls apart in three specific ways. First, GPUs are treated as an opaque integer resource.

Monitor your Amazon Bedrock workloads with Applications Manager

Organizations are increasingly integrating GenAI capabilities into their applications to deliver richer, more contextual user experiences—from AI-powered customer support and enterprise search to content generation, virtual assistants, and automated workflows. To build and scale these GenAI-powered experiences, they are turning to platforms such as Amazon Bedrock, which provides access to foundation models that developers can integrate into their applications.

Shipped: Catch a cost spike before it hits your bill

You’re probably already tracking the metrics that matter most in your Analytics dashboards like unit economics, AI ROI, and spend by team. Now you can put a target on any of them. Pick the metric, set the threshold, and CloudZero emails you when it’s crossed, with no ticket to us, no custom build.

How to measure AI ROI: metrics and a framework finance can actually run

To measure AI ROI, compare attributable value (revenue lift, cost savings, engineering time recovered, risk reduction) against fully loaded AI spend (API usage, subscriptions, infrastructure, people time) at the unit level: per initiative, per team, per task. The formula is simple. The instrumentation is the hard part, and it's where most organizations are failing: in CloudZero's 2026 survey, 34% of finance leaders couldn't produce a credible ROI number at all.
Sponsored Post

Building a Modern Cloud Outage Response Workflow in Slack and Microsoft Teams

On May 7 and 8, 2026, a thermal event in a single AWS data center hall knocked out power to EC2 instances and EBS volumes in a single Availability Zone in us-east-1. Within hours, more than 150 cloud services went down, including Coinbase, Reddit, HubSpot, and Atlassian's suite of tools, Jira, Confluence, and Trello among them. For teams without a structured cloud outage response workflow, the next several hours looked familiar: Slack DMs asking "is it down for you too?", tab-switching between status pages, and incident commanders repeating the same update in three different channels.

Why Config Changes Cause Most Cloud Outages in 2026

If you have watched the incident channels light up over the past few weeks, you already sense the theme of 2026: cloud outages are no longer rare, dramatic once a year events. They are a steady drumbeat, and most of them trace back to the same root cause. Not a data center fire, not a rogue backhoe severing a fiber line, but a routine configuration change that went out, behaved differently than expected, and cascaded.

Why Cloud Cost Visibility at Scale Fails (And How to Fix It) | Harness Blog

Cloud cost visibility at scale usually works great… until it suddenly doesn’t. At first, everything feels manageable. You can track spend by service. You know which team owns which resources. Reports are clean, and the numbers make sense. Then one day, there’s a $47,000 spike spread across three AWS accounts that no one noticed for eleven days. Leadership wants answers. Engineering wants context. And your carefully designed tagging strategy?

Azure Virtual Desktop Monitoring: Challenges, Metrics & Best Monitoring Solutions

Azure Virtual Desktop (AVD) is rapidly growing in popularity as modern way to deliver virtual desktops and apps to users, with Azure providing the infrastructure as alternative to on-prem VDI environments. As organizations expand their use of AVD in Azure, monitoring becomes critical.

Agent security starts with where the agent runs, not how it behaves

When engineering teams evaluate AI agents, the first questions are usually about capability. Which model performs best? How much faster can it write code? What's the return on investment? Security, if it enters the conversation at all, tends to come later. Patrick Dawkins, Principal Software Engineer at Upsun, thinks that's backward. Over the past year, he's been building the infrastructure that enables AI agents to operate safely within engineering teams.

Railway Mania, the birth of the S&P 500, and the lesson for the AI era

In 1846, Britain poured roughly 7% of its national income into railways, proportionally about three times what the U.S. spends on AI infrastructure today. The technology delivered everything it promised, and a generation of investors still lost their shirts. What sorted the winners from the wreckage wasn't conviction about the technology; it was whether ROI was measured or asserted. The man who fixed that problem gave his name to the S&P 500.

Migration playbook: escaping lock-in without disruption

Migration projects fail in a predictable sequence. The technical work gets scoped. The timeline gets set. The engineering team starts moving workloads. Somewhere in the middle, dependencies surface that weren't in the original assessment, the double-run period extends beyond the budget allocated for it, and the project either stalls or completes at significantly higher cost than planned.

AI cost reduction: tactics that preserve performance

AI cost reduction means lowering what you spend to run AI (tokens, inference, and compute) without sacrificing quality. The highest-leverage tactics, prompt caching, batching, and routing easy work to smaller models, cut spend 50 to 90% by removing waste, not capability. Somewhere right now, a finance leader is opening an AI bill that has quietly tripled, with no new product to show for it. Nobody approved it. No single decision caused it.

Shipped: Put every AI task on the cheapest model that can actually do it

If your team builds with AI, someone is defaulting to the biggest model available (say, Fable) because it feels like the safe pick, and the safe pick is almost always the most expensive one. One over-powered choice looks harmless on its own, but multiplied across every prompt, agent, and workflow, and you get a big number on the P&L. All that, yet nobody chose which model on purpose. As we like to say, using a default is not a decision.

Cloud Outage Response: Lessons From July 2026

Cloud outage response got a brutal stress test in July 2026. In the span of nine days, three separate cloud infrastructure failures took large chunks of the internet offline: AWS CloudFront on July 16, Microsoft Azure West US on July 23, and AWS us-west-2 on July 24. None of them were caused by a dramatic data center fire or a nation state attack. They were routing faults, configuration translation bugs, and a piece of networking hardware on the path between a region and a metro area.