Operations | Monitoring | ITSM | DevOps | Cloud

MTTR Is Not a Time Problem. It Is a Context Problem

Your Mean Time to Resolution (MTTR) has likely stayed flat for three or four quarters. The investment was real: scheduling tools, dispatch optimization, new training modules, and more technicians. Operations reviews still dissect response time, travel time, and wrench time. The metric still refuses to move. Most field service leaders measure MTTR from the start of the repair to the moment the asset returns to service.

AI cost allocation: how to attribute AI spend by team, product, and customer

AI cost allocation is the practice of attributing every dollar of AI spend to the team, product, feature, or customer that generated it. That spend includes API tokens, GPU compute, per-seat tools, and shared infrastructure. It's harder than cloud allocation because AI spend arrives untagged, spans vendors, and pools in shared resources. Four methods cover most cases: tag-based, key-based attribution, proportional split, and usage-telemetry.

Shipped: One CloudZero for everyone, starting October 1

On June 3, we made the new CloudZero experience the default for every customer. Since then, we’ve shipped around 30 improvements a week: side-by-side period comparisons in Explorer, budgets you can create and edit right in the app, threshold alerts on dashboard tiles, and Monitors, which flags AI and cloud spend that moves outside its normal pattern and shows you what changed. Pages load 28 to 61% faster. JavaScript execution is 85% faster.

Auto-Generate Richer Azure Architecture Diagrams

Azure architecture diagrams go out of date fast, and management-plane data alone misses the runtime connections that matter most. In v5.4, diagrams move to their own Diagrams tab in Azure Documenter. Alongside the enhanced network and workload diagrams, there is a new Resource Visualizer diagram. Scope it by subscription and resource group, or write your own custom Azure Resource Graph query to define exactly which resources to include.

Change the Cloud Cost Conversation from Spend to Margin

"Our Azure bill went up 20% last month." Without context, finance only sees a rising cost. Unit economics gives them the full picture. Turbo360 lets you overlay business KPIs on your Azure spend. Track units like orders, active users, document views, or monthly recurring revenue alongside cost, and see your cost per unit month over month. Now the conversation becomes: "Orders went up 150% and our cost per order came down." That is a story about efficiency, not overspend.

IT Service Continuity Management: How to Build an ITSCM Plan

Most IT teams have a recovery plan somewhere. It was written for a disruption that has not happened yet, and tested less often than anyone admits. The gap rarely sits in the technology. Nobody agreed which services come back first, or how fast. There was time to settle that calmly, and it went unused. IT service continuity management is the ITIL practice that settles those questions in advance. In this blog, you will: By the end you will know what belongs in an ITSCM plan and who has to agree to it.

How IT Infrastructure Management Keeps Services Reliable and Costs Predictable

When a business application slows down, how fast can your organization trace the cause to a server, a network link, storage or a cloud instance? Often it comes down to who's on call that day, since asset, observability and change data are scattered across separate systems. Engineers then check each tool one at a time while customers wait and the cost of the outage grows. IT infrastructure management solves this by keeping asset records, health data and change history in order before an incident starts.

Database change management on Databricks: migrations, environments, and AI-generated change

You wouldn’t ship untested SQL Server database changes – Why is Databricks different? Databricks is where the data estate is growing, and increasingly where AI workloads run and generate change. But schema change there still happens the hard way: views and stored procedures managed through manually versioned scripts, drift between workspaces discovered when a deployment fails, and no reliable record of what changed, when, where, or why. As AI raises the volume and speed of schema change, these gaps widen.

This AI agent finds your app's bottlenecks and suggests the fix

Most teams collect the profiles and traffic data that explain a slowdown. Almost nobody has time to read it before users notice. In this Product Highlights conversation, Sylvain Guittard, Senior Director of Product at Upsun who leads the team behind the Upsun console and CLI, breaks down the Upsun Cloud Performance Agent, the first background agent running on Upsun Cloud. His take: "We monitor everything, we feed that into an agent, and the agent will be capable of finding what the bottlenecks are in your application. And on top of it, it gives you a patch, or a way to fix it." We get into.