Operations | Monitoring | ITSM | DevOps | Cloud

The latest News and Information on Cloud monitoring, security and related technologies.

Shipped: Views now work for every access level

Most people who open CloudZero care about one slice of the spend, like their team, their product, or their region. A View gives them that slice in one click, with the grouping and filters already set, so nobody has to rebuild the same Explorer query every week. Views now work for everyone in your organization, including people with scoped access.

How to Guarantee a Website or Service Never Goes Down (And What You Can Actually Promise)

No one can guarantee that a website or service never goes down. What you can promise is a measured availability target, and with a multi-location, active-active design you can reach 99.999% (five nines), about 5 minutes 15 seconds of downtime a year. That takes redundancy at every layer, automatic health-based failover across regions and ideally providers, safe deployments, failure testing, and outside-in monitoring. Control Plane is built for that tier.

Azure in Bleemeo: your subscription next to your servers, with one read-only role

Most teams that run on Azure do not run only on Azure. There is a database on a VM nobody wants to move, a Kubernetes cluster somewhere else, a few servers in a rack, and a monitoring setup that grew around all of it. Azure Monitor sees the Azure part very well and nothing else, so the picture of an incident ends up split across two consoles, two alerting configurations and two sets of dashboards. Bleemeo now connects to Azure the same way it already connects to AWS.

18: Building an Agentic Future: AI and Optimization with Sachin Gharge

On today's episode, Andrew Hillier chats with Sachin Gharge, Head of Cloud Platform at Scandinavian Airlines (SAS). They discuss AI, agents, Kubernetes, MCP, and optimization. Sachin shares how he and his team are optimizing cloud costs, leveraging automation, and experimenting with agentic AI, including bots and Slack integrations, to make operations easier and more effective for developers and the business.

Run your first workflow in minutes, no sales call

The regression nobody catches passes a busy review and ships. An off-by-one, a change that reads as sensible and quietly breaks something, gets a nod from a tired reviewer and lands in production, where it erodes trust one small defect at a time. You can have an AI code reviewer running on your own repository in the time it takes to read this page. Get started without having to book a demo or contact sales.

Shipped: See what your AI spend is actually paying for

Most AI spend comes in with no tags and no owner attached. Your provider console shows total spend, maybe broken out by API key or model. It won’t tell you that the sales team spent $1,700 on Claude this week, let alone what the work was. And the problem is growing. McKinsey found that 56% of organizations now use AI in three or more business functions. More teams means more spend, and most companies respond with a spending cap. Set it too low and you slow down the work you wanted AI to help with.

Moving Mainframe Data to Snowflake and AWS Through Apache Kafka: Treehouse Software and meshIQ

Treehouse Dataflow Toolkit moves mainframe data from Db2, VSAM, and IMS through Apache Kafka pipelines into Snowflake and AWS targets. meshIQ keeps the streaming layer visible and stable. Together they give data science teams continuously updated enterprise data for AI and ML.

Get your agents off laptops and onto shared infrastructure

There's a specific, recognizable point where a team's use of AI agents changes shape. Not when they adopt agents; most teams already have. It's when agents stop running on someone's laptop and start running on infrastructure that the whole team can see. This is a real technical shift, not a policy change or a maturity score. Here's specifically what's different on each side of it.