Operations | Monitoring | ITSM | DevOps | Cloud

The latest News and Information on Cloud monitoring, security and related technologies.

Run your first workflow in minutes, no sales call

The regression nobody catches passes a busy review and ships. An off-by-one, a change that reads as sensible and quietly breaks something, gets a nod from a tired reviewer and lands in production, where it erodes trust one small defect at a time. You can have an AI code reviewer running on your own repository in the time it takes to read this page. Get started without having to book a demo or contact sales.

Shipped: See what your AI spend is actually paying for

Most AI spend comes in with no tags and no owner attached. Your provider console shows total spend, maybe broken out by API key or model. It won’t tell you that the sales team spent $1,700 on Claude this week, let alone what the work was. And the problem is growing. McKinsey found that 56% of organizations now use AI in three or more business functions. More teams means more spend, and most companies respond with a spending cap. Set it too low and you slow down the work you wanted AI to help with.

How to Guarantee a Website or Service Never Goes Down (And What You Can Actually Promise)

No one can guarantee that a website or service never goes down. What you can promise is a measured availability target, and with a multi-location, active-active design you can reach 99.999% (five nines), about 5 minutes 15 seconds of downtime a year. That takes redundancy at every layer, automatic health-based failover across regions and ideally providers, safe deployments, failure testing, and outside-in monitoring. Control Plane is built for that tier.

Azure in Bleemeo: your subscription next to your servers, with one read-only role

Most teams that run on Azure do not run only on Azure. There is a database on a VM nobody wants to move, a Kubernetes cluster somewhere else, a few servers in a rack, and a monitoring setup that grew around all of it. Azure Monitor sees the Azure part very well and nothing else, so the picture of an incident ends up split across two consoles, two alerting configurations and two sets of dashboards. Bleemeo now connects to Azure the same way it already connects to AWS.

How Honeycomb Private Cloud Drinks From the Fire Hose

Back in July, I wrote about how the Tenant team (the team behind Honeycomb Private Cloud (HPC)) has embraced the code review bottleneck to focus more of its work. One of the other challenges we have is that we're downstream of almost all the other teams at Honeycomb, meaning that we have to package up everyone's code and services, and how it gets provisioned! This is something impossible to handle through code review since there are so many engineers on other teams, and so few of us.

Shipped: Start every session where your work lives

Most people who use CloudZero spend their time in one or two places. For some it’s AI Signals, and for others it’s Optimize or Anomalies. If Explorer isn’t one of those places, every sign-in starts with a click to get where you need to be. Dates and numbers are another friction. A date like 04/07 means April 7 in the US and July 4 in much of Europe. When the platform shows a format your team doesn’t use, you end up having to convert each value before you can work with it.

Moving Mainframe Data to Snowflake and AWS Through Apache Kafka: Treehouse Software and meshIQ

Treehouse Dataflow Toolkit moves mainframe data from Db2, VSAM, and IMS through Apache Kafka pipelines into Snowflake and AWS targets. meshIQ keeps the streaming layer visible and stable. Together they give data science teams continuously updated enterprise data for AI and ML.