Operations | Monitoring | ITSM | DevOps | Cloud

Shipped: One CloudZero for everyone, starting October 1

On June 3, we made the new CloudZero experience the default for every customer. Since then, we’ve shipped around 30 improvements a week: side-by-side period comparisons in Explorer, budgets you can create and edit right in the app, threshold alerts on dashboard tiles, and Monitors, which flags AI and cloud spend that moves outside its normal pattern and shows you what changed. Pages load 28 to 61% faster. JavaScript execution is 85% faster.

IT Service Continuity Management: How to Build an ITSCM Plan

Most IT teams have a recovery plan somewhere. It was written for a disruption that has not happened yet, and tested less often than anyone admits. The gap rarely sits in the technology. Nobody agreed which services come back first, or how fast. There was time to settle that calmly, and it went unused. IT service continuity management is the ITIL practice that settles those questions in advance. In this blog, you will: By the end you will know what belongs in an ITSCM plan and who has to agree to it.

How IT Infrastructure Management Keeps Services Reliable and Costs Predictable

When a business application slows down, how fast can your organization trace the cause to a server, a network link, storage or a cloud instance? Often it comes down to who's on call that day, since asset, observability and change data are scattered across separate systems. Engineers then check each tool one at a time while customers wait and the cost of the outage grows. IT infrastructure management solves this by keeping asset records, health data and change history in order before an incident starts.

Realtime transaction fraud detection - with an LLM?

Conversational AI with a chatbot is great for drafting emails or debugging code, but it’s less ideal for real-time application middleware. If you’re trying to inspect a financial transaction for potential fraud in the middle of a checkout loop, you don’t need an LLM to write you an essay about why a credit card transaction looks suspicious – you just need a probability score, and you need it as fast as possible.

Right-size your analytics stack with Aiven for ClickHouse

TL;DR If most of your Snowflake or Databricks spend goes to dashboards and reports, you are paying for a platform built for much bigger problems. Aiven for ClickHouse runs those workloads on a fixed plan, so adding dashboard users does not add to your compute bill. Native integrations with PostgreSQL and Apache Kafka also replace most of the ingestion and orchestration tools around your current platform. Move one dashboard at a time, and keep Snowflake or Databricks for work such as model training.

Bring Your Own Key: Encryption sovereignty without the headache

Let's start with an uncomfortable question that tends to surface exactly once, usually in front of an auditor, a customer's security team, or your own CISO: who can actually decrypt your data right now? For most managed databases and message queues, the honest answer is "the provider, technically, if they really wanted to." That's not a scandal. It's just how managed encryption-at-rest normally works: the provider generates the key, holds the key, rotates the key, and you trust them not to misuse it.

Build The Future: Building a Startup Inside a Scaleup

TL;DR Stan's experience ranges from raising millions for his own startups to working across programming, marketing, and PR. Here is how that versatile founder mindset fuels his Product Director role at Aiven today. Stan has had a lot of ownership from day one, in the most literal sense. Aiven headhunted him to build a new product from the ground up. "My onboarding was thirty minutes with my boss, who gave me a handful of really good and original ideas, he says. "After that, I had to figure it out.

How Etsy gets its mobile apps ready for peak traffic

Most advice about surviving a traffic spike is about capacity. Scale the fleet, warm the caches, load test the checkout path. All of it assumes you can fix whatever breaks the moment you find it. Mobile apps don’t work that way. I spent an hour on a workshop with Jay Henry, a senior engineering manager at Etsy who owns engineering strategy across three teams covering CI, build, test, release, observe, and SRE. Jay’s take: a web team having a bad day can revert in minutes.

Show off your uptime with SLA badges.

Your uptime is probably better than your customers think. But the proof sits in your UptimeRobot dashboard, where only your team can see it. Starting today, you can put it on your website. SLA badges let you set an uptime target for your monitors, then embed a live badge that shows the uptime you actually delivered, colored against that target. Add it to your site footer, your docs or your GitHub README, and get the result by email at the end of each period.

How Unified Cloud Monitoring Closes Hybrid IT Blind Spots

Modern IT runs across public cloud, data centers, containers, SaaS, networks, and AI services. Cloud monitoring delivers more value when those environments share the same observability context, giving teams a clearer way to troubleshoot service issues, manage cloud spend, and support AI-assisted operations.