Operations | Monitoring | ITSM | DevOps | Cloud

Visualize Data Your Way, with Intelligence Dashboards Built for Your Stack

When something goes wrong in production, you do not want to spend the first five minutes rearranging charts. You want the error rate, throughput, queue depth, memory, and view metrics that show when customers are having issues while using your app.

What if your agent's hallucinations had a budget? How to start using SLOs for agent behavior

At Grafana Labs, observability is what we do. So as we started building AI agents, we naturally reached for the same instincts we bring to every system: measure it, set targets, and make reliability something you can reason about instead of hope for. That instinct led us somewhere unexpectedly useful. It turns out one of the oldest ideas in reliability engineering, the error budget, maps beautifully onto one of the newest problems in software: how do you know if an AI agent is actually any good?

You made coding faster. Guess where the bottleneck went next.

Somewhere in the last year, your team's code output went up. Pull requests are opened faster. The backlog of small fixes and routine changes started clearing quicker than it used to. If delivery still feels roughly as slow as it did before, that's what happens when you speed up one part of a process without touching anything downstream of it.

We stopped asking an LLM how much its own work would cost

There’s a specific kind of measurement problem worth naming precisely rather than dramatizing: this month we found that our model-routing agent was assigning a token budget to every unit of work, and that budget was noise in the strict sense. Fixing it meant improving a system that’s mostly right, not tearing one down.

Shipped: A customer support experience that starts with an answer

When you have a question about your cloud or AI spend, you want an answer quickly, not a ticket that disappears into a queue. Support should not mean waiting for business hours, repeating your account details to multiple people, or wondering whether anyone picked up your message. That changed this week for every CloudZero customer. You get answers to most product and account questions immediately, at any hour, and when your question needs a person, they already have context.

EU data residency for Hosted OpenSearch on Logit.io

EU buyers asking for Hosted OpenSearch with data residency usually mean something precise: indexes and cluster storage should land in a European data centre that lines up with GDPR expectations and the geography named in the DPA — not a US default that security later has to unwind. On Logit.io that choice is an account-level data storage region, not a free toggle on every stack. Get the first stack right and every later OpenSearch or log stack in that account follows.

AlloyScan vs. Alloy Discovery: Why AlloyScan Is the Way Forward

What does AlloyScan improve over Alloy Discovery? See how the two compare and what AlloyScan brings to modern IT environments. AlloyScan builds on Alloy Discovery’s proven discovery and audit capabilities and brings them to a modern, browser-based platform. It keeps the familiar collection approaches customers already know while making inventory easier to access, manage, analyze, and connect with other systems.

Incident Management Best Practices for Modern IT Teams

Incident management used to be easier to picture: an alert arrived, a ticket opened, a support team followed a process, and service returned. Today’s incidents move across cloud platforms, SaaS applications, networks, identity services, observability tools, ITSM queues, and engineering teams. The fundamentals still matter, though. Clear process, ownership, communication, and learning matter more when the environment becomes harder to understand. What has changed is how teams execute them.

What is a CRC Error and How to Find the Faulty Link Before It Slows the Business

Why does a switch port look healthy on every dashboard while people on that floor keep reporting dropped calls and slow file transfers? Often the cause is CRC errors, a port counter that most dashboards do not show by default. Each CRC error means a unit of data arrived damaged and the switch discarded it. Put simply, a CRC error tells you the receiving device noticed the data changed somewhere in transit.