Operations | Monitoring | ITSM | DevOps | Cloud

Tempo 3.1 release: new features for Kafka, TraceQL metrics updates, trace redaction, and more

Building on the major release of Tempo 3.0, Tempo 3.1 is here, delivering community-contributed Kafka client improvements, query-based trace redaction, sampling-aware TraceQL metrics, and more. Together, the updates in 3.1 make it easier to operate Tempo, get accurate insights from your trace data, and investigate issues more efficiently. You can continue reading and check out the video below to learn more about the latest features.

What if your agent's hallucinations had a budget? How to start using SLOs for agent behavior

At Grafana Labs, observability is what we do. So as we started building AI agents, we naturally reached for the same instincts we bring to every system: measure it, set targets, and make reliability something you can reason about instead of hope for. That instinct led us somewhere unexpectedly useful. It turns out one of the oldest ideas in reliability engineering, the error budget, maps beautifully onto one of the newest problems in software: how do you know if an AI agent is actually any good?

Grafana Alerting: Scale alert routing without scaling complexity using multiple notification policies

Alert routing often starts simple. A team creates a few contact points, adds some label matchers, and builds a notification policy tree that sends each alert to the right destination. But alerting configurations rarely stay simple. As an organization grows, its notification policy tree must accommodate more teams, services, and routing requirements. Changes for one team still require editing a global configuration, making ownership less clear and independent provisioning harder.

Digital Experience Monitoring with Grafana Cloud: Session Replay, synthetic checks, and faster investigations

When something breaks in production, the questions that matter most are also the toughest to answer from metrics alone: who was affected, what did they actually see, and is this worth waking someone up for? Answering those questions requires a fuller picture of the issue and its impact on your users. That’s where Digital Experience Monitoring (DEM) in Grafana Cloud comes in.

Custom labels in Grafana Cloud Synthetic Monitoring: New updates for consistency and ease-of-use

Labels are a powerful way to organize telemetry and define policies across Grafana Cloud, helping to streamline alerting, attribution, access control, and more. But traditionally, custom labels in Synthetic Monitoring have worked a little differently: they only lived on a single sm_check_info metric, and Grafana Cloud prefixed each one with label_.

How to monitor Cypress tests with Grafana Cloud

If your Cypress suite has tests that fail more often or run slower, you know it can be hard to figure out the pattern from a single job. It could be one spec that slowed down, or a single test that fails, or maybe the entire suite is trending slower. The root cause could be a bug in the app, or a flaky test, or something else.

How to measure and improve instrumentation quality for better full-stack observability

Modern engineering teams instrument everything, with metrics, logs, traces, and profiles flowing from hundreds of services at once. But full-stack observability isn’t really about collecting more telemetry; it's about having a single, unified picture of how your services connect to every layer beneath them, including their dependencies, the pods and nodes they run on, and the logs, traces, and profiles that explain their behavior.

The Grafana AI SDK for Go: a shared foundation for building AI applications

Starting an experiment with an LLM has never been easier. Keeping a growing collection of those experiments consistent is another matter. Earlier this year, as more teams began exploring AI features here at Grafana Labs, we repeatedly encountered the same pattern: a new experiment would start, move quickly, and build its own client for whichever model provider it needed. The next experiment would do the same, with a slightly different abstraction for streaming, tools, errors, or provider configuration.

How to scale Alloy as a central telemetry gateway: capacity planning, load testing, and production lessons

Running Alloy as a single-instance sidecar is simple. Running it as a centralized gateway that absorbs the full telemetry stream of an enterprise platform—tens of millions of active series, terabytes of logs per day, and tens of thousands of trace spans per second—is a different challenge altogether. To get it right, you need deliberate capacity planning, honest load testing, and a monitoring setup that doesn't rely on the very thing you're testing.

Grafana 13.2 release: easier ways to query and explore your data

Grafana 13.2 is here, bringing more improvements to help you and your team explore your data and get to insights faster. Download Grafana 13.2 In this post, we’ll highlight the latest updates to saved queries, a feature that lets teams share, discover, and reuse queries to get to trusted answers faster and help new teammates get up to speed. We’ll also explore how the new View panel sidebar makes exploring busy panels a breeze.