Operations | Monitoring | ITSM | DevOps | Cloud

AI Observability Deep Dive Demo | Grafana Cloud

Grafana AI Observability is our new database and platform for observing AI Agents. Over the past year at Grafana Labs, we built Agents and we needed a way to understand how they are performing, what are the costs associated with them, what's the error rate or time to the first token as well as how they are behaving. Grafana Staff Engineer, Ivana Hučková provides a deep dive demo on how Grafana AI Observability connects our experience building Agents with our experience building observability systems.

Grafana Assistant Context Offloading

Context Offloading is a pipeline solution for managing Observability with AI Agents. If you are building AI Agents that work with real data, the context window can very easily get filled with bloated context that the Agent does not really need. Sven demonstrates "Context Offloading", a solution that stores the JSON result and sends only the summary of the JSON blob, making the LLM loop performance much quicker and keeping your context window small.

iFrame Expands AI Infrastructure Offering With Hosted Inference Service for Open-Weight Models

Organizations looking to reduce AI operating costs while maintaining performance are increasingly turning to open-weight models. This trend accelerated throughout 2024 as businesses sought alternatives to expensive proprietary systems and greater control over their AI infrastructure.

Why AI Evaluation Is Becoming a Business Priority, Not Just a Technical Task

Artificial intelligence products are evolving at a pace that challenges traditional quality assurance and validation processes. As organizations race to release new AI-powered features, many product teams face the same question: how do they know a system is ready for real-world use? As reported by AI Journal, conversations with product leaders across different sectors reveal a growing focus on AI evaluation as a critical part of product development. Their experiences highlight the challenges of balancing innovation, risk management, customer expectations, and future regulatory requirements.

Running AI at Enterprise Scale w/ Anthropic, Descope, Port, Rootly and Twingate

The debate about whether AI can write production code is over. Companies are handing work to fleets of agents, and for many, they write most of the code that ships to production. The next challenge is everything that happens once an entire engineering organization runs this way, at full speed. Teams that generate code 10x faster still review it at human speed, and that mismatch is now the constraint. Code ownership is also becoming an issue, as developers learn to trust agentic processes a little too much. When an agent breaks production, who is responsible?

AI Dev Tools: What 100K Engineers at Google Really Taught Us

AI developer productivity, agentic workflows, and the lessons learned running engineering tools for 100,000+ software engineers at Google. John Montgomery, CCO at GitKraken, sits down with Asim Hussain, co-founder of Alterion AI and former Google VP of Engineering Productivity, to get real about what AI actually changes for engineering teams in 2025.

Autonomous Error Remediation in Cursor with Lightrun MCP

Lightrun's Gidi Freud demonstrates how your AI coding agent can now investigate and fix production errors, autonomously. Watch how Cursor, guided by Lightrun's Error Remediation skill, picks up a Sentry error, instruments the live service with a runtime snapshot, captures real evidence, and opens a validated PR for approval.

21 AI concepts every beginner should know before their first interview

If you’re prepping for your first AI or MLOps interview, the hardest part usually isn’t always the hands-on element. For me, it’s the vocabulary. Interviewers sometimes lob single-word concepts at you (“what’s quantization?”) and watch how far you can carry the thread. The questions sound clear-cut, but each one is really a doorway into a bigger topic, and the interviewer is judging how cleanly you walk through it.

CloudZero AI Hub: The nexus of autonomous AI cost control

CloudZero originated as a way to make sense of your cloud costs. Costs spread across bills with billions of line items belonging to resources that might or might not have been tagged (or taggable), spun up by engineers working across teams, on different microservices, features, and products, that served a wide range of customers. Kubernetes. Multi-cloud. Check, check, check.

AI ROI: How to measure and provide the return on AI investments in 2026

Every quarter, the same scene plays out in boardrooms across the Fortune 500. The CEO asks: “What is the return on everything the company is spending on AI?” The CTO talks about productivity gains and developer velocity. The CFO points at a cloud bill that doubled but cannot isolate which line items are AI. The board nods politely and tables the discussion until next quarter, when the same question will produce the same non-answer. (If this sounds familiar, you are not alone. Keep reading.)