Operations | Monitoring | ITSM | DevOps | Cloud

The latest News and Information on Monitoring for Websites, Applications, APIs, Infrastructure, and other technologies.

What Is LLM Observability? A Complete Guide

If you run LLM features in production, your most dangerous failures are the ones your monitoring never flags. Your LLM feature passed every test, and the demo went great. Three weeks after launch, a support ticket lands: the chatbot quoted a refund policy that does not exist. The dashboards are all green, and the same prompt answers correctly when you retry it. This is the blind spot LLM observability exists to close. Your existing tools saw the request come back fast with a clean status code.

Why AI-Generated Code Needs Monitoring More Than Handwritten Code

Like it or not, vibe coding is here to stay. It’s too easy to just go away. Maybe if the per token cost rises too much at some point that it becomes cheaper to hire a junior… But until then, you’d better get used to it. For now, tools like Cursor, Copilot, and Claude let developers (and plenty of non-devs) ship full-stack apps faster than a junior is able to completely grasp the concept of the app they’re working on. And that’s pretty neat.

Beyond performance monitoring: Understand the user experience with Grafana Cloud Frontend Observability

You've optimized your Largest Contentful Paint. Your Time to First Byte is under 200ms. Your Lighthouse scores are green. And yet, your checkout conversion rate is quietly dropping. A segment of users in Southeast Asia is churning. Your support team is fielding tickets about a form that "just doesn't work" and you have no idea which one. Traditional frontend performance monitoring tells you whether your application is fast. It doesn't tell you whether people are actually succeeding when using it.

Episode 13 - AI: The Hidden Layer (Part 1)

What does enterprise AI look like when the user is no longer human? In this episode of The Intelligent Enterprise, host Tom Stoneman sits down with Ash Ashutosh, CEO of Pinecone and a three-time founder, to explore the next evolution of AI infrastructure: from vector databases built for humans using chatbots, to knowledge engines designed for AI agents that need to understand context, take action, and get work done.

Cursor outage on July 16, 2026: high load errors worldwide and how to keep working

Cursor was hit by a global “high demand” outage on July 16, 2026, returning ERROR_RESOURCE_EXHAUSTED errors that blocked AI requests for just over two hours. StatusGator caught it early, sending an Early Warning Signal at 06:49 UTC, 12 minutes before Cursor acknowledged the incident at 07:01 UTC. Here is the full picture, including the workaround that kept many users coding.

Five worthy reads: Brains or bots-are we forgetting how to think?

Five worthy reads is a regular column on five noteworthy items we’ve discovered while researching trending and timeless topics. This week, we are exploring how prolonged dependence on AI could influence human beings' neural pathways, cognitive habits, and the behavioral changes that follows. As children, many of us would have watched the juggler at a circus in amazement. One ball became two, then three, and several more since it was a cumulative act.

Webinar: Halo & SquaredUp - Dashboards that deliver

SquaredUp and Halo teamed up for a live webinar exploring why visibility matters, the reporting challenges many organisations face, and how HaloPSA data can be transformed into engaging dashboards using SquaredUp. We also heard from a customer on how they're using dashboards to manage costs and improve service delivery. Whether your goal is sharper reporting, greater visibility for customers, or a clearer demonstration of the value your services deliver, this session offers practical insights and real-world examples to help you get there.