Operations | Monitoring | ITSM | DevOps | Cloud

The latest News and Information on Observabilty for complex systems and related technologies.

We Stuck Minecraft on a Kubernetes Cluster and Observed it with Open Source

Why did we do this? Not important: jump to 1:31 to see the Pis. TL;DR: We stuck Minecraft on kubernetes running on 4 Raspberry Pis in a 3D-Printed case, and monitored it with open source observability. Huge thanks to Percona DBA Ivan Zaitsev for putting the demo together for Percona University, Montevideo. Coroot automatically collects and visualizes all your telemetry data: logs, metrics, traces, profiles, and a complete map of your services. With the complete context of eBPF, it can diagnose the exact cause of an incident in seconds, and show you the exact commands to fix it.

OpenTelemetry Collector Configuration for LLM Observability

Your LLM application emits telemetry unlike anything else in your stack. Model calls, tool invocations, retrieval steps, and token usage arrive as spans whose attributes carry entire prompts and completions. That data is bulky, it's full of user content you may not be allowed to export, and depending on which instrumentation each service uses, the same fact can arrive under different attribute names.

Beyond Traditional Observability: Turning Technical Insight into Operational Intelligence

Observability has become a central part of modern IT operations and for good reason. Metrics, logs and traces give technical teams detailed evidence about how applications, infrastructure and services are behaving. Such evidence helps them investigate performance degradation, identify abnormal behavior and understand what changed around the time an issue occurred.

Observability vs Monitoring: Why Does IT Still Find Out After the Business Does?

✓ operational truth IT finds out late because traditional monitoring is built to detect what goes wrong, not what has quietly stopped happening. Closing that gap requires observability that validates business journeys end to end, detects missing activity, checks its own coverage, and predicts degradation before a threshold is ever crossed.

How Honeycomb Private Cloud Drinks From the Fire Hose

Back in July, I wrote about how the Tenant team (the team behind Honeycomb Private Cloud (HPC)) has embraced the code review bottleneck to focus more of its work. One of the other challenges we have is that we're downstream of almost all the other teams at Honeycomb, meaning that we have to package up everyone's code and services, and how it gets provisioned! This is something impossible to handle through code review since there are so many engineers on other teams, and so few of us.

The Year-2 Price Cliff: What Your Observability Stack Really Costs Over 3 Years

I’m not the one whose phone lights up at 3 a.m. when production breaks. But I’ve spent years working alongside the engineers who are, and I’ve noticed that observability migrations happen for two reasons. Either engineering needed one, or, far more often, a quote landed that looked too good to refuse. The engineers rarely regret the first kind. The second kind they tell me about in year two, usually with a renewal notice in hand.

Dr. Cat Hicks on the Psychology of Software Teams

What happens when a software engineer who has put their entire identity into being the resident expert of an obscure technology or language with ten years of experience and who knows the codebase like the back of their hand now has to compete with AI? Psychologically speaking, according to Dr. Cat Hicks, author of The Psychology of Software Teams and founder of Catharsis, that's called identity threat, and it's a surefire way to feel unsafe. We were honored Dr.