Operations | Monitoring | ITSM | DevOps | Cloud

The latest News and Information on Monitoring for Websites, Applications, APIs, Infrastructure, and other technologies.

A Look Inside Skylar AI: New Features and Updates

See what’s new in the latest Skylar AI release as Vice President of Skylar AI Douglas "DJ" James walks through updates designed to deliver smarter insights, stronger connectivity, more trustworthy answers, and greater deployment flexibility. New features and updates include: Explore the latest Skylar AI capabilities and see how ScienceLogic continues to deliver smarter insights, greater flexibility, and trusted AI for modern IT operations.

Episode 15 - Melissa M. Reeve (part-2)

When it comes to effectively rolling out AI in companies, middle managers may be more important than ever as the link between strategy and daily work. In Part 2 of this conversation on The Intelligent Enterprise, host Tom Stoneman continues his conversation with Melissa M. Reeve, author of Hyperadaptive: Rewiring to Become an AI-Native Organization, to dig into the practical side of how businesses can become AI-native. Melissa argues that AI should be treated less like a software rollout and more like the arrival of the personal computer in the 1990s.

AI is changing how organizations operate

AI is changing how organizations operate, but one thing has not changed: critical services cannot fail. Whether it is financial markets, healthcare, or other mission critical environments, organizations need observability that delivers value quickly, not weeks or months later. In this clip with theCube, Virtana CEO Paul Appleby explains how Virtana combines high fidelity telemetry with AI-driven intelligence to discover dependencies, correlate relationships, and deliver actionable insights within hours.

Cribl On Your Coffee Break Episode 2 - Setting up Syslog

In our second video Leon picks on Syslog (because honestly, it deserves it). Cribl is the perfect tool to whip that disorganized, loud, unruly mess of a data stream into shape. By the time the month is over, you will have a pretty good idea of what Cribl can do, and how to do it. You’ll also have consumed more caffeinated beverages than is strictly appropriate...

Application Metrics caught my broken size estimator

There’s a very specific kind of frustration that comes from waiting several minutes for a video to encode, dragging it into a message, and getting hit with a “file too large” error. Then you’re blindly trying to shave off a few more megabytes by re-encoding, maybe at a lower resolution or a smaller bitrate, hoping you won’t have to do it more than one or two more times. Here’s how I used Sentry’s Application Metrics to make a more accurate video size estimator.

Visualize how CUPED adjusts experiment results with Datadog

CUPED (Controlled-experiment Using Pre-Experiment Data) is a powerful tool that can reduce metric variance and help teams obtain precise experiment results with less data. However, the difference between an experiment’s CUPED-adjusted lift and raw lift can be difficult to explain, especially when an experiment uses many pre-exposure metrics and subject properties. The CUPED adjustments visualization in Datadog Experiments breaks the difference into a sequence of specific adjustments.

From traces to experiments: A loop for improving AI agents

Let’s say your team shipped a support agent last quarter. The launch demo went well, stakeholders were pleased, and everyone moved on. A few months later, things start to look off. Summaries of long conversations are truncated, and monitors show latency spikes on tool calls to the billing API. Your team’s first instinct is to ship fixes such as tweaking prompts or upgrading the model.

Bringing the Most Advanced Sampling to the OpenTelemetry Collector

Sampling is a core skill that everyone who runs an observability pipeline at scale will learn. There are lots of tradeoffs within the various decisions you'll make from reducing bandwidth, CPU, and memory, to reducing costs and making the observability backend's performance better for users. Historically, there have only been three mechanisms, each with their own tradeoffs: However, there is a secret fourth option: adaptive tail sampling—which changes those tradeoffs.

We Let AI Agents Rewrite a 92M-Message-a-Day Service in Go. Zero Incidents.

Our Results Daemon processes about 92 million messages a day. We recently rewrote it from Node.js to Go, and we let Claude Code write it. We wanted to know whether we could trust an agentic rewrite for a critical, high-throughput production service rather than a prototype. It shipped with zero incidents, a 70% reduction in running pods, and a lighter database load.

Best Storage Monitoring Software: 10 Tools Compared

Storage rarely fails loudly. A pool fills. Latency climbs on one LUN. The first to notice is a user whose application timed out. The best storage monitoring software catches it earlier. It watches capacity, IOPS, latency and drive health across your arrays, which is what storage resource monitoring is for. In this blog, you will see: By the end you will know which one fits. Storage monitoring software tracks the health, capacity and performance of your IT storage.