Operations | Monitoring | ITSM | DevOps | Cloud

Raygun APM Agent 3.1: async traces that stay with the right request

Raygun APM Agent 3.1 introduces more accurate asynchronous request tracing for Windows, Linux, and Azure App Service. Version 3.0 rebuilt the foundation of the Agent, profiler, installers, and release pipeline. Version 3.1 builds on that work with a focused improvement for ASP.NET Core: automatic request correlation that follows asynchronous execution without requiring developers to instrument their application.

Two cats, two dogs, four vendors, and the model the AI couldn't find (Tech Talk Companion)

Tech Talks went dark for a few months, and on episode 13 I finally got to ask why. Mathias Palmersheim’s answer, delivered completely straight, was that his users were unhappy with the availability and usability of their feeders and their litter box, and he wasn’t allowed back on stream until that got fixed. The users are two dogs and two cats, and they have titles. Maisie, a Shiba Inu who came to him through a rescue, is the recently promoted chief executive pawofficer.

Relational Query Superpowers

I'm investigating repeated errors in my e-commerce application, and I need to get enough context in a single Honeycomb query to piece the entire picture together. Each query returns events based on the event's WHERE clauses, but I want to know several things from outside of the event that recorded an error. Things like: Those attributes are all over the trace. That's going to make a single query tough, right? Wrong!

ITSM for Healthcare: IT Service Management in Hospitals and Health Systems

How long does a nurse stand at a workstation waiting for a record to load before the ward gives up and reaches for paper? In most hospitals nobody measures that number, and the ticket that reaches IT describes a symptom instead of a cause. Hospital IT support runs on a different clock from corporate IT. There is no quiet Sunday night and no safe window for maintenance.

Log Filtering: How to Cut Log Ingest Volume Without Losing Evidence

Every log estate reaches a point where volume grows faster than the value inside it. The usual response is to find the biggest source and drop it. Cutting volume is the easy part. Cutting the right half takes judgment. Log filtering is only one of four options for an expensive source, and the other three matter just as much. In this blog, you will: By the end you can defend every rule you write, including the ones that keep data. A volume cut fails in two directions.

From Monitoring to Prediction: How Fleet Data Is Changing Maritime Operations

Most maritime operators already collect more fleet data than their shore teams can meaningfully use. Positions appear on screens, real-time data arrives from onboard systems, and reports document vessel performance throughout a voyage. Tracking where a ship is has become the easy part. The harder question is what happens when that information starts indicating what the vessel will do next. The change in maritime operations comes down to shifting from reviewing events to anticipating them early enough to alter an operational decision.

Monitor HTTPS and SVCB Records with DNS Check

DNS Check now supports monitoring HTTPS records and SVCB records, DNS record types 65 and 64, both standardized in RFC 9460. They tell a client how to connect to a service rather than only where it is: which HTTP versions the endpoint speaks, which port it listens on, which addresses it can start connecting to, and which keys it needs for Encrypted ClientHello, all before it opens a connection.

Hybrid cloud management: 6 challenges IT teams need to solve in 2026

In 2026, a hybrid cloud is no longer something organizations are working toward; it's already where they are. According to Forrester's The State Of Cloud Series 2026, the vast majority of enterprises across major markets, including the United States, India, Australia and New Zealand, Canada, and the Asia-Pacific region, are running some form of a hybrid cloud, combining public cloud platforms with private infrastructure, colocation data centers, and sovereign cloud providers.

How to Cut SIEM Ingest by 90% Without Losing Detection Coverage

Every SOC team knows the trade-off. Send everything to the SIEM platform and pay for it. Or filter aggressively and risk missing something. Filter lists are written once, during onboarding. Detection content keeps moving after that. Smart Engine, the new core of the VirtualMetric DataStream pipeline, takes the guesswork out of that decision. It reduces SIEM ingest using your registered detection rules. An event that no registered detection could match is dropped.

Live Debugging for Critical Systems: MTBF, MTTR & MTTA

A critical system has to stay reliable without new failures or added downtime, and live debugging, confirming the root cause without stopping the system, is often the only way to do that. In practice, this means having runtime context: on-demand evidence generated at the point of failure rather than logging configured months earlier, which is what keeps MTBF up, MTTR, and MTTA down.

Can we live dangerously? Sandboxing Claude, and the Claude foreman that runs the rest

While logging into one’s LinkedIn will spew out endless talk of AI possibilities from “thought leaders” and the semi-disconnected alike, another pocket of the world spent the last few weeks watching the Shai-Hulud worm chew through npm. A self-propagating credential stealer that hit 400-plus packages and, delightfully, planted Claude Code and VS Code hooks so just opening the repo could run its payload.

Prompts, skills, and the AGENTS.md nobody wants to write (and how Anthropic writes theirs)

You’ve watched Claude Code compact a conversation. The context bar fills, it pauses, a summary appears, and it carries on like nothing happened. You probably assumed a housekeeping script trimmed the transcript in the background. It didn’t. The model compacted itself. When the window fills, Claude Code sends a long, specific prompt telling the model how to summarize its own conversation. Then it does, same model, same turn. The thing managing your context window is just another instruction.

Agent vs Agentless Monitoring and How to Decide What Goes Where

Why does half the infrastructure end up returning no monitoring data? The standard plan is to install collection software on everything, which moves quickly across servers and stops dead at the first device running closed firmware. Storage arrays, firewalls, and switches will never accept an install, and the rollout stalls there. That plan usually gets set once for the whole environment, with a single collection model applied to hardware it was never suited for.

Help Desk Software for Schools: Managing IT Support Across Campuses

How many support requests reach school IT staff each week without ever becoming a ticket? A teacher stops a technician in the corridor about a projector that will not connect. An office administrator sends a direct email about a locked account, and a student tells the librarian their laptop stopped charging during second period. Help desk software for schools collects those requests into one queue, routes them by site and category, and keeps a record of what was done.
Sponsored Post

Raygun APM Agent 3.0.14: faster, simpler, and ready for ARM64

Today we are releasing Raygun APM Agent 3.0.14 for Windows, Linux, and Azure App Service. This release is the result of a substantial modernization of the Agent, profiler, installers, and release pipeline. It makes Raygun APM easier to deploy, reduces overhead in several critical paths, adds native Linux ARM64 support, and lets developers investigate APM data through Raygun API v3 and the Raygun MCP server. If you are upgrading from version 2.3.0, there is much more here than a version-number change.

Top Tips: How to be a tech-savvy traveler

Top tips is a weekly column where we highlight what’s trending in the tech world and share ways to stay ahead. This week, let's look at a few ways you can be a tech-savvy traveler. Being a traveler is not easy, but with today's modern technology, it has become much easier. When we travel to places with no network, we sometimes forget about the ways we can use technology. Excluding the more familiar, I'm going to list some lesser-known tips. 1.

Automate Product Analytics reports with your agent and the CX CLI

Every page view, click, and session your RUM SDK captures lands in Coralogix as a log event under the cx_rum subsystem — the raw data behind how people actually use your product. You can turn it into a shareable report without writing a single query. Just ask your coding agent. Your agent queries that data through the CX CLI and writes the report for you: describe what you want in plain English, get a formatted report back — without leaving the terminal.

10 Top Website Monitoring Tools for Uptime, Page Speed and Real User Data

Your uptime tool reports the site as available all month. Support reports something else, because customers in one region spent the morning unable to complete a checkout. Both records are accurate, and that is the problem. An external check confirms the page answered. It cannot tell you that the answer took nine seconds for everyone routed through one CDN edge, or what caused the delay.

The pager shouldn't be what starts the investigation

Authored by Greg Janco, Engineering Manager at Mezmo I've always thought there was something backwards about incident response. An alert fires at 3 a.m. PagerDuty does its job. Somebody wakes up, grabs a laptop, connects to the VPN, opens the alert, and then starts answering the same basic questions we ask at the beginning of almost every incident. What changed? What else is broken? Have we seen this before? Some of those need a human eventually. A lot of the first pass doesn't.

Cribl On Your Coffee Break Episode 4 - Gathering REST data

In the 4th installment of our series, Leon looks at Cribl’s ability to collect REST API data. By the time the month (and the series) is over, you will have a pretty good idea of what Cribl can do, and how to do it. You’ll also have consumed your body weight in caffeinated beverages...

Build and run Datadog workflows from Bits Chat or AI agents

Teams use AI coding agents and Bits Chat to troubleshoot systems and handle complex tasks, often uncovering repetitive work worth automating. But turning those routines into workflows can still require switching tools and recreating context manually. Through the Datadog MCP Server, Workflow Automation now lets you build workflows from Bits Chat or AI coding agents like Claude Code, Cursor, and Codex.

August 2026 Early Warning Signals

August brought notable outages across developer platforms, SaaS tools, communications services, and cloud applications. StatusGator detected 789 Early Warning Signals during the month. Of those, 153 incidents (19.39%) were acknowledged by providers, while 636 (80.61%) were not officially acknowledged. StatusGator’s Early Warning Signals often surface service disruptions before providers post an official update.

SAP HANA Monitoring Tools 2026: How to Compare Options and Simplify Monitoring

If your team works with SAP HANA, the main challenge usually isn’t finding another dashboard. The real difficulty is identifying which tool can quickly help you trace vague complaints about slow transactions to their root cause. In many setups, SAP HANA monitoring is divided among native SAP interfaces, cloud monitoring, infrastructure dashboards, and the broader monitoring systems used by the rest of IT. This fragmented approach can slow down root cause analysis.

Cribl On Your Coffee Break Episode 3 - Configuring Prometheus Remote-Write

In day 3 of our coffee break series, Leon continues to explore common observability data types and how to get them into Cribl. Today, we’ll look at setting up a simple Prometheus ingestion. By the time the month (and the series) is over, you will have a pretty good idea of what Cribl can do, and how to do it. You’ll also have consumed your body weight in caffeinated beverages...

Why you should (not) build your own observability stack

If you are able to build it better than your vendor, then change your vendor. Not build it. Rishi builds large-scale observability systems at Last9, focusing on reliable and cost-efficient telemetry infrastructure, and writes about the practical lessons learned while operating ClickHouse, VictoriaMetrics, and OpenTelemetry in production.

Icinga Web SSO walkthrough

The ability to log into all corporate applications with one username and password is pretty convenient, even compared to a password manager. As a benefit, the IT department can centrally enforce one desired two-factor auth mechanism. Now we, Icinga, also provide a so-called OpenID Connect integration for single sign-on. By the end of this text you’ll know how to connect your Icinga Web instance to the ID provider of your choice.

6 Signs a Dedicated Log Tool Fits Better Than a Full Observability Platform

Most growing teams eventually consolidate onto a full observability platform, and for teams correlating logs, metrics, and traces across a complex system, that’s often the right call. But a dedicated log tool still wins for a specific set of teams: ones that need to move fast, keep costs simple, and get real answers from logs without carrying the weight of a platform they don’t fully need yet. Here’s when that’s you.

Log Processing: What Happens to a Log Line Before You Can Search It

A log line arrives as plain text and leaves as a record you can query. Six steps sit between those two states. Each one adds something useful, and each one costs you time, CPU, or storage. Most teams never look at that chain until a search comes back empty. Here is what log processing does to an event, step by step: By the end you can look at your own chain. You will know what each step buys you. Six steps turn a raw log line into a searchable record.

PII Redaction in Logs: Mask, Redact, Hash, or Drop?

Sensitive values reach your logs without anyone deciding they should. A debug line prints a whole request object. An error message carries the query string. A customer email address is suddenly stored in three systems. PII redaction in logs then gets treated as one setting to switch on. In practice it covers four separate treatments. The value is already inside the message before log ingestion finishes. In this blog, you will: By the end you can write a rule for each field and defend it.

Getting Started with InfluxDB 3 and Grafana Tutorial

Summary This guide walks through an end-to-end Grafana and InfluxDB 3 integration using a realistic dataset you generate yourself. The tutorial covers getting data in, transforming it, connecting Grafana, and building real dashboards. Table of Contents InfluxDB and Grafana are the most common pairing in time series monitoring, and division of labor between them is simple.

Monitor prompt caching to optimize your token usage

Datadog’s 2026 State of AI Engineering report showed organizations’ LLM inputs swelling rapidly as context engineering expands. In March 2026, 69% of all input tokens in Datadog customer traces were for system prompts: internal instructions, policy definitions, and tool guidance providing context and guardrails around the user input. This suggests that most context engineering spend among Datadog customers is going toward optimizing repeating system prompts in heavily scaffolded agent systems.

Azure integration now supports service principal authentication

We’ve released some improvements to our Azure status integration. StatusGator can now read your Azure Resource Health events via a service principal. Previously the only supported authentication mechanism was OAuth. Both pull the same data and produce the same alerts – the difference is who the connection belongs to, and what happens to it over time.

Application Metrics caught my broken size estimator

There’s a very specific kind of frustration that comes from waiting several minutes for a video to encode, dragging it into a message, and getting hit with a “file too large” error. Then you’re blindly trying to shave off a few more megabytes by re-encoding, maybe at a lower resolution or a smaller bitrate, hoping you won’t have to do it more than one or two more times. Here’s how I used Sentry’s Application Metrics to make a more accurate video size estimator.

From traces to experiments: A loop for improving AI agents

Let’s say your team shipped a support agent last quarter. The launch demo went well, stakeholders were pleased, and everyone moved on. A few months later, things start to look off. Summaries of long conversations are truncated, and monitors show latency spikes on tool calls to the billing API. Your team’s first instinct is to ship fixes such as tweaking prompts or upgrading the model.

Bringing the Most Advanced Sampling to the OpenTelemetry Collector

Sampling is a core skill that everyone who runs an observability pipeline at scale will learn. There are lots of tradeoffs within the various decisions you'll make from reducing bandwidth, CPU, and memory, to reducing costs and making the observability backend's performance better for users. Historically, there have only been three mechanisms, each with their own tradeoffs: However, there is a secret fourth option: adaptive tail sampling—which changes those tradeoffs.

We Let AI Agents Rewrite a 92M-Message-a-Day Service in Go. Zero Incidents.

Our Results Daemon processes about 92 million messages a day. We recently rewrote it from Node.js to Go, and we let Claude Code write it. We wanted to know whether we could trust an agentic rewrite for a critical, high-throughput production service rather than a prototype. It shipped with zero incidents, a 70% reduction in running pods, and a lighter database load.

Best Storage Monitoring Software: 10 Tools Compared

Storage rarely fails loudly. A pool fills. Latency climbs on one LUN. The first to notice is a user whose application timed out. The best storage monitoring software catches it earlier. It watches capacity, IOPS, latency and drive health across your arrays, which is what storage resource monitoring is for. In this blog, you will see: By the end you will know which one fits. Storage monitoring software tracks the health, capacity and performance of your IT storage.

How NIST Compliance Turns Observability Data Into Audit Evidence

Can you prove, on demand, which production systems were under continuous monitoring last quarter? Buyers, auditors, and insurers all ask a version of that question, and the answer decides contracts as often as audit findings. NIST compliance means aligning security controls and operations with standards from the National Institute of Standards and Technology, then holding evidence that the alignment stayed continuous. The frameworks are precise about outcomes and quiet about mechanics.

Compliance Doesn't Fail on Audit Day. It Drifts Every Day in Between.

For CIOs, compliance is no longer simply a box to check at audit time. It has become part of the operating standard for resilient, accountable, and well-managed enterprise IT. The reason is straightforward: enterprise technology environments change continuously. Infrastructure scales. Configurations change. Cloud resources move. Exceptions accumulate. Dependencies evolve across hybrid and distributed architectures.

Visualize how CUPED adjusts experiment results with Datadog

CUPED (Controlled-experiment Using Pre-Experiment Data) is a powerful tool that can reduce metric variance and help teams obtain precise experiment results with less data. However, the difference between an experiment’s CUPED-adjusted lift and raw lift can be difficult to explain, especially when an experiment uses many pre-exposure metrics and subject properties. The CUPED adjustments visualization in Datadog Experiments breaks the difference into a sequence of specific adjustments.

Cribl On Your Coffee Break Episode 2 - Setting up Syslog

In our second video Leon picks on Syslog (because honestly, it deserves it). Cribl is the perfect tool to whip that disorganized, loud, unruly mess of a data stream into shape. By the time the month is over, you will have a pretty good idea of what Cribl can do, and how to do it. You’ll also have consumed more caffeinated beverages than is strictly appropriate...

Data-Driven Decisions Accelerate IT Results

Modern IT teams, having moved beyond the traditional reliance on hunches and personal experience that once shaped their day-to-day choices, no longer operate on intuition, since every meaningful decision now rests upon measurable, verifiable evidence gathered from their systems and workflows. Every deployment, capacity change, and incident response now depends on measurable evidence, not guesswork. Companies that base their operations on concrete numbers ship faster, recover quicker, and allocate budgets with far greater accuracy.