Operations | Monitoring | ITSM | DevOps | Cloud

The latest News and Information on Monitoring for Websites, Applications, APIs, Infrastructure, and other technologies.

Searching Sentry Logs with Regex

We use Sentry's new Regex search for Logs to hunt for non-obvious bugs in our app. What do you do when traces show that Postgres jobs are backing up? Logs are trace-connected, so by looking at one of the affected traces, we can see all of the logs on that trace. In there, we can see Postgres has logged that it acquired a lock only after a long wait. Using regex, we can find all of these logs that describe acquiring a lock after a 10-second-plus wait. That narrows our search from thousands or hundreds of logs down to about a dozen.

Log in, look around, level up: Building a trial people want to explore

We turned the mic around for this episode, with Zoe Hawkins interviewing Adam White and Jake Lee about Sumo Logic’s redesigned free trial. The new 14-day sandbox comes preloaded with realistic data feeds, Cloud SIEM, and our AI agents, so no one has to build collectors or open tickets before finding out whether the platform fits. Adam and Jake walk through the guided “choose your own adventure” paths, the broad, open-ended questions that bring out the best in Mobot, and how the trial will keep pace with new releases.

We Stuck Minecraft on a Kubernetes Cluster and Observed it with Open Source

Why did we do this? Not important: jump to 1:31 to see the Pis. TL;DR: We stuck Minecraft on kubernetes running on 4 Raspberry Pis in a 3D-Printed case, and monitored it with open source observability. Huge thanks to Percona DBA Ivan Zaitsev for putting the demo together for Percona University, Montevideo. Coroot automatically collects and visualizes all your telemetry data: logs, metrics, traces, profiles, and a complete map of your services. With the complete context of eBPF, it can diagnose the exact cause of an incident in seconds, and show you the exact commands to fix it.

Reducing Android scope-sync overhead in Sentry Flutter

Our SDK adds work to the app that installs it. It records recent app activity as breadcrumbs and keeps user information and other diagnostic data up to date. Changes to that information are called scope updates. On Android, we send those updates to a worker isolate, which passes them to the Sentry Android SDK. We found that the calling isolate and the worker both normalized the same data. The encoding step also created a JSON string and an extra byte buffer that we could avoid.

Run incident response in your FedRAMP High environment

Earlier this year, Datadog for Government achieved FedRAMP High certification, extending our GovCloud environment (US1-FED) to the federal government’s most sensitive civilian workloads. That certification now covers Datadog Incident Response, bringing paging, incident coordination, automation, and postmortem workflows into US1-FED. When a government system goes down, responders need to reach the right people, coordinate a fix, and keep stakeholders informed.

Define user actions on your web app with visual labeling in Product Analytics

Adding or renaming a product event has traditionally meant creating an engineering ticket. Datadog Product Analytics uses the same SDKs and configuration as Real User Monitoring (RUM), and those SDKs autocapture actions such as clicks and taps. But an automatically generated action name describes the element rather than the user intent behind it.

Where Jev fits in ops

If you're using agents and MCPs to get a better understanding of your environment or work through an investigation, you can get a lot of useful information back. You can pull logs, look at recent changes, and check how services are configured, but you're still the one deciding what to do with all of it. That part of the process still lives in your head. To see where Jev might fit, look at decisions your team already makes and work backwards from them.