Operations | Monitoring | ITSM | DevOps | Cloud

Sponsored Post

5 Ways to Use Log Analytics and Telemetry Data for Fraud Prevention

As fraud continues to grow in prevalence, SecOps teams are increasingly investing in fraud prevention capabilities to protect themselves and their customers. One approach that's proved reliable is the use of log analytics and telemetry data for fraud prevention. By collecting and analyzing data from various sources, including server logs, network traffic, and user behavior, enterprise SecOps teams can identify patterns and anomalies in real time that may indicate fraudulent activity.

Closing the AI gap: How next-generation knowledge access unlocks mission outcomes for government

A recent IDC Spotlight report based on a survey of 685 public sector respondents found that 72% describe scaling AI from pilot to production as "very" or "somewhat" difficult.¹ Choosing the right model is only part of the challenge. Agencies also need to get their data ready for AI.

Stop Writing Log Lines: Use eBPF to Catch PII and Credentials

Tired of out-of-control log expenses and manual logging discipline? Discover next-generation observability with Speedscale. By using eBPF to record full-fidelity data right off the wire, you can instantly run full-text searches, track down leaking PII, and securely map out credentials across HTTP, Postgres, gRPC, and more—all without writing a single log line. Learn more: speedscale.com.

Tech Talk | From Insight to Action: Reducing Metrics Costs in Splunk Observability

Explore how to leverage observability insights to reduce metric costs with Span Observability. Led by Tomasz Romaniuk and Martyna Karbownik from Cisco, this tech talk covers the significance of metric time series in impacting cardinality and consequently driving metric costs. After a theoretical introduction, the presentation dives into practical real-life use cases demonstrating how to effectively utilize available tools for cost reduction. The session concludes with an overview of future developments and a Q&A segment for audience inquiries.

What Is a CBOM and Why Does It Matter?

Quantum computing is reshaping the future of cybersecurity. In this video, learn what a Cryptographic Bill of Materials (CBOM) is, why it matters for post-quantum security, and how it helps organizations identify vulnerable cryptography, prioritize migration, and improve cryptographic visibility. Perfect for security professionals, DevSecOps teams, IT leaders, and compliance teams preparing for the quantum era.

Graylog MCP Howto Webinar

In this video, we walk through connecting the Graylog MCP Server (introduced in Graylog v7.0) to Claude CLI, enabling natural language interaction with your Graylog instance through Claude Desktop. Topics covered: Whether you're a Graylog admin looking to speed up investigations or a security engineer curious about AI-assisted log analysis, this walkthrough gives you everything you need to get MCP running end to end.

How AI Agents Are Changing ITSM Faster Than We Think | Motadata Webinar

AI is transforming ITSM beyond co-pilots with autonomous AI agents that can understand, decide, and act. In this webinar, discover how AI agents are reducing manual workloads, accelerating incident resolution, improving service quality, and enabling intelligent service operations. Learn real-world use cases, key implementation strategies, and how organizations can move from reactive support to proactive, AI-driven IT service management. Watch now to explore the future of smarter, faster, and more efficient ITSM.

Install AURA to Debug Incidents Using an Open Source SRE Agent

AURA is a fully open-source agentic harness built for SRE and production operations work. In this walkthrough, Mezmo forward deployed engineer Jeff iinstalls AURA on a local desktop, runs `aura init` to generate the config and connect it to an Anthropic Sonnet model, then wires in a Grafana MCP server pointed at his homelab. He hands AURA a live incident: a set of addressable LED lights that stopped responding to Home Assistant.

9 Best Log File Analysis Tools for IT and DevOps Teams

An incident is open and the evidence is scattered. The application logs point to a connection timeout; the load balancer shows nothing unusual, and the container that produced the original error was replaced eighteen minutes ago. Three engineers are logged into three separate hosts running the same search, and the log line that would explain it has already rotated away. That is the moment most teams start shopping for a log file analysis platform.

Executive Roundtable: Guardrails for the Autonomous Era

Ask ten engineering leaders how far their organization has actually gotten with autonomous AI, and most will admit the same thing once the marketing language drops away: not nearly as far as it looks from the outside. That was the undercurrent of a roundtable Logz.io and Twingate hosted on July 22, 2026, bringing together VPs of engineering, CTOs, senior security directors, and several product leaders across different industries and company stages.

Elastic's new metrics capabilities will dramatically improve uptime for public sector IT

The new columnar metrics engine in Elastic Observability enables public sector IT teams to combine logging, metrics, and traces in one platform. As a result, SREs can improve uptime while protecting taxpayer dollars in the process. Public sector site reliability engineers (SREs) operate under a distinct set of pressures, whether that’s supporting a federal agency, a health department, a public university, or a transit authority.

On Release Days We Wear Teal Episode for release 4.19

In this episode, Leon explores some of the new features, functions, updates, and improvements in release 4.19, which includes a raft of AI-enabled features including the Cribl Apps, integrated MCP server, and the fact that AI features are now turned on by default. For more information, check out these links.

How to structure a log

You’ve decided to step up your logging game and start sending more valuable, structured logs that you can query, aggregate, and use for debugging in production. Go, you! Now, uh, how do you actually write them? We’re not going to spend much time on what you should log. We’ve covered that already, a few times before. What we will be covering is how to actually write those logs, answering questions like: What makes a log structured is not just pairing messages with arbitrary JSON objects.

I'll have my AI agent call your AI agent: Battle for your digital hub

On this episode of Masters of Data, we unpack what it actually means to expect AI to be the primary interface for everything we do. We dig into the pull toward centralizing work in a single hub like Claude versus staying spread across specialized tools like Slack, Asana and Zoom, and where the line sits between helpful automation and letting an agent speak on your behalf. We also get into the "chief of staff" agent workflow for daily roundups and why specialized, best-of-breed tools aren't going anywhere, even as hubs get smarter.

Why partners love working with Cribl

Hear directly from Cribl partners—including AWS—about what it’s really like to work together. This short is for technology and cloud partners, consulting firms, and customers who want a quick, human view of Cribl’s partner ecosystem and the value it delivers. In under two minutes, partners highlight Cribl’s partner program, the people they work with, and the outcomes they’re delivering for joint customers. You’ll hear about the FedRAMP opportunity, why “it’s all about the data” for AWS, and how Cribl helps get data where it needs to be for shared customers.

Prometheus Metrics Just Got a Cardinality Fix: What Native Histograms Change, and Why the Ecosystem Is Reacting

TL;DR: Prometheus’s biggest structural weakness has always been cardinality. A stable feature years in the making is finally addressing it, and the rest of the observability market is already responding. Native histograms allow for more efficient metrics storage, reducing cardinality strain and enabling faster, more cost-effective AI-powered observability. Ready to see how AI-powered observability can simplify your monitoring? Book a demo of the Open 360 platform.

What Is Alert Fatigue, and Why Do IT Teams Miss Critical Alerts?

Alert fatigue is one of the biggest reasons critical incidents get missed. In this video, learn what alert fatigue is, why it happens, and how reducing noisy, repetitive notifications helps IT teams respond faster to the alerts that actually matter. Whether you're an IT operations professional, SRE, DevOps engineer, NOC analyst, or IT manager, this video explains alert fatigue in simple terms and shares practical ways to reduce alert noise, prioritize critical issues, and improve incident response.

Top 12 Network Monitoring Tools in 2026: Complete Comparison & Reviews

Modern infrastructure is no longer a stack of routers, switches, and racks sitting in a single data center. Most teams now run a mix of Kubernetes clusters, virtual machines, managed cloud services, and SaaS dependencies spread across regions and providers. Knowing which device is up is not the same as knowing whether your application is healthy.

Coralogix | Magic Quadrant 2026

We are absolutely thrilled to share with you all that Coralogix has been recognized as a Leader in the Gartner Magic Quadrant for Observability Platforms. When we architected Coralogix around in-stream processing, open-format storage, and index-free query, we weren’t optimizing for where observability stood at the time. We were building for the world it was heading toward.

5 Things to Know About Context Engineering

Software systems are getting better at understanding themselves. The mix of richer telemetry, smarter pipelines, and agentic AI is shifting observability from a passive record of events into something more active and useful. That shift is what we mean by context engineering. We recently partnered with O’Reilly on a report by David Beale that introduces the discipline. Before you read it, here are five things worth knowing.

July 2026 at Bindplane: A new pipeline editor, friendlier pricing, and a Blueprints library

The new Advanced Pipeline Editor went live for all paid plans, we launched a public Blueprints library of ready-made pipeline patterns, we reworked plan pricing and raised the Free tier to 100 GB/day, and the source and processor catalog grew with an AWS Neuron source, an AWS CloudWatch metrics source, and a full set of XML processors. None of these are flashy on their own.

9 Best Log Aggregation Tools for 2026

Every on-call engineer has lost an evening to some version of this. Something breaks; the fix is usually somewhere in the logs, and the logs are scattered everywhere: a dozen servers, a few containers, a couple of cloud services, none of them in one place. So, you SSH into one box, grep, get nothing, move to the next, and an hour later there are fifteen terminal tabs open and still no clear sequence of events. Log aggregation tools kill that scramble.

ActiveMQ Log Analysis & Diagnostics: The Expert Guide

Senior engineers who are fast at diagnosing ActiveMQ incidents share one trait: they know exactly what they are looking for in the broker log before they open it. They know the PFC signature, the OOM warning pattern, the journal recovery sequence, and the connection drop format. For them, the log is not text to search through, it is a structured operational record that maps each entry to a specific broker state.

The Advanced Pipeline Editor Is Here: One View, Every Pipeline

The Advanced Pipeline Editor is now live for all paid Bindplane plans. It's a rebuilt configuration editing experience that puts your whole config in a single interactive graph: every source, processor, router, and destination, across logs, metrics, and traces, in one view you can search, pan, zoom, and edit directly. If you've ever bounced between pipeline tabs trying to figure out where a processor sits in a config with a dozen sources and three destinations, this release is for you.

An SRE agent for production

AI has changed how software gets built. It hasn't changed how software gets run. Most of the AI money in software has gone into the IDE: code generation, copilots, developer assistants, faster pull requests. That work matters. But writing software is one slice of the lifecycle. The harder problem, and the more expensive one, is running that software in production. Production is where systems fail in ways nobody predicted. Incidents don't stay inside one service.

We built an SRE bot on AURA. Here's what we learned.

PagerDuty fires. You open the incident. Title, timestamp, nothing else. Whatever context exists is in someone's head, in a Slack thread from two weeks ago, or in a runbook nobody has touched since the last reorg. We got tired of that. So we put an AURA agent behind a Slack bot and pointed it at our own production environment.

Life after SaaS: Enabling the System of Context

By: Tucker Callaway, CEO at Mezmo The market keeps saying “SaaS is dead.” That’s probably true, but it’s also incomplete. What’s actually dying is the idea that value lives inside a vendor-controlled black box. The next era is about utilities: unlimited coding capacity and unlimited analytical capability. And if those two utilities are real, then the vendor model has to change.

A new way to SIEM

For years, security teams have been sold the same bargain: send in more data, buy more tools, tune more rules, and you'll be better protected. In practice, a lot of teams have ended up with the opposite. They're carrying more cost and more complexity, and they still don't have much confidence that their detections are actually working the way they should. That's the backdrop for why Cribl is acquiring CardinalOps.

How Does a Configuration Item Fit Into Your CMDB?

In this video, you'll learn what a Configuration Item (CI) is, how it forms the foundation of a CMDB, and why connecting CIs helps IT teams understand dependencies, improve visibility, and resolve incidents faster. Discover how CIs transform scattered asset data into a complete, connected view of your IT environment. Whether you're an IT Manager, IT Administrator, Service Desk Analyst, ITSM Professional, Infrastructure Engineer, or IT Operations Leader, this video explains why Configuration Items are essential for effective IT Service Management.

Upgraded Alert AI Analysis: Automated Incident Investigation

TL;DR: OrionIQ has launched the next generation of its Alert AI Analysis agent within the Open 360 AI platform, designed to automate and accelerate incident investigation. Key features of this evolution include: Agent-Based Investigation: Instead of relying on a single prompt, the system coordinates specialized AI agents to correlate data across diverse sources like logs, metrics, deployments, and tickets.

What Is Synthetic Monitoring and Why Does It Matter?

A website can look healthy on your dashboard and still fail when customers try to use it. So how do you catch problems before anyone notices them? In this video, you'll learn what synthetic monitoring is, how it works, and why IT teams use it to detect website and application issues before they impact real users. Discover how automated user journeys help you monitor availability, performance, and critical business transactions 24/7.

Part II: Inside Alert AI Analysis: From a Single-Agent Prompt to an Agent Harness

TL;DR: This is the engineering companion to our announcement post, Upgraded Alert AI Analysis: Automated Incident Investigation, read that one for what the new generation does for your team; read on for how it works under the hood. Interested in hearing more? Book a demo to see the Alert AI Analysis Agent live. Root cause analysis is one of the harshest tests you can give an AI.

Unified Logs, Traces, and Errors: Why One Tool Beats Three

Last updated: July 2026 Your Rails app throws a 500. You open Sentry and find the exception. The stack trace points to a controller action, but it does not tell you why the database call failed. You switch to Datadog and search for the request trace. The trace shows a 3-second query, but you do not know what the application was logging at that moment. You open your log aggregator, paste in the request ID, and scroll through output until you find the slow query log line that explains the lock contention.

When and what should I be logging?

This is a follow-up to Sergiy’s post Errors, traces, logs, metrics: when to reach for what. Modern observability platforms, like Sentry, give developers a lot of choice. For a given problem, should you use traces, profiles, metrics, logs? If you take away one thing from this post, I hope it’s this: when in doubt, start by adding a few targeted log lines.

Build an SRE Agent Harness for AIOps Without Context Blowout

An agent harness for AIOps is the runtime layer that coding agents like Claude Code were never built to provide: context isolation, decision traceability, and gated execution for tools that touch production. Aura is Mezmo's open-source (Apache 2.0) agent harness, purpose-built for operations work rather than software development.

Claude Code Monitoring at Scale: Gateways and Routing With OpenTelemetry

Chelsea and I recently wrote a guide on how we monitor Claude Code usage internally with Bindplane. TLDR; We remotely manage a Bindplane Distribution of the OpenTelemetry Collector (BDOT) that runs on every engineer's laptop. This setup is great, but it has one downside. Sending to Google Cloud Monitoring, Swarmia, and any other destination directly from an engineer’s laptop is limited to local processing. You can’t get the benefit of centralized routing and processing on a gateway.

When Does a Self-Service Portal Actually Reduce Tickets?

A self-service portal is designed to reduce IT support tickets by enabling employees to solve common issues on their own. But if self-service is supposed to improve efficiency, why do so many portals remain unused while help desk queues continue to grow? In this video, you'll learn what a self-service portal is, why many organizations struggle with low adoption, and the three key factors that determine whether your portal actually reduces ticket volume.

The future of governing AI agents

How to build governance into autonomous security agents from the architecture up The industry has moved fast on capabilities. Agents now triage alerts, investigate endpoints, create detection rules, and enrich indicators, and they are even capable of performing most actions we as security operators can perform. The architecture patterns are maturing, as are the models, but governance is not keeping pace.

What Is Packet Loss? Causes, Symptoms & How to Fix It

In this video, learn what packet loss is, why it happens, and how it silently impacts your network performance even when monitoring dashboards appear healthy. Discover the most common causes of packet loss, how it affects applications like video calls and web services, and why identifying the root cause quickly is critical for maintaining a reliable network.

Called it (mostly): Checking in on 2026 predictions so far

On this episode of Masters of Data, we revisit the predictions Adam White, Zoe Hawkins, and David Girvin made at the end of last year, checking our own scorecard halfway through 2026. The hits: agents running amok and deleting databases, MCP becoming the backbone for tracking what agents actually do, growing security gaps around personal data, and a collective rejection of low-quality AI content. The misses: we underestimated how fast companies would cut staff for AI, then quietly start rehiring once the agents couldn't cover the work, and we're still arguing about whether token burn is a cost problem or a coming attack vector.

MCP vs CLI: Does it even make a difference? | Live Laugh Logs ep. 3

MCP vs CLI: does it even make a difference? Here’s everything you need to know. Welcome to Episode 3 of Live Laugh Logs, the podcast from the Coralogix Developer Relations team. This week Andre has made the move to the US, so Annie and Lewis are joined by George Pickers, Head of Solution Engineering for EMEA & APAC at Coralogix.

Q&A: How Elastic and Anyshift are bringing AI-powered context to incident response

Incident response often depends on connecting two kinds of context: what changed in the environment and what the logs say happened next. Through a new integration with Elastic, Anyshift’s AI agent, Annie, can read from a customer’s Elasticsearch deployment to search logs, surface error and warning spikes, and correlate log evidence with infrastructure change history.

SLA vs SLO vs SLI Explained: What Should You Track?

In this video, learn the difference between SLA, SLO, and SLI and why understanding each one is essential for delivering reliable IT services. Discover how these three service level metrics work together and why tracking the right one helps improve service reliability, customer satisfaction, and operational performance. Whether you're an IT operations professional, SRE, DevOps engineer, or service manager, this video explains SLA, SLO, and SLI in simple terms so you can build measurable goals and realistic service commitments.

Tech Talk: Observability Simplified, APM and Network Behavior

Participants are welcomed to a session titled "Observability Simplified," focusing on user experience, application performance, and network behavior. This second part of a three-part series highlights how the Splunk Observability Cloud and Cisco ThousandEyes can create a unified view of applications, infrastructure, and network performance. Key discussions include addressing siloed troubleshooting, enhancing visibility, and a live demo showcasing how to identify network issues affecting application performance. Attendees are encouraged to participate in the Q&A and are reminded that the session will be recorded for future reference.

What Is NetFlow, and How Does It Reveal Where Traffic Goes?

In this video, learn what NetFlow is and why it's one of the most effective technologies for understanding network traffic. Discover how NetFlow goes beyond basic bandwidth monitoring by showing who is using your network, what applications are consuming bandwidth, and how traffic patterns change over time. Whether you're a network administrator, IT operations engineer, or infrastructure manager, this video explains NetFlow in simple terms and shows how it helps identify bandwidth hogs, troubleshoot slow networks, and make smarter capacity planning decisions.

From Alert Noise to Automated Action: The Case for Workflow-Driven Monitoring

TL;DR: Modern monitoring platforms face a “workflow problem”: engineers are drowning in telemetry but lack tools that connect detection to resolution, often leading to fragmented, manual incident investigations. Most organizations have mastered data collection but fail at incident response. Engineers waste precious time manually stitching together logs, metrics, and traces across siloed tools. The Solution: Workflow-driven monitoring acts as a guide, not just a dashboard.

A Four-Step Blueprint for Faster Root Cause Analysis: A Logz.io Webinar

Incident investigations take so long not because the fix is hard, but because finding the right fix is. Most engineers spend 20 to 60 minutes just understanding what’s wrong before they can act, not fixing anything, just trying to see the full picture. The framework that changes this has four steps: Orient, Isolate, Hypothesize, and Verify, and the order matters more than the tools.