Operations | Monitoring | ITSM | DevOps | Cloud

The latest News and Information on Monitoring for Websites, Applications, APIs, Infrastructure, and other technologies.

Top tips: Find the right answer in a sea of search results

Top tips is a weekly column where we highlight what's trending in the tech world and list practical ways to explore these trends. This week, we're looking at how to search the web more effectively and find the information you need faster. Searching the web can feel a little like playing hide-and-seek. You know what you're looking for is somewhere out there, but the internet has an impressive number of places to hide it.

Free vs. paid website monitoring: When should you upgrade?

Every growing business hits a moment of truth somewhere between "we just launched" and "we just lost a customer because our site was down for 40 minutes." It's then that you realize the free monitoring tool you set up six months ago—the one that felt more than adequate at the time—is no longer pulling its weight.

Building an AI-Powered Time Series Dashboard for Data Center Ops

AI demand is pushing data centers to get bigger, denser, and more distributed, while the data that runs them still often sits in different systems, slowing down operations and hampering efficiency. By leveraging a time series database like InfluxDB to create a single, unified telemetry layer, you can simplify your data stack, comfortably handle the high volume of data, and solve problems faster and more efficiently, helping to minimize waste and maximize efficiency, saving time, electricity, equipment, and money.

Telemetry Talks ep 7 - Beyond OpenTelemetry with anomaly detection

In this episode, we continue to dive into the workshop we hosted at Cloud Native Days Romania in May, together with our guest, Fred Navruzov, correlating OpenTelemetry with anomaly detection. Furthermore we explore how the VictoriaMetrics MCP server and skills bring AI-powered observability to your workflows. Learn how to detect anomalies faster and interact with your metrics, logs and traces using natural language.

High Bandwidth Usage on Firewalls: How to Diagnose & Fix It

If your network has been feeling sluggish, connections keep dropping, or you're getting alerts that don't point to an obvious cause, high bandwidth usage is one of the most common culprits and one of the hardest to pin down without the right visibility. The tricky part is that "slow network" can mean a dozen different things depending on where the congestion is actually happening.

Cribl On Your Coffee Break Episode 14 - Routes, part 3: One-to-Many

Welcome to the 14th installment in our series to help you get started with the Cribl platform. Here, we continue our conversation about Cribl routes and routing techniques By the time the month is over, you will have a pretty good idea of what Cribl can do, and how to do it. You’ll also have consumed more caffeinated beverages than is strictly appropriate...

Digital Signature vs Electronic Signature and When ITSM Approvals Need Each

Would a single click on Approve in your service desk satisfy an auditor reviewing a high-risk change, a purchase order, or a vendor contract? Often the answer only becomes clear when someone requests a signed copy and the ticket has nothing to show. Much of this traces back to terminology. Many organizations treat electronic signatures, digital signatures, and approvals as interchangeable, and ITSM tools tend to call every sign-off an approval.

UPS Monitoring for Data Centers That Cannot Afford an Unplanned Stop

How many minutes of battery runtime are left in the unit protecting your primary rack right now? Most monitoring deployments can answer a question like that for every switch, server and virtual machine in the building. Ask about power and the dashboard goes quiet. Backup power tends to fall between two owners. Facilities buys the hardware and books the service visits, while IT owns everything plugged into it.