Operations | Monitoring | ITSM | DevOps | Cloud

The latest News and Information on Monitoring for Websites, Applications, APIs, Infrastructure, and other technologies.

Vulnerability Fatigue: When Discovery Outpaces Remediation Capacity

Recent findings from Anthropic’s Project Glasswing offer a useful indication of where vulnerability discovery may be heading. Anthropic reported that it and its partners had used Claude Mythos Preview to identify more than 10,000 high- or critical-severity vulnerabilities across the software they reviewed. More significantly, Anthropic reported that the bottleneck had shifted from finding vulnerabilities to having the capacity to verify, disclose, and patch them.

Inputs Demystified - Connect Anything with the Input Wizard Webinar

Getting logs into Graylog should not require a PhD in syslog. Part of the Getting the Most out of Graylog Open series, this session educates Open users on the full Inputs framework in Open, what input types are available, when to use each, and how to use the Input Wizard to get new sources connected faster. Open users need to be on the latest version of Graylog. We also cover the revamped Inputs page and how to validate that your data is arriving clean.

Cribl On Your Coffee Break Episode 18 - All About AI

With 3 more days to go, we’ve finally arrived at the AI episode in the Cribl on your coffee break series. Today we’ll touch on a few of the many ways we’ve enabled Cribl to use AI, and also to help you manage the data generated by AI-enabled tools. By the time the month is over, you will have a pretty good idea of what Cribl can do, and how to do it. You’ll also have consumed more caffeinated beverages than is strictly appropriate...

Correlating Business and IT Events: The Path to Business Process Observability

At 9:40 on a Tuesday morning, an order sits unconfirmed in the fulfillment process. Two systems away, a queue depth ticks upward in the integration layer. Both events are recorded. Neither is connected to the other, so nobody escalates nothing is technically down. By 2 PM, order confirmations have stalled across a region. The CFO is asking why the daily revenue number looks soft. Customer service is fielding calls.

Monitor Third Party Services With UptimeRobot.

Your checkout might run on Stripe, your files on AWS, and your images on a CDN. When one of those providers has an incident, part of your product breaks with it, and your own monitors can stay green the whole time. Starting today, UptimeRobot monitors beyond your own infrastructure. Third party monitoring lets you add the services your product depends on, select the components you actually use, and get an alert through your existing channels when their status changes.

Deduplicate logs at the edge: Same insights, a fraction of the volume

Ask a platform team why their observability bill keeps growing and you'll often get a one-sentence answer: And that's usually where it ends. The application teams own the log output, the platform team owns the bill, and nobody has the leverage to change what gets emitted. A single retry loop can print the same error thousands of times a minute. Every one of those lines is ingested, indexed, and stored. You pay for all of them, and they tell you exactly one thing: this error happened, a lot.

Status Pages: Publish Post-Mortems on Your Incidents

Status pages now have a place for the last step of an incident: the post-mortem. Once an incident is resolved, you can write what happened, why it happened, and what you are changing, then publish it on the incident itself. Until now, the updates you posted during an outage ended with "Resolved", and the explanation lived somewhere else: a blog post, a PDF sent to a few customers, or an email thread. Customers who read the incident on your status page never saw it.

Configure RUM SDKs remotely from Datadog

Datadog Real User Monitoring (RUM) SDK settings live in your application code, so changing how the SDK collects RUM data has traditionally required shipping a new application version. These configuration changes can include adjusting sampling rates, enabling Session Replay, or changing which events the SDK collects. For mobile teams, this means that updates often sit in app store review for days or weeks before users start adopting the new version. Full user adoption can take weeks or months longer.