Operations | Monitoring | ITSM | DevOps | Cloud

The latest News and Information on Monitoring for Websites, Applications, APIs, Infrastructure, and other technologies.

API update: Better visibility into rate limits

We’ve updated the StatusGator API v3 to help you see how many requests your integration can make and handle rate limits more smoothly. Responses that count toward your organization’s per-minute limit now include three headers: If a request exceeds your limit, the API returns 429 Too Many Requests and a Retry-After header telling your integration how many seconds to wait before trying again.

Debugging a checkout latency spike with Sentry metrics

Checkout suddenly got a lot slower for a chunk of users, and the only clue was a spike in an application metric. In this walkthrough, we go from that spike to a linked trace that shows exactly where the time went. Using Codex alongside Sentry, we dig into the trace data, isolate the slowdown to the European backend, and follow it down to a single Stripe request that's dragging everyone down with it. Along the way, we talk through what actually helps here: caching, timeouts, async calls, and other ways to keep one slow dependency from becoming everyone's problem.

Every 404 in Your Rails App Might Be Allocating 13 MB

A bot requests /wp-login.php on your Rails app. Rails can’t route it, raises ActionController::RoutingError, and returns a 404. That should cost almost nothing. On a Rails 8.1 app with a few thousand compiled templates, it can cost 13 MB of allocations and 27 ms of CPU. A detailed report on rails/rails#58887 traces the cost to one method, ActionDispatch::ExceptionWrapper#build_backtrace.

Tempo 3.1 release: new features for Kafka, TraceQL metrics updates, trace redaction, and more

Building on the major release of Tempo 3.0, Tempo 3.1 is here, delivering community-contributed Kafka client improvements, query-based trace redaction, sampling-aware TraceQL metrics, and more. Together, the updates in 3.1 make it easier to operate Tempo, get accurate insights from your trace data, and investigate issues more efficiently. You can continue reading and check out the video below to learn more about the latest features.

Telemetry Talks ep 8 - Fireside chat with OpenTelemetry maintainers

Telemetry Talks episode 8 is here We sat down with OTel maintainers to talk about the future of the community, GenAI semantic conventions, contributing beyond code, OTel in Practice, and what they’re currently building, writing, organizing, and experimenting with across the CloudNative and OpenSource ecosystem. A conversation about where OTel is today and what comes next. Playlist Resources for Further Learning.

IT Service Continuity Management: How to Build an ITSCM Plan

Most IT teams have a recovery plan somewhere. It was written for a disruption that has not happened yet, and tested less often than anyone admits. The gap rarely sits in the technology. Nobody agreed which services come back first, or how fast. There was time to settle that calmly, and it went unused. IT service continuity management is the ITIL practice that settles those questions in advance. In this blog, you will: By the end you will know what belongs in an ITSCM plan and who has to agree to it.

How IT Infrastructure Management Keeps Services Reliable and Costs Predictable

When a business application slows down, how fast can your organization trace the cause to a server, a network link, storage or a cloud instance? Often it comes down to who's on call that day, since asset, observability and change data are scattered across separate systems. Engineers then check each tool one at a time while customers wait and the cost of the outage grows. IT infrastructure management solves this by keeping asset records, health data and change history in order before an incident starts.

How Etsy gets its mobile apps ready for peak traffic

Most advice about surviving a traffic spike is about capacity. Scale the fleet, warm the caches, load test the checkout path. All of it assumes you can fix whatever breaks the moment you find it. Mobile apps don’t work that way. I spent an hour on a workshop with Jay Henry, a senior engineering manager at Etsy who owns engineering strategy across three teams covering CI, build, test, release, observe, and SRE. Jay’s take: a web team having a bad day can revert in minutes.