Operations | Monitoring | ITSM | DevOps | Cloud

How to Send Critical Alerts to the OnPage App | 3 Ways

What are the different ways to send a critical alert to the OnPage app? This video shows three ways people and external systems can trigger a high-priority OnPage mobile alert that's also HIPAA compliant (secure or healthcare use cases): These options are particularly useful for users with OnPage mobile licenses who do not have a Silver or Gold plan.

Best LLM inference providers 2026: 16+ on cost per outcome

An LLM inference provider hosts open-weight models like Llama, DeepSeek, and Qwen behind a pay-per-token API, handling GPUs, scaling, and serving for you. The same Llama 3.3 70B model ranges from $0.10 to $1.04 per million input tokens depending on who serves it, so provider choice is a pricing decision. Top picks as of September 2026: Groq and Cerebras for speed, DeepInfra for price, Together and Fireworks for breadth, Baseten for custom models.

Distributed Tracing Is Now in Beta for Ruby, PHP, and Python

A request comes in, enqueues a job, and returns. Twenty seconds later the job runs, and it’s slow. You have a trace of the request and a trace of the job, and nothing joining them. Time Detective has always helped you reconstruct what happened. Now we join it up for you, across applications, services, background jobs and infrastructure, even when they’re built in different languages.

Self-Improving Agents: A Practical Guide to Continuous Learning

We build agents to take work off engineers’ plates. Then we give those engineers a new manual job: reading failed runs and babysitting prompts. Agents will improve themselves automatically. We’re not there yet, but this is the future I’m betting on. We’ve been working on this ourselves at Komodor over the past year. We know how hard it is to turn a failure into an improvement that holds up beyond a few examples.

What Is a Network Topology Diagram? Types, Examples and How to Build One That Stays Current

Most network diagrams are accurate exactly once: the day they are finished. The network keeps changing, the drawing does not, and the gap shows up during the next outage. According to the Uptime Institute Annual Outage Analysis 2026, failure to follow established procedures remains the leading driver of human-error outages. A wrong diagram is how a right procedure hits the wrong port. The fix is a network topology diagram that matches the live network topology.

Incident Management System: What It Is and How to Choose One

An alert fires. A ticket opens. Someone gets paged. Then the real work begins: gathering context, finding the affected service, deciding who owns the issue, running diagnostics, applying a fix, validating recovery, and documenting the result. Many IT teams assume that an incident management system is simply the application that opens and tracks the ticket. That is part of the job, but it’s not the whole operating model.

When AI Agents Attacked Their Own Evaluators, the Industry's Own Leaders Started Asking for Guardrails

When AI agents attacked their own evaluators in July 2026, it exposed a gap no policy commitment can close. The OpenAI Hugging Face incident revealed that enterprise agent governance requires in-flow runtime controls, not retrospective auditing or industry safety agreements.

ilert now supports a native Bleemeo integration

Bleemeo monitoring now connects natively to ilert, linking threshold detection to on-call management and alerting. DevOps, SRE, and IT operations teams get a direct path from a breached threshold to the phone of the engineer who can fix it, and back to a clean slate once the problem is gone.

Run Your GitHub Actions Workflows on CircleCI (Open Preview)

CircleCI can now run supported GitHub Actions workflows directly on CircleCI infrastructure, using the YAML you already have. In this quick demo, we’ll walk through setting up a CircleCI project with an existing GitHub Actions workflow and running your first build. You’ll see how to: GitHub Actions compatibility is currently available in open preview on Linux. Not all GitHub Actions features are supported yet, so check the documentation for current compatibility.

Bleemeo and ilert: two European companies, one alerting chain

Some alerts only need to reach a Slack channel. Some need to reach one specific person, at 3am, and keep trying until they answer. For the second kind, we are partnering with ilert — an incident response platform covering the full lifecycle: from the moment an alert arrives, through paging the right responder, coordinating the response, telling customers what’s happening, and learning from it afterwards. The integration is live today, on both sides.