Operations | Monitoring | ITSM | DevOps | Cloud

AMA Recap: More Answers From the Observability Engineering Authors

Last week, we sat down with the authors of Observability Engineering for a live AMA. We ended up getting so many questions (pre-submitted and live) that we couldn't get through them all. Charity, Liz, George, and Austin kindly stuck around afterward to answer more, ranging from low-hanging observability fruits and telemetry to AI and what software engineers can do that Claude can't. Missed the live session? Watch it on demand now.

From Vision to Value: New Splunk Platform Innovations Supporting Cisco Data Fabric Are Generally Available

At.conf25, we announced our vision for Cisco Data Fabric, an architecture designed to help organizations unlock the value of machine data, fuel AI with trusted context, and support more intelligent and resilient operations. Today, that vision has become reality. Key Splunk Platform innovations including Machine Data Lake, Catalog, and Agent Launchpad, together with expanded Federated Search and Data Management capabilities, are now generally available.

Why Config Changes Cause Most Cloud Outages in 2026

If you have watched the incident channels light up over the past few weeks, you already sense the theme of 2026: cloud outages are no longer rare, dramatic once a year events. They are a steady drumbeat, and most of them trace back to the same root cause. Not a data center fire, not a rogue backhoe severing a fiber line, but a routine configuration change that went out, behaved differently than expected, and cascaded.

On-Call in 2026: Preparing for Cascading Failures

The biggest outages of 2026 are not being caused by a single server dying or one bad deploy. They are being caused by cascading failures, where healthy systems interact in ways nobody planned for and take each other down. That shift changes what good on-call looks like. If your incident response still assumes that "something broke" and one team owns the fix, you are going to be slow exactly when speed matters most.

Selecting the Best Enterprise Session Border Controller

A Session Border Controller (SBC) is the gatekeeper for modern real-time communications. It's not just a network device. It is a core component of your security, interoperability, and service quality strategy. An SBC secures the boundary between your enterprise network and the outside world, protecting against threats, enabling connectivity between disparate systems, and managing traffic to deliver a high-quality user experience. Think of it as a firewall optimized for voice communications.

Where Historians Fall Short for Physical AI

Summary Physical AI—machines and industrial systems that sense conditions, reason, and act in the real world—needs two things from operational data: detailed history for training, and real-time telemetry for inference. Traditional data historians weren’t built for either at the speed Physical AI requires. Four gaps result: limited real-time access, compression that strips model-relevant signal, IT/OT fragmentation, and site-by-site architectures.

How to Select the Right PCS for Commercial and Industrial Energy Storage Systems

When companies evaluate commercial and industrial energy storage systems, the battery usually receives most of the attention. Capacity, cycle life, and battery chemistry are often the first specifications discussed. However, the Power Conversion System (PCS) plays an equally important role because it determines how effectively stored energy can be used.