Operations | Monitoring | ITSM | DevOps | Cloud

EVOLVE 2026 Recap: Operational Excellence for the AI-Native SDLC

Last week we hosted our annual conference, EVOLVE, where global engineering leaders came to talk through the realities of building AI-native SDLCs. When asked to name the biggest friction point in their software lifecycle, attendees gave telling answers: ‘reviewing changes’ took 36 percent, ‘measuring impact’ took 35, and ‘building’ drew zero votes.

AI didn't kill tech debt, it just changed the currency you pay it in.

Ganesh Datta on why every team still has a finite budget, now it's tokens instead of headcount. $500 to spend: ship the feature or fix the P2? The orgs building a framework for that call now will have it a lot easier when the CFO puts a cap on spend. From Braintrust by Cortex. Full episode out next Thursday.

You don't know what your model is going to do when you tell it what to do.

Give a model permission to act on your computer and you're trusting it will behave the way you expect. It might not. In this clip from our Braintrust conversation, Adam Berman, engineering leader at Semgrep, breaks down why a backdoored model is a different threat model than backdoored software. You can't fuzz-test your way to finding it, and you often can't detect it until it's already acting the way it shouldn't.

What AI compresses, and What it Amplifies

Adam Berman, VP of Engineering at Semgrep, on the double edge of AI tools for engineering leaders: they compress the distance between an idea and a working prototype, letting him get from exploration to a demoable POC in the gaps between meetings. But that same leverage amplifies risk. One person can spin up 1,000 unowned problems just as fast as they can spin up 1,000 wins. From a Braintrust by Cortex conversation on how AI is changing the job of engineering leadership.

DRIVE vs DX Core 4: What each framework measures and when to use them

Engineering organizations spent the past decade learning to measure developer and team productivity. Measurement of the organization itself did not keep pace. Now that agents are responsible for writing most of the code across many engineering organizations, it’s more important than ever to have an effective methodology for measuring productivity.

KPI cards: build a reliability dashboard that doesn't force tradeoffs

This week's Feature Friday: Principal Product Manager Christine Byun walks through KPI cards, a new way to build custom dashboards in Engineering Intelligence. KPI cards pull key metrics, like change failure rate and rollback frequency, into compact tiles so they stay visible without taking up chart space. That means the metric you're actively working, incidents, in this demo, gets full-size room, without losing sight of the rest of your system.

Garbage in, garbage out: Splunk's Steve Flanders on why AI can't fix your bad telemetry

Cortex co-founder and CTO Ganesh Datta sits down with Steve Flanders, who leads AI transformation at Splunk and wrote the book on OpenTelemetry, to talk about why AI acceleration without strong observability foundations creates more problems than it solves.

DRIVE vs SPACE: What each framework measures and when to use them

When Nicole Forsgren, Margaret-Anne Storey, and their coauthors published "The SPACE of Developer Productivity" in 2021, they settled an argument the industry had been losing for years. Productivity is not one number, and it is not a proxy like commits or story points. It is multidimensional, and any attempt to flatten it into a single metric will mislead you. Most of what came after in developer productivity measurement builds on SPACE. SPACE and DRIVE were built for different jobs.

AI can't correlate what was never standardized

Steve Flanders (Senior Director of Engineering, Splunk) makes the case that AI can't save an observability stack that never agreed on a standard. Mix formats across metrics and logs, and AI stops correlating and starts guessing, which means you either make the wrong call or miss the answer you actually needed. OpenTelemetry is one fix, but Prometheus and Fluentd work too. The standard matters more than which one you pick.