Block AI Agent Regressions Before They Ship | SAO Pre-Push Eval Gate Demo

Oct 3, 2026

Every engineering team has unit tests. They tell you the code still works. They tell you nothing about what the model started saying.

This demo wires a single eval gate script into a git pre-push hook, so Splunk Agent Observability scores every agent's output before the push is allowed through. Luna, an on-premise small language model, runs as a synchronous judge against fixed thresholds. Fail one, and the push is blocked.

In this walkthrough:

  • The app: a Multi-Agent Risk Consensus pipeline where a Growth Analyst and a Risk Analyst debate a ticker and converge on a shared risk score
  • The gate: one script that fetches live market data, invokes both agents, and sends their outputs to Splunk Agent Observability, where Luna scores them synchronously against fixed thresholds
  • A clean baseline run first: context adherence, toxicity, PII and sexism evaluated on both agents, all four checks passing in 12.6 seconds
  • A regression is introduced. The app still runs, the build stays green, and nothing in the diff looks alarming
  • git push fires the pre-push hook and the gate re-runs the identical pipeline
  • The Risk Analyst fails on PII, the script exits non-zero, 1/4 CHECKS FAILED, and the commit never reaches the remote
  • No log-diving and no staging deploy: one eval, one agent, the failing value printed in the terminal
  • The same gate runs locally as a pre-push hook and in GitHub Actions on every pull request

Same evals, same thresholds, binary verdict. How many silent regressions has your team shipped this month?

Try Splunk Agent Observability: https://www.splunk.com/en_us/download/observability-cloud-free-edition.html

Docs: https://agent-observability-docs.splunk.com/what-is-splunk-agent-observability

0:00 Unit tests pass. The agent still leaks PII.

0:14 The app, and the eval gate script wired into a pre-push hook

0:22 Inside the script: market data, both agents, then Luna scores the output

1:00 Clean run: establishing the baseline

1:16 All four checks pass

1:23 A regression is introduced

1:49 git push fires the pre-push hook

2:01 The gate re-runs the identical pipeline

2:12 PII detected: the Risk Analyst fails and the push is blocked

2:30 Nothing to debug: one eval, one agent, the value in the terminal

2:41 Under two minutes, and the same gate runs in GitHub Actions