Turn Production Failures Into Test Datasets | SAO Dataset Curation
Your agent breaks in production. You fix it and move on. But the input that actually broke it is gone and two months later the same failure quietly comes back, because there was never anything to test against.
That's not a debugging problem. It's a missing dataset.
This demo turns low-scoring production traces into a regression suite you can run against every prompt and model change, without writing a single test case by hand.
In this walkthrough:
- A live logstream from a production agent, with metrics scoring every trace
- One filter for context adherence below 70%, surfacing 67 struggling traces
- Creating a new dataset, by selecting the traces worth keeping
- The Dataset Store, where every failure, edge case and response worth protecting lives
- A regression suite built entirely from real production traffic, growing every time someone debugs
Agents don't stay still. New prompts, new models, and with every change something can quietly break. Curated datasets are how teams stay ahead of it.
Turn every production failure into a permanent regression suite with Agent Observability.
Try Splunk Agent Observability: https://www.splunk.com/en_us/download/observability-cloud-free-edition.html
Docs: https://agent-observability-docs.splunk.com/what-is-splunk-agent-observability
0:00 The input that broke your agent is already gone
0:22 A live logstream with metrics on every trace
0:30 One filter: every trace below 70% context adherence
0:45 Inside the trace: the agent's price vs the tool's
1:02 Select the traces that matter, Copy to Dataset
1:08 Create a new dataset and name it
1:20 The Dataset Store: every failure worth protecting
1:28 Datasets that grow every time you debug