Turn every bad run into a fix someone owns
A bad run becomes an owned ticket and a test case.
When monitoring spots a bad run, the run goes to a person who can fix the part that failed. The team saves the run as a test case and checks the fix on cases the fix wasn’t built from before calling the problem solved.
Deep dive
What a trace contains, four tracing tools worth trying, what to log beyond the model call, and a review routine a small team can start this week.
Why the pattern matters#
A monitor flags a bad run, someone changes the agent, and the flagged example now passes. Passing that example is debugging, not proof. The fix was shaped by the example, so the team needs other cases, including ones where nothing should go wrong.
When to use the pattern#
Use the pattern when
- Automated checks flag problems in live runs.
- A team is deciding whether to ship a change to the agent.
- One AI is grading another AI’s work.
- A team is planning what the agent should do next from failed runs or feature requests.
Skip the pattern when
- An average score would hide a critical failure, like a change the agent wasn’t allowed to make.
- A case used to build the fix would count as an independent test.
- Every permission denial would be treated as a bug to remove. A denial may be the right boundary.
Build for the job users couldn’t finish#
A request for a new tool is a clue, not a spec. Three people ask for calendar rescheduling. One never found the existing action, one lacked permission, and one rejected an approval that didn’t say who’d be affected. Counting all three as demand for a new tool sends the roadmap toward something that already exists.
Read what each attempt shows before choosing what to build. Count distinct customers, not tool calls: more calls can mean deeper work or wasted retries.
| What the attempt shows | Possible explanations | What to test next |
|---|---|---|
| No relevant tool was called | The capability is hidden, unavailable, or not needed | Discovery, and what the person actually wanted |
| The action was denied | The right boundary, or a missing grant | The policy, with the person who owns the policy |
| The tool wasn’t available | A rollout limit, or an unsupported action | Eligible versus ineligible demand |
| The person abandoned the approval | A bad proposal, an unclear consequence, or an interruption | The proposal, and why the person stopped |
| The run finished, then the person redid the work by hand | The result was wrong, incomplete, or hard to trust | The output and the follow-up work |
| Few requests at all | Low demand, or people stopped trying | Non-users and their workarounds |
In real products#
What each product documents, strongest example first.
- Intercom Fin
A Fin Procedure simulation passes or fails against criteria the builder writes, and a QA review can open an issue ticket, with an assignee, straight from a conversation.
Run simulations for Fin Procedures - HubSpot Customer Agent
HubSpot labels each coaching opportunity as a knowledge gap, a knowledge conflict, or an action gap, and shows why the conversation was flagged. Adding help-center text fixes only the first kind.
Analyze the customer agent’s performance - Atlassian Rovo
Rovo’s live conversation review shows managers where users get stuck or drop off.
Review live conversations with a Rovo agent - Agentforce Service
Agent Optimization helps teams dig into unresolved interactions, find knowledge gaps, and analyze agent sessions.
About Agent Optimization
Sources checked September 28, 2026.
Checklist#
Yes-or-no checks for a design review.
- Each flagged run goes to an owner who can change the part that failed.
- The saved case holds the starting state and the expected result, and the check inspects the changed record, not the agent’s words.
- The fix is tested on cases the fix wasn’t built from, including cases where nothing should go wrong.
- A critical failure, like an unauthorized change, blocks a release even when the average score improves.
- An AI grader is checked against human-reviewed cases, including ones the grader didn’t flag.
- A small random sample of ordinary runs is reviewed next to the flagged ones.
Try it on your product
Trace one bad run to its owner
Take one run a customer complained about. Write the expected result. Find the first step that went wrong: the context supplied, the search result, the tool call, or the backend’s response. Open the affected record to confirm what changed. Name who owns that step, and save the run as a test that checks the record, not the agent’s reply.
Before the agent does something that matters, the product shows the person exactly what will change: which items, which fields, who is affected, and what can’t be undone.



