Skip to content

Turn every bad run into a fix someone owns

A bad run becomes an owned ticket and a test case.

When monitoring spots a bad run, the run goes to a person who can fix the part that failed. The team saves the run as a test case and checks the fix on cases the fix wasn’t built from before calling the problem solved.

Deep dive

Deep dive · 7 min readAgent tracing in practice

What a trace contains, four tracing tools worth trying, what to log beyond the model call, and a review routine a small team can start this week.

Why the pattern matters#

A monitor flags a bad run, someone changes the agent, and the flagged example now passes. Passing that example is debugging, not proof. The fix was shaped by the example, so the team needs other cases, including ones where nothing should go wrong.

When to use the pattern#

Use the pattern when

  • Automated checks flag problems in live runs.
  • A team is deciding whether to ship a change to the agent.
  • One AI is grading another AI’s work.
  • A team is planning what the agent should do next from failed runs or feature requests.

Skip the pattern when

  • An average score would hide a critical failure, like a change the agent wasn’t allowed to make.
  • A case used to build the fix would count as an independent test.
  • Every permission denial would be treated as a bug to remove. A denial may be the right boundary.

Build for the job users couldn’t finish#

A request for a new tool is a clue, not a spec. Three people ask for calendar rescheduling. One never found the existing action, one lacked permission, and one rejected an approval that didn’t say who’d be affected. Counting all three as demand for a new tool sends the roadmap toward something that already exists.

Read what each attempt shows before choosing what to build. Count distinct customers, not tool calls: more calls can mean deeper work or wasted retries.

What the attempt showsPossible explanationsWhat to test next
No relevant tool was calledThe capability is hidden, unavailable, or not neededDiscovery, and what the person actually wanted
The action was deniedThe right boundary, or a missing grantThe policy, with the person who owns the policy
The tool wasn’t availableA rollout limit, or an unsupported actionEligible versus ineligible demand
The person abandoned the approvalA bad proposal, an unclear consequence, or an interruptionThe proposal, and why the person stopped
The run finished, then the person redid the work by handThe result was wrong, incomplete, or hard to trustThe output and the follow-up work
Few requests at allLow demand, or people stopped tryingNon-users and their workarounds

In real products#

What each product documents, strongest example first.

Sources checked September 28, 2026.

Checklist#

Yes-or-no checks for a design review.

  • Each flagged run goes to an owner who can change the part that failed.
  • The saved case holds the starting state and the expected result, and the check inspects the changed record, not the agent’s words.
  • The fix is tested on cases the fix wasn’t built from, including cases where nothing should go wrong.
  • A critical failure, like an unauthorized change, blocks a release even when the average score improves.
  • An AI grader is checked against human-reviewed cases, including ones the grader didn’t flag.
  • A small random sample of ordinary runs is reviewed next to the flagged ones.

Try it on your product

Trace one bad run to its owner

Take one run a customer complained about. Write the expected result. Find the first step that went wrong: the context supplied, the search result, the tool call, or the backend’s response. Open the affected record to confirm what changed. Name who owns that step, and save the run as a test that checks the record, not the agent’s reply.

Next patternShow the exact change before it happens

Before the agent does something that matters, the product shows the person exactly what will change: which items, which fields, who is affected, and what can’t be undone.