Skip to content

Measure cost per result that actually worked

Divide spend by checked results, not by attempts.

Count every request the agent should have handled, charge the agent for every attempt including failures, and compare only checked results against the manual cost of those same results. Keep unknowns marked as unknown.

Why the pattern matters#

An agent can look good when a team counts only successes, drops hard requests, or compares 20 manual jobs with 16 agent ones. In the site’s made-up example, counting only completed jobs quietly drops 67 minutes spent on attempts that didn’t succeed.

Measure cost per result that actually workedIllustration
This monthExample numbers
Requests the agent should handle100
Attempts, at $0.40 each130 · $52
Requests with a result that worked72 of 100
Per attempt$0.40
Per result that worked ($52 ÷ 72)$0.72
Measure cost per result that actually worked. Illustration.
  1. Count every request the agent should have handled: 100 this month.
  2. Charge every attempt, failures and retries included.
  3. Only checked results count: 72 of 100 requests got a result that worked.
  4. Cost per attempt looks cheap. Cost per result that worked is the real number.

When to use the pattern#

Use the pattern when

  • A team is claiming the agent saves time or money.
  • A team is comparing releases, vendors, or agent versus manual work.
  • Someone is reading a vendor’s “resolution” or “deflection” rate.

Skip the pattern when

  • A missing time measurement would be treated as zero.
  • Two vendors’ percentages would sit side by side before anyone checks that both count the same things over the same time window.

In real products#

What each product documents, strongest example first.

Sources checked September 28, 2026.

Checklist#

Yes-or-no checks for a design review.

  • Every eligible request is counted before anyone looks at results, including requests the agent declined.
  • Each request has one ID; retries link to the request instead of counting as new ones.
  • A result counts only when checked against a written definition of done.
  • Failed and abandoned attempts are charged to the results the attempts helped produce.
  • The manual baseline covers exactly the same results the agent delivered.
  • A missing measurement stays marked unknown, never zero.
  • Every rate shows what the rate is divided by, the time window, and what was excluded.

Try it on your product

Define the job in four sentences

Pick one real user job your agent handles. Write four sentences: the starting context, the expected result, the changes the agent may make, and the evidence that proves the job is done. The fourth sentence is what you count.

Deep dive#

Deep dive · 3 min readMeasure cost per result that actually worked

Four counting rules for cost per checked result, worked through one made-up example, plus what to record beside every metric.

Next patternTurn every bad run into a fix someone owns

When monitoring spots a bad run, the run goes to a person who can fix the part that failed.