Measure cost per result that actually worked
Divide spend by checked results, not by attempts.
Count every request the agent should have handled, charge the agent for every attempt including failures, and compare only checked results against the manual cost of those same results. Keep unknowns marked as unknown.
Why the pattern matters#
An agent can look good when a team counts only successes, drops hard requests, or compares 20 manual jobs with 16 agent ones. In the site’s made-up example, counting only completed jobs quietly drops 67 minutes spent on attempts that didn’t succeed.
- Count every request the agent should have handled: 100 this month.
- Charge every attempt, failures and retries included.
- Only checked results count: 72 of 100 requests got a result that worked.
- Cost per attempt looks cheap. Cost per result that worked is the real number.
When to use the pattern#
Use the pattern when
- A team is claiming the agent saves time or money.
- A team is comparing releases, vendors, or agent versus manual work.
- Someone is reading a vendor’s “resolution” or “deflection” rate.
Skip the pattern when
- A missing time measurement would be treated as zero.
- Two vendors’ percentages would sit side by side before anyone checks that both count the same things over the same time window.
In real products#
What each product documents, strongest example first.
- Linear Agent
Linear bills AI credits for failed runs, retries, and partial completions, not only for successes. Cost per useful result has to count all of them.
AI credits - Intercom Fin
One Intercom page counts the outcome rate over all conversations, another over Fin-involved conversations. With made-up numbers, the same 300 outcomes read as 30% or 50%.
Fin AI Agent reporting - HubSpot Customer Agent
HubSpot sets a conversation’s resolution status 72 hours after the visitor’s last reply, and says deflections don’t always mean resolution.
Understand the customer agent - Agentforce Service
Salesforce now has a model judge deflection, and fixed a bug that counted some sessions as deflected.
LLM-based deflection and abandonment metrics in Agent Analytics
Sources checked September 28, 2026.
Checklist#
Yes-or-no checks for a design review.
- Every eligible request is counted before anyone looks at results, including requests the agent declined.
- Each request has one ID; retries link to the request instead of counting as new ones.
- A result counts only when checked against a written definition of done.
- Failed and abandoned attempts are charged to the results the attempts helped produce.
- The manual baseline covers exactly the same results the agent delivered.
- A missing measurement stays marked unknown, never zero.
- Every rate shows what the rate is divided by, the time window, and what was excluded.
Try it on your product
Define the job in four sentences
Pick one real user job your agent handles. Write four sentences: the starting context, the expected result, the changes the agent may make, and the evidence that proves the job is done. The fourth sentence is what you count.
Deep dive#
Four counting rules for cost per checked result, worked through one made-up example, plus what to record beside every metric.
When monitoring spots a bad run, the run goes to a person who can fix the part that failed.



