Skip to content

Measure cost per result that actually worked

Four counting rules for cost per checked result, worked through one made-up example, plus what to record beside every metric.

By Daniel Ternyak3 min readSources checked September 28, 2026

Twenty-four eligible requests came in. Four were declined before work started. Of the twenty attempts, twelve finished cleanly, four needed rework, two failed, and two were abandoned. Each number answers a different question about whether the agent is useful.

The question I want a measurement to answer is narrower: what did it take to produce a result someone could actually use? The answer includes the person checking the result, the failed attempt that ate an afternoon, and the correction found after the agent declared success. The numbers here are a made-up example that shows accounting choices, not any product’s performance.

Four counting rules

1. Count every eligible request before looking at results. Otherwise an agent can look better by attempting easier work, declining hard work, or losing the people who can’t get it to finish. In the example, 20 of 24 requests started, and 16 of those 20 attempts produced a verified output. Only 16 of the 24 eligible requests got a verified result from the agent.

2. Count a result only when the result is checked against a definition of done. For a record update, done might mean the requested fields changed on the intended records and nothing outside the scope changed. A polite final answer isn’t a completed account change. Store the evidence: the changed object, a receipt, or a person’s review against a rubric. When the evidence is missing, keep the result unknown.

3. Charge every attempt, including the failures. Cost per verified result puts the cost of all attempts on top and the verified results underneath. In the example, all twenty attempts took 217.5 minutes of active effort. Counting only completed jobs shows 150.5 minutes: that view quietly drops 67 minutes spent on attempts that didn’t succeed. With zero verified results, cost per result is undefined; show the total spend next to the zero instead.

4. Compare the same results, and keep unknowns unknown. Compare the agent’s all-attempt cost with the manual cost of the same sixteen verified outputs. Comparing twenty manual jobs with sixteen agent ones credits the agent for work the agent didn’t finish. Three timing values are missing in the example, and missing review time doesn’t mean nobody reviewed. Keep a measured zero, an unknown, and an estimate as three different things, and label any estimate that fills a gap.

QuestionCalculationWhat the number leaves open
How often did eligible work start?20 / 24 = 83%Whether declining was the right call
How often did started work produce a verified output?16 / 20 = 80%Unmet demand among the declined requests
How much eligible demand got a verified output from the agent?16 / 24 = 67%Whether the rest was done another way
How many verified outputs needed a person’s correction?4 / 16 = 25%Setup and review effort on every job

The last row describes the same sixteen outputs a second way. The row doesn’t add four more completions. And “clean” doesn’t mean autonomous: the twelve clean completions still include a person’s setup and review.

Keep the meaning beside the metric

Vendor metric names can hide different definitions. One Intercom page counts Fin’s outcome rates over all conversations, and another counts over Fin-involved conversations only. A configured handoff can also count as an outcome. (Fin AI Agent reporting, Reporting metrics) With made-up numbers, the same 300 outcomes read as 30% of 1,000 conversations or 50% of 600. I would check the actual report, filters, and export before comparing those rates.

Record five things beside every metric: what’s counted, what the count is divided by, what’s excluded, the time window, and the rule for evidence. An approval rate deserves the same treatment. High acceptance can mean good proposals or weak review, so link acceptance to later errors, corrections, and review effort before optimizing for more approvals.

Some rework arrives late. A campaign can be scheduled correctly and still contain a mistake someone notices the next morning. Attach the late correction to the original job. And don’t compare a mature group of jobs with a new release whose results haven’t had the same time to fail.

The pattern#