A finance team pilots an agent for invoice reminders. Engineering sees more successful tool calls. Finance wants to know whether the team spends less effort collecting the same invoices, with fewer mistakes and no extra customer confusion.
The primary unit is a business job with a checked result. Tool calls, tokens, approvals, and notification opens explain the job; those events do not replace the result.
Agree on the result before the pilot
For reminder preparation, done could mean a correct draft for every eligible invoice, using the approved contact and current balance. For sending, done includes the agreed provider or delivery evidence and no duplicate reminder. State whether delivery confirmation is required; otherwise label accepted and delivered separately.
Write exclusions before seeing outcomes. Paid invoices, disputes, unsupported currencies, and missing contacts may be outside the first release. Keep those cases visible as unmet demand rather than quietly dropping the cases from the report.
Finance owns the definition of useful work. Engineering owns reliable event collection. Product owns the cohort and release interpretation. One owner should reconcile the weekly numbers across all three teams.
Keep a small scorecard with explicit counting rules
| Measure | Calculation | Decision the measure supports |
|---|---|---|
| Coverage | Eligible jobs with checked agent results / all eligible jobs | Whether the release supports enough real demand |
| Quality | Reviewed outputs meeting the rubric / reviewed outputs | Whether completed work is usable |
| Rework | Jobs needing correction / checked jobs | Whether apparent success creates hidden effort |
| Active effort | Setup + review + intervention + correction minutes | Whether people actually spend less time |
| Full cost | Model, tools, platform allocation, human effort, and maintenance | Whether the operation is economically useful |
| Cost per checked result | Full attempt cost / checked results | Whether failures consume the expected value |
| Safety and control | Unauthorized effects, wrong recipients, duplicates, stale writes | Whether rollout should continue or pause |
Show unknown timing and unknown outcomes as separate counts. Report quality sampling coverage beside the quality rate. A small reviewed sample does not automatically establish the quality of every completed job.
The measurement guide works through the counting rules. The back-office calculator includes exceptions, retries, monitoring, and human cleanup in the cost model.
Connect product events to the same job
The following event envelope is an illustrative application schema, not a vendor API. Extend the fields for the specific job, and keep raw customer content in access-controlled product systems.
{
"event": "agent.result_checked",
"eventId": "event_902",
"occurredAt": "2026-10-02T17:02:00Z",
"jobId": "job_184",
"tenantId": "workspace_A",
"jobType": "billing.reminder",
"cohort": "pilot_reviewed",
"capabilityVersion": "2",
"skillVersion": "4",
"traceId": "trace_184",
"result": "passed",
"evidenceId": "receipt_184",
"checkMethod": "finance_review",
"activeHumanSeconds": 95
}Use separate events for eligibility, preparation, approval requested, approved or declined, execution, result checked, correction, and handoff. Record provider cost and tool cost against the same job. Deduplicate event IDs in the analytics pipeline.
An elapsed ten-minute wait for review is not ten minutes of active human effort. Capture both elapsed time and active effort. When active effort is estimated, record the method and keep the estimate distinct from measured time.
Choose a comparison that matches the decision
A randomized eligible cohort can compare the agent-assisted workflow with the existing manual workflow when assignment is practical. Keep access, training, and result criteria comparable, then report differences by job type and complexity.
If random assignment is impractical, use a matched comparison or staged rollout. Explain the limitations: early users may be more motivated, easier jobs may be selected, and the underlying work may change over time. Historical averages alone do not establish that the agent caused the improvement.
Give both groups the same time for late errors to appear. A newly completed agent job and a month-old manual job have different opportunities to show corrections. Join later corrections back to the original job and release.
Use approval rate as a diagnostic
Track approval requested, accepted, declined, edited, expired, and ignored. Break the rates down by capability and consequence.
High acceptance can reflect accurate proposals, hidden details, habituation, or pressure to clear a queue. Compare acceptance with later corrections, wrong selections, and review effort before optimizing the rate.
A declining reviewer should be able to record a useful reason: wrong recipient, unnecessary work, stale data, poor message, or unclear consequence. Those reasons become product research and capability priorities.
Separate capacity from cash savings
Freed hours create capacity. Cash savings require a change in spending or avoided spending. Report the measured hours first, then state how the operating team used the capacity.
An illustrative calculation: 100 comparable checked jobs take five minutes each manually and two minutes each with the agent, including review and correction. The difference is 300 minutes, or five hours. Failed agent attempts, monitoring, and setup outside those jobs still need to be charged before claiming net savings. The numbers are made up to explain the method.
Revenue effects need their own measurement. Faster reminders do not by themselves prove higher collections. Compare payment outcomes under an appropriate design, and keep the operational time result separate.
Turn the weekly review into decisions
Review the scorecard with representative successful, declined, failed, and corrected jobs. Assign an owner and saved evaluation case to each material failure. Define rollout limits before launching, including any consequence that requires an immediate pause.
Use product analytics for cohort and funnel events, a warehouse for joining outcomes and costs, and a tracing tool for inspecting individual runs. Tracing vendors and setup covers LangSmith, Langfuse, Braintrust, and Phoenix. Avoid making the trace dashboard the sole business report; the authoritative invoice result lives in the product.
A useful pilot decision says which jobs passed, how much effort changed, which costs remain uncertain, and whether the next release should improve quality, expand coverage, or increase autonomy.