A team planning a finance agent has dozens of possible tools: invoice search, contact updates, reminders, refunds, account notes, and escalation. Historical cases can show which tools unlock useful jobs before the team builds every endpoint.
Label the result people wanted and the step that blocked completion. A request for “follow-up help” does not automatically justify a sending tool. The person might need a correct contact, a disputed balance explained, or a reminder drafted.
Gather cases from more than conversations
Use support tickets, operator notes, product events, and agent traces. Add the final record state and any later correction. A tool returning successfully establishes execution, not that the customer’s problem ended.
Start with a bounded period and preserve the selection rule. Sample by customer, job type, and consequence. Include jobs outside the current agent’s capability and jobs where a person never tried the agent. Otherwise the analysis overrepresents work the current product already supports.
Remove credentials and unnecessary personal data before copying cases into a research tool. Keep stable internal case IDs so a reviewer can check the source in the authorized system.
Use a labeling sheet a reviewer can audit
| Field | Example value | Why the field matters |
|---|---|---|
| Case ID and date | CASE-184, September 12 | Find the original case and assign a time period |
| Customer or segment | Small business, workspace A | Count breadth rather than repeated requests from one account |
| Requested result | Prepare a reminder for an unpaid invoice | Group jobs by outcome |
| Starting objects | Invoice and customer IDs | Identify context the agent needs |
| Human steps | Check dispute, confirm contact, draft message | Separate reasoning from required actions |
| Completion evidence | Reviewed draft attached to the invoice | Define a result someone can verify |
| Blocker | No contact-role lookup | Find the missing capability |
| Consequence | External email if approved | Set review and permission requirements |
| Active effort | Measured minutes, or unknown | Estimate value without treating gaps as zero |
| Later correction | Recipient replaced before sending | Preserve quality costs |
Keep several labels for one case when the job needs several capabilities. Count the case once when estimating job coverage. Ten tool calls in one difficult case are not ten customers asking for the tool.
Separate tool gaps from other failures
A missing action, a bad selection, weak context, and a confusing review screen need different fixes.
| Observed case | Likely next investigation |
|---|---|
| The agent cannot read dispute status | Add or repair a permitted data lookup |
| The agent reads dispute status but ignores the dispute | Improve the skill and add a saved evaluation case |
| The recipient tool returns an old billing contact | Fix source freshness and contact-role semantics |
| Finance declines correct proposals because the card hides recipients | Improve the review card before optimizing acceptance |
| A reminder sends but no receipt is available | Fix execution reporting and history |
| One customer requests unusual payment negotiations | Investigate the segment before generalizing the capability |
The original Fin session provides a smaller example. Fin returned a working numbered citation and a written URL that opened a 404. Adding more knowledge would not, by itself, establish that every generated link is usable. The observed gap suggests checking all displayed destinations.
Turn labeled cases into a capability backlog
The table below is an illustration with made-up counts, not a benchmark. Assume 100 sampled finance jobs.
| Candidate | Relevant jobs | Additional checks | Decision |
|---|---|---|---|
| Explain invoice status | 45 | Source fields, citation accuracy, access | Good read-only starting point if the cases pass |
| Find the correct billing contact | 30 | Role, freshness, workspace boundary | Enables drafting and safer sending |
| Prepare reminder drafts | 25 | Dispute exclusion, message correctness | Pilot with finance review |
| Send approved reminders | 18 | Recipient approval, duplicate prevention, delivery receipt | Add after drafts pass review |
| Issue refunds | 4 | Policy, amount limits, accounting effects | Separate job and permission model |
Relevant jobs overlap. Adding the rows does not produce 122 independent opportunities. Estimate the union of cases supported by a complete tool set, then inspect remaining blockers.
An expected-hours estimate can use job volume multiplied by the measured difference in active effort. Adjust for the share of jobs completed at the required quality, review, and correction. Treat the estimate as a planning assumption until a pilot measures the difference.
Keep training and evaluation cases apart
Use one case set to refine instructions, tool descriptions, and schemas. Hold out a second set for comparison. Replaying a case after tailoring the skill to that case measures a different question from handling unseen work.
For a historical replay, capture the facts available at the time of the original request. Reading today’s paid status into a months-old collections case gives the agent an answer unavailable to the original operator.
Require two reviewers for ambiguous high-consequence cases during label calibration. Record disagreements and the final rule. A model can propose labels, but a sample of those labels still needs human review before the labels determine the roadmap.
Review priorities after launch
Join declined proposals, missing-tool requests, handoffs, corrections, and abandoned jobs to the same taxonomy. Review one job family at a time with product, engineering, and the operating team.
The useful backlog item names the desired result, affected customers, missing step, quality checks, consequence, owner, and evidence links. “Add CRM tools” is too broad to estimate or evaluate.
Agent tracing in practice compares tracing vendors and a weekly review routine. The BizOps scorecard explains how the pilot converts a priority estimate into a measured business result.