Skip to content

So you’re building a product agent

The eight decisions every team hits after the demo works, in the order they arrive, with the company that got each one right and the one that shows what getting it wrong costs.

By Daniel Ternyak5 min readSources checked September 28, 2026

Someone just told your team to put an agent in the product.

The model will take a week. You’ll pick one, wire up a few tools, and have a demo that answers questions and changes a record. Everything after that week is product work: whose access the agent uses, what a person sees before anything changes, what happens at 9:00 AM on a Monday when nobody is watching, and what number you show leadership.

I’ve taken apart twelve agents from companies that already shipped one: Shopify, Notion, Linear, Intercom, Salesforce, HubSpot, Asana, Figma, Airtable, Microsoft, Atlassian and Zapier. They hit the same eight decisions, roughly in this order. Most of them got some right and at least one wrong. Here they are, with the team that got each one right and the one that shows what it costs to get it wrong.

1. Write down what “done” means

Before you pick a model, pick one job and write what finished looks like. Which records change? What stays a draft? How does the person hear about the parts that didn’t work?

Linear shows why this comes first. Its agent sessions end in a state called “Complete,” and Complete means the session ended. It says nothing about whether anyone accepted the work. Figma has the better definition for its design agent: done is when the selected area changed and the new objects still use the named component. That’s a check you can run, and it catches a screen that looks right but was built wrong.

A final message from the agent is a claim. Changed records and receipts are the evidence.

2. Decide whose authority the agent uses

This is the decision teams most often make by accident. There are two choices, and both are defensible.

The agent can act with the person’s own access. Asana does this: a Teammate answers only from what the requester may see, and runs actions on the requester’s own connection.

Or the agent can carry its own grants. Notion’s Custom Agents do this, and sharing the agent shares its access. A department lead can ask a finance agent about pages they can’t open. That’s powerful on purpose, and it surprises admins who think in page permissions.

Atlassian makes this a visible setting on Rovo: the person, or the agent’s own account. Whatever you choose, show it.

For customer-facing agents there’s a second question hiding here. Salesforce binds the booking action to a verified customer ID in configuration, so a name typed into the chat can’t redirect it. Who the customer is and whose permissions run the action are two separate answers.

3. Show the change where it will land

Don’t ask people to approve a chat summary. Put the agent’s draft in the screen they already trust.

Shopify’s Sidekick is the best example. It fills Shopify’s own discount form and tints every value it chose. That’s how a merchant can see that “15% off this weekend” became a discount that runs through Monday, before saving it.

Microsoft does the same in Excel: Copilot’s work lands in Excel’s own formulas, tables and undo stack, where you can check it. Change an input and see whether the answer moves.

4. Make approval bind to the exact change

An approval only means something if what runs is exactly what the person approved.

Zapier shows how easy this is to get wrong. With Request Approval set to continue, the step reports “Success” for approved, declined and timed-out runs alike. Put the approval after the AI step and it approves a write that already happened.

Intercom gets it right. Fin’s Procedures put human approval in the program as a real branch with a timeout that escalates. Branch on the reviewer’s decision, never on whether the step finished. The approvals guide covers the rest.

5. Plan for the job that outlives the chat

Your agent will graduate from answering in a chat to running on a schedule. When it does, the safety you designed for chat may quietly disappear.

Rovo is the clearest case: the same agent that asks before writing in chat can write directly when a rule runs it every Monday at 9:00 AM. Treat moving a job from chat to a schedule as a new permission review.

Airtable’s field agents show the same thing in a spreadsheet. Approving the first value approves every future value, because the field reruns whenever its inputs change.

Two habits help. Make long tasks saved jobs with a status, so people can leave and come back. And keep who owns the work separate from who does it. Linear keeps the human assignee while the agent works the issue as a delegate, so a stalled agent never takes work out of someone’s queue. The guide on background work goes deeper.

6. Check who will read the output

Read permission covers the sources. It says nothing about the audience.

Asana gets the first part right and leaves the second to the person asking. The Teammate answers from what the requester may see, then posts the answer as a comment everyone on the task can read. Check who will read the output separately from who may read the sources.

7. Keep a record and an honest way back

Undo is several different operations, and teams usually build only one.

Intercom separates rolling back the agent’s configuration from undoing what it already changed in other systems. Shopify’s weak spot is right here: after Save, no single undo covers what Sidekick changed. And Airtable gets one detail exactly right: a cell someone edited is never overwritten. Treat a person’s correction to agent output as a lock.

Give every unattended run a log a teammate can read without the original chat. You’ll need it the first time someone asks what happened. The traces guide shows what to record.

8. Name the metric for exactly what it counts

The last decision is the one leadership will ask about first.

HubSpot’s Customer Agent reports “resolved,” and resolved means no qualifying handoff within 72 hours of the visitor’s last reply. Intercom documents Fin’s outcome rate over two different sets of conversations, so one set of outcomes can produce two different rates.

Neither number is wrong. Both are easy to misread. Agree with finance on what counts before launch, then name the metric for exactly that. The guide on proving the pilot has the scorecard.

The one-page version

Before your first pilot, you should be able to answer these in a sentence each:

  1. What does done mean for this job, and what evidence proves it?
  2. Does the agent act with the person’s access or its own, and where does the person see which?
  3. Which screen shows the proposed change, with the agent’s choices marked?
  4. What exactly does an approval bind to, and what happens on decline or timeout?
  5. What changes when this job runs on a schedule instead of in a chat?
  6. Who will read the output, and may they?
  7. What can be undone, what can’t, and who can read the run log?
  8. What does the headline metric count, and who agreed to it?

If any answer is “we’ll figure it out,” that’s the next thing to build. Start with the product is the harness, then read the teardowns of the companies closest to your own product.

Get the weekly teardown

One agent, taken apart, with the decisions you can copy.

One teardown a week. Unsubscribe anytime.

Next guideThe product is the harness

The model picks the next step. Your product decides what the agent may touch, what a person sees before it changes, and what proves it happened. That is where the work is.