The browser says two updates are complete. The worker also knows about two. The destination has already accepted a third.
Nobody is necessarily wrong. The third reply disappeared after the write landed. A person returning to the job needs to understand what’s known, what’s uncertain, and what can safely happen next. A spinner can’t show all three.
Background work starts with one distinction: the operation, the process doing the operation, and the connection showing progress all last for different lengths of time. An agent that keeps working after a tab closes needs a product built around those differences.
A job is not a chat
One conversation can contain several jobs. Mara asks an agent to move twelve launch tasks to Friday. The agent prepares a proposal against those twelve task versions. Before approving, Mara also asks for a launch announcement: a second job in the same conversation, with a different output and no authority to change dates.
A teammate, Luis, then opens the project and asks to exclude the legal-review task. Meanwhile the announcement draft assumes every task moves to Friday. Mara approves the original date proposal from another device.
Which instruction wins? “Use the latest chat context” isn’t enough. Luis might have permission to edit the project without permission to amend Mara’s proposal. The announcement depends on a schedule nobody has committed. There are several jobs here, even if the interface shows one friendly assistant.
A useful product makes the jobs explicit. The date job owns a proposal version and a list of tasks. The announcement job refers to either the current schedule or the proposal, labeled as not yet approved. Luis’s request becomes a proposed amendment, a separate edit, or a rejected instruction, depending on Luis’s actual permissions. The request doesn’t quietly rewrite what Mara approved.
Give the work an ID before giving it workers
A job record names the requester, the objective, the affected resources, the authority, and what counts as done. Messages link to the job. Several conversations can refer to the same job without becoming separate runs.
A notification saying “Review schedule” should reopen the schedule job and the exact proposal version. Cancelling the announcement shouldn’t suggest the approved date changes stopped too.
I would start with one clear completion owner, whether the work uses one loop or several agents. Someone has to reconcile the request, the evidence, and the claim that the work is finished. Parallel reads, like checking holidays, dependencies, and availability, don’t need parallel authority. A dependency checker can report conflicts against a fixed proposal without permission to change dates. A copywriting agent can draft the announcement without permission to publish the announcement. Splitting the work pays off only when checked results, speed, or review effort improve, counting the time spent merging what the agents return.
Keep three accounts of the same work
Now give each intended update a stable ID of its own, one that survives a replaced worker or a new browser session. Three accounts of the work then exist side by side:
| Account | What the account can establish | What the account can’t establish alone |
|---|---|---|
| Destination | A change the destination accepted and, where supported, a receipt | What the worker or the user has heard |
| Run record | Confirmed results, pending work, uncertain requests, and the current worker | Whether an unacknowledged request landed |
| Browser | The latest state delivered to this viewer | That nothing happened while the viewer was disconnected |
The product usually can’t see the destination perfectly: a receipt API, a search, or a person who can check may or may not exist. Record the gap in the run as uncertainty, not as an assumed failure.
A progress count should say what the count counts. “Two confirmed; one being checked” is useful when a third update may already exist. “Two of five” leaves the person to guess whether the rest is safe to retry.
A timeout changes knowledge, not history
Suppose update three reached the destination, landed, and lost its reply. Retrying with a new ID asks for new work, and the update might add an audit entry or trigger another effect even when the value looks the same.
A stable ID helps only when the destination honors the ID. Can the destination recognize a repeat? Can the destination return the original result, and for how long? Does the destination reject the same ID with different inputs? Amazon’s Builders’ Library describes caller-supplied request IDs saved together with the change, and responses that let a retrying caller recognize the original result. Making retries safe with idempotent APIs
After a lost reply, a product has three possible next actions:
- Retrieve the saved result of the same operation, when the destination supports a lookup.
- Check the affected object directly, when the evidence can identify this operation’s effect.
- Keep the outcome unresolved, when neither route gives enough evidence.
The third is inconvenient but honest. Finding no matching record at one moment isn’t enough while an old request is still in flight and could still land. Replacing a worker doesn’t recall the old worker’s requests either: “the new worker owns the run” and “the old worker can no longer cause an effect” need separate evidence.
Cancellation is a request with an outcome
Say update four has started when the person asks to stop. Cancellation can prevent update five from being sent. Update four has already crossed a different boundary: the outcome depends on whether the request can be withdrawn before delivery or the destination accepts the request first.
“Stopping” is a useful state while that settles. “Cancelled” should come with the work already kept and any remaining uncertainty. Cancellation doesn’t restore the first three updates; recovery is a separate operation.
Linear’s agent protocol makes the boundary explicit: after a stop signal, an agent must halt and send a final activity describing the session’s state. The rule doesn’t promise that a submitted request can be rolled back. Linear agent signals
Completion belongs to the job
Imagine nine date changes succeed, two conflict, and one is still uncertain after a lost reply. An agent returning “done” mustn’t close the job. The completion owner reconciles each intended change and keeps the uncertain one visible.
The announcement job can now be honest. The announcement might describe the confirmed subset, wait for the uncertain change to settle, or ask Mara to revise the objective. The announcement can’t safely say every date moved just because the writing task succeeded.
Reconnect to evidence, not another request
When a viewer reconnects, the product should replay every event the viewer missed, in order, or fetch a fresh snapshot. When old events have expired, say so and fetch the current state instead of pretending the old position still works. Reconnecting the transport doesn’t create durable work by itself. HTML Standard: Last-Event-ID
The run’s own page is the return address: the scope, confirmed results, uncertain effects, pending decisions, and links to the affected objects. A notification can lead there but shouldn’t be the only record.
For scheduled work, add the standing instruction, the owner, and the current authority that justified starting. A schedule doesn’t answer what happened after a timeout any better than a chat message does. Underneath all of it sits one requirement: one durable job ID, with one owner.