An agent can prepare a billing reminder under version 1, wait for review, and resume after billing has changed the capability. A different team may still need its support tools while billing is paused.
This experiment checks a small shared execution service with two capability domains, saved proposals, and a replayable event stream. Fifteen cases passed in Node 24.17.0 on September 30, 2026. The source runs locally without a model, external provider, or account.
The question and setup
Does pending work retain the capability version and object state the person reviewed? Can a headless worker finish while the viewer is disconnected? Does a returning viewer get only the new events and undergo a fresh access check?
The example uses SQLite tables for capabilities, grants, objects, proposals, effects, and events. Billing and support have separate capability owners and reviewer grants. A child process opens the database for each commit. The event endpoint serves actual local HTTP server-sent events.
The effects table is a made-up provider ledger with a unique operation ID. The HTTP headers supply synthetic test identities. The example is not an authentication system or a deployable agent platform.
Run the experiment
Download the single-file source, then run:
node --no-warnings capability-trial.mjs --report results.jsonNode 24 is required for the built-in SQLite module. The experiment creates its own temporary database, binds a server only to localhost, checks the cases, and removes its temporary state. The published raw result records every case and the execution time.
What passed
| Boundary | Executed result |
|---|---|
| Persistence and duplicate review | A fresh process committed saved work; a repeat decision created no second ledger effect |
| Tenant and domain grants | An unrelated actor and cross-domain reviewer were rejected; the matching support reviewer completed support work |
| Current policy | Revoked access and a paused billing capability blocked prepared billing work |
| Domain isolation | Pausing billing left support work available |
| Pending versions | A capability upgrade and a changed source revision rejected the old proposal |
| Object scope | A same-tenant proposal targeting another tenant’s object was rejected |
| No side effects on denial | The blocked proposals left no entries in the synthetic effect ledger |
| Viewer return | A worker completed after the first viewer detached; replay after the cursor returned only completion |
| Stream access | Wrong-tenant and revoked viewers could not reconnect |
These rows group the fifteen individual cases in the raw report. The trial tests the application contract, not a model’s ability to select the correct tool.
The contracts worth carrying into a product
A capability registration needs an owner, version, and execution policy. A saved proposal pins the capability version and source revisions. The execution service checks the current grant and rollout state before committing the saved proposal.
The atomic transition and synthetic effect happen in one SQLite transaction. A real remote provider needs a different contract: durable operation ownership, an outbox or equivalent, provider-supported idempotency where available, and reconciliation of unknown outcomes. The SDK restart trial isolates that lost-reply question using a separate synthetic provider.
The viewer’s cursor belongs to an event log for a named job. Closing the viewer does not cancel the job. Reconnection checks access again and returns events after the cursor. The background-run guide explains the product and UI responsibilities.
What this trial does not establish
The SSE endpoint returns a finite replay; the example does not test a continuously open connection, production backpressure, retention gaps, mobile delivery, or proxy timeouts. Reviewer identities and grants are synthetic. The example does not integrate SSO, secret storage, external providers, or a real agent framework. Concurrent reviewer races are not tested by the sequential harness.
Use the source to inspect the boundary and add cases for the actual product. A production design also needs incident ownership, capability rollback policy, stream retention, and a reconciliation screen for unknown effects. Read skills and a shared capability platform for the team and rollout contract.