Research
Experiments with inspectable results. Tools for checking a decision.
Executed experiments#
Each experiment names the question, setup, actual result, and limits.
- Capability versions, pause controls, and stream returnFifteen executed application cases: domain grants, pending versions, paused tools, headless work, and HTTP event replay. Synthetic identities and effects.Source + results ↗
- Approval and restart: an executed framework trialThe same invoice operation through AI SDK, OpenAI Agents SDK, and LangGraph: saved approval, replacement workers, policy changes, and a lost provider reply.Source + results ↗
Calculators#
Editable assumptions, not measured customer outcomes.
- Agent pilot impactCompare verified results, review, rework, and full operating cost.Calculator ↗
- Support agent cost modelCoverage, checked resolution, vendor billing, handoffs, and human review.Calculator ↗
- In-product copilot cost modelAccepted changes, review effort, corrections, and customer value.Calculator ↗
- Operations agent cost modelScheduled work, exceptions, operating cost, and checked results.Calculator ↗
How the research is reported
A hands-on product session uses independent inputs and checks the resulting work. A documentation review explains published behavior. A scripted tour shows the vendor’s chosen path. The product pages identify the evidence available.
A runnable implementation can demonstrate its own permission, persistence, or recovery contract. A synthetic model or provider does not establish a vendor’s behavior. Illustrations explain mechanisms and are labeled separately from original screenshots.
Cross-product comparisons will be published after the same useful task has been checked in the relevant products. Sidekick promotion and Notion/Airtable feedback sessions are still awaiting suitable test environments.