A product agent needs a model loop, tools, saved state, and interfaces for people reviewing work. A framework can supply some of those pieces. The product team still needs to decide which pieces belong to the application.
Choose the framework using one representative job: prepare invoice reminders, wait for approval, send once, and recover after a worker restart. A weather-tool demo does not test the requirements of that job.
Decide what must survive before comparing SDKs
Write down the required behavior first:
- A browser can close while the job continues.
- A reviewer can return tomorrow to the same saved proposal.
- A worker can restart without sending the same reminder twice.
- Revoked access blocks execution even after the reviewer accepted.
- A support engineer can inspect the actual inputs, tool results, and product receipt.
A streaming response is a delivery mechanism. A database, checkpointer, or workflow runtime stores progress. A queue starts or resumes work. The framework choice should account for all three responsibilities.
Compare a small set of credible starting points
The table describes documented building blocks and an editorial recommendation about fit. The comparison is not a latency benchmark or a claim that other frameworks lack the features.
| Starting point | Good first fit | Documented building blocks | Product work still required |
|---|---|---|---|
| AI SDK | A TypeScript product with an interactive UI and an existing backend | Model and tool calls, streaming, tool approval requests and responses | Saved jobs and proposals, current authorization, restart handling, review UI |
| OpenAI Agents SDK | A code-owned agent loop with specialist handoffs and integrated tracing | TypeScript/Python agent definitions, orchestration, tool integration, runtime state and observability | Deployment, application storage, business tools, approval decisions and product policy |
| LangGraph | A workflow needing explicit graph state and pauses between steps | Persistent checkpoints, thread IDs, interrupts and resumption | Production checkpointer, job ownership, authorized resume endpoints, safe side effects |
| Direct model API plus the existing job system | A small bounded loop where the team already owns reliable orchestration | The team controls the request, tool dispatch, and stop conditions | Every loop, stream, replay, error, and approval convention the team chooses to build |
Read the primary docs alongside the row: AI SDK tool calling, OpenAI Agents SDK, and LangGraph interrupts.
Check approval behavior rather than the feature name
AI SDK’s current v7 documentation configures approval through toolApproval. Older examples use needsApproval, which the current docs mark deprecated. A manual approval produces request parts; the application collects the decision and makes the next call. Provider-executed tools are outside that local approval setting. AI SDK approval semantics
The practical question is where the pending proposal lives between those calls. A browser-only message array does not establish durable review. The application should save the exact proposal, bind the review to the authorized person, and enforce current permissions when the tool commits.
LangGraph’s interrupt() needs a checkpointer and a stable thread ID. Resuming can rerun the containing node, so effects before the interrupt need duplicate prevention or separation from the approval node. A graph checkpoint does not make an email send transactional. LangGraph interrupts and side effects
An SDK approval flag is useful plumbing. The operation contract remains the product’s responsibility.
Run a framework trial with the same acceptance tests
Give each candidate the same tools and case set. Keep the model and inputs fixed for the initial comparison so changes in orchestration are easier to inspect.
| Trial | Evidence to inspect |
|---|---|
| Prepare a correct draft | Exact selected objects and output, not only the final answer |
| Decline a proposed send | No external send; a durable declined decision |
| Restart while waiting for review | Same proposal and authorized resume path |
| Restart after the provider accepts a send | One send and a reconciled receipt |
| Revoke access before resumption | Blocked operation with a readable reason |
| Reconnect the UI stream | Current job state without restarting the model loop |
| Exhaust the tool budget | Bounded spend and a useful partial result |
Record implementation effort, missing infrastructure, operational debugging, active review time, and verified output quality. Framework popularity does not answer those questions.
Keep business operations outside framework-specific glue
Define the invoice operation once in the product service. Thin SDK tools call that service using server-bound context. Native UI, background rules, and MCP can use the same service with different entry adapters.
The separation makes a future SDK change smaller: the model loop may change, while permission checks, proposal IDs, receipts, and recovery remain stable. Tight product coupling and replaceable orchestration can coexist.
An existing job runtime is often enough for a first agent. Add another runtime when a measured requirement justifies the operational cost. Durable graph state helps when the workflow really has graph state to recover; several specialist agents help when the specialists have distinct work and ownership.
Write down the framework decision
A short decision record should name the first job, required pauses and recovery, selected SDK and pinned version, storage, worker runtime, tracing destination, and the tests that passed. Include the conditions that would make the team reconsider.
The recommendation follows the team’s constraints: start with the existing application stack, choose the smallest loop that supports the job, and prove the restart and approval behavior before increasing autonomy.
Documentation checked September 29, 2026. Framework behavior and API names should be rechecked when upgrading the pinned version.