Skip to content

Agent tracing in practice

What a trace contains, four tracing tools worth trying, what to log beyond the model call, and a review routine a small team can start this week.

By Daniel Ternyak7 min readSources checked September 28, 2026

A trace is a record of an agent handling one request. A trace connects the model calls, retrieved context, tool calls, tool results, and application code involved in producing a response. Each recorded step is called a span. A conversation can contain several traces, grouped into a session or thread. Langfuse’s data model explains the structure.

Tracing becomes necessary when reading the final answer stops being enough to debug the agent. Why did the agent choose the wrong campaign? Which tool used most of the request’s 18 seconds? Did the model receive the selected customer’s ID? Did the agent retry a write after a timeout? A trace gives the intermediate inputs and outputs needed to answer those questions, as long as the team recorded those steps.

Tools such as Langfuse, LangSmith, Braintrust, and Arize Phoenix collect and display traces. Below is what I would look for in a tool, what I would log, and how I would use traces with a product team.

What a trace actually contains

Consider an email marketing agent. The user has a campaign open and asks, “Pause the Summer Sale email scheduled for Friday.” The example is made up; the IDs, tool names, and timings show the records worth inspecting.

text
agent.request                                          4.2s
  input: "Pause the Summer Sale email scheduled for Friday"
  selected_campaign_id: cmp_104
  llm.choose_action                                    1.6s
  campaigns.search("Summer Sale")                      0.3s
    result: cmp_087 (Oct 2), cmp_104 (Sep 25)
  llm.choose_campaign                                  1.1s
    tool_call: campaigns.pause({ id: "cmp_087" })
  campaigns.pause                                      0.4s
    result: { id: "cmp_087", status: "paused", operation: "op_55" }
  llm.respond                                          0.8s
    output: "I've paused Friday's Summer Sale email."

The useful detail is the mismatch between selected_campaign_id: cmp_104 and the tool argument id: cmp_087. A support ticket with only the final answer wouldn’t show where the mismatch happened.

Next, inspect the input to llm.choose_campaign. When cmp_104 was missing from the input, fix how the application supplies the selected context. When the model received cmp_104 and still chose cmp_087, test the selection instructions and the tool contract. When the user had no campaign selected, the product may need to ask which Friday the user means. Each case needs a different fix.

Also inspect the saved campaign records. Imagine cmp_087 is now paused and cmp_104 is still scheduled: the wrong campaign changed. Keep that result, or a restricted link to the result, with the trace. A tracing SDK won’t discover the application’s saved state unless the team records or links the state.

Which tracing tool to try

Four options are worth evaluating. The capabilities below come from each tool’s documentation, checked September 21, 2026. The advice is my recommendation, not a hands-on ranking.

ToolWhat to evaluate the tool for
LangfuseFollowing a whole customer conversation: sessions group traces across turns and support human scoring. Langfuse can also be self-hosted.
LangSmithEngineers and product reviewers working from the same runs, with annotation queues for structured human review.
BraintrustTurning production failures into regression cases: examples move from traces into versioned evaluation datasets.
Arize PhoenixStarting locally, or building on existing OpenTelemetry instrumentation: a tracing UI a team can run itself, with a collector that accepts OTLP.

The four overlap a lot. I would try the same failed request in two candidates before committing. Can the tool find the request’s exact model input, tool arguments, tool result, and slowest step? Can a reviewer follow the next customer turn, attach a review, and save the case? Have both the engineer and the product reviewer try each tool.

Then check the requirements that are expensive to change later: supported SDKs, access controls, redaction, data location, retention, export, and the cost at the expected volume. Self-hosting also means running storage, upgrades, and backups. OpenTelemetry sits underneath many of these integrations, but the standard doesn’t provide a review workflow; the collector and the analysis tool are separate choices.

What to log beyond the model call

Start with the tool’s supported integration for model calls. Then add spans around the product’s own tools, retrieval, and background jobs. An integration that captures the prompt and completion can still miss the product state that explains a failure.

For the campaign example, I would capture:

RecordQuestion the record answers
Request, selected object ID, relevant time zoneWhat did the user ask for, and what context did the application supply?
Model input, model name, prompt version, available tool schemasWhat information and actions were available at that step?
Retrieved records and their IDs or restricted referencesDid the search return the right campaign?
Tool name, arguments, result, duration, and errorWhat did the application actually attempt, and where did the time go?
Acting account, resource scope, permission decision, and approval ID when there is oneWhich account could act, and which operation had approval? Record references and decisions, never credentials.
Stable operation ID, affected object ID, receipt, and verification resultWhich saved change belongs to this request, and what remains unverified?
Application release, capability version, session ID, and task IDDid the problem start after a release? Are several requests retries of the same job?
Token usage and available cost estimatesWhich model calls are expensive? Keep estimated cost labeled as an estimate.

Keep a stable task ID when work moves to a queue or spans several requests, and link later traces when the job continues separately. A new browser request shouldn’t make an existing job impossible to find.

Don’t send API keys, authorization headers, or whole customer records just because a wrapper can serialize them. Decide which fields to capture, apply redaction, set retention, and test one known request to check that the trace is useful and appropriately limited.

A review routine a small team can start with

Put a trace link beside support feedback and internal error reports. Someone investigating “the campaign is still scheduled” should be able to open the matching run without lining up timestamps across three systems.

For each reported problem:

  1. Write down the expected result. In this case, cmp_104 should become paused and the other campaign should stay unchanged.
  2. Find the first incorrect step. Inspect the selected context, the search result, the model’s tool call, and the backend response, in that order. Don’t start by rewriting the final-answer prompt.
  3. Check the actual result. Open the affected object or retrieve the operation’s receipt. When neither is available, record exactly what can’t be verified.
  4. Assign the fix. Missing selected context goes to the integration owner. A wrong ID from otherwise complete input calls for a change to tool selection or the tool contract. A backend writing a different ID than the backend received is a bug in the tool itself.
  5. Keep a regression case. Save the relevant input, the starting state, and the expected result. Run the case against test tools when changing the agent; replaying a write against production could repeat the action.

In the campaign example, the regression test checks that cmp_104 is paused and cmp_087 is unchanged. Checking whether the answer contains “paused” would miss the original bug. Keep separate cases for ambiguous dates, a missing selection, and a legitimate permission denial. Before calling a fix done, test the fix on cases the fix wasn’t built from; Turn every bad run into a fix someone owns explains why.

Alongside reported problems, review a small random sample of ordinary runs: say, ten flagged runs and ten random ones each week. The routine is a habit, not a statistical estimate, so keep the two groups separate when reporting.

Use traces to find product work

Once runs carry a capability and a task ID, useful review views become possible. I would start with these:

Saved viewWhat to investigate
Repeated attempts at the same taskDid the first attempt fail, give an unclear answer, or leave the user unable to find the result?
Requests for unsupported actionsWhich specific action is missing, and how many distinct customers need it?
Tool calls denied by permissionsIs the denial expected policy, a missing connection, or a confusing setup flow?
Runs with unusually many tool calls or high latencyIs the agent searching repeatedly, retrying an error, or waiting on one slow service?
Results customers later correctedWhich field or choice needed repair, and can the correction become a test case?

Count distinct customers and underlying tasks before prioritizing anything: twenty retries by one person tell a different story from twenty customers asking independently. Open several examples to see whether the customers need a new feature or clearer access to one that already exists.

Traces cover people who used the agent. Traces don’t explain why someone avoided the agent or finished the job elsewhere. Take those questions into customer conversations instead of treating trace counts as the whole roadmap.

A first useful tracing setup

Instrument one capability end to end before expanding. In a test environment, open a successful run and a deliberately failed run. Check that both include model inputs, tool arguments and results, the application version, and links to the affected objects.

Give a teammate the failed run and ask the teammate to explain what happened. When the teammate has to ask which campaign was selected or whether the change persisted, add the missing context. Save the failure as a regression case and link the trace from the support workflow. A setup like that is enough to start using tracing in everyday product work.

The pattern#