A trace is a record of an agent handling one request. A trace connects the model calls, retrieved context, tool calls, tool results, and application code involved in producing a response. Each recorded step is called a span. A conversation can contain several traces, grouped into a session or thread. Langfuse’s data model explains the structure.
Tracing becomes necessary when reading the final answer stops being enough to debug the agent. Why did the agent choose the wrong campaign? Which tool used most of the request’s 18 seconds? Did the model receive the selected customer’s ID? Did the agent retry a write after a timeout? A trace gives the intermediate inputs and outputs needed to answer those questions, as long as the team recorded those steps.
Tools such as Langfuse, LangSmith, Braintrust, and Arize Phoenix collect and display traces. Below is what I would look for in a tool, what I would log, and how I would use traces with a product team.
What a trace actually contains
Consider an email marketing agent. The user has a campaign open and asks, “Pause the Summer Sale email scheduled for Friday.” The example is made up; the IDs, tool names, and timings show the records worth inspecting.
agent.request 4.2s
input: "Pause the Summer Sale email scheduled for Friday"
selected_campaign_id: cmp_104
llm.choose_action 1.6s
campaigns.search("Summer Sale") 0.3s
result: cmp_087 (Oct 2), cmp_104 (Sep 25)
llm.choose_campaign 1.1s
tool_call: campaigns.pause({ id: "cmp_087" })
campaigns.pause 0.4s
result: { id: "cmp_087", status: "paused", operation: "op_55" }
llm.respond 0.8s
output: "I've paused Friday's Summer Sale email."The useful detail is the mismatch between selected_campaign_id: cmp_104 and the tool argument id: cmp_087. A support ticket with only the final answer wouldn’t show where the mismatch happened.
Next, inspect the input to llm.choose_campaign. When cmp_104 was missing from the input, fix how the application supplies the selected context. When the model received cmp_104 and still chose cmp_087, test the selection instructions and the tool contract. When the user had no campaign selected, the product may need to ask which Friday the user means. Each case needs a different fix.
Also inspect the saved campaign records. Imagine cmp_087 is now paused and cmp_104 is still scheduled: the wrong campaign changed. Keep that result, or a restricted link to the result, with the trace. A tracing SDK won’t discover the application’s saved state unless the team records or links the state.
Which tracing tool to try
Four options are worth evaluating. The capabilities below come from each tool’s documentation, checked September 21, 2026. The advice is my recommendation, not a hands-on ranking.
| Tool | What to evaluate the tool for |
|---|---|
| Langfuse | Following a whole customer conversation: sessions group traces across turns and support human scoring. Langfuse can also be self-hosted. |
| LangSmith | Engineers and product reviewers working from the same runs, with annotation queues for structured human review. |
| Braintrust | Turning production failures into regression cases: examples move from traces into versioned evaluation datasets. |
| Arize Phoenix | Starting locally, or building on existing OpenTelemetry instrumentation: a tracing UI a team can run itself, with a collector that accepts OTLP. |
The four overlap a lot. I would try the same failed request in two candidates before committing. Can the tool find the request’s exact model input, tool arguments, tool result, and slowest step? Can a reviewer follow the next customer turn, attach a review, and save the case? Have both the engineer and the product reviewer try each tool.
Then check the requirements that are expensive to change later: supported SDKs, access controls, redaction, data location, retention, export, and the cost at the expected volume. Self-hosting also means running storage, upgrades, and backups. OpenTelemetry sits underneath many of these integrations, but the standard doesn’t provide a review workflow; the collector and the analysis tool are separate choices.
What to log beyond the model call
Start with the tool’s supported integration for model calls. Then add spans around the product’s own tools, retrieval, and background jobs. An integration that captures the prompt and completion can still miss the product state that explains a failure.
For the campaign example, I would capture:
| Record | Question the record answers |
|---|---|
| Request, selected object ID, relevant time zone | What did the user ask for, and what context did the application supply? |
| Model input, model name, prompt version, available tool schemas | What information and actions were available at that step? |
| Retrieved records and their IDs or restricted references | Did the search return the right campaign? |
| Tool name, arguments, result, duration, and error | What did the application actually attempt, and where did the time go? |
| Acting account, resource scope, permission decision, and approval ID when there is one | Which account could act, and which operation had approval? Record references and decisions, never credentials. |
| Stable operation ID, affected object ID, receipt, and verification result | Which saved change belongs to this request, and what remains unverified? |
| Application release, capability version, session ID, and task ID | Did the problem start after a release? Are several requests retries of the same job? |
| Token usage and available cost estimates | Which model calls are expensive? Keep estimated cost labeled as an estimate. |
Keep a stable task ID when work moves to a queue or spans several requests, and link later traces when the job continues separately. A new browser request shouldn’t make an existing job impossible to find.
Don’t send API keys, authorization headers, or whole customer records just because a wrapper can serialize them. Decide which fields to capture, apply redaction, set retention, and test one known request to check that the trace is useful and appropriately limited.
A review routine a small team can start with
Put a trace link beside support feedback and internal error reports. Someone investigating “the campaign is still scheduled” should be able to open the matching run without lining up timestamps across three systems.
For each reported problem:
- Write down the expected result. In this case,
cmp_104should become paused and the other campaign should stay unchanged. - Find the first incorrect step. Inspect the selected context, the search result, the model’s tool call, and the backend response, in that order. Don’t start by rewriting the final-answer prompt.
- Check the actual result. Open the affected object or retrieve the operation’s receipt. When neither is available, record exactly what can’t be verified.
- Assign the fix. Missing selected context goes to the integration owner. A wrong ID from otherwise complete input calls for a change to tool selection or the tool contract. A backend writing a different ID than the backend received is a bug in the tool itself.
- Keep a regression case. Save the relevant input, the starting state, and the expected result. Run the case against test tools when changing the agent; replaying a write against production could repeat the action.
In the campaign example, the regression test checks that cmp_104 is paused and cmp_087 is unchanged. Checking whether the answer contains “paused” would miss the original bug. Keep separate cases for ambiguous dates, a missing selection, and a legitimate permission denial. Before calling a fix done, test the fix on cases the fix wasn’t built from; Turn every bad run into a fix someone owns explains why.
Alongside reported problems, review a small random sample of ordinary runs: say, ten flagged runs and ten random ones each week. The routine is a habit, not a statistical estimate, so keep the two groups separate when reporting.
Use traces to find product work
Once runs carry a capability and a task ID, useful review views become possible. I would start with these:
| Saved view | What to investigate |
|---|---|
| Repeated attempts at the same task | Did the first attempt fail, give an unclear answer, or leave the user unable to find the result? |
| Requests for unsupported actions | Which specific action is missing, and how many distinct customers need it? |
| Tool calls denied by permissions | Is the denial expected policy, a missing connection, or a confusing setup flow? |
| Runs with unusually many tool calls or high latency | Is the agent searching repeatedly, retrying an error, or waiting on one slow service? |
| Results customers later corrected | Which field or choice needed repair, and can the correction become a test case? |
Count distinct customers and underlying tasks before prioritizing anything: twenty retries by one person tell a different story from twenty customers asking independently. Open several examples to see whether the customers need a new feature or clearer access to one that already exists.
Traces cover people who used the agent. Traces don’t explain why someone avoided the agent or finished the job elsewhere. Take those questions into customer conversations instead of treating trace counts as the whole roadmap.
A first useful tracing setup
Instrument one capability end to end before expanding. In a test environment, open a successful run and a deliberately failed run. Check that both include model inputs, tool arguments and results, the application version, and links to the affected objects.
Give a teammate the failed run and ask the teammate to explain what happened. When the teammate has to ask which campaign was selected or whether the change persisted, add the missing context. Save the failure as a regression case and link the trace from the support workflow. A setup like that is enough to start using tracing in everyday product work.