Skip to content
Blog

Root Cause Analysis Playbooks for Agent Failures: From Symptom to Fix

A systematic method for debugging agent failures: classify the symptom, walk a trace from the failing span upward, and apply a playbook for each failure class — routing, tool, context, hallucination, and infrastructure — with durable execution and state inspection as first-class techniques.

Published on • October 10, 2026

AI Assistant

“The agent did the wrong thing” is not a bug report. It is a symptom with at least six unrelated causes behind it, and until you separate them you will be guessing — fixing the prompt when the tool was slow, adding retries when the router was misrouting, rebuilding context management when the real problem was a stale cache.

Agents make root cause analysis harder than ordinary software for one reason: the failure is rarely where the error appears. The user sees a wrong answer; the cause is three tool calls back, in an argument the model misread. Traditional RCA — bisect the stack trace, find the throwing frame — does not apply, because nothing threw.

What replaces it is a classified playbook: map the symptom to a failure class, then run the diagnostic procedure for that class against the trace. This post is that playbook.

In this tutorial, you will learn how to:

  • Classify agent failures into six operational classes from the symptom alone
  • Walk a trace from the failing span upward instead of reading logs chronologically
  • Apply a specific diagnostic procedure to each failure class
  • Use durable execution state to reproduce a failure deterministically
  • Distinguish model problems from harness problems — and fix the right one
  • Build an RCA practice that improves the system rather than just closing incidents

Key technologies: LangGraph (durable execution, persistence, interrupts), LangSmith tracing and observability, OpenTelemetry spans, structured evaluation sets.

Prerequisites

  • An agent system with tracing enabled end to end
  • Access to run state or conversation history for the failing run
  • An evaluation set, even a small one

Step one: classify the symptom

Before opening a single trace, assign the failure to one of these classes. It takes a minute and it determines which procedure you run.

ClassSymptomQuestion to answer
RoutingThe wrong agent or tool handled the requestDid control go where it should have?
ToolThe right agent, wrong action or resultDid the tool do what was asked?
ContextThe agent forgot or ignored known informationWas the right information in the prompt?
ReasoningRight inputs, wrong conclusionDid the model draw a valid inference?
InteractionCorrect answer, wrong turn — extra loops, dead endsDid the control flow terminate properly?
InfrastructureSlow, timed out, errored, or produced garbageDid the machinery work at all?

Two habits make classification reliable:

Classify from the transcript, not from the user’s description. Users report outcomes (“it charged me twice”), not mechanisms. Read the actual run before deciding.

Classify the earliest observable failure, not the loudest one. A hallucinated answer caused by a tool returning truncated data is a tool failure, not a reasoning failure. The later symptom will mislead you if you classify from it.

If you genuinely cannot tell, you are missing trace data — jump to the infrastructure procedure first, because untraced runs are their own failure class.

Step two: walk the trace from the failure upward

Once classified, open the trace for the failing run. The order you read it in matters.

1. Find the span where reality diverged from expectation. Not the last span — the first one that did something you can point at and say “that is wrong.” In a trace, that is usually a tool.* span with surprising arguments or a surprising result.

2. Read the parent chain upward from there. Why was this tool called with those arguments? Follow the llm.call span that produced them, then its parent. Each hop answers “what led to that.”

3. Stop when the decision looks correct. The span whose inputs were all reasonable but whose output was wrong is your root cause candidate. If every span looks reasonable and the outcome is still wrong, the defect is in your expectation, not the code — which usually means an evaluation problem.

4. Check what was in the prompt. For every llm.call along the way, confirm the context actually contained what the decision depended on. This is the single most common finding: the model did not have the information, and the failure looks like reasoning.

The reason to work upward rather than chronologically forward is that agent traces are mostly noise — dozens of spans doing exactly what they should. The divergence point is the signal, and reading toward it avoids drowning in the parts that worked.

The playbooks

Class: routing

Diagnostic. Compare the handoff span’s handoff.reason attribute against what the request actually needed. Then inspect the llm.call that produced the routing decision and the tool/agent list it was offered.

Common causes, in order of frequency:

  • Ambiguous tool or agent descriptions. The model chose between two candidates whose descriptions overlapped. This is the most common cause by a wide margin, and it is a text problem, not a model problem.
  • Missing candidate. The right agent was not in the list at all — a registration bug, or a handoff that was gated behind a feature flag nobody turned on in that environment.
  • Prompt overrides routing. The system prompt says “always escalate to a human” and the router dutifully does, for everything.
  • Stale routing context. The routing decision was made from turn 1’s context and never revisited after the user clarified.

Fix. Rewrite descriptions so they are mutually exclusive — if two descriptions could both be true for the same request, the model has no basis to choose. Verify candidate registration with a test that asserts the expected agents are present. Move routing instructions out of the general system prompt.

Verify. Build a routing evaluation set: 20–50 (request, expected agent) pairs. Run it after every prompt change. Routing regressions are invisible in casual testing because most requests still work.

Class: tool

Diagnostic. Start at the tool.* span and check four things in order: were the arguments correct, did it return an error, was the result truncated, and did it take abnormally long?

  • Wrong arguments, correct call intent → the model misread or hallucinated a parameter. Check whether the schema’s descriptions are precise enough to disambiguate.
  • Tool error returned to the model → the model usually handles this gracefully if the error message is specific. A stack trace or a bare 500 produces confusion; Invalid invoice number: expected format INV-YYYY-NNNN, got "12345" produces a correction.
  • Truncated result → the tool returned a payload larger than the context window allows, and the model reasoned over the visible part. Look for tool.result.size attributes near your context limit.
  • Slow tool → not a correctness bug, but it triggers retries and timeouts that then look like correctness bugs.

Fix. Tighten parameter descriptions. Make tool errors model-actionable — this is a code change in your error handling, not a prompt change. Add explicit truncation markers to large results rather than silently cutting them.

Verify. Unit-test each tool’s error paths with the actual inputs that failed. Tools are the one part of an agent system that is conventional software and should have conventional tests.

Class: context

Diagnostic. Open the llm.call span for the decision that went wrong and read the context that was actually sent. Three questions:

  1. Was the relevant information present at all?
  2. If present, was it buried — past the middle of the prompt, under three layers of nested JSON?
  3. Was contradictory information also present?

Causes and fixes:

Never loaded. A retrieval step failed silently, returned nothing, or queried the wrong collection. Check the retrieval span’s result count — a search span returning zero hits in a domain where results should exist is your bug, and it will present as a reasoning failure.

Loaded but lost to the window. Context management dropped it during compaction, or an aggressive truncation strategy kept the recent turns and discarded the older fact. Inspect what survived compaction; a summarization step that loses entities and dates is the usual culprit.

Contradictory. The prompt contains both “the refund window is 30 days” and an older “refund window is 90 days.” Models resolve conflicts unpredictably. Dedupe at the source.

Positional. The fact sits at token 14,000 of a 16,000-token prompt. Long-context models are measurably worse at mid-context recall than at the beginning and end. Move it up.

Verify. Write context assertions: given a fixture conversation, assert the compiled prompt contains the required facts. These are cheap, deterministic, and they catch regressions that no behavioural test catches reliably.

Class: reasoning

Diagnostic. This is the last class you should conclude, because it is the hardest to fix. Before blaming the model, confirm:

  • The context contained the correct, non-contradictory information (context playbook passed)
  • The tools returned complete, correct data (tool playbook passed)
  • The routing was right (routing playbook passed)
  • The model was not asked to do something it demonstrably cannot — multi-step arithmetic, exact string matching, counting items in a list

If all four hold, you have a genuine reasoning failure.

Fix, in order of leverage:

  1. Give it a tool instead. Models are good at deciding what to compute and bad at computing. If the task is arithmetic, date math, or counting, emit a function call and use the result.
  2. Break the step apart. A single llm.call asked to retrieve, reason, and format will fail more often than three calls each doing one. Chain them and make the intermediate state inspectable.
  3. Add a verifier. A second pass that checks the output against the inputs catches a meaningful fraction of reasoning errors, at the cost of another call.
  4. Change the model. Only after the above. This is the most expensive lever and usually the least durable.

Verify. Keep a labelled set of the specific reasoning failures you have seen, as regression cases. Reasoning quality is not something you can assert in a unit test, but you can assert that a known-bad case now produces a known-good answer.

Class: interaction

Symptom: the agent produces the right answer but gets there badly — three tool calls where one would do, a loop that never terminates, an early exit that skips required steps.

Diagnostic. Look at the span sequence rather than any individual span. Count calls, look for repeated identical calls with identical arguments, and check where the run terminated.

  • Repeated identical calls → the model is not seeing the previous result, or the result did not satisfy what it was waiting for. Check whether tool results are actually entering the conversation history.
  • Unbounded loop → missing a termination condition. Set a hard cap on iterations per run; a framework-level max_turns is the standard control.
  • Premature termination → the model produced a final answer instead of calling the required tool. Usually a prompt issue, or an output schema that does not force the tool choice when one is mandatory.

Fix. Cap iterations. Make loop-breaking explicit in the prompt (“if you have already called this tool twice with the same arguments, stop and report the failure”). Use structured output or forced tool choice for flows where the path is fixed.

Verify. Assert on span counts in your evaluation harness: this input should produce exactly one search span. It is a blunt assertion, but it catches loops and redundant calls precisely.

Class: infrastructure

Diagnostic. Everything about the run looks wrong, or you cannot find the run at all. Check in this order: was the trace captured completely, did any span time out, were there rate limits or provider errors, and did the model return a truncated or malformed response.

Causes: provider rate limits producing silent degradations; a timeout mid-stream leaving a half-written response; a context window exceeded at the API level (the request never ran); a tracing exporter dropping spans so the trace you are reading is incomplete.

Fix. Make failures loud. A truncated stream should raise, not return partial content as if it were complete. Rate-limit responses should be classified and retried with backoff, not swallowed.

Verify. Chaos-test deliberately: inject a timeout and a malformed response, and assert your handling produces a clean, classified error rather than a plausible-looking wrong answer.

Durable execution: reproduce instead of guess

The reason many teams cannot do RCA at all is that the failure does not reproduce. It depended on a specific conversation state, a specific tool response, and a specific model output, and by the time anyone looks, all three are gone.

Frameworks with durable execution solve this by checkpointing run state at each step, so a paused or failed run can be resumed from an exact point:

  • Persist state at every step, not just at the end. A checkpoint store lets you inspect what the agent knew at the moment it decided.
  • Replay from a checkpoint with the same state and see the decision re-occur. To make this deterministic, record and replay the model responses rather than calling the provider again.
  • Use interrupts for interactive steps so a human can inspect and modify state at any point in the run — which is exactly what an RCA session needs.

The practical version: keep the checkpoint for failed runs for a defined retention window, and make “load run X and show me state at step N” a one-command operation. If reconstructing a failure takes an afternoon, nobody will do it and your RCA practice dies.

From one incident to a system that improves

A playbook that fixes the incident and stops has failed. The output of an RCA should be one of four artefacts:

A new class or subclass. Your first ten failures will not fit the six classes cleanly. Adding “retrieval returned zero results” as an explicit sub-case, with its own diagnostic, makes the next occurrence a two-minute diagnosis.

A regression case. Every diagnosed failure becomes an input to your evaluation set with the expected behaviour attached. This is how the failure stops recurring — not by the fix, but by the test that fails when someone reverts it.

A trace attribute you were missing. Most RCA sessions end with “we wish we had logged X.” Add the attribute. The next incident is cheaper.

A classification change. If the tool playbook keeps firing on what everyone assumed were reasoning failures, your consent or routing classification is wrong, and the fix belongs elsewhere than where you were looking.

Track three numbers: mean time to classify (should fall as playbooks accrete), regression cases in the eval set (should rise), and repeat incidents (should fall). The first is a proxy for whether the practice is being used; the other two are whether it is working.

A worked sequence

For a concrete failure — “the agent refunded the wrong order”:

  1. Classify. Read the transcript. The agent called issue_refund with order 4470 when the user asked about 4471. Wrong arguments, right tool, right agent → tool class, with a reasoning flavour.
  2. Find the divergence. In the trace, the tool.issue_refund span has arguments = {"order_id": 4470}. Its parent llm.call produced that. That span’s context contains two orders: 4470 and 4471.
  3. Read upward. The search_orders span returned both records, and the user’s message says “4471”. The model transposed a digit. Reasoning failure — but check the context first: both orders are in the prompt, unmarked, with similar fields.
  4. Confirm the class. Is this reasoning or context? The context contains a distractor that the model should have disambiguated. The robust fix is not “make the model more careful” — it is to have the tool return a single order when the query is unambiguous, so the model never chooses between two.
  5. Fix at the tool layer. search_orders should filter by the ID from the message rather than returning a list.
  6. Add the regression case. Fixture with two similar orders and a request for one; assert exactly one issue_refund span with the correct ID.
  7. Add the attribute. Log tool.arguments.order_id explicitly so the next similar failure is a one-line diagnosis.

Note where the fix landed: not in the prompt, and not in the model. In the tool contract. That is the outcome of a good RCA — the change goes where the leverage is.

A checklist

  • Classify from the transcript, using the earliest observable failure
  • Walk the trace from the divergence point upward, not chronologically
  • Check context contents before concluding a reasoning failure
  • Prefer a tool change over a prompt change over a model change
  • Keep failed-run checkpoints long enough to reproduce
  • Turn every diagnosed failure into a regression case
  • Track time-to-classify and repeat-incident rate
  • Update the playbook when a failure does not fit

Further reading