Skip to content
Blog

Visualizing Agent Trajectories: Debugging Loops and Dead Ends

Agent logs lie by omission. Learn to render trajectories as graphs, spot infinite loops and dead ends, and turn run visualization into a debugging practice.

Published on • October 6, 2026

AI Assistant

Ask an agent framework for its logs and you get a timeline: messages in order, tool calls in order, tokens in order. That’s exactly the wrong shape for debugging. Agent failures are structural - the agent circled back to a tool it had already exhausted, it took a branch that could never reach the goal, it dead-ended in a tool that returned nothing usable. You can’t see a loop in a list. You see it in a graph.

Trajectory visualization means rendering a run as nodes (steps, tool calls, decisions) and edges (control flow), then interrogating that graph for the shapes of failure.

Anatomy of a trajectory

A trajectory is the complete record of one agent run:

{
  "run_id": "run-7f3a",
  "steps": [
    {"id": 1, "type": "llm", "output": "search for refund policy"},
    {"id": 2, "type": "tool", "name": "search", "args": {"q": "refund"}, "result": "0 hits"},
    {"id": 3, "type": "llm", "output": "search again with different terms"},
    {"id": 4, "type": "tool", "name": "search", "args": {"q": "returns"}, "result": "0 hits"},
    {"id": 5, "type": "llm", "output": "search again..."}
  ]
}

As a timeline, this reads as “tried a few things.” As a graph, the pattern jumps out: the LLM node keeps feeding the same tool, the tool keeps returning nothing, and there’s no edge to a fallback. That’s not a hard problem - it’s a missing edge.

Building the graph

Whatever framework you use, the visualization pipeline is the same three steps:

1. Capture spans, not just messages

Every step becomes a span with: type (llm / tool / retrieval / control), name, arguments, result, duration, parent. Parent-child relationships give you the graph edges for free - a tool called during a sub-agent run hangs off the sub-agent’s node, not the root.

Standards matter here: OpenTelemetry’s GenAI semantic conventions give you gen_ai.system, gen_ai.request.model, and tool-call attributes that multiple backends already understand.

2. Render with structure-preserving views

Three views cover 90% of debugging sessions:

  • Directed graph (default) - nodes = steps, edges = “called during” or “followed”. Color by type (LLM blue, tool orange, error red); size by duration.
  • Flame/timeline graph - the same tree, laid out left-to-right by time. Shows where time went while preserving hierarchy.
  • State-diff view - for state-machine agents (LangGraph-style), render the state at each node rather than the raw messages. Often the bug is “the state key never got populated.”

LangGraph’s ecosystem is built around this: the library redirects its docs to the LangGraph Platform documentation where Studio provides exactly this graph view of a run - nodes for each super-step, edges for transitions, and the full state payload on click. The point isn’t the tool; it’s the discipline of looking at runs structurally.

3. Aggregate across runs

One run’s graph shows an incident. A thousand runs’ graphs show systemic shapes: heat-map the edges and you’ll find the tool everyone fails through, the decision node that flips, the branch that never terminates.

The failure shapes to hunt

Once you’re looking at graphs, failures fall into recognizable patterns:

The loop. A cycle: the same node (or node pair) visited repeatedly without progress. Classic cause: a tool fails, the LLM retries with a near-identical prompt, the tool fails identically. Detection is trivial on a graph (cycle detection), impossible on a timeline without counting by hand.

The dead end. A node with no outgoing edges that isn’t a terminal state - the agent hit a “no results” tool result and simply stopped producing steps. In the graph: an out-degree-0 node that isn’t labeled done/error.

The thrash. Not a cycle, but near-repetition: 12 searches with paraphrased queries, each returning marginal results. Graph signature: a fan-out of same-type nodes under one LLM parent with monotonic diminishing returns.

The wrong branch. The agent committed to a path early and the visualization shows the correct branch was available one decision earlier. This is a prompt/planning issue you can only see with the full tree, including the options the LLM didn’t take (log the candidate actions at each decision point).

The escalating context. Node payloads grow monotonically until the run dies on context overflow. Render node size by token count and the trajectory becomes a visual budget warning.

A minimal instrumentation pattern

You don’t need a platform to start. A tracer that emits structured steps:

async def traced_step(kind, name, fn, **attrs):
    span_id = uuid4().hex
    tracer.emit({"id": span_id, "kind": kind, "name": name,
                 "attrs": attrs, "ts": time.time(), "parent": CURRENT.get()})
    start = time.time()
    try:
        result = await fn()
        tracer.emit({"id": span_id, "event": "ok",
                     "duration": time.time() - start,
                     "result_summary": summarize(result)})
        return result
    except Exception as e:
        tracer.emit({"id": span_id, "event": "error",
                     "error": repr(e), "duration": time.time() - start})
        raise
    finally:
        ...

Then a renderer: dump spans to JSON, and either (a) feed them to an OTel backend with graph views, or (b) render client-side with any graph library - dot for a static PNG during development is genuinely enough.

Minimum viable queries once you have the data:

  • Find cycles: standard SCC detection over the step graph.
  • Find dead ends: nodes with no children and non-terminal status.
  • Find repeated tool calls with identical normalized args.
  • Compute the “distance to success” distribution across failed runs - if most failures die at depth 2, the early planning is broken, not the tools.

Live visualization during development

For interactive work, streaming views beat post-hoc analysis: render the graph as the run grows, gray out completed nodes, highlight the node being executed. Two uses:

  1. Loop alarm - highlight on the first repeated node rather than after the run dies. The cheapest intervention is stopping a loop at step 3, not post-mortem at step 30.
  2. Branch inspection - when a run goes off the rails, pause and inspect the current state node’s payload. In state-graph frameworks this is a first-class operation.

From visualization to guardrails

Every pattern the graph reveals becomes a runtime check:

Graph patternRuntime guardrail
CycleMax-visit counter per node/tool; force fallback edge
Dead endRequire every non-terminal step to emit next-step intent
ThrashDeduplicate normalized tool args within a run
Context growthToken budget per branch; summarize-and-truncate
Wrong branchLog decision candidates; add re-plan trigger at depth N

The loop you visualize this week is the loop you can prevent next week - but only if the trajectory data survives the run. Persist trajectories keyed by run ID alongside your evaluation suite, and you can regression-test the shape of agent behavior, not just its final answer.

Key Takeaways

  1. Timelines hide structural failures; render trajectories as graphs (nodes = steps, edges = control flow) to see loops, dead ends, and thrash instantly.
  2. Capture structured spans with parent-child relationships - the graph falls out of the data model, whatever framework you use.
  3. Learn the five failure shapes: loop, dead end, thrash, wrong branch, escalating context - each has a graph signature and a corresponding guardrail.
  4. Aggregate across runs with edge heat maps to find systemic weak points, not just incidents.
  5. Turn every visualized pattern into a runtime check, and persist trajectories so agent behavior itself is regression-tested.

References: