Skip to content
Blog

Caching Tool Results to Shorten Agent Trajectories

Agent trajectories grow expensive fast because every tool call costs a model round trip. Learn what is safely cacheable, how to key the cache, where to place it in a LangGraph graph, and how to measure the wins.

Published on • October 4, 2026

AI Assistant

Why Agent Trajectories Get Long and Expensive

An agent trajectory is the full sequence of model calls, tool invocations, and state updates produced while solving one task. A simple “what changed in this repo?” question can easily produce a dozen steps: list files, read file A, read it again with a different range, grep for a symbol, re-read the file after the grep, check the config twice because the first answer was truncated.

Each tool call is not just a function invocation — it is an extra model round trip. The model emits a tool call, your runtime executes it, the result goes back into the message history, and the model runs again to decide what to do next. That round trip costs input tokens for the entire accumulated history, output tokens for the decision, and latency from tool execution plus a full prefill over the growing context.

The compounding part is what hurts. Step 10 pays for steps 1–9 in its prompt. A trajectory that repeats the same read_file call three times pays for that file’s contents three separate times, at three progressively larger prompt sizes.

Caching does not shorten the logical plan, but it shortens the expensive parts: tool executions become instant, redundant model round trips can be skipped, and prompt prefixes stop being billed at full price when the provider supports prefix caching.

What Is Safely Cacheable

The first rule of tool-result caching: cache reads, never writes.

Cacheable candidates are deterministic (or near-deterministic) and read-only:

  • File reads — read_file, list_dir, git diff at a pinned ref
  • Searches — grep, glob, semantic search over an unchanging index
  • Code analysis — AST parses, type inference, dependency graphs for an unchanged commit
  • API GETs — issue details, package metadata, documentation pages, CI status at a fixed run ID

Not cacheable:

  • Stateful mutations — write_file, run_command, send_email, create_ticket, DB writes
  • Anything with side effects — even a “read” that increments a counter or logs a billable event
  • Non-deterministic calls — current time, random numbers, live stock prices, “latest” records
  • User-context-dependent reads — a get_profile call whose result depends on who is asking

Quick test: would replaying this exact call tomorrow against the same inputs give an identical answer, with nothing mutated? If yes, it is a candidate. If the answer is “identical except for who asked,” it is still cacheable — provided caller identity is part of the key.

Key Design: Tool Name + Normalized Args + Context Version

A cache key has three parts, and omitting any of them creates a correctness bug:

  1. Tool name — read_file and grep must never collide.
  2. Normalized arguments — canonicalized so semantically identical calls hit the same entry.
  3. Relevant context version — the thing that could make the answer change: repo commit, file mtime/hash, user ID, index version, prompt/schema version.
import hashlib
import json
from pathlib import Path

def cache_key(tool: str, args: dict, context: dict) -> str:
    payload = {"tool": tool, "args": _normalize(args), "ctx": context}
    raw = json.dumps(payload, sort_keys=True, separators=(",", ":"))
    return hashlib.sha256(raw.encode()).hexdigest()

def _normalize(args: dict) -> dict:
    return {
        k: str(Path(v).as_posix()).lstrip("./") if isinstance(v, str) and k in {"path", "file", "directory"} else v
        for k, v in args.items()
    }

key = cache_key("read_file", {"path": "src/app.py"},
                {"commit": repo_commit, "schema": "v3"})

The context version is what makes invalidation cheap. Key by repo_commit and entries for old commits become unreachable the moment the commit moves — garbage-collect them lazily.

Where to Place the Cache

There are two natural seams:

1. In front of tool execution (tool-result cache). A wrapper around the tool node checks the key, returns the stored output on a hit, and executes + stores on a miss. This eliminates redundant tool work and slow-call latency (big grep, cold API), but the model still runs.

2. In the model-call layer (response/prefix cache). Memoize (messages_hash) -> model_output, or rely on provider-side prompt caching for a stable prefix. This eliminates redundant model round trips — the expensive part — but requires the history to match byte-for-byte up to the cached prefix.

Most systems want both: provider prefix caching for the history, plus a tool-result cache so the history stops growing with duplicate tool messages.

from langgraph.graph import StateGraph

def make_cached_tool_node(tools, store, ttl_seconds=900):
    async def run_tools(state):
        last = state["messages"][-1]
        outputs = []
        for call in last.tool_calls:
            key = cache_key(call["name"], call["args"], state.get("cache_ctx", {}))
            cached = store.get(key)
            if cached is not None:
                outputs.append(cached)          # hit: no execution, no network
                continue
            result = await tools[call["name"]].ainvoke(call["args"])
            store.set(key, result, ttl=ttl_seconds)
            outputs.append(result)
        return {"messages": outputs}
    return run_tools

builder = StateGraph(AgentState)
builder.add_node("tools", make_cached_tool_node(TOOLS, store))
builder.add_edge("tools", "agent")

A memoized-reducer variant works too: hash the tool call inside a custom state reducer and skip appending a duplicate tool message when the same key was already resolved this run.

TTL, Invalidation, and Staleness Budgets

TTL alone is a blunt instrument. Combine three levers:

  • Content-addressed invalidation — key by file hash or commit; changes invalidate naturally.
  • Explicit version tags — bump a schema or index_version byte when your tool’s output format changes.
  • Staleness budgets per tool class — repo file contents can be fresh for one session; package metadata can be stale for hours; a CI status may be stale for 60 seconds.

Tune TTLs against how wrong you can afford to be: a stale file read that makes the model edit the wrong lines is a hard failure, a 10-minute-old docs result is fine. Set the budget from the failure cost.

Correctness Risks (and How They Bite)

Mutable data. The classic bug: cache run_command("git status") and the agent spends the session looking at a pre-edit snapshot. Anything another tool in the same trajectory can change is either keyed to an invalidation signal or not cached.

User context leaking across keys. get_user_settings(user_id) cached under a key that omits user_id serves one user’s preferences to another — a privacy incident, not just staleness. Put tenant, user, and permission scope in the key, or refuse to cache user_scoped tools without it.

Cache poisoning via malformed args. If you normalize by dropping unknown fields, {"path": "a.py", "mode": "write"} collides with {"path": "a.py", "mode": "read"}. Normalize by canonicalizing, not truncating: distinctly key any argument you do not understand, and validate types before hashing so 1 and "1" do not collide.

Truncated or shaped results. If a tool truncates output at N characters or paginates, include limits, offsets, and page tokens in the key — two calls with different output shaping must not collide.

Cross-run contamination. A cache shared between tests and production, or across two repos in the same working directory, needs repo identity in the context block too.

The LangGraph Pattern

In LangGraph terms, the pieces map cleanly:

  • Tool node wrapper — the run_tools node above; the standard place to intercept.
  • Cache context in state — carry commit, user_id, index_version on graph state so the key builder never reaches into globals.
  • Memoized reducer — a custom add_messages-style reducer that dedupes tool messages by key, keeping transcripts short.
  • Checkpointers — durable snapshots for resuming a trajectory without re-executing paid-for tool calls.

Measuring Success

Do not ship a cache you cannot see. Instrument four numbers before and after:

  1. Tool call count per trajectory — total calls and unique calls; the gap is your theoretical hit rate.
  2. Tokens — prompt and completion tokens, split into cached vs. uncached input.
  3. Latency — p50/p95 wall time, split into tool time vs. model time.
  4. Hit rate and stale-hit rate — hits divided by lookups, plus a sampled audit of hits that should have been misses.

Log those alongside usage from each response — most providers report a cached_tokens detail field you can sum directly.

Healthy early results: 30–50% of tool executions eliminated on repeat-heavy workloads, prompt tokens down as duplicate tool messages disappear, and p95 latency dropping most on trajectories that touch slow tools. If your hit rate is high but latency did not move, you cached the cheap part — move the cache in front of the model call layer instead.

Caching is not a substitute for tight agent design. Fewer, better-chosen tool calls still win. But for the redundant half of every trajectory — the re-reads, the re-greps, the re-fetches — a well-keyed cache is the cheapest performance work you will do on an agent.

Sources