Skip to content
Blog

Caching for Agentic RAG: Index Refresh and Query Caching

Agentic RAG loops re-retrieve constantly. Layer query caching, response caching, and index refresh policies to cut latency and LLM bills without serving stale answers.

Published on • October 6, 2026

AI Assistant

A classic RAG pipeline retrieves once and answers. An agentic RAG pipeline retrieves many times: the agent reformulates the query, calls search, inspects results, decides it needs more, and calls search again - sometimes a dozen times for one user question. Each hop pays embedding + vector-search + (often) LLM re-ranking costs, and every tool call adds latency the user feels.

Caching is the highest-leverage optimization for this pattern - but naive caching answers yesterday’s questions with today’s data. This post covers the cache layers that work for agentic RAG, and the refresh policies that keep them honest.

The four cacheable layers

User query
   │
   ▼
[1] Query semantic cache ──hit──► cached answer
   │ miss
   ▼
[2] Retrieval cache (embeddings + search results)
   │ miss
   ▼
[3] Re-rank / synthesis LLM calls (provider prompt cache)
   │
   ▼
Answer + [4] Index freshness policy

Layer 1: Semantic query cache

Exact-match caching rarely hits in conversational agents - “what’s our refund policy” and “how do refunds work” are the same intent. A semantic cache embeds the incoming query and serves the cached answer if cosine similarity exceeds a threshold (typically 0.92–0.97).

import litellm

litellm.cache = litellm.Cache()  # in-memory by default

# Redis-backed for multi-instance deployments
litellm.cache = litellm.Cache(type="redis", host="localhost", port=6379)

LiteLLM supports both exact and semantic caching; the gateway-level implementation means you can enable it for all agent traffic without touching application code.

Tuning risks:

  • Threshold too low → subtly wrong answers (“refund window is 30 days” served for a “60 days?” question).
  • Threshold too high → cache never hits, you pay for nothing.
  • Cache the wrong scope → tenant A’s answer served to tenant B. Always include tenant/user ID in the cache key.

Layer 2: Retrieval cache

Even when the final answer must be fresh, the retrieval step is cacheable:

  • Embedding cache: embedding the same (or normalized) query repeatedly is pure waste. Key = model + query hash. Hit rates in agent loops are high because reformulated queries converge.
  • Search result cache: key = (index version, filter, query embedding bucket). Store the top-K chunks with scores, not the rendered prompt.

This layer tolerates staleness better than Layer 1 because the synthesis step re-reads the chunks - if you version the key by index generation, a refresh automatically invalidates.

Layer 3: Provider-side prompt caching

For the LLM calls themselves, most providers now offer prompt caching: a long shared prefix (system prompt, tool definitions, few-shot examples, static context) is cached server-side and billed at a discount. Agentic loops are the ideal shape for this - the prefix stays stable across turns while only the tail changes.

Practical moves:

  • Order prompts static-to-dynamic: system prompt and tools first, conversation history last.
  • Avoid churning the prefix - inserting a timestamp at position 0 destroys the cache.
  • Confirm your gateway passes provider cache controls through (LiteLLM exposes cache_control breakpoints for Anthropic-style caching, and passes OpenAI/Bedrock cache directives).

Layer 4: Index refresh policy (the staleness contract)

Caching answers while your corpus changes underneath is the classic failure. Define refresh per content class:

Content classPolicy
Docs/FAQ (changes daily+)TTL 1–24h + event-driven invalidation on publish
Tickets/CRM (changes hourly)TTL 5–15min, or invalidate on write
Real-time (prices, inventory)No answer-level cache; cache only embeddings
Generated summariesKey = source doc hash + model version

Two implementation styles:

  1. Versioned keys. cache_key = {index_version}:{normalized_query}. On re-index, bump the version - every entry is invalidated atomically with zero eviction code.
  2. Event-driven eviction. A publish webhook or CDC stream deletes related keys (tag/namespace-based). Lower latency for targeted invalidation, more moving parts.

For agentic loops specifically: cache the plan, not just the answer. A successful retrieval plan (“search docs for X, then filter by version”) can be replayed with a fresh index in milliseconds - the expensive part was the LLM deciding the plan, not executing it.

Putting it together with LiteLLM

The cited gateway makes layering concrete - caching, guardrails, and spend tracking in one place:

import litellm

litellm.cache = litellm.Cache(
    type="redis",
    host="localhost",
    port=6379,
    namespace="agent-rag",
)

async def retrieve(query: str, tenant: str):
    # Application-level retrieval cache with explicit staleness control
    key = f"{tenant}:{index_version}:{normalize(query)}"
    if (cached := await redis.get(key)):
        return json.loads(cached)

    results = await vector_store.search(query, top_k=8)
    await redis.setex(key, ttl_seconds, json.dumps(results))
    return results

async def synthesize(query: str, chunks: list) -> str:
    resp = await litellm.acompletion(
        model="gpt-4o-mini",
        messages=[
            {"role": "system", "content": STATIC_RAG_PROMPT},  # cached prefix
            {"role": "user", "content": build_prompt(query, chunks)},
        ],
    )
    return resp.choices[0].message.content

Operating notes:

  • Measure hit rate by layer. A 60% query-cache hit rate with a 5% semantic false-hit rate is a problem; you need both numbers.
  • Tag cached entries with generation. When serving, log cache_hit, cache_layer, and index_version - you’ll need them for incident forensics.
  • Budget the fallback. When the cache is cold (new deployment, post-refresh), requests stampede the vector DB. Add jittered TTLs or a short request-coalescing window.

What NOT to cache

  • Answers containing per-request authorization (user-specific data) without the user in the key.
  • Negative results indefinitely - “no documents found” often means the index wasn’t ready.
  • Anything above the semantic threshold but with materially different freshness requirements (prices vs policies).

Key Takeaways

  1. Agentic RAG multiplies retrieval calls - caching each layer (query, embedding, search, prompt prefix) compounds savings.
  2. Semantic query caching needs tuned thresholds and tenant-scoped keys to avoid wrong-but-similar answers.
  3. Version your cache keys by index generation; refresh becomes an atomic invalidation instead of an eviction chore.
  4. Provider prompt caching rewards static-first prompt ordering - don’t churn the prefix.
  5. Instrument hit rate and staleness per layer; an unmeasured cache is a correctness liability.

References: