Skip to content
Blog

Multi-Hop Retrieval: Answering Questions That Span Multiple Sources

How multi-hop retrieval extends RAG with an iterative retrieve, reason, and re-query loop, and how to implement it as a bounded LangGraph workflow.

Published on • October 4, 2026

AI Assistant

The question single-shot RAG cannot answer

Classic RAG is a single hop: embed the user’s question, fetch the top-k chunks, stuff them into the prompt, generate. It works when one passage contains the whole answer, and fails predictably the moment the answer is assembled from several places.

Consider: “How did the 2024 acquisition of Bolt affect the shipping fees described in the partner pricing document?” Answering requires finding the acquisition announcement, extracting Bolt’s new terms, then locating the pricing document and applying those terms to its fee table. No single chunk holds both facts. A single-shot retriever returns whichever half scores highest against the raw question, and the model either answers from that half or hallucinates the missing bridge.

This is multi-hop retrieval: an iterative loop of retrieve → reason → re-query in which each pass produces findings that reshape the next query. The technique predates modern frameworks — multi-hop QA datasets like HotpotQA were built to expose single-hop failure — but it became practical to ship once graph runtimes gave us somewhere to put the loop.

Why one-shot retrieval breaks down

Three failure modes show up consistently:

Insufficient context. The needed evidence is split across documents, or the linking fact (“Bolt acquired X”) lives in a document that shares no vocabulary with the question about shipping fees. Vector similarity ranks it low because similarity is not entailment.

Lost-in-the-middle. Long-context models do not read uniformly: accuracy on evidence parked in the middle of a large window degrades relative to evidence at the start or end. Retrieving eight documents and hoping is not a strategy; you pay for context you will not use reliably.

No chance to correct course. A single retrieval is committed before the model has reasoned at all. If the query was ambiguous or the top-k is dominated by a distractor, nothing in the pipeline can notice. Retrieval quality is downstream of understanding, and understanding takes a first look at the data.

Multi-hop fixes this by making retrieval a decision rather than a precondition.

The loop: retrieve, reason, decide, re-query

The cycle is small enough to hold in your head:

  1. Retrieve — run the current query (initially the user’s question) against the vector index.
  2. Reason — ask the model: given everything gathered so far, what have we learned, and what is still missing?
  3. Decide — is the evidence sufficient to answer? If not, and the hop budget is not exhausted, continue. Otherwise synthesize.
  4. Rewrite — generate a new query targeted at the gap: a sub-question, a named entity, or a reformulation that terms the missing fact in retrievable language.

Each iteration appends to a growing set of gathered facts, each tagged with its source document. The final answer is synthesized only from that set, which is what makes citation anchoring tractable.

Two query strategies dominate in practice:

  • Decomposition — split the original question into explicit sub-questions (“Who acquired Bolt?” → “What are Bolt’s 2026 terms?” → “How do those alter the fee table?”) and answer them in order. Deterministic and easy to debug, but brittle when the decomposition misses a dependency.
  • Rewrite / step-back — keep one query and let the model rewrite it each hop, optionally abstracting to a higher-level question first. More flexible, harder to evaluate.

Many production systems combine both: decompose once, then rewrite within each branch.

Implementing it in LangGraph

LangGraph is a low-level orchestration runtime for stateful workflows, and it fits here because multi-hop retrieval is a state machine with a conditional cycle. Its headline strengths — mixing deterministic with model-driven steps, durable execution, persistence, human-in-the-loop interrupts — map onto the pattern: retrieve is deterministic code, reason is an LLM call, the conditional edge is plain Python.

Model the shared state first:

from typing import TypedDict, Annotated, Optional
from operator import add as merge
from langgraph.graph import StateGraph, START, END

class HopState(TypedDict):
    question: str                 # original user question
    query: str                    # current retrieval query
    docs: list[str]               # raw chunks from the last retrieve
    facts: Annotated[list[str], merge]   # accumulated, cited findings
    hops: int                     # iterations used so far
    max_hops: int                 # hard budget
    answer: Optional[str]

Annotated[list[str], merge] gives you reducer-based accumulation: each node returns only the new facts and LangGraph merges them, so nodes stay independent and the state history stays inspectable.

Now the nodes:

def retrieve(state: HopState):
    docs = retriever.invoke(state["query"])
    return {"docs": [d.page_content for d in docs]}

def reason(state: HopState):
    prompt = (
        f"Question: {state['question']}\n"
        f"Known facts: {state['facts']}\n"
        f"Retrieved: {state['docs']}\n"
        "List supported new facts with source ids, then what is missing."
    )
    return {"facts": llm.invoke(prompt), "hops": state["hops"] + 1}

def rewrite(state: HopState) -> dict:
    query = llm.invoke(
        f"Rewrite to retrieve the missing evidence: {state['question']}"
    )
    return {"query": query}

def should_continue(state: HopState) -> str:
    if state["hops"] >= state["max_hops"]:
        return "answer"
    return "continue" if missing_evidence(state) else "answer"

def synthesize(state: HopState) -> dict:
    return {"answer": llm.invoke(
        f"Answer using only these facts, citing sources: {state['facts']}"
    )}

Wire it into a graph with a bounded loop — the conditional edge is what stops this from becoming an unbounded agent run:

b = StateGraph(HopState)
b.add_node("retrieve", retrieve)
b.add_node("reason", reason)
b.add_node("rewrite", rewrite)
b.add_node("synthesize", synthesize)

b.add_edge(START, "retrieve")
b.add_edge("retrieve", "reason")
b.add_conditional_edges("reason", should_continue, {
    "continue": "rewrite",
    "answer": "synthesize",
})
b.add_edge("rewrite", "retrieve")
b.add_edge("synthesize", END)

graph = b.compile()
graph.invoke({"question": q, "query": q, "facts": [],
              "docs": [], "hops": 0, "max_hops": 3, "answer": None})

Three details matter more than they look:

  • max_hops is the cost ceiling. Set it to 2 or 3 for most corpora; each hop costs embedding queries plus a full LLM call.
  • Distinguish “not found” from “not needed.” Make missing_evidence an LLM judgment with structured yes/no output, not a substring test on the gathered facts.
  • Persist state. LangGraph’s checkpointer lets a run survive a crash or pause for human review between hops — useful when a retrieval returns nothing and you would rather ask a clarifying question than guess.

If you do not need custom control flow, LangGraph also offers prebuilt agent loops; multi-hop retrieval is precisely the case where the extra control earns its keep.

When not to use multi-hop

More hops is not more better. Skip the loop when:

  • The corpus fits in the window. If your documentation set is under ~100k tokens, retrieve generously or stuff it all in and skip the retrieval complexity.
  • Questions are single-hop. Fact lookups and “what does setting X do” questions get no lift from iteration and only pay the tax.
  • Latency or cost is binding. Expect roughly N retrievals and N LLM calls instead of one of each — easily several seconds and an order of magnitude more spend per query.
  • You have no eval harness. A loop you cannot measure is a loop that silently degrades after a prompt tweak.

A practical rule: ship single-shot RAG, measure where it fails on a labelled set, and only introduce hops for genuinely compositional failures.

Evaluating multi-hop retrieval

Single “did it answer correctly” scoring hides where the pipeline broke. Break metrics down per hop:

  • Retrieval recall@k per hop — did hop i return the document needed for sub-question i? A miss is a retrieval problem, not a generation problem.
  • Intermediate fact correctness — are the accumulated facts supported by their cited chunks? The highest-signal metric: garbage facts propagate to every later hop.
  • Answer correctness and faithfulness — the final synthesis scored against a reference answer.
  • Hop efficiency — the distribution of hops used. If 80% of queries hit max_hops, your budget or rewrite prompt is wrong.

Citation anchoring ties it together: store the document ID and chunk offsets with each fact rather than letting the model free-form cite. The final answer then renders as claims mapped to retrievable spans, and unanchored claims can be flagged before the response reaches a user. LangSmith makes the trace of every node — each query, fact, and conditional branch — inspectable, which is the difference between debugging a loop and staring at one.

Conclusion

Multi-hop retrieval trades a bounded amount of extra compute for the ability to answer compositional questions correctly. The winning shape is unglamorous: explicit state, a small number of well-defined nodes, one conditional edge, a hard hop budget, and per-hop evaluation. Start with single-shot RAG, measure, and add the loop only where the evidence demands it.

Sources