Skip to content
Blog

Review Queues: Human Triage of Agent-Produced Work

LlamaIndex Workflows implement review queues with InputRequiredEvent and HumanResponseEvent — pause, persist, and resume agent runs for human triage.

Published on • October 7, 2026

AI Assistant

An agent that drafts an email, merges a PR, or files a refund shouldn’t just fire and forget. Somebody has to look at the output — and the hard part isn’t the approval button, it’s the pause. A review queue is a system that stops a run at a decision point, hands the artifact to a human, and later resumes the exact same execution with the human’s verdict.

LlamaIndex Workflows implements this with an event pair. The docs put it plainly: “Human-in-the-loop workflows need to pause, tell the caller what input is needed, and continue when the caller sends a response. Workflows support that with normal events.”

The Event Pair

A step that needs human input returns an InputRequiredEvent. A later step consumes a HumanResponseEvent. Between them, the workflow is suspended:

from workflows import Workflow, step
from workflows.events import StartEvent, StopEvent, InputRequiredEvent, HumanResponseEvent

class NumberWorkflow(Workflow):
    @step
    async def ask(self, ev: StartEvent) -> InputRequiredEvent:
        return InputRequiredEvent(prefix="Enter a number: ")

    @step
    async def answer(self, ev: HumanResponseEvent) -> StopEvent:
        return StopEvent(result=ev.response)

The caller watches the event stream and replies into the same handler:

workflow = NumberWorkflow()
handler = workflow.run()

async for event in handler.stream_events():
    if isinstance(event, InputRequiredEvent):
        response = input(event.prefix)   # input(), websocket reply, web form...
        await handler.send_event(HumanResponseEvent(response=response))

final_result = await handler

The transport is deliberately unspecified — input() for a script, a WebSocket frame for a chat UI, a form POST for a dashboard. The workflow doesn’t care; it only sees HumanResponseEvent.

Events can be subclassed for structured payloads, which is what you want for a real review: instead of a bare string, pass the draft body, a diff, a risk score, and the allowed verdicts.

class CodeReviewRequested(InputRequiredEvent):
    pr_title: str
    diff: str
    risk_level: str

class CodeReviewResponse(HumanResponseEvent):
    approved: bool
    comments: str = ""

Building an Actual Review Queue

The simple pattern assumes the human answers immediately. A real review queue can’t: reviewers are busy, requests arrive overnight, and the process hosting the workflow may restart. The docs show how to stop between responses by snapshotting the run’s context:

handler = workflow.run()

async for event in handler.stream_events():
    if isinstance(event, InputRequiredEvent):
        # Persist the run so it can be picked up later — by another
        # process, another server, or tomorrow morning.
        await db.save("run-123", json.dumps(handler.ctx.to_dict()))
        await handler.cancel_run()
        break

# ... hours later, when the reviewer responds ...

ctx_dict = json.loads(await db.load("run-123"))
restored_ctx = Context.from_dict(workflow, ctx_dict)
handler = workflow.run(ctx=restored_ctx)
await handler.send_event(HumanResponseEvent(response=response))

Three details matter:

  1. Serialize handler.ctx.to_dict() — that’s the entire suspended execution: which step is waiting, what it already computed, where it’s headed.
  2. Cancel the original handler after snapshotting. Otherwise the in-memory run sits waiting for a response that’s now going to the restored run.
  3. Restore with Context.from_dict(workflow, ...) and re-enter through workflow.run(ctx=...), then deliver the response.

This is what turns a demo into a queue: the run is now a durable row in a database, the review request is a row in a work list, and the response is a message that revives it. Multiple reviewers, priorities, and SLAs are application logic on top.

The Single-Step Alternative: wait_for_event()

If the waiting happens inside one step rather than across steps, use ctx.wait_for_event():

@step
async def ask_user(self, ctx: Context, ev: StartEvent) -> StopEvent:
    response = await ctx.wait_for_event(
        HumanResponseEvent,
        waiter_event=InputRequiredEvent(prefix="Enter a number: "),
        waiter_id="get_number",
    )
    return StopEvent(result=response.response)

Supporting details:

  • waiter_id distinguishes multiple waits within the same step.
  • requirements={"request_id": ...} routes a response to the right waiter when several are outstanding.

The caveat the docs flag: wait_for_event replays all preceding code each time the step resumes. Everything before the call must be repeatable — idempotent side effects, no non-deterministic ordering. For anything with external effects, prefer the separate-steps event approach, which resumes at a step boundary instead.

Where Review Queues Fit in the Workflow Toolkit

LlamaIndex positions Workflows as “event-driven, multi-step processes that combine agents, data connectors and tools, with branching, retries and human-in-the-loop review.” The review step is one specialization of primitives already in the box:

  • Branching — route to auto_approve or human_review based on risk score.
  • Retries — a rejected submission re-enters the draft step with reviewer comments as context.
  • Streaming events — drive a live queue UI from handler.stream_events().
  • Durable workflows — DBOS-backed execution so pending reviews survive deploys.
  • Error handling (retry_steps) — a reviewer who never responds doesn’t wedge the system; escalation steps handle timeouts.

Design Lessons for Any Framework

Even if you’re not on LlamaIndex, the pattern transfers:

  1. Represent “needs a human” as an explicit event type, not as a blocking call. Events serialize; thread stacks don’t.
  2. Snapshot before waiting. Persist the execution context at the pause point so the queue is durable across processes and restarts.
  3. Cancel the in-flight handler when you hand off to a restored run — one response, one consumer.
  4. Prefer step boundaries over mid-step waits when the preceding code has side effects; mid-step waits replay code.
  5. Make the response structured. approved: bool plus free-text comments beats a raw string — it lets routing logic branch without parsing prose.

The approval UI is the easy 10%. The pause, persist, and resume machinery is what makes agent-produced work safe to actually ship.

Further Reading