Skip to content
Blog

Streaming vs. Polling: Response Patterns for Async Agent Clients

When to stream agent responses, when polling is the right answer, and how to combine them with server-sent events, WebSockets, and durable run state.

Published on • October 8, 2026

AI Assistant

Agent runs take seconds to minutes. Users will not wait behind a spinner for either. The client-side response pattern you choose — streaming, polling, or a hybrid — determines whether your product feels alive or broken.

Both patterns have real engineering costs. This post works through when each wins, and how to implement the hybrid that most production agent systems converge on.

The two patterns

Streaming — the server pushes updates as they happen over a long-lived connection: Server-Sent Events (SSE), WebSockets, or gRPC streaming.

Polling — the client repeatedly asks “done yet?” against a resource URL, typically with a run/job id.

Everything else (long-polling, WebSub, message queues with client subscriptions) is a variation on these two.

When streaming wins

Streaming is the right default for interactive, single-session agent work:

  • Chat with token-by-token output — watching text arrive is part of the perceived speed
  • Multi-step tool runs where the user should see “searching… reading… drafting…”
  • Long generations where an early-abort or steer is possible
  • Any UI where the first token matters more than the last

LangChain’s documentation makes streaming a core primitive of its agent harness rather than an add-on: the harness wraps the model loop and exposes streamed output as first-class, because rendering partial output is how agent UIs are built.

Streaming’s advantages:

  • Perceived latency drops to time-to-first-token, not time-to-complete
  • No request amplification — one connection, many updates
  • Immediate failure signaling — a dropped connection is detectable instantly
  • Natural fit for agent events — tool started/finished, state deltas, and text chunks are all just more events on the same stream

When polling wins

Polling is the right answer more often than streaming advocates admit:

  • The work outlives the connection. A run that takes 10 minutes will not survive a mobile network switch, a laptop sleep, or an app backgrounded by the OS. iOS and Android aggressively kill long-lived sockets.
  • Many clients per run. A run three teammates are watching shouldn’t hold three sockets.
  • Serverless or proxied infrastructure. Some platforms and CDNs have hard limits on connection duration.
  • Fire-and-forget jobs. Batch runs, exports, document generation — the user wants a status, not a feed.
  • Offline-tolerant clients. Polling resumes naturally after a network gap; a broken stream does not.

Polling’s advantages:

  • Simplicity. A GET /runs/{id} endpoint is trivially cacheable, testable, and debuggable.
  • Resumability. Every poll returns complete current state — no event replay logic.
  • Infrastructure honesty. Nothing pretends a connection is durable when it isn’t.
  • Easy fan-out. Any number of clients can watch the same URL.

The costs of each

StreamingPolling
Server loadOne long-lived connectionN requests per interval per client
Latency to updateImmediateUp to poll interval
Reconnect logicRequired (event ids, resume)Implicit
Proxy/CDN issuesBuffering, timeoutsCache staleness
Mobile backgroundingConnection killedResumes cleanly
Ordering guaranteesMust be explicitTrivial (full state)
DebuggingNeeds event logsJust look at the response

Streaming’s hidden cost is reconnect semantics. If your events aren’t numbered and resumable, a single dropped packet means the client missed a tool result and will render a run that doesn’t make sense. Polling’s hidden cost is latency and load — a 1-second poll from 10,000 clients is 10,000 requests/second against a run table.

The hybrid that production converges on

Most mature agent systems end up with the same architecture:

Client
  │
  ├─ POST /runs            → create run, get run_id (202)
  │
  ├─ GET  /runs/{id}/events (SSE)   → live updates while connected
  │     Last-Event-ID: <resume>
  │
  └─ GET  /runs/{id}       → authoritative state, polled on
                              reconnect / background / cold start

The rules that make it work:

  1. Run state lives in a durable store, not in the connection. The stream is a view of state, never the source of truth.
  2. Every event has a monotonically increasing id. SSE’s Last-Event-ID (or a WebSocket resume token) lets the client ask for “everything after N.”
  3. The full-state endpoint is always correct. On reconnect, cold start, or app resume, fetch GET /runs/{id} first, then attach to the stream from that sequence number.
  4. Poll as a fallback, not as the primary. If SSE fails repeatedly for a client (corporate proxy, for example), degrade to 2s polling automatically.
  5. Exponential backoff on reconnect with jitter, capped — otherwise every client that dropped reconnects simultaneously.

This is the “state snapshot + incremental deltas” pattern: polling gives you snapshots, streaming gives you deltas, and the two keep each other honest.

Event design

Whichever you choose, define the event vocabulary once:

{"seq": 12, "type": "run.started",   "ts": "2026-10-08T03:14:07Z"}
{"seq": 13, "type": "step.started",  "step": "retrieve", "tool": "search_docs"}
{"seq": 14, "type": "step.finished", "step": "retrieve", "docs": 8, "ms": 412}
{"seq": 15, "type": "text.delta",    "delta": "The pipeline "}
{"seq": 16, "type": "text.delta",    "delta": "has three stages"}
{"seq": 17, "type": "run.finished",  "usage": {"in": 4820, "out": 214}}
{"seq": 18, "type": "run.failed",    "error": {"code": "tool_timeout"}}

Design rules:

  • seq is mandatory. It’s what makes reconnect safe and out-of-order detectable.
  • Include both step.started and step.finished. Users need “it’s doing X” and “X took 412ms.”
  • Never put the only copy of important state in a delta. If a text.delta is lost, the snapshot on GET /runs/{id} must still contain the final text.
  • Emit terminal events exactly once. run.finished and run.failed are idempotent at the client — ignore duplicates.
  • Separate user-visible events from diagnostics. Not every internal step needs to reach the UI; log those instead.

Transport choice

SSE is the default for HTTP-friendly, server→client streaming: no new protocol, works through most proxies, automatic reconnect in EventSource, and text-oriented — ideal for token streams and JSON events. Its asymmetry is a feature: client→server still goes over normal POSTs.

WebSockets earn their cost when you need true bidirectionality — mid-run human steering, collaborative multi-client sessions, or very high event rates. You take on connection lifecycle management, auth-on-connect semantics, and proxy configuration yourself.

HTTP/2 server push is generally not what you want here; it’s for pushing resources, not application events.

gRPC streaming works well inside a service mesh but requires HTTP/2 end-to-end and is awkward from browsers.

Rule of thumb: start with SSE + a snapshot endpoint. Add WebSockets only when you have a concrete bidirectional requirement.

Mobile and offline

Mobile is where streaming-only designs fail:

  • Backgrounding kills the socket
  • Network transitions (Wi-Fi → LTE) drop it
  • Users expect to close the app and reopen it to find the run

Treat mobile as a polling-first, streaming-when-foregrounded client. On foreground, sync via the snapshot endpoint, then attach to SSE. Persist seq locally so the app knows what it has seen.

For web, the same principle applies to tab suspension and bfcache restores: re-sync on visibilitychange rather than trusting a stale stream.

Server-side considerations

  • Timeouts. Set idle timeouts on streams (e.g., send a heartbeat every 15s) so proxies don’t kill a quiet connection.
  • Fan-out. If many clients watch one run, publish events once to a channel (Redis pub/sub, a message bus) and have each connection subscribe — don’t fan out from the run executor.
  • Backpressure. A slow client must not block the run. Drop-to-snapshot: if a client’s buffer grows past a threshold, stop sending deltas and let it resync from the snapshot.
  • Auth. Authorize on connect and periodically; a long-lived stream shouldn’t outlive the credential that opened it.

Observability

You can’t tune what you don’t measure. Track per pattern:

  • Streaming: time-to-first-event, events/sec, reconnect rate, bytes buffered/dropped
  • Polling: request rate, p95 poll→update latency, cache hit rate
  • Both: end-to-end time from run completion to client render — the number users actually feel

If your polling p95 is 400ms and your stream reconnect rate is 15%, you have an infrastructure problem, not a pattern problem.

Decision summary

SituationChoose
Interactive chat, tokens matterSSE streaming
Long runs, mobile clientsPoll + snapshot; stream only in foreground
Multiple viewers of one runPoll, or fan-out stream from a bus
Human steering mid-runWebSocket
Serverless/short request limitsPoll
Background batch jobsPoll (or webhook)
Offline-tolerant desktop appHybrid: snapshot on launch, stream while open

Wrapping up

Streaming and polling aren’t rivals — they’re complementary views over the same durable run state. Keep state in a store, make snapshots authoritative, give every event a sequence number, stream while it’s cheap, and poll when the world gets in the way. Systems that get this right feel instant on a good network and merely correct on a bad one — which is exactly the trade you want.

References