Streaming vs. Polling: Response Patterns for Async Agent Clients
When to stream agent responses, when polling is the right answer, and how to combine them with server-sent events, WebSockets, and durable run state.
Published on • October 8, 2026
AI Assistant

Agent runs take seconds to minutes. Users will not wait behind a spinner for either. The client-side response pattern you choose — streaming, polling, or a hybrid — determines whether your product feels alive or broken.
Both patterns have real engineering costs. This post works through when each wins, and how to implement the hybrid that most production agent systems converge on.
The two patterns
Streaming — the server pushes updates as they happen over a long-lived connection: Server-Sent Events (SSE), WebSockets, or gRPC streaming.
Polling — the client repeatedly asks “done yet?” against a resource URL, typically with a run/job id.
Everything else (long-polling, WebSub, message queues with client subscriptions) is a variation on these two.
When streaming wins
Streaming is the right default for interactive, single-session agent work:
- Chat with token-by-token output — watching text arrive is part of the perceived speed
- Multi-step tool runs where the user should see “searching… reading… drafting…”
- Long generations where an early-abort or steer is possible
- Any UI where the first token matters more than the last
LangChain’s documentation makes streaming a core primitive of its agent harness rather than an add-on: the harness wraps the model loop and exposes streamed output as first-class, because rendering partial output is how agent UIs are built.
Streaming’s advantages:
- Perceived latency drops to time-to-first-token, not time-to-complete
- No request amplification — one connection, many updates
- Immediate failure signaling — a dropped connection is detectable instantly
- Natural fit for agent events — tool started/finished, state deltas, and text chunks are all just more events on the same stream
When polling wins
Polling is the right answer more often than streaming advocates admit:
- The work outlives the connection. A run that takes 10 minutes will not survive a mobile network switch, a laptop sleep, or an app backgrounded by the OS. iOS and Android aggressively kill long-lived sockets.
- Many clients per run. A run three teammates are watching shouldn’t hold three sockets.
- Serverless or proxied infrastructure. Some platforms and CDNs have hard limits on connection duration.
- Fire-and-forget jobs. Batch runs, exports, document generation — the user wants a status, not a feed.
- Offline-tolerant clients. Polling resumes naturally after a network gap; a broken stream does not.
Polling’s advantages:
- Simplicity. A
GET /runs/{id}endpoint is trivially cacheable, testable, and debuggable. - Resumability. Every poll returns complete current state — no event replay logic.
- Infrastructure honesty. Nothing pretends a connection is durable when it isn’t.
- Easy fan-out. Any number of clients can watch the same URL.
The costs of each
| Streaming | Polling | |
|---|---|---|
| Server load | One long-lived connection | N requests per interval per client |
| Latency to update | Immediate | Up to poll interval |
| Reconnect logic | Required (event ids, resume) | Implicit |
| Proxy/CDN issues | Buffering, timeouts | Cache staleness |
| Mobile backgrounding | Connection killed | Resumes cleanly |
| Ordering guarantees | Must be explicit | Trivial (full state) |
| Debugging | Needs event logs | Just look at the response |
Streaming’s hidden cost is reconnect semantics. If your events aren’t numbered and resumable, a single dropped packet means the client missed a tool result and will render a run that doesn’t make sense. Polling’s hidden cost is latency and load — a 1-second poll from 10,000 clients is 10,000 requests/second against a run table.
The hybrid that production converges on
Most mature agent systems end up with the same architecture:
Client
│
├─ POST /runs → create run, get run_id (202)
│
├─ GET /runs/{id}/events (SSE) → live updates while connected
│ Last-Event-ID: <resume>
│
└─ GET /runs/{id} → authoritative state, polled on
reconnect / background / cold start
The rules that make it work:
- Run state lives in a durable store, not in the connection. The stream is a view of state, never the source of truth.
- Every event has a monotonically increasing id. SSE’s
Last-Event-ID(or a WebSocket resume token) lets the client ask for “everything after N.” - The full-state endpoint is always correct. On reconnect, cold start, or app resume, fetch
GET /runs/{id}first, then attach to the stream from that sequence number. - Poll as a fallback, not as the primary. If SSE fails repeatedly for a client (corporate proxy, for example), degrade to 2s polling automatically.
- Exponential backoff on reconnect with jitter, capped — otherwise every client that dropped reconnects simultaneously.
This is the “state snapshot + incremental deltas” pattern: polling gives you snapshots, streaming gives you deltas, and the two keep each other honest.
Event design
Whichever you choose, define the event vocabulary once:
{"seq": 12, "type": "run.started", "ts": "2026-10-08T03:14:07Z"}
{"seq": 13, "type": "step.started", "step": "retrieve", "tool": "search_docs"}
{"seq": 14, "type": "step.finished", "step": "retrieve", "docs": 8, "ms": 412}
{"seq": 15, "type": "text.delta", "delta": "The pipeline "}
{"seq": 16, "type": "text.delta", "delta": "has three stages"}
{"seq": 17, "type": "run.finished", "usage": {"in": 4820, "out": 214}}
{"seq": 18, "type": "run.failed", "error": {"code": "tool_timeout"}}
Design rules:
seqis mandatory. It’s what makes reconnect safe and out-of-order detectable.- Include both
step.startedandstep.finished. Users need “it’s doing X” and “X took 412ms.” - Never put the only copy of important state in a delta. If a
text.deltais lost, the snapshot onGET /runs/{id}must still contain the final text. - Emit terminal events exactly once.
run.finishedandrun.failedare idempotent at the client — ignore duplicates. - Separate user-visible events from diagnostics. Not every internal step needs to reach the UI; log those instead.
Transport choice
SSE is the default for HTTP-friendly, server→client streaming: no new protocol, works through most proxies, automatic reconnect in EventSource, and text-oriented — ideal for token streams and JSON events. Its asymmetry is a feature: client→server still goes over normal POSTs.
WebSockets earn their cost when you need true bidirectionality — mid-run human steering, collaborative multi-client sessions, or very high event rates. You take on connection lifecycle management, auth-on-connect semantics, and proxy configuration yourself.
HTTP/2 server push is generally not what you want here; it’s for pushing resources, not application events.
gRPC streaming works well inside a service mesh but requires HTTP/2 end-to-end and is awkward from browsers.
Rule of thumb: start with SSE + a snapshot endpoint. Add WebSockets only when you have a concrete bidirectional requirement.
Mobile and offline
Mobile is where streaming-only designs fail:
- Backgrounding kills the socket
- Network transitions (Wi-Fi → LTE) drop it
- Users expect to close the app and reopen it to find the run
Treat mobile as a polling-first, streaming-when-foregrounded client. On foreground, sync via the snapshot endpoint, then attach to SSE. Persist seq locally so the app knows what it has seen.
For web, the same principle applies to tab suspension and bfcache restores: re-sync on visibilitychange rather than trusting a stale stream.
Server-side considerations
- Timeouts. Set idle timeouts on streams (e.g., send a heartbeat every 15s) so proxies don’t kill a quiet connection.
- Fan-out. If many clients watch one run, publish events once to a channel (Redis pub/sub, a message bus) and have each connection subscribe — don’t fan out from the run executor.
- Backpressure. A slow client must not block the run. Drop-to-snapshot: if a client’s buffer grows past a threshold, stop sending deltas and let it resync from the snapshot.
- Auth. Authorize on connect and periodically; a long-lived stream shouldn’t outlive the credential that opened it.
Observability
You can’t tune what you don’t measure. Track per pattern:
- Streaming: time-to-first-event, events/sec, reconnect rate, bytes buffered/dropped
- Polling: request rate, p95 poll→update latency, cache hit rate
- Both: end-to-end time from run completion to client render — the number users actually feel
If your polling p95 is 400ms and your stream reconnect rate is 15%, you have an infrastructure problem, not a pattern problem.
Decision summary
| Situation | Choose |
|---|---|
| Interactive chat, tokens matter | SSE streaming |
| Long runs, mobile clients | Poll + snapshot; stream only in foreground |
| Multiple viewers of one run | Poll, or fan-out stream from a bus |
| Human steering mid-run | WebSocket |
| Serverless/short request limits | Poll |
| Background batch jobs | Poll (or webhook) |
| Offline-tolerant desktop app | Hybrid: snapshot on launch, stream while open |
Wrapping up
Streaming and polling aren’t rivals — they’re complementary views over the same durable run state. Keep state in a store, make snapshots authoritative, give every event a sequence number, stream while it’s cheap, and poll when the world gets in the way. Systems that get this right feel instant on a good network and merely correct on a bad one — which is exactly the trade you want.
References
- LangChain documentation — agent harness, model streaming, and LangGraph integration
- LangGraph documentation — durable execution and state checkpoints
- MDN: Server-Sent Events