Progress Visibility: Streaming Agent Steps and Thoughts to Users
Why silent multi-second agent runs destroy UX, and how to stream step lifecycles, tool calls, and reasoning to the client with the OpenAI Agents SDK, SSE, and MCP-style progress notifications.
Published on • October 4, 2026
AI Assistant

Progress Visibility: Streaming Agent Steps and Thoughts to Users
An agent that takes eleven seconds to answer is not a broken product. An agent that takes eleven seconds and shows nothing is. Users do not know whether the request was sent, whether the model is thinking, whether a tool is running, or whether the tab should be closed. Multi-second silence reads as failure, and the reflex is to hit refresh — which starts a second run, doubles your cost, and leaves the first one orphaned.
The fix is not just token streaming. Tokens tell you the answer is arriving; they do not tell you that the agent planned, called a tool, waited on it, and is now drafting. Progress visibility means surfacing the step lifecycle — planning, tool call, tool result, final answer — while the run is still in flight.
Key technologies: OpenAI Agents SDK for Python (Runner.run_streamed, RunResultStreaming, StreamEvent), Server-Sent Events, Model Context Protocol JSON-RPC notifications (notifications/progress), and aria-live regions.
Why a Silent Run Feels Broken
Interface latency research has long shown that perceived wait time depends on feedback, not duration. A spinner is the weakest form of feedback because it carries no information: it looks identical whether the agent is waiting on a model, retrying a failed tool, or wedged. A step timeline answers the questions users actually have — what is it doing now, what did it already do, how much is left — and converts dead time into visible work.
There is a second reason: step events are where errors surface. A tool that fails at second eight of a ten-second run is diagnosable only if the client saw tool_called with no matching tool_output. Silent runs hide the failure and its cause in the same black box.
Two Streams, Not One: Tokens vs Steps
The OpenAI Agents SDK draws a clean line between two event families, and conflating them is the most common design mistake.
from agents import Agent, Runner
result = Runner.run_streamed(agent, input="Please tell me 5 jokes.")
async for event in result.stream_events():
if event.type == "raw_response_event":
# Token-level: wrap raw LLM events such as
# response.created or response.output_text.delta
...
elif event.type == "run_item_stream_event":
# Step-level: an item has been fully generated
...
elif event.type == "agent_updated_stream_event":
# The current agent changed (e.g. a handoff)
print(f"Now running: {event.new_agent.name}")
- Raw response events (
RawResponsesStreamEvent) wrap events straight from the OpenAI Responses API, includingresponse.output_text.delta. They are the right source for a typewriter effect. - Run item events (
RunItemStreamEvent) fire when an item is complete — “message generated”, “tool ran” — which is the right granularity for a progress timeline. - Agent updated events (
AgentUpdatedStreamEvent) tell you the active agent changed after a handoff.
Rendering tokens as your only progress signal collapses these layers. A tool call produces no text deltas at all: for two seconds the token stream is empty while the agent is doing its most interesting work.
Important semantics to honor: a streamed run is not complete when the last token appears. Keep consuming result.stream_events() until the iterator ends — session persistence, approval bookkeeping, and history compaction can all finish afterwards. result.is_complete reflects the final state once the loop exits.
An Event Schema You Can Render
RunItemStreamEvent.name uses a fixed set of semantic names, which makes it a stable contract for UI state machines:
| Event type | Payload signal | Renders as |
|---|---|---|
raw_response_event | event.data is a Responses API event (response.output_text.delta) | Token-by-token text |
run_item_stream_event | name: message_output_created | Final answer bubble |
run_item_stream_event | name: tool_called | Tool chip in “running” state |
run_item_stream_event | name: tool_output | Tool chip resolves with result |
run_item_stream_event | name: reasoning_item_created | Collapsed “thought process” row |
run_item_stream_event | name: handoff_requested / handoff_occured | ”Switching to specialist” transition |
run_item_stream_event | name: mcp_approval_requested | Blocking approval prompt |
agent_updated_stream_event | event.new_agent | Header/agent label update |
(handoff_occured is intentionally misspelled in the SDK for backward compatibility.)
The item payload itself carries a type you can switch on:
from agents import ItemHelpers
async for event in result.stream_events():
if event.type != "run_item_stream_event":
continue
if event.item.type == "tool_call_item":
ui.add_tool_chip(status="running")
elif event.item.type == "tool_call_output_item":
ui.add_tool_chip(status="done", output=event.item.output)
elif event.item.type == "message_output_item":
ui.set_answer(ItemHelpers.text_message_output(event.item))
If you render partial text from raw deltas, do not then treat the completed message_output_item as a new message — replace the partial buffer with the final item to avoid duplicated paragraphs.
Carrying Progress Over the Wire
For a browser client, Server-Sent Events is the pragmatic transport: one-way, automatic reconnect, works through plain HTTP. The server side is a thin adapter from agent events to named SSE events:
async def event_stream(request, agent, user_input):
result = Runner.run_streamed(agent, input=user_input)
seq = 0
async for event in result.stream_events():
if event.type == "raw_response_event":
payload = {"type": "text.delta", "data": getattr(event.data, "delta", "")}
elif event.type == "run_item_stream_event":
payload = {"type": "step", "name": event.name, "item": event.item.type}
elif event.type == "agent_updated_stream_event":
payload = {"type": "agent", "name": event.new_agent.name}
else:
continue
seq += 1
yield f"id: {seq}\nevent: {payload['type']}\ndata: {json.dumps(payload)}\n\n"
yield f"event: done\ndata: {json.dumps({'complete': result.is_complete})}\n\n"
The browser side is equally small:
const source = new EventSource("/api/run?input=" + encodeURIComponent(q));
source.addEventListener("step", (e) => {
const step = JSON.parse((e as MessageEvent).data);
timeline.upsert({ id: `${step.name}-${step.item}`, ...step });
});
source.addEventListener("done", () => source.close());
MCP as a reference for progress semantics
The Model Context Protocol shows what a standardized progress channel looks like. MCP runs over JSON-RPC 2.0 with two transports: stdio for local processes and Streamable HTTP (HTTP POST plus optional Server-Sent Events) for remote servers.
When a client wants progress for a long-running request, it attaches a progressToken in the request metadata:
{"jsonrpc": "2.0", "id": 1, "method": "tools/call",
"params": {"_meta": {"progressToken": "abc123"}, "name": "reindex"}}
The server may then push notifications with no response expected (JSON-RPC notifications carry no id):
{"jsonrpc": "2.0", "method": "notifications/progress",
"params": {"progressToken": "abc123", "progress": 50, "total": 100,
"message": "Reticulating splines..."}}
The rules are the rules your UI needs: progress must increase even when total is unknown, total may be omitted, message should be human-readable, notifications must stop after completion, and both sides should rate-limit. MCP also states plainly that notifications are best-effort across reconnects — clients should poll to preserve freshness.
UI Patterns That Make Progress Legible
- Step timeline. A vertical list of completed and in-flight steps with timestamps. Completed steps stay visible; the run reads as a story, not a spinner.
- Tool-call chips. Compact pills:
search_docsrunning, done in 420 ms, 3 results. Failed chips stay red and expandable instead of disappearing. - Reasoning disclosure.
reasoning_item_createdcontent in a collapsed<details>row — available for the curious, invisible to everyone else. - Optimistic partial renders. Show token deltas into a “draft” bubble immediately, then swap in the authoritative
message_output_item. - Skeleton states. Before the first event arrives, render the expected shape of the answer rather than an empty panel.
- Elapsed time and cancellation. Show seconds ticking and a stop button wired to
result.cancel(), orresult.cancel(mode="after_turn")when you want the current turn to finish cleanly.
Out-of-Order Events, Buffering, and Backpressure
Networks reorder, retries duplicate, and browsers render slower than models emit tokens. Three defenses:
- Sequence and dedupe. Give every event a monotonic
seq(or use the SSEid:field). Apply only events withseq > lastApplied, and key timeline items by a stable ID so a retriedtool_calledupdates rather than duplicates. - Buffer and coalesce. Append deltas into a string buffer and flush to the DOM on a frame clock (roughly 30–60 fps), not per token. Ten thousand tiny
textContentwrites will make the UI thread the bottleneck of an otherwise fast run. - Backpressure by design. Never
awaita render inside the read loop. Push into a queue the render loop drains, and past a cap drop intermediate deltas while keeping the latest snapshot. On the server, flush SSE frames as you go and disable response buffering — a proxy that buffers until the run ends reintroduces the silent run you set out to eliminate.
Accessibility: Live Regions That Help
Streaming text is hostile to screen readers if handled naively. Put one polite live region on the status line — “Searching documentation”, “Running tool: reindex”, “Answer complete” via role="status" — and leave the answer body itself out of it. Announcing every token turns a screen reader into noise.
Use aria-live="polite" for step transitions; reserve assertive announcements for genuine errors and approval requests. Mark decorative spinners aria-hidden="true". Keep the timeline as a real list, and respect prefers-reduced-motion by dropping typing animations.
When the Client Drops: Graceful Degradation
SSE reconnects automatically, but a reconnect is a gap: the client missed events, and MCP’s warning about best-effort delivery applies to your transport too.
- Send every event with an
id:and supportLast-Event-IDon reconnect so the server can replay from that sequence. - Make events idempotent, so replayed history reconstructs state instead of corrupting it.
- If replay is impossible, expose a
GET /runs/{id}snapshot endpoint that returns current step list plus final output, and reconcile after reconnect. - Persist the run server-side regardless of the connection. A disconnected client must be able to come back and see the answer, not a dead spinner.
- Never treat “stream closed” as “run finished”. The final frame should say so explicitly, and the client should fall back to polling the snapshot endpoint if
donenever arrives.
Best Practices
- Stream steps, not just tokens — tool calls are where seconds go and where tokens go silent.
- Define a small, stable event schema (
text.delta,step,agent,error,done) and version it from the first release. - Keep the run alive to completion: drain the event iterator, then report
is_complete. - Include a monotonic sequence and stable item IDs in every event.
- Coalesce renders on a frame clock; never render per token.
- Treat notifications as lossy: support replay, snapshot fallback, and polling.
- Announce step transitions politely; never announce tokens.
- Give users elapsed time, a cancel path, and errors that name the failed step.
Wrapping Up
Silent agent runs fail users at exactly the moment the system is working hardest. The Agents SDK already emits the lifecycle you need — raw deltas for texture, run item events for structure, agent updates for handoffs — and MCP shows the discipline a progress channel should have: monotonic progress, human-readable messages, explicit completion, honest best-effort semantics. Wire those into a timeline, buffer aggressively, dedupe by ID, announce sparingly, and plan for the connection to drop. Progress visibility is the difference between a system that looks intelligent and one that looks stuck.