Skip to content
Blog

Parallel Tool Calls: Batching and Dependency Ordering in Agent Loops

How multi-tool turns actually execute in the OpenAI Agents SDK: model-side parallel_tool_calls vs SDK-side concurrency, capping fan-out, and patterns for enforcing dependency order safely.

Published on • October 3, 2026

AI Assistant

When a model emits search_inventory() and check_shipping_estimate() in the same turn, do they run at once — or one after the other? The answer depends on two different knobs, and confusing them is the most common source of both slow agents and surprise race conditions.

Sources: OpenAI Agents SDK — Running agents, OpenAI Agents SDK — Tools.

The loop first

The runner’s shape is simple: call the LLM → if the response contains tool calls, execute them, append results → call the LLM again → repeat until there’s a final answer, a handoff, or max_turns is hit. All ordering questions live inside that “execute them” step.

Two knobs, two layers

1. ModelSettings.parallel_tool_calls — decides whether the model is allowed to emit more than one tool call in a single response. This is a provider-side capability flag.

2. RunConfig.tool_execution.max_function_tool_concurrency — decides how the SDK executes the calls it received. From the docs:

By default (max_function_tool_concurrency=None), when a model emits multiple function tool calls in a turn, the SDK starts all emitted local function tool calls.

So the default is full fan-out: every local function tool in the turn starts concurrently. To cap it:

from agents import RunConfig, ToolExecutionConfig

run_config = RunConfig(
    tool_execution=ToolExecutionConfig(max_function_tool_concurrency=2),
)

If your tools are async, they interleave on the event loop — five aiohttp calls that each take 1 s finish in ~1 s total, not 5 s. If they’re def (sync) functions, the SDK still starts them together from its perspective, but blocking work inside a sync tool blocks the loop: make anything I/O-bound async or you’ll serialize by accident.

Why concurrency creates dependency bugs

Turn-level batching means the model can emit calls whose semantics only hold in a particular order:

# Emitted together in one turn — order between them is NOT guaranteed
book_flight(flight_id)      # writes
apply_coupon(coupon)        # assumes the booking exists

The model intended a sequence; the runtime delivered a set. Defenses, in increasing order of robustness:

  1. Make tools order-independent by design. get_price() and get_eta() can’t conflict. Prefer read-only, commutative tools and push mutations into one explicit step.
  2. Encode dependencies in data, not in prose. Return results with IDs and let the next turn’s tool call consume them, instead of hoping the model chains calls within a turn.
  3. Serialize mutations on the server. A QueuedInterceptor-style single-flight lock (or an idempotency key derived from the logical operation) makes duplicate/racy calls safe regardless of ordering.
  4. Cap concurrency when fan-out is useful but dangerous. max_function_tool_concurrency=2 keeps genuine batch reads fast while limiting blast radius.
  5. Disable model-side parallelism for strict workflows. Setting ModelSettings(parallel_tool_calls=False) asks the model to emit one call at a time — a blunt but effective ordering lever when tool semantics demand it.

Note that the SDK has no dependency graph or batching scheduler: ordering is model-driven. If you need guaranteed sequencing, put it in the tool layer or the workflow engine, not in hope.

Error handling inside a batch

One failing call in a parallel turn shouldn’t nuke the whole batch — and the SDK gives you levers:

  • Tools raise → the error is fed back to the model, which can retry or route around it.
  • RunConfig(tool_not_found_behavior=...) — default raises ModelBehaviorError; "return_error_to_model" keeps the run alive and lets the model recover.
  • tool_error_formatter customizes how that error is described to the model (verbose for debugging, terse to save tokens).
  • Per-call @function_tool(timeout=...) works for async handlers — the natural guard against one hung tool stalling a fan-out.

Design rule: failed reads degrade; failed writes stop. If charge_card() errors mid-batch, don’t let the model plow on to ship_order() — return an explicit “payment pending, escalate” result.

Batching for throughput

Beyond same-turn fan-out, batch explicitly when the workload is a set:

  • Pre-declare the batch in one tool (get_prices(sku_list)) instead of N round-trips — fewer tokens, fewer loop iterations.
  • Group by side-effect class: all reads in one turn, all writes in a later turn.
  • Bound the work per turn. Uncapped concurrency plus large batches is how you earn rate-limit errors; chunk and let the loop drive the pace.
  • For model-controlled branching at scale, the SDK’s experimental ProgrammaticToolCallingTool lets the model generate JavaScript in a hosted V8 sandbox that can loop and issue parallel tool calls — a different trade-off (the model writes the orchestration code instead of emitting calls).

A practical checklist

  1. Decide deliberately: parallel reads (default) or serialized turns (parallel_tool_calls=False).
  2. Set max_function_tool_concurrency based on your backend’s tolerance, not on optimism.
  3. Make I/O tools async so concurrency is real.
  4. Give every mutating tool an idempotency key; never rely on turn ordering for correctness.
  5. Add timeouts to async tools and errors the model can recover from.
  6. Watch your traces: if tool calls in one turn overlap less than expected, a sync tool or a low concurrency cap is silently serializing you.

Parallel tool calls are the agent equivalent of Promise.all: gloriously fast when the operations are independent, subtly wrong when they aren’t. Do the dependency audit once, up front, and the loop gets both quicker and safer.