Parallel Tool Calls: Batching and Dependency Ordering in Agent Loops
How multi-tool turns actually execute in the OpenAI Agents SDK: model-side parallel_tool_calls vs SDK-side concurrency, capping fan-out, and patterns for enforcing dependency order safely.
Published on • October 3, 2026
AI Assistant

When a model emits search_inventory() and check_shipping_estimate() in the same turn, do they run at once — or one after the other? The answer depends on two different knobs, and confusing them is the most common source of both slow agents and surprise race conditions.
Sources: OpenAI Agents SDK — Running agents, OpenAI Agents SDK — Tools.
The loop first
The runner’s shape is simple: call the LLM → if the response contains tool calls, execute them, append results → call the LLM again → repeat until there’s a final answer, a handoff, or max_turns is hit. All ordering questions live inside that “execute them” step.
Two knobs, two layers
1. ModelSettings.parallel_tool_calls — decides whether the model is allowed to emit more than one tool call in a single response. This is a provider-side capability flag.
2. RunConfig.tool_execution.max_function_tool_concurrency — decides how the SDK executes the calls it received. From the docs:
By default (
max_function_tool_concurrency=None), when a model emits multiple function tool calls in a turn, the SDK starts all emitted local function tool calls.
So the default is full fan-out: every local function tool in the turn starts concurrently. To cap it:
from agents import RunConfig, ToolExecutionConfig
run_config = RunConfig(
tool_execution=ToolExecutionConfig(max_function_tool_concurrency=2),
)
If your tools are async, they interleave on the event loop — five aiohttp calls that each take 1 s finish in ~1 s total, not 5 s. If they’re def (sync) functions, the SDK still starts them together from its perspective, but blocking work inside a sync tool blocks the loop: make anything I/O-bound async or you’ll serialize by accident.
Why concurrency creates dependency bugs
Turn-level batching means the model can emit calls whose semantics only hold in a particular order:
# Emitted together in one turn — order between them is NOT guaranteed
book_flight(flight_id) # writes
apply_coupon(coupon) # assumes the booking exists
The model intended a sequence; the runtime delivered a set. Defenses, in increasing order of robustness:
- Make tools order-independent by design.
get_price()andget_eta()can’t conflict. Prefer read-only, commutative tools and push mutations into one explicit step. - Encode dependencies in data, not in prose. Return results with IDs and let the next turn’s tool call consume them, instead of hoping the model chains calls within a turn.
- Serialize mutations on the server. A
QueuedInterceptor-style single-flight lock (or an idempotency key derived from the logical operation) makes duplicate/racy calls safe regardless of ordering. - Cap concurrency when fan-out is useful but dangerous.
max_function_tool_concurrency=2keeps genuine batch reads fast while limiting blast radius. - Disable model-side parallelism for strict workflows. Setting
ModelSettings(parallel_tool_calls=False)asks the model to emit one call at a time — a blunt but effective ordering lever when tool semantics demand it.
Note that the SDK has no dependency graph or batching scheduler: ordering is model-driven. If you need guaranteed sequencing, put it in the tool layer or the workflow engine, not in hope.
Error handling inside a batch
One failing call in a parallel turn shouldn’t nuke the whole batch — and the SDK gives you levers:
- Tools raise → the error is fed back to the model, which can retry or route around it.
RunConfig(tool_not_found_behavior=...)— default raisesModelBehaviorError;"return_error_to_model"keeps the run alive and lets the model recover.tool_error_formattercustomizes how that error is described to the model (verbose for debugging, terse to save tokens).- Per-call
@function_tool(timeout=...)works for async handlers — the natural guard against one hung tool stalling a fan-out.
Design rule: failed reads degrade; failed writes stop. If charge_card() errors mid-batch, don’t let the model plow on to ship_order() — return an explicit “payment pending, escalate” result.
Batching for throughput
Beyond same-turn fan-out, batch explicitly when the workload is a set:
- Pre-declare the batch in one tool (
get_prices(sku_list)) instead of N round-trips — fewer tokens, fewer loop iterations. - Group by side-effect class: all reads in one turn, all writes in a later turn.
- Bound the work per turn. Uncapped concurrency plus large batches is how you earn rate-limit errors; chunk and let the loop drive the pace.
- For model-controlled branching at scale, the SDK’s experimental
ProgrammaticToolCallingToollets the model generate JavaScript in a hosted V8 sandbox that can loop and issue parallel tool calls — a different trade-off (the model writes the orchestration code instead of emitting calls).
A practical checklist
- Decide deliberately: parallel reads (default) or serialized turns (
parallel_tool_calls=False). - Set
max_function_tool_concurrencybased on your backend’s tolerance, not on optimism. - Make I/O tools
asyncso concurrency is real. - Give every mutating tool an idempotency key; never rely on turn ordering for correctness.
- Add timeouts to async tools and errors the model can recover from.
- Watch your traces: if tool calls in one turn overlap less than expected, a sync tool or a low concurrency cap is silently serializing you.
Parallel tool calls are the agent equivalent of Promise.all: gloriously fast when the operations are independent, subtly wrong when they aren’t. Do the dependency audit once, up front, and the loop gets both quicker and safer.