Error Handling for Tool Calls: Retryable vs. Fatal Errors
Classifying agent tool failures into retryable and fatal buckets — how the OpenAI Agents SDK surfaces errors to the model, configures retries, and when never to retry at all.
Published on • October 9, 2026
AI Assistant

When a tool call fails inside an agent run, you face a classification decision that determines whether the run recovers, retries forever, or crashes: is this error retryable, or fatal? Get the bucket wrong in one direction and you burn tokens on doomed retries; wrong in the other and transient network blips kill runs that should have survived.
The OpenAI Agents SDK has settled this taxonomy — it is a useful map even if you use a different framework.
The SDK’s core move: errors become model-visible output
By default, the SDK converts tool failures into output strings the model can read. A timeout does not crash the run — the model receives:
Tool 'slow_lookup' timed out after 2 seconds.
and can adapt: pick another tool, narrow the query, or admit it cannot answer. The tool layer’s failure becomes the model’s information.
from agents import function_tool
@function_tool(timeout=5.0, timeout_behavior="error_as_result")
def slow_lookup(query: str) -> str:
...
# On timeout the model receives a readable message and can self-correct.
# Flip to timeout_behavior="raise_exception" to raise ToolTimeoutError instead.
The companion control is failure_error_function — the default returns a model-visible error message; set it to None to propagate the exception instead.
The retryable class
Transient, infrastructure-shaped failures worth trying again:
- HTTP 408, 409, 429, 500, 502, 503, 504
- Network errors (connection resets, DNS hiccups)
- Explicit
retry_aftersignals - Provider-suggested retries
- Timeouts — if you want the model to see them rather than raise
The SDK’s retry system is opt-in via ModelRetrySettings:
from agents import ModelSettings
from agents.retry_policies import (
any as retry_any, provider_suggested, retry_after,
network_error, http_status,
)
settings = ModelSettings(retry=ModelRetrySettings(
max_retries=4,
backoff={"initial_delay": 0.5, "max_delay": 5.0, "multiplier": 2.0, "jitter": True},
policy=retry_any(
provider_suggested(),
retry_after(),
network_error(),
http_status([408, 409, 429, 500, 502, 503, 504]),
),
))
Two properties worth internalizing:
- Jitter and capped exponential backoff — without jitter, concurrent agents retry in lockstep and amplify the outage.
- Replay-safety: no retry once events have been emitted. Streaming consumers must never see duplicate events, so the SDK suppresses retries after the event stream has begun.
The fatal class
Deterministic failures where retrying only burns turns:
| Failure | Why it is fatal |
|---|---|
tool_not_found | The tool will not appear next call either |
| Schema/collision violations | Input is wrong, not the transport |
model_refusal | Policy outcome, not an error |
max_turns exceeded | Budget exhausted by design |
invalid_final_output | Structured output contract broken |
MCPToolCancellationError | Deliberate cancellation, not a failure |
These are routed to run-level error handlers, not retry loops:
# error handler kinds: "max_turns", "model_refusal", "invalid_final_output"
# plus tool_not_found_behavior and ToolErrorFormatter on RunConfig
ToolErrorFormatter with ToolErrorFormatterArgs (kinds "approval_rejected", "tool_not_found") lets you shape the error message the model or human sees — a fatal error should still produce a useful message, because the model may need to explain the failure to the user.
The decision table
| Error | Class | Action |
|---|---|---|
| 429 / 5xx | Retryable | SDK retry with backoff + jitter |
| Network drop | Retryable | SDK retry |
| Timeout (transient dep) | Retryable | error_as_result so model adapts, or raise |
| Timeout (model wants fallback) | Retryable | Let the error string reach the model |
| Unknown tool name | Fatal | Handler + formatter; never retry |
| Invalid arguments | Fatal-ish | Return error to model once so it can fix args — guard with max_turns |
| Policy refusal | Fatal | Run handler |
| Invalid structured output | Fatal | Run handler |
| Cancellation | Not an error | Propagate |
Notice invalid arguments sit in the middle: they are deterministic, but returning the error to the model is often the right recovery because the model can regenerate correct arguments. The guardrail is a max_turns cap so a confused model cannot loop forever.
Designing error messages the model can use
If an error string reaches the model, write it like a tool contract, not a stack trace:
- Say what failed —
Tool 'get_quote' failed: symbol 'MSFTT' not found. - Say what to try —
Did you mean a valid ticker? Available symbols: ... - Do not leak secrets — no tokens, no internal hostnames.
A great error message is a second chance for the model; a raw exception is a dead end.
Gotchas
- The default
failure_error_functionand timeout-as-result hide real bugs — in development, flip toraise_exception/Noneand log tool errors independently. - Retries being opt-in means a naive setup never retries — configure
ModelRetrySettingsexplicitly. - Never retry the fatal class: it deterministically fails and just consumes turns.
- Concurrency matters too:
max_function_tool_concurrencycaps parallel calls, and tool namespaces withallowed_callersscope who can invoke what.
Wrapping up
Split every tool failure into two buckets: transient infrastructure (retry with capped, jittered backoff — or surface to the model so it can route around the failure) and deterministic contract violations (format a clear error, hand it to a run handler, never blind-retry). Then make sure the defaults you ship are the ones you mean: models that see useful errors recover, and agents that never retry doomed calls stay within budget.
References
- OpenAI Agents SDK tools documentation - function_tool, timeouts, failure_error_function, concurrency
- Exceptions reference - ToolTimeoutError, ModelBehaviorError, MCPToolCancellationError
- Run error handlers - max_turns, model_refusal, invalid_final_output
- OpenAI Agents SDK home - retry policies and ModelRetrySettings