Skip to content
Blog

Error Handling for Tool Calls: Retryable vs. Fatal Errors

Classifying agent tool failures into retryable and fatal buckets — how the OpenAI Agents SDK surfaces errors to the model, configures retries, and when never to retry at all.

Published on • October 9, 2026

AI Assistant

When a tool call fails inside an agent run, you face a classification decision that determines whether the run recovers, retries forever, or crashes: is this error retryable, or fatal? Get the bucket wrong in one direction and you burn tokens on doomed retries; wrong in the other and transient network blips kill runs that should have survived.

The OpenAI Agents SDK has settled this taxonomy — it is a useful map even if you use a different framework.

The SDK’s core move: errors become model-visible output

By default, the SDK converts tool failures into output strings the model can read. A timeout does not crash the run — the model receives:

Tool 'slow_lookup' timed out after 2 seconds.

and can adapt: pick another tool, narrow the query, or admit it cannot answer. The tool layer’s failure becomes the model’s information.

from agents import function_tool

@function_tool(timeout=5.0, timeout_behavior="error_as_result")
def slow_lookup(query: str) -> str:
    ...

# On timeout the model receives a readable message and can self-correct.
# Flip to timeout_behavior="raise_exception" to raise ToolTimeoutError instead.

The companion control is failure_error_function — the default returns a model-visible error message; set it to None to propagate the exception instead.

The retryable class

Transient, infrastructure-shaped failures worth trying again:

  • HTTP 408, 409, 429, 500, 502, 503, 504
  • Network errors (connection resets, DNS hiccups)
  • Explicit retry_after signals
  • Provider-suggested retries
  • Timeouts — if you want the model to see them rather than raise

The SDK’s retry system is opt-in via ModelRetrySettings:

from agents import ModelSettings
from agents.retry_policies import (
    any as retry_any, provider_suggested, retry_after,
    network_error, http_status,
)

settings = ModelSettings(retry=ModelRetrySettings(
    max_retries=4,
    backoff={"initial_delay": 0.5, "max_delay": 5.0, "multiplier": 2.0, "jitter": True},
    policy=retry_any(
        provider_suggested(),
        retry_after(),
        network_error(),
        http_status([408, 409, 429, 500, 502, 503, 504]),
    ),
))

Two properties worth internalizing:

  1. Jitter and capped exponential backoff — without jitter, concurrent agents retry in lockstep and amplify the outage.
  2. Replay-safety: no retry once events have been emitted. Streaming consumers must never see duplicate events, so the SDK suppresses retries after the event stream has begun.

The fatal class

Deterministic failures where retrying only burns turns:

FailureWhy it is fatal
tool_not_foundThe tool will not appear next call either
Schema/collision violationsInput is wrong, not the transport
model_refusalPolicy outcome, not an error
max_turns exceededBudget exhausted by design
invalid_final_outputStructured output contract broken
MCPToolCancellationErrorDeliberate cancellation, not a failure

These are routed to run-level error handlers, not retry loops:

# error handler kinds: "max_turns", "model_refusal", "invalid_final_output"
# plus tool_not_found_behavior and ToolErrorFormatter on RunConfig

ToolErrorFormatter with ToolErrorFormatterArgs (kinds "approval_rejected", "tool_not_found") lets you shape the error message the model or human sees — a fatal error should still produce a useful message, because the model may need to explain the failure to the user.

The decision table

ErrorClassAction
429 / 5xxRetryableSDK retry with backoff + jitter
Network dropRetryableSDK retry
Timeout (transient dep)Retryableerror_as_result so model adapts, or raise
Timeout (model wants fallback)RetryableLet the error string reach the model
Unknown tool nameFatalHandler + formatter; never retry
Invalid argumentsFatal-ishReturn error to model once so it can fix args — guard with max_turns
Policy refusalFatalRun handler
Invalid structured outputFatalRun handler
CancellationNot an errorPropagate

Notice invalid arguments sit in the middle: they are deterministic, but returning the error to the model is often the right recovery because the model can regenerate correct arguments. The guardrail is a max_turns cap so a confused model cannot loop forever.

Designing error messages the model can use

If an error string reaches the model, write it like a tool contract, not a stack trace:

  • Say what failed — Tool 'get_quote' failed: symbol 'MSFTT' not found.
  • Say what to try — Did you mean a valid ticker? Available symbols: ...
  • Do not leak secrets — no tokens, no internal hostnames.

A great error message is a second chance for the model; a raw exception is a dead end.

Gotchas

  • The default failure_error_function and timeout-as-result hide real bugs — in development, flip to raise_exception/None and log tool errors independently.
  • Retries being opt-in means a naive setup never retries — configure ModelRetrySettings explicitly.
  • Never retry the fatal class: it deterministically fails and just consumes turns.
  • Concurrency matters too: max_function_tool_concurrency caps parallel calls, and tool namespaces with allowed_callers scope who can invoke what.

Wrapping up

Split every tool failure into two buckets: transient infrastructure (retry with capped, jittered backoff — or surface to the model so it can route around the failure) and deterministic contract violations (format a clear error, hand it to a run handler, never blind-retry). Then make sure the defaults you ship are the ones you mean: models that see useful errors recover, and agents that never retry doomed calls stay within budget.

References