Skip to content
Blog

Circuit Breakers for LLM and Tool Dependencies

Apply the circuit-breaker pattern specifically to LLM providers and agent tool endpoints — failure classification, per-dependency state machines, LiteLLM gateway wiring, and fault-injection tests that prove the breaker trips when it should.

Published on • October 11, 2026

AI Assistant

At 02:14 UTC, a mid-size support automation started burning money instead of saving it. A major LLM provider began returning intermittent 503s for roughly nine minutes. The agent platform had retries — three per call, exponential backoff — but no breaker anywhere. Every queued conversation re-issued its completion call, each retry fanned out across the provider’s already-struggling endpoint, and worker threads piled up waiting on requests that would never return. By the time a human paged in, the retry storm had consumed a day’s token budget and starved the healthy secondary provider of concurrency too.

The fix was not more retries. The fix was a component that stops calling a dependency that has already proven it cannot serve: a circuit breaker. When the provider’s failure rate crossed the trip threshold, the breaker opened, calls fast-failed in microseconds instead of seconds, and a fallback model absorbed the traffic. Nine minutes later the breaker probed, the provider answered, and traffic flowed again — with the whole outage visible as three state transitions in the metrics.

This post is narrowly about that component applied to the two dependency classes that dominate agent systems: LLM provider deployments and tool endpoints. It complements our earlier discussion of orchestrator-level resilience; here we go deep on the breaker itself — what trips it, what must not, how it layers with retries and fallbacks, and how gateways like LiteLLM already implement the pattern for you.

In this tutorial, you will learn how to:

  • Model closed, open, and half-open breaker states as an explicit state machine for LLM and tool calls
  • Classify failures correctly — 5xx, exhausted 429s, and timeouts trip the breaker; content errors do not
  • Choose failure thresholds, open windows, and half-open probe limits that fit provider traffic shapes
  • Scope breakers per dependency (per provider, per model deployment, per tool endpoint) instead of globally
  • Separate concerns: retries inside the closed state, breakers across unhealthy deployments, fallback chains across providers
  • Wire breaker behavior into an LLM gateway in the LiteLLM style — cooldowns, retries, timeouts, and fallbacks as configuration
  • Instrument and alert on breaker signals: error rate, latency p95, and state transitions
  • Prove the breaker works with fault injection and deterministic state-machine tests

Key technologies: Python 3.11+, LiteLLM Router and Proxy (cooldowns, retries, fallbacks), Prometheus-style metrics, fault-injection testing.

Prerequisites

  • Comfortable writing concurrent Python (asyncio, timeouts, exception hierarchies)
  • Basic familiarity with calling LLM providers through an OpenAI-compatible API
  • A provider key (OpenAI, Anthropic, or any OpenAI-compatible endpoint) if you want to run the gateway examples live; the breaker unit tests need no network

Why LLM and tool dependencies fail differently

Remote calls can fail or hang — that is the classic argument for breakers. LLM and tool dependencies add three wrinkles that change how you tune the pattern:

Rate limits are normal operation, not incidents. A 429 from a provider is often just correct backpressure. If every 429 trips the breaker, a legitimate quota ceiling looks identical to an outage and you fail over to a provider you do not have capacity with either. The breaker must distinguish “we exceeded our rate limit” (retry later, maybe same deployment) from “this deployment is sick” (stop sending traffic).

Slow tails matter as much as hard errors. LLM endpoints degrade latently: p99 climbs from four seconds to ninety while the error rate stays flat. A breaker that counts only exceptions never trips, and your agent’s per-request deadline does the cascading for you. Latency-based trip conditions — p95 above a budget over a rolling window — are part of a modern breaker for this dependency class.

Failures are heterogeneous across a model group. One Azure deployment of a model can be auth-broken (401) while its twin is healthy. Treating the model name as the failure domain masks the sick deployment inside a pool that still has capacity. LiteLLM’s Router documents this exact shape: non-retryable errors like 401, 404, and 408 put the specific deployment into a cooldown (5 seconds by default) while the group keeps serving.

The breaker state machine

Martin Fowler’s write-up of the pattern (crediting Michael Nygard’s Release It!) defines three states. Concretely for an LLM or tool call:

StateEntry conditionCall behaviorExit condition
ClosedInitial state; also after recoveryCalls pass through; failures recordedFailure count or rate crosses threshold → Open
OpenThreshold crossedFast-fail immediately; no upstream call madeOpen window (reset timeout) elapses → Half-open
Half-openOpen window elapsedLimited probe calls onlyProbe succeeds → Closed; probe fails → Open (window restarts, often grown)

Fowler’s self-resetting variant is the one you want in production: the breaker itself decides when to try again after the reset timeout, rather than requiring an operator to flip it back. A compact async implementation:

import enum, time, asyncio
from dataclasses import dataclass, field

class BreakerState(enum.Enum):
    CLOSED = "closed"
    OPEN = "open"
    HALF_OPEN = "half_open"

@dataclass
class CircuitBreaker:
    name: str
    failure_threshold: int = 5          # consecutive failures to trip
    open_window_s: float = 30.0         # how long to stay open before probing
    half_open_max_probes: int = 1       # concurrent probes allowed in half-open
    state: BreakerState = BreakerState.CLOSED
    _failures: int = 0
    _opened_at: float = 0.0
    _probes: int = 0
    _lock: asyncio.Lock = field(default_factory=asyncio.Lock)

    def allow(self) -> bool:
        if self.state is BreakerState.CLOSED:
            return True
        if self.state is BreakerState.OPEN:
            if time.monotonic() - self._opened_at >= self.open_window_s:
                self.state = BreakerState.HALF_OPEN
                self._probes = 0
                return True
            return False
        # HALF_OPEN: only a bounded number of probes pass
        if self._probes < self.half_open_max_probes:
            self._probes += 1
            return True
        return False

    def record_success(self) -> None:
        self._failures = 0
        self.state = BreakerState.CLOSED

    def record_failure(self) -> None:
        if self.state is BreakerState.HALF_OPEN:
            self.state = BreakerState.OPEN
            self._opened_at = time.monotonic()
            self.open_window_s = min(self.open_window_s * 2, 300.0)  # back off probes
            return
        self._failures += 1
        if self._failures >= self.failure_threshold:
            self.state = BreakerState.OPEN
            self._opened_at = time.monotonic()

Two details from Fowler’s article that matter here: not every error should trip the breaker — some reflect normal failures and belong in regular call logic — and a count that resets on success is the simple trip rule, while frequency-based rules (“50% failure rate over the current minute”) handle bursty traffic better.

What counts as a failure (and what does not)

The trip predicate is the single most consequential decision. Counting the wrong exceptions turns your breaker into either a no-op or a flap machine.

OutcomeTrip breaker?Why
HTTP 5xx from providerYesProvider-side failure; sustained 5xx means unhealthy deployment
Connection error / DNS failureYesDependency unreachable
Timeout (total or first-token while streaming)YesSlow tails consume agent deadlines; treat as failure
HTTP 429 after retry budget exhaustedYesRate limiting that survived backoff indicates a capacity problem at current send rate
HTTP 400 context-window exceededNoThat specific request is wrong, not the provider; use a context-window fallback instead
HTTP 401 / 403 / 404Yes, but alarm differentlyMisconfiguration will not self-heal; trip fast and page — LiteLLM’s Router likewise cools down deployments on these non-retryable errors rather than retrying
Model refusal / content-policy responseNoA valid response; route via content-policy fallbacks if needed
Malformed tool arguments from the modelNoModel-behavior problem; repair or re-prompt, never penalize the provider
Tool business error (e.g. record not found)NoThe tool worked; the call was just wrong

Encode this as a single classifier that both the breaker and the retry policy share, so the two mechanisms never disagree:

import httpx

RETRIABLE_STATUS = {408, 429, 500, 502, 503, 504}

def is_dependency_failure(exc: Exception) -> bool:
    if isinstance(exc, httpx.TimeoutException):
        return True
    if isinstance(exc, httpx.TransportError):
        return True
    if isinstance(exc, httpx.HTTPStatusError):
        return exc.response.status_code in RETRIABLE_STATUS
    # Unknown transport-level exceptions: fail safe, count them
    return isinstance(exc, (ConnectionError, OSError))

Keep business and content outcomes entirely outside this function. A refusal is a 200 with a grumpy body, and a record-not-found is a 200 with an empty payload — neither ever reaches the classifier.

Thresholds, open windows, and half-open probes

Defaults worth shipping, then tuning with your own traffic:

  • Consecutive-failure threshold: 5. Good for low-volume tools. For a gateway doing hundreds of requests per minute per deployment, prefer a rolling-window failure rate (50%+ failures over the current minute) so a single bad minute trips without waiting for five sequential failures.
  • Open window: 30 seconds, doubling to a 5-minute cap on repeated probe failures. Too short and you probe a struggling provider every few seconds; too long and you keep traffic off a recovered dependency. LiteLLM’s Router defaults to short cooldowns (5 seconds) precisely because a model group with healthy deployments can absorb the failover while the sick one recovers.
  • Half-open probes: 1 concurrent, 3–5 successes to fully close. One probe protects a provider that is still flapping; requiring multiple successes prevents a single lucky response from reopening the floodgates. In the implementation above, require half_open_max_probes successes before treating recovery as real, or decay the failure counter across probes.
  • 429s get their own handling: honor Retry-After inside the closed state via retries; only sustained 429s after backoff should influence the breaker. A useful disambiguation from LiteLLM’s error reference: a 429 carrying retry-after from your own gateway’s limits is a you-problem, while a bare 429 is the provider’s throttle.

Scope breakers per dependency

A breaker protects a dependency, so its key must be the failure domain, not the abstract capability:

Key granularityProtects againstCost of too coarse
Per provider (e.g. openai)Whole-vendor outageOne sick model deployment trips the breaker for every model
Per deployment (e.g. azure/eastus/gpt-4o-mini-02)Deployment-level auth, quota, regional issuesRecommended default for gateways with multiple keys/endpoints
Per model groupSustained group-level failureHides individual deployment health inside the pool
Per tool endpoint (e.g. crm.search)That tool’s API being downA global tools breaker lets a dead search API disable billing lookups

In an agent that calls a dozen tools, per-tool breakers are non-negotiable for exactly the reason Fowler gives: a broken circuit should stop the calls that are likely to fail, not the calls that are not. Your CRM search being down says nothing about your calculator.

Separating breakers from retries and fallbacks

These three mechanisms answer different questions and must not be smeared into one retry loop:

QuestionMechanismWhere it lives
Was this a transient blip?Retry with backoff (and jitter)Inside the closed state, per attempt
Is this specific deployment unhealthy?Circuit breakerWraps each deployment or tool endpoint
Is the whole provider unavailable?Fallback chainAbove breakers: next model group or provider

The ordering is: check the breaker, run retries only if the breaker is closed, and consult fallbacks only when calls actually fail or the breaker is open. LiteLLM’s Router documents this layering explicitly: function_with_fallbacks wraps function_with_retries, which wraps the provider call — retries re-attempt within the same model group, fallbacks move to a different group. A deployment in cooldown (the gateway’s breaker) is excluded from selection entirely.

The failure mode to avoid: retries outside the breaker. With the breaker open, each incoming request would still burn its full retry budget before fast-failing — the exact resource pile-up breakers exist to prevent. Conversely, breakers without retries trip on single blips; Fowler’s original count-based example is fine for teaching, but production traffic needs the retry layer to filter noise before it reaches the trip counter.

Wiring into an LLM gateway (LiteLLM-style)

You rarely need to hand-roll the state machine for model traffic. LiteLLM’s Router implements the breaker equivalent — cooldowns — alongside retries, timeouts, and fallbacks, all configurable:

from litellm import Router

model_list = [
    {
        "model_name": "support-fast",                       # model group
        "litellm_params": {
            "model": "azure/eastus/gpt-4o-mini",
            "api_key": "sk-...", "api_base": "https://eastus.example.openai.azure.com",
            "timeout": 30, "stream_timeout": 10,            # total + first-token deadlines
        },
    },
    {
        "model_name": "support-fast",
        "litellm_params": {
            "model": "anthropic/claude-haiku",
            "api_key": "sk-ant-...",
        },
    },
]

router = Router(
    model_list=model_list,
    num_retries=3,                       # retries live inside each attempt
    timeout=30,
    fallbacks=[{"support-fast": ["support-premium"]}],  # cross-group failover
    context_window_fallbacks=[{"support-fast": ["support-longctx"]}],
)

Key behaviors to rely on (from LiteLLM’s routing and reliability docs):

  • Cooldowns are the breaker. Deployments hitting a 429, exceeding a ~50% failure rate in the current minute, or returning non-retryable errors (401/404/408) are cooled down (5 seconds by default) and excluded from routing while the rest of the group serves.
  • Retries use exponential backoff for rate limits and immediate re-attempt for generic errors; num_retries bounds them.
  • Fallbacks try each backup once, in order, and raise a final error — “All fallback attempts failed” — when the chain is exhausted; the success response carries x-litellm-attempted-fallbacks for observability. Router-level fallbacks add cooldown-awareness that the plain SDK completion(fallbacks=[...]) pass does not.
  • Errors you should never retry (400 malformed request, 401, 403, 404, 422) are classified in the error reference; only 408 and 429 (honoring retry-after) are retry candidates.
  • Health-check driven routing (documented under Routing & Load Balancing) runs background probes on an interval and removes failing deployments from the pool before user traffic hits them — proactive breakers layered on reactive cooldowns.

On the proxy side, the same concerns appear as litellm_settings (num_retries, request_timeout, default_fallbacks, plus context_window_fallbacks and content_policy_fallbacks on the model list), so application code inherits breaker behavior by pointing its OpenAI SDK at the gateway rather than at providers.

Breakers in agent tool loops

Model traffic is only half the surface. An agent loop that executes tool calls has the same dependency shape, and the breaker belongs around each tool’s outbound call:

import json

BREAKERS = {name: CircuitBreaker(name) for name in TOOLS}

async def execute_tool(name: str, args: dict) -> str:
    breaker = BREAKERS[name]
    if not breaker.allow():
        # Fast-fail: return a structured tool error so the model can adapt
        return json.dumps({"error": "tool_unavailable",
                           "detail": f"{name} circuit open; try a different approach"})
    try:
        result = await call_tool_endpoint(name, args, timeout=15)
    except Exception as exc:
        if is_dependency_failure(exc):
            breaker.record_failure()
        return json.dumps({"error": "tool_failed", "detail": str(exc)})
    breaker.record_success()
    return json.dumps(result)

Two agent-specific points. First, feed the open-breaker result back to the model as a tool error rather than aborting the run — models handle “tool unavailable, try another route” well, and the run degrades instead of dying. Second, respect the failure classifier: a tool returning “no matching records” is a success for the breaker even though the agent’s plan has to change.

Metrics, observability, and operations

Fowler is explicit that breakers are a prime monitoring point: any state change should be logged, and operators should be able to trip or reset breakers manually. Minimum viable instrumentation:

  • breaker_state{dependency} — gauge (0 closed / 1 half-open / 2 open) for dashboards
  • breaker_transitions_total{dependency, from, to} — counter; alert on any transition to open
  • dependency_error_rate{dependency} — rolling error percentage, the trip input
  • dependency_latency_p95{dependency} — histogram; latency-based trip conditions
  • breaker_fast_fail_total{dependency} — calls rejected without an upstream request (your savings, made visible)
  • fallback_attempted_total{from, to} and the gateway’s own x-litellm-attempted-fallbacks / x-litellm-attempted-retries headers for tracing failover paths

Pair transition alerts with runbooks: which dependency, which fallback engages, and when to trip manually. LiteLLM’s error reference makes a related operational point worth adopting: expose routing debug detail (expose_router_debug_in_errors is on by default) in gateway responses so an exhausted retry/fallback chain is diagnosable from the error itself.

Testing breakers with fault injection

A breaker you have never seen trip is a breaker you do not have. Three layers of tests:

  1. Deterministic state-machine tests. Inject outcomes into the breaker with a fake clock: N failures trip; success resets; open blocks; the window elapses into half-open; a failed probe re-opens with a grown window. No network, no sleeps beyond the fake clock.
  2. Fault injection against the transport. Wrap the tool HTTP client with a chaos layer that returns 503s, hangs past the timeout, or returns 429s with Retry-After. LiteLLM supports this natively for the model path — mock_timeout=True on /chat/completions exercises your retry and timeout handling without real latency.
  3. Scenario tests at the agent level. Kill one provider for five simulated minutes; assert the breaker opened within a bounded time, the fallback carried traffic, the healthy dependency’s breaker never moved, and total spend stayed inside a budget.
def test_open_breaker_blocks_and_recovers():
    cb = CircuitBreaker("crm.search", failure_threshold=3, open_window_s=10)
    for _ in range(3):
        cb.record_failure()
    assert cb.state is BreakerState.OPEN
    assert cb.allow() is False            # fast-fail, no upstream call
    time.sleep(0)                         # (use a fake clock in real tests)
    cb._opened_at -= 11                   # window elapsed
    assert cb.allow() is True             # probe admitted
    cb.record_success()
    assert cb.state is BreakerState.CLOSED

Extend the scenario suite to the failure you actually fear: two dependencies failing at once (does the fallback chain have a sane terminal state?), and a breaker stuck half-open (does anything alert?).

Common misconfigurations

  • One global breaker for all tools. The search API outage disables the calendar. Scope per dependency.
  • Counting content and business errors as failures. Refusals, context-window overflows, and empty result sets trip the breaker while the provider is perfectly healthy. Classify first.
  • Retrying outside the breaker. Open circuit, full retry budget, pile-up — the original incident all over again.
  • Thresholds tuned for the demo, not the traffic. Five consecutive failures is nothing at 200 rpm per deployment; use rolling-window failure rates there.
  • Open window shorter than typical recovery. You probe every few seconds into a provider that needs minutes; probe storms become a self-inflicted outage.
  • No latency trip condition. Error-rate-only breakers sleep through p95 blowups that still destroy agent UX.
  • Breaking fallbacks out of cooldown logic. Failing over to a deployment that is itself in cooldown just moves the error; prefer gateways (like LiteLLM’s Router) that skip cooled-down deployments, and reserve “explicit model-id fallback” for cases where you deliberately want to bypass the cooldown check.
  • Silent transitions. If nobody gets paged when a breaker opens, you will learn about the outage from customers.

Pre-production checklist

  • Breaker keys match real failure domains (deployment-level for gateways, endpoint-level for tools)
  • Failure classifier excludes content, business, and context-window errors
  • 429s honor Retry-After and only trip the breaker after backoff is exhausted
  • Retries are nested inside the closed state; fallbacks sit above breakers
  • Half-open admits a bounded number of probes; repeated probe failure grows the open window
  • Latency (p95) trip condition exists alongside error-rate tripping
  • State transitions emit metrics and alert; fast-fails are counted
  • Degraded-mode tool results are structured so the agent can adapt instead of dying
  • Fault-injection suite covers: provider 5xx storm, hang-past-timeout, 429 storm, dual-dependency failure, stuck half-open
  • Runbook exists for manual trip/reset of each critical dependency

Further reading