Skip to content
Blog

Chaos Testing Agent Workflows: Breaking Things on Purpose

How to apply chaos engineering to agent workflows — fault injection for model outages, tool failures, and stalls, plus the observability needed to prove resilience.

Published on • October 8, 2026

AI Assistant

Chaos engineering exists because production is where systems actually run, and the failure modes that hurt you are the ones you never tested. Agent workflows add a new category of brittleness on top of distributed systems: nondeterministic models, tools with flaky dependencies, retry loops that multiply cost, and “successful” runs that silently produce garbage.

This post is about deliberately breaking agent workflows — in controlled, observable ways — to find the failure modes before your users do.

Why agents need chaos testing too

Classical chaos engineering (kill a pod, drop packets, saturate a disk) still applies. Agents inherit all of that plus failures unique to their shape:

  • Model provider outages and rate limits — 429s, 503s, degraded quality windows
  • Tool failures — timeouts, auth expiry, schema drift, rate limits on third-party APIs
  • Hallucinated-but-valid outputs — the run “succeeds” with wrong data
  • Retry storms — a transient failure multiplied into 10× cost
  • Stalls — a tool call or sub-agent that never returns
  • Context overflow — a retrieval step that returns too much, blowing the window mid-run
  • Partial failures in multi-agent graphs — one branch fails while others complete

The point of chaos testing is the same as ever: build confidence that your system can handle the failure modes you’ve imagined, and discover the ones you haven’t.

The principles, adapted for agents

Follow the classic principles (from the Netflix lineage that popularized the discipline):

  1. You’re running a steady-state hypothesis experiment, not an outage drill. Define what “normal” looks like in measurable terms before injecting anything.
  2. Run it in production-ish conditions. A mock model in a mock environment proves your mocks work.
  3. Automate it and run it continuously. A quarterly manual chaos drill gets forgotten; a nightly experiment in staging gets fixed.
  4. Minimize blast radius. Target one run, one tenant, one canary — never the fleet.
  5. Stop if you learn nothing new. Experiments are hypothesis tests, not punishment.

Define steady state first

For agent workflows, steady state is richer than uptime. Candidate metrics:

  • Run completion rate by task type (p95 baseline)
  • Cost per run within a band
  • Time-to-completion within a band
  • Tool error rate below a threshold
  • Output quality score (eval/judge) at or above baseline
  • Zero orphaned runs (runs stuck without a terminal state)

If you can’t measure these, fix observability before injecting faults — a chaos experiment without a steady-state signal is just an outage.

Fault injection catalog for agent workflows

Model-layer faults

FaultHow to injectWhat you’re testing
Rate limitingReturn 429 with Retry-AfterBackoff respects server hints, no storm
Hard outageReturn 503 for N minutesFallback model or graceful failure to user
Slow responsesAdd latency to 80% of callsTimeout policy, user-facing stall handling
Truncated outputCut mid-streamClient handles partial content, run state
Degraded qualityRoute to a weaker model silentlyEval/drift detection catches the change
Invalid outputMalformed JSON against schemaParser recovery, retry with repair prompt

The degraded quality fault is the one teams skip. If you can’t tell a GPT-class response from a toy model’s response in your telemetry, you also can’t tell a provider regression from your own bug.

Tool-layer faults

CHAOS_FAULTS = {
    "tool.timeout":       lambda req: raise ToolTimeout(req, after=30),
    "tool.auth_expired":  lambda req: raise ToolAuthError(401),
    "tool.rate_limited":  lambda req: raise ToolError(429),
    "tool.bad_schema":    lambda req: inject_bad_field(req, field="amount"),
    "tool.empty_result":  lambda req: replace_result(req, []),
    "tool.slow":          lambda req: sleep_then(req, 5),
    "tool.flaky":         lambda req: fail_if(f"hash:{req.id}", p=0.3),
}

Each maps to a real production incident you’ve either had or will have. Notably:

  • empty_result tests whether the agent correctly concludes “nothing found” vs. hallucinating a plausible answer
  • bad_schema tests validation at the boundary — does a wrong-typed field propagate into a decision?
  • flaky tests retry logic under partial reliability, which is where backoff policies usually break

Workflow-layer faults

  • Kill a sub-agent mid-run — does the orchestrator mark the run failed, or leave it hanging forever?
  • Drop a queued event — does state reconciliation recover?
  • Duplicate an event — are handlers idempotent?
  • Reorder events — does the state machine tolerate out-of-order delivery?
  • Starve one branch of a graph — do fan-out/fan-in paths time out symmetrically?
  • Inject context overflow — does the summarization/truncation path fire before the provider rejects?

Cost faults

  • Amplify tool-call loops — force an agent into a repeated-call pattern and verify circuit breakers trip before the budget does
  • Latency-induced retries — slow responses cause client-side timeouts, which cause duplicate submissions; verify idempotency keys on run creation

Experiment design

Pick experiments the way you’d pick tests — hypothesis, method, abort criteria:

experiment:
  id: model-outage-fallback
  hypothesis: "If the primary model returns 503 for 3 minutes,
               runs complete on the fallback within 120% of baseline
               completion time, with quality ≥ baseline − 5%."
  steady_state:
    - completion_rate >= 0.95
    - p95_duration_ms <= baseline * 1.2
    - eval_score >= baseline * 0.95
  fault:
    type: model.503
    target: primary_model
    duration: 3m
    scope: canary_tenant
  abort_if:
    - completion_rate < 0.70
    - spend_rate > baseline * 3
  verify:
    - fallback_model_invoked = true
    - alert.fired within 60s
    - zero_runs_in_non_terminal_state

Two things worth emphasizing: abort criteria (stop before you cause a real incident) and alert verification (your alerting is part of what you’re testing — see our post on alerting on agent anomalies).

Where to run it

  • Local/dev with a fault proxy. Wrap model and tool traffic through a proxy you control (or dependency injection at the client layer). Cheapest place to develop experiments.
  • Staging with production-shaped traffic. Replay recorded runs through the fault-injected pipeline. Best signal-to-noise for most teams.
  • Production, scoped. Canary tenant, percentage rollout, or time-boxed window. The only place that tests real conditions — start here only once the first two are boring.

Production experiments need: a kill switch, a single-command rollback, on-call awareness (even for a “safe” experiment), and a clear owner.

Observability is the prerequisite

You cannot run chaos experiments without:

  1. Run-level traces — every model call, tool call, and retry attributed to a run.id
  2. Decision/event logs with sequence numbers — so you can reconstruct what the agent chose under fault
  3. Alerting on stalls and anomalies — the experiment should trigger your alerts; if nothing pages during a 3-minute model outage, your alerting has a gap
  4. Cost accounting per run — to catch retry storms as they happen
  5. A steady-state dashboard — the hypothesis’s metrics, visible live during the experiment

If you’re missing any of these, the experiment’s first finding will be an observability gap — which is a perfectly good outcome, as long as you write it down.

Anti-patterns

Chaos without a hypothesis. “Let’s break stuff and see what happens” produces anecdotes, not confidence.

Only testing the happy-path faults. Everyone tests “what if the model is down.” Fewer test “what if the model succeeds but returns confidently wrong structured data” — which is the failure that actually ships.

One-shot experiments. Resilience decays. Run the suite nightly in staging; run a rotating subset in production.

Testing in an environment that doesn’t match production. Mock models hide retry-storm and timeout bugs entirely.

Ignoring the cost axis. A workflow that survives a fault by retrying 20× has survived at $40 a run. That’s a failure.

Not writing it down. Every experiment — especially the ones that found nothing — belongs in a runbook entry: hypothesis, result, date, follow-ups.

The payoff

The most valuable output of a chaos program isn’t the bugs you find; it’s the calibrated confidence you gain. After you’ve watched the model outage experiment pass — fallback engaged, alerts fired in 52 seconds, zero orphaned runs — the 3 a.m. page stops being an unknown. Teams that have practiced failure respond to failure. Teams that haven’t improvise.

Start with three experiments: model outage with fallback, tool timeout with bounded retry, and stalled-run detection. Run them in staging until they’re boring, then scope them into production. Everything else in the fault catalog is iteration on that foundation.

References