Contract Testing Your Agents: Asserting Behavior Before Deployment
Stop asserting on prose. Define the inputs, tool schemas, output shapes, and side effects an agent must honor, then test those contracts deterministically in CI.
Published on • October 4, 2026
AI Assistant

Contract Testing Your Agents: Asserting Behavior Before Deployment
Most agent test suites are a folder of prompt strings and hopeful substring checks: send “I want a refund,” assert the answer contains the word “refund.” They pass on Tuesday, fail on Wednesday when the provider silently updates a model, and tell you nothing when they fail — you cannot tell a regression from sampling variance, so you learn to ignore them.
The fix is to stop testing the prose and test the contract: the structured parts of the agent’s behavior that downstream code depends on — inputs, tool schemas, the shape of the structured output, and the side effects a run is allowed to produce. Those are assertable, and once they are asserted, prompt changes and model swaps stop being scary.
Key technologies: Pydantic v2 (models, JSON Schema, strict mode), Pydantic AI (output_type, TestModel, FunctionModel, capture_run_messages), OpenAI Agents SDK (ScriptedModel, ensure_strict_json_schema), pytest, hypothesis.
Why Prompt-Only Tests Are Brittle
A prompt assertion couples your test to three things you do not control: the model’s wording, the provider’s model version, and temperature. The failures that actually ship are structural — a field renamed in the output schema, a tool argument that became optional, a refund() call that now fires twice — and a substring test is blind to all of them, because none of them change the prose.
Contract testing inverts the priority: assert the invariants that would break an integration, and leave “is this answer good?” to the eval layer.
What a Contract Is for an Agent
| Contract facet | Example assertion | Breaks when |
|---|---|---|
| Inputs | Prompt payload validates against the request model | Callers start passing partial data |
| Tool schemas | Generated JSON Schema matches a committed golden file | A signature, docstring, or default changes |
| Structured output shape | Every run returns a model that validates, strictly | A field is renamed, retyped, or made optional |
| Side-effect invariants | Exactly one refund() call, amount within bounds | The model retries, loops, or hallucinates a call |
| Turn shape | The scripted sequence of model calls is consumed exactly | The workflow gains or loses a step |
The first four are ordinary unit and contract tests. The fifth is where agent frameworks earn their keep.
Structured Output Is the Strongest Contract You Have
Pydantic is the natural place to write it down: per the Pydantic docs, “schema validation and serialization are controlled by type annotations,” and models emit JSON Schema directly — exactly what the LLM is told to produce.
# contracts/output.py
from typing import Literal
from pydantic import BaseModel, ConfigDict, Field
class RefundDecision(BaseModel):
model_config = ConfigDict(strict=True)
refund: bool
amount_usd: float = Field(ge=0, le=10_000)
reason: Literal["duplicate", "not_received", "damaged", "none"]
order_id: str = Field(pattern=r"^ord_[a-z0-9]+$")
Strict mode is the part people skip. By default Pydantic coerces: '123' becomes 123. With strict=True — settable per call via model_validate(payload, strict=True), per field with Field(strict=True), or model-wide with ConfigDict(strict=True) — it raises instead. That is the difference between “the agent produced a refund amount” and “the agent produced something we could coerce into a refund amount.”
# tests/test_output_contract.py
import pytest
from pydantic import ValidationError
from contracts.output import RefundDecision
def test_valid_decision_round_trips():
decision = RefundDecision.model_validate(
{"refund": True, "amount_usd": 12.5, "reason": "duplicate",
"order_id": "ord_7"},
strict=True,
)
assert decision.amount_usd == 12.5
@pytest.mark.parametrize("payload", [
{"refund": "yes", "amount_usd": 1, "reason": "none", "order_id": "ord_1"},
{"refund": True, "amount_usd": -1, "reason": "none", "order_id": "ord_1"},
{"refund": True, "amount_usd": 1, "reason": "maybe", "order_id": "ord_1"},
{"refund": True, "amount_usd": 1, "reason": "none", "order_id": "7"},
])
def test_shape_is_a_hard_contract(payload):
with pytest.raises(ValidationError):
RefundDecision.model_validate(payload, strict=True)
Wire the same model to the agent as its output_type and the guarantee becomes runtime, not just test-time: Pydantic AI states “the run is guaranteed to return a RefundDecision,” and if validation fails the agent is prompted to try again. On the OpenAI Agents SDK side, agents.strict_schema.ensure_strict_json_schema(schema) mutates a JSON schema so it conforms to the strict standard the OpenAI API expects — worth asserting before a provider call ever happens.
Tool-Schema Golden Tests
Tool schemas are a public API: the model reads them to decide what to call, and Pydantic AI passes your signature and docstring through as the schema — “the rest of its signature and its docstring become the tool schema, arguments are validated before your code runs.” A renamed parameter or dropped docstring changes model behavior with zero type errors.
Snapshot it:
# tests/test_tool_schema_golden.py
import json
import pathlib
from contracts.tools import LookupOrderArgs
GOLDEN = pathlib.Path(__file__).parent / "golden" / "lookup_order.json"
def test_lookup_order_schema_has_not_drifted():
actual = LookupOrderArgs.model_json_schema()
if not GOLDEN.exists():
GOLDEN.parent.mkdir(exist_ok=True)
GOLDEN.write_text(json.dumps(actual, indent=2, sort_keys=True))
raise AssertionError("golden file created — review and commit it")
expected = json.loads(GOLDEN.read_text())
assert actual == expected, "schema drifted; regenerate if intentional"
Now a schema change is a red diff in review instead of a quiet behavior change in production.
Replaying Recorded Tool Calls
The second half of the contract is the sequence: does the agent call the tool it should, once, with the right arguments, then stop? Both SDKs let you script the model boundary so the answer is deterministic.
# tests/test_refund_flow.py
import pytest
from agents import Agent, RunConfig, Runner
from agents.decorators import tool
from agents.testing import ScriptedModel, assistant_message, function_call
refunds: list[tuple[str, float]] = []
@tool
def refund(order_id: str, amount_usd: float) -> str:
"""Issue a refund for an order."""
refunds.append((order_id, amount_usd))
return f"Refunded ${amount_usd:.2f}."
@pytest.mark.asyncio
async def test_refund_flow_calls_tool_once_then_answers():
refunds.clear()
model = ScriptedModel([
[function_call("refund", {"order_id": "ord_7", "amount_usd": 12.5},
call_id="c1")],
[assistant_message("Refunded $12.50.")],
])
agent = Agent(name="Support", model=model, tools=[refund])
result = await Runner.run(
agent, "refund ord_7",
run_config=RunConfig(tracing_disabled=True),
)
assert result.final_output == "Refunded $12.50."
assert refunds == [("ord_7", 12.5)] # side-effect invariant
assert len(model.calls) == 2
model.assert_complete()
ScriptedModel records every call (calls, first_call, last_call), and assert_complete() fails if the workflow ended before consuming every scripted step. The SDK raises structured errors — UnexpectedModelCall for an extra request, UnconsumedModelSteps for an early exit — so drift shows up as a typed failure, not a timeout. Under the hood the real tool pipeline runs: schema generation, argument validation, execution, handoff of the result to the next model turn.
Pydantic AI exposes the same idea from the message side: capture_run_messages() records the ModelRequest / ModelResponse exchange, so you can assert the ToolCallPart name and args, the ToolReturnPart, and the final TextPart while TestModel or FunctionModel stands in for the model.
Property-Based Tests for Tool Inputs
Example-based tests only probe the inputs you thought of. A tool that parses dates, computes money, or builds SQL deserves generated inputs:
# tests/test_refund_properties.py
from hypothesis import given, settings, strategies as st
from tools.money import apply_refund
@given(
amount_usd=st.floats(min_value=0, max_value=10_000, allow_nan=False),
fee_rate=st.floats(min_value=0, max_value=0.25, allow_nan=False),
)
@settings(max_examples=300)
def test_refund_is_always_within_bounds(amount_usd: float, fee_rate: float):
result = apply_refund(amount_usd, fee_rate)
assert 0 <= result <= amount_usd
@given(order_id=st.text(max_size=200))
@settings(max_examples=300)
def test_lookup_never_raises_on_arbitrary_ids(order_id: str):
assert lookup_order(order_id) is not None
The point is not coverage of your example table — it is that a tool the model can call with arbitrary generated text must not raise, must not return a negative refund, and must not accept a ') OR 1=1 -- order ID as valid. Hypothesis hands you each counterexample as a minimal repro you can paste into a fixture.
Stubbing the Model for Deterministic CI
None of the above should touch a real provider. Pydantic AI recommends the standard setup: use TestModel or FunctionModel in place of your real model “to avoid the usage, latency and variability of real LLM calls,” swap it in with agent.override(model=...), and set ALLOW_MODEL_REQUESTS = False globally so an accidental real request fails the run instead of quietly costing money.
# tests/conftest.py
import pytest
from pydantic_ai import models
from pydantic_ai.models.test import TestModel
models.ALLOW_MODEL_REQUESTS = False
@pytest.fixture
def stubbed_agent(agent):
with agent.override(model=TestModel()):
yield
TestModel calls every registered tool, then produces a response shaped by your output_type — procedural schema-driven data, no model. When you need a specific turn sequence (call the tool first, return a future date, fail on the second attempt), FunctionModel takes a plain function that receives the message list and returns the scripted ModelResponse. On the Agents SDK side, ScriptedModel plus RunConfig(tracing_disabled=True) does the same job, and the testing modules make no model or network requests.
The Test Pyramid for Agents
| Layer | What it asserts | Real model | Runtime | When |
|---|---|---|---|---|
| Unit | Pure functions, tool bodies, money and date math | No | Milliseconds | Every commit |
| Contract | Schemas, output shape, side effects, scripted turns | Stubbed | Under a second | Every commit |
| Eval | Task success and quality scores on a fixed dataset | Yes | Minutes | Nightly |
| E2E | Provider auth, network, retries, streaming, real tools | Yes | Minutes to hours | Pre-release |
Pydantic Evals is built for the third row — testing “agent behavior the way pytest tests code.” The rule of thumb: unit through contract must be free, deterministic, and runnable on a laptop; only eval and e2e may be expensive or stochastic, and they report scores, not booleans.
Best Practices
- Assert structure, not wording. Compare parsed output and recorded tool calls, never the sentences around them.
- Fail loudly if CI reaches a provider.
ALLOW_MODEL_REQUESTS = Falseinconftest.py,tracing_disabled=Truein scripted runs. - Golden files are reviewed artifacts. A schema diff belongs in the pull request, regenerated in its own commit so reviewers see exactly what behavior changed.
- One invariant per side effect. Exactly once, never out of bounds, never on the no-op path — name each so a failure tells you which contract broke.
- Script the model boundary, not the tool. Calling your tool function directly bypasses argument validation, hooks, and result conversion — the code you are trying to protect.
- Keep stochastic assertions in the eval layer. Anything depending on a model’s judgment is a score with a threshold, not a
==in CI. - Test the failure path. Inject a model error and a tool exception; assert the agent degrades rather than retries into a three-minute loop.
Wrapping Up
A prompt is prose; a contract is a test. Pin the four things downstream systems depend on — inputs, tool schemas, output shape, side effects — with Pydantic models in strict mode, golden schema files, scripted model boundaries, and property-based probes over tool inputs, all running offline with the model stubbed. That leaves evals to measure quality and e2e to prove the wiring, which is where a real model belongs. Ship the contract tests in CI and a prompt rewrite becomes a routine change instead of a deploy with your fingers crossed.