Skip to content
Blog

Contract Testing Your Agents: Asserting Behavior Before Deployment

Stop asserting on prose. Define the inputs, tool schemas, output shapes, and side effects an agent must honor, then test those contracts deterministically in CI.

Published on • October 4, 2026

AI Assistant

Contract Testing Your Agents: Asserting Behavior Before Deployment

Most agent test suites are a folder of prompt strings and hopeful substring checks: send “I want a refund,” assert the answer contains the word “refund.” They pass on Tuesday, fail on Wednesday when the provider silently updates a model, and tell you nothing when they fail — you cannot tell a regression from sampling variance, so you learn to ignore them.

The fix is to stop testing the prose and test the contract: the structured parts of the agent’s behavior that downstream code depends on — inputs, tool schemas, the shape of the structured output, and the side effects a run is allowed to produce. Those are assertable, and once they are asserted, prompt changes and model swaps stop being scary.

Key technologies: Pydantic v2 (models, JSON Schema, strict mode), Pydantic AI (output_type, TestModel, FunctionModel, capture_run_messages), OpenAI Agents SDK (ScriptedModel, ensure_strict_json_schema), pytest, hypothesis.

Why Prompt-Only Tests Are Brittle

A prompt assertion couples your test to three things you do not control: the model’s wording, the provider’s model version, and temperature. The failures that actually ship are structural — a field renamed in the output schema, a tool argument that became optional, a refund() call that now fires twice — and a substring test is blind to all of them, because none of them change the prose.

Contract testing inverts the priority: assert the invariants that would break an integration, and leave “is this answer good?” to the eval layer.

What a Contract Is for an Agent

Contract facetExample assertionBreaks when
InputsPrompt payload validates against the request modelCallers start passing partial data
Tool schemasGenerated JSON Schema matches a committed golden fileA signature, docstring, or default changes
Structured output shapeEvery run returns a model that validates, strictlyA field is renamed, retyped, or made optional
Side-effect invariantsExactly one refund() call, amount within boundsThe model retries, loops, or hallucinates a call
Turn shapeThe scripted sequence of model calls is consumed exactlyThe workflow gains or loses a step

The first four are ordinary unit and contract tests. The fifth is where agent frameworks earn their keep.

Structured Output Is the Strongest Contract You Have

Pydantic is the natural place to write it down: per the Pydantic docs, “schema validation and serialization are controlled by type annotations,” and models emit JSON Schema directly — exactly what the LLM is told to produce.

# contracts/output.py
from typing import Literal

from pydantic import BaseModel, ConfigDict, Field


class RefundDecision(BaseModel):
    model_config = ConfigDict(strict=True)

    refund: bool
    amount_usd: float = Field(ge=0, le=10_000)
    reason: Literal["duplicate", "not_received", "damaged", "none"]
    order_id: str = Field(pattern=r"^ord_[a-z0-9]+$")

Strict mode is the part people skip. By default Pydantic coerces: '123' becomes 123. With strict=True — settable per call via model_validate(payload, strict=True), per field with Field(strict=True), or model-wide with ConfigDict(strict=True) — it raises instead. That is the difference between “the agent produced a refund amount” and “the agent produced something we could coerce into a refund amount.”

# tests/test_output_contract.py
import pytest
from pydantic import ValidationError

from contracts.output import RefundDecision


def test_valid_decision_round_trips():
    decision = RefundDecision.model_validate(
        {"refund": True, "amount_usd": 12.5, "reason": "duplicate",
         "order_id": "ord_7"},
        strict=True,
    )
    assert decision.amount_usd == 12.5


@pytest.mark.parametrize("payload", [
    {"refund": "yes", "amount_usd": 1, "reason": "none", "order_id": "ord_1"},
    {"refund": True, "amount_usd": -1, "reason": "none", "order_id": "ord_1"},
    {"refund": True, "amount_usd": 1, "reason": "maybe", "order_id": "ord_1"},
    {"refund": True, "amount_usd": 1, "reason": "none", "order_id": "7"},
])
def test_shape_is_a_hard_contract(payload):
    with pytest.raises(ValidationError):
        RefundDecision.model_validate(payload, strict=True)

Wire the same model to the agent as its output_type and the guarantee becomes runtime, not just test-time: Pydantic AI states “the run is guaranteed to return a RefundDecision,” and if validation fails the agent is prompted to try again. On the OpenAI Agents SDK side, agents.strict_schema.ensure_strict_json_schema(schema) mutates a JSON schema so it conforms to the strict standard the OpenAI API expects — worth asserting before a provider call ever happens.

Tool-Schema Golden Tests

Tool schemas are a public API: the model reads them to decide what to call, and Pydantic AI passes your signature and docstring through as the schema — “the rest of its signature and its docstring become the tool schema, arguments are validated before your code runs.” A renamed parameter or dropped docstring changes model behavior with zero type errors.

Snapshot it:

# tests/test_tool_schema_golden.py
import json
import pathlib

from contracts.tools import LookupOrderArgs

GOLDEN = pathlib.Path(__file__).parent / "golden" / "lookup_order.json"


def test_lookup_order_schema_has_not_drifted():
    actual = LookupOrderArgs.model_json_schema()
    if not GOLDEN.exists():
        GOLDEN.parent.mkdir(exist_ok=True)
        GOLDEN.write_text(json.dumps(actual, indent=2, sort_keys=True))
        raise AssertionError("golden file created — review and commit it")
    expected = json.loads(GOLDEN.read_text())
    assert actual == expected, "schema drifted; regenerate if intentional"

Now a schema change is a red diff in review instead of a quiet behavior change in production.

Replaying Recorded Tool Calls

The second half of the contract is the sequence: does the agent call the tool it should, once, with the right arguments, then stop? Both SDKs let you script the model boundary so the answer is deterministic.

# tests/test_refund_flow.py
import pytest

from agents import Agent, RunConfig, Runner
from agents.decorators import tool
from agents.testing import ScriptedModel, assistant_message, function_call

refunds: list[tuple[str, float]] = []


@tool
def refund(order_id: str, amount_usd: float) -> str:
    """Issue a refund for an order."""
    refunds.append((order_id, amount_usd))
    return f"Refunded ${amount_usd:.2f}."


@pytest.mark.asyncio
async def test_refund_flow_calls_tool_once_then_answers():
    refunds.clear()
    model = ScriptedModel([
        [function_call("refund", {"order_id": "ord_7", "amount_usd": 12.5},
                       call_id="c1")],
        [assistant_message("Refunded $12.50.")],
    ])
    agent = Agent(name="Support", model=model, tools=[refund])

    result = await Runner.run(
        agent, "refund ord_7",
        run_config=RunConfig(tracing_disabled=True),
    )

    assert result.final_output == "Refunded $12.50."
    assert refunds == [("ord_7", 12.5)]          # side-effect invariant
    assert len(model.calls) == 2
    model.assert_complete()

ScriptedModel records every call (calls, first_call, last_call), and assert_complete() fails if the workflow ended before consuming every scripted step. The SDK raises structured errors — UnexpectedModelCall for an extra request, UnconsumedModelSteps for an early exit — so drift shows up as a typed failure, not a timeout. Under the hood the real tool pipeline runs: schema generation, argument validation, execution, handoff of the result to the next model turn.

Pydantic AI exposes the same idea from the message side: capture_run_messages() records the ModelRequest / ModelResponse exchange, so you can assert the ToolCallPart name and args, the ToolReturnPart, and the final TextPart while TestModel or FunctionModel stands in for the model.

Property-Based Tests for Tool Inputs

Example-based tests only probe the inputs you thought of. A tool that parses dates, computes money, or builds SQL deserves generated inputs:

# tests/test_refund_properties.py
from hypothesis import given, settings, strategies as st

from tools.money import apply_refund


@given(
    amount_usd=st.floats(min_value=0, max_value=10_000, allow_nan=False),
    fee_rate=st.floats(min_value=0, max_value=0.25, allow_nan=False),
)
@settings(max_examples=300)
def test_refund_is_always_within_bounds(amount_usd: float, fee_rate: float):
    result = apply_refund(amount_usd, fee_rate)
    assert 0 <= result <= amount_usd


@given(order_id=st.text(max_size=200))
@settings(max_examples=300)
def test_lookup_never_raises_on_arbitrary_ids(order_id: str):
    assert lookup_order(order_id) is not None

The point is not coverage of your example table — it is that a tool the model can call with arbitrary generated text must not raise, must not return a negative refund, and must not accept a ') OR 1=1 -- order ID as valid. Hypothesis hands you each counterexample as a minimal repro you can paste into a fixture.

Stubbing the Model for Deterministic CI

None of the above should touch a real provider. Pydantic AI recommends the standard setup: use TestModel or FunctionModel in place of your real model “to avoid the usage, latency and variability of real LLM calls,” swap it in with agent.override(model=...), and set ALLOW_MODEL_REQUESTS = False globally so an accidental real request fails the run instead of quietly costing money.

# tests/conftest.py
import pytest

from pydantic_ai import models
from pydantic_ai.models.test import TestModel

models.ALLOW_MODEL_REQUESTS = False


@pytest.fixture
def stubbed_agent(agent):
    with agent.override(model=TestModel()):
        yield

TestModel calls every registered tool, then produces a response shaped by your output_type — procedural schema-driven data, no model. When you need a specific turn sequence (call the tool first, return a future date, fail on the second attempt), FunctionModel takes a plain function that receives the message list and returns the scripted ModelResponse. On the Agents SDK side, ScriptedModel plus RunConfig(tracing_disabled=True) does the same job, and the testing modules make no model or network requests.

The Test Pyramid for Agents

LayerWhat it assertsReal modelRuntimeWhen
UnitPure functions, tool bodies, money and date mathNoMillisecondsEvery commit
ContractSchemas, output shape, side effects, scripted turnsStubbedUnder a secondEvery commit
EvalTask success and quality scores on a fixed datasetYesMinutesNightly
E2EProvider auth, network, retries, streaming, real toolsYesMinutes to hoursPre-release

Pydantic Evals is built for the third row — testing “agent behavior the way pytest tests code.” The rule of thumb: unit through contract must be free, deterministic, and runnable on a laptop; only eval and e2e may be expensive or stochastic, and they report scores, not booleans.

Best Practices

  1. Assert structure, not wording. Compare parsed output and recorded tool calls, never the sentences around them.
  2. Fail loudly if CI reaches a provider. ALLOW_MODEL_REQUESTS = False in conftest.py, tracing_disabled=True in scripted runs.
  3. Golden files are reviewed artifacts. A schema diff belongs in the pull request, regenerated in its own commit so reviewers see exactly what behavior changed.
  4. One invariant per side effect. Exactly once, never out of bounds, never on the no-op path — name each so a failure tells you which contract broke.
  5. Script the model boundary, not the tool. Calling your tool function directly bypasses argument validation, hooks, and result conversion — the code you are trying to protect.
  6. Keep stochastic assertions in the eval layer. Anything depending on a model’s judgment is a score with a threshold, not a == in CI.
  7. Test the failure path. Inject a model error and a tool exception; assert the agent degrades rather than retries into a three-minute loop.

Wrapping Up

A prompt is prose; a contract is a test. Pin the four things downstream systems depend on — inputs, tool schemas, output shape, side effects — with Pydantic models in strict mode, golden schema files, scripted model boundaries, and property-based probes over tool inputs, all running offline with the model stubbed. That leaves evals to measure quality and e2e to prove the wiring, which is where a real model belongs. Ship the contract tests in CI and a prompt rewrite becomes a routine change instead of a deploy with your fingers crossed.

Sources