Skip to content
Blog

PydanticAI vs. OpenAI Agents SDK: A Side-by-Side Build Comparison

Build the same triage agent twice - once in PydanticAI, once in the OpenAI Agents SDK - and compare typing, tools, guardrails, sessions, and observability.

Published on • October 6, 2026

AI Assistant

Both PydanticAI and the OpenAI Agents SDK let you build production agents in Python, and both ship batteries included. Yet they optimize for different things, and the difference only shows once you build. So let’s build - the same support-triage agent, twice, then compare where the code lands.

The spec: an agent that takes a support request, classifies urgency, can look up an order, either answers or hands off to a specialist agent, and refuses anything outside policy. Structured output required.

Round 1: PydanticAI

from pydantic import BaseModel, Field
from pydantic_ai import Agent, RunContext

class Triage(BaseModel):
    urgency: Literal['low', 'medium', 'high']
    needs_human: bool = Field(description='Escalate to a human agent')
    summary: str

@dataclass
class Deps:
    db: DatabaseConn
    user: User

triage_agent = Agent(
    'anthropic:claude-fable-5-1',
    deps_type=Deps,
    output_type=Triage,          # structured output, validated
    instructions='Classify support requests. Never guess order status.',
)

@triage_agent.tool
def lookup_order(ctx: RunContext[Deps], order_id: str) -> Order | None:
    """Look up an order by ID for the current user."""
    return ctx.deps.db.find_order(order_id, ctx.deps.user.id)

Key PydanticAI moves visible immediately:

  • Types are the contract. output_type=Triage means every run returns a validated Triage - the Pydantic model generates the JSON schema for the LLM and validates the response. A bad response triggers a repair-and-retry loop automatically.
  • Dependency injection is typed. RunContext[Deps] carries the database and user; the type checker catches a wrong deps_type at authoring time.
  • The tool signature is the tool schema. Docstring → description, parameter annotations → JSON schema, RunContext stripped out.

Guardrails and structure stay declarative:

@triage_agent.output_guardrail
async def no_empty_summary(ctx: RunContext[Deps], out: Triage) -> None:
    if len(out.summary.strip()) < 10:
        raise OutputGuardrailException('summary too short')

Round 2: OpenAI Agents SDK

from agents import Agent, Runner, function_tool
from pydantic import BaseModel

class Triage(BaseModel):
    urgency: Literal['low', 'medium', 'high']
    needs_human: bool
    summary: str

@function_tool
def lookup_order(order_id: str, user_id: str) -> Order | None:
    """Look up an order by ID for the current user."""
    return db.find_order(order_id, user_id)

triage_agent = Agent(
    name='Triage',
    instructions='Classify support requests. Never guess order status.',
    tools=[lookup_order],
    output_type=Triage,
)

result = Runner.run_sync(triage_agent, 'My order never arrived')
print(result.final_output)  # Triage

The SDK’s philosophy is “few primitives, Python-first”: agents (instructions + tools), handoffs (delegate to another agent), guardrails (input/output checks run in parallel), plus built-in tracing, sessions, and human-in-the-loop - but orchestration is plain Python control flow, not a framework abstraction.

Handoff to a specialist is the SDK’s signature:

billing_agent = Agent(name='Billing', instructions='Handle billing disputes...', tools=[...])

triage_agent = Agent(
    name='Triage',
    instructions='Route billing issues onward.',
    tools=[lookup_order],
    handoffs=[billing_agent],
)

In PydanticAI, the equivalent is either an explicit call to another agent’s run(), an agent-as-tool, or Pydantic Graph for typed multi-step control flow - more explicit, more structure, more code.

The comparison, dimension by dimension

DimensionPydanticAIOpenAI Agents SDK
Core philosophyTyped everything; correctness by constructionFew primitives; Python control flow
Structured outputoutput_type= + validation + auto-repairoutput_type= (Pydantic supported)
Dependency injectionTyped deps_type / RunContext[T]Function args / closures / context objects
ToolsFunction decorators, schema from signature@function_tool, schema from signature
GuardrailsOutput/input guardrail decoratorsParallel input/output guardrail agents
Multi-agentExplicit calls, agent-as-tools, Pydantic Graphhandoffs=[...], agents-as-tools
Sessions/memoryPersistence APIs + durable execution enginesBuilt-in Sessions (SQLite/SQLAlchemy/Redis…)
Human-in-the-loopDeferred tools (human approval on tool execution)First-class human_in_the_loop module
ObservabilityOpenTelemetry-native → LogfireBuilt-in tracing → OpenAI platform tools
TestingTestModel (offline), Pydantic Evalstesting module with ModelTest
Provider supportVirtually every provider, string-swappableMulti-provider, OpenAI-first defaults (Responses API)
DurabilityTemporal/DBOS/Prefect/… integrationsNot the focus (sessions ≠ durable execution)
ExtrasCapabilities (web search, MCP, image gen), HarnessSandbox agents, voice/realtime pipelines, MCP

Where PydanticAI pulls ahead

  1. Type-driven correctness. If your team lives in mypy/pyright and PR reviews catch type errors, PydanticAI’s RunContext[Deps], generic Agent[Deps, Output], and validated outputs move whole error classes from runtime to compile time.
  2. Structured output as a guarantee. Validation failure → the model is re-prompted with the error until the output conforms (or the run fails loudly). “The JSON was fine but the field was wrong” disappears.
  3. Provider neutrality with a typed surface. Swapping openai:gpt-6-sol for anthropic:claude-fable-5-1 is a string change; the typed agent contract doesn’t move.
  4. Durability. Need an agent that survives restarts for days? Attach Temporal/DBOS durability and every model/tool call becomes a durable activity. This is a category the SDK doesn’t compete in.
  5. Observability posture. OTel-native instrumentation means any OTLP backend works; Logfire one-liners light up traces with costs attached.

Where the OpenAI Agents SDK pulls ahead

  1. Zero-ceremony start. Agent(...) + Runner.run_sync(...) - hello-world in five lines, and the “when in doubt, use Python” model means no new abstraction to learn for orchestration.
  2. Handoffs as a primitive. Multi-agent routing (handoffs=[...]) is a first-class, platform-traced concept rather than code you write.
  3. Batteries for the OpenAI surface. Voice pipelines, realtime agents with interruption handling, sandbox agents with isolated workspaces, and evaluation/fine-tuning hooks into the OpenAI platform - these ship in the SDK.
  4. Parallel guardrails. Input guardrails run concurrently with the agent, failing fast without blocking the whole run.
  5. Sessions out of the box. Session implementations (SQLite, SQLAlchemy, Redis, encrypted variants) give working memory across turns with a one-line choice.

Decision matrix

Choose PydanticAI when:

  • Type safety and validated outputs are non-negotiable (fintech, health, data pipelines).
  • You’re multi-provider today or plausibly tomorrow.
  • You need durable execution, Graph-shaped workflows, or heavyweight evals as a first-class part of the design.
  • Observability must go to your existing OTel stack.

Choose the OpenAI Agents SDK when:

  • You’re building on OpenAI and want the shallowest possible learning curve.
  • Your architecture is handoffs/human-in-the-loop/voice - the SDK’s native vocabulary.
  • You want built-in tracing and sessions with minimal setup.
  • You may need sandboxed code execution or realtime voice without adding more dependencies.

Common ground worth noting: both use Pydantic for schemas, both support MCP tools, both treat structured outputs and guardrails as core rather than plugins, and both are production-tested. The “wrong” choice is rarely fatal; the migration cost is real but bounded.

Key Takeaways

  1. The same triage agent takes similar shape in both - the divergence is in typing, orchestration primitives, and platform batteries.
  2. PydanticAI optimizes for type-driven correctness: RunContext[Deps], validated output_type with auto-repair, OTel observability, durable execution.
  3. The OpenAI Agents SDK optimizes for few-primitives Python: handoffs, parallel guardrails, built-in sessions/tracing, voice and sandbox agents included.
  4. Pick by constraint: multi-provider + types + durability → PydanticAI; OpenAI-centric + handoffs + batteries → Agents SDK.
  5. Both support MCP and Pydantic schemas - prototype in either, but decide before your orchestration logic hardens around one model’s abstractions.

References: