PydanticAI vs. OpenAI Agents SDK: A Side-by-Side Build Comparison
Build the same triage agent twice - once in PydanticAI, once in the OpenAI Agents SDK - and compare typing, tools, guardrails, sessions, and observability.
Published on • October 6, 2026
AI Assistant

Both PydanticAI and the OpenAI Agents SDK let you build production agents in Python, and both ship batteries included. Yet they optimize for different things, and the difference only shows once you build. So let’s build - the same support-triage agent, twice, then compare where the code lands.
The spec: an agent that takes a support request, classifies urgency, can look up an order, either answers or hands off to a specialist agent, and refuses anything outside policy. Structured output required.
Round 1: PydanticAI
from pydantic import BaseModel, Field
from pydantic_ai import Agent, RunContext
class Triage(BaseModel):
urgency: Literal['low', 'medium', 'high']
needs_human: bool = Field(description='Escalate to a human agent')
summary: str
@dataclass
class Deps:
db: DatabaseConn
user: User
triage_agent = Agent(
'anthropic:claude-fable-5-1',
deps_type=Deps,
output_type=Triage, # structured output, validated
instructions='Classify support requests. Never guess order status.',
)
@triage_agent.tool
def lookup_order(ctx: RunContext[Deps], order_id: str) -> Order | None:
"""Look up an order by ID for the current user."""
return ctx.deps.db.find_order(order_id, ctx.deps.user.id)
Key PydanticAI moves visible immediately:
- Types are the contract.
output_type=Triagemeans every run returns a validatedTriage- the Pydantic model generates the JSON schema for the LLM and validates the response. A bad response triggers a repair-and-retry loop automatically. - Dependency injection is typed.
RunContext[Deps]carries the database and user; the type checker catches a wrongdeps_typeat authoring time. - The tool signature is the tool schema. Docstring → description, parameter annotations → JSON schema,
RunContextstripped out.
Guardrails and structure stay declarative:
@triage_agent.output_guardrail
async def no_empty_summary(ctx: RunContext[Deps], out: Triage) -> None:
if len(out.summary.strip()) < 10:
raise OutputGuardrailException('summary too short')
Round 2: OpenAI Agents SDK
from agents import Agent, Runner, function_tool
from pydantic import BaseModel
class Triage(BaseModel):
urgency: Literal['low', 'medium', 'high']
needs_human: bool
summary: str
@function_tool
def lookup_order(order_id: str, user_id: str) -> Order | None:
"""Look up an order by ID for the current user."""
return db.find_order(order_id, user_id)
triage_agent = Agent(
name='Triage',
instructions='Classify support requests. Never guess order status.',
tools=[lookup_order],
output_type=Triage,
)
result = Runner.run_sync(triage_agent, 'My order never arrived')
print(result.final_output) # Triage
The SDK’s philosophy is “few primitives, Python-first”: agents (instructions + tools), handoffs (delegate to another agent), guardrails (input/output checks run in parallel), plus built-in tracing, sessions, and human-in-the-loop - but orchestration is plain Python control flow, not a framework abstraction.
Handoff to a specialist is the SDK’s signature:
billing_agent = Agent(name='Billing', instructions='Handle billing disputes...', tools=[...])
triage_agent = Agent(
name='Triage',
instructions='Route billing issues onward.',
tools=[lookup_order],
handoffs=[billing_agent],
)
In PydanticAI, the equivalent is either an explicit call to another agent’s run(), an agent-as-tool, or Pydantic Graph for typed multi-step control flow - more explicit, more structure, more code.
The comparison, dimension by dimension
| Dimension | PydanticAI | OpenAI Agents SDK |
|---|---|---|
| Core philosophy | Typed everything; correctness by construction | Few primitives; Python control flow |
| Structured output | output_type= + validation + auto-repair | output_type= (Pydantic supported) |
| Dependency injection | Typed deps_type / RunContext[T] | Function args / closures / context objects |
| Tools | Function decorators, schema from signature | @function_tool, schema from signature |
| Guardrails | Output/input guardrail decorators | Parallel input/output guardrail agents |
| Multi-agent | Explicit calls, agent-as-tools, Pydantic Graph | handoffs=[...], agents-as-tools |
| Sessions/memory | Persistence APIs + durable execution engines | Built-in Sessions (SQLite/SQLAlchemy/Redis…) |
| Human-in-the-loop | Deferred tools (human approval on tool execution) | First-class human_in_the_loop module |
| Observability | OpenTelemetry-native → Logfire | Built-in tracing → OpenAI platform tools |
| Testing | TestModel (offline), Pydantic Evals | testing module with ModelTest |
| Provider support | Virtually every provider, string-swappable | Multi-provider, OpenAI-first defaults (Responses API) |
| Durability | Temporal/DBOS/Prefect/… integrations | Not the focus (sessions ≠ durable execution) |
| Extras | Capabilities (web search, MCP, image gen), Harness | Sandbox agents, voice/realtime pipelines, MCP |
Where PydanticAI pulls ahead
- Type-driven correctness. If your team lives in
mypy/pyrightand PR reviews catch type errors, PydanticAI’sRunContext[Deps], genericAgent[Deps, Output], and validated outputs move whole error classes from runtime to compile time. - Structured output as a guarantee. Validation failure → the model is re-prompted with the error until the output conforms (or the run fails loudly). “The JSON was fine but the field was wrong” disappears.
- Provider neutrality with a typed surface. Swapping
openai:gpt-6-solforanthropic:claude-fable-5-1is a string change; the typed agent contract doesn’t move. - Durability. Need an agent that survives restarts for days? Attach Temporal/DBOS durability and every model/tool call becomes a durable activity. This is a category the SDK doesn’t compete in.
- Observability posture. OTel-native instrumentation means any OTLP backend works; Logfire one-liners light up traces with costs attached.
Where the OpenAI Agents SDK pulls ahead
- Zero-ceremony start.
Agent(...)+Runner.run_sync(...)- hello-world in five lines, and the “when in doubt, use Python” model means no new abstraction to learn for orchestration. - Handoffs as a primitive. Multi-agent routing (
handoffs=[...]) is a first-class, platform-traced concept rather than code you write. - Batteries for the OpenAI surface. Voice pipelines, realtime agents with interruption handling, sandbox agents with isolated workspaces, and evaluation/fine-tuning hooks into the OpenAI platform - these ship in the SDK.
- Parallel guardrails. Input guardrails run concurrently with the agent, failing fast without blocking the whole run.
- Sessions out of the box.
Sessionimplementations (SQLite, SQLAlchemy, Redis, encrypted variants) give working memory across turns with a one-line choice.
Decision matrix
Choose PydanticAI when:
- Type safety and validated outputs are non-negotiable (fintech, health, data pipelines).
- You’re multi-provider today or plausibly tomorrow.
- You need durable execution, Graph-shaped workflows, or heavyweight evals as a first-class part of the design.
- Observability must go to your existing OTel stack.
Choose the OpenAI Agents SDK when:
- You’re building on OpenAI and want the shallowest possible learning curve.
- Your architecture is handoffs/human-in-the-loop/voice - the SDK’s native vocabulary.
- You want built-in tracing and sessions with minimal setup.
- You may need sandboxed code execution or realtime voice without adding more dependencies.
Common ground worth noting: both use Pydantic for schemas, both support MCP tools, both treat structured outputs and guardrails as core rather than plugins, and both are production-tested. The “wrong” choice is rarely fatal; the migration cost is real but bounded.
Key Takeaways
- The same triage agent takes similar shape in both - the divergence is in typing, orchestration primitives, and platform batteries.
- PydanticAI optimizes for type-driven correctness:
RunContext[Deps], validatedoutput_typewith auto-repair, OTel observability, durable execution. - The OpenAI Agents SDK optimizes for few-primitives Python: handoffs, parallel guardrails, built-in sessions/tracing, voice and sandbox agents included.
- Pick by constraint: multi-provider + types + durability → PydanticAI; OpenAI-centric + handoffs + batteries → Agents SDK.
- Both support MCP and Pydantic schemas - prototype in either, but decide before your orchestration logic hardens around one model’s abstractions.
References: