Versioning Tools: Backward-Compatible Function Evolution for Agents
Treat LLM function-calling tool schemas as a public API you cannot break — additive-only evolution rules, deprecation with dual registration, model behavior on schema changes, and golden tool-call fixtures that catch regressions before deploy.
Published on • October 11, 2026
AI Assistant

A payments team shipped what looked like a cleanup. The tool create_invoice had a customer_id parameter, and a parallel project had standardized on account_id everywhere, so a one-hour refactor renamed the parameter in the tool schema and the handler. Unit tests passed — the new name was wired end to end. Within a day, invoice creation in production dropped by roughly 60%. The models kept emitting customer_id. Sessions that began before the deploy were still executing against the schema fixed at their start; long-lived system prompts cached the old declaration; and eval fixtures pinned to the old recordings never exercised the new path. The schema had changed, but the callers — models, cached contexts, and pinned clients — had not.
That incident is the defining constraint of agent tool design: your function-calling schema is a public API, and your callers include probabilistic consumers you do not control. An HTTP API at least tells you the client version via a header. A tool schema is sent to a model that may have seen similar shapes thousands of times, embedded in contexts that may be cached, and mirrored by clients that pin exactly what they send. You cannot force a coordinated upgrade. You can only evolve additively and deprecate loudly.
This post is about that evolution discipline for LLM function-calling schemas in the Gemini/OpenAI style — function declarations with JSON-Schema parameters — distinct from MCP server versioning we covered earlier. If you treat every parameter rename as the breaking change it is, and you build the telemetry to prove when it is safe to remove the old path, your tools can evolve for years without stranding a single agent.
In this tutorial, you will learn how to:
- Apply additive-only evolution rules to tool schemas — what is safe to change and what silently breaks callers
- Run a naming and deprecation strategy with version suffixes, dual registration, and sunset windows
- Keep schema options open with enum hygiene, sensible defaults, and server-side defaulting
- Predict and test how models behave when schemas change — re-selection, stale calls, and cached contexts
- Execute migrations with parallel tools, adapter shims, and capability negotiation
- Detect old-schema usage with telemetry before you turn anything off
- Protect compatibility with golden tool-call fixtures and contract tests
- Treat descriptions, examples, and docs as part of the behavioral contract
Key technologies: Gemini API function calling (ai.google.dev/gemini-api/docs/function-calling), OpenAI-style function declarations, JSON Schema, Python (TypedDict/Pydantic schema generation), prompt/response fixture testing.
Prerequisites
- Experience defining function/tool declarations for at least one LLM provider
- Familiarity with JSON Schema object/property constructs
- A test runner (pytest) for the fixture examples
Why a tool schema is a public API
The Gemini function-calling model makes the contract explicit: you define a function declaration (name, parameters, and purpose), send it to the model with the user’s prompt, the model decides whether to call it and supplies arguments, and your application executes the function. The model never runs the code; it only ever sees the schema. That means four populations hold references to your schema that you do not deploy when you ship a change:
- In-flight sessions. Multi-turn tool loops carry the declarations they started with. In Gemini’s stateless mode (
store=false), you must resend full history — including the function declarations — on every turn, so a mid-session schema swap produces exactly the stale-call pattern from the incident above. - Provider-side caches. Prompt caches are keyed on request prefixes; tool declarations are part of that prefix. Changing your schema invalidates caches — a cost regression that appears even when calls keep working.
- Pinned clients and recordings. Eval suites, regression recordings, and any client that serializes “the” tool list at build time keep calling the shape you shipped months ago.
- The model’s own selection behavior. Each request, the model re-selects among the tools it is given, guided by names, descriptions, and parameter semantics — a rename is not a rename; it is a new tool competing for selection against a habit the model may not have.
Gemini’s docs note that the API now generates a unique id for every function call — useful for tracing, but it does not version your schema. There is no negotiation protocol that says “caller expects v1”. Evolution discipline is entirely on your side of the wire.
Additive vs breaking changes
Internalize this table before touching any live tool:
| Change | Safe? | Effect on callers |
|---|---|---|
| Add new optional parameter with a default | Yes | Old calls keep working; model may start using it once it appears in the schema |
| Add new required parameter | No | Every existing call site and recorded fixture is now invalid |
| Rename a parameter | No | Models emit the old name (habit, cached context, pinned fixtures); handler misses it |
| Re-type a parameter (e.g. int → string) | No | Serialized calls from old habits fail validation or coerce wrongly |
| Remove a parameter | No | Calls carrying it error; models mid-session still send it |
| Add a new enum value | Careful | Callers with old schemas never emit it (they never see it); your handler must accept old values forever; models with refreshed schemas need a period where both work |
| Remove an enum value | No | Old schemas still emit it; handler must keep accepting or you break live traffic |
Loosen validation (drop a maximum) | Yes | Strict old calls remain valid |
Tighten validation (add a maximum) | No | Previously valid old calls start failing |
| Change a description | Yes, deliberately | Safe mechanically, but it steers model behavior — treat as a behavior change with evals |
| Add a new tool | Yes, with monitoring | New tools compete for selection and can cannibalize calls from existing tools |
| Remove a tool | No | Models mid-session still call it; results in “unknown tool” failures |
| Add output fields | Yes | Agents tolerate extra fields if written defensively |
| Remove or rename output fields | No | Agent parsers and follow-on prompts depend on them |
Two subtleties worth a sentence each. Renaming is two changes, not one: even if you update every caller you own, you cannot update the model’s in-context habits or a partner’s pinned client. And defaults are behavior: changing a server-side default (or adding one where there was none) alters outcomes for every old call that omits the field, with zero schema diff to review.
Schema hygiene that keeps your options open
Design every declaration so future changes stay additive:
from typing import TypedDict, NotRequired
class CreateInvoiceParams(TypedDict):
customer_id: str # stable identifier — never rename for fashion
amount: float
currency: NotRequired[str]
memo: NotRequired[str] # free text: future fields can live here if needed
idempotency_key: NotRequired[str] # lets you evolve semantics safely
CREATE_INVOICE_DECLARATION = {
"name": "create_invoice",
"description": (
"Create an invoice for a customer. "
"currency defaults to the customer's default currency when omitted. "
"Always pass idempotency_key to make retries safe."
),
"parameters": {
"type": "object",
"properties": {
"customer_id": {"type": "string", "description": "Customer identifier"},
"amount": {"type": "number", "description": "Invoice amount (> 0)"},
"currency": {"type": "string", "description": "ISO 4217 code; defaults server-side"},
"memo": {"type": "string", "description": "Optional free-text note"},
"idempotency_key": {"type": "string", "description": "Client-generated key; deduplicates retries"},
},
"required": ["customer_id", "amount"],
},
}
Hygiene rules behind that declaration:
- Prefer strings over enums for open-ended domains. Enums are fine for closed sets the model must pick from (
"priority": {"enum": ["low", "normal", "high"]}), but every value you list is a value you can never remove, and every value you add only reaches models whose schema refresh happened. For anything a backend might extend, accept a string and validate server-side. - Server-side defaults beat schema-required parameters. Omitted optional fields give you a knob that does not exist in the schema: you can change server-side behavior, or interpret “absent” differently per caller, without any schema change at all.
- Keep
requiredminimal and stable. Every parameter you move intorequiredlater is a breaking change; every parameter you start required can never quietly relax without old callers breaking when the server enforces it. - Reserve an extension slot. A free-text
memo/notesfield and anidempotency_keygive models a place to express new intent and give you safe retry semantics — both reduce pressure for impulsive schema changes. - Stable identifiers over readable ones.
customer_idvsaccount_idis not a naming preference once models and clients depend on the old name; if you must align names, add the new one alongside, do not swap.
Naming and deprecation strategy
When a change is genuinely breaking, do not mutate the old tool — add a sibling and retire the old one on a schedule:
# Phase 1: both registered; v1 marked deprecated in description
CREATE_INVOICE_V1 = {
"name": "create_invoice",
"description": (
"[DEPRECATED: use create_invoice_v2] Create an invoice by customer_id. "
"Remains available until 2026-11-30."
),
"parameters": CREATE_INVOICE_DECLARATION["parameters"],
}
CREATE_INVOICE_V2 = {
"name": "create_invoice_v2",
"description": (
"Create an invoice for an account. Preferred over create_invoice. "
"currency defaults server-side; idempotency_key makes retries safe."
),
"parameters": { ... }, # account_id replaces customer_id; additive extras allowed
}
The lifecycle that makes this safe:
- Dual registration. Register v1 and v2 together. Models handle this well — parallel and compositional calling are first-class in the Gemini API, so two similarly named tools do not confuse selection as long as descriptions disambiguate (“preferred over”, “[DEPRECATED: use …]”).
- Deprecation lives in the description. The description is the only instruction channel you have with the model. State the replacement name and the date, in that order.
- Sunset window driven by telemetry, not the calendar. Announce a date, measure old-tool call volume daily, and refuse to remove while the curve is nonzero (see the telemetry section). A 30-day zero-traffic window is a reasonable removal gate.
- Removal is a two-step. First stop offering v1 to models while the handler still accepts it (catches stragglers from pinned clients); then remove the handler. The handler outlives the declaration.
Capability negotiation, when you need something richer than dates: keep a get_capabilities tool or a version field in outputs so agents (or your own orchestrator) can probe what is available and choose v1 or v2 explicitly. For OpenAI- and Gemini-style function calling there is no protocol-level negotiation — the tools you send are the negotiation.
How models behave when schemas change
Plan for three observable behaviors:
- Re-selection. On each turn the model chooses among currently offered tools using names, descriptions, and argument semantics. A new, well-described tool will gradually take traffic; a renamed clone of an old tool splits traffic unpredictably until descriptions force a winner. Measure selection share after any addition — cannibalization is real.
- Stale calls. Mid-session tool calls follow the declarations the session started with, and pinned clients replay old shapes indefinitely. Expect a long tail of old-name calls after every breaking change; if your change was additive, that tail is harmless, which is the entire argument for additive-only evolution.
- Schema-shaped hallucinations. When descriptions contradict parameters, or two tools have near-identical names, models compensate by inventing plausible arguments or picking the wrong sibling. Gemini’s function-calling docs emphasize keeping declarations precise (name, parameters, purpose); ambiguity that a human would tolerate becomes tool-call noise at scale.
One more behavior with a budget impact: because tool declarations are part of the request prefix, schema churn invalidates provider prompt caches. Batch schema changes into scheduled releases rather than continuous cosmetic edits.
Migration patterns for real breaking changes
Four patterns, increasingly heavy, for when additive evolution is not enough:
- Parallel tools (default). Ship
tool_v2besidetool_v1, steer via descriptions, migrate traffic, retire v1 on telemetry. Zero coordination cost; temporary tool-list bloat. - Adapter shim. Keep the old tool name as a thin translation layer over the new implementation:
def handle_create_invoice(args: dict) -> dict:
if "customer_id" in args and "account_id" not in args:
args = {**args, "account_id": lookup_account_id(args.pop("customer_id"))}
return create_invoice_v2_impl(args)
The shim is your compatibility insurance: it absorbs stale calls from cached contexts, pinned clients, and mid-session habits. Budget a removal date, and instrument it (next section) so removal is a data decision. 3. Alias acceptance. Accept both parameter names in the handler without advertising both in the schema. Weaker than a shim (the model will not learn the new name from the schema) but zero tool-list pollution; useful when a rename leaked into a partner integration you do not control. 4. Coordinated cutover with dual-write. For output-breaking changes, return both old and new output shapes for one release, update agents to read the new shape and ignore the old, then drop the old. Additive outputs are always safe; this is how you make removals safe too.
Telemetry: know who is still on the old schema
You cannot remove v1 on a hunch. Instrument at the gateway or handler boundary — log tool name plus parameter key set (never values — that is customer data), tool-call id where the API provides one (Gemini 3 function calls carry unique ids), and session age:
import time
def audit_tool_call(tool: str, args: dict, schema_version: str) -> None:
metrics.increment("tool_calls_total", tags={
"tool": tool,
"schema_version": schema_version, # "v1", "v2"
"param_keys": ",".join(sorted(args.keys())), # fingerprint, not content
"args_match_v2": str(uses_new_shape(args)), # shim hit rate
})
The dashboards that gate removals:
- v1 call volume → must be zero (or shim-only from known pinned clients with a partner exit plan) for a full retention window before deprecation ends.
- Unknown-parameter rate — calls carrying parameters not in the advertised schema, the clearest possible stale-schema signal; a spike after a deploy means something is still pinned.
- Shim translation rate — how often the adapter shim fires; the direct measure of how much life is left in the old shape.
- Per-client (virtual key / team) tool version mix — lets you email the last three holdouts instead of guessing.
Testing backward compatibility
Two complementary test layers, both cheap:
Golden tool-call fixtures. Record real (or realistic) model tool calls as JSON fixtures — tool name plus arguments — and keep them forever:
import json, pathlib
import pytest
FIXTURE_DIR = pathlib.Path("fixtures/tool_calls")
@pytest.mark.parametrize("fixture", sorted(FIXTURE_DIR.glob("*.json")))
def test_all_recorded_calls_still_valid(fixture):
record = json.loads(fixture.read_text())
schema = TOOL_SCHEMAS[record["tool"]] # the schema advertised today
args = record["arguments"]
validate(instance=args, schema=schema["parameters"]) # jsonschema validation
result = execute_tool(record["tool"], args) # handler must not raise
assert "error" not in result
Any additive change passes these by construction; any accidental breaking change fails them before deploy. Where a shim is in place, run each fixture twice: once against the current schema (shim absorbing old shapes) and once re-recorded in the new shape.
Contract tests for the evolution rules. Assert the invariants mechanically: every parameter of the previous release’s schema still exists with a compatible type and sits outside required unless it was already there; no advertised enum value was removed; every deprecated tool still registers while its sunset is pending. Store the previous release’s schema snapshot in-repo as the contract baseline — schema diffs in pull requests then show exactly what a model will see change.
Also worth an eval, not just a contract test: a small suite of prompts that should select the new tool and complete end-to-end, run against both v1-only and v1+v2 registration. That catches selection regressions — the new tool nobody calls — which no schema validator can see.
Documentation and examples are part of the contract
Three artifacts ship with every schema change:
- Descriptions written as instructions to the model, because that is what they are: what the tool does, when to prefer it over a sibling, what the defaults are, what is deprecated and what to use instead.
- Examples in the description or docs showing canonical calls with the new parameters. Models imitate the shapes they are shown; a stale example in your docs is a stale-call generator.
- A changelog entry per tool, in the same repo as the schema, referencing the contract baseline diff. When an unknown-parameter spike appears at 3 a.m., the first question is always “what changed, and when?”
Common pitfalls
- Rename-as-cleanup.
customer_id→account_idfor consistency. The most expensive two-word diff in agent systems; add, never swap. - Adding a required parameter to “force correctness.” Every old call is now invalid. Make it optional with a server-side default that enforces the new correctness.
- Tightening validation on a live tool. Dropping a max, requiring a stricter enum, adding a pattern — old calls that used to pass now fail, and nothing in your diff says “breaking.”
- Removal by calendar. The sunset date arrives, v1 still receives 200 calls a day from a pinned client you forgot, and support opens ten tickets. Calendar informs; telemetry decides.
- Forgetting output compatibility. Additive inputs with breaking outputs just moves the outage from model to agent parser.
- Ignoring selection dynamics. Registering a better-named v2 and assuming traffic moves; measure selection share, and fix descriptions when it does not.
- Schema churn invalidating caches. Cosmetic description edits every deploy quietly multiply prompt-cache misses.
- Testing only the happy path. If your fixtures only contain calls your current code emits, they can never catch a regression against callers who have not upgraded.
Pre-removal checklist
- Change is additive, or a v2 tool + shim exists and both are registered
- Previous release’s schema snapshot updated; PR shows the contract diff
- Descriptions updated: replacement named, deprecation marker and date included
- Golden fixtures from the old shape still execute without error (via shim)
- Contract tests cover: no removed params, no new required params, no removed enum values
- Selection eval confirms the new tool is actually being picked
- Telemetry shows v1 (or old-param) traffic at zero for the full retention window
- Unknown-parameter rate flat post-change; shim hit rate trending to zero
- Output changes are additive, or dual-shape outputs shipped one release ahead
- Changelog entry published; cache-cost impact of the schema change considered
Further reading
- Gemini API — Function calling (declaration flow, parallel and compositional calling, per-call ids,
function_calling_config) - Gemini API documentation hub (interactions, stateless mode, tool integrations)
- Model Context Protocol — Tools specification (input/output schemas and
list_changednotifications — the same contract instincts, protocol-level) - Model Context Protocol
- Related on this blog: Versioning MCP Servers: Backward-Compatible Tool Interfaces