Skip to content
Blog

Idempotent Tool Design: Making Every Agent Call Safe to Repeat

Agents retry constantly, and naive tools double-charge or duplicate rows when they do. Apply Stripe-style idempotency keys, natural keys, and dedupe patterns so every tool call is safe to repeat.

Published on • October 4, 2026

AI Assistant

Introduction

Your agent called charge_customer twice. The customer was billed twice. The support ticket says “the AI messed up,” but the real bug is older than the model: you wrote a tool that assumed exactly-once execution in an at-least-once world.

Agents are compulsive retryers: LLM loops re-issue calls after malformed responses, HTTP clients retry on timeouts and 500s, orchestrators replay steps after crashes, and models sometimes just call the same tool twice because it seemed like a good idea. Every path turns one logical action into multiple physical requests. If your tool has side effects and no dedupe strategy, retries become duplicates.

Stripe’s idempotency model is the cleanest template available for agent tool design. Let’s apply it.

Why Agents Retry

Three retry sources, all normal:

  • LLM loops: a tool returns output the model doesn’t like, or the orchestrator enforces a max-iteration policy — the model calls again with slightly different arguments.
  • Timeouts: the request succeeded server-side, but the response never arrived. The client sees a timeout and retries. The work already happened.
  • Transient errors: 429s, 502s, connection resets. Any sane SDK retries these.

A naive POST /orders tool executed under any of these paths inserts two rows. A naive send_email tool mails the customer twice. The failure mode is invisible in development and catastrophic in production.

Idempotency Keys, Stripe’s Model

Stripe’s docs describe the pattern precisely: an idempotency key lets you “safely repeat requests without accidentally performing the same operation twice.” The mechanics worth stealing:

  1. Client generates the key. A unique string — Stripe suggests V4 UUIDs or random strings with enough entropy — sent alongside the request. Keys are capped at 255 characters, and you should avoid putting sensitive data (emails, IDs of people) in them.
  2. Server saves the first result. Stripe saves “the resulting status code and body of the first request… regardless of whether it succeeds or fails.” Subsequent requests with the same key return that stored response — including 500 errors. Replay, not re-execute.
  3. Request fingerprinting. If a key is reused with different parameters, the layer errors rather than silently returning the wrong response. Same key must mean same request.
  4. Retention window. Keys can be pruned once “at least 24 hours old.” After pruning, the same key is treated as new. The window is your retry horizon: any retry longer than 24h needs a different strategy.
  5. POST only. “All POST requests accept idempotency keys. Don’t send idempotency keys in GET and DELETE requests” — those are already idempotent by definition.

One subtlety that matters: Stripe only saves results after execution begins. Validation failures and concurrent conflicts store nothing, so those remain safely retryable.

Designing Tools with Client-Generated Keys

The key insight for agents: the agent framework (client) generates the key, the tool (server) honors it. The key must be minted once per logical intent — not once per attempt.

import uuid

# Minted when the agent forms the intent; reused across every retry.
idem_key = str(uuid.uuid4())

result = toolhouse.call(
    "create_order",
    args={"sku": "PRO-1", "qty": 2},
    idempotency_key=idem_key,
)

Inside the tool:

def create_order(args: dict, idempotency_key: str) -> dict:
    existing = db.query(
        "SELECT response FROM idempotency_log WHERE key = %s",
        (idempotency_key,),
    )
    if existing:
        return json.loads(existing.response)  # replay, don't re-execute

    fingerprint = hash_request(args)
    order = insert_order(args)                 # the side effect
    response = serialize(order)

    db.execute(
        "INSERT INTO idempotency_log (key, fingerprint, response) "
        "VALUES (%s, %s, %s)",
        (idempotency_key, fingerprint, json.dumps(response)),
    )
    return response

The fingerprint check catches the dangerous case: the model retries with the same key but mutated arguments (say, a different quantity). You reject with a clear error instead of returning a response that doesn’t match what the model thinks it did.

Natural Keys and Upserts

Not every tool call comes with a convenient UUID. Often the domain itself provides a natural key — an order reference, a slug, an email plus date. When it does, prefer making the operation idempotent structurally:

INSERT INTO orders (customer_id, sku, qty, status)
VALUES (%s, %s, %s, 'pending')
ON CONFLICT (customer_id, sku, created_on) DO NOTHING;

Or, in databases without upsert, guard explicitly:

INSERT INTO orders (...)
SELECT %s, %s, %s
WHERE NOT EXISTS (
    SELECT 1 FROM orders
    WHERE customer_id = %s AND sku = %s AND created_on = CURRENT_DATE
);

This is often better than an idempotency table: the uniqueness constraint lives with the data and survives schema migrations you forget about. The rule of thumb — if the domain knows what “the same thing” means, use a natural key; if only your workflow knows, use an idempotency key.

PUT vs POST Semantics for Agent Tools

Your tool’s HTTP verb encodes a contract about repetition:

  • POST /orders — “create something new.” Not idempotent. Repeat it and you get two orders. Requires a key or natural key.
  • PUT /orders/{id} — “make the resource look like this.” Repeat it and the state is identical. Inherently safe.
  • PATCH with partial updates — idempotent only if you’re setting absolute values, not increments (balance = balance + 10 is not idempotent; balance = 100 is).

For agent tools, prefer PUT-shaped operations whenever the domain allows it: set_order_status(order_id, "shipped") instead of advance_order(order_id); upsert_contact(email, fields) instead of create_contact(fields). Agents will re-run things — design tools where re-running is a no-op.

When you genuinely need POST, route it through the idempotency key pattern.

Dedupe Tables and the Outbox Pattern

For flows that span multiple systems — insert a row, then charge a card, then email a receipt — single-row idempotency isn’t enough. You need two pieces:

A dedupe table recording which logical operations have completed:

CREATE TABLE processed_events (
    idem_key     TEXT PRIMARY KEY,
    payload_hash TEXT NOT NULL,
    status       TEXT NOT NULL,      -- 'started' | 'done'
    response     JSONB,
    created_at   TIMESTAMPTZ DEFAULT now()
);

The outbox pattern: write the intent and the side effect in the same database transaction, then let a worker drain the outbox to external systems (payment gateway, email provider). Each outbox row carries the idempotency key, so the worker can retry the external call safely forever. This closes the classic gap where your database committed but the payment call died — or vice versa.

Partial Failures and the Exactly-Once Illusion

Be honest with your team: exactly-once delivery does not exist across a network. What you can build is at-least-once delivery plus idempotent processing, which is observationally equivalent to exactly-once for the consumer:

  1. Execute the side effect at least once.
  2. Record completion atomically with the effect (or dedupe on read).
  3. Make every downstream consumer tolerant of replays.

Stripe’s choice to replay the original response — including errors — is what makes this hold: the client can’t distinguish a retry from the first call, so the retry never changes observable state.

Test Patterns: Prove It’s Safe to Repeat

Idempotency claims are worthless untested. Two tests catch nearly every regression:

def test_create_order_tool_is_idempotent():
    key = str(uuid.uuid4())
    first  = create_order({"sku": "PRO-1", "qty": 2}, idempotency_key=key)
    second = create_order({"sku": "PRO-1", "qty": 2}, idempotency_key=key)

    assert first["order_id"] == second["order_id"]
    assert count_orders(sku="PRO-1") == 1   # exactly one side effect

def test_key_reuse_with_different_args_rejected():
    key = str(uuid.uuid4())
    create_order({"sku": "PRO-1", "qty": 2}, idempotency_key=key)
    with pytest.raises(IdempotencyConflict):
        create_order({"sku": "PRO-1", "qty": 9}, idempotency_key=key)

Add a third: simulate a crash between the side effect and the log write, then retry — your design should recover (this is where the single-transaction approach earns its keep). Run these in CI against a real Postgres, not mocks; the bug lives in the constraint.

Worked Example: create_order

Pulling it together for a typical agent tool:

def create_order(customer_id: str, sku: str, qty: int,
                 idempotency_key: str) -> Order:
    with db.transaction():                       # 1. atomic scope
        row = db.get("processed_events", idempotency_key)
        if row:
            if row.payload_hash != hash_args(customer_id, sku, qty):
                raise IdempotencyConflict("key reused with new payload")
            return Order.from_row(row.response)  # 2. replay

        order = Order.insert(                    # 3. side effect
            customer_id=customer_id, sku=sku, qty=qty, status="pending"
        )
        db.insert("outbox", {                    # 4. intent + effect together
            "idem_key": idempotency_key,
            "action": "charge_card",
            "payload": {"order_id": order.id, "sku": sku},
        })
        db.insert("processed_events", {
            "idem_key": idempotency_key,
            "payload_hash": hash_args(customer_id, sku, qty),
            "status": "done",
            "response": order.serialize(),
        })
    dispatcher.flush()                           # 5. at-least-once, key-safe
    return order

The agent can now crash, time out, or loop on this tool all it likes — it sees the same order object every time, the card is charged once, and the receipt goes out once.

Sources