Idempotency Keys in Agent Payment Loops: Retry Without Double-Charge
The full lifecycle of a payment idempotency key inside an agentic retry loop: mint it per logical intent, persist it before the tool call, replay stored results, scope and fingerprint keys, and reconcile for double charges.
Published on • October 11, 2026
AI Assistant

At 02:14 a checkout agent hit a tool timeout. The card had been charged; the response never arrived. The orchestrator retried the charge_customer tool, the model regenerated a fresh UUID for the retry, and the customer saw two identical charges on their statement. By the time reconciliation flagged the pair, the refund workflow had already opened three tickets — one for each attempt the loop made after the payment provider’s rate limiter kicked in.
This is not a model-intelligence problem. Agents retry by design: LLM loops re-issue calls after malformed tool output, HTTP clients retry on timeouts and 5xx, orchestrators replay steps after crashes, and planners re-derive “pay the invoice” from scratch after a context compaction. Every path turns one logical payment intent into multiple physical API requests. The fix is old and well-proven — Stripe-style idempotency keys — but the hard part in an agent system is not the API header. It is the key lifecycle: where the key is born, when it is persisted, how it is scoped, and what the loop does on replay. That is what this post covers.
In this tutorial, you will learn how to:
- Explain why agent loops make payment retries inevitable
- Describe Stripe-style idempotency semantics: 24-hour replay, parameter fingerprinting, and result storage
- Decide where an idempotency key must originate — per logical intent, never per attempt
- Generate and durably persist keys before the payment tool call
- Wrap a payment tool in a state machine: intent → key minted → call → record → replay
- Scope keys per account and endpoint, and fingerprint payloads to catch misuse
- Detect double charges with a reconciliation job
- Test idempotency so regressions fail in CI, not in production
Key technologies: Stripe idempotent requests (Idempotency-Key header), Python 3.12, Postgres, Pydantic, webhook-driven reconciliation, durable execution patterns.
Prerequisites
- A payment provider account with test keys (Stripe test mode is assumed)
- Python 3.12+ and a Postgres database for the key store
- An agent framework that lets you intercept tool calls (any tool-calling loop with middleware or wrapper hooks)
- Comfort with state machines and transactional outbox concepts
Why Agent Payment Loops Multiply Requests
Three retry sources, all normal operation:
- LLM non-determinism. A tool result the model dislikes, a schema violation, or a re-plan produces a second call. Nothing in the loop guarantees the model reuses the same request metadata — unless you force it to.
- Tool and network timeouts. The charge succeeded server-side; the response was lost. The client sees a timeout. The work already happened.
- Orchestration replay. Crash recovery and durable executors re-run the last step. A step that is not idempotent becomes a duplicate side effect.
A payment tool without key discipline converts each of these into a double charge. Worse, the failure is silent: both calls return 200 with different charge IDs.
Anatomy of an Idempotency Key: Stripe’s Semantics
Stripe’s idempotent requests documentation is the reference implementation. The semantics worth internalizing:
| Property | Behavior |
|---|---|
| Header | Idempotency-Key on the request |
| Key generation | Client-generated; V4 UUIDs or random strings with sufficient entropy suggested |
| Length | Up to 255 characters |
| Content | Avoid sensitive data (emails, personal identifiers) in the key |
| Replay window | Keys may be pruned after at least 24 hours; reuse after pruning creates a new request |
| Result storage | API v1 saves the status code and body of the first request once endpoint execution begins; retries replay it, including saved 500 errors |
| Fingerprinting | Reusing a key with different parameters errors instead of returning a mismatched response |
| Non-executed failures | Validation failures and concurrent conflicts store nothing — those requests remain safely retryable |
| Methods | All POST requests accept keys; GET and DELETE are already idempotent and should not carry them |
Two consequences dominate agent design:
- The 24-hour window is your retry horizon. Any loop that may retry beyond a day needs an application-level key store, not just the provider’s.
- The key must be stable across attempts. If the agent mints a new UUID per attempt, the provider sees N distinct requests. The key is the only link between retries of one logical intent.
Where the Key Must Originate
The rule: the key is minted once per logical payment intent and reused across every attempt. Not per tool call, not per model turn, not per HTTP request.
import uuid
from dataclasses import dataclass
@dataclass
class PaymentIntent:
intent_id: str # your domain ID: invoice, order, cart
account_id: str # scoping: whose money
amount: int
currency: str
idempotency_key: str # minted ONCE, stored, reused forever
def new_intent(intent_id: str, account_id: str, amount: int, currency: str) -> PaymentIntent:
return PaymentIntent(
intent_id=intent_id,
account_id=account_id,
amount=amount,
currency=currency,
idempotency_key=f"pay_{intent_id}_{uuid.uuid4()}", # ≤255 chars, no PII
)
Minting is cheap; the discipline is persisting the intent row — including the key — in the same transaction that records the business event, before any network call. If you mint in memory and crash before the call, the retry mints a different key and you are exposed. If you persist first, every retry path (model loop, orchestrator replay, manual re-run) reads the same key from durable state.
CREATE TABLE payment_intents (
intent_id TEXT PRIMARY KEY,
account_id TEXT NOT NULL,
amount INTEGER NOT NULL,
currency CHAR(3) NOT NULL,
idempotency_key TEXT NOT NULL UNIQUE,
state TEXT NOT NULL DEFAULT 'key_minted',
charge_id TEXT,
response_json JSONB,
created_at TIMESTAMPTZ NOT NULL DEFAULT now(),
updated_at TIMESTAMPTZ NOT NULL DEFAULT now()
);
The Payment Tool State Machine
Wrap the provider call in an explicit state machine. The states exist so that every retry hits the correct branch without consulting the model’s memory:
| State | Meaning | On retry |
|---|---|---|
key_minted | Intent persisted; no call attempted | Call the provider with the stored key |
in_flight | Call started; result not recorded | Call again with the same key (safe — provider replays) |
recorded | Provider response stored | Replay response_json; never re-charge |
failed_retryable | Network error before execution began | Re-call with the same key |
failed_terminal | Declined / validation error | Do not auto-retry; surface to the planner or a human |
reconciling | Ambiguous outcome (timeout after possible execution) | Query by key/charge, do not blind-retry |
def execute_payment(intent_id: str, stripe_client) -> dict:
intent = db.get_payment_intent(intent_id) # 1. load durable state
if intent.state == "recorded": # 2. replay branch
return intent.response_json
if intent.state == "failed_terminal":
raise TerminalPaymentError(intent.response_json)
db.set_state(intent_id, "in_flight") # 3. mark before network I/O
try:
charge = stripe_client.PaymentIntent.create(
amount=intent.amount,
currency=intent.currency,
metadata={"intent_id": intent.intent_id},
idempotency_key=intent.idempotency_key, # 4. the stable key
)
except StripeError as exc:
if is_retryable(exc):
db.set_state(intent_id, "failed_retryable")
raise RetryableToolError(str(exc)) from exc
db.record_failure(intent_id, str(exc)) # declined → terminal
raise
db.record_success(intent_id, charge) # 5. atomic with state
return charge_to_tool_result(charge)
The agent-facing tool wrapper is then trivially safe to loop:
@mcp_tool()
def charge_customer(intent_id: str) -> dict:
"""Charge the stored payment intent. Safe to call repeatedly:
the idempotency key is loaded from durable state, not generated here."""
return execute_payment(intent_id, stripe_client)
Note what is not in the tool signature: an amount, a card token, or a freshly generated UUID. Arguments that vary per attempt are exactly how fingerprint mismatches and duplicate charges start. The model supplies intent_id; everything else comes from the persisted intent.
Scoping and Fingerprinting Keys
Keys need a scope and a fingerprint, or you trade double charges for wrong-response bugs:
- Scope: account × endpoint. Never reuse a key across accounts or across operations.
pay_inv-1041_...(create charge) andrefund_inv-1041_...(refund) must be different keys even for the same invoice. Include the operation in the key or, better, in a dedicated column used for uniqueness:UNIQUE (account_id, operation, idempotency_key). - Fingerprint the payload. Hash the canonical request parameters (
amount,currency,intent_id, provider-specific fields) alongside the key. On replay, compare fingerprints; a mismatch is a bug in your loop and must raise loudly rather than return the old response.
import hashlib, json
def fingerprint(amount: int, currency: str, intent_id: str) -> str:
payload = json.dumps(
{"amount": amount, "currency": currency, "intent_id": intent_id},
sort_keys=True,
)
return hashlib.sha256(payload.encode()).hexdigest()
def record_success(intent_id: str, charge, fp: str):
with db.transaction():
intent = db.get_payment_intent(intent_id, for_update=True)
if intent.fingerprint and intent.fingerprint != fp:
raise IdempotencyConflict("key reused with different parameters")
db.update(
intent_id,
state="recorded",
charge_id=charge.id,
response_json=serialize(charge),
fingerprint=fp,
)
This mirrors Stripe’s own guard: the idempotency layer compares incoming parameters to the original request and errors if they differ.
Common Pitfalls
- Regenerating the key on retry. The single most common agent bug. Any
uuid4()inside the retry path defeats the entire mechanism. Mint only in the intent-creation path. - Sharing keys across operations. Reusing a create-charge key for a refund or a capture produces wrong replays or provider errors. One logical operation, one key.
- Key leaks. Keys are not secrets, but they are capability-adjacent: log them as identifiers, never place PII inside them (Stripe explicitly warns against sensitive data in keys), and keep them out of model-visible tool arguments when possible — reference an
intent_idinstead. - Retrying past the 24-hour window. After pruning, the provider treats the key as new. For long-horizon agents, keep your own
payment_intentstable authoritative and reconcile rather than replaying. - Retrying validation failures. Stripe stores nothing when validation fails or a concurrent conflict occurs — those requests are safe to re-run. Declines are terminal: auto-retrying a declined card in a loop is a fraud signal and a UX disaster.
- Ignoring in-flight ambiguity. A timeout after
in_flightmay mean the charge exists. Query the provider byintent_idmetadata (or list charges with the same key semantics) before re-calling; the same key makes the re-call safe, but reconciliation is faster and cheaper.
Reconciliation and Double-Charge Detection
Idempotency prevents duplicates; reconciliation proves it. Run a periodic job that compares your intents to provider truth:
-- Candidate double charges: same intent recorded with two provider IDs
-- (should be impossible if record_success is atomic; this is your tripwire)
SELECT intent_id, COUNT(DISTINCT charge_id) AS charge_count
FROM payment_intents
WHERE charge_id IS NOT NULL
GROUP BY intent_id
HAVING COUNT(DISTINCT charge_id) > 1;
Plus a webhook consumer: on payment_intent.succeeded, look up the local intent by metadata intent_id. If none exists, you have an orphan charge — an execution that ran without a durable record — and that is an incident, not noise. Alert on three signals: orphan charges, intents stuck in in_flight past a timeout threshold, and fingerprint conflicts.
Testing Idempotency
Idempotency claims are worthless untested. Four tests catch nearly every regression; run them against a real database and a provider’s test mode, not mocks:
def test_same_key_same_intent_single_charge(stripe_mock):
intent = new_intent("inv-1", "acct-9", 500, "usd")
db.save(intent)
first = execute_payment(intent.intent_id, stripe_mock)
second = execute_payment(intent.intent_id, stripe_mock) # replay
assert first.id == second.id
assert stripe_mock.calls == 1
def test_retry_after_timeout_single_charge(stripe_mock):
intent = new_intent("inv-2", "acct-9", 500, "usd")
db.save(intent)
stripe_mock.fail_next_with_timeout()
with pytest.raises(RetryableToolError):
execute_payment(intent.intent_id, stripe_mock)
result = execute_payment(intent.intent_id, stripe_mock) # recovery
assert stripe_mock.calls == 2 # same key both times
assert result["state"] == "recorded"
def test_fingerprint_mismatch_raises(stripe_mock):
intent = new_intent("inv-3", "acct-9", 500, "usd")
db.save(intent)
execute_payment(intent.intent_id, stripe_mock)
db.update(intent.intent_id, amount=999) # simulate tampering
with pytest.raises(IdempotencyConflict):
execute_payment(intent.intent_id, stripe_mock)
def test_new_intent_gets_new_key():
a, b = new_intent("inv-4", "acct-9", 500, "usd"), new_intent("inv-5", "acct-9", 500, "usd")
assert a.idempotency_key != b.idempotency_key
Wire these into CI. The bug always lives in the state machine and the key store, not in the model.
Hardening Checklist
- Key minted once per logical intent, never inside a retry path
- Intent row (with key) committed before the first provider call
- All retries read the key from durable state, not from model context
- Tool signature exposes only an
intent_id— no per-call amounts or fresh UUIDs - Keys scoped per account and operation; payload fingerprinted and checked
- Declines are terminal; only network-level failures auto-retry
- In-flight timeouts trigger provider-side lookup, not blind new charges
- Reconciliation job detects orphans, stuck intents, and duplicate charge IDs
- Retry horizon respects the provider’s 24-hour replay window
- Idempotency tests run in CI against a real database