Skip to content
Blog

Correlation IDs Across Agent Handoffs

Keep a single identifier alive as control moves between agents: why log IDs break at handoffs, how OpenTelemetry context propagation and baggage carry correlation across boundaries, the security rules for propagating IDs, and a concrete schema for multi-agent trace correlation.

Published on • October 10, 2026

AI Assistant

A user reports that their refund request failed. The triage agent handled it, handed off to the billing agent, which called a tool, which errored. Three services produced logs. Somewhere in there is the answer, and you cannot find it — because the triage agent’s logs have one request ID, the billing agent’s have another, and the tool’s have a third. Nothing joins them.

Handoffs are where correlation usually breaks. Each hop is a fresh execution context, often a fresh process, sometimes a different vendor’s framework — and each one mints a new identifier because nobody told it not to. The fix is not a cleverer ID; it is making the existing one part of the execution context that propagates automatically.

This post is about keeping one correlation ID alive from the moment a user’s request enters your system until the last agent stops talking about it — using OpenTelemetry’s context propagation and baggage as the transport, and a disciplined schema as the contract.

In this tutorial, you will learn how to:

  • Understand why per-service request IDs fail at agent handoffs
  • Distinguish trace IDs, span IDs, request IDs, and business correlation IDs
  • Carry a correlation ID across handoffs with OpenTelemetry baggage
  • Inject and extract context across HTTP, queues, and in-process calls
  • Follow the security rules for propagating identifiers across trust boundaries
  • Define a correlation schema that lets you jump between logs, traces, and metrics
  • Handle the cases where propagation genuinely cannot happen

Key technologies: OpenTelemetry context propagation, baggage, W3C TraceContext, traceparent, structured logging, correlation IDs.

Prerequisites

  • An agent system with at least one handoff boundary (service, framework, or vendor)
  • An OpenTelemetry SDK with a backend that indexes on trace ID
  • A structured logging setup that emits fields, not interpolated strings

Why the obvious approach breaks

Most systems start with a request ID: something the gateway mints, puts in a header, and every service logs. It works right up until the handoff.

At a handoff, one of several things happens:

  • The next agent is a separate service with its own middleware, which mints a new request ID because the incoming one is not in the header set it knows about.
  • The next agent is a different framework that does not read your custom header at all.
  • The next agent is a hosted or third-party agent that accepts a metadata field you did not pass, so it generates its own.
  • The handoff happens in-process, and the correlation lives in a request-scoped object that goes out of scope when the calling agent’s stack unwinds.

In every case the ID chain is severed, and the logs on either side of the seam cannot be joined by anything except timestamps — which, in a system running concurrent requests, join the wrong things.

The deeper problem is architectural: a request ID is infrastructure metadata owned by the transport layer, but an agent handoff is application control flow. Transport-layer metadata does not survive application-layer control flow changes unless something explicitly carries it.

Four identifiers, four jobs

Before fixing anything, separate the things being conflated:

IdentifierScopeCarries across handoffs?Used for
Trace IDOne distributed traceYes, by designJoining spans across services
Span IDOne operationNo — new span per operationParent/child relationships inside a trace
Request IDOne inbound HTTP requestUsually noCorrelating logs within one service
Business correlation IDOne user intentYes, if you propagate itJoining across retries, replans, and sessions

The trace ID is already the thing you want — OpenTelemetry propagates it across service boundaries as a matter of course. The mistake is treating your custom request ID as the primary correlation mechanism when a trace ID exists and is standardized.

The business correlation ID is the one that does not exist by default and that you usually need. Trace IDs identify one execution attempt. If the user retries, or the agent replans and restarts, you get a new trace — but it is the same user intent. The business ID (“this is refund request 4471”) is what joins attempts.

Use both: the trace ID to navigate one execution, the business ID to aggregate across attempts.

Carrying context with OpenTelemetry propagation

Context propagation has two halves:

Context is an object holding the identifiers for the current execution unit — trace ID, span ID, and any baggage. When service A calls service B, A includes a trace ID and a span ID; B uses them to create a span that belongs to the same trace, with A’s span as its parent.

Propagation is the mechanism that serializes that context across process and network boundaries.

The default propagator uses the W3C TraceContext traceparent header:

<version>-<trace-id>-<parent-id>-<trace-flags>
00-a0892f3577b34da6a3ce929d0e0e4736-f03067aa0ba902b7-01

For HTTP this is nearly free: instrumented clients inject it, instrumented servers extract it. Your job is to make sure the handoff path is an instrumented one.

On the sender side you inject context into a carrier — HTTP headers in the common case, but any metadata store you choose. On the receiving side you extract it from that carrier.

from opentelemetry import propagate, trace

# Sender: attach current context to outbound headers
headers = {}
propagate.inject(carrier=headers)
await http_post(next_agent_url, headers=headers, json=payload)

# Receiver: restore context from inbound headers
carrier = dict(request.headers)
propagate.extract(carrier)

Once extracted, the next agent’s spans automatically parent under the previous agent’s span. The handoff becomes a link in one trace instead of the start of a new one.

Baggage: the part that carries your own ID

Trace context carries the trace and span IDs. It does not carry your correlation ID — unless you put it in baggage.

Baggage propagates arbitrary key-value pairs alongside the trace context, across every service boundary in the call path. That makes it exactly the right transport for a business correlation ID:

from opentelemetry import baggage

# Set once, at the entry point where the user intent is known.
ctx = baggage.set_baggage("correlation_id", refund_request_id)
ctx = baggage.set_baggage("tenant_id", tenant)
attach(ctx)

# Read it anywhere downstream — the next agent, the tool, the DB layer.
correlation_id = baggage.get_baggage("correlation_id")

Baggage is attached to the current context, so it flows through propagation automatically alongside traceparent. Every downstream service sees correlation_id without you adding a second header.

Combine that with structured logging and the join becomes trivial:

logger.info(
    "refund tool executed",
    extra={
        "correlation_id": baggage.get_baggage("correlation_id"),
        "trace_id": format(trace.get_current_span().get_span_context().trace_id, "032x"),
        "agent": "billing_agent",
        "tool": "issue_refund",
    },
)

Now every log line in every agent carries both the trace ID and the business ID. Searching either one gets you the full story.

The security rules for propagation

Propagation sends your identifiers across service boundaries, and that has consequences. OpenTelemetry’s guidance is explicit on three points:

Incoming context can be forged. A malicious actor can send a fabricated traceparent to manipulate your tracing data or exploit a parser. Consider ignoring or sanitizing incoming context from untrusted sources — for example, regenerating the trace ID at your edge for external traffic rather than adopting it.

Outgoing context can leak. Internal trace IDs, span IDs, and baggage items may reveal architecture or business logic. Configure your propagators to withhold context from external or public-facing endpoints. In practice: propagate freely inside your trust boundary, strip at the edge.

Baggage must never hold sensitive data. It is propagated across every service boundary and may be logged anywhere along the path. No credentials, no API keys, no PII. A correlation_id is safe because it is an opaque handle; the thing it points to stays behind in a store you control.

One more rule specific to non-HTTP carriers: if you propagate context over a protocol with no dedicated metadata field — a message queue, a WebSocket frame, a file — the receiving side must extract and remove the context before processing the payload. Leaving it in produces undefined behaviour downstream.

A concrete schema

Define this once and treat it as a contract across your agent system:

correlation_id   opaque, globally unique, assigned at intent intake
trace_id         OpenTelemetry, 128-bit, one per execution attempt
span_id          OpenTelemetry, 64-bit, per operation
agent.name       the agent currently holding control
handoff.from     previous agent, on the first span of the new agent
handoff.to       target agent, on the last span of the outgoing agent
handoff.reason   why control transferred
session.id       the conversation/session the intent belongs to
tenant_id        multi-tenant isolation

Two entries do real work:

handoff.reason — without it you can see that control moved but not why, and “why did the triage agent escalate?” is the first question in any multi-agent incident review.

agent.name on every span — in a handoff architecture the active agent changes mid-trace. A tool span with no agent name is ambiguous the moment you have more than one agent with a search tool.

Emit all of it on every span and in every log line. Duplicated deliberately: traces and logs are different stores, and the whole point is being able to start from either.

Instrumenting the handoff itself

The handoff is an operation. Give it a span.

def handoff(target_agent, reason: str, payload):
    with tracer.start_as_current_span(f"handoff.{target_agent.name}") as span:
        span.set_attribute("handoff.from", current_agent.name)
        span.set_attribute("handoff.to", target_agent.name)
        span.set_attribute("handoff.reason", reason)
        span.set_attribute("session.id", session_id)

        headers = {}
        propagate.inject(carrier=headers)
        return target_agent.run(payload, headers=headers)

Because the span is created before propagation injects, the outgoing context parents the next agent’s work under this handoff span. Your waterfall becomes:

agent.run
├── llm.call (triage)
├── handoff.billing_agent          12ms
│   ├── llm.call (billing)         [billing_agent]
│   └── tool.issue_refund          [billing_agent]
└── agent.run (billing)            ← linked, same trace ID

The handoff span is also where you measure transfer cost. A handoff that takes 400 ms because the next agent re-reads its full context is a real performance problem, and it is invisible unless the transition is instrumented.

When propagation cannot happen

Three situations require something other than automatic propagation:

The next agent is a third-party service that does not accept your context. Pass the correlation ID as an explicit field in the payload, and have that agent attach it to its own spans as an attribute. You lose automatic trace parenting — compensate by recording the outbound trace ID in the payload and the inbound one on the receiving side, then link the traces.

The handoff crosses a queue rather than a direct call. Message queues do not carry HTTP headers. Either use a messaging propagator, or serialize the context into the message envelope yourself and extract it in the consumer.

The handoff is a full page reload or a new user session. There is no live context to propagate. The business correlation ID is what survives here — persist it client-side or in your session store, and re-establish context from it on the next entry.

For all three, the fallback is the same: put the business correlation ID in the payload, propagate it as baggage where you can, and accept that some edges will need a manual join. Document those edges — an undocumented gap is indistinguishable from a bug.

Verifying it works

Do not assume propagation is happening. Test it:

  1. Trigger a handoff in a staging environment with a unique correlation ID — something like test-<uuid> you can search for unambiguously.
  2. Search your log store for that ID. You should find entries from every agent in the chain, including the tool’s own logs.
  3. Search your trace backend for the trace ID from the first span. Confirm the waterfall crosses agent boundaries with correct parenting.
  4. Check for orphan spans. Any span in the trace with no agent.name is a code path you forgot to instrument.
  5. Check the edge. Confirm external traffic does not inherit a caller-supplied trace ID.

Automate steps 2 and 3 as a smoke test. Correlation degrades silently — someone adds a new handoff path with a raw HTTP client that does not inject headers, and the break only surfaces during the incident you are trying to debug.

A checklist

  • Trace ID is the primary correlation mechanism, not a custom request ID
  • A business correlation ID is assigned at intent intake and propagated as baggage
  • Every handoff injects and extracts OpenTelemetry context
  • The handoff itself is a span with from, to, and reason
  • Every span and log line carries agent.name
  • Baggage contains no secrets or PII
  • Context is stripped or regenerated at the trust boundary
  • Non-HTTP carriers extract and remove context before processing
  • Cross-vendor edges pass the correlation ID explicitly in the payload
  • A smoke test proves correlation end to end after every new handoff path

Further reading