Trace Sampling in High-Volume Agent Systems
Head-based, tail-based, and parent-based sampling with OpenTelemetry — how to keep 1% of traces that still tell the whole story.
Published on • October 7, 2026
AI Assistant

An agent system that chains model calls, tool invocations, vector searches, and HTTP fan-outs can easily generate thousands of traces per second. At that volume, “store everything” isn’t a strategy — it’s a bill. The OpenTelemetry docs frame the decision crisply: sampling is worth it when you generate 1000+ traces/second, most of your traffic is healthy and low-variation, and you have error, latency, or domain criteria that matter more than completeness.
Their rule of thumb: “For high-volume systems, it is quite common for a sampling rate of 1% or lower to very accurately represent the other 99%.”
Here’s how to choose and configure it.
Head-Based vs. Tail-Based
Head sampling decides as early as possible — at the root span — without inspecting the rest of the trace. The dominant form is consistent probability sampling: the decision derives from trace ID + desired percentage, so every service makes the same call for the same trace. If you want 5%, you get whole 5% traces, consistently, wherever they run.
- Upsides: simple to understand and configure, cheap, works anywhere in the pipeline.
- Downside: it can’t see the future. A trace that starts fine and then hits an error was already rejected at the head. Head sampling alone cannot guarantee “keep all errors.”
Tail sampling decides after seeing most or all spans. Use it to:
- Always sample traces containing an error
- Sample based on overall latency
- Sample by span attributes (e.g., traffic to a newly deployed service)
- Apply different rates to high- vs. low-volume services
The costs are real: rules are hard to write and keep current; the processor must be stateful, sometimes across “dozens or even hundreds of compute nodes”; and implementations are often vendor-specific.
The practical answer is usually both: head-sample a small percentage to protect the pipeline, then tail-sample downstream to catch the traces that matter.
Collector Processors
Two OpenTelemetry Collector processors handle this:
- Probabilistic Sampling Processor —
open-telemetry/opentelemetry-collector-contrib/tree/main/processor/probabilisticsamplerprocessor - Tail Sampling Processor —
open-telemetry/opentelemetry-collector-contrib/tree/main/processor/tailsamplingprocessor
Head-based probability sampling in the collector; tail-based rules later in the pipeline, where state and compute can be provisioned deliberately.
SDK Configuration
The fastest lever is two environment variables:
export OTEL_TRACES_SAMPLER="traceidratio"
export OTEL_TRACES_SAMPLER_ARG="0.5" # probability in [0..1], default 1.0
Accepted values for OTEL_TRACES_SAMPLER:
| Value | Sampler |
|---|---|
always_on | AlwaysOnSampler |
always_off | AlwaysOffSampler |
traceidratio | TraceIdRatioBased |
parentbased_always_on | ParentBased(root=AlwaysOnSampler) — the default |
parentbased_always_off | ParentBased(root=AlwaysOffSampler) |
parentbased_traceidratio | ParentBased(root=TraceIdRatioBased) |
jaeger_remote / parentbased_jaeger_remote | Jaeger remote sampling |
xray | AWS X-Ray |
The default — parentbased_always_on — exists for a reason: sampling decisions must propagate. Once a trace is sampled or dropped at the head, every downstream service should honor that decision rather than re-rolling the dice. That’s what the parentbased_* wrappers do: respect the parent’s decision, and only fall back to the root sampler for new traces.
For jaeger_remote, the arg takes a richer form:
export OTEL_TRACES_SAMPLER_ARG="endpoint=...,pollingIntervalMs=5000,initialSamplingRate=0.25"
Context Propagation: How the Decision Travels
Sampling only works if the decision crosses process boundaries. OpenTelemetry’s default propagator is W3C TraceContext, which encodes everything in one header:
traceparent: 00-a0892f3577b34da6a3ce929d0e0e4736-f03067aa0ba902b7-01
│ │ │ │
│ trace-id parent-id trace-flags
version
That last field — trace-flags — carries the sampled bit. 01 means sampled; 00 means not. A service receiving traceparent with flags 00 knows not to emit spans for that trace, no matter what its local sampler is configured to do.
The mechanics:
- Sender injects context into a carrier (HTTP headers) via the Propagators API.
- Receiver extracts it before processing.
OTEL_PROPAGATORSdefaults totracecontext,baggageand also acceptsb3,b3multi,jaeger,xray,ottrace,nonefor interoperating with other ecosystems.
Security notes from the docs worth repeating:
- Sanitize incoming context from untrusted sources.
- Don’t propagate internal trace IDs or baggage to external services.
- Never put credentials or PII in baggage — it travels in headers.
When NOT to Sample
The docs are direct: skip sampling if you’re at tens of traces per second or lower, if you only need aggregate metrics, or if regulation prohibits dropping data. Also weigh the three costs of sampling they name: compute (the tail-sampler proxy), engineering maintenance, and the opportunity cost of missing the one critical trace you’d have wanted.
A Practical Recipe for Agent Systems
Agent workloads have a useful property: most traces are boring (a happy-path RAG turn), and the interesting ones are identifiable by cheap signals available at the head.
- Head-sample at 1–5% with
parentbased_traceidratioacross all services. This bounds volume immediately and keeps whole traces consistent. - Always keep error and slow traces. In the collector, run the Tail Sampling Processor with rules like:
status == ERROR→ always sample; duration > p99 → always sample; agent tool name in [refund, delete] → always sample. - Tag what you might want later. Attach span attributes for model, tool, token counts, and session id — tail rules and later queries key off them.
- Propagate deliberately. Keep
tracecontextenabled everywhere; verify the sampled flag survives your message queues (Kafka headers, gRPC metadata) — queue bridges are the classic place the decision silently resets. - Revisit the ratio quarterly. As volume grows, 1% may hold; as you add services, propagation gaps may show up as broken traces.
The Metrics Side
Sampling drops traces, not metrics. Keep counters and histograms at full resolution — request counts, error rates, latency distributions — so aggregate dashboards stay accurate while the trace store stays small. The correlation payoff: metrics tell you that checkout got slow; the 1% of traces you kept tell you why.