Skip to content
Blog

Metrics for Agent Health: Error Rates, Tokens, and Turn Length

A concrete OpenTelemetry metric set for LLM agents — golden signals, instrument choice, cardinality discipline, exemplars, and a dashboard layout you can ship.

Published on • October 4, 2026

AI Assistant

Metrics for Agent Health: Error Rates, Tokens, and Turn Length

An agent failure rarely looks like a failure. The API returns 200, the model answers, and the run quietly burns four extra turns retrying a tool call that was never wired up correctly. Logs do not catch this: there is nothing to log, because the system did exactly what it was told. The only thing that notices is a graph — error rate creeping up, time to first token doubling, token spend climbing while successful runs stay flat.

Metrics are the right signal for that job. OpenTelemetry’s metrics concept draws the line: traces capture request lifecycles, metrics are “intended to provide statistical information in aggregate.” Only metrics answer “is the agent getting worse?” at a glance.

Key technologies: OpenTelemetry Metrics API and SDK for Python (opentelemetry-api, opentelemetry-sdk), OTLP export, GenAI semantic conventions (gen_ai.* instruments), Counter / Histogram / UpDownCounter / Gauge, Views, exemplars.

The Four Golden Signals, Applied to Agents

The four golden signals — latency, traffic, errors, saturation — were written for request/response services. Translated to an agent they land like this:

Golden signalWhat it means for an agent
LatencyEnd-to-end run duration, plus time to first token (TTFT) and time per output chunk
TrafficRuns started per second, inference calls per run, tool calls per run
ErrorsFailed runs, provider errors, tool-call failures, structured-output validation failures
SaturationQueue depth waiting for execution, concurrency against provider rate limits, context growth per turn

Two classic lenses organize these. RED (Rate, Errors, Duration) is the user-facing view: how often we are called, how often it breaks, how long it takes. USE (Utilization, Saturation, Errors) is the resource view: where work is piling up. A healthy dashboard has a RED row for the person answering the page and a USE row for the person diagnosing it.

The Concrete Metric Set

Everything below maps to a real instrument. Names in the first column are from the GenAI semantic conventions, which moved out of the main semconv repo into their own project; that YAML marks all of them stability: development, so treat them as the direction of travel rather than a frozen contract.

MetricInstrumentUnitWhy
Run request rateCounter{run}Numerator for the RED rate
Run error rateCounter{error}Same counter filtered by outcome, never a separate gauge
gen_ai.client.operation.durationHistogramsProvider-facing call latency; percentiles come from here
gen_ai.client.operation.time_to_first_chunkHistogramsTTFT for streaming calls — should not be reported for non-streaming
gen_ai.client.operation.time_per_output_chunkHistogramsSustained generation speed after the first chunk
gen_ai.execute_tool.durationHistogramsOne bucket per tool execution; tool-call failure rate is this metric split by outcome
gen_ai.invoke_agent.inference_callsHistogram{inference_call}Turn count per agent invocation — the “turn length” distribution
gen_ai.invoke_agent.tool_callsHistogram{tool_call}Tool calls per invocation, including the ones that failed
gen_ai.client.inference.usage.input_tokensCounter{token}Total prompt tokens; carries a required gen_ai.token.modality attribute
gen_ai.client.inference.usage.output_tokensCounter{token}Total completion tokens, including reasoning tokens
gen_ai.client.inference.operation.output_tokensHistogram{token}Per-operation output tokens — the distribution behind cost surprises
Queue depthUpDownCounter{run}Rises and falls with pending work; the classic UpDownCounter example
Cost per runHistogramUSDComputed in your code from token counts and a price table

Note the shape of the list: rates are counters, distributions are histograms, current values are gauges, and anything that goes up and down is an UpDownCounter. The docs describe a Counter as “an odometer on a car — it only ever goes up” and a Histogram as “a client-side aggregation of values.” Getting this wrong is the common bug: an in-process “error rate” gauge loses samples between scrapes and cannot be aggregated across instances.

Cost is deliberately derived: the GenAI conventions count tokens, they do not price them. Multiply your token counters by a per-model price table and record the result per run, and cost percentiles land next to latency percentiles.

Instrumenting a Run in Python

from opentelemetry import metrics
from opentelemetry.exporter.otlp.proto.grpc.metric_exporter import OTLPMetricExporter
from opentelemetry.sdk.metrics import MeterProvider
from opentelemetry.sdk.metrics.export import PeriodicExportingMetricReader
from opentelemetry.sdk.resources import Resource

resource = Resource.create({
    "service.name": "support-agent",
    "deployment.environment": "production",
})
reader = PeriodicExportingMetricReader(
    OTLPMetricExporter(endpoint="http://otel-collector:4317"),
    export_interval_millis=15_000,
)
metrics.set_meter_provider(
    MeterProvider(resource=resource, metric_readers=[reader])
)

meter = metrics.get_meter("redlinesoft.agents", version="1.4.0")

runs = meter.create_counter("agent.runs", unit="{run}", description="Agent runs")
run_duration = meter.create_histogram("gen_ai.client.operation.duration", unit="s")
ttfc = meter.create_histogram(
    "gen_ai.client.operation.time_to_first_chunk", unit="s"
)
tool_duration = meter.create_histogram("gen_ai.execute_tool.duration", unit="s")
turns = meter.create_histogram(
    "gen_ai.invoke_agent.inference_calls", unit="{inference_call}"
)
in_tokens = meter.create_counter(
    "gen_ai.client.inference.usage.input_tokens", unit="{token}"
)
out_tokens = meter.create_counter(
    "gen_ai.client.inference.usage.output_tokens", unit="{token}"
)
queue_depth = meter.create_up_down_counter("agent.queue.depth", unit="{run}")
cost = meter.create_histogram("agent.run.cost", unit="USD")

Then record at exactly three boundaries: run start/end, model call, and tool call.

def finish_run(agent: str, model: str, outcome: str, *, seconds: float,
               first_chunk: float, inference_calls: int,
               input_tokens: int, output_tokens: int, usd: float) -> None:
    attrs = {
        "gen_ai.agent.name": agent,
        "gen_ai.response.model": model,
        "agent.outcome": outcome,          # ok | provider_error | tool_error
    }
    runs.add(1, attrs)
    run_duration.record(seconds, attrs)
    ttfc.record(first_chunk, {k: attrs[k] for k in
                              ("gen_ai.agent.name", "gen_ai.response.model")})
    turns.record(inference_calls, {"gen_ai.agent.name": agent})
    in_tokens.add(input_tokens, attrs)
    out_tokens.add(output_tokens, attrs)
    cost.record(usd, attrs)

Keep agent.outcome an enumerated string, never an exception message: error rate becomes a readable by outcome sum over a handful of values, and it still reads well a year later.

Cardinality: Labeling Without Exploding

The metrics SDK keeps a separate aggregation state for every unique attribute combination, so cardinality — not request volume — drives memory cost. It enforces a cardinality limit of 2000 unique combinations per metric stream by default. When the limit is hit, measurements are not dropped: they fold into a single overflow data point tagged otel.metric.overflow=true. Totals stay correct, memory stays bounded, and one query detects overflow everywhere.

The catch is that overflow replaces the entire attribute set. A counter recording {model, agent.outcome} that overflows loses agent.outcome too — and an error-rate alert built on agent.outcome silently stops firing while the metric’s total keeps looking healthy. That is the failure mode to design against.

Rules that hold up in production:

  • Low cardinality on the metric, high cardinality on the trace. Model, agent name, tool name, outcome enum, region — fine. Session ID, user ID, run ID, raw URL path, prompt text — never; the docs call out user IDs and raw URL paths as attributes that “can cause unbounded memory growth.”
  • Process-lifetime context belongs on the Resource. service.name, deployment.environment and service.instance.id are Resource attributes, exempt from the cardinality limit and queryable even during overflow.
  • Temporality matters. Delta temporality resets state each cycle, so the limit bounds only what is active that cycle; cumulative holds state until process restart, so once you hit the ceiling you keep overflowing.
  • Use Views to trim before the limit. A View’s attribute_keys allow-list drops unwanted keys, and enforcement runs after filtering — filtering is how you stay under the ceiling.

Exemplars: Metrics That Point at Traces

Exemplars are the bridge. Per the metrics SDK spec they “allow correlation between aggregated metric data and the original API calls where measurements are recorded,” and for synchronous instruments they carry the trace ID and span ID of the active span at recording time. You alert on the histogram, click the spike, land on the exact run.

Two things worth knowing:

  • Exemplar sampling SHOULD be turned on by default. The SDK exposes an ExemplarFilter (AlwaysOn, AlwaysOff, TraceBased — eligibility) plus an ExemplarReservoir that decides storage: histograms default to AlignedHistogramBucketExemplarReservoir, everything else to SimpleFixedSizeExemplarReservoir.
  • Attributes removed from a stream by a View can still show up on exemplars as filtered attributes. If you strip a sensitive key from a metric stream, decide deliberately whether you also want it reachable through exemplars.

Dashboard and Alert Layout

One dashboard, four rows, in this order:

  1. RED row. Runs/sec, error rate by agent.outcome, p50/p95/p99 of gen_ai.client.operation.duration, p95 TTFT.
  2. Turn and tool row. Turn count distribution, tool calls per run, tool failure rate, gen_ai.execute_tool.duration p95.
  3. Tokens and cost row. Input/output token counters, tokens per operation histogram, cost per run p95.
  4. USE row. Queue depth, in-flight runs, per-model token burn as a saturation signal.

Alerting that stays quiet:

  • Page on error rate and tail latency, never averages: error ratio over 5 minutes above threshold, p99 over SLO for two windows.
  • Warn on turn-count creep. A rising p95 of gen_ai.invoke_agent.inference_calls with flat traffic usually means a tool started failing and the model is retrying — visible before users notice.
  • Warn on cost per run. Same shape: the run still succeeds, it just costs 40% more.
  • Page on saturation, and treat sustained queue depth with flat throughput as the actionable combination.
  • Alert on otel.metric.overflow=true. It is one bit, and it means at least one dashboard above is lying to you.

Common Pitfalls

  1. Gauging a rate in-process. Anything that is a count over time is a Counter. Let the backend compute the rate.
  2. Labeling with session, user, or run IDs. This is the fastest route to the 2000-combination limit and a permanently undercounted error rate.
  3. Averaging latency. p95 and p99 come from histograms; an average hides exactly the runs your users complain about.
  4. Recording TTFT for non-streaming calls. The convention says time-to-first-chunk belongs to streaming calls only — inventing values breaks cross-service comparison.
  5. Forgetting failed work. Count inference and tool calls “including failed ones,” and record error outcomes on the same instruments. A metric that only sees success is a dashboard that only sees good days.
  6. Environment on a metric attribute, or no exemplars. Environment belongs on the Resource, where it survives overflow; without exemplars you get a red graph with no run to open.

Wrapping Up

Agent health is not one number; it is a small, stable set of instruments — counts as counters, latency, TTFT, tool duration, turn count and token distributions as histograms, queue depth as an UpDownCounter — labeled with a handful of low-cardinality attributes and carrying exemplars back to traces. Get the instrument kind right, keep the label set small, put process context on the Resource, and the four golden signals fall out of data you already emit. The dashboard is the easy part; the discipline is refusing the one attribute that makes it interesting.

Sources