Skip to content
Blog

Trace Spans for Tool Calls: Instrumenting Function Execution

Instrument agent tool calls with OpenTelemetry spans: span structure, trace and span IDs, parent relationships, attributes versus events, status codes, wrapping sync and async tools, and the schema that makes tool spans queryable instead of just present.

Published on • October 10, 2026

AI Assistant

When an agent misbehaves, the questions you actually need answered are mundane: which tool ran, with what arguments, how long it took, what it returned, and which of the three tools in that turn was the slow one. Logs give you fragments with no ordering. Metrics give you aggregates with no specifics. Traces give you the shape of a single execution — and the unit of a trace is the span.

A span is a named, timed operation with a parent, a set of attributes, and a status. Wrap each tool call in one and you can answer every one of those questions by clicking a bar in a waterfall. Skip the attribute schema and you get spans that exist in your backend and tell you nothing.

This post is about the second half: what a tool span should contain to be useful.

In this tutorial, you will learn how to:

  • Model an agent turn as a trace and each tool call as a child span
  • Choose span names that group and filter well
  • Separate attributes (queryable) from events (a timeline within a span)
  • Set status codes correctly, including recording exceptions
  • Wrap synchronous, asynchronous, and streaming tools without changing their signatures
  • Handle retries, timeouts, and cancellations as their own spans
  • Define a tool-span attribute schema your team can build dashboards on

Key technologies: OpenTelemetry traces, spans, attributes, events, status codes, W3C TraceContext, opentelemetry-instrumentation conventions.

Prerequisites

  • An agent framework with a tool-calling loop
  • An OpenTelemetry SDK for your language and a backend to export to (Jaeger, Grafana Tempo, Honeycomb, Datadog, or a vendor of your choice)
  • Comfort with the idea of a parent/child relationship between operations

The model: one turn is one trace

OpenTelemetry’s trace model is a tree. For an agent, the natural mapping is:

trace: agent.run                              ← the whole turn
├── span: llm.call (planner)                  ← model inference
│   └── event: tokens: {"input": 812, "output": 64}
├── span: tool.search_docs                    ← tool execution
│   └── span: http GET /search?q=...          ← what the tool itself called
├── span: tool.search_docs
├── span: llm.call (responder)
└── span: tool.issue_refund
    └── span: http POST /refunds              ← nested external call

A trace represents one agent run end to end. Every span in the trace shares a single trace ID. A span represents one operation within it: a model call, a tool call, or a downstream HTTP request made by that tool.

Two identifiers tie it together:

  • Trace ID — 128 bits, identical for every span in the run. This is what you search on when someone says “show me run 4f2a…”.
  • Span ID — 64 bits, unique per span. The parent span ID field on a child points at the operation that caused it.

The span that starts the trace is the root span; it has no parent. Everything else hangs off it.

That parent pointer is the entire value proposition. Without it, you have a bag of durations. With it, you can ask “of the 4.2 seconds this run took, how much was model inference versus tool execution versus waiting on HTTP” — and get an answer computed from real timings rather than guessed from log timestamps.

Span naming: the decision you cannot easily undo

Span names end up as the default grouping key in most backends. Get this wrong and you will have 50,000 unique span names and no aggregation.

The rule: name by operation, put the instance in attributes.

# Bad: a new span name for every argument value
tracer.start_as_current_span(f"tool.search_docs?q={query}")

# Good: one span name, query is an attribute
tracer.start_as_current_span("tool.search_docs")

tool.search_docs will aggregate across every query your agent has ever run. The broken version creates a new series per query, which will also quietly explode your backend’s cardinality bill.

Recommended conventions:

Span nameMeaning
agent.runRoot span for one turn
llm.callA model inference request
tool.<name>Execution of a specific tool
handoff.<target>Control transferred to another agent

If your tools come from multiple namespaces (in-process functions, MCP servers, host-provided tools), encode the namespace: tool.mcp.jira.create_issue. Keep it a fixed shape — the moment you start embedding arguments, you lose aggregation.

Attributes versus events

This distinction is the difference between a span you can query and a span you can only look at.

Attributes are key-value pairs attached to the span. They are indexed and searchable: you can filter on tool.name = "issue_refund" and get every refund ever issued. They should be your primary vehicle.

Events are timestamped entries on the span’s own timeline. They are great for things that happen during the operation in sequence — a retry, a chunk of streamed output, a checkpoint — but they are not the right place for the operation’s defining characteristics.

span.set_attribute("tool.name", "search_docs")
span.set_attribute("tool.call_id", call_id)
span.set_attribute("gen_ai.request.model", "gpt-5.2")
span.set_attribute("http.response.status_code", 200)

span.add_event("retry", {"attempt": 2, "reason": "rate_limited"})
span.add_event("chunk", {"index": 14, "bytes": 512})

Ask yourself: would I want to filter or group by this? If yes, it is an attribute. If it is a moment within the span, it is an event.

A few hard rules on attributes:

  • Values must be primitives or arrays of primitives. A dict or an object gets dropped or stringified inconsistently.
  • Cardinality matters. user.id is fine if your backend handles it; a raw prompt with 4,000 tokens is not an attribute — truncate it, hash it, or attach it as an event.
  • Never put secrets in attributes. They are exported to your observability backend and retained. Redact API keys, tokens, passwords, and payment details at the instrumentation layer, not downstream.

Status codes and exceptions

A span has a status: UNSET, OK, or ERROR. The rules that trip people up:

  • An uncaught exception automatically sets ERROR and records the exception. If you catch the exception yourself, none of that happens — you have to do it explicitly.
  • UNSET is the default and is not the same as OK. Most spans will and should remain UNSET.
  • Set OK deliberately when the operation completed and you want to distinguish “succeeded” from “we never got around to marking it.”

For a tool that throws:

try:
    result = await tool.fn(**arguments)
except Exception as exc:
    span.set_status(Status(StatusCode.ERROR, str(exc)))
    span.record_exception(exc)
    raise
else:
    span.set_attribute("tool.result.size", len(result))

record_exception writes the exception type, message, and stack trace as an event on the span. The raise matters — re-throwing preserves the failure for your framework’s retry and error handling. Swallowing it to keep the trace green produces traces that lie.

Wrapping tools without changing their signatures

The cleanest approach is a decorator, because it works the same for sync and async tools and does not require your tool implementations to know tracing exists:

import functools
from opentelemetry import trace

tracer = trace.get_tracer("agent.tools")

def traced_tool(fn):
    @functools.wraps(fn)
    async def wrapper(*args, **kwargs):
        with tracer.start_as_current_span(f"tool.{fn.__name__}") as span:
            span.set_attribute("tool.name", fn.__name__)
            span.set_attribute("tool.call_id", kwargs.pop("_call_id", "unknown"))
            span.set_attribute("tool.argument_count", len(kwargs))
            # Redact, do not log raw arguments.
            span.set_attribute("tool.arguments", redact(kwargs))

            try:
                result = await fn(*args, **kwargs)
            except Exception as exc:
                span.set_status(Status(StatusCode.ERROR, str(exc)))
                span.record_exception(exc)
                raise

            span.set_attribute("tool.result.size", size_of(result))
            return result

    return wrapper

start_as_current_span does two things: it creates the span and makes it the active span for the duration of the block. Any child span created inside — an HTTP call the tool makes, a database query it runs — automatically becomes a child of this one. You get the nested structure for free as long as your HTTP and database clients are instrumented.

Three details worth flagging:

functools.wraps preserves the original name and docstring, so your tool registry and schema generation still see the real function.

Pass the call ID through, do not read it from a global. In concurrent runs, globals are wrong. Thread it through the call context or the tool’s execution context.

Pop _call_id before forwarding **kwargs — or better, accept it as an explicit parameter rather than smuggling it in the argument dict.

For synchronous tools, the same shape works with with tracer.start_as_current_span(...) directly; there is no await to worry about.

Retries, timeouts, and cancellations

Do not fold retries into a single span. A tool that retried three times and succeeded in 900 ms is materially different from one that succeeded on the first attempt in 300 ms, and a single span hides that.

tool.search_docs                    ← the logical operation, 900ms
├── attempt 1                       ← 120ms, ERROR
├── attempt 2                       ← 150ms, ERROR
└── attempt 3                       ← 630ms, OK

You can model this as child spans per attempt, or as events on the outer span if you do not need per-attempt timings. Child spans are better — they give you a latency distribution for attempts, which is what you need when tuning retry backoff.

Timeouts deserve the same treatment. Record tool.timeout_ms as an attribute on the tool span, and when it fires, set status ERROR with a reason. A timeout and a genuine failure look identical in an aggregate error rate if you do not label them.

Cancellation is not an error. If the user navigated away or the run was superseded, mark the span accordingly — either leave it UNSET with a tool.cancelled attribute, or set ERROR with a distinct status description. Lumping cancellations into your failure rate makes every post-incident review noisier.

Making it queryable: an attribute schema

Span names get you a waterfall. A consistent attribute schema gets you answers. Define it once, document it, and enforce it in code review.

A workable minimum for tool spans:

AttributeTypeExample
tool.namestringissue_refund
tool.namespacestringbuiltin / mcp / host
tool.call_idstringcall_8f3a…
tool.argument_countint3
tool.argumentsstringredacted JSON
tool.result.sizeint512
tool.duration_msint412
tool.approval.requiredbooltrue
tool.approval.decisionstringapproved / rejected / auto
gen_ai.agent.namestringrefund_agent
gen_ai.request.modelstringgpt-5.2
http.response.status_codeint200
retry.countint2

Two entries there are worth calling out because they are agent-specific and routinely missed:

tool.approval.decision — if you have a consent flow, the approval outcome is one of the most valuable things to slice on. “Refunds are rejected 40% of the time” is a product insight that only exists if you recorded it.

gen_ai.agent.name — in a multi-agent system, tool spans are ambiguous without knowing which agent invoked them. In a handoff architecture the agent name changes mid-trace, and you need it on every span to reconstruct who did what.

Align with the GenAI semantic conventions where they exist rather than inventing llm.* keys. The conventions are not exhaustive for tool calls yet, but the model and agent naming pieces are settled, and consistency with them saves a migration later.

Streaming tools

A streaming tool produces output over time, and the temptation is to emit one span per chunk. Resist it — you will create thousands of spans and lose the operation boundary.

Keep one span for the whole stream and record chunks as events:

with tracer.start_as_current_span("tool.stream_response") as span:
    span.set_attribute("tool.name", "stream_response")
    chunk_count = 0

    async for chunk in tool.stream(**arguments):
        chunk_count += 1
        span.add_event("chunk", {"index": chunk_count, "bytes": len(chunk)})
        yield chunk

    span.set_attribute("tool.chunk_count", chunk_count)

The span’s duration is the total stream time — which is the number a user notices. The events give you the shape of the stream inside it: whether it stalled at chunk 40, or whether the first token took three seconds.

Instrumenting across process boundaries

Tools frequently call other services. If the tool does an HTTP request, you want that request to appear as a child span in the same trace — not as a separate trace in your HTTP service’s dashboard.

That is context propagation. When your service A calls service B, A includes a trace ID and span ID in the context; B creates a new span whose parent is A’s span. The default propagator uses the W3C TraceContext headers:

traceparent: 00-<trace-id>-<parent-id>-<trace-flags>

For example:

00-a0892f3577b34da6a3ce929d0e0e4736-f03067aa0ba902b7-01

Instrumented HTTP clients inject this automatically. If a tool talks to a service through a path your instrumentation does not cover — a raw socket, a queue, a WebSocket — you have to inject and extract manually via the Propagators API. When you do, remember the security note: on protocols without a dedicated metadata field, the receiving side must extract and remove the context before processing, or you risk undefined behaviour.

Two other propagation rules worth internalizing:

  • Do not send internal trace IDs or baggage to external, untrusted endpoints. They can reveal architecture and business logic.
  • Do not put sensitive data in baggage. Baggage propagates across every service boundary in the call path and may be logged anywhere along it.

What to build dashboards on

Once your spans are consistent, the questions get easy:

  • Latency distribution per tool — p50/p95 on tool.name. The slowest tool is usually not the one anyone suspects.
  • Error rate per tool — filter status = ERROR grouped by tool.name, split by retry.count > 0 to separate transient from real failures.
  • Model time versus tool time — sum durations of llm.call versus tool.* within agent.run. The ratio tells you whether optimizing prompts or optimizing tool implementations is the higher-leverage work.
  • Approval rejection rate — tool.approval.decision = rejected grouped by tool.name. Feed this back into your consent classification.
  • Trace completeness — how many agent.run traces contain at least one tool.* span. A run that made no tool calls when it should have is a routing bug you can only see in the trace.

The prerequisite for all five is the schema. Spans without attributes produce waterfalls you can stare at; spans with attributes produce dashboards you can act on.

A common failure mode: instrumenting only the happy path

Wrap the whole tool body, not just the success return. If your decorator sets attributes after the tool returns, a throwing tool produces a span with a name, a duration, and nothing else — which is exactly the case you needed the data for.

# The attribute assignment after this line never runs on failure.
result = await tool.fn(**arguments)
span.set_attribute("tool.result.size", len(result))

Set your identity attributes (tool.name, tool.call_id, arguments) before invoking the tool, and result attributes after — inside a try/except that records the failure first.

Further reading