Decision Logs as First-Class Data: Post-Mortems for Multi-Agent Systems
Treat agent decisions as structured, queryable data using OpenTelemetry logs and traces — and turn post-mortems into an automated, evidence-driven practice.
Published on • October 8, 2026
AI Assistant

When an autonomous agent takes a wrong action, the question is never “what did the code do” — the code did exactly what it was written to do. The question is why the agent decided that, and answering it requires the decision itself to be recorded as data, not as a prose log line someone wrote with print().
That’s the shift this post is about: promoting agent decisions from incidental debug output to first-class, structured, correlated data — and using that data to run post-mortems that don’t depend on someone remembering to add logging.
Why print() fails as a decision record
OpenTelemetry’s logs documentation draws a distinction that maps directly onto agent systems: logs can be structured, semistructured, or unstructured, and the difference matters more than the encoding.
A log encoded as JSON is not automatically “structured” in the sense of having a stable schema — it may be semistructured. Structured logs imply a consistent schema or well-defined typed fields that downstream processing can reliably depend on.
So {"msg": "agent chose tool X because ...", "run": 123} is not structured in the sense that matters. It’s a JSON blob with free-text in it, and you can’t aggregate it, filter it reliably, or join it against anything.
Production observability needs more. OpenTelemetry is explicit: “Structured logs are preferred in production because their stable schema makes them straightforward to validate, parse, correlate with traces and metrics, and analyze at scale.”
What a decision record contains
Give every agent decision a fixed schema. At minimum:
{
"timestamp": "2026-10-08T02:14:07.881Z",
"severityText": "INFO",
"eventName": "agent.decision",
"attributes": {
"agent.id": "billing_triage",
"agent.version": "1.4.2",
"run.id": "run_9f3a12",
"session.id": "sess_77b1",
"decision.id": "dec_004",
"decision.type": "tool_selection",
"decision.selected": "lookup_invoice",
"decision.candidates": ["lookup_invoice", "issue_refund", "escalate_human"],
"decision.confidence": 0.71,
"decision.reasoning_summary": "Invoice found; refund threshold not exceeded",
"input.token_count": 4820,
"output.token_count": 214,
"cost.usd": 0.0132,
"outcome": "tool_executed",
"human_override": false
}
}
A few field choices worth defending:
decision.idwithinrun.id— decisions are ordered and addressable, so a post-mortem can say “the failure was atdec_004.”decision.candidates— the alternatives considered. Without this you can never distinguish “the agent picked the only plausible option” from “the agent ignored the obvious option.”agent.version— drift investigations are impossible without knowing which build made the call.outcomeandhuman_override— links decision to consequence, which is what turns a log into evidence.
Keep reasoning_summary short and templated. Full chain-of-thought is expensive, often unreliable, and frequently shouldn’t be retained at all; a structured summary of what factors were considered is what post-mortems actually need.
Correlate logs with traces automatically
OpenTelemetry’s logs signal was designed for exactly this: “When you add autoinstrumentation or activate an SDK, OpenTelemetry will automatically correlate your existing logs with any active trace and span, wrapping the log body with their IDs.”
A LogRecord carries these top-level fields:
| Field | Purpose in a post-mortem |
|---|---|
Timestamp / ObservedTimestamp | When it happened vs. when it was ingested |
TraceId / SpanId | Ties the decision to the exact span that produced it |
TraceFlags | Sampling decision — keep unsampled incident traces |
SeverityText / SeverityNumber | Filtering and routing |
Body | The structured payload |
Resource | Which service/agent emitted it |
Attributes | Your decision schema |
EventName | agent.decision, agent.tool_call, agent.escalation |
Because TraceId rides along with every record, you can pivot from an anomalous span to every decision inside it, and from a decision to the retrieval results, model call, and tool responses that surrounded it — without having manually threaded IDs through your code.
Practical tip: emit decision logs at an appropriate severity (INFO for normal choices, WARN for low-confidence or fallback choices, ERROR for failed decisions), and set your collector’s transformprocessor to enforce the schema on ingest. Failing schema validation at the collector turns a malformed logger call into a visible error rather than a silent gap three weeks later.
Choosing what to sample
Decision logs are high-volume. The failure mode is sampling away the one record you needed.
- Never sample decisions that resulted in an error, override, or escalation. Force-sample them regardless of the trace’s sample decision.
- Sample ordinary decisions at a rate that keeps weekly distributions stable — you need enough volume for drift comparisons.
- Keep
TraceFlagsso you know whether a “missing” record was never emitted or was sampled out. This distinction prevents false post-mortem conclusions.
OpenTelemetry supports tail sampling at the collector, where you can decide after seeing the whole trace whether to keep it — the right place for “keep any trace where an agent decision was overridden.”
From records to post-mortems
A post-mortem is an evidence reconstruction. With decision data, the reconstruction becomes queryable.
Step 1 — Anchor on the incident. Pull all records for the affected run.id, or all runs for an agent.id in a time window.
SELECT timestamp, attributes['decision.id'], attributes['decision.selected'],
attributes['decision.confidence'], attributes['outcome']
FROM otel_logs
WHERE attributes['agent.id'] = 'billing_triage'
AND timestamp BETWEEN @start AND @end
ORDER BY timestamp;
Step 2 — Find the divergence point. In a timeline of decisions, the divergence point is the first decision whose outcome differs from comparable successful runs. Compare against a golden run or a cohort of similar sessions.
Step 3 — Classify. Was it a bad decision with good information (model problem), good decision on bad input (retrieval/tool problem), or a decision the schema didn’t allow a good option for (design problem)? The decision.candidates field is what lets you tell these apart.
Step 4 — Quantify blast radius. Count affected runs, cost, and users from the same window:
SELECT attributes['run.id'], COUNT(*), SUM(attributes['cost.usd'])
FROM otel_logs
WHERE eventName = 'agent.decision'
AND attributes['outcome'] = 'escalated_human'
GROUP BY attributes['run.id'];
Step 5 — Write the timeline. Five entries maximum: first divergence, detection, containment, remediation, prevention. Every entry cites a run.id/decision.id — that’s what separates evidence from narrative.
What to measure after a post-mortem
The record schema should make these cheap:
- Override rate — how often humans corrected the agent. Rising override rate is leading-indicator drift.
- Confidence distribution — a bimodal shift after a deploy is a red flag even if accuracy looks flat.
- Decision mix — the share of runs selecting each tool. A tool that quietly stops being selected is often a broken tool, not an agent improvement.
- Time-to-detect — gap between the first anomalous decision and the alert firing. This is the number your alerting work actually optimizes.
Governance and retention
Decision records contain reasoning about real users. Treat them as sensitive:
- Redact PII at the collector (
transformprocessorcan drop or hash attributes) - Set retention deliberately — post-mortem data doesn’t need to live forever, but the window should cover your slowest-to-detect failure class
- Restrict access to reasoning fields to the roles that run post-mortems
- Document what’s recorded, in the same way you’d document what’s in your access logs
Anti-patterns
Logging decisions as prose. "Decided to refund because it seemed right" is unfalsifiable and unqueryable.
Only logging failures. You need successful decisions as the comparison cohort, or you can’t tell “wrong choice” from “unusual but correct choice.”
Post-mortems from memory. If the timeline comes from Slack scrollback rather than records, your logging isn’t first-class yet.
Schema drift. Renaming an attribute without a migration quietly breaks every query. Version the schema (decision.schema_version) and keep old fields readable.
One giant log line. Emit one record per decision, not one per run. Runs contain dozens of decisions and post-mortems need to address them individually.
Wrapping up
A multi-agent system’s failure modes live in its decisions, not its stack traces. Give those decisions a stable schema, correlate them to traces with OpenTelemetry’s built-in log-trace association, sample conservatively around incidents, and the post-mortem writes itself from queries instead of from memory. The first time you can answer “why did it do that?” in one SQL query, you’ll wonder how you operated without it.