Ops series: observability for agents with traces, not transcripts
AI agent observability works best as distributed tracing: model and tool calls become spans with tokens, cost and decisions. Use OpenTelemetry, avoid lock-in.
This post is also available in Deutsch.
AI agent observability is distributed tracing with a few extra attributes. A single agent run is a tree of work: a plan, several model calls, tool calls, hand-overs and approvals. Record each as a span, attach tokens, cost and policy decisions as attributes, and you can answer the questions that matter in operations: what happened, how long did it take, what did it cost and why did it do that. Chat transcripts alone cannot do this, because they show the conversation but not the timing, the failures, or the decisions around it.
This post is part of the ops series that treats agents as services you run, not demos you watch.
Why transcripts are not enough
A transcript is a useful artefact for reading what a model said. As an operations tool it has gaps:
- It has no timing. You cannot tell whether a slow run waited on the model, a tool or a human approval.
- It is hard to aggregate. Counting failed tool calls across a thousand runs means parsing free text.
- It mixes the sensitive and the structural. Prompts and outputs often contain personal or confidential data, while the structure you need for monitoring is metadata.
- It does not link across systems. When an agent calls a service that has its own logs, nothing connects the two.
Traces solve each of these. They have a duration per step, typed attributes you can query, and a trace ID that can travel with the request.
An agent run as a trace
Map the pieces of a run to spans:
| Part of the run | Span | Useful attributes |
|---|---|---|
| The whole run | root span agent.run | run ID, agent name, version, outcome |
| Planning | agent.plan | number of steps, plan size |
| A model call | model.call | model, input and output tokens, finish reason, cost |
| A tool call | tool.call | tool name, status, duration, argument size |
| Policy check | event on the tool span | decision (allow, deny, approval), rule ID |
| Waiting for a human | approval.wait | approver role, wait time, result |
| Hand-over to another agent | agent.handover | from, to, contract ID |
Keep content out of attributes by default. Store sizes, hashes and IDs, and keep the raw prompt and output in a separate, access-controlled store that the trace links to. That way, anyone who can read dashboards does not automatically read customer data.
Map the golden signals to agents
The classic monitoring vocabulary still applies. The Google SRE book puts it in one sentence:
The four golden signals of monitoring are latency, traffic, errors, and saturation.
— Site Reliability Engineering, Chapter 6: Monitoring Distributed Systems, Google
For an agent service, a workable translation is:
- Latency. Time per run, split into model time, tool time and approval wait. Alert on the model and tool parts; approval wait is a business metric, not a fault.
- Traffic. Runs started per minute, by agent and trigger.
- Errors. Failed runs, failed tool calls, policy denials and schema validation failures. Also count wrong answers if you have evals that detect them, which a later post in this series covers under SLOs for agents.
- Saturation. Rate limits hit, queue depth, concurrent runs, token budget consumed.
Add one signal that classic services lack: cost. Tokens and money per run are first-class metrics, and a loop that spends ten times the usual amount is an incident even if every call succeeded. See cost is a platform concern for budgets and attribution.
Use OpenTelemetry conventions, avoid lock-in
Agent frameworks and observability vendors each invent their own attribute names. Then every dashboard, alert and export is tied to one product. The OpenTelemetry project argues for common naming for exactly this reason:
establish standards around the shape of the telemetry generated by agent apps to avoid lock-in
— AI Agent Observability - Evolving Standards and Best Practices, OpenTelemetry
The project’s generative-AI work defines what that shape is for model calls:
For generative AI, these conventions streamline monitoring, troubleshooting, and optimizing AI models by standardizing attributes such as model parameters, response metadata, and token usage.
— OpenTelemetry for Generative AI, OpenTelemetry
Practical advice:
- Emit OpenTelemetry spans from the agent runtime and use the GenAI attribute names where they exist for model calls and token usage. Check the current semantic conventions before you fix names; the area has been evolving.
- Namespace your own attributes (for example
agent.run.id,policy.decision) for what the standard does not cover, and document them. - Export through a collector. The application sends to an OpenTelemetry Collector, and the collector decides where data goes. Changing a backend then does not require changing agents.
- Propagate context. Pass the trace context in calls to tools and other services so their spans join the same trace.
A minimal configuration for the collector side:
receivers:
otlp:
protocols:
http: {}
processors:
batch: {}
attributes/redact:
actions:
- key: gen_ai.prompt
action: delete
exporters:
otlphttp:
endpoint: https://telemetry.example.org
service:
pipelines:
traces:
receivers: [otlp]
processors: [attributes/redact, batch]
exporters: [otlphttp]
The redaction processor in the collector is a safety net, not a substitute for not emitting content in the first place.
What to alert on
Start with a handful of alerts that point to a person who can act:
- Run failure rate above a threshold for an agent over 15 minutes.
- p95 tool-call latency above normal for a given tool.
- Cost per run above a multiple of the trailing median, or a budget burning faster than planned.
- A spike in policy denials, which can mean a misconfigured agent or an attempted abuse. See policy decides, audit proves.
- Approvals waiting longer than the agreed time.
Link each alert to a runbook and to a query that shows the offending traces. The wider point, that none of this is special, is made in agent ops is just ops.
Limits
Traces tell you what the agent did, not whether it was right. Pair them with evals and sampled human review. Sampling is also a trade-off: keep every failed or expensive run, and sample the healthy ones. And be honest about what your conventions cover; semantic conventions for agents are still maturing, so expect to adjust attribute names over time.
Key takeaways
- Model an agent run as a trace: spans for plan, model calls, tool calls, approvals and hand-overs.
- Put tokens, cost and policy decisions on spans as attributes; keep raw content out of them.
- Translate the four golden signals and add cost as a fifth.
- Use OpenTelemetry and a collector to avoid backend lock-in.
- Alert on failures, saturation, cost and policy denials, each with a runbook.