Ops series: audit trails are not logs
An AI audit trail must be complete, attributable and tamper-evident; logs are for debugging. How to separate them on agent platforms and what goes where.
This post is also available in Deutsch.
An AI audit trail and an application log answer different questions. Logs and traces help engineers find out why something broke, so they can be sampled, filtered and rotated. An audit trail has to prove what happened, so it must be complete, attributable to an identity and tamper-evident. Agent platforms need both, fed from the same run but stored, protected and retained differently. Mixing them gives you a log that is too thin to trust or an audit store that is too noisy to read.
Two jobs, two sets of requirements
Ask what each record is for.
- Telemetry (logs, traces, metrics) exists to debug and to operate. It is useful when it is cheap, fast to query and tolerant of loss. Dropping 90% of debug lines in a healthy system is a feature.
- An audit trail exists to answer questions later, often from someone who was not there: who authorised this action, what exactly was it, what did the policy decide, what changed. It is only useful if nothing is missing and nobody could have quietly edited it.
| Property | Telemetry | Audit trail |
|---|---|---|
| Main reader | Engineer debugging | Auditor, security, owner, sometimes a regulator or customer |
| Completeness | Sampling is fine | Every relevant event, no gaps |
| Attribution | Often a service name | A specific identity: user, agent, delegation chain |
| Integrity | Best effort | Tamper-evident (for example hash-chained) |
| Retention | Days or weeks, rotated | Defined by policy, often months or years |
| Content | Verbose, technical | Structured, minimal, stable schema |
| Failure mode | A missing line is annoying | A missing record is a finding |
What telemetry looks like for agents
Agent telemetry has a natural shape: a run is a tree of steps. OpenAI’s Agents SDK documentation describes the top-level unit this way:
Traces represent a single end-to-end operation of a “workflow”.
Source: OpenAI Agents SDK docs, Tracing.
Inside a trace you find spans for model calls, tool calls and hand-offs, with timings and often the payloads. That is exactly what you need to see why a run was slow or took a wrong turn. It is also why telemetry makes a poor audit record: it is detailed in the wrong way (large prompts, noisy), usually sampled, and often kept in systems that many engineers can edit or purge.
The wider ecosystem is working on common formats. The OpenTelemetry project writes that its goal is to:
establish standards around the shape of the telemetry generated by agent apps to avoid lock-in
Source: OpenTelemetry, AI Agent Observability - Evolving Standards and Best Practices (2025-03-06).
Use such standards for telemetry, so you can switch backends. For the practical side of tracing agents, see observability for agents. For audit you need a different design.
What an audit trail must contain
An audit event should be small, structured and boring. For an agent platform, a useful minimum is:
- When: a timestamp from a trusted clock.
- Who: the acting identity, including on whose behalf and through which delegation chain. See agent identity and delegation.
- What: the action, such as a tool call with its (redacted) arguments, an approval, a policy change, a deployment.
- Decision: what the policy gate decided (allow, deny, require approval) and which rule applied.
- Outcome: success or failure, and a reference to the result.
- Correlation: the run and trace id, so you can jump to telemetry for detail.
- Integrity data: a hash linking this record to the previous one.
An example event:
{
"seq": 18422,
"time": "2026-07-28T09:14:03Z",
"actor": { "agent": "refund-triage", "run": "run_7f3a", "on_behalf_of": "user:support-lead" },
"action": { "tool": "tickets.comment", "args": { "ticket": "SUP-1841", "comment_sha256": "9c1e..." } },
"decision": { "result": "allow", "rule": "comment-after-review" },
"outcome": "ok",
"trace_id": "4bf92f3577b34da6a3ce929d0e0e4736",
"prev_hash": "a41c...",
"hash": "e07b..."
}
Note the hash of the comment text instead of the text: an audit record should prove what was done without becoming a second copy of sensitive data.
Which events go where
One run produces both streams. A simple rule for routing:
| Event | Telemetry | Audit |
|---|---|---|
| Model call latency and token counts | yes | no (aggregate cost may be audited) |
| Full prompt and response text | sampled, with redaction | no, only hash or reference |
| Tool call that reads data | yes | yes, if the data is sensitive |
| Tool call that changes state | yes | yes, always |
| Policy decision (allow, deny, approve) | yes | yes, always |
| Human approval or rejection | yes | yes, always |
| Change to agent, skill or policy | yes | yes, always |
| Retry, cache hit, debug message | yes | no |
The test: if someone asks in a year “who allowed this and what exactly happened”, could you answer from this stream alone? If yes, it belongs in audit.
Making an audit trail tamper-evident
Tamper-evident does not mean tamper-proof. It means that changing or deleting a record becomes detectable. The usual building blocks:
- Append-only writes. The platform component that records events can add but not modify or delete.
- Hash chaining. Each record contains the hash of the previous one, so removing or altering one breaks every later hash.
- Periodic anchoring. Store the latest hash somewhere the platform’s administrators cannot rewrite, such as a separate system or write-once storage.
- Separate access. The people who operate agents are not the people who can administer the audit store.
- Verification. A job recomputes the chain on a schedule and alerts on a break.
Without anchoring, an attacker with full control of the store can rebuild the whole chain, so access control still matters.
Retention and privacy
Completeness has a cost, and not only in storage. Audit records can contain personal data. Keep them minimal: identifiers and references instead of content, redacted arguments, hashes where you need to prove content without keeping it. Set retention by policy and by the rules that apply to you, and delete on schedule. A long retention period is not automatically better.
For frameworks that ask for evidence of oversight, the audit trail is the raw material. The NIST AI RMF, for instance, says of itself:
The NIST AI Risk Management Framework (AI RMF) is intended for voluntary use and to improve the ability to incorporate trustworthiness considerations
Source: NIST, AI Risk Management Framework.
It does not prescribe a log format, so the design choices above are yours. What it asks, in effect, is that you can show how your system is governed and measured, and a complete, attributable record is how you show it. For how policy and audit fit together, see policy decides, audit proves.
Common mistakes
- Using the application log as the audit trail. It is sampled, rotated, editable and unstructured.
- Auditing everything. A trail that includes every debug event hides the ones that matter and drives up cost.
- Anonymous actors. “agent-service” is not an identity. Record the run, the agent and the human or system it acted for.
- Storing secrets or full prompts. Audit stores are read by more people than you expect.
- Never verifying. A chain that nobody checks is decoration.
Key takeaways
- Telemetry helps you debug and may be sampled and rotated; an audit trail must be complete, attributable and tamper-evident.
- Feed both from the same run, but store, protect and retain them separately.
- Audit every state change, every policy decision, every approval and every change to agents, skills and policies.
- Hash-chain records, anchor the latest hash elsewhere and verify the chain regularly.
- Keep audit records minimal to protect privacy and cost.
Sources
- AI Agent Observability - Evolving Standards and Best Practices, OpenTelemetry, 2025-03-06.
- Tracing, OpenAI Agents SDK documentation.
- AI Risk Management Framework, NIST, 2023-01-26.