How to Build an Audit Trail for AI Agent Actions
What an AI agent audit trail should record, how to structure audit events with trace and correlation IDs, and how to keep them tamper-resistant and private.
Quick answer
An AI agent audit trail records every consequential thing an agent does in a form you can trust later: who requested the task, which agent acted and on whose behalf, what it decided, which tools it called with which inputs and outputs, who approved what, and what changed in business systems, all linked by a trace ID and written to append-only storage the agent cannot alter. Design it before the agent can change data, call paid APIs or contact customers, because reconstructing those actions afterwards from scattered application logs rarely works.
Why agents need a dedicated audit trail
An AI agent becomes hard to account for the moment it can change records, spend money or message customers without a person approving every step. When something goes wrong, three questions come up immediately: what exactly did it do, why did it do it, and who allowed it? Application logs answer fragments of those questions across many systems. A purpose-built audit trail answers them in one place.
Audit trails also make other controls work. Rollback depends on knowing precisely which actions to reverse (see AI agent rollback); accountability depends on evidence (see who is responsible when an AI agent makes a mistake); and regulated uses may require automatic logging. The EU AI Act, for example, requires high-risk AI systems to support automatic recording of events over their lifetime and providers to keep those logs for at least six months unless other law provides otherwise.
Audit trail vs observability vs application logs
| Audit trail | Observability traces | Application logs | |
|---|---|---|---|
| Purpose | Accountability, investigation, compliance | Debugging, performance, quality | Troubleshooting a service |
| Coverage | Every consequential action, complete | Often sampled | Whatever developers log |
| Retention | Months to years, by policy | Days to weeks | Days to weeks |
| Integrity | Append-only, tamper-evident | Best effort | Best effort |
| Access | Restricted, audited | Engineering | Engineering |
Key takeaway
Use observability to understand agents and an audit trail to account for them. Share IDs between them so an investigator can jump from an audit event to the detailed trace.
What to record
Record events at the level of decisions and actions, not every token. For each agent run:
| Field group | Contents |
|---|---|
| Identity | Agent identity and version; user it acted for; authorizing session or delegation; tenant |
| Task | Request or trigger, channel, task type, autonomy level in force |
| Context references | Sources retrieved (document IDs, record IDs, versions), not necessarily full text |
| Decisions | Chosen action and a short structured reason; model and prompt version |
| Tool calls | Tool name and version, validated inputs, outputs or output references, status, duration |
| Approvals | Approver identity, what they were shown, decision, timestamp |
| Changes | System, record, before/after or diff reference, external IDs (refund ID, message ID) |
| Timing and linkage | Timestamps, trace ID, event ID, parent event ID, correlation IDs to business systems |
Trace, event and correlation IDs
IDs turn scattered records into a story. Use a trace ID for the whole agent run (W3C Trace Context's trace ID works well if you already use distributed tracing), a unique event ID for each audit record with a parent event ID to show sequence, and correlation IDs that connect to business systems: the order number, ticket ID or payment reference. Pass the trace ID into every tool call and store it with every change, so an auditor can start from a refund in your payment system and find the agent run, the approval and the inputs that led to it.
If you trace AI workloads with OpenTelemetry, its generative AI semantic conventions define attributes for model calls, tools and agent runs. They are still marked as in development, so treat attribute names as subject to change.
An example audit event
{
"event_id": "evt_01J9Z6K4Q8",
"parent_event_id": "evt_01J9Z6K3X2",
"trace_id": "4bf92f3577b34da6a3ce929d0e0e4736",
"timestamp": "<ISO 8601 UTC timestamp>",
"event_type": "tool_call.completed",
"agent": { "id": "support-agent", "version": "3.4.1", "model": "provider/model-name@version" },
"on_behalf_of": { "user_id": "cust_88231", "delegation_id": "del_5521" },
"task": { "type": "refund_request", "channel": "chat", "autonomy_level": 4 },
"tool": { "name": "issue_store_credit", "version": "2" },
"input": { "order_id": "ORD-104552", "amount": 18.00, "currency": "GBP", "reason_code": "late_delivery" },
"policy": { "rule": "refund_under_25_auto", "result": "allowed" },
"result": { "status": "success", "external_id": "cr_77310" },
"correlation": { "order_id": "ORD-104552", "ticket_id": "T-55102" },
"sources": ["policy:refunds@v12", "order:ORD-104552@rev7"]
}Pro tip
Store the policy rule that allowed the action and the versions of the sources used. When policies change, you can still explain why an older action was permitted at the time.
Tamper resistance, privacy and retention
- Append-only storage the agent and application cannot modify or delete
- Separate permissions for writing, reading and administering logs; log access itself is audited
- Hash chaining or signing so gaps and edits are detectable
- Data minimization: store references and hashes for large or sensitive content; redact secrets and payment data
- Tiered retention: decision and change records kept longer than raw prompts and outputs
- Deletion by policy only, through an audited process that respects legal holds and privacy rights
Putting agents in front of real systems?
ZSpace Labs designs agent audit trails, identity and approval records alongside the agents themselves, so every action can be explained and reversed. See AI automation services.
Common mistakes
- Relying on model provider logs or chat transcripts as the audit record
- Logging the agent's actions without the user it acted for
- No link between the agent run and the change in the business system
- Storing full prompts with personal data indefinitely
- Letting the same service that acts also edit or delete its own logs
- Designing the audit trail after an incident
Conclusion
An audit trail is what makes an agent accountable: identity, task, decisions, tool calls, approvals and changes, linked by IDs and stored where they cannot be quietly altered. Build it alongside access control and agent authentication, connect it to observability for detail, and use it as the foundation for rollback and incident response.
Common questions.
A tamper-resistant record of what an AI agent did: who asked, which agent acted and on whose behalf, which tools it called with which inputs, what came back, what was approved and what changed in business systems, linked by IDs so any action can be traced end to end.