What is an AI agent audit trail?
An AI agent audit trail is the record you can hand to a reviewer to show what an agent did and under whose authority. It is narrower than "everything the agent logged" and stricter than a transcript: each row describes one change to something that matters, with enough context to judge it later.
Teams often have three other records and mistake one of them for an audit trail. Each is useful. None of them answers the accountability question on its own.
| Record | Answers | Weakness as audit evidence |
|---|---|---|
| Chat transcript | What was said | Narrated by the model; long; the reason and the action are mixed together |
| Traces (for example OpenTelemetry GenAI spans) | Which agent and tool calls ran, how long, with what arguments | Built for debugging and cost; rarely ties a call to an approved process step |
| Application logs | What the system did internally | Unstructured; identity is whatever the caller sent |
| Audit trail | What changed, who was accountable, at which step, and why | Only as good as its identity model and its scope |
Why is auditing AI agent actions harder than auditing people?
Because an agent acts under someone else's authority and describes itself. When a person edits a record, the session belongs to them. When an agent edits it, two parties are involved: the automation that performed the change and the human whose credential it ran under. Blend them into one "user" field and the row can no longer answer either question.
Agents also narrate. Ask a model what it did and it will tell you, fluently and sometimes wrongly. Anything the agent writes about itself, its name, its role, its claim that a check passed, is a claim. A trail built on claims is a story, not evidence.
Finally, agents act through many tools in one session, and their output varies run to run. That makes the process step the most useful anchor: "this change happened at the review step of the release workflow" is stable even when the model's wording is not.
Delegation adds a further layer. A lead agent hands a sub-task to another agent, which calls a tool, which changes a record. If each hop overwrites the identity of the one before, the trail shows only the last automation in the chain. Record the chain where you can, and at minimum keep the human anchor fixed: whoever started the work, the credential that authorised it does not change as agents pass the task along.
What should an AI agent audit trail record?
Record the change, the context and the two identities. The fields below are a practical minimum; regulated settings, such as high-risk systems under the EU AI Act, will add their own.
Granularity matters as much as the field list. One row per meaningful change is reviewable; one row per token or per tool call is noise that hides the change you are looking for. Keep the full conversation, if you keep it at all, as linked supporting material. And treat the trail itself as sensitive: never write secrets into it, and keep personal data out unless your data rules require it, because an audit log is usually kept longer and read more widely than the system it describes.
| Field | Why it matters | Example |
|---|---|---|
| Timestamp | Ordering and retention | 2026-10-04T10:05:12Z |
| Entity and action | What was touched, and how | task 4f2a… updated |
| Before and after | What actually changed, field by field | status: review → completed |
| Agent label | Which automation claims it acted | implementer |
| Verified human | Whose authority it ran under, from the credential | the signed-in account |
| Process step | Where in the process it happened | release workflow · approve |
| Decision context | Why, in a sentence | tests green, reviewer signed off |
| Tenant or workspace | Isolation and scoping | workspace: payments |
What is dual attribution in an audit trail?
Dual attribution records two separate identities on every action: the agent's self-declared label, and the human derived on the server from the verified credential. The first answers "which automation says it did this". The second answers "whose authority was this done under". Keep both, and never let one stand in for the other.
The property that makes the human half trustworthy is an absence: no request field, header or tool parameter can carry it. If the identity can only come from the authenticated credential, a forged value has nowhere to land. Test any audit design by asking where the human identity enters. If the answer is "the caller sends it", it is a label, not evidence.
Which standards and rules apply to AI agent logging?
No single standard governs AI agent audit trails yet, but several sources shape what a good one looks like. Treat them as inputs to your design, not as boxes a product can tick for you.
| Source | What it says | Status |
|---|---|---|
| EU AI Act, Article 12 | High-risk AI systems must technically allow automatic recording of events (logs) over their lifetime, for traceability and monitoring | Regulation (EU) 2024/1689; high-risk dates moved by the 2026 omnibus |
| EU AI Act, Article 26(6) | Deployers keep the logs under their control for a period appropriate to the purpose, at least six months unless other law says otherwise | Same regulation |
| IETF draft-sharif-agent-audit-trail | A JSON audit record format for agents with mandatory fields, an action taxonomy and SHA-256 hash chaining for tamper evidence | Individual Internet-Draft, version 06 (29 September 2026); not an IETF standard |
| OpenTelemetry GenAI semantic conventions | Span names and attributes for agent runs and tool calls, such as invoke_agent and execute_tool | Observability conventions, still evolving |
- Use Article 12 to decide which events must be captured at all, then use your audit trail for the ones that change state
- Use the IETF draft as a field checklist and as a reference if you need tamper evidence
- Use OpenTelemetry for the debugging and cost view, and join it to the audit row through a run or conversation id
sourcesEU AI Act, Article 12 (record-keeping)EU AI Act, Article 26 (deployer obligations)IETF draft-sharif-agent-audit-trailOpenTelemetry GenAI agent spans
How do you build an audit trail for AI agents?
Build it where the agent changes state, not where it talks. The steps below work with any stack.
- 1
Choose the surfaces that matter
List the records agents mutate: tickets, tasks, deployments, customer records. Audit those completely before auditing anything else.
- 2
Derive the human on the server
Take the person from the verified token or session. Remove every input path through which a caller could supply it.
- 3
Let the agent label itself, and call it a label
Store the agent name the caller declares in its own field. Useful for filtering, never used as proof.
- 4
Write field-level before and after
A row that says "updated" without the change is not reviewable. Record each changed field with its old and new value.
- 5
Anchor every change to a process step
If the work runs through a defined workflow, store the step and the reason given when it advanced.
- 6
Decide retention and integrity explicitly
Write down how long rows are kept, who can read them, and whether you need tamper evidence such as hash chaining. Then check the tool does it, rather than assuming.
- 7
Test it like an auditor
Pick a finished run and reconstruct it from the trail alone: what changed, at which step, why, and under whose credential. Then try to forge the human identity through every input you expose. Both tests should be boring.
What does ConvOps record, and what does it not?
ConvOps writes an audit row for every change to tasks and workflow instances, with field-level before and after. The actor field holds the label the agent declares. The human is stamped by the server from the verified credential; no tool parameter can carry it. Each workflow advance also lands in the step history with the reason given, and a high-level activity feed tells the story of each task.
The limits matter as much. The audit covers tasks and workflow instances, not every table in the product. Approval is passed explicitly on the advance call, and the row records the verified account behind that call; there is no separate approver field or signed approval. The audit log is read in the web app and through the REST API (GET /api/v1/audit-log, filterable by entity, actor and action), not from inside the chat client. There is no built-in export, retention setting or hash chaining today. If your regime needs those, plan for them outside ConvOps.
{
"actor": "implementer",
"actor_user_name": "Sam Lee (verified)",
"action": "updated",
"entity_type": "task",
"changes": {
"status": { "old": "review", "new": "completed" }
}
}Frequently asked questions
Is a chat transcript an audit trail?
Not on its own. A transcript shows what was said, narrated by the model. An audit trail records what changed, field by field, with the verified human and the process step. Keep transcripts as supporting context.
Can an agent forge the human identity in a ConvOps audit row?
No request field carries it. The server derives the human from the verified credential the call ran under, so there is no parameter through which an agent could supply a different person.
How long should AI agent logs be kept?
It depends on your purpose and the law that applies. For deployers of high-risk AI systems, Article 26(6) of the EU AI Act sets at least six months unless other Union or national law says otherwise. This is general information, not legal advice.
Does ConvOps audit every action an agent takes?
No, and it does not claim to. It audits changes to tasks and workflow instances, the places where agents act on the work, plus step history and an activity feed. Actions an agent takes in other systems belong in those systems' logs.
Does the IETF have a standard for agent audit trails?
Not yet. draft-sharif-agent-audit-trail is an individual Internet-Draft (version 06, September 2026). It is a useful reference for fields and tamper evidence, but it has no formal IETF standing.