the audit guide

AI agent audit trail: what to record, and why

An AI agent audit trail is a durable record of what an agent changed, when, at which step of which process, and on whose authority. A useful one keeps two identities apart: the label the agent declares, and the person verified from the credential. Only the second one is evidence.

the route

7 sections · 10 min

  1. What it is
  2. Why agents differ
  3. What to record
  4. Dual attribution
  5. Standards
  6. How to build it
  7. ConvOps, and limits
01What it is

What is an AI agent audit trail?

An AI agent audit trail is the record you can hand to a reviewer to show what an agent did and under whose authority. It is narrower than "everything the agent logged" and stricter than a transcript: each row describes one change to something that matters, with enough context to judge it later.

Teams often have three other records and mistake one of them for an audit trail. Each is useful. None of them answers the accountability question on its own.

Records an agent leaves behind
RecordAnswersWeakness as audit evidence
Chat transcriptWhat was saidNarrated by the model; long; the reason and the action are mixed together
Traces (for example OpenTelemetry GenAI spans)Which agent and tool calls ran, how long, with what argumentsBuilt for debugging and cost; rarely ties a call to an approved process step
Application logsWhat the system did internallyUnstructured; identity is whatever the caller sent
Audit trailWhat changed, who was accountable, at which step, and whyOnly as good as its identity model and its scope
Transcriptwhat was said
Traces and logswhich calls ran
Audit trailwhat changed, on whose authority
fig 1 Transcripts and traces help. Only the audit row answers who was accountable.
02Why agents differ

Why is auditing AI agent actions harder than auditing people?

Because an agent acts under someone else's authority and describes itself. When a person edits a record, the session belongs to them. When an agent edits it, two parties are involved: the automation that performed the change and the human whose credential it ran under. Blend them into one "user" field and the row can no longer answer either question.

Agents also narrate. Ask a model what it did and it will tell you, fluently and sometimes wrongly. Anything the agent writes about itself, its name, its role, its claim that a check passed, is a claim. A trail built on claims is a story, not evidence.

Finally, agents act through many tools in one session, and their output varies run to run. That makes the process step the most useful anchor: "this change happened at the review step of the release workflow" is stable even when the model's wording is not.

Delegation adds a further layer. A lead agent hands a sub-task to another agent, which calls a tool, which changes a record. If each hop overwrites the identity of the one before, the trail shows only the last automation in the chain. Record the chain where you can, and at minimum keep the human anchor fixed: whoever started the work, the credential that authorised it does not change as agents pass the task along.

03What to record

What should an AI agent audit trail record?

Record the change, the context and the two identities. The fields below are a practical minimum; regulated settings, such as high-risk systems under the EU AI Act, will add their own.

Granularity matters as much as the field list. One row per meaningful change is reviewable; one row per token or per tool call is noise that hides the change you are looking for. Keep the full conversation, if you keep it at all, as linked supporting material. And treat the trail itself as sensitive: never write secrets into it, and keep personal data out unless your data rules require it, because an audit log is usually kept longer and read more widely than the system it describes.

A practical minimum, per row
FieldWhy it mattersExample
TimestampOrdering and retention2026-10-04T10:05:12Z
Entity and actionWhat was touched, and howtask 4f2a… updated
Before and afterWhat actually changed, field by fieldstatus: review → completed
Agent labelWhich automation claims it actedimplementer
Verified humanWhose authority it ran under, from the credentialthe signed-in account
Process stepWhere in the process it happenedrelease workflow · approve
Decision contextWhy, in a sentencetests green, reviewer signed off
Tenant or workspaceIsolation and scopingworkspace: payments
Audit · exampletask history
changeagentperson
statusin reviewdoneimplementerdeclared by the clientdana@example.com from the verified sign-in
fig 2 Illustration. Field-level before and after, with both identities.
04Dual attribution

What is dual attribution in an audit trail?

Dual attribution records two separate identities on every action: the agent's self-declared label, and the human derived on the server from the verified credential. The first answers "which automation says it did this". The second answers "whose authority was this done under". Keep both, and never let one stand in for the other.

The property that makes the human half trustworthy is an absence: no request field, header or tool parameter can carry it. If the identity can only come from the authenticated credential, a forged value has nowhere to land. Test any audit design by asking where the human identity enters. If the answer is "the caller sends it", it is a label, not evidence.

Agent labeldeclared by the caller
Verified humanfrom the credential only
fig 3 Two identities, never merged.
05Standards

Which standards and rules apply to AI agent logging?

No single standard governs AI agent audit trails yet, but several sources shape what a good one looks like. Treat them as inputs to your design, not as boxes a product can tick for you.

Sources to design against, as of October 2026
SourceWhat it saysStatus
EU AI Act, Article 12High-risk AI systems must technically allow automatic recording of events (logs) over their lifetime, for traceability and monitoringRegulation (EU) 2024/1689; high-risk dates moved by the 2026 omnibus
EU AI Act, Article 26(6)Deployers keep the logs under their control for a period appropriate to the purpose, at least six months unless other law says otherwiseSame regulation
IETF draft-sharif-agent-audit-trailA JSON audit record format for agents with mandatory fields, an action taxonomy and SHA-256 hash chaining for tamper evidenceIndividual Internet-Draft, version 06 (29 September 2026); not an IETF standard
OpenTelemetry GenAI semantic conventionsSpan names and attributes for agent runs and tool calls, such as invoke_agent and execute_toolObservability conventions, still evolving
  1. Use Article 12 to decide which events must be captured at all, then use your audit trail for the ones that change state
  2. Use the IETF draft as a field checklist and as a reference if you need tamper evidence
  3. Use OpenTelemetry for the debugging and cost view, and join it to the audit row through a run or conversation id

sourcesEU AI Act, Article 12 (record-keeping)EU AI Act, Article 26 (deployer obligations)IETF draft-sharif-agent-audit-trailOpenTelemetry GenAI agent spans

06How to build it

How do you build an audit trail for AI agents?

Build it where the agent changes state, not where it talks. The steps below work with any stack.

  1. 1

    Choose the surfaces that matter

    List the records agents mutate: tickets, tasks, deployments, customer records. Audit those completely before auditing anything else.

  2. 2

    Derive the human on the server

    Take the person from the verified token or session. Remove every input path through which a caller could supply it.

  3. 3

    Let the agent label itself, and call it a label

    Store the agent name the caller declares in its own field. Useful for filtering, never used as proof.

  4. 4

    Write field-level before and after

    A row that says "updated" without the change is not reviewable. Record each changed field with its old and new value.

  5. 5

    Anchor every change to a process step

    If the work runs through a defined workflow, store the step and the reason given when it advanced.

  6. 6

    Decide retention and integrity explicitly

    Write down how long rows are kept, who can read them, and whether you need tamper evidence such as hash chaining. Then check the tool does it, rather than assuming.

  7. 7

    Test it like an auditor

    Pick a finished run and reconstruct it from the trail alone: what changed, at which step, why, and under whose credential. Then try to forge the human identity through every input you expose. Both tests should be boring.

surfaceswhat agents mutate
identityapproval
before / after
step + reason
retentiondecided, checked
fig 4 Audit where state changes, not where the agent talks.
07ConvOps, and limits

What does ConvOps record, and what does it not?

ConvOps writes an audit row for every change to tasks and workflow instances, with field-level before and after. The actor field holds the label the agent declares. The human is stamped by the server from the verified credential; no tool parameter can carry it. Each workflow advance also lands in the step history with the reason given, and a high-level activity feed tells the story of each task.

The limits matter as much. The audit covers tasks and workflow instances, not every table in the product. Approval is passed explicitly on the advance call, and the row records the verified account behind that call; there is no separate approver field or signed approval. The audit log is read in the web app and through the REST API (GET /api/v1/audit-log, filterable by entity, actor and action), not from inside the chat client. There is no built-in export, retention setting or hash chaining today. If your regime needs those, plan for them outside ConvOps.

an audit row, as stored (example values)
{
  "actor": "implementer",
  "actor_user_name": "Sam Lee (verified)",
  "action": "updated",
  "entity_type": "task",
  "changes": {
    "status": { "old": "review", "new": "completed" }
  }
}

Frequently asked questions

Is a chat transcript an audit trail?

Not on its own. A transcript shows what was said, narrated by the model. An audit trail records what changed, field by field, with the verified human and the process step. Keep transcripts as supporting context.

Can an agent forge the human identity in a ConvOps audit row?

No request field carries it. The server derives the human from the verified credential the call ran under, so there is no parameter through which an agent could supply a different person.

How long should AI agent logs be kept?

It depends on your purpose and the law that applies. For deployers of high-risk AI systems, Article 26(6) of the EU AI Act sets at least six months unless other Union or national law says otherwise. This is general information, not legal advice.

Does ConvOps audit every action an agent takes?

No, and it does not claim to. It audits changes to tasks and workflow instances, the places where agents act on the work, plus step history and an activity feed. Actions an agent takes in other systems belong in those systems' logs.

Does the IETF have a standard for agent audit trails?

Not yet. draft-sharif-agent-audit-trail is an individual Internet-Draft (version 06, September 2026). It is a useful reference for fields and tamper evidence, but it has no formal IETF standing.