Logging CLI Agent Actions for Audit
Logging what the model said is not an audit trail. Real cli agent action logging records what your code executed — here's how to rebuild narration logging into structured, queryable audit records.
def run_tool(name, args):
print(f"Running {name}...")
result = execute(name, args)
print(f"Got result: {result[:80]}")
return resultMost agent logging is useless for the one job it's supposed to do. Teams sprinkle print statements and log what the model said — "I'll check the deploy status now" — and call it an audit trail. But when something goes wrong and you need to know what actually happened, the model's narration tells you its intentions, not its actions. Real cli agent action logging records what your code executed: which tool ran, with what arguments, against what, with what result. This teardown takes the logging almost everyone writes first and rebuilds it into something you can actually audit.
The gap between "what the model said" and "what the code did" is exactly where incidents hide. An audit log that captures only the former is theater.
Before: Logging the Narration
Here's the logging most agents ship with. It prints the model's text and maybe a note that a tool ran.
[object Object], ,[object Object],(,[object Object],):
,[object Object],(,[object Object],)
result = execute(name, args)
,[object Object],(,[object Object],)
,[object Object], resultWhat this does: Prints a human-readable line before and after each tool call. It looks like logging and reads fine in the terminal during a session, which is exactly why teams mistake it for an audit trail. It is not one.
Why It Fails
This logging fails the moment you need it, and it fails in several ways at once. It's unstructured text, so you can't query it — answering "did the agent ever run a delete on the payments service" means grepping freeform strings and hoping the phrasing was consistent. It's ephemeral, printed to a terminal that scrolls away, so there's no record after the session ends.
It logs intent mixed with action, so you can't tell what the model proposed from what the code executed — and those differ constantly, because the dispatch layer blocks, modifies, and rejects requests. A log that shows "I'll delete the old logs" tells you nothing about whether the delete was allowed, ran, or was refused by your allowlist. For an audit, that's the only question that matters.
And it captures nothing about the actor or the context: no timestamp you can trust, no session identifier to correlate a sequence of actions, no record of rejected requests. When someone asks "what did the agent do at 2 a.m. that broke staging," this logging can't answer, because it was never designed to be read after the fact. Effective cli agent action logging has to be built for the interrogation, not the demo.
⚠️ Common mistake: Treating the model's narration as an audit log. What the model says it will do is not what your code did — the dispatch layer sits between them, allowing some actions, blocking others, altering arguments. An audit trail must record the executed action from the code's perspective, or it's recording fiction.
After: Structured, Append-Only Action Records
The rebuilt logging captures each executed action as a structured record — machine-readable, timestamped, correlated, and written to durable, append-only storage. Every tool dispatch produces one line of queryable JSON.
[object Object], json, time, uuid, pathlib
SESSION = uuid.uuid4().,[object Object],[:,[object Object],]
LOG = pathlib.Path.home() / ,[object Object], / ,[object Object],
,[object Object], ,[object Object],(,[object Object],):
LOG.parent.mkdir(exist_ok=,[object Object],)
record = {
,[object Object],: time.time(),
,[object Object],: SESSION,
,[object Object],: action,
,[object Object],: redact(args), ,[object Object],
,[object Object],: allowed, ,[object Object],
,[object Object],: outcome, ,[object Object],
}
,[object Object], LOG.,[object Object],(,[object Object],) ,[object Object], f: ,[object Object],
f.write(json.dumps(record) + ,[object Object],)What this does: Writes one structured JSON record per action to an append-only log, capturing the timestamp, a session ID that ties a run's actions together, the exact tool and arguments, whether it was allowed, and how it turned out. Secrets are redacted before writing. Every field exists to answer a question an auditor will actually ask.
Crucially, this logs at the dispatch layer — around the code that runs the action — not around the model's output, and it logs before and regardless of execution.
[object Object], ,[object Object],(,[object Object],):
allowed = is_allowed(tool, args)
,[object Object], ,[object Object], allowed:
audit(tool, args, outcome=,[object Object],, allowed=,[object Object],) ,[object Object],
,[object Object], ,[object Object],
,[object Object],:
result = execute(tool, args)
audit(tool, args, outcome=,[object Object],, allowed=,[object Object],)
,[object Object], result
,[object Object], Exception ,[object Object], e:
audit(tool, args, outcome=,[object Object],, allowed=,[object Object],)
,[object Object],What this does: Records every path through the dispatcher — allowed-and-ran, blocked, and errored — so the audit log reflects reality completely. A blocked delete is now a logged event, which is exactly the kind of thing a security review needs to see.
Breaking Down Each Element
The structure is what makes the log useful. Because each record is JSON with consistent fields, you can answer questions with a query instead of a grep: every destructive action this week, every blocked request by session, every error from a given tool. Freeform text can't do that; structured records can. The whole point of logging for audit is being able to ask questions later, and only structure enables that.
The session ID is the thread that ties a run together. One action in isolation rarely tells the story; a sequence does. When staging breaks, you pull the session that was active and read its actions in order — read the config, ran the migration, restarted the service — reconstructing exactly what happened as a coherent timeline rather than scattered fragments.
Logging denials is the non-obvious piece that most implementations skip. A record of what the agent was stopped from doing is as valuable as what it did — it shows your guardrails working, reveals what capability users keep reaching for, and in a security review, proves the boundary held. An audit log that only records successes is missing half the evidence.
⚡ Pro tip: Redact secrets before they reach the log, not after. An audit log is durable and often shipped to a central system, so a secret written into it is a secret leaked to everywhere the log goes. Run arguments through your redaction step in the audit function itself, so no code path can write a raw credential to disk.
⚡ Pro tip: Use JSON Lines — one JSON object per line — rather than a single JSON array. Append-only writing is trivial (just add a line), the file stays valid even if a write is interrupted mid-session, and standard tools stream it without loading the whole thing into memory. It's the format built for exactly this job.
How Do You Actually Use the Audit Log?
A log you write but never read is just disk usage. The payoff of cli agent action logging comes when you can interrogate it, and because the records are structured JSON Lines, the queries are simple even with plain command-line tools.
[object Object],
grep ,[object Object], ~/.agent/audit.jsonl | jq ,[object Object],
,[object Object],
jq -r ,[object Object], ~/.agent/audit.jsonl | ,[object Object], | ,[object Object], -cWhat this does: Answers real audit questions directly from the log — what destructive actions actually ran, what the guardrails blocked and how often — using the structure of each record. The same queries feed a dashboard or an alert when you graduate beyond ad-hoc investigation.
Retention is the other half of using the log well. Audit records are cheap to keep and expensive to have missed, so err toward keeping them longer than feels necessary — a compliance question or an incident postmortem often reaches back weeks or months. Rotate the file by size or date so it doesn't grow without bound, but archive rotated logs rather than deleting them, and treat the retention window as a policy decision rather than an accident of when the file got too big.
⚡ Pro tip: Add a monotonic sequence number to each record alongside the timestamp. Wall-clock time can jump backward across daylight-saving changes or clock corrections, which scrambles a time-sorted audit trail; a sequence number gives you an unambiguous order of events no matter what the clock does.
⚡ Pro tip: Log a schema version field in every record. The day you add or rename a field, old and new records coexist in the same file, and a version tag lets your queries and dashboards handle both instead of silently misreading the old format. Audit logs outlive the code that wrote them, so plan for their format to evolve.
Variations for Different Contexts
A compliance-focused fintech ships each audit record to a central, write-once store and includes a hash of the previous record in each new one, so the log is tamper-evident — an altered or deleted entry breaks the chain and is detectable.
A DevOps team correlates the agent's session ID with their existing observability stack, so an agent action shows up on the same timeline as the deploy and the alert it triggered, making incident reconstruction a single query across systems.
A security team running agents across an org centralizes all audit logs and alerts on patterns — a spike in blocked destructive actions, an agent reaching for tools outside its normal set — turning the audit trail from a forensic record into an active monitor.
Each extends the same structured, append-only foundation. The format is fixed; what you do with the stream is yours — a forensic archive for one team, a live security monitor for another, a compliance record for a third, all built on the same handful of well-chosen fields written once per action.
Save and Reuse This
The audit schema — the fields, the redaction rules, the decision to log denials — is a design you get right once and want identical on every agent, because inconsistent audit logs across tools are nearly as useless as no logs at all. Rebuilding the schema from scratch each time risks forgetting the field that turns out to matter most in an incident.
Keeping your audit record format and its accompanying guidance in a library like PromptABCD means every agent you build logs the same way, so a security review or an incident postmortem works identically across your whole fleet of tools. The narration you can throw away; the structured record of what the code actually did is the thing worth keeping forever, because it's the only artifact that can answer the question every incident eventually asks.
Continue Reading
Save the prompts from this post
PromptABCD is a free prompt manager. Paste, organize, and reuse your best AI prompts — no more hunting through chat history.
