Undo and Rollback in a CLI Agent
What happens when your agent makes a change you didn't want? Good cli agent undo changes support turns a scramble into one command — here's why in-memory undo fails and git-backed checkpoints work.
_undo_stack = []
def edit_file(path, new_content):
old = read(path) if exists(path) else None
_undo_stack.append((path, old)) # remember prior content in memory
write(path, new_content)
def undo():
path, old = _undo_stack.pop()
write(path, old) if old else delete(path)What happens when your agent makes a change you didn't want? Not a catastrophic one it should have blocked — just a wrong one. It edited the right file with the wrong logic, reformatted something you liked, refactored a function into three. Without a way to reverse it, your only recourse is to notice, remember what it touched, and fix each change by hand. Good cli agent undo changes support turns that scramble into a single command. This is the story of a team that added undo the hard way, after learning why the obvious approach doesn't work.
Undo sounds trivial until you build it. The naive version breaks in ways that teach you what a real rollback actually requires.
The Problem the Team Faced
A six-person product team gave their engineers a coding agent that could edit files across a repo. It was genuinely useful and genuinely nerve-wracking, because a single request might touch eight files, and if two of those edits were wrong, untangling them meant reading every diff and reverting by hand. Engineers started copying files to .bak before every agent run — a manual, error-prone ritual that told the team something was missing.
The fear had a cost beyond the annoyance. People used the agent for small, safe edits and did the big, valuable refactors by hand, precisely because the big changes were the ones too risky to reverse. The tool's usefulness was capped by the absence of an undo, and everyone knew it. They needed a way to make any agent change reversible, so that trying something bold carried no more risk than trying something small.
The Wrong Approach
The first attempt stored the previous content of each file in memory before editing, so undo could restore it. It worked in a demo and failed the moment real work got complicated.
_undo_stack = []
,[object Object], ,[object Object],(,[object Object],):
old = read(path) ,[object Object], exists(path) ,[object Object], ,[object Object],
_undo_stack.append((path, old)) ,[object Object],
write(path, new_content)
,[object Object], ,[object Object],():
path, old = _undo_stack.pop()
write(path, old) ,[object Object], old ,[object Object], delete(path)What this does: Pushes each file's prior content onto an in-memory stack before overwriting, and restores it on undo. The flaw isn't the idea — it's that the memory dies with the process, so quitting the agent erases all undo history, and it only tracks file writes, missing every other kind of change.
The team hit the limits fast. Undo vanished when they closed the terminal, so the reversal they wanted most — "undo what it did in yesterday's session" — was impossible. It didn't cover shell commands that moved or deleted files outside the edit_file path. And with multiple changes across many files, popping one entry at a time gave no way to say "undo that whole request as a unit." An in-memory stack of file contents was too narrow and too fragile.
⚠️ Common mistake: Building undo as an in-memory stack of file contents. It doesn't survive a restart, it only captures the one mutation path you thought of, and it can't group a multi-file change into a single reversible unit. Durable, complete rollback needs a checkpoint of state, not a list of remembered strings.
The Correct Approach
The rewrite used the tool already built for exactly this problem: version control. Before executing a batch of changes, the agent creates a checkpoint commit on a scratch branch; undo becomes a reset to that checkpoint. Git handles the durability, the completeness, and the grouping for free.
[object Object], subprocess
,[object Object], ,[object Object],(,[object Object],):
subprocess.run([,[object Object],, ,[object Object],, ,[object Object],], check=,[object Object],)
subprocess.run([,[object Object],, ,[object Object],, ,[object Object],, ,[object Object],,
,[object Object],, ,[object Object],], check=,[object Object],)
,[object Object], subprocess.run([,[object Object],, ,[object Object],, ,[object Object],],
capture_output=,[object Object],, text=,[object Object],).stdout.strip()
,[object Object], ,[object Object],(,[object Object],):
subprocess.run([,[object Object],, ,[object Object],, ,[object Object],, commit], check=,[object Object],)What this does: Snapshots the entire working tree as a commit before the agent acts, capturing every file at once, and rolls back by resetting to that commit. Because it's a real commit, the checkpoint survives restarts, covers any change to tracked files regardless of which tool made it, and reverts a whole batch as a single unit.
The agent wraps each user request in a checkpoint, so every request becomes an atomic, reversible transaction.
[object Object], ,[object Object],(,[object Object],):
cp = checkpoint(goal[:,[object Object],]) ,[object Object],
,[object Object],:
result = agent_loop(goal)
,[object Object],(,[object Object],)
,[object Object], result
,[object Object], Exception:
rollback_to(cp) ,[object Object],
,[object Object],What this does: Creates a checkpoint before each request, reports the undo handle when it finishes, and automatically rolls back if the request errors out partway. A half-completed change that crashed never gets left behind — the working tree returns to its last known-good state.
Results and What Changed
With git-backed checkpoints, the team's behavior changed within a week. The .bak ritual disappeared. Engineers started using the agent for the big refactors they'd been avoiding, because "undo the whole thing" was now one command and they trusted it. The willingness to attempt bold changes — the thing that made the agent actually valuable — came directly from knowing every change was reversible.
The auto-rollback on failure removed a subtler anxiety too. Before, a request that died halfway left a mess of partial edits that were worse than no change at all. Now a failed request cleaned up after itself, returning the tree to exactly where it started. The agent went from "powerful but scary" to "powerful and safe to experiment with," and experimentation is where its value had been hiding.
⚡ Pro tip: Use a dedicated scratch branch or git stash entries for agent checkpoints so they don't clutter the real history. The user's actual commits stay clean; the agent's checkpoints live in a parallel track they can prune later. Undo infrastructure shouldn't pollute the project's story.
⚡ Pro tip: Label each checkpoint with the request that created it. "agent-checkpoint: refactor auth to use tokens" is a legible undo history — a user scanning their checkpoints sees what each one was, so choosing how far back to roll is obvious instead of a guess between anonymous hashes.
What About Actions That Can't Be Undone?
Git checkpoints solve undo for anything on disk, but the hardest part of cli agent undo changes support is being honest about what can't be reversed. Some actions leave the local world entirely — an email sent, a payment charged, a deploy triggered, a row deleted from a production database. No checkpoint brings those back, and pretending otherwise is worse than admitting it.
The right design classifies each action by reversibility and treats the two classes differently. Reversible actions get a checkpoint and a cheap undo. Irreversible ones get a confirmation before they run, because "are you sure" beforehand is the only protection you have when "undo" afterward isn't an option.
IRREVERSIBLE = {,[object Object],, ,[object Object],, ,[object Object],, ,[object Object],}
,[object Object], ,[object Object],(,[object Object],):
,[object Object], tool ,[object Object], IRREVERSIBLE:
,[object Object], confirm(,[object Object],) ,[object Object],
,[object Object], ,[object Object], ,[object Object],What this does: Routes irreversible actions through a mandatory pre-execution confirmation while letting reversible ones rely on the checkpoint. The system's safety net shifts from "undo later" to "confirm first" exactly where undo can't reach, so no action is left unprotected.
This split reframes undo as one half of a pair. Reversibility is a spectrum: checkpoint the actions you can take back, gate the ones you can't, and never let an irreversible action hide behind an undo promise you can't keep.
⚡ Pro tip: Label every tool with its reversibility in the tool definition itself, right next to its schema. Then undo behavior and confirmation gating derive automatically from that one attribute, instead of being maintained in a separate list that drifts out of sync with your actual toolset.
⚡ Pro tip: When an irreversible action runs, still record what it did in a journal — not to undo it, but to know it happened. "Sent invoice email to 40 customers at 14:03" is not reversible, but it's exactly what you'll want to see when someone asks what the agent did, and it turns an untraceable side effect into an accountable one.
How to Apply This to Your Situation
The checkpoint pattern adapts to whatever kind of state your agent changes.
A backend developer whose agent edits code uses git checkpoints directly, exactly as above, getting atomic per-request undo across the whole repo for the cost of a commit.
A data engineer whose agent modifies datasets can't use git for gigabytes of data, so they snapshot differently — a copy-on-write filesystem snapshot or a versioned object-store prefix — but the pattern is identical: checkpoint before, roll back to.
A DevOps engineer whose agent changes infrastructure records the inverse operation for each action instead of a state snapshot, because you can't "reset" a cloud resource — creating one is undone by deleting it, so the agent journals a reverse-operation for every forward one.
That last case reveals the deeper principle: some actions have a state you can snapshot, and some only have an inverse you can record. Know which kind each of your agent's actions is, because it decides whether undo is a checkpoint or a journal — and whether "undo" is even the right promise to make, or whether the honest answer for that action is a confirmation before it ever runs.
Next Steps
Wrap each request in a checkpoint, report the undo handle, and auto-roll-back on failure. If your agent edits files in a git repo, you can have this working in an afternoon. For actions git can't capture — network calls, external side effects — start a journal of inverse operations and be honest in your UI about which actions are reversible and which aren't.
The checkpoint-and-rollback logic, plus the rules for which actions are undoable versus which need an inverse recorded, are reusable across every agent you build. Keeping that pattern and its accompanying prompts in a library like PromptABCD means the next agent gets reliable cli agent undo changes support from day one, and your users get to be bold with it, instead of copying files to .bak and hoping.
Continue Reading
Save the prompts from this post
PromptABCD is a free prompt manager. Paste, organize, and reuse your best AI prompts — no more hunting through chat history.
