Handling Long-Running Tasks in a CLI Agent
A twenty-minute agent job you can't watch, interrupt, or resume is a gamble. Handling cli agent long running tasks well means checkpointing, safe interruption, and resumability — here's how a team did it.
from tqdm import tqdm
def run_job(files):
results = []
for f in tqdm(files): # shows progress, but that's all
results.append(process(f))
return resultsPicture this: you're a machine learning engineer, you kick off an agent task that has to process a few hundred files, and twenty minutes in you have no idea whether it's working, stuck, or dead. You don't dare touch it — a stray Ctrl-C might lose everything it's done so far. So you sit and wait, and when your laptop sleeps and the process dies at file 340, you start over from zero. That helpless feeling is what handling cli agent long running tasks well is designed to eliminate. This is the story of a team that turned fragile, all-or-nothing runs into tasks you can watch, interrupt, and resume.
Long-running tasks break the comfortable assumptions of a quick chat loop. The moment work takes minutes instead of seconds, a whole set of problems that didn't exist before become the difference between a usable tool and an infuriating one.
The Problem the Team Faced
An ML tooling team built an agent that ran multi-step data-processing jobs — cleaning files, running validations, generating reports across large batches. Each job took anywhere from five to forty minutes. The agent worked, but using it was miserable in three specific ways. There was no visible progress, so users couldn't tell a slow job from a hung one. There was no safe interruption, so stopping meant killing the process and losing all in-progress work. And there was no resumability, so any crash — a closed laptop, a network blip, an out-of-memory kill — meant starting the entire job over.
The cost showed up as abandoned jobs and wasted compute. Users would kill a job that turned out to be seconds from finishing, because they couldn't tell. Others would lose thirty minutes of work to a crash and simply give up on the big jobs, using the agent only for tasks small enough to survive in one shot. The tool's reach was capped by its fragility, and everyone worked around it rather than trusting it with real work.
The Wrong Approach
The first fix was to add a progress bar. It helped the visibility problem and did nothing for the two that actually mattered.
[object Object], tqdm ,[object Object], tqdm
,[object Object], ,[object Object],(,[object Object],):
results = []
,[object Object], f ,[object Object], tqdm(files): ,[object Object],
results.append(process(f))
,[object Object], resultsWhat this does: Wraps the work in a progress bar so users can see how far along a job is. It's a real improvement to visibility, but the work still lives entirely in memory and vanishes the instant the process stops — a crash at 90% still loses everything, and Ctrl-C is still catastrophic.
The team learned that visibility alone was treating a symptom. Users could now watch their thirty minutes of work disappear when the laptop slept, which was arguably worse. The real problems were that progress existed only in memory and that interruption was unsafe. A progress bar made the failure legible without making it survivable.
⚠️ Common mistake: Treating long-running task support as just a progress indicator. Visibility matters, but the harder and more valuable problems are safe interruption and resumability. A task that shows beautiful progress and then loses all of it on a crash has solved the easy part and skipped the important ones.
The Correct Approach
The rewrite made the task's progress durable. Instead of holding results in memory, the agent checkpoints completed work to disk as it goes, so the state of the job survives anything that stops the process. Resuming means skipping what's already done.
[object Object], json, pathlib
,[object Object], ,[object Object],(,[object Object],):
ckpt = pathlib.Path.home() / ,[object Object], / ,[object Object], / ,[object Object],
ckpt.parent.mkdir(parents=,[object Object],, exist_ok=,[object Object],)
done = json.loads(ckpt.read_text()) ,[object Object], ckpt.exists() ,[object Object], {}
,[object Object], item ,[object Object], items:
,[object Object], item ,[object Object], done: ,[object Object],
,[object Object],
done[item] = process(item)
ckpt.write_text(json.dumps(done)) ,[object Object],
,[object Object], doneWhat this does: Records each completed item to a checkpoint file as it finishes, and on startup skips anything already recorded. A job that dies at item 340 resumes from 340 on the next run instead of restarting — the work done is work kept, regardless of why the process stopped.
The second half was making interruption safe. Rather than letting Ctrl-C kill the process mid-item, the agent catches the signal, finishes the current step, saves, and exits cleanly.
[object Object], signal
,[object Object], ,[object Object],:
stop = ,[object Object],
,[object Object], ,[object Object],(,[object Object],):
signal.signal(signal.SIGINT, ,[object Object],._handle)
,[object Object], ,[object Object],(,[object Object],):
,[object Object],(,[object Object],)
,[object Object],.stop = ,[object Object],
,[object Object], ,[object Object],(,[object Object],):
,[object Object], item ,[object Object], items:
,[object Object], guard.stop: ,[object Object],
,[object Object],
process_and_checkpoint(item)What this does: Intercepts Ctrl-C and sets a flag instead of tearing the process down immediately, so the loop stops at a clean boundary with its checkpoint intact. Interruption becomes a safe, deliberate pause rather than a data-losing crash. The user can stop a job and resume it later with nothing lost.
Results and What Changed
Once jobs were checkpointed and interruptible, the team's usage pattern inverted. The big jobs people had been avoiding became the ones they ran most, because a forty-minute job was no longer a forty-minute gamble — they could stop it for a meeting, resume it after, and survive a crash without losing progress. Wasted compute from restarts effectively disappeared, because a failure cost only the single in-flight item, not the whole run.
The behavioral shift mirrored what the team had seen with undo and confirmation features: reducing the risk of an action is what unlocks people actually using it. When cli agent long running tasks stopped being fragile, users stopped treating them as fragile, and the agent finally got used for the substantial work it had always been capable of but never trusted with.
⚡ Pro tip: Checkpoint at the natural unit of work, not on a timer. Saving after each processed file or completed step means resumption is always clean and at a sensible boundary. Timer-based checkpoints can capture a half-finished item, which is messier to resume from than simply persisting after each discrete piece completes.
⚡ Pro tip: Give every job a stable ID and print it at startup. "Job a1b2c3 started — resume with agent resume a1b2c3" tells the user exactly how to pick up where they left off. A resumable job the user can't find the handle for isn't really resumable, so make the reattach command obvious from the first line.
How Do You Run a Task in the Background?
Checkpointing and safe interruption make a long job survivable; running it in the background makes it convenient. A user who kicks off a forty-minute job usually doesn't want their terminal held hostage for forty minutes — they want to start it, get their prompt back, and check on it later. Detaching cli agent long running tasks into a background job with a reattach handle delivers exactly that.
[object Object], subprocess, sys
,[object Object], ,[object Object],(,[object Object],):
,[object Object],
log = pathlib.Path.home() / ,[object Object], / ,[object Object], / ,[object Object],
subprocess.Popen([sys.executable, ,[object Object],, ,[object Object],, ,[object Object],, goal],
stdout=,[object Object],(log, ,[object Object],), stderr=subprocess.STDOUT,
start_new_session=,[object Object],) ,[object Object],
,[object Object],(,[object Object],)What this does: Launches the job in a detached session that keeps running after the terminal closes, redirecting its output to a log file the user can tail. The user regains their prompt immediately and can check progress or reattach whenever they want, rather than babysitting the run.
The background pattern only works because the checkpointing came first. A detached job that crashes is fine precisely because its progress is on disk — reattaching or restarting picks up from the last checkpoint. Background execution and resumability are complementary: one frees the user's terminal, the other protects the user's work, and together they make a long job feel effortless.
⚡ Pro tip: Write background job output to a log file and give users a logs <job-id> command to tail it. A detached job the user can't observe is unsettling — a simple log-tailing command restores the visibility that detaching took away, so background doesn't mean blind.
⚡ Pro tip: Keep a small registry of active jobs — their IDs, goals, and status — so a user can list what's running. Someone who started three background jobs yesterday needs a way to see them all today; a jobs command that lists them turns detached tasks from fire-and-forget into something manageable.
How to Apply This to Your Situation
The checkpoint-and-resume pattern fits any task too long to trust to a single uninterrupted run.
A data engineer processing millions of records checkpoints by batch offset, so a pipeline that dies at batch 800 resumes there rather than reprocessing the first 799 from scratch.
A security researcher running a long scan across many hosts records results per-host as they complete, so an interrupted scan resumes at the unscanned hosts and never redoes finished work.
A DevOps engineer running a multi-stage migration checkpoints after each stage, so a failure in stage four resumes at stage four with the first three stages' results intact and verified.
The through-line is the same in every case: persist completed work at a natural boundary, make interruption save rather than destroy, and let resumption skip what's already done — so the length of a task stops being a measure of how much you stand to lose.
Next Steps
Add resumability to your longest tasks first: pick a stable unit of work, checkpoint completed units to disk, and skip completed units on restart. Then add a SIGINT handler that finishes the current unit and exits cleanly. Those two changes convert an all-or-nothing gamble into a task users can watch, pause, and recover — which is what makes them willing to run the big jobs at all, and the big jobs are usually the ones worth running.
The checkpoint format, the job-ID scheme, and the resume logic are reusable across every long-running agent you build, along with the prompts that structure how the agent breaks work into checkpointable units. Keeping that machinery in a library like PromptABCD means your next agent handles long jobs durably from day one, and your users stop losing thirty minutes of work to a closed laptop.
Continue Reading
Save the prompts from this post
PromptABCD is a free prompt manager. Paste, organize, and reuse your best AI prompts — no more hunting through chat history.
