Building a CLI Agent With Python
A teardown of a weak Python agent, rebuilt into a CLI agent Python developers can trust: real tool calls, a command allowlist, growing memory, and a stop condition that actually works.
import anthropic
client = anthropic.Anthropic()
def agent(goal):
while True:
resp = client.messages.create(
model="claude-sonnet-5",
max_tokens=1024,
messages=[{"role": "user", "content": goal}],
)
print(resp.content[0].text)
if "done" in resp.content[0].text.lower():
breakShould you reach for LangChain the moment you want to build a CLI agent Python beginners can actually understand? That's the question everyone asks, and the honest answer is no — not at first. A CLI agent Python developers can read top to bottom teaches you more than a framework that hides the loop. This teardown takes a weak first attempt and rebuilds it into something you'd actually trust in a terminal.
We'll start with the version most people write on day one, break down exactly why it falls over, then fix it piece by piece.
Before: The Weak Version
Here's the classic first draft — a loop that calls the model, prints whatever comes back, and calls it an agent.
[object Object], anthropic
client = anthropic.Anthropic()
,[object Object], ,[object Object],(,[object Object],):
,[object Object], ,[object Object],:
resp = client.messages.create(
model=,[object Object],,
max_tokens=,[object Object],,
messages=[{,[object Object],: ,[object Object],, ,[object Object],: goal}],
)
,[object Object],(resp.content[,[object Object],].text)
,[object Object], ,[object Object], ,[object Object], resp.content[,[object Object],].text.lower():
,[object Object],What this does: Loops on model calls and stops when the reply happens to contain the word "done." It has no tools, no memory of prior turns, and a stopping condition made of wet paper.
Why It Fails
Count the bugs. First, it rebuilds messages from scratch every iteration, so the model has amnesia — turn two doesn't know what turn one said. Second, the loop never grows the conversation, so if the model did have tools, their results would vanish. Third, the stop condition is a substring check; the moment the model writes "not done yet," the loop exits early.
But the deepest failure is conceptual: there are no tools, so this isn't an agent at all. It's a chatbot in a while loop. An agent needs to act — read a file, run a command — and observe the result. Without that, you've built an expensive echo.
⚠️ Common mistake: Using a substring like "done" as your stopping condition. Models are chatty and unpredictable with phrasing. Stop on a structural signal — the absence of a tool call — not on matching words in prose.
After: The Improved CLI Agent Python Version
Here's the rebuilt loop. It accumulates messages, declares a tool, dispatches it safely, and stops on a real signal.
[object Object], anthropic, subprocess, shlex
client = anthropic.Anthropic()
TOOLS = [{
,[object Object],: ,[object Object],,
,[object Object],: ,[object Object],,
,[object Object],: {
,[object Object],: ,[object Object],,
,[object Object],: {,[object Object],: {,[object Object],: ,[object Object],}},
,[object Object],: [,[object Object],],
},
}]
ALLOW = {,[object Object],, ,[object Object],, ,[object Object],, ,[object Object],, ,[object Object],, ,[object Object],, ,[object Object],}
,[object Object], ,[object Object],(,[object Object],):
parts = shlex.split(cmd)
,[object Object], ,[object Object], parts ,[object Object], parts[,[object Object],] ,[object Object], ,[object Object], ALLOW:
,[object Object], ,[object Object],
,[object Object],:
r = subprocess.run(parts, capture_output=,[object Object],, text=,[object Object],, timeout=,[object Object],)
,[object Object], (r.stdout ,[object Object], r.stderr)[:,[object Object],]
,[object Object], subprocess.TimeoutExpired:
,[object Object], ,[object Object],What this does: Defines a read-only command tool and an executor that parses safely with shlex, checks an allowlist, and enforces a timeout. shlex.split gives you a real argument list instead of hoping a shell parses it correctly.
[object Object], ,[object Object],(,[object Object],):
messages = [{,[object Object],: ,[object Object],, ,[object Object],: goal}]
,[object Object], _ ,[object Object], ,[object Object],(max_steps):
resp = client.messages.create(
model=,[object Object],, max_tokens=,[object Object],,
tools=TOOLS, messages=messages,
)
messages.append({,[object Object],: ,[object Object],, ,[object Object],: resp.content})
calls = [b ,[object Object], b ,[object Object], resp.content ,[object Object], b.,[object Object], == ,[object Object],]
,[object Object], ,[object Object], calls:
text = ,[object Object],((b.text ,[object Object], b ,[object Object], resp.content ,[object Object], b.,[object Object], == ,[object Object],), ,[object Object],)
,[object Object], text
results = [{
,[object Object],: ,[object Object],,
,[object Object],: c.,[object Object],,
,[object Object],: run_cmd(c.,[object Object],[,[object Object],]),
} ,[object Object], c ,[object Object], calls]
messages.append({,[object Object],: ,[object Object],, ,[object Object],: results})
,[object Object], ,[object Object],What this does: Runs the real plan-act-observe loop. It keeps the growing messages list as memory, stops when the model emits no tool call, and caps iterations so a stuck model can't run forever.
Breaking Down Each Element
The messages list is the agent's memory. Every assistant turn and every tool result appends to it, so the model sees the full history on each call. Reset it and you get amnesia; grow it and you get context.
The allowlist is the security boundary. A CLI agent Python teams deploy internally lives or dies on this — the model can request anything, but ALLOW decides what actually runs. Keep it small and read-only until you've earned trust.
The max_steps cap is your circuit breaker. Twelve is generous for most triage tasks; if the model can't finish in twelve tool calls, it's usually stuck, and you want to fail loudly rather than silently burn tokens.
⚡ Pro tip: Print each tool call before running it, prefixed with something like → read_cmd: git log -5. In a terminal, that running trace is your debugger — you can see the model's plan unfold and spot a bad step the instant it happens.
⚡ Pro tip: Set timeout on every subprocess call. A single hung command — a git log on a huge repo, a grep over a massive tree — will otherwise freeze your whole agent with no way out but Ctrl-C.
Streaming the Agent's Thinking
A terminal agent that goes silent for thirty seconds feels broken even when it's working perfectly. Python's SDK streams tokens as they arrive, so you can print the model's reasoning live while still collecting the full message for tool dispatch.
[object Object], ,[object Object],(,[object Object],):
,[object Object], client.messages.stream(
model=,[object Object],, max_tokens=,[object Object],,
tools=TOOLS, messages=messages,
) ,[object Object], stream:
,[object Object], text ,[object Object], stream.text_stream:
,[object Object],(text, end=,[object Object],, flush=,[object Object],)
,[object Object], stream.get_final_message()What this does: Prints the model's text as it's generated using text_stream, then returns the complete message object so your loop can still extract tool calls. Users see thinking happen in real time instead of staring at a frozen cursor.
The difference is entirely psychological and entirely real. The same forty-second task feels fast when you watch it reason and slow when you don't. For a CLI agent Python users run interactively, streaming isn't a nice-to-have — it's the line between "responsive tool" and "did it crash?"
⚡ Pro tip: Print a newline and a short separator after each streamed response, before the tool output. Without visual breaks, streamed reasoning and command output blur into one wall of text that's miserable to read back.
Testing the Loop Without Burning Tokens
Here's the part almost no tutorial covers: you can — and should — test your agent loop without calling the real model at all. The loop logic (memory growth, tool dispatch, the stop condition) is deterministic. Only the model is not. So you fake the model.
[object Object], ,[object Object],:
,[object Object], ,[object Object],(,[object Object],):
,[object Object],.scripted = ,[object Object],(scripted)
,[object Object], ,[object Object],:
,[object Object],
,[object Object], ,[object Object],():
,[object Object],
fake = FakeResponses([
tool_use_block(,[object Object],, {,[object Object],: ,[object Object],}),
text_block(,[object Object],),
])
result = agent(,[object Object],, client=fake, max_steps=,[object Object],)
,[object Object], ,[object Object], ,[object Object], result
,[object Object], fake.calls == ,[object Object], ,[object Object],What this does: Feeds the loop a scripted sequence of responses instead of a live model, so you can assert on behavior — that it dispatches the tool, appends the result, and stops when no tool call comes back. Fast, free, and deterministic.
This catches the bugs that matter: an off-by-one in the step limit, a tool result appended to the wrong role, a stop condition that never triggers. Those are logic errors, not model errors, and logic errors are exactly what unit tests are for. I run these on every commit; they take milliseconds and cost nothing.
⚡ Pro tip: Write one test per failure mode you've actually hit — infinite loop, empty tool result, malformed tool input. Your test suite then becomes a record of every way the loop has broken before, which is the cheapest insurance a growing agent can have.
Making Tool Errors Recoverable
There's a subtle failure mode the improved loop still has, and it separates a toy from a tool: when a command fails, the model should see the failure and adapt, not crash the process. The trick is to feed errors back as observations rather than raising exceptions.
[object Object], ,[object Object],(,[object Object],):
parts = shlex.split(cmd)
,[object Object], ,[object Object], parts ,[object Object], parts[,[object Object],] ,[object Object], ,[object Object], ALLOW:
,[object Object], {,[object Object],: ,[object Object],,
,[object Object],: ,[object Object],}
,[object Object],:
r = subprocess.run(parts, capture_output=,[object Object],, text=,[object Object],, timeout=,[object Object],)
out = (r.stdout ,[object Object], r.stderr)[:,[object Object],]
,[object Object], {,[object Object],: out, ,[object Object],: r.returncode != ,[object Object],}
,[object Object], Exception ,[object Object], e:
,[object Object], {,[object Object],: ,[object Object],, ,[object Object],: ,[object Object],}What this does: Returns a structured result carrying both the text and an is_error flag, which you pass straight through to the tool_result block. A blocked command, a non-zero exit, or a thrown exception all become normal observations the model reads on its next turn.
Why this matters: when the model asks to cat missing.txt and gets back ERROR: No such file, a well-prompted model will try ls to find the real name and recover on its own. Raise an exception instead and you've turned a two-second self-correction into a stack trace and a dead session. The loop's whole value is that the model can observe the world and adjust — errors are just another observation.
⚠️ Common mistake: Swallowing failures silently by returning an empty string on error. The model then thinks the command produced no output and confidently moves on with a wrong assumption. Always return the actual error text; a model that can read the failure can route around it.
Variations for Different Contexts
The same skeleton adapts across roles with only the toolset changing.
A bioinformatics researcher swaps the shell tool for a pandas-query tool, so "how many samples failed QC" runs a DataFrame filter instead of a command, all inside the terminal where the analysis scripts already live.
A DevOps engineer adds a kubectl read-only tool (get, describe, logs only) so "why is the api pod unhealthy" pulls real cluster state — with writes like delete and apply kept firmly out of the allowlist.
A financial analyst who lives in Python notebooks wants a terminal companion that reads CSVs and summarizes anomalies, using a file-read tool capped at, say, the first 500 rows so a giant export can't swamp the context window.
Each is the identical loop. Only the tools and the system prompt change — which is exactly why understanding the raw loop pays off before any framework.
When Should You Reach for a Framework?
Back to the opening question. You've now seen the whole loop, so here's the honest answer: reach for a framework when you need something the raw loop doesn't give you cheaply, and not a moment before.
The raw loop wins when you're learning, when you want to understand exactly what's sent to the model, and when your agent is small enough that you can hold it in your head. That covers a surprising number of real tools. A CLI agent Python teams run internally often never outgrows the forty-line loop plus a handful of tools, and every abstraction you add is a layer you have to debug through when something breaks at 2 a.m.
Frameworks earn their weight when you need retry policies, structured tracing across many agents, pluggable memory backends, or a team of people who shouldn't each reinvent the loop. At that scale, the framework's conventions save more than they cost. The mistake is starting there — importing a heavy dependency to write your first agent means you're debugging the framework's abstractions before you understand the problem they solve.
My rule of thumb: build the raw loop first, ship it, and let real pain tell you which abstraction you actually need. You'll adopt a framework for a concrete reason instead of a vague sense that serious projects use one. And because you understand the underlying loop, you'll read the framework's docs with real comprehension rather than cargo-culting its examples. The forty lines aren't throwaway scaffolding — they're the mental model that makes everything above them make sense.
Save and Reuse This
Once your loop works, the parts worth saving aren't the plumbing — they're the tool schemas and the system prompt that shapes behavior. Those get tuned over many sessions, and rewriting them from memory each time is how good behavior gets lost.
Keeping your tool definitions and system prompts in a library like PromptABCD means your next CLI agent starts from the schemas you already refined. The forty lines of loop code you can retype in your sleep; the prompt that finally got the model to stop over-explaining is the thing you actually want to reuse.
Continue Reading
Save the prompts from this post
PromptABCD is a free prompt manager. Paste, organize, and reuse your best AI prompts — no more hunting through chat history.
