PromptABCD
FeaturesLearnGuideBlogContext Blocks
Sign inGet started free
Sign inSign up
PromptABCD

A calm home for your best AI prompts. Save them once, find them in seconds, reuse them forever.

Product

  • Features
  • Chrome Extension
  • Free Courses
  • How it works
  • Use cases
  • Blog
  • Context Blocks
  • Export Anywhere
  • FAQ

Resources

  • User guide
  • Learn prompting
  • Sign in
  • Get started free

© 2026 PromptABCD. All rights reserved.

AboutPrivacy PolicyTerms and Conditions
Home/Blog/CLI AI Agents/Building a CLI Agent With Python
CLI AI Agents

Building a CLI Agent With Python

A teardown of a weak Python agent, rebuilt into a CLI agent Python developers can trust: real tool calls, a command allowlist, growing memory, and a stop condition that actually works.

September 16, 2026·10 min read
ShareShare
⚡Featured Prompt— copy and use right now
import anthropic
client = anthropic.Anthropic()

def agent(goal):
    while True:
        resp = client.messages.create(
            model="claude-sonnet-5",
            max_tokens=1024,
            messages=[{"role": "user", "content": goal}],
        )
        print(resp.content[0].text)
        if "done" in resp.content[0].text.lower():
            break

Should you reach for LangChain the moment you want to build a CLI agent Python beginners can actually understand? That's the question everyone asks, and the honest answer is no — not at first. A CLI agent Python developers can read top to bottom teaches you more than a framework that hides the loop. This teardown takes a weak first attempt and rebuilds it into something you'd actually trust in a terminal.

We'll start with the version most people write on day one, break down exactly why it falls over, then fix it piece by piece.

Before: The Weak Version

Here's the classic first draft — a loop that calls the model, prints whatever comes back, and calls it an agent.

python
[object Object], anthropic
client = anthropic.Anthropic()

,[object Object], ,[object Object],(,[object Object],):
    ,[object Object], ,[object Object],:
        resp = client.messages.create(
            model=,[object Object],,
            max_tokens=,[object Object],,
            messages=[{,[object Object],: ,[object Object],, ,[object Object],: goal}],
        )
        ,[object Object],(resp.content[,[object Object],].text)
        ,[object Object], ,[object Object], ,[object Object], resp.content[,[object Object],].text.lower():
            ,[object Object],

What this does: Loops on model calls and stops when the reply happens to contain the word "done." It has no tools, no memory of prior turns, and a stopping condition made of wet paper.

Why It Fails

Count the bugs. First, it rebuilds messages from scratch every iteration, so the model has amnesia — turn two doesn't know what turn one said. Second, the loop never grows the conversation, so if the model did have tools, their results would vanish. Third, the stop condition is a substring check; the moment the model writes "not done yet," the loop exits early.

But the deepest failure is conceptual: there are no tools, so this isn't an agent at all. It's a chatbot in a while loop. An agent needs to act — read a file, run a command — and observe the result. Without that, you've built an expensive echo.

⚠️ Common mistake: Using a substring like "done" as your stopping condition. Models are chatty and unpredictable with phrasing. Stop on a structural signal — the absence of a tool call — not on matching words in prose.

After: The Improved CLI Agent Python Version

Here's the rebuilt loop. It accumulates messages, declares a tool, dispatches it safely, and stops on a real signal.

python
[object Object], anthropic, subprocess, shlex

client = anthropic.Anthropic()

TOOLS = [{
    ,[object Object],: ,[object Object],,
    ,[object Object],: ,[object Object],,
    ,[object Object],: {
        ,[object Object],: ,[object Object],,
        ,[object Object],: {,[object Object],: {,[object Object],: ,[object Object],}},
        ,[object Object],: [,[object Object],],
    },
}]

ALLOW = {,[object Object],, ,[object Object],, ,[object Object],, ,[object Object],, ,[object Object],, ,[object Object],, ,[object Object],}

,[object Object], ,[object Object],(,[object Object],):
    parts = shlex.split(cmd)
    ,[object Object], ,[object Object], parts ,[object Object], parts[,[object Object],] ,[object Object], ,[object Object], ALLOW:
        ,[object Object], ,[object Object],
    ,[object Object],:
        r = subprocess.run(parts, capture_output=,[object Object],, text=,[object Object],, timeout=,[object Object],)
        ,[object Object], (r.stdout ,[object Object], r.stderr)[:,[object Object],]
    ,[object Object], subprocess.TimeoutExpired:
        ,[object Object], ,[object Object],

What this does: Defines a read-only command tool and an executor that parses safely with shlex, checks an allowlist, and enforces a timeout. shlex.split gives you a real argument list instead of hoping a shell parses it correctly.

python
[object Object], ,[object Object],(,[object Object],):
    messages = [{,[object Object],: ,[object Object],, ,[object Object],: goal}]
    ,[object Object], _ ,[object Object], ,[object Object],(max_steps):
        resp = client.messages.create(
            model=,[object Object],, max_tokens=,[object Object],,
            tools=TOOLS, messages=messages,
        )
        messages.append({,[object Object],: ,[object Object],, ,[object Object],: resp.content})
        calls = [b ,[object Object], b ,[object Object], resp.content ,[object Object], b.,[object Object], == ,[object Object],]
        ,[object Object], ,[object Object], calls:
            text = ,[object Object],((b.text ,[object Object], b ,[object Object], resp.content ,[object Object], b.,[object Object], == ,[object Object],), ,[object Object],)
            ,[object Object], text
        results = [{
            ,[object Object],: ,[object Object],,
            ,[object Object],: c.,[object Object],,
            ,[object Object],: run_cmd(c.,[object Object],[,[object Object],]),
        } ,[object Object], c ,[object Object], calls]
        messages.append({,[object Object],: ,[object Object],, ,[object Object],: results})
    ,[object Object], ,[object Object],

What this does: Runs the real plan-act-observe loop. It keeps the growing messages list as memory, stops when the model emits no tool call, and caps iterations so a stuck model can't run forever.

Breaking Down Each Element

The messages list is the agent's memory. Every assistant turn and every tool result appends to it, so the model sees the full history on each call. Reset it and you get amnesia; grow it and you get context.

The allowlist is the security boundary. A CLI agent Python teams deploy internally lives or dies on this — the model can request anything, but ALLOW decides what actually runs. Keep it small and read-only until you've earned trust.

The max_steps cap is your circuit breaker. Twelve is generous for most triage tasks; if the model can't finish in twelve tool calls, it's usually stuck, and you want to fail loudly rather than silently burn tokens.

⚡ Pro tip: Print each tool call before running it, prefixed with something like → read_cmd: git log -5. In a terminal, that running trace is your debugger — you can see the model's plan unfold and spot a bad step the instant it happens.

⚡ Pro tip: Set timeout on every subprocess call. A single hung command — a git log on a huge repo, a grep over a massive tree — will otherwise freeze your whole agent with no way out but Ctrl-C.

Streaming the Agent's Thinking

A terminal agent that goes silent for thirty seconds feels broken even when it's working perfectly. Python's SDK streams tokens as they arrive, so you can print the model's reasoning live while still collecting the full message for tool dispatch.

python
[object Object], ,[object Object],(,[object Object],):
    ,[object Object], client.messages.stream(
        model=,[object Object],, max_tokens=,[object Object],,
        tools=TOOLS, messages=messages,
    ) ,[object Object], stream:
        ,[object Object], text ,[object Object], stream.text_stream:
            ,[object Object],(text, end=,[object Object],, flush=,[object Object],)
        ,[object Object], stream.get_final_message()

What this does: Prints the model's text as it's generated using text_stream, then returns the complete message object so your loop can still extract tool calls. Users see thinking happen in real time instead of staring at a frozen cursor.

The difference is entirely psychological and entirely real. The same forty-second task feels fast when you watch it reason and slow when you don't. For a CLI agent Python users run interactively, streaming isn't a nice-to-have — it's the line between "responsive tool" and "did it crash?"

⚡ Pro tip: Print a newline and a short separator after each streamed response, before the tool output. Without visual breaks, streamed reasoning and command output blur into one wall of text that's miserable to read back.

Testing the Loop Without Burning Tokens

Here's the part almost no tutorial covers: you can — and should — test your agent loop without calling the real model at all. The loop logic (memory growth, tool dispatch, the stop condition) is deterministic. Only the model is not. So you fake the model.

python
[object Object], ,[object Object],:
    ,[object Object], ,[object Object],(,[object Object],):
        ,[object Object],.scripted = ,[object Object],(scripted)
    ,[object Object], ,[object Object],:
        ,[object Object],

,[object Object], ,[object Object],():
    ,[object Object],
    fake = FakeResponses([
        tool_use_block(,[object Object],, {,[object Object],: ,[object Object],}),
        text_block(,[object Object],),
    ])
    result = agent(,[object Object],, client=fake, max_steps=,[object Object],)
    ,[object Object], ,[object Object], ,[object Object], result
    ,[object Object], fake.calls == ,[object Object],   ,[object Object],

What this does: Feeds the loop a scripted sequence of responses instead of a live model, so you can assert on behavior — that it dispatches the tool, appends the result, and stops when no tool call comes back. Fast, free, and deterministic.

This catches the bugs that matter: an off-by-one in the step limit, a tool result appended to the wrong role, a stop condition that never triggers. Those are logic errors, not model errors, and logic errors are exactly what unit tests are for. I run these on every commit; they take milliseconds and cost nothing.

⚡ Pro tip: Write one test per failure mode you've actually hit — infinite loop, empty tool result, malformed tool input. Your test suite then becomes a record of every way the loop has broken before, which is the cheapest insurance a growing agent can have.

Making Tool Errors Recoverable

There's a subtle failure mode the improved loop still has, and it separates a toy from a tool: when a command fails, the model should see the failure and adapt, not crash the process. The trick is to feed errors back as observations rather than raising exceptions.

python
[object Object], ,[object Object],(,[object Object],):
    parts = shlex.split(cmd)
    ,[object Object], ,[object Object], parts ,[object Object], parts[,[object Object],] ,[object Object], ,[object Object], ALLOW:
        ,[object Object], {,[object Object],: ,[object Object],,
                ,[object Object],: ,[object Object],}
    ,[object Object],:
        r = subprocess.run(parts, capture_output=,[object Object],, text=,[object Object],, timeout=,[object Object],)
        out = (r.stdout ,[object Object], r.stderr)[:,[object Object],]
        ,[object Object], {,[object Object],: out, ,[object Object],: r.returncode != ,[object Object],}
    ,[object Object], Exception ,[object Object], e:
        ,[object Object], {,[object Object],: ,[object Object],, ,[object Object],: ,[object Object],}

What this does: Returns a structured result carrying both the text and an is_error flag, which you pass straight through to the tool_result block. A blocked command, a non-zero exit, or a thrown exception all become normal observations the model reads on its next turn.

Why this matters: when the model asks to cat missing.txt and gets back ERROR: No such file, a well-prompted model will try ls to find the real name and recover on its own. Raise an exception instead and you've turned a two-second self-correction into a stack trace and a dead session. The loop's whole value is that the model can observe the world and adjust — errors are just another observation.

⚠️ Common mistake: Swallowing failures silently by returning an empty string on error. The model then thinks the command produced no output and confidently moves on with a wrong assumption. Always return the actual error text; a model that can read the failure can route around it.

Variations for Different Contexts

The same skeleton adapts across roles with only the toolset changing.

A bioinformatics researcher swaps the shell tool for a pandas-query tool, so "how many samples failed QC" runs a DataFrame filter instead of a command, all inside the terminal where the analysis scripts already live.

A DevOps engineer adds a kubectl read-only tool (get, describe, logs only) so "why is the api pod unhealthy" pulls real cluster state — with writes like delete and apply kept firmly out of the allowlist.

A financial analyst who lives in Python notebooks wants a terminal companion that reads CSVs and summarizes anomalies, using a file-read tool capped at, say, the first 500 rows so a giant export can't swamp the context window.

Each is the identical loop. Only the tools and the system prompt change — which is exactly why understanding the raw loop pays off before any framework.

When Should You Reach for a Framework?

Back to the opening question. You've now seen the whole loop, so here's the honest answer: reach for a framework when you need something the raw loop doesn't give you cheaply, and not a moment before.

The raw loop wins when you're learning, when you want to understand exactly what's sent to the model, and when your agent is small enough that you can hold it in your head. That covers a surprising number of real tools. A CLI agent Python teams run internally often never outgrows the forty-line loop plus a handful of tools, and every abstraction you add is a layer you have to debug through when something breaks at 2 a.m.

Frameworks earn their weight when you need retry policies, structured tracing across many agents, pluggable memory backends, or a team of people who shouldn't each reinvent the loop. At that scale, the framework's conventions save more than they cost. The mistake is starting there — importing a heavy dependency to write your first agent means you're debugging the framework's abstractions before you understand the problem they solve.

My rule of thumb: build the raw loop first, ship it, and let real pain tell you which abstraction you actually need. You'll adopt a framework for a concrete reason instead of a vague sense that serious projects use one. And because you understand the underlying loop, you'll read the framework's docs with real comprehension rather than cargo-culting its examples. The forty lines aren't throwaway scaffolding — they're the mental model that makes everything above them make sense.

Save and Reuse This

Once your loop works, the parts worth saving aren't the plumbing — they're the tool schemas and the system prompt that shapes behavior. Those get tuned over many sessions, and rewriting them from memory each time is how good behavior gets lost.

Keeping your tool definitions and system prompts in a library like PromptABCD means your next CLI agent starts from the schemas you already refined. The forty lines of loop code you can retype in your sleep; the prompt that finally got the model to stop over-explaining is the thing you actually want to reuse.

cli agentpythonai agenttool useterminalagent loop

Continue Reading

Managing Reusable Prompts for Terminal Workflows
CLI AI Agents

Managing Reusable Prompts for Terminal Workflows

Retyping your best prompt from memory loses its refinements every time. Managing cli agent reusable prompts as named, parameterized, versioned assets keeps the prompt quality you earned — and lets you share it.

September 19, 2026·9 min read
Distributing System Prompts With Your CLI Tool
CLI AI Agents

Distributing System Prompts With Your CLI Tool

Hardcoding your agent's system prompt as a string is the wrong place for it. Treating cli agent system prompt distribution as content — versioned, overridable, updatable — is how prompts evolve independently of code.

September 19, 2026·9 min read
Building a Plugin System for Your CLI Agent
CLI AI Agents

Building a Plugin System for Your CLI Agent

How do you let people add tools to your agent without forking it? A cli agent plugin system lets users extend the agent with their own tools. Here's how to rebuild a hardcoded tool list into a real plugin system.

September 19, 2026·8 min read

Save the prompts from this post

PromptABCD is a free prompt manager. Paste, organize, and reuse your best AI prompts — no more hunting through chat history.

Start free →
← PreviousBuilding a CLI Agent With Node.jsNext →Parsing Natural Language Into CLI Commands
Share this post:
ShareShare