How to Build Your Own CLI AI Agent
The core loop to build a CLI AI agent is about 40 lines — the real engineering is everything around it. Here's how to build a terminal agent that reads your world and acts on it safely.
import anthropic
client = anthropic.Anthropic()
TOOLS = [{
"name": "run_shell",
"description": "Run a read-only shell command and return stdout.",
"input_schema": {
"type": "object",
"properties": {"cmd": {"type": "string"}},
"required": ["cmd"],
},
}]Here's a number that surprises most engineers: the core loop of a working CLI AI agent is about 40 lines of code. Not 400. Not a framework. Forty lines. Everything else you add — retries, confirmation prompts, logging — is scaffolding around that tiny loop. So if you've been putting off learning how to build a CLI AI agent because it sounds like a weekend-eating project, the actual agent brain is smaller than most React components you've shipped.
The hard part isn't the model call. It's what happens between model calls: parsing tool requests, running them safely, and feeding results back without blowing your token budget. That's where real engineering lives, and it's what this guide focuses on.
What Is a CLI AI Agent?
A CLI AI agent is a terminal program that takes a plain-English goal, asks a language model what to do, executes the tools the model picks (running a command, reading a file, calling an API), feeds the results back, and repeats until the job is done. The terminal is the interface. The model is the planner. Your code is the muscle that actually touches the system.
The distinction that matters: a chatbot answers you, an agent acts for you. When you build a CLI AI agent, you're wiring a model to real capabilities — the filesystem, shell, network — and taking responsibility for what it does with them.
[object Object], anthropic
client = anthropic.Anthropic()
TOOLS = [{
,[object Object],: ,[object Object],,
,[object Object],: ,[object Object],,
,[object Object],: {
,[object Object],: ,[object Object],,
,[object Object],: {,[object Object],: {,[object Object],: ,[object Object],}},
,[object Object],: [,[object Object],],
},
}]What this does: Declares a single tool the model is allowed to request. The model never runs anything itself — it emits a structured request, and your code decides whether to honor it.
What Does It Take to Build a CLI AI Agent?
Three pieces, in order of difficulty. The model call is easy. The tool dispatch is medium. The loop control is where people underestimate the work.
Here's the whole loop:
[object Object], ,[object Object],(,[object Object],):
messages = [{,[object Object],: ,[object Object],, ,[object Object],: goal}]
,[object Object], step ,[object Object], ,[object Object],(max_steps):
resp = client.messages.create(
model=,[object Object],,
max_tokens=,[object Object],,
tools=TOOLS,
messages=messages,
)
messages.append({,[object Object],: ,[object Object],, ,[object Object],: resp.content})
tool_uses = [b ,[object Object], b ,[object Object], resp.content ,[object Object], b.,[object Object], == ,[object Object],]
,[object Object], ,[object Object], tool_uses:
,[object Object], resp.content[,[object Object],].text ,[object Object],
results = []
,[object Object], tu ,[object Object], tool_uses:
out = dispatch(tu.name, tu.,[object Object],)
results.append({
,[object Object],: ,[object Object],,
,[object Object],: tu.,[object Object],,
,[object Object],: out[:,[object Object],], ,[object Object],
})
messages.append({,[object Object],: ,[object Object],, ,[object Object],: results})
,[object Object], ,[object Object],What this does: Runs the plan-act-observe cycle. Notice max_steps — that guard is not optional. Without it, a confused model can loop forever burning tokens and money.
That out[:4000] truncation is the single most overlooked line. Tool output — a directory listing, a stack trace, an API response — can be enormous. Feed the raw blob back and you'll hit context limits by step three, and every subsequent call costs more. I truncate aggressively and let the model ask for more if it needs it.
⚡ Pro tip: Truncate tool results from the middle, not the end. The start of a stack trace and the final error line both carry signal; the repetitive middle usually doesn't. A head + tail truncation keeps more useful information per token than a plain cutoff.
Why Build a CLI AI Agent Instead of Using a Web Chat?
Because the terminal is where the work already is. A DevOps engineer triaging a failing deploy doesn't want to copy-paste logs into a browser tab — they want to type agent "why is the staging pod crashlooping" and have the tool pull the logs itself. The context is right there.
Three concrete scenarios where a CLI agent beats a web chat:
A data engineer running a nightly ETL wants an agent that can grep yesterday's failure logs, check the row counts in a staging table, and summarize what broke — all without leaving the shell where the pipeline runs.
A security analyst reviewing a repo wants read-only reconnaissance: list dependencies, flag ones with known advisories, summarize the auth flow. A terminal agent with a locked-down toolset does this in one command.
A backend developer on call at 2 a.m. wants to ask "what changed in the last three deploys" and get a git-log summary correlated with the incident timeline — no dashboard hunting.
The common thread: the data lives near the terminal, and moving it to a browser is friction. When you build a CLI AI agent, you eliminate that round trip.
[object Object], ,[object Object],(,[object Object],):
,[object Object], name == ,[object Object],:
,[object Object], subprocess
allowed = (,[object Object],, ,[object Object],, ,[object Object],, ,[object Object],, ,[object Object],, ,[object Object],, ,[object Object],)
,[object Object], ,[object Object], args[,[object Object],].split()[,[object Object],] ,[object Object], allowed:
,[object Object], ,[object Object],
r = subprocess.run(args[,[object Object],], shell=,[object Object],,
capture_output=,[object Object],, text=,[object Object],, timeout=,[object Object],)
,[object Object], r.stdout ,[object Object], r.stderr
,[object Object], ,[object Object],What this does: Executes only allowlisted commands with a timeout. The allowlist is your safety boundary — the model proposes, but this function disposes.
⚡ Pro tip: Start every agent read-only. An allowlist of ls, cat, grep, git status lets you test the loop, the prompting, and the truncation logic with zero risk of a destructive command. Add write capabilities only once the read path is boring.
How Do You Keep the Agent From Going Off the Rails?
Three guardrails, and you want all three. The step limit stops infinite loops. The allowlist stops dangerous actions. And a system prompt that states the agent's scope stops it from wandering into tasks it shouldn't attempt.
SYSTEM = ,[object Object],What this does: Sets the behavioral contract. The model is far better at staying in bounds when the bounds are explicit, and this prompt pairs with the code-level allowlist as defense in depth.
Here's the insight you won't find in most tutorials: the system prompt and the code allowlist must agree, but the code is the real boundary. Prompts can be talked around; an allowlist cannot. I've watched a model cheerfully explain why it "needs" to run rm — and the allowlist simply returned an error, and the agent adapted. Treat the prompt as a hint and the code as the law.
⚡ Pro tip: Log every tool request before you execute it, including ones you reject. When something surprising happens, that pre-execution log is the only record of what the model actually asked for versus what ran.
How Do You Give the Agent Memory Within a Task?
The messages list is the agent's working memory, and it grows on every turn — one entry for the model's response, one for each batch of tool results. That growth is a feature until it isn't. Around the point where a long task has run eight or ten tool calls, you're resending a large transcript on every request, and each call gets slower and more expensive because you pay for the whole history each time.
The naive fix is to trim the oldest messages, but that throws away decisions the model made early and needs later. A better pattern keeps the conclusions and drops the raw data. Once the conversation crosses a token threshold, replace old tool results with one-line summaries of what they contained, while keeping the assistant's reasoning intact.
[object Object], ,[object Object],(,[object Object],):
,[object Object], ,[object Object],(messages) <= keep_last + ,[object Object],:
,[object Object], messages
head, tail = messages[:-keep_last], messages[-keep_last:]
note = {,[object Object],: ,[object Object],, ,[object Object],:
,[object Object],
,[object Object],}
,[object Object], [messages[,[object Object],], note] + tailWhat this does: Collapses old turns into a short summary note while preserving the first user goal and the most recent exchanges. The model keeps the thread of what happened without re-reading every byte of every earlier result.
In practice you'd generate that summary with a quick model call rather than hardcoding it, but the shape is the point: memory isn't "keep everything," it's "keep what the next decision needs." An agent that manages its own context this way can run far longer tasks on the same budget.
⚡ Pro tip: Track your running token count locally and log it per step. The moment you can see context growth as a number, you'll notice which tools return bloated results — and fixing those at the source (a tighter grep, a summarized API response) beats compacting after the fact every time.
Common Mistakes When Building Your First Agent
⚠️ Common mistake: Letting the model run arbitrary shell with no allowlist "just for the prototype." Prototypes leak into production, and a shell-shaped hole is the one you'll regret. Build the allowlist first — it's ten lines and it changes how safe you can be while iterating.
The second trap is feeding entire files back as tool results. A 2,000-line config file eats your context window and teaches the model nothing it needs. Return a summary or a grep, not the whole thing.
The third is skipping the step counter. It feels unnecessary until the night a model gets stuck asking to cat a file that doesn't exist, over and over, until your API bill notices.
And the subtle one: not distinguishing "the model is done" from "the model produced no tool call by accident." Check for a final text block explicitly, rather than assuming an empty tool list means success.
Then there's the multi-tool turn people don't plan for. A single model response can contain several tool_use blocks at once — the model may ask to git log and cat README in the same turn. If your loop grabs only the first block, the model's request goes half-answered and it gets confused on the next turn. The loop above handles this correctly by iterating over every tool_use block and returning a result for each, but it's an easy detail to miss when you first write it, and the symptom — an agent that seems to "forget" what it just asked for — is baffling until you spot the cause.
One more that bites teams later: hardcoding the model name in five places. Put it in a single config value from day one. When you want to test a cheaper model for routine tasks or a stronger one for hard reasoning, you'll change one line instead of hunting through the codebase. The same goes for max_steps and your truncation limit — they're tuning knobs, and tuning knobs belong in config, not scattered as magic numbers through the loop.
Conclusion
The agent loop is small. The engineering is in the edges — truncation, allowlists, step limits, and a system prompt that agrees with your code. Get those four right and you have something genuinely useful: a terminal tool that reads your world and reasons about it, safely.
Once you've got a loop you trust, the next challenge is managing the prompts that drive it — the system prompt, the tool descriptions, the task templates you reuse across agents. Keeping those in scattered text files gets messy fast. A prompt library like PromptABCD lets you version and reuse those building blocks, so the next agent you build starts from prompts you've already tuned instead of a blank file. Build the loop once, then keep the good prompts forever.
Continue Reading
Save the prompts from this post
PromptABCD is a free prompt manager. Paste, organize, and reuse your best AI prompts — no more hunting through chat history.
