Cost Control for CLI Coding Agents
Good cli agent cost control isn't using the agent less — it's removing invisible waste. Here are the levers, biggest first: scope context (re-billed every turn), cap turns, route models, and cache the stable parts.
claude -p "$TASK" --output-format json | jq '{cost: .total_cost_usd, turns: .num_turns, in: .usage.input_tokens, out: .usage.output_tokens}'Why is your agent bill three times what you expected? It's the question everyone asks in month two, after the novelty wears off and the invoice arrives. The honest answer is almost never "the agent used too many words." It's that a few specific, invisible habits — dumping whole repos into context, letting sessions run unbounded, using an expensive model for cheap work — quietly multiply your token spend by five. Good cli agent cost control isn't about using the agent less. It's about removing the waste you can't see, so the same work costs a fraction of what it did.
This guide walks through the actual levers, biggest first, so you can find where your money is going and fix it.
Quick-Start: The One Number to Watch
Before optimizing anything, measure. Every serious agent can report the cost of a run, and you can't control what you don't see.
claude -p ,[object Object], --output-format json | jq ,[object Object],What this does: Reports the exact dollar cost, turn count, and token breakdown for a run. The input-token number is usually the revelation — for coding agents, input tokens (the context you feed the model each turn) typically dwarf output tokens, which tells you immediately where the money actually goes.
That last point is the whole game. People assume the cost is in what the agent writes. It's overwhelmingly in what the agent reads, every turn, over and over.
⚡ Pro tip: Log total_cost_usd for every run for a week before you optimize. The distribution is almost always surprising — a handful of runs account for most of the spend, and they're usually the ones that pulled huge context or looped many turns. Optimize those specific patterns, not your usage in general.
Understanding the Variables
Four levers control agent cost, and they're wildly unequal in impact.
Context size is the biggest by far. Every turn, the agent re-sends its context to the model, so a bloated context isn't a one-time cost — it's a tax on every single turn of the session. Dumping your whole repo into context, or letting a long session accumulate irrelevant history, is the number-one cause of surprise bills.
Turn count multiplies context cost. A session that takes twenty turns pays the context cost twenty times. An agent stuck retrying a failing approach burns turns — and therefore money — with nothing to show for it.
Model choice sets the per-token rate. Frontier models cost multiples of lighter ones per token. Using the most expensive model for mechanical work you could do with a cheaper one is pure waste.
Caching cuts repeated-context cost. If your provider supports prompt caching, the stable parts of your context (system prompt, project files) can be cached so you're not paying full rate to re-send them every turn. This directly attacks the biggest lever.
⚡ Pro tip: The cheapest token is the one you never send. Before reaching for a cheaper model, cut the context you're feeding the agent — scope it to the files the task needs instead of the whole repo. Context discipline saves more than model-switching for most workloads, because it attacks the cost that recurs every turn.
CLI Agent Cost Control, Step by Step
Work the levers in order of impact.
Step one: scope the context. Point the agent at the specific files or subdirectory the task needs, not the repo root. This is the highest-impact change because it reduces the per-turn cost that everything else multiplies.
[object Object],
,[object Object], services/billing && claude ,[object Object],What this does: Narrows the agent's working context to one service, so every turn sends a fraction of the tokens a repo-root session would. On a large monorepo this single change can cut a session's cost by an order of magnitude.
Step two: cap the turns. A --max-turns limit stops a stuck agent from paying the context cost indefinitely.
claude -p ,[object Object], --max-turns 8 ,[object Object],What this does: Bounds the session so a failing approach costs at most eight turns instead of an open-ended spiral. The cap is a cost ceiling as much as a safety one.
Step three: route models to match the work. Use a cheaper model for mechanical tasks and reserve the expensive one for genuine reasoning. Model-agnostic agents make this easy; even single-vendor agents let you switch tiers per session.
Step four: turn on caching if your provider offers it, so the stable context isn't re-billed every turn.
⚠️ Common mistake: Reaching for a cheaper model while still dumping the whole repo into context every turn. You've cut the per-token rate but left the token volume — the bigger lever — untouched, so the bill barely moves. Scope the context first; it's the change that actually pays off, and only then optimize the rate.
The Hidden Cost of Retries and Re-Runs
One cost source hides from every per-run measurement: the runs you throw away. When an agent's first attempt is wrong and you re-run the task, you paid for the failed attempt too. A task you run three times before it works costs roughly three times a task that works first try — and that multiplier is invisible if you only look at the run that finally succeeded.
This reframes where prompt effort pays off. A vague prompt that needs three attempts to get right isn't just slower; it's three times the token cost of a precise prompt that works the first time. The minutes you spend making the instruction complete and specific aren't overhead — they're directly cutting the retry multiplier on the bill.
[object Object],
claude -p ,[object Object], --max-turns 5What this does: Gives a specific, bounded, complete instruction that's likely to succeed on the first pass, avoiding the re-run tax. The precision that makes a prompt safe and correct also makes it cheap, because it collapses the number of expensive attempts needed to get a usable result.
The same logic applies to exploration versus execution. Open-ended "figure out how to do X" sessions are expensive because they wander, re-reading context as they explore. When you already know the approach, telling the agent the approach instead of making it discover one cuts both turns and cost.
⚡ Pro tip: When a task keeps failing and you keep re-running it, stop and fix the prompt instead of re-rolling. Each re-run costs full price for the same likely-to-fail instruction. Investing one minute in a clearer, more constrained prompt is almost always cheaper than the third and fourth attempts at the vague one.
Pro-Level Variations
A startup watching burn rate sets a hard per-run cost ceiling in every automated job, parsing total_cost_usd and failing the job if a single run exceeds the limit. Runaway spend is capped structurally, not caught after the fact on the invoice.
A platform team running many agents routes a cheap model for the mechanical stages (formatting, simple edits, extraction) and an expensive one only for planning and hard reasoning, cutting aggregate spend without touching output quality on the parts that matter.
A solo developer scopes every session tightly and starts fresh for each new task rather than continuing one long session, so context never balloons with accumulated irrelevant history. Short, scoped sessions are cheap sessions.
⚡ Pro tip: Beware the sub-agent cost multiplier. Spawning parallel sub-agents is powerful, but each one carries its own context and burns its own tokens, so a task fanned out to five sub-agents can cost several times a single-agent run. Sub-agents are worth it for genuine parallelism on big tasks — just know you're trading money for speed, and don't fan out work that didn't need it.
Troubleshooting Common Issues
If your bill is high and you don't know why, sort your logged runs by cost and look at the top few. They'll almost always share a cause — huge context or a long turn count — and fixing that one pattern fixes most of the spend. Don't optimize the average run; optimize the expensive tail.
If costs are creeping up over a session rather than spiking on one run, it's context accumulation. Long-running sessions drag an ever-growing history that's re-billed every turn. Start fresh sessions for new tasks, and use whatever context-compaction your agent offers.
⚡ Pro tip: Set a monthly spend alert at your provider, not just per-run caps. Per-run caps catch a single runaway; a monthly alert catches the slow drift of many slightly-too-expensive runs that no single cap would flag. You want both the circuit breaker and the trend line.
If a single task is surprisingly expensive no matter what you do, check whether it's genuinely a big task or just an under-scoped one. Asking an agent to "improve the codebase" pulls enormous context and wanders for many turns because the goal is unbounded. A bounded, specific version of the same intent costs a fraction, because the agent isn't re-reading the world to figure out what you meant.
Your Turn
Effective cli agent cost control works the levers in order of impact: scope the context first because it's re-billed every turn, cap the turns so nothing spirals, route models to match the work, and cache the stable context. The surprise bill almost never comes from the agent being too wordy — it comes from feeding it too much, too many times, on too expensive a model. Measure first, then fix the expensive tail.
The context-scoping habits, the cost-ceiling scripts, the model-routing rules — these apply to every agent you'll run. Save them in PromptABCD so your next project starts cost-controlled by default instead of teaching you where the money went through a month-two invoice.
Continue Reading
Save the prompts from this post
PromptABCD is a free prompt manager. Paste, organize, and reuse your best AI prompts — no more hunting through chat history.
