Cost Runaway: The Autonomous Agent's Biggest Risk
Ever gotten a bill for an agent that ran overnight and did nothing useful? Autonomous agent cost runaway is the most common expensive surprise in agent work. Here's how it happens and how to stop it.
def run_agent(goal, max_steps=25, max_spend=5.00, max_context_tokens=32000):
spend = 0.0
for step in range(max_steps): # 1. step cap
ctx = build_context(goal, history)
ctx = truncate(ctx, max_context_tokens) # 2. context cap
result, cost = agent.step(ctx)
spend += cost
if spend >= max_spend: # 3. spend cap
return finalize(history, reason="budget reached")
history.append(result)
return finalize(history, reason="step cap")Ever opened a billing dashboard to find an agent ran all night, made a few thousand model calls, and accomplished nothing you'd pay a dollar for? Autonomous agent cost runaway is the most common expensive surprise in agent work, and almost everyone who ships agents hits it at least once. The frustrating part is that the biggest cost driver isn't the one people watch for - and once you see it, the fixes are straightforward.
Autonomous agent cost runaway is when an agent's spending balloons far beyond what its task justified, usually silently, usually overnight. It happens because agents loop, retry, and accumulate context, and every one of those multiplies calls. This guide breaks down exactly where the money goes - including the sneaky driver most people miss - and the specific caps that stop it.
Quick-start: the caps that stop a runaway right now
Drop these three limits into any agent before you run it unattended:
[object Object], ,[object Object],(,[object Object],):
spend = ,[object Object],
,[object Object], step ,[object Object], ,[object Object],(max_steps): ,[object Object],
ctx = build_context(goal, history)
ctx = truncate(ctx, max_context_tokens) ,[object Object],
result, cost = agent.step(ctx)
spend += cost
,[object Object], spend >= max_spend: ,[object Object],
,[object Object], finalize(history, reason=,[object Object],)
history.append(result)
,[object Object], finalize(history, reason=,[object Object],)What this does: it bounds an agent on three independent axes at once - how many steps it takes, how much it spends, and how large its context can grow - so no single runaway path can produce a surprise bill.
The third cap, on context size, is the one most people leave out - and it's the one that addresses the cost driver almost nobody watches. More on that next.
Understanding the variables: where the money actually goes
Four things drive agent cost, and they're not equally obvious.
Loop count is the obvious one. An agent that runs 400 steps costs roughly 16 times one that runs 25. Step caps handle this, and most people add them.
Retries multiply quietly. An agent that retries failed actions three times is making up to three times the calls on every failure. A run with lots of failures can cost several times a clean run for the same task.
Context growth is the sneaky one, and usually the biggest. Here's the trap: most agents re-send their entire history on every step. So step 1 sends a little context, but step 50 sends everything from steps 1 through 49. Cost per step grows as the run goes on, which means a long run is superlinearly expensive - the later steps each cost far more than the early ones. A 50-step agent doesn't cost 50 times a 1-step agent; it can cost far more, because the average step is carrying a huge accumulated context.
Parallel branching multiplies fastest of all. An agent that spawns sub-agents, each spawning more, can fan out into an enormous number of calls from one goal if nothing caps the branching.
Branching deserves a specific warning because its cost is the least intuitive. An agent that spawns three sub-agents, each of which spawns three more, is at nine agents after two levels and twenty-seven after three - the cost grows geometrically with depth. A branching bug that looks like "the agent spawned a few helpers" can be thousands of calls two levels down. Unlike loops and context growth, which climb predictably, branching can explode from a single unlucky decision, which is why depth caps on sub-agent spawning aren't optional for any agent that can delegate to itself.
⚡ Pro tip: Watch context growth, not just step count. The step cap limits how many steps run, but context cap limits how expensive each step becomes - and on a long run, accumulated context is usually the largest single cost, quietly turning a 50-step run into something that costs like a 200-step one.
Step-by-step: stopping autonomous agent cost runaway
Step 1 - Cap steps and spend, both. A step cap alone fails when steps are expensive; a spend cap alone fails when cheap steps loop forever. You need both, because either can bind first depending on the run.
Step 2 - Cap and compact context. Don't let history grow unbounded. Truncate to a window, or summarize old history into a compact digest, so cost per step stops climbing. This attacks the superlinear driver directly.
Step 3 - Bound retries and branching. Limit retries per action and depth of sub-agent spawning. Uncapped branching is the fastest path to a genuinely shocking bill.
Step 4 - Estimate before you run. For expensive tasks, do a quick pre-flight cost estimate - expected steps times expected cost per step - and refuse to start if it exceeds a threshold. Catching a runaway before it starts beats catching it after.
[object Object], ,[object Object],(,[object Object],):
est_steps = estimate_steps(goal)
est_cost = est_steps * avg_cost_per_step(goal)
,[object Object], est_cost > HARD_LIMIT:
,[object Object], CostError(,[object Object],)What this does: it estimates a task's cost before the agent runs and refuses to start tasks whose projected spend exceeds a hard limit - stopping the most expensive runaways before they consume anything.
⚠️ Common mistake: Setting a step cap and assuming you're protected from cost runaway. Step caps don't limit cost per step, so an agent with a generous step cap and unbounded context can still run up a large bill inside the step limit - the later steps are just quietly expensive.
What does autonomous agent cost runaway actually cost per step?
It helps to see the numbers, because the shape surprises people. Imagine a naive agent that re-sends its full history each step, adding roughly 500 tokens of new context per step. Step 1 carries ~500 tokens. Step 20 carries ~10,000. Step 50 carries ~25,000. The cost of the run isn't 50 times the first step - it's the sum of a growing series, which grows with the square of the step count, not linearly.
That quadratic shape is the mechanism behind most autonomous agent cost runaway that isn't a simple infinite loop. A run that looks like it "only ran 50 steps" can cost like a run of 200 flat-cost steps, because the later steps are each carrying an enormous accumulated context. People budget for the step count and get blindsided by the per-step growth.
⚡ Pro tip: Model your agent's cost as the sum of a growing series, not step count times a flat rate. If context accumulates, cost grows roughly quadratically in steps - so doubling the step cap can quadruple the bill, which is why the step cap alone never protects you.
The fix follows directly from the mechanism. If cost grows because context grows, cap the context. Truncating to a recent window keeps per-step cost flat - turning that quadratic curve back into a straight line. Summarizing old history into a short digest does the same while preserving more of what mattered. Either way, you're attacking the growth, not just the count.
⚡ Pro tip: Flatten the cost curve by capping context, and you often save more than by cutting steps. A step you remove saves one step's cost; capping context saves a little on every remaining step - and on a long run, that compounding saving is far larger.
Pro-level variations
For a startup running agents on user requests, add a per-user daily spend cap on top of per-run caps, because one user (or one bug in a loop) triggering many runs can blow the budget even if every single run stays under its own cap.
For a data team running scheduled batch agents, add pre-flight estimation on the batch size, since the cost scales with input volume and a data spike can turn a routine nightly job into an expensive one without any code change.
For a platform team running many agents, add an aggregate circuit breaker - if total spend across all agents in an hour crosses a threshold, halt and alert - because per-run caps don't catch a fleet-wide runaway where thousands of individually-cheap runs sum to a large number.
⚡ Pro tip: Cap at multiple scopes - per run, per user, per hour across the fleet. A single-run cap is necessary but not sufficient; the runaway that gets you is often a thousand cheap runs, not one expensive one, and only an aggregate cap catches that.
Troubleshooting common issues
If costs are high despite a step cap, your context is growing unbounded - cap and compact it. If a single user or job runs up the bill, you're missing per-user or per-job caps - add scoped limits. If costs spike suddenly across everything, you have no aggregate circuit breaker - add one. If a run costs far more than similar runs, check its retry count - a failure-heavy run multiplies calls invisibly.
The core insight: autonomous agent cost runaway is a multiplication problem. Loops, retries, context growth, and branching each multiply your base cost, and they stack. Every cap you add divides one of those multipliers back down.
Your turn
Take one agent and add just the context cap from the quick-start - the driver most people miss. Run a long task and compare the cost per step early versus late. On most agents the later steps are dramatically more expensive, and capping context flattens that curve - often the single biggest saving available.
The cap configurations and pre-flight estimation prompts are reusable across every agent you run. Keeping them in PromptABCD means your next agent ships with multi-scope cost protection from the start, instead of teaching you about cost runaway the way it teaches most people - through one memorable overnight bill for a task that should have cost pennies.
Continue Reading
Save the prompts from this post
PromptABCD is a free prompt manager. Paste, organize, and reuse your best AI prompts — no more hunting through chat history.
