Reducing Token Costs in Multi-Agent Systems
Picture opening a bill 8x your estimate for an agent team that barely did anything. Multi agent token cost hides in re-sent context and chatty coordination. Here's the teardown.
def run_task(task):
context = build_full_context(task) # ~8k tokens, sent to everyone
plan = coordinator.run(context) # sees full context
results = []
for worker in workers:
# each worker re-receives full context + plan + prior results
r = worker.run(context + plan + results)
results.append(r)
return coordinator.run(context + plan + results) # final synthesisPicture opening your provider bill on a Monday and finding it's eight times your estimate for an agent team that, on paper, handled a few hundred tasks. No bug, no runaway loop, nothing crashed. The system "worked." It just quietly cost a fortune. This is the most common way multi agent token cost surprises people - not through a dramatic failure, but through steady, invisible waste spread across thousands of ordinary calls.
Let's tear down a real high-cost agent team and find where the tokens actually went, because the culprits are almost never where people look first. Everyone blames "the model is expensive." The real waste is architectural.
Before: The Expensive Design
Here's the shape of the system, simplified. A coordinator and four worker agents, collaborating on document processing.
[object Object], ,[object Object],(,[object Object],):
context = build_full_context(task) ,[object Object],
plan = coordinator.run(context) ,[object Object],
results = []
,[object Object], worker ,[object Object], workers:
,[object Object],
r = worker.run(context + plan + results)
results.append(r)
,[object Object], coordinator.run(context + plan + results) ,[object Object],What this does: It builds a large shared context and passes it, plus the growing results list, to every agent on every call. Each worker re-receives the full context and everything prior workers produced - so the same 8,000 tokens get sent five, six, seven times in a single task.
Why It Fails
The waste is multiplicative, and that's what makes it brutal. That 8,000-token context isn't sent once - it's sent to the coordinator, then to each of four workers, then to the final synthesis. Six times, minimum. And the results list grows, so late workers receive early workers' full outputs too. A task that should cost the price of processing a document once ends up paying to re-read the same context six times plus re-read intermediate results.
Break it down and three cost centers dominate. Re-sent context is the biggest - the same tokens billed repeatedly because there's no mechanism to avoid re-sending what an agent already effectively knows. Chatty coordination is second - agents that converse back and forth ("Are you done?" "Not yet." "How about now?") each turn costing a full round-trip. And model over-provisioning is third - using a top-tier model for every agent, including the ones doing trivial formatting or routing that a small model handles perfectly.
⚠️ Common mistake: Optimizing the prompt wording to save tokens while ignoring the architecture that re-sends context six times. Trimming 200 tokens from a prompt that gets sent 6,000 times saves something; fixing the re-send saves 5x that with one change. People obsess over prompt golf because it's visible and skip the architecture because it's not. Measure where the tokens go before you trim a single word.
⚡ Pro tip: Before optimizing anything, multiply each prompt's token count by how many times per task it's sent. That product - not the raw prompt size - is your real cost per component. A 500-token prompt sent twelve times costs more than a 2,000-token prompt sent once, and yet teams reliably attack the big prompt and ignore the small-but-repeated one. The multiplication reorders your priorities correctly.
After: The Efficient Design
The rewrite attacks all three cost centers. Context gets sent once and referenced thereafter, coordination becomes event-based instead of conversational, and each agent gets right-sized to its job.
[object Object], ,[object Object],(,[object Object],):
ctx_ref = store.put(build_full_context(task)) ,[object Object],
plan = coordinator.run(ctx_ref) ,[object Object],
results = []
,[object Object], worker ,[object Object], workers:
,[object Object],
r = worker.run(ctx_ref, plan_ref=plan.ref, new_since=results[-,[object Object],:])
results.append(r)
,[object Object], synthesizer.run(plan.ref, results_ref=store.put(results))What this does: It stores the big context once and passes references, so the 8,000 tokens are paid for a single time instead of six. Workers see only recent results rather than the entire growing list, and a right-sized synthesizer replaces a second full coordinator call. Same behavior, a fraction of the tokens.
Breaking Down Each Element
The reference-passing change is the single biggest lever on multi agent token cost, and it maps directly to the re-send problem. By storing context once and passing a handle, you convert "context cost times number of agents" into "context cost times one." On a five-agent pipeline that's roughly an 80% cut on context tokens alone, and the savings grow with team size - the bigger the team, the more times you were re-sending, so the more you save by sending once.
The new_since slice fixes the growing-results problem. Instead of every worker receiving all prior results, each sees only the last few. Most workers only need recent context, and the ones that need older results can fetch them by reference. This stops the quadratic growth where the last worker pays for every earlier worker's full output.
Right-sizing models is the third lever and the most underused. Not every agent needs the flagship model. Routing, classification, formatting, and validation often run fine on a small, cheap model. Reserve the expensive model for genuine reasoning steps. A mixed fleet - cheap models for mechanical work, expensive models for hard thinking - routinely cuts cost by half again with no quality loss on the mechanical tasks.
AGENT_MODELS = {
,[object Object],: ,[object Object],, ,[object Object],
,[object Object],: ,[object Object],, ,[object Object],
,[object Object],: ,[object Object],, ,[object Object],
,[object Object],: ,[object Object],, ,[object Object],
}What this does: It assigns each agent the smallest model that does its job well, instead of defaulting everything to the flagship. The router and formatter run on a cheap model because their work is mechanical; only the analyzer, which genuinely reasons, gets the expensive model.
⚡ Pro tip: Audit model assignment by asking, per agent, "would a smaller model produce an acceptable output here?" For routing, classification, and formatting the answer is almost always yes. Teams leave enormous savings on the table by defaulting every agent to their best model out of caution, when most agents are doing work a cheap model aces.
Variations for Different Contexts
For high-volume, latency-tolerant workloads, add a caching layer keyed on operation inputs. Identical retrievals and identical sub-computations return cached results at near-zero cost. In systems with repetitive tasks, cache hits alone can cut multi agent token cost by a third.
Before any of these variations, though, you need to see the cost, and this is where most teams go wrong. They optimize by intuition. Intuition is reliably incorrect about where tokens go, because the expensive parts are invisible - re-sent context doesn't show up as a line item, it hides inside every prompt. So instrument first: record tokens per agent, per task, broken out by prompt versus completion.
[object Object], ,[object Object],(,[object Object],):
by_agent = defaultdict(,[object Object],: {,[object Object],: ,[object Object],, ,[object Object],: ,[object Object],})
,[object Object], call ,[object Object], trace:
by_agent[call.agent][,[object Object],] += call.prompt_tokens
by_agent[call.agent][,[object Object],] += call.completion_tokens
,[object Object], ,[object Object],(by_agent)What this does: It attributes every token to a specific agent and splits prompt from completion cost. The split matters enormously - if an agent's prompt tokens dwarf its completion tokens, you're paying to send it context it barely uses, which points straight at reference passing. If completion dominates, the agent is genuinely verbose and needs a tighter output spec.
The prompt-versus-completion ratio is the most diagnostic number in the whole exercise and almost nobody looks at it. A healthy reasoning agent has substantial completion tokens - it's doing work. An agent whose bill is 95% prompt tokens is mostly being fed context it drowns in. That ratio tells you which optimization applies to which agent, so you fix the right thing instead of trimming prompts everywhere and hoping.
For interactive systems where latency matters, focus on the reference passing and model right-sizing rather than caching, since caching adds a lookup that can hurt tail latency. The right optimization depends on whether you're optimizing for throughput or responsiveness - they pull in slightly different directions.
⚡ Pro tip: Compute the prompt-to-completion token ratio per agent. Above roughly 10:1, the agent is context-heavy and reference passing is your lever. Below 2:1, the agent is output-heavy and a tighter output spec or a cheaper model is your lever. This one ratio routes you to the correct fix per agent instead of applying the same optimization blindly across the whole team.
⚡ Pro tip: Set a per-task token budget and have the system report tasks that blow past it. The outliers - the 5% of tasks consuming 40% of tokens - are where your savings concentrate. Optimizing the median task is far less productive than finding and fixing the handful of pathological ones.
Save and Reuse This
The reusable pattern is the trio: pass references not payloads, send only new results not the full history, and right-size models per agent. Applied together, they typically cut an unoptimized agent team's cost by 70-80% with no capability loss - the "8x bill" becomes roughly a 1.5x bill over a single-agent baseline, which is a fair price for parallelism.
Keep your per-agent model assignments and your reference-passing prompts versioned so you can reuse a known-cheap configuration. I store the model-routing map and the reference-aware prompts in PromptABCD, because the work of figuring out which agents can run cheap without quality loss is real experimentation, and once you've done it, you want to reapply the exact configuration to the next team rather than re-running the whole audit.
One last framing worth internalizing: multi agent token cost is not a single problem, it's three problems wearing a trench coat - re-sent context, chatty coordination, and over-provisioned models. Each has a different fix, and applying the wrong fix to the wrong cause is why so much optimization effort produces so little savings. Measure first, identify which of the three dominates your bill, and apply the matching lever. A team that trims prompt wording when their real problem is re-sent context will work hard and save almost nothing; a team that passes references when their real problem is an oversized model will do the same. The measurement is what routes you to the fix that actually moves the number.
Continue Reading
Save the prompts from this post
PromptABCD is a free prompt manager. Paste, organize, and reuse your best AI prompts — no more hunting through chat history.
