Passing Context Between Agents Efficiently
One team's agents passed 40,000 tokens of context between each step - and 90% was never used. Multi agent context passing done right moves references, not payloads.
# Attempt 1: summarize everything between every step
def handoff(output):
return summarizer.run(f"Summarize for the next agent:\n{output}")When I profiled a struggling agent pipeline last year, one number stopped me cold: each handoff between agents carried about 40,000 tokens of context, and instrumentation showed the receiving agent actually referenced under 10% of it. The other 90% was pure tax - paid on every step, every run, forever. That's the pattern behind most multi agent context passing problems. Teams pass everything "just in case," and the just-in-case tax dwarfs the actual work.
Efficient multi agent context passing isn't about clever compression. It's about moving references instead of payloads, and being deliberate about what each agent genuinely needs. Here's the case study, because the fix took a system that was slow and expensive and made it both fast and cheaper without losing any capability.
The Problem This Data Team Faced
A data analytics team built a five-agent pipeline: an extractor pulled raw records, a cleaner normalized them, an analyzer ran statistics, an interpreter explained the results, and a writer produced a report. Reasonable design. The problem was how context flowed.
Each agent received the entire accumulated history - every prior agent's full output - concatenated into its prompt. So the writer, the last agent, received the raw extracted records, the full cleaned dataset, all the statistical output, and the interpretation. Most of that it never touched; it needed the interpretation and a summary, not 30,000 rows of raw data serialized into its context.
The symptoms were textbook. Cost scaled quadratically with pipeline length, because each new agent re-paid for all prior context. Latency climbed as prompts ballooned. And - the subtle one - quality dropped on longer runs, because burying the relevant interpretation under 35,000 tokens of raw data made the writer miss it. More context made the agent dumber, not smarter.
The Wrong Approach
The team's first fix was aggressive summarization: after each agent, run a summarizer to compress its output before passing it on.
[object Object],
,[object Object], ,[object Object],(,[object Object],):
,[object Object], summarizer.run(,[object Object],)What this does: It compresses each agent's output before the next agent sees it, shrinking the context. It also adds a model call at every handoff, loses detail that a later agent might have needed, and can introduce summarization errors that then propagate downstream as facts.
This helped cost but hurt correctness. The interpreter sometimes needed a specific statistic that the summarizer had rounded away. And summarization is lossy in an unpredictable direction - you can't know in advance which detail a downstream agent will want, so you either keep too much (no savings) or cut too much (broken pipeline). The team was trading a cost problem for a correctness problem, which is rarely a good trade.
⚡ Pro tip: Summarization is a tempting fix for context bloat and usually the wrong first move. It adds latency, adds a failure mode, and destroys information you can't recover. Reach for it only after you've tried passing references, because references are lossless and free.
The Correct Pattern
The fix was a shared artifact store plus reference passing. Instead of concatenating outputs into prompts, each agent writes its output to a store and passes forward a small handle - an ID, a schema, and a one-line description. The next agent reads only the specific artifacts it needs, by reference.
[object Object], ,[object Object],:
,[object Object], ,[object Object],(,[object Object],):
,[object Object],._store = {}
,[object Object], ,[object Object],(,[object Object],) -> ,[object Object],:
ref_id = hash_id(artifact)
,[object Object],._store[ref_id] = artifact
,[object Object], Ref(,[object Object],=ref_id, kind=artifact.kind,
summary=artifact.one_line, schema=artifact.schema)
,[object Object], ,[object Object],(,[object Object],):
,[object Object], ,[object Object],._store[ref_id]What this does: It stores each agent's full output once and hands back a lightweight reference - just an ID, a type, a one-line summary, and a schema. Downstream agents see the cheap reference in their prompt and pull the full artifact only when they actually need it, so nobody pays for data they don't use.
The agent prompts changed to work with references. Instead of "here is all prior output," an agent now sees a manifest of available artifacts and fetches what it needs:
WRITER_PROMPT = ,[object Object],What this does: It shows the writer a compact list of what's available rather than the data itself, and instructs it to pull artifacts on demand. The writer's prompt now carries a few hundred tokens of manifest instead of tens of thousands of tokens of raw data it would ignore anyway.
Results and What Changed
Average context per handoff dropped from roughly 40,000 tokens to about 3,000 - the manifest plus whatever the agent deliberately fetched. Cost per run fell by more than half, and because prompts were leaner, latency improved and the quality regression on long runs disappeared entirely. The writer stopped missing the interpretation because it was no longer buried.
The result I didn't predict: reference passing made the pipeline debuggable. Because every artifact was stored with an ID, I could see exactly which agent fetched which artifact - who used the raw data, who ignored it. That fetch log became the clearest map of the pipeline's real information flow I'd ever had, and it revealed two agents that fetched artifacts they never used, which we then trimmed.
⚡ Pro tip: Log every fetch. The set of artifacts an agent actually pulls is the ground truth of what it depends on - far more reliable than what its prompt claims it needs. Prune the manifest down to what agents actually fetch and you'll find another round of savings almost every time.
⚡ Pro tip: When you convert a handoff to references, keep the old full-context path behind a flag for one release so you can A/B the quality. Reference passing occasionally starves an agent that was quietly relying on data you assumed it ignored, and the flag lets you catch that instantly instead of shipping a subtle regression.
How to Apply This to Your Situation
Start by measuring your real context-usage ratio. Instrument one pipeline to record how many tokens each agent receives versus how many it references in its output. If the ratio is bad - and for most pipelines passing full history, it's terrible - references will pay off immediately.
Then introduce an artifact store and convert handoffs one at a time. Begin with your largest payload, usually raw data or a big intermediate result, and replace it with a reference. Give agents an explicit fetch tool and a prompt that tells them to fetch narrowly.
The design principle underneath all of this: context is a resource with a cost, and passing it should be a deliberate act, not a default of "include everything." When you make each agent declare and fetch what it needs, multi agent context passing stops being a hidden quadratic cost and becomes a small, legible line item.
⚡ Pro tip: Pass a schema alongside every reference, not just an ID. When an agent knows an artifact's shape without fetching it, it can often decide it doesn't need the data at all - answering from the schema and summary. The schema is the cheapest possible context, and it prevents a large fraction of fetches outright.
What About Context an Agent Genuinely Needs Every Time?
References work beautifully for large, occasionally-needed payloads. But some context is small and needed by everyone - the task goal, the user's constraints, the output format. Passing those by reference is silly; the fetch costs more than just including them. So multi agent context passing is really a tiering decision, and getting the tiers right is where the craft lives.
I sort context into three tiers. Tier one is ambient - tiny and universally needed, like the goal and constraints. Include it inline in every agent, always. It's cheap and its absence causes agents to drift off-task. Tier two is referenced - large and selectively needed, like datasets and full intermediate outputs. Pass these by reference and let agents fetch. Tier three is derived - things an agent can compute from what it already has, which shouldn't be passed at all.
[object Object], ,[object Object],(,[object Object],):
,[object Object], {
,[object Object],: task.goal_and_constraints, ,[object Object],
,[object Object],: store.manifest(scope=agent.role), ,[object Object],
,[object Object],
}What this does: It assembles each agent's input from the two tiers that belong there - ambient context inline, everything large as references - and deliberately omits derived context. The agent always has its goal, can reach anything big on demand, and isn't handed things it could work out itself.
The mistake I see most is treating all context as one undifferentiated blob and then arguing about whether to "pass it or not." That's the wrong question. Different context wants different handling, and once you tier it, most of the debate evaporates - ambient is obviously inline, big stuff is obviously referenced, and derived stuff obviously shouldn't travel at all.
⚠️ Common mistake: Passing an agent the conversation history of other agents as context. Another agent's full back-and-forth is almost never what a receiver needs - it needs the conclusion, not the deliberation. Passing raw agent transcripts as context is one of the largest and most useless token sinks in multi-agent systems, because deliberation is verbose and the useful signal is a tiny fraction of it. Pass the conclusion as an artifact; discard the transcript.
Next Steps
Instrument the usage ratio first, convert your biggest handoff to references second, and log fetches so you can keep trimming. Expand the artifact store to cover every handoff once you've proven it on one. One practical sequencing note: don't convert every handoff at once. Each conversion changes an agent's prompt, and changing several prompts simultaneously makes it hard to attribute a regression. Convert one handoff, confirm quality holds, then move to the next - the boring, incremental path is far faster overall than a big-bang rewrite that leaves you bisecting five changes when something breaks.
Keep the reference-aware agent prompts - the manifest format, the "fetch narrowly" instruction - versioned so you can reuse them. I store these in PromptABCD and drop the same manifest-and-fetch prompt into every new pipeline, because teaching an agent to fetch on demand instead of demanding everything up front is the reusable skill, and the exact wording that produces disciplined fetching is worth keeping intact.
Continue Reading
Save the prompts from this post
PromptABCD is a free prompt manager. Paste, organize, and reuse your best AI prompts — no more hunting through chat history.
