Single Agent vs Multi-Agent: When to Split
Single vs multi agent isn't a question of which is more sophisticated — it's about which fits the task. Here's a practical framework for making the right call before you build.
# A single-agent approach that hits its ceiling
from anthropic import Anthropic
client = Anthropic()
def single_agent_pipeline(user_request: str, context_data: str) -> str:
# Works fine for simple tasks. Starts degrading past ~8 reasoning steps.
response = client.messages.create(
model="claude-sonnet-4-6",
max_tokens=4096,
system="""You are a research analyst. Given data and a user request, you will:
1. Extract relevant data points
2. Identify trends and patterns
3. Cross-reference with industry benchmarks
4. Write a structured analysis
5. Recommend three specific actions
6. Estimate ROI for each action
7. Flag data quality issues
8. Summarize for executive audience
Do all of these thoroughly.""",
messages=[{
"role": "user",
"content": f"Request: {user_request}\n\nData:\n{context_data}"
}]
)
return response.content[0].textPicture this: you're a backend engineer building an internal tool that generates API documentation from OpenAPI specs. You wire it up as a multi-agent system — a parser agent, a writing agent, a formatter agent, a reviewer agent. It takes six weeks to build, costs four times more per run than expected, and still occasionally produces documentation that contradicts itself. Your colleague builds the same thing as a single well-prompted agent in two days and it works fine.
This isn't a hypothetical. The single vs multi agent decision is one of the most commonly botched choices in AI engineering, and it almost always fails in one direction: developers reach for multi-agent complexity when a single agent would have served them better.
Before: The Single-Everything Approach (and Its Real Limits)
A single agent is powerful. Modern LLMs can handle genuinely complex tasks in one call — summarize a 50-page report, write a complete Python module, plan a multi-step project. The temptation is to just keep adding to a single agent's instructions until it handles everything.
This works until it doesn't. The failure modes are specific:
Context window saturation: Past around 8–10 dense reasoning steps, a single agent's ability to maintain constraint coherence degrades measurably. It's not that it forgets — the information is still in context — but attention becomes diluted. Early decisions carry less weight. Constraint drift sets in.
Task type heterogeneity: A single agent asked to simultaneously perform legal analysis, code generation, and customer-facing copywriting will do all three worse than an agent optimized for each. The system prompt becomes a compromise that serves none of the tasks well.
Lack of validation: A single agent cannot meaningfully evaluate its own output for errors it didn't anticipate. Self-critique prompts help with surface problems but don't catch the logical errors the agent already committed to in its reasoning.
[object Object],
,[object Object], anthropic ,[object Object], Anthropic
client = Anthropic()
,[object Object], ,[object Object],(,[object Object],) -> ,[object Object],:
,[object Object],
response = client.messages.create(
model=,[object Object],,
max_tokens=,[object Object],,
system=,[object Object],,
messages=[{
,[object Object],: ,[object Object],,
,[object Object],: ,[object Object],
}]
)
,[object Object], response.content[,[object Object],].textWhat this does: This single-agent approach works well for requests that need 3-4 of these steps. Once you need all 8 with high accuracy on each, the agent's constraint coherence starts to slip — and it's hard to tell which step degraded because the output is monolithic.
Why the Single-Agent-Only Approach Fails at Scale
The real ceiling isn't token count — it's reasoning fidelity across multiple distinct cognitive modes.
Legal analysis requires different reasoning patterns than code generation. Code generation requires different evaluation criteria than customer copywriting. When you ask a single agent to switch between cognitive modes inside one call, you're asking it to maintain multiple, sometimes conflicting heuristics simultaneously.
The result is output that's competent at each task individually but incoherent across them. The legal language appears in the code comments. The code logic influences the customer communication. The agent finds compromises that satisfy no task well.
⚠️ Common mistake: Treating the multi-agent decision as a question of "can the single agent do this?" rather than "does the single agent do this reliably enough for production?" Many tasks are technically possible in a single call but fail at a rate that's unacceptable for real users. The question isn't capability — it's reliability under production conditions.
After: The Multi-Agent Approach (and When It Actually Helps)
A well-designed multi-agent split solves specific problems: it gives each agent a focused cognitive mode, it enables parallel execution of independent subtasks, and it introduces validation checkpoints that a single agent can't provide for itself.
[object Object], anthropic ,[object Object], Anthropic
,[object Object], concurrent.futures
client = Anthropic()
,[object Object], ,[object Object],(,[object Object],) -> ,[object Object],:
response = client.messages.create(
model=,[object Object],,
max_tokens=,[object Object],,
system=,[object Object],,
messages=[{,[object Object],: ,[object Object],, ,[object Object],: ,[object Object],}]
)
,[object Object], response.content[,[object Object],].text
,[object Object], ,[object Object],(,[object Object],) -> ,[object Object],:
response = client.messages.create(
model=,[object Object],,
max_tokens=,[object Object],,
system=,[object Object],,
messages=[{,[object Object],: ,[object Object],, ,[object Object],: ,[object Object],}]
)
,[object Object], response.content[,[object Object],].text
,[object Object], ,[object Object],(,[object Object],) -> ,[object Object],:
response = client.messages.create(
model=,[object Object],,
max_tokens=,[object Object],,
system=,[object Object],,
messages=[{,[object Object],: ,[object Object],, ,[object Object],: ,[object Object],}]
)
,[object Object], response.content[,[object Object],].text
,[object Object], ,[object Object],(,[object Object],) -> ,[object Object],:
findings = data_analyst_agent(request, data)
strategy = strategy_agent(findings, request)
summary = writer_agent(strategy, audience)
,[object Object], {,[object Object],: findings, ,[object Object],: strategy, ,[object Object],: summary}What this does: Each agent operates in a single cognitive mode. The analyst never writes executive prose. The strategist never extracts raw data patterns. The writer never reasons about statistical confidence. Each agent's system prompt can be optimized for its specific mode without compromising any other.
⚡ Pro tip: The clearest signal that you need a multi-agent split is when you find yourself writing "but first check X, then switch to Y mode, but remember Z from earlier" inside a single system prompt. That's the orchestrator thinking leaking into an agent's instructions. Extract the logic into actual orchestration code.
Breaking Down the Decision
The single vs multi agent decision comes down to four diagnostic questions:
How many distinct reasoning modes does the task require? Count the cognitive shifts — "analyze data" to "write strategy" to "communicate to executives" is three distinct modes. One or two modes: single agent. Three or more: consider splitting.
Are there independent subtasks? If parts of the task don't depend on each other's outputs, parallelism is available. Running five competitive analyses simultaneously instead of sequentially changes a 10-minute job into a 2-minute job. A single agent can't do that.
How long is the full task chain? Under 6-8 reasoning steps with moderate complexity: single agent. Over that threshold, or with high-stakes requirements on each step: multi-agent.
Do you need validation you can trust? If output quality errors have real consequences — wrong code ships, bad advice goes to clients, errors reach customers — you need a critic agent. Self-critique from the generating agent is not the same thing.
⚡ Pro tip: Run the same task through a single agent and a two-agent (generator + critic) setup ten times each. Count errors in both sets. That empirical comparison is worth more than any theoretical framework for deciding which architecture fits your specific task.
Variations for Different Contexts
For content teams: Single agent handles first drafts, outlines, and research summaries well. Split to multi-agent when the workflow involves both creation and fact-checking, or when you need brand-voice compliance checked separately from content quality.
For data engineering teams: Single agent works for schema documentation, query explanation, and simple transformations. Multi-agent becomes necessary when you need simultaneous validation of data quality, SQL correctness, and business logic alignment.
For customer service automation: Single agent handles standard FAQ responses adequately. Multi-agent is worth the cost when you need sentiment classification, policy compliance checking, and response personalization happening independently before synthesis.
Save and Reuse This Framework
The single vs multi agent decision isn't a one-time architectural choice — it's a recurring question as your requirements evolve. Workflows that start simple enough for a single agent often grow into multi-agent territory over months.
Keep the four diagnostic questions above in your decision toolkit. When a single-agent workflow starts showing quality degradation, run the diagnostics before assuming more agents is the answer — sometimes the issue is the system prompt, not the architecture.
The Real Cost of Multi-Agent Architecture
Before committing to a multi-agent design, engineers typically estimate two costs: the API call overhead and the implementation time. These are real but often underestimated. The costs that are consistently overlooked are operational — and they compound over the system's lifetime.
Latency compounds with agent count. A five-agent sequential pipeline running at three seconds per call takes a minimum of fifteen seconds per end-to-end request. In practice it's longer, because the slowest agent on any given input stalls all downstream agents. If your users expect sub-five-second responses, sequential multi-agent pipelines require either aggressive caching, parallel execution, or genuine architectural re-evaluation.
Debugging surface expands nonlinearly. When a single-agent system produces wrong output, you look at one prompt and one response. When a five-agent system produces wrong output, you inspect five prompts and five responses, tracking where the error was introduced and how it propagated. Each additional agent multiplies the debugging surface. Teams that don't log every agent's intermediate output (inputs, outputs, latency, model used) spend three times as long diagnosing failures in multi-agent systems as in equivalent single-agent systems.
Context window management becomes explicit work. In a single agent, context is one conversation. In a multi-agent pipeline, every agent operates on partial context — what it explicitly received, not everything that happened earlier in the pipeline. Designing what information each agent needs, serializing it compactly, and ensuring downstream agents have the right context without being overloaded is real engineering work that single-agent systems don't require.
⚡ Pro tip: Before building a multi-agent system, write down the specific failure modes of the single-agent version that motivate the split. "The quality isn't good enough" is not a specific failure mode. "The single agent generates research that contradicts itself because it doesn't distinguish between finding information and evaluating information" is specific enough to design an agent split around. Specific failure modes produce targeted architectures; vague dissatisfaction produces over-engineered ones.
The Minimal Multi-Agent Starting Point
If you're convinced multi-agent is right for your task, start with the minimum viable split: two agents.
One agent generates. One agent validates. The generator produces output as if it will be used directly; the validator checks it against specific criteria you define. This two-agent pattern handles the most common motivation for multi-agent architecture — that a single agent can't both generate freely and evaluate critically at the same time — without introducing routing complexity, coordination overhead, or multiple points of failure.
If two agents don't solve the problem, add a third for a specific reason. The decision to add each subsequent agent should be grounded in a measured quality improvement over the N-1 architecture, not architectural ambition.
If you're maintaining a library of agent system prompts for different cognitive modes, PromptABCD makes it straightforward to version and iterate on each one independently, which is exactly how effective multi-agent prompt libraries get built.
Continue Reading
Save the prompts from this post
PromptABCD is a free prompt manager. Paste, organize, and reuse your best AI prompts — no more hunting through chat history.
