The Router Pattern for Multi-Agent Systems
A 2025 analysis of production multi-agent failures found that misrouting — sending a task to the wrong specialist — was responsible for more quality failures than individual agent errors. The router is the most critical agent in your system.
from anthropic import Anthropic
client = Anthropic()
def naive_router(ticket: str) -> str:
response = client.messages.create(
model="claude-sonnet-4-6",
max_tokens=128,
system="""Classify customer support tickets. Respond with ONLY one word:
- BILLING: payment, invoice, subscription, refund
- TECHNICAL: bug, error, not working, broken
- FEATURE: request, suggestion, would be nice""",
messages=[{"role": "user", "content": ticket}]
)
return response.content[0].text.strip()
# This router will classify ANYTHING as one of three options
# A spam message gets classified as one of three categories
# An abusive message gets classified as one of three categories
# A message in a language the agents don't support gets classified
# There's no "none of the above" optionA 2025 analysis of production multi-agent system failures across several enterprise deployments found that misrouting — sending a task to the wrong specialist agent — was responsible for more output quality failures than individual agent errors combined. Individual agents, when given the right input, performed well. The breakdown happened before they received any input at all.
The router is the most critical agent in a multi-agent system, and it gets the least design attention. Most implementations treat it as a simple classifier: analyze the input, pick a category, dispatch to the matching agent. That works until you get an input that doesn't match any category, or matches multiple categories, or shouldn't be routed to any agent at all.
The agent router pattern done correctly routes to "none" as a first-class option, handles ambiguous inputs explicitly, and logs routing decisions in a way that makes misrouting diagnosable.
The Problem: A Customer Support Router That Couldn't Say No
A SaaS company built a customer support multi-agent system with three specialist agents: billing, technical support, and feature requests. The router classified each incoming ticket into one of these three categories and dispatched accordingly.
Three months in, they discovered an unexpected failure mode: phishing attempts and spam that reached their support email were being routed to their technical support agent, which would helpfully attempt to address "technical issues" described in the phishing message. The router had never been given the option to reject a ticket — it was trained to classify everything as billing, technical, or feature request.
The fix required a conceptual change: the router needed a "none of the above" category, plus an explicit escalation path for inputs that didn't belong in any automated workflow.
The Wrong Approach: Classification Without Rejection
[object Object], anthropic ,[object Object], Anthropic
client = Anthropic()
,[object Object], ,[object Object],(,[object Object],) -> ,[object Object],:
response = client.messages.create(
model=,[object Object],,
max_tokens=,[object Object],,
system=,[object Object],,
messages=[{,[object Object],: ,[object Object],, ,[object Object],: ticket}]
)
,[object Object], response.content[,[object Object],].text.strip()
,[object Object],
,[object Object],
,[object Object],
,[object Object],
,[object Object],What this does: Forces every input into one of three categories, regardless of whether any category is appropriate. Ambiguous tickets (billing AND technical issues simultaneously) are forced into one bucket arbitrarily. Invalid inputs — spam, abuse, unsupported languages — are classified into legitimate workflows.
⚠️ Common mistake: Building a router that treats classification as the only output. Routing decisions should include: which agent to route to, confidence in the routing decision, whether multi-agent routing is needed, and whether the input should be rejected entirely. A router that can only classify and dispatch is missing three of its four responsibilities.
The Correct Approach: Router With Rejection and Confidence Scoring
[object Object], json
,[object Object], anthropic ,[object Object], Anthropic
client = Anthropic()
AGENT_REGISTRY = {
,[object Object],: {
,[object Object],: ,[object Object],,
,[object Object],: [,[object Object],, ,[object Object],, ,[object Object],]
},
,[object Object],: {
,[object Object],: ,[object Object],,
,[object Object],: [,[object Object],, ,[object Object],, ,[object Object],]
},
,[object Object],: {
,[object Object],: ,[object Object],,
,[object Object],: [,[object Object],, ,[object Object],, ,[object Object],]
}
}
,[object Object], ,[object Object],(,[object Object],) -> ,[object Object],:
agent_descriptions = json.dumps(AGENT_REGISTRY, indent=,[object Object],)
response = client.messages.create(
model=,[object Object],,
max_tokens=,[object Object],,
system=,[object Object],,
messages=[{,[object Object],: ,[object Object],, ,[object Object],: ,[object Object],}]
)
,[object Object],:
routing = json.loads(response.content[,[object Object],].text)
,[object Object], routing
,[object Object], json.JSONDecodeError:
,[object Object], {
,[object Object],: ,[object Object],,
,[object Object],: ,[object Object],,
,[object Object],: ,[object Object],,
,[object Object],: ,[object Object],,
,[object Object],: ,[object Object],,
,[object Object],: ,[object Object],
}
,[object Object], ,[object Object],(,[object Object],) -> ,[object Object],:
routing = intelligent_router(ticket)
,[object Object], routing[,[object Object],] == ,[object Object],:
,[object Object], {
,[object Object],: ,[object Object],,
,[object Object],: routing.get(,[object Object],, ,[object Object],),
,[object Object],: ,[object Object],
}
,[object Object], routing[,[object Object],] < ,[object Object],:
,[object Object], {
,[object Object],: ,[object Object],,
,[object Object],: routing,
,[object Object],: ,[object Object],
}
,[object Object],
primary_result = dispatch_to_agent(routing[,[object Object],], ticket)
,[object Object],
,[object Object], routing[,[object Object],] ,[object Object], routing[,[object Object],]:
secondary_result = dispatch_to_agent(routing[,[object Object],], ticket)
,[object Object], {
,[object Object],: ,[object Object],,
,[object Object],: routing,
,[object Object],: primary_result,
,[object Object],: secondary_result
}
,[object Object], {,[object Object],: ,[object Object],, ,[object Object],: routing, ,[object Object],: primary_result}
,[object Object], ,[object Object],(,[object Object],) -> ,[object Object],:
agent_prompts = {
,[object Object],: ,[object Object],,
,[object Object],: ,[object Object],,
,[object Object],: ,[object Object],
}
response = client.messages.create(
model=,[object Object],,
max_tokens=,[object Object],,
system=agent_prompts[agent_name],
messages=[{,[object Object],: ,[object Object],, ,[object Object],: ticket}]
)
,[object Object], response.content[,[object Object],].textWhat this does: The router returns a structured decision with confidence score, multi-agent routing support, and rejection handling. Tickets with confidence below 0.7 go to human review rather than being forced into an uncertain dispatch. Spam, abuse, and out-of-scope tickets route to "none" rather than being sent to legitimate agents. Multi-category tickets can be dispatched to two agents simultaneously.
Results and What Changed
After implementing the agent router pattern with rejection handling:
Quality improvement: The routing accuracy on a 500-ticket test set improved from 78% (naive classifier) to 94% (structured router with confidence scoring). The 16% improvement came almost entirely from correctly identifying ambiguous and invalid tickets rather than from better classification of clear-cut tickets.
Spam containment: The company identified that 4% of incoming "tickets" were spam or phishing attempts. The "none" routing option now catches these before they reach any automated agent. Previously, all 4% were being processed by the technical support agent.
Debugging tractability: Routing decisions logged with reasons and confidence scores made misrouting diagnosable. The team could filter their logs to low-confidence routings and find systematic gaps in their agent registry descriptions.
⚡ Pro tip: Treat your agent registry as living documentation. When the router makes a wrong routing decision — and it will — the fix is almost always to improve the agent description in the registry, not to retrain the router. The router uses agent descriptions to reason about which agent to choose. Better descriptions produce better routing.
How to Apply This to Your Situation
The agent router pattern applies whenever you have more than two specialist agents and the routing decision requires understanding the input content rather than parsing structured metadata.
Before implementing, answer three questions: What are the valid rejection reasons for your domain? (Spam, out-of-scope, unsupported input type, ambiguous.) What confidence threshold makes sense? (0.7 is a reasonable default; adjust based on the cost of misrouting.) How do you handle multi-category inputs? (Dual routing, ask for clarification, or route to a generalist that can handle both.)
Next Steps
Build the rejection cases before building the routing cases. Knowing what should not route to any agent forces clarity about what your agents are actually responsible for. A router that handles rejection well will handle routing well — the hard cognitive work is the same.
Multi-Step and Cascading Routing
Simple routing maps one input to one agent. Real routing problems often require cascading decisions: the initial router makes a broad classification, then a second router makes a narrower one within that category.
Consider a legal document processing system: the first router classifies the document type (contract, litigation, compliance, IP). A second router within the "contract" category classifies the contract type (employment, vendor, NDA, service agreement). Each router at each tier has a smaller, better-defined decision space. A single router trying to classify all document types and subtypes in one step produces a much harder routing problem than two sequential routers each with a narrow scope.
[object Object], json
,[object Object], anthropic ,[object Object], Anthropic
client = Anthropic()
TIER1_ROUTER_SYSTEM = (
,[object Object],
,[object Object],
,[object Object],
)
TIER2_CONTRACT_ROUTER_SYSTEM = (
,[object Object],
,[object Object],
,[object Object],
)
,[object Object], ,[object Object],(,[object Object],) -> ,[object Object],:
,[object Object],
t1_response = client.messages.create(
model=,[object Object],, max_tokens=,[object Object],,
system=TIER1_ROUTER_SYSTEM,
messages=[{,[object Object],: ,[object Object],, ,[object Object],: document_text[:,[object Object],]}]
)
tier1 = json.loads(t1_response.content[,[object Object],].text)
route = {,[object Object],: tier1[,[object Object],], ,[object Object],: tier1[,[object Object],]}
,[object Object],
,[object Object], tier1[,[object Object],] == ,[object Object], ,[object Object], tier1[,[object Object],] > ,[object Object],:
t2_response = client.messages.create(
model=,[object Object],, max_tokens=,[object Object],,
system=TIER2_CONTRACT_ROUTER_SYSTEM,
messages=[{,[object Object],: ,[object Object],, ,[object Object],: document_text[:,[object Object],]}]
)
tier2 = json.loads(t2_response.content[,[object Object],].text)
route[,[object Object],] = tier2[,[object Object],]
route[,[object Object],] = tier2[,[object Object],]
,[object Object], routeWhat this does: The cascade router makes two sequential classification decisions. The first determines the broad document category. If classification confidence exceeds the threshold and the category is "contract," a second router narrows to contract type. Each router operates on the same document text but with a system prompt scoped to its specific classification task. The result includes both tier classifications and confidence scores.
⚡ Pro tip: Use confidence thresholds to decide when to run a second-tier router. If the tier-1 router is uncertain (confidence below 0.7), escalating to a tier-2 classifier within the uncertain category is likely to produce another uncertain result. For low-confidence inputs, route to a human review queue rather than cascading into increasingly uncertain classifications.
Testing Router Accuracy
Router accuracy is measurable in a way that most LLM task quality isn't. For any routing system, you can build a labeled test set: inputs with known correct routing decisions. Running the router against this set gives you an accuracy score that tracks changes to the routing prompt over time.
Build the test set before tuning the router prompt. Including test cases in your prompt tuning process risks over-fitting the router to the test set — the router learns to classify your examples correctly without improving on novel inputs. Keep test and development data separate.
A minimum viable router test set covers: unambiguous inputs for each route (the easy cases), ambiguous inputs that sit on the boundary between routes (the hard cases), out-of-scope inputs that should trigger rejection, and multi-intent inputs that need multi-agent routing. If your router handles all four categories correctly, it will handle most production inputs correctly.
⚡ Pro tip: Log every routing decision in production, including the agent selected and the confidence score. After two weeks, review the low-confidence decisions — those where the router chose but wasn't certain. These are the inputs where routing is most likely to be wrong, and they reveal which classification boundaries in your agent registry descriptions need refinement. Low-confidence cases are your highest-value data for improving routing accuracy.
The Router as System Boundary
A router in a multi-agent system is more than a classification component — it is the system boundary between the outside world and the specialized agents behind it. Everything that enters the system passes through the router. This gives the router two responsibilities beyond routing:
Input normalization. The router is the right place to normalize inputs before they reach specialized agents. Stripping irrelevant context, standardizing terminology, correcting obvious formatting issues — these preprocessing steps improve agent performance and should happen at the routing layer, not inside each individual agent.
Capacity management. The router can enforce rate limits, priority queues, and load balancing across its agent pool. High-priority inputs route to dedicated capacity; low-priority inputs wait. When a specialist agent is overloaded, the router can queue additional inputs rather than stacking requests on an already-saturated agent. These concerns belong at the routing layer, not distributed across all agents.
The agent router pattern's scope is the full boundary function: normalize inputs, classify intent, manage capacity, route to specialists, handle rejections. Systems that treat the router as only a classifier implement only a fraction of the pattern's value.
Implementing systematic input normalization at the router also makes the specialist agents simpler. Each specialist receives a cleaned, well-scoped input and doesn't need to handle the variety of raw input formats that arrive from external callers. That simplification pays off in prompt quality — specialist system prompts can focus on domain expertise rather than input handling.
When you've tuned your router's system prompt and your agent registry descriptions, save them both to PromptABCD. The registry descriptions evolve as your agent capabilities evolve — keeping them versioned alongside the router prompt makes updates systematic rather than ad hoc.
Continue Reading
Save the prompts from this post
PromptABCD is a free prompt manager. Paste, organize, and reuse your best AI prompts — no more hunting through chat history.
