Multi-Agent Systems for Customer Support: A Prompt Teardown
Ever wonder why your support bot escalates the easy stuff and confidently botches the hard stuff? A multi agent customer support setup fixes the routing — here's the weak prompt, why it fails, and the rebuild.
You are a helpful customer support agent for {company}.
Answer customer questions accurately and politely. You can
reset passwords, check order status, explain policies, and
process refunds. If you can't help, escalate to a human.
Always be empathetic and concise.Why does your support bot handle a password reset perfectly and then confidently give a wrong answer about a refund policy it never actually read? If you've run any support automation, you've watched this exact split. The bot is great at the mechanical stuff and dangerous at the judgment stuff — and cranking up the instructions makes both worse. A multi agent customer support design fixes it, but only if you understand why the single prompt fails first.
This is a teardown. We'll start with the prompt almost everyone ships, take it apart, and rebuild it as a small team of agents that route by capability instead of pretending one agent handles everything.
Before: the weak prompt
Here's the single-agent support prompt that ships on day one across thousands of companies:
You are a helpful customer support agent for {company}.
Answer customer questions accurately and politely. You can
reset passwords, check order status, explain policies, and
process refunds. If you can't help, escalate to a human.
Always be empathetic and concise.
What this does: it hands one model four jobs with wildly different risk profiles and one vague escape hatch, so it treats a refund decision with the same confidence as a password reset.
On the surface this looks reasonable. It's polite, it lists capabilities, it mentions escalation. In production it produces a very specific failure pattern that support leads know well.
Why it fails
The prompt collapses four categories of task that should never share a decision boundary. Password resets and order lookups are deterministic — there's a right answer and a clear procedure. Policy questions are retrieval problems — the answer exists in a document and must be quoted, not improvised. Refund decisions are judgment calls with money attached. Escalation is a routing decision. One agent, one confidence level, applied to all four.
The measurable symptom is dangerous confidence on the retrieval and judgment tasks. The agent "knows" the refund policy the way it "knows" anything — as a plausible-sounding guess. It doesn't distinguish between "I retrieved this from the policy doc" and "this sounds like a reasonable policy." So it states invented policy with the same tone it uses for a genuine order status.
⚠️ Common mistake: adding "only answer from official policy" to the prompt and assuming it worked. Without a retrieval tool wired in and a separate agent whose only job is grounded answers, that instruction is a suggestion the model overrides whenever a fluent guess is available. The instruction and the capability are different things.
The escalation logic fails too, but in the opposite direction. Because escalation is framed as "if you can't help," the agent escalates when it's confused — which is precisely when it's mechanical tasks it should handle — and doesn't escalate when it's confidently wrong, which is exactly when a human is needed. The trigger is backwards.
After: the improved multi agent customer support design
The rebuild splits support into a router plus three specialists, each with a matched confidence posture and its own tools. The router classifies intent and hands off. It never answers.
ROUTES = {
,[object Object],: ,[object Object],, ,[object Object],
,[object Object],: ,[object Object],, ,[object Object],
,[object Object],: ,[object Object],, ,[object Object],
,[object Object],: ,[object Object],, ,[object Object],
}
,[object Object], ,[object Object],(,[object Object],):
system = (,[object Object],
,[object Object],
,[object Object],
,[object Object],
,[object Object],)
label = llm(system=system, user=msg).strip()
,[object Object], ROUTES.get(label, ,[object Object],)
,[object Object], ,[object Object],(,[object Object],):
doc = retrieve_policy(msg) ,[object Object],
system = (,[object Object],
,[object Object],
,[object Object],
,[object Object],)
,[object Object], llm(system=system, user=,[object Object],)What this does: it turns "answer everything" into "classify, then route to the agent whose posture matches the task," so the policy agent is physically constrained to quote a retrieved document instead of improvising.
The refund agent is where the design earns its keep. It never approves money on its own. It gathers the facts, checks them against explicit rules, and either proposes an action a human confirms or auto-approves only inside tight, pre-set limits.
[object Object], ,[object Object],(,[object Object],):
facts = gather_order_facts(ticket) ,[object Object],
system = (,[object Object],
,[object Object],
,[object Object],
,[object Object],
,[object Object],)
,[object Object], json.loads(llm(system=system, user=json.dumps(facts)))What this does: it converts a judgment call into a bounded decision with an explicit human-in-the-loop threshold, so the agent auto-resolves cheap clear-cut cases and escalates anything expensive or ambiguous — the reverse of the single agent's backwards trigger.
⚡ Pro tip: route by risk, not by topic. The instinct is to classify tickets by subject (billing, shipping, technical). Classify by what happens if the agent is wrong instead. A wrong order-status lookup costs a re-query; a wrong refund costs money and trust. Your router's job is to protect the expensive mistakes, so its categories should map to consequences.
Breaking down each element
The router's power comes from what it can't do: answer. By stripping its ability to respond directly, you remove the single agent's core failure — answering everything at one confidence level. A router that only classifies can't confidently invent a policy because it never speaks to the customer.
The policy agent's power comes from grounding. It's handed retrieved text and forbidden from going beyond it. This is the difference between a system that quotes your actual refund window and one that guesses "usually 30 days" because that sounds right.
The refund agent's power comes from the human threshold. It's not trying to be autonomous; it's trying to clear the easy volume and surface the rest. That framing — clear volume, surface judgment — is what makes support automation safe to deploy.
⚡ Pro tip: log every router decision with the message and the chosen route, then audit the "unknown -> human" bucket weekly. That bucket is your roadmap. The messages landing there repeatedly are the next specialist agent you should build. Most teams guess at what to automate next; the router's misclassification log tells you directly.
Variations for different contexts
A SaaS company weights the design toward the policy and account agents, since most tickets are how-do-I and billing questions. An e-commerce team weights it toward refund and shipping specialists, with a tighter auto-approve limit during peak returns season. A healthcare provider adds a compliance agent that reviews any response touching patient data before it sends, because the cost of a wrong answer is regulatory, not just annoyed customers.
The router stays the same in all three. Only the specialists and thresholds change. That's the reusable core: classify by risk, route to a matched posture, keep money and compliance behind a human threshold.
⚠️ Common mistake: letting the specialists talk to each other freely. A refund agent that can query the policy agent mid-decision sounds efficient and creates a loop where each agent trusts the other's guess. Keep handoffs one-directional through the router. If the refund agent needs policy, the router fetches it — the agents don't negotiate.
How do you measure whether it's working?
The single-agent bot fails silently — a wrong refund policy looks identical to a right one until a customer complains. A multi agent customer support system gives you measurement points a single agent can't, and using them is the difference between a system that improves and one that quietly rots.
Track three numbers. First, router accuracy: sample classified tickets weekly and check whether the route matched the true intent. A router drifting below ninety percent accuracy is silently sending refund questions to the account agent, and every downstream metric will lie to you until you fix it. Second, escalation precision: of the tickets escalated to humans, what fraction genuinely needed a person? Too high means your thresholds are too cautious and you're drowning humans in easy tickets; too low means the bot is resolving things it should have escalated. Third, grounded-answer rate for the policy agent: what fraction of its answers actually quote a retrieved source versus improvising? This is your fabrication guardrail, and it should sit near a hundred percent.
⚡ Pro tip: shadow-run the new multi-agent system against your existing bot before cutting over. Route real tickets to both, serve only the old bot's answers to customers, and compare. You'll see exactly where the multi-agent design disagrees with the single agent — and the disagreements are almost always cases where the single agent was confidently wrong. Cutting over on evidence beats cutting over on hope.
The migration pattern that avoids disaster is capability-by-capability, not all-at-once. Start by routing only account tickets — the deterministic, low-risk ones — through the new system while everything else stays on the old bot. Prove the router classifies correctly on the safe category, then add policy, then refund last. Refund goes last precisely because it's the highest-risk route, and you want maximum confidence in the router before money is on the line.
⚠️ Common mistake: launching all specialists at once and losing the ability to isolate a regression. If quality drops after a full launch, you can't tell which agent caused it. Rolling out one route at a time means any regression is attributable to the route you just added — a debugging property worth the slower rollout.
Save and reuse this
The router, policy, and refund prompts above are the skeleton of a working multi agent customer support system. The specifics — your thresholds, your route categories, your escalation rules — are what make it yours, and they're exactly the parts you'll tune repeatedly.
One more design note before you save anything: the router is the piece you'll tune most, because customer language shifts constantly. New products create new intents, seasonal patterns change the mix, and a phrase that meant "account help" last quarter might signal a refund this one. Treat the router prompt as a living document with a changelog, not a set-and-forget classifier. The specialists are comparatively stable; the router is where the ongoing maintenance lives.
Save these prompts as a versioned set in PromptABCD rather than pasting them into a config file that drifts. When you adjust the refund threshold or add a specialist, you want one source of truth your whole support team pulls from — not four slightly different copies quietly disagreeing about when to involve a human.
Continue Reading
Save the prompts from this post
PromptABCD is a free prompt manager. Paste, organize, and reuse your best AI prompts — no more hunting through chat history.
