PromptABCD
FeaturesLearnGuideBlogContext Blocks
Sign inGet started free
Sign inSign up
PromptABCD

A calm home for your best AI prompts. Save them once, find them in seconds, reuse them forever.

Product

  • Features
  • Chrome Extension
  • Free Courses
  • How it works
  • Use cases
  • Blog
  • Context Blocks
  • Export Anywhere
  • FAQ

Resources

  • User guide
  • Learn prompting
  • Sign in
  • Get started free

© 2026 PromptABCD. All rights reserved.

AboutPrivacy PolicyTerms and Conditions
Home/Blog/Multi-Agent Systems/Supervisor Agents Explained
Multi-Agent Systems

Supervisor Agents Explained

Supervisor agents sound simple — one agent managing others — but most implementations fail for the same reason: the supervisor is asked to do too much reasoning and not enough deciding. Here's how to get it right.

September 21, 2026·10 min read
ShareShare
⚡Featured Prompt— copy and use right now
from anthropic import Anthropic
import json

client = Anthropic()

WORKER_DESCRIPTIONS = {
    "billing_agent": "Handles payment issues, invoice disputes, subscription changes, and refund requests.",
    "technical_agent": "Diagnoses and resolves product bugs, integration issues, API errors, and configuration problems.",
    "account_agent": "Manages account settings, user permissions, team administration, and enterprise contract questions."
}

def supervisor_agent(ticket: str, worker_results: dict = None) -> dict:
    context = f"Customer ticket:\n{ticket}"
    if worker_results:
        context += f"\n\nWork completed so far:\n{json.dumps(worker_results, indent=2)}"

    response = client.messages.create(
        model="claude-sonnet-4-6",
        max_tokens=512,
        system=f"""You are a support supervisor. You route tickets to specialist agents and evaluate their responses.

Available workers:
{json.dumps(WORKER_DESCRIPTIONS, indent=2)}

You must respond with a JSON object containing:
- "action": one of "route", "approve", "retry", "escalate"
- "worker": (if action is "route" or "retry") which worker to use
- "reason": brief explanation of your decision
- "ready_to_send": (if action is "approve") true/false

Evaluate whether the customer's full issue has been addressed before approving.""",
        messages=[{"role": "user", "content": context}]
    )
    return json.loads(response.content[0].text)

A startup tried to build a customer support automation system using a supervisor agent pattern. The supervisor's job: route incoming tickets to the right specialist agent (billing, technical support, or account management), then review the specialist's response before sending it to the customer.

The system worked for three days. Then a customer sent a ticket that was simultaneously a billing complaint and a technical issue. The supervisor, instructed to "route to the most appropriate specialist," chose billing. The billing agent answered the billing part, but the customer's technical problem — which was actually causing the billing issue — went unaddressed. The supervisor reviewed the billing response, saw that it correctly addressed the billing complaint, and approved it.

The customer escalated. The support team investigated and found that the supervisor agent pattern had a gap: the supervisor was optimized to route to one specialist, not to coordinate across multiple specialists when the problem required it.

This is the most common failure mode for supervisor agents, and it points to a specific design flaw: the supervisor's decision space was too narrow.

What is a Supervisor Agent?

A supervisor agent pattern is an architecture where one agent — the supervisor — manages the behavior of one or more worker agents. The supervisor is responsible for task decomposition, worker selection, output evaluation, and escalation decisions. Workers are responsible for executing assigned subtasks.

The pattern exists to solve a specific problem: complex tasks that require multiple specialists where the routing logic is too dynamic to express as fixed code. If your routing logic is a static sequence — always research, then write, then review — you want a code-based orchestrator, not a supervisor agent. The supervisor pattern adds value when the decision about which worker runs next requires reasoning that changes based on content or context.

python
[object Object], anthropic ,[object Object], Anthropic
,[object Object], json

client = Anthropic()

WORKER_DESCRIPTIONS = {
    ,[object Object],: ,[object Object],,
    ,[object Object],: ,[object Object],,
    ,[object Object],: ,[object Object],
}

,[object Object], ,[object Object],(,[object Object],) -> ,[object Object],:
    context = ,[object Object],
    ,[object Object], worker_results:
        context += ,[object Object],

    response = client.messages.create(
        model=,[object Object],,
        max_tokens=,[object Object],,
        system=,[object Object],,
        messages=[{,[object Object],: ,[object Object],, ,[object Object],: context}]
    )
    ,[object Object], json.loads(response.content[,[object Object],].text)

What this does: The supervisor receives the ticket and any work completed so far, then returns a structured decision rather than free-form text. The structured output — JSON with explicit action types — makes the supervisor's decisions interpretable and the orchestration logic downstream clean and testable.

Why Supervisor Agents Matter for Dynamic Workflows

The supervisor agent pattern solves a specific gap between static pipelines and fully autonomous agents.

Static pipelines are predictable but rigid. If every ticket needs the same three agents in the same order, a pipeline works fine. But what if some tickets need one agent, some need two, and occasionally one needs all three plus a human escalation?

Fully autonomous agents that decide their own next steps are flexible but unpredictable. They're hard to debug, hard to constrain, and tend to produce surprising behavior when they encounter input outside their training distribution.

The supervisor pattern occupies the middle ground: the supervisor reasons about coordination and sequencing, while workers execute specific, bounded tasks. The supervisor decides what work to do; the workers decide how to do their specific piece.

⚡ Pro tip: Give your supervisor explicit decision criteria, not just available options. Instead of "route to the most appropriate agent," write "route to billing_agent if the ticket mentions charges, payment, invoice, or subscription; route to technical_agent if the ticket describes an error, bug, or behavior that doesn't match expected; route to multiple agents if the ticket contains both." Specific criteria produce consistent routing.

The Core Components of a Working Supervisor

A supervisor agent that works in production has four specific responsibilities that must be explicitly represented in its system prompt:

Task analysis: Before routing, the supervisor needs to understand what kind of work the ticket or request actually requires. This is more than classification — it's identifying whether the work requires one specialist or several, and whether those specialists need to work sequentially or can work in parallel.

Worker selection: The supervisor must know what each worker is capable of and what their limitations are. System prompts that list worker names without describing worker capabilities produce random routing. Worker descriptions need to be specific enough that the supervisor can match task requirements to worker strengths.

Output evaluation: After workers produce results, the supervisor evaluates whether the customer's or user's full request has been addressed. This is different from evaluating output quality — the supervisor is checking for completeness, not style. Did the response actually answer what was asked?

Escalation policies: What does the supervisor do when workers fail, when tasks exceed worker capabilities, or when confidence is low? Supervisors without explicit escalation policies handle edge cases unpredictably. Define escalation thresholds explicitly.

python
[object Object], ,[object Object],(,[object Object],) -> ,[object Object],:
    system_prompts = {
        ,[object Object],: ,[object Object],,
        ,[object Object],: ,[object Object],,
        ,[object Object],: ,[object Object],
    }
    response = client.messages.create(
        model=,[object Object],,
        max_tokens=,[object Object],,
        system=system_prompts[worker_type],
        messages=[{
            ,[object Object],: ,[object Object],,
            ,[object Object],: ,[object Object],
        }]
    )
    ,[object Object], response.content[,[object Object],].text

,[object Object], ,[object Object],(,[object Object],) -> ,[object Object],:
    worker_results = {}
    iteration = ,[object Object],

    ,[object Object], iteration < max_iterations:
        decision = supervisor_agent(ticket, worker_results ,[object Object], worker_results ,[object Object], ,[object Object],)
        action = decision.get(,[object Object],)

        ,[object Object], action == ,[object Object],:
            worker = decision[,[object Object],]
            result = worker_agent(worker, ,[object Object],, ticket)
            worker_results[worker] = result
        ,[object Object], action == ,[object Object],:
            ,[object Object], {,[object Object],: ,[object Object],, ,[object Object],: worker_results}
        ,[object Object], action == ,[object Object],:
            ,[object Object], {,[object Object],: ,[object Object],, ,[object Object],: decision.get(,[object Object],), ,[object Object],: worker_results}
        ,[object Object], action == ,[object Object],:
            worker = decision[,[object Object],]
            result = worker_agent(worker, ,[object Object],, ticket)
            worker_results[,[object Object],] = result

        iteration += ,[object Object],

    ,[object Object], {,[object Object],: ,[object Object],, ,[object Object],: worker_results}

What this does: The supervisor loop runs until the supervisor approves, escalates, or hits the maximum iteration limit. The worker_results dictionary accumulates as agents complete their work, so the supervisor always has full context about what's been done before making its next decision.

Common Mistakes in Supervisor Agent Implementations

⚠️ Common mistake: Building a supervisor that can only route to one worker per ticket. Real customer support requests — and most complex business tasks — often require multiple specialists. Design your supervisor to support multi-worker coordination from the start: let it route to worker A, then evaluate that output, then route to worker B if needed, then synthesize for the final response.

Four other mistakes that appear regularly:

Not passing the original request to workers: Workers that only receive the supervisor's routing decision — not the original customer ticket — lack context for doing their job well. Always pass both the original input and the specific subtask to every worker.

Supervisor reviewing its own routing decisions: If the supervisor approved routing to worker A, it has a built-in bias toward approving worker A's output. Consider adding an independent review step outside the supervisor's responsibility for high-stakes outputs.

No visibility into supervisor reasoning: If the supervisor's decisions are black-box text outputs, you can't tell whether it's routing correctly without reading every ticket manually. Structured JSON output from the supervisor — with explicit action types and reasons — makes the system auditable.

Over-trusting the supervisor on edge cases: Supervisors are LLMs. They make mistakes, especially on input that falls outside the distribution they were tuned against. Treat supervisor escalations as ground truth — if the supervisor says "I don't know how to handle this," believe it and build a human handoff for that case.

⚡ Pro tip: Supervisors should use slightly higher temperature (0.3–0.5) than workers (0.0–0.2). Workers benefit from deterministic, consistent execution of their specific task. Supervisors benefit from slightly more flexible reasoning when matching novel task requirements to available worker capabilities. This is an underappreciated configuration difference between the two roles.

Scaling the Pattern

The supervisor agent pattern scales well to complex organizations of workers. A supervisor can manage five workers as easily as two, as long as its worker descriptions stay precise and its evaluation criteria stay specific.

Where the pattern struggles is when the supervisor itself becomes overloaded. A supervisor managing 15 different specialist workers, with complex routing logic, high evaluation standards, and real-time performance pressure, will degrade in quality. The fix is hierarchical supervision — a meta-supervisor that routes to department-level supervisors, each of which manages a smaller team of workers.

The supervisor agent pattern — properly implemented with clear decision criteria, structured outputs, explicit escalation policies, and independent review steps — handles a significant majority of real-world dynamic routing requirements without the unpredictability of fully autonomous agent systems.

For storing and maintaining the system prompts that define each worke## Testing Your Supervisor's Routing Logic

Supervisor routing is the most failure-prone component of a supervisor-worker system. When the supervisor routes incorrectly — sends a billing question to the technical agent, or routes to a single worker when two specialists are needed — the downstream agents produce confident, well-formatted responses to the wrong question. These failures are harder to detect than technical errors because the output looks reasonable.

Test routing logic explicitly before deploying, with a labeled test set:

python
[object Object], ,[object Object],():
    test_cases = [
        {
            ,[object Object],: ,[object Object],,
            ,[object Object],: [,[object Object],],
            ,[object Object],: [,[object Object],]
        },
        {
            ,[object Object],: ,[object Object],,
            ,[object Object],: [,[object Object],, ,[object Object],],
            ,[object Object],: []
        },
        {
            ,[object Object],: ,[object Object],,
            ,[object Object],: [,[object Object],],
            ,[object Object],: [,[object Object],, ,[object Object],]
        }
    ]
    
    passed = ,[object Object],
    ,[object Object], tc ,[object Object], test_cases:
        route = supervisor.route(tc[,[object Object],])
        
        correct = ,[object Object],(a ,[object Object], route ,[object Object], a ,[object Object], tc[,[object Object],])
        no_false_routes = ,[object Object],(a ,[object Object], ,[object Object], route ,[object Object], a ,[object Object], tc[,[object Object],])
        
        ,[object Object], correct ,[object Object], no_false_routes:
            passed += ,[object Object],
        ,[object Object],:
            ,[object Object],(,[object Object],)
            ,[object Object],(,[object Object],)
    
    ,[object Object],(,[object Object],)
    ,[object Object], passed == ,[object Object],(test_cases)

What this does: Each test case specifies the expected routing decision and which agents the supervisor should NOT route to. Testing negative routes — confirming the supervisor doesn't over-route — is as important as testing positive ones. A supervisor that routes everything to every agent is technically "never wrong" but defeats the purpose of specialization.

The Minimal Supervisor Architecture

If you're building a supervisor system for the first time, start with the minimum configuration: one supervisor, two workers, and a single routing criterion. The supervisor reads the task description and decides which worker handles it. Worker A handles one category. Worker B handles the other. The supervisor synthesizes their outputs or routes to only one, depending on the task.

This minimal configuration teaches you the real failure modes of supervisor routing — ambiguous inputs that could go to either worker, inputs that need both workers, inputs that match neither — without the complexity of a large worker pool. Once you understand how your supervisor handles edge cases on the minimal configuration, expand the worker pool based on the specific routing failures you observe.

⚡ Pro tip: Run supervisor routing tests on every prompt change, not just on initial deployment. A supervisor prompt update intended to improve handling of one ticket category can inadvertently change routing for another. Automated routing tests with a labeled dataset catch these regressions in minutes rather than in production.

r's specialization and the supervisor's routing criteria, PromptABCD makes it easy to version these prompts and iterate on them independently. The supervisor's routing criteria and each worker's specialty definition are the most important prompts in the system — they deserve careful version control.

multi-agent-systemssupervisor-agentagent-patternsai-agentsllm-orchestration

Continue Reading

A Reusable Prompt Kit for Agent Teams
Multi-Agent Systems

A Reusable Prompt Kit for Agent Teams

A team rebuilt their agent prompts from memory every project, and every project drifted a little worse. That failure is why a multi agent prompt kit matters. Here's the reusable set of role prompts every team should keep.

October 2, 2026·9 min read
Multi-Agent Systems on a Budget: An Interactive Guide
Multi-Agent Systems

Multi-Agent Systems on a Budget: An Interactive Guide

Most multi-agent tutorials assume you'll burn tokens freely. That's wrong for anyone shipping on real constraints. A cheap multi agent system can match an expensive one with the right moves. Here's how to build one.

October 2, 2026·9 min read
The Manager Agent Anti-Pattern: A Teardown
Multi-Agent Systems

The Manager Agent Anti-Pattern: A Teardown

Why does your orchestrator agent keep becoming a bottleneck that mangles every handoff? You've hit the manager agent anti pattern — one agent trying to coordinate everything. Here's why it fails and what replaces it.

October 2, 2026·9 min read

Save the prompts from this post

PromptABCD is a free prompt manager. Paste, organize, and reuse your best AI prompts — no more hunting through chat history.

Start free →
← PreviousThe Orchestrator-Worker PatternNext →Hierarchical Multi-Agent Systems
Share this post:
ShareShare