PromptABCD
FeaturesLearnGuideBlogContext Blocks
Sign inGet started free
Sign inSign up
PromptABCD

A calm home for your best AI prompts. Save them once, find them in seconds, reuse them forever.

Product

  • Features
  • Chrome Extension
  • Free Courses
  • How it works
  • Use cases
  • Blog
  • Context Blocks
  • Export Anywhere
  • FAQ

Resources

  • User guide
  • Learn prompting
  • Sign in
  • Get started free

© 2026 PromptABCD. All rights reserved.

AboutPrivacy PolicyTerms and Conditions
Home/Blog/Multi-Agent Systems/How to Orchestrate Multiple AI Agents
Multi-Agent Systems

How to Orchestrate Multiple AI Agents

Knowing how to orchestrate multiple AI agents is the difference between a collection of API calls and a working system. Here's how to build coordination logic that holds up in production.

September 21, 2026·11 min read
ShareShare
⚡Featured Prompt— copy and use right now
from anthropic import Anthropic

client = Anthropic()

def run_agent(system_prompt: str, user_message: str, max_tokens: int = 1024) -> str:
    response = client.messages.create(
        model="claude-sonnet-4-6",
        max_tokens=max_tokens,
        system=system_prompt,
        messages=[{"role": "user", "content": user_message}]
    )
    return response.content[0].text

def orchestrate(user_request: str) -> str:
    # Step 1: Plan
    plan = run_agent(
        "You are a task planner. Break the user's request into 3-5 concrete subtasks. Return a numbered list.",
        user_request
    )
    # Step 2: Execute
    result = run_agent(
        "You are an executor. Complete the subtasks listed, producing high-quality output for each.",
        f"Original request: {user_request}\n\nSubtasks to complete:\n{plan}"
    )
    # Step 3: Validate
    validated = run_agent(
        "You are a quality reviewer. Check this output against the original request. Flag any gaps or errors.",
        f"Original request: {user_request}\n\nOutput to review:\n{result}"
    )
    return validated

How do you actually coordinate multiple AI agents without ending up with a spaghetti mess of API calls and conditional logic that nobody can debug three months later? That question comes up constantly once developers move past their first two-agent experiment.

The answer isn't a specific framework — it's a set of orchestration patterns that work regardless of which library you use. Understanding these patterns is what separates AI engineers who build systems that scale from those who rebuild the same system every six months.

This is a practical walkthrough. You'll leave with working code for the most common orchestration scenarios and a clear mental model for when to use each one.

Quick-Start: Orchestrate Multiple AI Agents in Five Lines

Before the deep dive, here's the minimal working orchestrator:

python
[object Object], anthropic ,[object Object], Anthropic

client = Anthropic()

,[object Object], ,[object Object],(,[object Object],) -> ,[object Object],:
    response = client.messages.create(
        model=,[object Object],,
        max_tokens=max_tokens,
        system=system_prompt,
        messages=[{,[object Object],: ,[object Object],, ,[object Object],: user_message}]
    )
    ,[object Object], response.content[,[object Object],].text

,[object Object], ,[object Object],(,[object Object],) -> ,[object Object],:
    ,[object Object],
    plan = run_agent(
        ,[object Object],,
        user_request
    )
    ,[object Object],
    result = run_agent(
        ,[object Object],,
        ,[object Object],
    )
    ,[object Object],
    validated = run_agent(
        ,[object Object],,
        ,[object Object],
    )
    ,[object Object], validated

What this does: Three-agent pipeline — planner, executor, reviewer — assembled in plain Python. No framework required. The orchestrator is just function calls with string passing. This pattern handles a surprisingly large percentage of real-world use cases.

Understanding the Orchestration Variables

Before you wire agents together, you need to understand the four decisions that define every orchestrator:

Sequencing: What order do agents run in? Does Agent B need Agent A's output before it can start, or can they run simultaneously? Most orchestrators start sequential and add parallelism only where the task graph actually permits it.

Data passing: What information does each agent receive? At minimum: its subtask. Ideally: the original user request, relevant context from previous agents, and any constraints the user specified. Agents that only see the immediately preceding agent's output tend to drift from the user's original intent.

Error policy: What happens when an agent returns garbage, times out, or produces output that fails validation? A production orchestrator needs retry logic, fallback behaviors, and escalation paths. An orchestrator without error handling is a demo, not a system.

Termination: When does the orchestration end? For pipelines, it's when the last agent finishes. For loops with critic agents, you need an explicit stopping condition — usually a maximum iteration count or a quality threshold. Without a clear termination condition, you can end up in an infinite improvement loop that burns tokens indefinitely.

⚡ Pro tip: Keep a log of every agent's input, output, model used, latency, and timestamp. This isn't optional for production systems — it's how you diagnose which agent introduced an error when the final output is wrong. Log everything from day one, not after the first production incident.

Step-by-Step: Building a Production Orchestrator

Here's a complete orchestrator that handles sequencing, state management, error handling, and logging:

python
[object Object], time
,[object Object], json
,[object Object], anthropic ,[object Object], Anthropic
,[object Object], typing ,[object Object], ,[object Object],

client = Anthropic()

,[object Object], ,[object Object],:
    ,[object Object], ,[object Object],(,[object Object],):
        ,[object Object],.task_id = task_id
        ,[object Object],.state = {,[object Object],: ,[object Object],, ,[object Object],: {}, ,[object Object],: []}
        ,[object Object],.logs = []

    ,[object Object], ,[object Object],(,[object Object],) -> ,[object Object],:
        ,[object Object], attempt ,[object Object], ,[object Object],(max_retries + ,[object Object],):
            start = time.time()
            ,[object Object],:
                response = client.messages.create(
                    model=,[object Object],,
                    max_tokens=,[object Object],,
                    system=system,
                    messages=[{,[object Object],: ,[object Object],, ,[object Object],: message}]
                )
                output = response.content[,[object Object],].text
                ,[object Object],.logs.append({
                    ,[object Object],: agent_name,
                    ,[object Object],: attempt + ,[object Object],,
                    ,[object Object],: ,[object Object],((time.time() - start) * ,[object Object],),
                    ,[object Object],: ,[object Object],(output),
                    ,[object Object],: ,[object Object],
                })
                ,[object Object], output
            ,[object Object], Exception ,[object Object], e:
                ,[object Object],.state[,[object Object],].append(,[object Object],)
                ,[object Object], attempt == max_retries:
                    ,[object Object],
                time.sleep(,[object Object], ** attempt)

    ,[object Object], ,[object Object],(,[object Object],) -> ,[object Object],:
        ,[object Object],.state[,[object Object],] = ,[object Object],

        ,[object Object],
        research = ,[object Object],.run_agent(
            ,[object Object],,
            ,[object Object],,
            ,[object Object],
        )
        ,[object Object],.state[,[object Object],][,[object Object],] = research

        ,[object Object],
        draft = ,[object Object],.run_agent(
            ,[object Object],,
            ,[object Object],,
            ,[object Object],
        )
        ,[object Object],.state[,[object Object],][,[object Object],] = draft

        ,[object Object],
        review = ,[object Object],.run_agent(
            ,[object Object],,
            ,[object Object],,
            ,[object Object],
        )
        ,[object Object],.state[,[object Object],][,[object Object],] = review
        ,[object Object],.state[,[object Object],] = ,[object Object],

        ,[object Object], ,[object Object],.state

What this does: The orchestrator class handles state persistence, retry logic with exponential backoff, structured logging, and clean error propagation. Each agent receives both its immediate input and the original user request, preventing the intent drift that's one of the most common multi-agent failure modes.

⚡ Pro tip: The most important architectural decision in your orchestrator isn't which agents to include — it's whether to use code-based routing or LLM-based routing. Code-based orchestrators (if-else, routing tables, state machines) are deterministic, don't consume tokens for routing decisions, and are much easier to debug. Use an LLM as your orchestrator only when routing decisions are genuinely too complex for code — which is rarer than you'd think.

Parallel Execution for Independent Tasks

When subtasks don't depend on each other's outputs, parallel execution reduces wall-clock time from O(n) to O(1):

python
[object Object], concurrent.futures

,[object Object], ,[object Object],(,[object Object],) -> ,[object Object],[,[object Object],]:
    ,[object Object],
    ,[object Object], ,[object Object],(,[object Object],) -> ,[object Object],:
        ,[object Object],:
            output = run_agent(task[,[object Object],], task[,[object Object],])
            ,[object Object], {,[object Object],: task[,[object Object],], ,[object Object],: output, ,[object Object],: ,[object Object],}
        ,[object Object], Exception ,[object Object], e:
            ,[object Object], {,[object Object],: task[,[object Object],], ,[object Object],: ,[object Object],, ,[object Object],: ,[object Object],(e), ,[object Object],: ,[object Object],}

    ,[object Object], concurrent.futures.ThreadPoolExecutor(max_workers=,[object Object],) ,[object Object], executor:
        futures = [executor.submit(execute_task, task) ,[object Object], task ,[object Object], tasks]
        results = [f.result() ,[object Object], f ,[object Object], concurrent.futures.as_completed(futures)]

    ,[object Object],
    name_to_result = {r[,[object Object],]: r ,[object Object], r ,[object Object], results}
    ,[object Object], [name_to_result[task[,[object Object],]] ,[object Object], task ,[object Object], tasks]

What this does: Uses Python's ThreadPoolExecutor to fire all agent calls simultaneously. Five competitor analyses that would take 50 seconds sequentially complete in roughly 10–12 seconds in parallel. The results are re-sorted to match input order so downstream logic doesn't have to handle non-deterministic ordering.

Troubleshooting Common Orchestration Issues

Agents diverging from the original request: The fix is always the same — pass the original user request to every agent in the chain, not just the first one. An agent at step 4 that only sees step 3's output is flying blind relative to what the user actually wanted.

Infinite loops in critic-generator pairs: Your critic needs an explicit exit condition. Either give it a maximum number of iterations (3 is usually enough for quality improvement), or have it return a structured signal like {"approved": true} when the output meets standards. Don't rely on the critic spontaneously deciding it's done.

Slow orchestrators: Profile before optimizing. Most orchestrator slowness comes from unnecessary sequential execution of tasks that could run in parallel. Mapping your task graph to identify independent subtasks usually reveals more parallelism than you expected.

Inconsistent output quality: If your orchestrator sometimes produces excellent output and sometimes mediocre output on similar inputs, the problem is almost always in your agent system prompts, not the orchestration logic. Add few-shot examples to the agents that are producing inconsistent results.

⚠️ Common mistake: Building error handling for agent failures while ignoring errors in the orchestrator's own routing logic. The orchestrator itself can produce wrong decisions — routing the wrong input to the wrong agent, skipping a validation step because a condition was written incorrectly. Orchestrator logic needs unit tests, not just end-to-end tests.

Pro-Level Variations

Dynamic agent selection: Instead of a fixed pipeline, have a planning agent emit a JSON list of agent names and inputs, then execute them in the specified order. This lets the orchestration path vary based on the input — a customer service request might need different agents than a technical support request.

Conditional fan-out: Route to different agent sets based on input classification. A document that contains code gets the code-review agent added to its pipeline. A document in a language other than English gets the translation agent inserted before analysis.

Confidence-weighted validation: Give your critic agent a structured output schema that includes a confidence score. If confidence is below a threshold, the orchestrator triggers a second generation pass. Above threshold, it passes immediately. This avoids always-run validation overhead on straightforward tasks.

Your Turn

Start by picking one workflow you're currently running through a single agent and apply the three-agent pipeline from the Quick-Start section. Measure output quality on 20 examples before and after the split. The improvement — or lack thereof — will tell you more than any theoretical framework.

State Management Between Agents

One of the underexamined challenges in multi-agent orchestration is state management: how do you track what each agent received and produced in a way that is inspectable and recoverable?

The minimal viable approach uses a dictionary that flows through the pipeline and is updated at each step:

python
[object Object], time

,[object Object], ,[object Object],(,[object Object],) -> ,[object Object],:
    state = {
        ,[object Object],: initial_input,
        ,[object Object],: {},
        ,[object Object],: [],
        ,[object Object],: time.time()
    }
    
    research_output = research_agent(state[,[object Object],])
    state[,[object Object],][,[object Object],] = research_output
    
    analysis_output = analysis_agent(
        research=research_output,
        original_query=state[,[object Object],]
    )
    state[,[object Object],][,[object Object],] = analysis_output
    
    final_output = synthesis_agent(
        research=research_output,
        analysis=analysis_output
    )
    state[,[object Object],][,[object Object],] = final_output
    state[,[object Object],] = time.time()
    
    ,[object Object], state

What this does: Each agent receives only the inputs it needs — not the full state, which would bloat context windows unnecessarily — and the orchestrator writes outputs back to the state dict. When the pipeline produces unexpected output, you inspect state["agent_outputs"] to see exactly what each agent received and produced, without re-running the pipeline.

⚡ Pro tip: Store pipeline state to disk or a database before and after each agent run, not just at the end. If your orchestrator crashes during agent 4 of 6, you want to resume from agent 4, not restart. Persistent state makes multi-agent pipelines resilient to infrastructure failures and cuts debugging time significantly.

Handling Agent Failures Gracefully

Production orchestrators fail at the agent level more often than expected. Network timeouts, rate limits, and unexpectedly long responses that exceed token limits all produce agent failures. Handling them in the orchestrator — not by crashing — is essential for production systems:

python
[object Object], tenacity ,[object Object], retry, stop_after_attempt, wait_exponential

,[object Object],
,[object Object], ,[object Object],(,[object Object],) -> ,[object Object],:
    response = client.messages.create(
        model=,[object Object],,
        max_tokens=,[object Object],,
        system=system_prompt,
        messages=[{,[object Object],: ,[object Object],, ,[object Object],: user_content}]
    )
    ,[object Object], response.content[,[object Object],].text

,[object Object], ,[object Object],(,[object Object],) -> ,[object Object],:
    results = {}
    
    ,[object Object],:
        results[,[object Object],] = run_agent_with_retry(RESEARCHER_PROMPT, query, client)
    ,[object Object], Exception ,[object Object], e:
        results[,[object Object],] = ,[object Object],
        results[,[object Object],] = ,[object Object],(e)
    
    ,[object Object], results.get(,[object Object],):
        ,[object Object],:
            results[,[object Object],] = run_agent_with_retry(
                SYNTHESIS_PROMPT,
                ,[object Object],,
                client
            )
        ,[object Object], Exception ,[object Object], e:
            results[,[object Object],] = ,[object Object],
    ,[object Object],:
        results[,[object Object],] = ,[object Object],
    
    ,[object Object], results

What this does: Each agent call uses exponential backoff retry for transient failures. If an agent fails after retries, the orchestrator logs the error and continues with partial results rather than crashing. Downstream agents receive graceful fallback content rather than exceptions.

⚡ Pro tip: Build your orchestrator's error handling before you build the happy path. It takes thirty minutes to write solid retry logic and graceful degradation. It takes days to retrofit it into a production system that crashes on the first agent timeout. Treating failure handling as an afterthought is one of the most expensive multi-agent engineering mistakes.

When you've dialed in the system prompts for each agent in your pipeline, save them to PromptABCD. ## Choosing Your First Orchestration Pattern

Two patterns cover the majority of real multi-agent orchestration use cases. Understanding which fits your task before writing any code saves significant rework.

Sequential with state passing is the right default for tasks with clear stage dependencies — research, then analyze, then write. Each agent in the chain receives the prior agent's output and adds its contribution. The pipeline is easy to debug, easy to test, and handles the 80% of cases where the task order is fixed.

Fan-out with aggregation is right for tasks with genuinely parallel subtasks — analyzing multiple documents simultaneously, researching multiple topics in parallel, or getting independent perspectives on the same input. Fan-out sends multiple agents the same (or similar) input concurrently. The aggregator collects all outputs and synthesizes them into a final result.

The mistake is defaulting to fan-out for sequential tasks (adding unnecessary parallelism complexity) or defaulting to sequential for independent tasks (introducing unnecessary latency). Match the pattern to the task's actual dependency structure.

Reusing battle-tested prompts across projects is how you build effective multi-agent systems faster over time, without starting the prompt engineering process from scratch each time.


multi-agent-systemsagent-orchestrationai-agentsllm-engineeringpython-ai

Continue Reading

A Reusable Prompt Kit for Agent Teams
Multi-Agent Systems

A Reusable Prompt Kit for Agent Teams

A team rebuilt their agent prompts from memory every project, and every project drifted a little worse. That failure is why a multi agent prompt kit matters. Here's the reusable set of role prompts every team should keep.

October 2, 2026·9 min read
Multi-Agent Systems on a Budget: An Interactive Guide
Multi-Agent Systems

Multi-Agent Systems on a Budget: An Interactive Guide

Most multi-agent tutorials assume you'll burn tokens freely. That's wrong for anyone shipping on real constraints. A cheap multi agent system can match an expensive one with the right moves. Here's how to build one.

October 2, 2026·9 min read
The Manager Agent Anti-Pattern: A Teardown
Multi-Agent Systems

The Manager Agent Anti-Pattern: A Teardown

Why does your orchestrator agent keep becoming a bottleneck that mangles every handoff? You've hit the manager agent anti pattern — one agent trying to coordinate everything. Here's why it fails and what replaces it.

October 2, 2026·9 min read

Save the prompts from this post

PromptABCD is a free prompt manager. Paste, organize, and reuse your best AI prompts — no more hunting through chat history.

Start free →
← PreviousSingle Agent vs Multi-Agent: When to SplitNext →The Orchestrator-Worker Pattern
Share this post:
ShareShare