How to Orchestrate Multiple AI Agents
Knowing how to orchestrate multiple AI agents is the difference between a collection of API calls and a working system. Here's how to build coordination logic that holds up in production.
from anthropic import Anthropic
client = Anthropic()
def run_agent(system_prompt: str, user_message: str, max_tokens: int = 1024) -> str:
response = client.messages.create(
model="claude-sonnet-4-6",
max_tokens=max_tokens,
system=system_prompt,
messages=[{"role": "user", "content": user_message}]
)
return response.content[0].text
def orchestrate(user_request: str) -> str:
# Step 1: Plan
plan = run_agent(
"You are a task planner. Break the user's request into 3-5 concrete subtasks. Return a numbered list.",
user_request
)
# Step 2: Execute
result = run_agent(
"You are an executor. Complete the subtasks listed, producing high-quality output for each.",
f"Original request: {user_request}\n\nSubtasks to complete:\n{plan}"
)
# Step 3: Validate
validated = run_agent(
"You are a quality reviewer. Check this output against the original request. Flag any gaps or errors.",
f"Original request: {user_request}\n\nOutput to review:\n{result}"
)
return validatedHow do you actually coordinate multiple AI agents without ending up with a spaghetti mess of API calls and conditional logic that nobody can debug three months later? That question comes up constantly once developers move past their first two-agent experiment.
The answer isn't a specific framework — it's a set of orchestration patterns that work regardless of which library you use. Understanding these patterns is what separates AI engineers who build systems that scale from those who rebuild the same system every six months.
This is a practical walkthrough. You'll leave with working code for the most common orchestration scenarios and a clear mental model for when to use each one.
Quick-Start: Orchestrate Multiple AI Agents in Five Lines
Before the deep dive, here's the minimal working orchestrator:
[object Object], anthropic ,[object Object], Anthropic
client = Anthropic()
,[object Object], ,[object Object],(,[object Object],) -> ,[object Object],:
response = client.messages.create(
model=,[object Object],,
max_tokens=max_tokens,
system=system_prompt,
messages=[{,[object Object],: ,[object Object],, ,[object Object],: user_message}]
)
,[object Object], response.content[,[object Object],].text
,[object Object], ,[object Object],(,[object Object],) -> ,[object Object],:
,[object Object],
plan = run_agent(
,[object Object],,
user_request
)
,[object Object],
result = run_agent(
,[object Object],,
,[object Object],
)
,[object Object],
validated = run_agent(
,[object Object],,
,[object Object],
)
,[object Object], validatedWhat this does: Three-agent pipeline — planner, executor, reviewer — assembled in plain Python. No framework required. The orchestrator is just function calls with string passing. This pattern handles a surprisingly large percentage of real-world use cases.
Understanding the Orchestration Variables
Before you wire agents together, you need to understand the four decisions that define every orchestrator:
Sequencing: What order do agents run in? Does Agent B need Agent A's output before it can start, or can they run simultaneously? Most orchestrators start sequential and add parallelism only where the task graph actually permits it.
Data passing: What information does each agent receive? At minimum: its subtask. Ideally: the original user request, relevant context from previous agents, and any constraints the user specified. Agents that only see the immediately preceding agent's output tend to drift from the user's original intent.
Error policy: What happens when an agent returns garbage, times out, or produces output that fails validation? A production orchestrator needs retry logic, fallback behaviors, and escalation paths. An orchestrator without error handling is a demo, not a system.
Termination: When does the orchestration end? For pipelines, it's when the last agent finishes. For loops with critic agents, you need an explicit stopping condition — usually a maximum iteration count or a quality threshold. Without a clear termination condition, you can end up in an infinite improvement loop that burns tokens indefinitely.
⚡ Pro tip: Keep a log of every agent's input, output, model used, latency, and timestamp. This isn't optional for production systems — it's how you diagnose which agent introduced an error when the final output is wrong. Log everything from day one, not after the first production incident.
Step-by-Step: Building a Production Orchestrator
Here's a complete orchestrator that handles sequencing, state management, error handling, and logging:
[object Object], time
,[object Object], json
,[object Object], anthropic ,[object Object], Anthropic
,[object Object], typing ,[object Object], ,[object Object],
client = Anthropic()
,[object Object], ,[object Object],:
,[object Object], ,[object Object],(,[object Object],):
,[object Object],.task_id = task_id
,[object Object],.state = {,[object Object],: ,[object Object],, ,[object Object],: {}, ,[object Object],: []}
,[object Object],.logs = []
,[object Object], ,[object Object],(,[object Object],) -> ,[object Object],:
,[object Object], attempt ,[object Object], ,[object Object],(max_retries + ,[object Object],):
start = time.time()
,[object Object],:
response = client.messages.create(
model=,[object Object],,
max_tokens=,[object Object],,
system=system,
messages=[{,[object Object],: ,[object Object],, ,[object Object],: message}]
)
output = response.content[,[object Object],].text
,[object Object],.logs.append({
,[object Object],: agent_name,
,[object Object],: attempt + ,[object Object],,
,[object Object],: ,[object Object],((time.time() - start) * ,[object Object],),
,[object Object],: ,[object Object],(output),
,[object Object],: ,[object Object],
})
,[object Object], output
,[object Object], Exception ,[object Object], e:
,[object Object],.state[,[object Object],].append(,[object Object],)
,[object Object], attempt == max_retries:
,[object Object],
time.sleep(,[object Object], ** attempt)
,[object Object], ,[object Object],(,[object Object],) -> ,[object Object],:
,[object Object],.state[,[object Object],] = ,[object Object],
,[object Object],
research = ,[object Object],.run_agent(
,[object Object],,
,[object Object],,
,[object Object],
)
,[object Object],.state[,[object Object],][,[object Object],] = research
,[object Object],
draft = ,[object Object],.run_agent(
,[object Object],,
,[object Object],,
,[object Object],
)
,[object Object],.state[,[object Object],][,[object Object],] = draft
,[object Object],
review = ,[object Object],.run_agent(
,[object Object],,
,[object Object],,
,[object Object],
)
,[object Object],.state[,[object Object],][,[object Object],] = review
,[object Object],.state[,[object Object],] = ,[object Object],
,[object Object], ,[object Object],.stateWhat this does: The orchestrator class handles state persistence, retry logic with exponential backoff, structured logging, and clean error propagation. Each agent receives both its immediate input and the original user request, preventing the intent drift that's one of the most common multi-agent failure modes.
⚡ Pro tip: The most important architectural decision in your orchestrator isn't which agents to include — it's whether to use code-based routing or LLM-based routing. Code-based orchestrators (if-else, routing tables, state machines) are deterministic, don't consume tokens for routing decisions, and are much easier to debug. Use an LLM as your orchestrator only when routing decisions are genuinely too complex for code — which is rarer than you'd think.
Parallel Execution for Independent Tasks
When subtasks don't depend on each other's outputs, parallel execution reduces wall-clock time from O(n) to O(1):
[object Object], concurrent.futures
,[object Object], ,[object Object],(,[object Object],) -> ,[object Object],[,[object Object],]:
,[object Object],
,[object Object], ,[object Object],(,[object Object],) -> ,[object Object],:
,[object Object],:
output = run_agent(task[,[object Object],], task[,[object Object],])
,[object Object], {,[object Object],: task[,[object Object],], ,[object Object],: output, ,[object Object],: ,[object Object],}
,[object Object], Exception ,[object Object], e:
,[object Object], {,[object Object],: task[,[object Object],], ,[object Object],: ,[object Object],, ,[object Object],: ,[object Object],(e), ,[object Object],: ,[object Object],}
,[object Object], concurrent.futures.ThreadPoolExecutor(max_workers=,[object Object],) ,[object Object], executor:
futures = [executor.submit(execute_task, task) ,[object Object], task ,[object Object], tasks]
results = [f.result() ,[object Object], f ,[object Object], concurrent.futures.as_completed(futures)]
,[object Object],
name_to_result = {r[,[object Object],]: r ,[object Object], r ,[object Object], results}
,[object Object], [name_to_result[task[,[object Object],]] ,[object Object], task ,[object Object], tasks]What this does: Uses Python's ThreadPoolExecutor to fire all agent calls simultaneously. Five competitor analyses that would take 50 seconds sequentially complete in roughly 10–12 seconds in parallel. The results are re-sorted to match input order so downstream logic doesn't have to handle non-deterministic ordering.
Troubleshooting Common Orchestration Issues
Agents diverging from the original request: The fix is always the same — pass the original user request to every agent in the chain, not just the first one. An agent at step 4 that only sees step 3's output is flying blind relative to what the user actually wanted.
Infinite loops in critic-generator pairs: Your critic needs an explicit exit condition. Either give it a maximum number of iterations (3 is usually enough for quality improvement), or have it return a structured signal like {"approved": true} when the output meets standards. Don't rely on the critic spontaneously deciding it's done.
Slow orchestrators: Profile before optimizing. Most orchestrator slowness comes from unnecessary sequential execution of tasks that could run in parallel. Mapping your task graph to identify independent subtasks usually reveals more parallelism than you expected.
Inconsistent output quality: If your orchestrator sometimes produces excellent output and sometimes mediocre output on similar inputs, the problem is almost always in your agent system prompts, not the orchestration logic. Add few-shot examples to the agents that are producing inconsistent results.
⚠️ Common mistake: Building error handling for agent failures while ignoring errors in the orchestrator's own routing logic. The orchestrator itself can produce wrong decisions — routing the wrong input to the wrong agent, skipping a validation step because a condition was written incorrectly. Orchestrator logic needs unit tests, not just end-to-end tests.
Pro-Level Variations
Dynamic agent selection: Instead of a fixed pipeline, have a planning agent emit a JSON list of agent names and inputs, then execute them in the specified order. This lets the orchestration path vary based on the input — a customer service request might need different agents than a technical support request.
Conditional fan-out: Route to different agent sets based on input classification. A document that contains code gets the code-review agent added to its pipeline. A document in a language other than English gets the translation agent inserted before analysis.
Confidence-weighted validation: Give your critic agent a structured output schema that includes a confidence score. If confidence is below a threshold, the orchestrator triggers a second generation pass. Above threshold, it passes immediately. This avoids always-run validation overhead on straightforward tasks.
Your Turn
Start by picking one workflow you're currently running through a single agent and apply the three-agent pipeline from the Quick-Start section. Measure output quality on 20 examples before and after the split. The improvement — or lack thereof — will tell you more than any theoretical framework.
State Management Between Agents
One of the underexamined challenges in multi-agent orchestration is state management: how do you track what each agent received and produced in a way that is inspectable and recoverable?
The minimal viable approach uses a dictionary that flows through the pipeline and is updated at each step:
[object Object], time
,[object Object], ,[object Object],(,[object Object],) -> ,[object Object],:
state = {
,[object Object],: initial_input,
,[object Object],: {},
,[object Object],: [],
,[object Object],: time.time()
}
research_output = research_agent(state[,[object Object],])
state[,[object Object],][,[object Object],] = research_output
analysis_output = analysis_agent(
research=research_output,
original_query=state[,[object Object],]
)
state[,[object Object],][,[object Object],] = analysis_output
final_output = synthesis_agent(
research=research_output,
analysis=analysis_output
)
state[,[object Object],][,[object Object],] = final_output
state[,[object Object],] = time.time()
,[object Object], stateWhat this does: Each agent receives only the inputs it needs — not the full state, which would bloat context windows unnecessarily — and the orchestrator writes outputs back to the state dict. When the pipeline produces unexpected output, you inspect state["agent_outputs"] to see exactly what each agent received and produced, without re-running the pipeline.
⚡ Pro tip: Store pipeline state to disk or a database before and after each agent run, not just at the end. If your orchestrator crashes during agent 4 of 6, you want to resume from agent 4, not restart. Persistent state makes multi-agent pipelines resilient to infrastructure failures and cuts debugging time significantly.
Handling Agent Failures Gracefully
Production orchestrators fail at the agent level more often than expected. Network timeouts, rate limits, and unexpectedly long responses that exceed token limits all produce agent failures. Handling them in the orchestrator — not by crashing — is essential for production systems:
[object Object], tenacity ,[object Object], retry, stop_after_attempt, wait_exponential
,[object Object],
,[object Object], ,[object Object],(,[object Object],) -> ,[object Object],:
response = client.messages.create(
model=,[object Object],,
max_tokens=,[object Object],,
system=system_prompt,
messages=[{,[object Object],: ,[object Object],, ,[object Object],: user_content}]
)
,[object Object], response.content[,[object Object],].text
,[object Object], ,[object Object],(,[object Object],) -> ,[object Object],:
results = {}
,[object Object],:
results[,[object Object],] = run_agent_with_retry(RESEARCHER_PROMPT, query, client)
,[object Object], Exception ,[object Object], e:
results[,[object Object],] = ,[object Object],
results[,[object Object],] = ,[object Object],(e)
,[object Object], results.get(,[object Object],):
,[object Object],:
results[,[object Object],] = run_agent_with_retry(
SYNTHESIS_PROMPT,
,[object Object],,
client
)
,[object Object], Exception ,[object Object], e:
results[,[object Object],] = ,[object Object],
,[object Object],:
results[,[object Object],] = ,[object Object],
,[object Object], resultsWhat this does: Each agent call uses exponential backoff retry for transient failures. If an agent fails after retries, the orchestrator logs the error and continues with partial results rather than crashing. Downstream agents receive graceful fallback content rather than exceptions.
⚡ Pro tip: Build your orchestrator's error handling before you build the happy path. It takes thirty minutes to write solid retry logic and graceful degradation. It takes days to retrofit it into a production system that crashes on the first agent timeout. Treating failure handling as an afterthought is one of the most expensive multi-agent engineering mistakes.
When you've dialed in the system prompts for each agent in your pipeline, save them to PromptABCD. ## Choosing Your First Orchestration Pattern
Two patterns cover the majority of real multi-agent orchestration use cases. Understanding which fits your task before writing any code saves significant rework.
Sequential with state passing is the right default for tasks with clear stage dependencies — research, then analyze, then write. Each agent in the chain receives the prior agent's output and adds its contribution. The pipeline is easy to debug, easy to test, and handles the 80% of cases where the task order is fixed.
Fan-out with aggregation is right for tasks with genuinely parallel subtasks — analyzing multiple documents simultaneously, researching multiple topics in parallel, or getting independent perspectives on the same input. Fan-out sends multiple agents the same (or similar) input concurrently. The aggregator collects all outputs and synthesizes them into a final result.
The mistake is defaulting to fan-out for sequential tasks (adding unnecessary parallelism complexity) or defaulting to sequential for independent tasks (introducing unnecessary latency). Match the pattern to the task's actual dependency structure.
Reusing battle-tested prompts across projects is how you build effective multi-agent systems faster over time, without starting the prompt engineering process from scratch each time.
Continue Reading
Save the prompts from this post
PromptABCD is a free prompt manager. Paste, organize, and reuse your best AI prompts — no more hunting through chat history.
