PromptABCD
FeaturesLearnGuideBlogContext Blocks
Sign inGet started free
Sign inSign up
PromptABCD

A calm home for your best AI prompts. Save them once, find them in seconds, reuse them forever.

Product

  • Features
  • Chrome Extension
  • Free Courses
  • How it works
  • Use cases
  • Blog
  • Context Blocks
  • Export Anywhere
  • FAQ

Resources

  • User guide
  • Learn prompting
  • Sign in
  • Get started free

© 2026 PromptABCD. All rights reserved.

AboutPrivacy PolicyTerms and Conditions
Home/Blog/Multi-Agent Systems/The Planner-Executor Split in Multi-Agent Systems
Multi-Agent Systems

The Planner-Executor Split in Multi-Agent Systems

Mixing planning and execution in the same agent is one of the most common sources of multi-agent failure. The planner-executor split gives each concern its own agent — and the difference in reliability is significant.

September 23, 2026·11 min read
ShareShare
⚡Featured Prompt— copy and use right now
import json
from anthropic import Anthropic

client = Anthropic()

PLANNER_SYSTEM = """You are a strategic planning agent. Your job is to decompose complex tasks into concrete, executable steps.

For each task, produce a structured plan with:
- A sequence of steps, each with an action type and specific parameters
- Dependencies between steps (which steps must complete before others begin)
- Expected output format for each step
- Fallback instructions if a step fails

Return your plan as JSON matching this schema:
{
  "objective": "string",
  "steps": [
    {
      "id": "step_1",
      "action": "web_search | code_execution | data_analysis | text_generation | api_call",
      "description": "What this step does",
      "inputs": {"param": "value"},
      "expected_output": "Description of expected result",
      "depends_on": [],
      "fallback": "What to do if this step fails"
    }
  ],
  "success_criteria": "How to know the objective was achieved"
}

Do not execute any steps. Only plan."""

def plan_task(task: str, available_tools: list[str]) -> dict:
    response = client.messages.create(
        model="claude-opus-4-5",
        max_tokens=2048,
        system=PLANNER_SYSTEM,
        messages=[{
            "role": "user",
            "content": f"Task: {task}\n\nAvailable tools: {', '.join(available_tools)}"
        }]
    )
    return json.loads(response.content[0].text)

Ask a single agent to both figure out how to solve a problem and then solve it, and you'll frequently see a well-known failure mode: the agent commits early to a plan and then executes it even when intermediate results suggest the plan is wrong. It conflates two distinct cognitive tasks — strategic reasoning about what to do and tactical execution of specific steps — and performance suffers on both.

The planner executor agents pattern separates these concerns into distinct roles. The planner reasons about the overall approach and produces a structured plan. Executors carry out individual steps in that plan. Neither agent does the other's job. The result is a system that's more reliable, more debuggable, and more adaptable when plans need revision.

Why the Split Matters

Planning and execution require different behaviors from an LLM:

Planning requires breadth. A planner needs to consider multiple approaches, anticipate obstacles, and reason about what resources are available. It benefits from a wider context window focused on the problem statement and available tools — not on the details of how to execute any specific step.

Execution requires depth. An executor needs to focus narrowly on one step, have access to specific tools, and produce precise output. It doesn't need to think about the overall strategy — it needs to do one thing well.

When the same agent handles both, it often does neither optimally. The planning gets shallow because the agent is already thinking about execution details. The execution gets sloppy because the agent is simultaneously managing strategic considerations.

Step 1: Build the Planner Agent

python
[object Object], json
,[object Object], anthropic ,[object Object], Anthropic

client = Anthropic()

PLANNER_SYSTEM = ,[object Object],

,[object Object], ,[object Object],(,[object Object],) -> ,[object Object],:
    response = client.messages.create(
        model=,[object Object],,
        max_tokens=,[object Object],,
        system=PLANNER_SYSTEM,
        messages=[{
            ,[object Object],: ,[object Object],,
            ,[object Object],: ,[object Object],
        }]
    )
    ,[object Object], json.loads(response.content[,[object Object],].text)

What this does: The planner receives a task description and the list of available tools. It returns a structured execution plan with explicit step dependencies, expected outputs, and fallback strategies. Critically, its system prompt ends with "Do not execute any steps. Only plan." — this instruction prevents the planner from attempting to execute steps in its response, keeping its role boundary clean.

⚡ Pro tip: Give the planner explicit knowledge of what each tool can and cannot do. A planner that doesn't know the limitations of its executors produces plans with steps that are impossible to execute. The cleaner the tool capability description, the fewer invalid plans you'll receive.

Step 2: Build the Executor Agent

python
EXECUTOR_SYSTEM = ,[object Object],

,[object Object], ,[object Object],(,[object Object],) -> ,[object Object],:
    ,[object Object],
    enriched_inputs = step[,[object Object],].copy()
    ,[object Object], dep_id ,[object Object], step.get(,[object Object],, []):
        ,[object Object], dep_id ,[object Object], prior_outputs:
            enriched_inputs[,[object Object],] = prior_outputs[dep_id][,[object Object],]
    
    step_context = {
        ,[object Object],: step[,[object Object],],
        ,[object Object],: step[,[object Object],],
        ,[object Object],: step[,[object Object],],
        ,[object Object],: enriched_inputs,
        ,[object Object],: step[,[object Object],]
    }
    
    response = client.messages.create(
        model=,[object Object],,
        max_tokens=,[object Object],,
        system=EXECUTOR_SYSTEM,
        messages=[{
            ,[object Object],: ,[object Object],,
            ,[object Object],: ,[object Object],
        }]
    )
    ,[object Object], json.loads(response.content[,[object Object],].text)

What this does: The executor receives one step at a time, enriched with the outputs of any steps it depends on. It executes only that step, returning a structured result with status and output. The system prompt explicitly prohibits the executor from attempting other steps or modifying the plan — keeping its scope narrow.

Step 3: The Orchestration Loop

python
[object Object], ,[object Object],(,[object Object],) -> ,[object Object],:
    plan = plan_task(task, available_tools)
    ,[object Object],(,[object Object],)
    
    completed = {}
    failed = []
    replan_count = ,[object Object],
    
    ,[object Object],
    ,[object Object], ,[object Object],(,[object Object],):
        ready = []
        ,[object Object], step ,[object Object], plan_steps:
            ,[object Object], step[,[object Object],] ,[object Object], completed_ids ,[object Object], step[,[object Object],] ,[object Object], failed_ids:
                ,[object Object],
            deps_met = ,[object Object],(d ,[object Object], completed_ids ,[object Object], d ,[object Object], step.get(,[object Object],, []))
            ,[object Object], deps_met:
                ready.append(step)
        ,[object Object], ready
    
    ,[object Object], ,[object Object],:
        ready = get_ready_steps(plan[,[object Object],], completed.keys(), failed)
        
        ,[object Object], ,[object Object], ready:
            ,[object Object],  ,[object Object],
        
        ,[object Object], step ,[object Object], ready:
            ,[object Object],(,[object Object],)
            result = execute_step(step, completed)
            
            ,[object Object], result[,[object Object],] == ,[object Object],:
                completed[step[,[object Object],]] = result
            ,[object Object],:
                failed.append(step[,[object Object],])
                ,[object Object],(,[object Object],)
                
                ,[object Object],
                ,[object Object], replan_count < max_replans:
                    replan_count += ,[object Object],
                    revised_task = ,[object Object],
                    plan = plan_task(revised_task, available_tools)
                    completed_ids_backup = ,[object Object],(completed.keys())
                    ,[object Object],
                    failed = []
                    ,[object Object],
    
    ,[object Object], {
        ,[object Object],: plan,
        ,[object Object],: ,[object Object],(completed),
        ,[object Object],: failed,
        ,[object Object],: replan_count,
        ,[object Object],: completed.get(plan[,[object Object],][-,[object Object],][,[object Object],], {}).get(,[object Object],)
    }

,[object Object],
result = run_planner_executor(
    task=,[object Object],,
    available_tools=[,[object Object],, ,[object Object],, ,[object Object],]
)
,[object Object],(result[,[object Object],])

What this does: The orchestration loop resolves step dependencies, executes ready steps in order, and handles failures by optionally replanning. When a step fails, the replan request includes the failure context and prior successful outputs so the planner can produce a better alternative — it doesn't start from scratch. The loop terminates when all steps are done, or when remaining steps are blocked by unresolved failures.

⚡ Pro tip: Implement parallel execution for independent steps. The dependency resolution logic above finds all "ready" steps — steps whose dependencies are met. Rather than running them one at a time, batch the ready steps and use asyncio.gather() to execute them concurrently. For tasks with natural parallelism (researching three topics simultaneously, for example), parallel executor calls reduce total latency significantly.

When to Use Planner-Executor vs Other Patterns

The planner executor agents split is most valuable when:

Tasks are complex enough to benefit from upfront planning. For three-step sequential tasks with known structure, explicit planning adds overhead with minimal benefit. For eight-plus-step tasks with conditional dependencies, upfront planning prevents the mid-execution confusion that happens when a single agent tries to think through next steps while also executing current steps.

The execution environment has real tools. Planner-executor systems shine when executors have access to web search, code execution, file I/O, or external APIs — capabilities where execution requires specific parameters and produces specific outputs. For tasks that are purely language generation, the split adds less value.

Failure recovery matters. The replanning loop handles failures gracefully because the planner has the context to produce a revised plan. A single agent that fails mid-execution has to restart or muddle through with degraded context.

Keeping Plans Aligned With Reality

The most common failure mode in planner executor agents systems is plan-reality divergence: the planner assumes tool capabilities that don't exist, or produces plans that depend on information only available after execution. Two practices reduce this:

First, give the planner explicit tool schemas — not just names, but what inputs each tool takes and what outputs it produces. A planner that knows web_search returns a list of snippets, not full page text, will not produce steps that assume full page content.

Second, inject execution feedback into replanning requests. When you replan after a failure, the revised plan should include what actually happened in prior steps — not what the original plan expected to happen.

⚠️ Common mistake: Using the same model for planning and execution without adjusting temperature. Planners benefit from slightly higher temperature (0.3–0.5) to explore the solution space more broadly. Executors benefit from lower temperature (0.0–0.1) to produce precise, consistent output. The same model can fill both roles well if temperature is set appropriately for each role.

Building a Reusable Planner-Executor Library

The planner and executor system prompts above are reusable across task types with minor modifications to the tool list and output schema. Once you've tuned prompts that produce reliable plans and clean executor outputs for your domain, those prompts are your most valuable infrastructure asset.

Handling Plan Failures and Partial Completion

The planner-executor system's replanning loop handles failures, but the replan strategy matters. Two approaches:

Replan from current state. The planner receives the current state — what succeeded, what failed — and produces a revised plan that works around the failures. This is the most flexible approach and handles cases where the original plan's assumptions were wrong.

Replan only the failed step. Rather than replanning the entire pipeline, the planner receives just the failed step and its context and produces an alternative approach for that step. This is faster and appropriate when the overall plan structure is sound but one specific step needs a different execution strategy.

python
[object Object], ,[object Object],(,[object Object],) -> ,[object Object],:
    ,[object Object],
    
    context = {
        ,[object Object],: original_step,
        ,[object Object],: failure_reason,
        ,[object Object],: ,[object Object],(prior_outputs.keys()),
        ,[object Object],: {
            k: ,[object Object],(v[,[object Object],])[:,[object Object],]  ,[object Object],
            ,[object Object], k, v ,[object Object], prior_outputs.items()
        }
    }
    
    response = client.messages.create(
        model=,[object Object],,
        max_tokens=,[object Object],,
        system=(
            ,[object Object],
            ,[object Object],
            ,[object Object],
            ,[object Object],
        ),
        messages=[{,[object Object],: ,[object Object],, ,[object Object],: ,[object Object],}]
    )
    ,[object Object], json
    data = json.loads(response.content[,[object Object],].text)
    alt = data[,[object Object],]
    alt[,[object Object],] = original_step[,[object Object],]  ,[object Object],
    alt[,[object Object],] = original_step.get(,[object Object],, [])
    ,[object Object], alt

What this does: The step-level replanner receives only the failed step and its context, not the full original plan. It produces an alternative execution approach for that specific step — trying a different tool, a different query strategy, or a simplified objective that can be achieved with available information. This is faster than full replanning and preserves all the work completed before the failure.

⚡ Pro tip: Log the replan trigger and the revised step for every replan event. After a week of production operation, review the replan log. Steps that trigger replanning frequently are candidates for improvement — either the planner consistently underspecifies them, or the executor consistently fails on them, or the tool they depend on is unreliable. Frequent replans on the same step are a signal to address the root cause rather than continuing to rely on the replanning mechanism.

Planner Context Window Management

The planner's context window is a resource to manage carefully. As the plan execution proceeds and prior step outputs accumulate, the replanning context can grow large. Large contexts slow the planner and can dilute its focus on the current failure.

For long pipelines, summarize prior step outputs before passing them to the planner rather than passing the full outputs. A concise summary of what five steps produced ("research found X, analysis identified Y, formatting produced Z") gives the planner sufficient context without flooding its input with full output text. The planner needs to understand what was accomplished — not the full content of each step's output.

Establish a maximum context budget for planner calls: prior step summaries capped at 200 tokens each, current step description capped at 500 tokens, tool list capped at 100 tokens. Staying within a consistent budget keeps planner latency predictable across different pipeline lengths.

When Planner-Executor Is Overkill

The planner-executor split earns its overhead for complex multi-step tasks with many tools and possible failure paths. For simpler tasks, the overhead of explicit planning may exceed its value.

A five-step task with a fixed, known sequence doesn't benefit from dynamic planning. The executor can follow a hardcoded sequence without a planner's involvement. Three-step pipelines with no failure modes that require replanning are better served by a simple sequential orchestrator than a full planner-executor system.

Use planner-executor when: the task has eight or more steps, the step sequence depends on intermediate outputs, tool failures are common enough to require adaptive replanning, or the task type varies significantly across runs. For predictable, short pipelines, explicit planning adds latency without adding flexibility. The right architecture is always the simplest one that reliably solves the problem — and for many tasks, that's a sequential script, not a planner-executor system. The planner executor agents pattern is a tool for a specific class of problem, not a template for all agentic work.

Store them in PromptABCD with version tracking. When you modify the planner system prompt to support a new tool type, keep the prior version available — the old version may work better for certain task categories than the new one.

multi-agent-systemsagent-patternsplanner-executorllm-engineeringagent-design

Continue Reading

A Reusable Prompt Kit for Agent Teams
Multi-Agent Systems

A Reusable Prompt Kit for Agent Teams

A team rebuilt their agent prompts from memory every project, and every project drifted a little worse. That failure is why a multi agent prompt kit matters. Here's the reusable set of role prompts every team should keep.

October 2, 2026·9 min read
Multi-Agent Systems on a Budget: An Interactive Guide
Multi-Agent Systems

Multi-Agent Systems on a Budget: An Interactive Guide

Most multi-agent tutorials assume you'll burn tokens freely. That's wrong for anyone shipping on real constraints. A cheap multi agent system can match an expensive one with the right moves. Here's how to build one.

October 2, 2026·9 min read
The Manager Agent Anti-Pattern: A Teardown
Multi-Agent Systems

The Manager Agent Anti-Pattern: A Teardown

Why does your orchestrator agent keep becoming a bottleneck that mangles every handoff? You've hit the manager agent anti pattern — one agent trying to coordinate everything. Here's why it fails and what replaces it.

October 2, 2026·9 min read

Save the prompts from this post

PromptABCD is a free prompt manager. Paste, organize, and reuse your best AI prompts — no more hunting through chat history.

Start free →
← PreviousBlackboard Architecture for AI AgentsNext →Specialist vs Generalist Agents in Multi-Agent Systems
Share this post:
ShareShare