PromptABCD
FeaturesLearnGuideBlogContext Blocks
Sign inGet started free
Sign inSign up
PromptABCD

A calm home for your best AI prompts. Save them once, find them in seconds, reuse them forever.

Product

  • Features
  • Chrome Extension
  • Free Courses
  • How it works
  • Use cases
  • Blog
  • Context Blocks
  • Export Anywhere
  • FAQ

Resources

  • User guide
  • Learn prompting
  • Sign in
  • Get started free

© 2026 PromptABCD. All rights reserved.

AboutPrivacy PolicyTerms and Conditions
Home/Blog/Multi-Agent Systems/Message Passing Between AI Agents
Multi-Agent Systems

Message Passing Between AI Agents

How you pass messages between AI agents determines whether your multi-agent system is debuggable, reliable, and scalable — or a tangle of string-parsing logic that breaks under pressure.

September 21, 2026·11 min read
ShareShare
⚡Featured Prompt— copy and use right now
# The fragile approach that fails silently in production
from anthropic import Anthropic

client = Anthropic()

def bad_message_pipeline(user_request: str) -> str:
    # No schema validation, no error handling, no delivery confirmation
    agent_a_output = client.messages.create(
        model="claude-sonnet-4-6",
        max_tokens=1024,
        system="You are an analyst. Analyze the request.",
        messages=[{"role": "user", "content": user_request}]
    ).content[0].text

    # Agent B silently receives garbage if Agent A failed
    agent_b_output = client.messages.create(
        model="claude-sonnet-4-6",
        max_tokens=1024,
        system="You are a writer. Write a report based on the analysis.",
        messages=[{"role": "user", "content": agent_a_output}]  # Could be an error message
    ).content[0].text

    return agent_b_output

How does your multi-agent system handle a message that no agent can process? That question might seem academic until the night a customer's request falls between the cracks of your agent routing logic — consumed by the pipeline, acknowledged by nobody, and answered by silence.

That's what happened to a fintech team building an automated financial report system. Their orchestrator passed a task to an analysis agent. The analysis agent failed on a malformed data input. The error was swallowed by a bare except clause in the orchestrator. The downstream report-writing agent received an empty string as its input, generated a report that was entirely fabricated, and the fabricated report made it to a client.

The root cause wasn't the agent failure — it was the message passing architecture. There was no dead letter queue, no delivery confirmation, and no schema validation between agents. A failed message was invisible.

The Problem: Invisible Message Failures

Agent message passing fails in production for one of three reasons: the message content is invalid, the receiving agent is unavailable, or the receiving agent's response doesn't meet the expected schema. Most multi-agent tutorials handle none of these cases.

python
[object Object],
,[object Object], anthropic ,[object Object], Anthropic

client = Anthropic()

,[object Object], ,[object Object],(,[object Object],) -> ,[object Object],:
    ,[object Object],
    agent_a_output = client.messages.create(
        model=,[object Object],,
        max_tokens=,[object Object],,
        system=,[object Object],,
        messages=[{,[object Object],: ,[object Object],, ,[object Object],: user_request}]
    ).content[,[object Object],].text

    ,[object Object],
    agent_b_output = client.messages.create(
        model=,[object Object],,
        max_tokens=,[object Object],,
        system=,[object Object],,
        messages=[{,[object Object],: ,[object Object],, ,[object Object],: agent_a_output}]  ,[object Object],
    ).content[,[object Object],].text

    ,[object Object], agent_b_output

What this does: This works in the happy path. When Agent A fails with a network error, times out, or returns a truncated response, Agent B receives that broken content as if it were valid analysis. Agent B has no way to distinguish a real analysis from an error message, so it writes a report based on whatever it received.

⚠️ Common mistake: Adding error handling to individual agent calls but not to the messages being passed between them. You can catch API errors without ever catching the case where the API succeeds but returns content that's semantically invalid for the next stage of your pipeline. Both kinds of failure need handling.

The Wrong Approach: String-Chained Pipelines

The most common poor approach to agent message passing is direct string chaining: one agent's full text output becomes the next agent's full text input, with no structure, validation, or acknowledgment in between.

This breaks as soon as:

  • An agent's output is longer than expected and inflates the next agent's context
  • An agent's output uses a format the next agent doesn't know how to interpret
  • A message needs to be delivered to multiple agents, some of which might fail
  • You need to debug which message introduced an error in a long pipeline

The Correct Approach: Structured Message Passing With Validation

python
[object Object], json
,[object Object], time
,[object Object], dataclasses ,[object Object], dataclass, field, asdict
,[object Object], typing ,[object Object], ,[object Object],, ,[object Object],
,[object Object], anthropic ,[object Object], Anthropic

client = Anthropic()

,[object Object],
,[object Object], ,[object Object],:
    ,[object Object],
    sender: ,[object Object],
    recipient: ,[object Object],
    task_id: ,[object Object],
    payload: ,[object Object],
    schema_version: ,[object Object], = ,[object Object],
    sent_at: ,[object Object], = field(default_factory=time.time)
    retry_count: ,[object Object], = ,[object Object],
    max_retries: ,[object Object], = ,[object Object],

,[object Object],
,[object Object], ,[object Object],:
    ,[object Object],
    sender: ,[object Object],
    original_task_id: ,[object Object],
    success: ,[object Object],
    payload: ,[object Object],
    error: ,[object Object],[,[object Object],] = ,[object Object],
    completed_at: ,[object Object], = field(default_factory=time.time)

,[object Object], ,[object Object],:
    ,[object Object], ,[object Object],(,[object Object],):
        ,[object Object],.dead_letters: ,[object Object],[AgentMessage] = []
        ,[object Object],.delivery_log: ,[object Object],[,[object Object],] = []

    ,[object Object], ,[object Object],(,[object Object],) -> AgentResponse:
        ,[object Object],
        attempt = ,[object Object],
        last_error = ,[object Object],

        ,[object Object], attempt <= message.max_retries:
            ,[object Object],:
                response = handler(message)
                ,[object Object],.delivery_log.append({
                    ,[object Object],: message.task_id,
                    ,[object Object],: message.sender,
                    ,[object Object],: message.recipient,
                    ,[object Object],: ,[object Object],,
                    ,[object Object],: attempt + ,[object Object],
                })
                ,[object Object], response
            ,[object Object], Exception ,[object Object], e:
                last_error = ,[object Object],(e)
                attempt += ,[object Object],
                message.retry_count = attempt
                ,[object Object], attempt <= message.max_retries:
                    time.sleep(,[object Object], ** attempt)

        ,[object Object],
        ,[object Object],.dead_letters.append(message)
        ,[object Object],.delivery_log.append({
            ,[object Object],: message.task_id,
            ,[object Object],: message.recipient,
            ,[object Object],: ,[object Object],,
            ,[object Object],: last_error,
            ,[object Object],: ,[object Object],
        })
        ,[object Object], AgentResponse(
            sender=message.recipient,
            original_task_id=message.task_id,
            success=,[object Object],,
            payload={},
            error=,[object Object],
        )

,[object Object], ,[object Object],(,[object Object],) -> AgentResponse:
    ,[object Object],
    ,[object Object], ,[object Object], ,[object Object], ,[object Object], message.payload:
        ,[object Object], ValueError(,[object Object],)

    response = client.messages.create(
        model=,[object Object],,
        max_tokens=,[object Object],,
        system=,[object Object],,
        messages=[{,[object Object],: ,[object Object],, ,[object Object],: ,[object Object],}]
    )

    ,[object Object],:
        result = json.loads(response.content[,[object Object],].text)
        ,[object Object],
        required = [,[object Object],, ,[object Object],, ,[object Object],, ,[object Object],]
        ,[object Object], ,[object Object], ,[object Object],(k ,[object Object], result ,[object Object], k ,[object Object], required):
            ,[object Object], ValueError(,[object Object],)

        ,[object Object], AgentResponse(
            sender=,[object Object],,
            original_task_id=message.task_id,
            success=,[object Object],,
            payload=result
        )
    ,[object Object], (json.JSONDecodeError, ValueError) ,[object Object], e:
        ,[object Object], ValueError(,[object Object],)

,[object Object],
bus = MessageBus()

analysis_msg = AgentMessage(
    sender=,[object Object],,
    recipient=,[object Object],,
    task_id=,[object Object],,
    payload={,[object Object],: ,[object Object],}
)

result = bus.deliver(analysis_msg, analysis_handler)

,[object Object], result.success:
    ,[object Object],(,[object Object],)
,[object Object],:
    ,[object Object],(,[object Object],)
    ,[object Object],(,[object Object],)

What this does: Every message has a typed envelope with sender, recipient, task ID, schema version, and retry tracking. The message bus handles retries with exponential backoff and routes failed messages to a dead letter queue rather than silently dropping them. Both the message payload and the agent response payload are validated against expected schemas before the pipeline continues.

Results and What Changed

When the fintech team implemented structured agent message passing with dead letter queues:

Invisible failures became visible: The dead letter queue caught 3–5 failed messages per 1,000 processed that the original system had silently swallowed. Each of those was a case where fabricated output had previously reached clients.

Debugging became tractable: Task IDs in every message envelope meant that when a report was wrong, the team could trace exactly which message failed and which agent was responsible. Mean debugging time dropped from 4 hours to 20 minutes.

Retry logic reduced errors overall: 60% of agent failures were transient — network timeouts, brief API rate limits, model context overflows. Exponential backoff retry handling meant those failures resolved without human intervention rather than triggering dead letters.

⚡ Pro tip: The least obvious benefit of structured message passing is what it does for your monitoring. When every inter-agent message is logged with a task ID, sender, recipient, and timestamp, you can build a visual trace of every task through your system. That traceability is worth more than the schema validation in most production environments.

How to Apply This to Your Situation

Start with the three-question diagnostic for your current message passing setup:

Do your agents validate incoming messages? If Agent B just takes whatever Agent A returned and tries to use it, you have no protection against Agent A's failures propagating as wrong data rather than explicit errors. Add input validation to every agent that receives messages from other agents.

What happens when a message fails delivery? If the answer is "I don't know" or "it silently fails," implement a dead letter queue. It doesn't need to be a full message queue system — a simple list that captures failed messages with their error context is a dramatic improvement over silent failures.

Can you trace a specific message through your pipeline? If not, add task IDs to your messages today. A unique identifier that follows a request through the entire system is the foundation of debugging, auditing, and monitoring in production agent systems.

Next Steps

Once you have structured message passing in place, the natural extension is schema evolution. As your agents improve, their output schemas change. Add schema_version to your message envelopes from day one — even if you only have one version — so that version negotiation between agents with different schema expectations becomes manageable rather than a breaking change event.

Advanced Message Routing Patterns

Production agent message passing systems require routing logic that handles more than simple point-to-point delivery. Three routing patterns appear repeatedly in production multi-agent systems:

Topic-based routing delivers messages to all agents subscribed to a particular message category. When a research agent publishes a "findings_ready" message, all agents that need research findings receive it — the sender doesn't need to know who the recipients are.

Content-based routing inspects the message payload and routes based on field values. An orchestrator that reads message.priority == "urgent" and routes to an expedited pipeline, while routing message.priority == "normal" to a standard pipeline, is using content-based routing.

Dead letter queues capture messages that fail to be delivered or processed. Rather than silently dropping failed messages, a dead letter queue stores them for inspection. This is essential for debugging: when a downstream agent rejects a message (wrong format, missing required field), the dead letter queue preserves the message so you can understand what went wrong.

python
[object Object], json
,[object Object], dataclasses ,[object Object], dataclass, field
,[object Object], typing ,[object Object], ,[object Object],, ,[object Object],
,[object Object], enum ,[object Object], Enum

,[object Object], ,[object Object],(,[object Object],):
    LOW = ,[object Object],
    NORMAL = ,[object Object],
    HIGH = ,[object Object],
    URGENT = ,[object Object],

,[object Object],
,[object Object], ,[object Object],:
    task_id: ,[object Object],
    sender: ,[object Object],
    recipient: ,[object Object],
    message_type: ,[object Object],
    payload: ,[object Object],
    priority: MessagePriority = MessagePriority.NORMAL
    correlation_id: ,[object Object],[,[object Object],] = ,[object Object],  ,[object Object],
    reply_to: ,[object Object],[,[object Object],] = ,[object Object],        ,[object Object],

,[object Object], ,[object Object],:
    ,[object Object], ,[object Object],(,[object Object],):
        ,[object Object],.routes: ,[object Object], = {}
        ,[object Object],.dead_letter_queue: ,[object Object], = []
        ,[object Object],.message_log: ,[object Object], = []
    
    ,[object Object], ,[object Object],(,[object Object],):
        ,[object Object],.routes[,[object Object],] = handler
    
    ,[object Object], ,[object Object],(,[object Object],) -> ,[object Object],:
        route_key = ,[object Object],
        handler = ,[object Object],.routes.get(route_key)
        
        ,[object Object],.message_log.append({
            ,[object Object],: message.task_id,
            ,[object Object],: message.sender,
            ,[object Object],: message.recipient,
            ,[object Object],: message.message_type,
            ,[object Object],: handler ,[object Object], ,[object Object], ,[object Object],
        })
        
        ,[object Object], ,[object Object], handler:
            ,[object Object],.dead_letter_queue.append(message)
            ,[object Object],(,[object Object],)
            ,[object Object], ,[object Object],
        
        ,[object Object],:
            handler(message)
            ,[object Object], ,[object Object],
        ,[object Object], Exception ,[object Object], e:
            message.payload[,[object Object],] = ,[object Object],(e)
            ,[object Object],.dead_letter_queue.append(message)
            ,[object Object], ,[object Object],

router = MessageRouter()

,[object Object], ,[object Object],(,[object Object],):
    findings = msg.payload.get(,[object Object],, [])
    ,[object Object],(,[object Object],)
    ,[object Object],

router.register_handler(,[object Object],, ,[object Object],, handle_research_result)

,[object Object],
msg = AgentMessage(
    task_id=,[object Object],,
    sender=,[object Object],,
    recipient=,[object Object],,
    message_type=,[object Object],,
    payload={,[object Object],: [,[object Object],, ,[object Object],], ,[object Object],: ,[object Object],},
    priority=MessagePriority.NORMAL,
    correlation_id=,[object Object],
)

success = router.route(msg)
,[object Object],(,[object Object],)
,[object Object],(,[object Object],)

What this does: The MessageRouter maps message type plus recipient to a handler function. Messages that don't have a registered handler go to the dead letter queue rather than disappearing silently. The message log captures every routing attempt, providing a complete audit trail even when routing fails. The correlation_id field enables tracking request-response pairs across multiple agents.

⚡ Pro tip: Set a dead letter queue alert threshold. When the dead letter queue exceeds N messages (five is a reasonable starting point), log a warning or trigger a monitoring alert. A growing dead letter queue in production is an early warning signal of a schema mismatch or routing configuration error — catching it early prevents compounding failures.

Message Replay for Debugging

One of the most valuable properties of structured agent message passing is that it enables message replay: re-running a failing pipeline from a known-good checkpoint by replaying the messages from that point forward.

When a pipeline fails at step 4 of 6, message replay means you don't have to re-run steps 1-3 (which may involve expensive tool calls or slow LLM calls). You load the step 3 message from your message log and replay it into step 4.

This requires that all inter-agent messages are stored durably before delivery, not just logged after. The storage-before-delivery pattern is more complex to implement but makes production debugging dramatically faster.

⚡ Pro tip: Add a schema_version field to every message type. As your message schemas evolve, versioning ensures that replay of older messages still works — the handler can deserialize based on the schema version in the message, rather than assuming all messages use the current schema. This one field prevents an entire class of replay failures.

Store your agent message schemas and the system prompts that produce them in PromptABCD. A well-defined schema paired with the system prompt that reliably generates it is one of the most valuable reusable artifacts in multi-agent development.

Message Batching for High-Volume Systems

At high throughput — hundreds of messages per minute between agents — individual message delivery becomes a bottleneck. Message batching collects multiple messages and delivers them in a single call, reducing the overhead of per-message processing:

python
[object Object], time
,[object Object], collections ,[object Object], defaultdict

,[object Object], ,[object Object],:
    ,[object Object], ,[object Object],(,[object Object],):
        ,[object Object],.batches: ,[object Object], = defaultdict(,[object Object],)
        ,[object Object],.batch_size = batch_size
        ,[object Object],.flush_interval = flush_interval
        ,[object Object],.last_flush = time.time()
    
    ,[object Object], ,[object Object],(,[object Object],) -> ,[object Object],:
        ,[object Object],.batches[recipient].append(message)
        flushed = []
        
        ,[object Object],
        ,[object Object], (,[object Object],(,[object Object],.batches[recipient]) >= ,[object Object],.batch_size ,[object Object], 
                time.time() - ,[object Object],.last_flush > ,[object Object],.flush_interval):
            flushed = ,[object Object],.flush(recipient)
        
        ,[object Object], flushed
    
    ,[object Object], ,[object Object],(,[object Object],) -> ,[object Object],:
        batch = ,[object Object],.batches.pop(recipient, [])
        ,[object Object],.last_flush = time.time()
        ,[object Object], batch

What this does: Messages for each recipient accumulate in a queue. When the batch reaches the maximum size or the flush interval elapses, the batch is delivered to the recipient in one call. The recipient agent processes all messages in the batch, reducing per-message overhead at the cost of slightly increased latency on the first messages in a batch. For most multi-agent systems, message batching is unnecessary until throughput exceeds a few hundred messages per minute — optimize for clarity first and add batching only when profiling shows message delivery overhead is a measurable bottleneck. The queue's flush_interval parameter controls the latency-throughput tradeoff: shorter intervals reduce latency at the cost of smaller average batch sizes; longer intervals improve throughput at the cost of higher latency for individual messages.


multi-agent-systemsmessage-passingagent-communicationai-agentsllm-engineering

Continue Reading

A Reusable Prompt Kit for Agent Teams
Multi-Agent Systems

A Reusable Prompt Kit for Agent Teams

A team rebuilt their agent prompts from memory every project, and every project drifted a little worse. That failure is why a multi agent prompt kit matters. Here's the reusable set of role prompts every team should keep.

October 2, 2026·9 min read
Multi-Agent Systems on a Budget: An Interactive Guide
Multi-Agent Systems

Multi-Agent Systems on a Budget: An Interactive Guide

Most multi-agent tutorials assume you'll burn tokens freely. That's wrong for anyone shipping on real constraints. A cheap multi agent system can match an expensive one with the right moves. Here's how to build one.

October 2, 2026·9 min read
The Manager Agent Anti-Pattern: A Teardown
Multi-Agent Systems

The Manager Agent Anti-Pattern: A Teardown

Why does your orchestrator agent keep becoming a bottleneck that mangles every handoff? You've hit the manager agent anti pattern — one agent trying to coordinate everything. Here's why it fails and what replaces it.

October 2, 2026·9 min read

Save the prompts from this post

PromptABCD is a free prompt manager. Paste, organize, and reuse your best AI prompts — no more hunting through chat history.

Start free →
← PreviousHow Agents Communicate With Each OtherNext →Shared Memory in Multi-Agent Systems
Share this post:
ShareShare