How Agents Communicate With Each Other
Agent to agent communication seems simple — one agent's output becomes another's input — until you need to pass structured data, handle failures, or coordinate between parallel agents. Here's how each approach works.
from anthropic import Anthropic
import json
client = Anthropic()
# Pattern 1: String passing (simplest, most fragile)
def string_passing_example():
output_a = "The product launch was delayed by 3 weeks due to supply chain issues."
input_b = f"Summarize this update for the CEO: {output_a}"
# Agent B receives raw string — works when format doesn't matter
return input_b
# Pattern 2: Structured JSON passing (recommended for most systems)
def structured_passing_example():
output_a = {
"event": "product_launch_delay",
"delay_weeks": 3,
"root_cause": "supply_chain",
"affected_skus": ["SKU-001", "SKU-002"],
"confidence": 0.92
}
# Agent B receives structured data — can reliably access any field
return json.dumps(output_a)
# Pattern 3: Shared state object (for stateful multi-step workflows)
shared_state = {"task_id": "t_001", "results": {}, "status": "pending"}
def shared_state_example(agent_name: str, result: str):
shared_state["results"][agent_name] = result
shared_state["status"] = f"{agent_name}_complete"
return shared_state
# Pattern 4: Event-driven (for async, parallel, or long-running agents)
event_queue = []
def event_driven_example(event_type: str, payload: dict):
event_queue.append({"type": event_type, "payload": payload, "processed": False})
return event_queuePicture this: you're a software engineer who just got a multi-agent system working in development. Two agents coordinate on a content task — the first does research, the second writes. It works perfectly. Then you try to add a third agent that needs specific structured data from the first, and suddenly you're parsing free-text responses with regex, praying the research agent always formats its output the same way.
Agent to agent communication is where multi-agent systems most commonly fall apart in real development. The failure is almost never in the individual agents — it's in the channel between them.
This guide covers the four main communication patterns, when to use each, and how to avoid the formatting and reliability issues that break agent-to-agent data exchange in production.
Quick-Start: The Four Communication Patterns
Before diving deep, here's a working example of each pattern:
[object Object], anthropic ,[object Object], Anthropic
,[object Object], json
client = Anthropic()
,[object Object],
,[object Object], ,[object Object],():
output_a = ,[object Object],
input_b = ,[object Object],
,[object Object],
,[object Object], input_b
,[object Object],
,[object Object], ,[object Object],():
output_a = {
,[object Object],: ,[object Object],,
,[object Object],: ,[object Object],,
,[object Object],: ,[object Object],,
,[object Object],: [,[object Object],, ,[object Object],],
,[object Object],: ,[object Object],
}
,[object Object],
,[object Object], json.dumps(output_a)
,[object Object],
shared_state = {,[object Object],: ,[object Object],, ,[object Object],: {}, ,[object Object],: ,[object Object],}
,[object Object], ,[object Object],(,[object Object],):
shared_state[,[object Object],][agent_name] = result
shared_state[,[object Object],] = ,[object Object],
,[object Object], shared_state
,[object Object],
event_queue = []
,[object Object], ,[object Object],(,[object Object],):
event_queue.append({,[object Object],: event_type, ,[object Object],: payload, ,[object Object],: ,[object Object],})
,[object Object], event_queueWhat this does: Shows the four communication patterns side by side. String passing is zero setup but unreliable for structured data. JSON passing is the right default for most systems. Shared state works for multi-step workflows where agents build on each other's results. Events work for async or parallel systems.
Understanding the Variables
The right communication pattern depends on three factors:
Data structure: Does the receiving agent need specific fields from the sending agent's output, or can it work with free text? If Agent B needs to access delay_weeks as a number to compute a new date, you need structured JSON, not a string. If Agent B just needs to understand the content and rephrase it, string passing is fine.
Timing: Do your agents run sequentially (one finishes before the next starts), or in parallel? Sequential agents can safely share a mutable state object. Parallel agents writing to the same state object need locking or, better, an event queue where each agent appends rather than overwrites.
Failure behavior: What should happen to downstream agents if an upstream agent fails or returns malformed output? String passing fails silently — Agent B just gets an error message and tries to work with it. JSON passing fails explicitly — if parsing fails, you get an exception you can handle deliberately.
⚡ Pro tip: Your agent to agent communication schema is your system's API contract. Define it before you write agents, not after. If you decide mid-build that Agent A's output format needs to change, every downstream agent that parsed that output needs updating. Treat inter-agent schemas with the same discipline as external API schemas.
Step-by-Step: Implementing Structured Agent Communication
Here's a complete example of agent-to-agent communication using structured JSON — the pattern that works reliably in production:
[object Object], anthropic ,[object Object], Anthropic
,[object Object], json
,[object Object], typing ,[object Object], ,[object Object],
client = Anthropic()
,[object Object], ,[object Object],(,[object Object],) -> ,[object Object],:
,[object Object],
response = client.messages.create(
model=,[object Object],,
max_tokens=,[object Object],,
system=,[object Object],,
messages=[{,[object Object],: ,[object Object],, ,[object Object],: ,[object Object],}]
)
,[object Object],:
,[object Object], json.loads(response.content[,[object Object],].text)
,[object Object], json.JSONDecodeError:
,[object Object],
,[object Object], {
,[object Object],: topic,
,[object Object],: [response.content[,[object Object],].text],
,[object Object],: [],
,[object Object],: [,[object Object],],
,[object Object],: ,[object Object],
}
,[object Object], ,[object Object],(,[object Object],) -> ,[object Object],:
,[object Object],
,[object Object],
key_points = ,[object Object],.join(,[object Object], ,[object Object], f ,[object Object], research[,[object Object],])
data_points = ,[object Object],.join(,[object Object],
,[object Object], d ,[object Object], research[,[object Object],])
response = client.messages.create(
model=,[object Object],,
max_tokens=,[object Object],,
system=,[object Object],,
messages=[{
,[object Object],: ,[object Object],,
,[object Object],: ,[object Object],
}]
)
,[object Object],:
,[object Object], json.loads(response.content[,[object Object],].text)
,[object Object], json.JSONDecodeError:
,[object Object], {,[object Object],: ,[object Object],, ,[object Object],: [{,[object Object],: ,[object Object],, ,[object Object],: response.content[,[object Object],].text}], ,[object Object],: ,[object Object],, ,[object Object],: []}
,[object Object],
research = research_agent(,[object Object],)
article = writing_agent(research, ,[object Object],)
,[object Object],(,[object Object],)
,[object Object],(,[object Object],)
,[object Object],(,[object Object],)What this does: Both agents return JSON, both include a fallback for when the model deviates from the schema, and the writing agent can access research fields by name — research["key_findings"], research["data_points"] — without any string parsing. Adding a third agent that needs the article's key_claims is trivial.
⚡ Pro tip: Include a JSON schema in your agent's system prompt, not just a description. "Return your findings as a JSON object" produces inconsistent formatting. "Return valid JSON only — no markdown, no explanation — matching this exact schema: {..." produces reliable, parseable output. The schema is the most important sentence in your system prompt for agent-to-agent communication.
Shared State for Multi-Step Workflows
When agents run in a sequential pipeline and each agent needs to read what previous agents have done, shared state is cleaner than daisy-chaining outputs:
[object Object], time
,[object Object], ,[object Object],:
,[object Object], ,[object Object],(,[object Object],):
,[object Object],.state = {
,[object Object],: task,
,[object Object],: time.time(),
,[object Object],: {},
,[object Object],: ,[object Object],,
,[object Object],: []
}
,[object Object], ,[object Object],(,[object Object],):
,[object Object],.state[,[object Object],][agent_name] = {
,[object Object],: result,
,[object Object],: time.time()
}
,[object Object],.state[,[object Object],] = ,[object Object],
,[object Object], ,[object Object],(,[object Object],) -> ,[object Object],[,[object Object],]:
entry = ,[object Object],.state[,[object Object],].get(agent_name)
,[object Object], entry[,[object Object],] ,[object Object], entry ,[object Object], ,[object Object],
,[object Object], ,[object Object],(,[object Object],) -> ,[object Object],:
,[object Object],
pipeline = AgentPipeline(,[object Object],.state[,[object Object],])
research_output = research_agent(,[object Object],.state[,[object Object],])
pipeline.add_result(,[object Object],, research_output)
,[object Object],
writing_output = writing_agent(
pipeline.get_result(,[object Object],),
,[object Object],
)
pipeline.add_result(,[object Object],, writing_output)
,[object Object],.state = pipeline.state
,[object Object], ,[object Object],.stateWhat this does: Each agent reads from and writes to shared state. The pipeline is auditable — you can inspect state["agent_outputs"] at any point to see exactly what each agent produced and when. Adding an agent to the middle of the pipeline only requires reading from and writing to the same state dictionary.
Troubleshooting Common Communication Issues
Agent returns non-JSON when you need JSON: Increase the specificity of your schema instruction. Add: "Do not include markdown formatting, code fences, or explanatory text. Return only the JSON object." If the model still deviates, add a validation step that reprompts once if parsing fails.
Downstream agents receive stale context: When agents read from shared state that was written by an earlier agent many steps ago, they might receive context that's no longer accurate given what's happened since. Pass relevant recent state alongside the original task — don't assume agents will naturally read back far enough.
Parallel agents overwriting each other's state: Use agent names as keys in shared state dictionaries rather than overwriting a single "result" field. If agent_a and agent_b both write state["result"] = output, the second one overwrites the first. Use state["agent_a_result"] and state["agent_b_result"] instead.
Communication overhead slowing the system: Long outputs from one agent that become inputs to the next agent inflate token costs and slow processing. Design agents to produce concise, structured outputs. If Agent A's full output is 2,000 words but Agent B only needs the key findings, have Agent A return a structured object with a key_findings field — not the full prose.
⚠️ Common mistake: Passing an agent's entire output to the next agent when only part of it is relevant. If your research agent produces a 3,000-word report and your writing agent only needs the five key findings, extract and pass only the key findings. Sending the full report as context inflates token usage and introduces noise that degrades the writing agent's output quality.
Pro-Level Variations
Typed schemas: Use Python dataclasses or Pydantic models to define your inter-agent communication schemas. This catches type errors before they reach your agents and makes your pipeline self-documenting.
Schema versioning: As your system evolves, agents' output schemas change. Version your schemas explicitly — {"schema_version": "v2", ...} — so you can run old and new agents side by side during transitions without breaking the pipeline.
Dead letter handling: For asynchronous systems, implement a dead letter queue for messages that no agent successfully processed. This is especially important for event-driven agent architectures where a message might be routed to an agent that's unavailable or overloaded.
Your Turn
Start by auditing your current agent-to-agent communication. Are you passing free text and parsing it downstream? Replace those string-to-string connections with structured JSON schemas. Add error handling for schema violations. Measure the reduction in downstream agent failures.
Event-Driven Communication Patterns
For asynchronous multi-agent systems — where agents run on different schedules or where one agent's output triggers multiple downstream agents — event-driven communication is more appropriate than direct message passing:
[object Object], json
,[object Object], time
,[object Object], collections ,[object Object], defaultdict
,[object Object], typing ,[object Object], ,[object Object],
,[object Object], ,[object Object],:
,[object Object], ,[object Object],(,[object Object],):
,[object Object],._subscribers: ,[object Object], = defaultdict(,[object Object],)
,[object Object],._event_log: ,[object Object], = []
,[object Object], ,[object Object],(,[object Object],):
,[object Object],._subscribers[event_type].append(handler)
,[object Object], ,[object Object],(,[object Object],):
event = {
,[object Object],: event_type,
,[object Object],: payload,
,[object Object],: source_agent,
,[object Object],: time.time()
}
,[object Object],._event_log.append(event)
,[object Object],
,[object Object], handler ,[object Object], ,[object Object],._subscribers.get(event_type, []):
handler(event)
,[object Object], ,[object Object],(,[object Object],) -> ,[object Object],:
,[object Object], event_type:
,[object Object], [e ,[object Object], e ,[object Object], ,[object Object],._event_log ,[object Object], e[,[object Object],] == event_type]
,[object Object], ,[object Object],._event_log
,[object Object],
bus = EventBus()
,[object Object], ,[object Object],(,[object Object],):
,[object Object],
research_data = event[,[object Object],]
,[object Object],
analysis = {,[object Object],: [,[object Object], + research_data[,[object Object],]]}
bus.publish(,[object Object],, analysis, ,[object Object],)
bus.subscribe(,[object Object],, research_agent_handler)
,[object Object],
bus.publish(
,[object Object],,
{,[object Object],: ,[object Object],, ,[object Object],: [,[object Object],, ,[object Object],]},
,[object Object],
)What this does: The event bus decouples producers from consumers. The research agent doesn't know which agents will process its output — it publishes an event and the bus notifies all subscribers. This makes it easy to add new downstream agents (subscribe them to existing events) or remove agents without modifying the producers. The event log provides a complete audit trail of all inter-agent communication.
⚡ Pro tip: Always include a source_agent field in event payloads. When debugging a multi-agent event-driven system, knowing which agent produced the event that triggered a given behavior is essential. Without source attribution in the event log, tracing a behavior back to its origin requires correlating timestamps across separate agent logs — which is much slower and less reliable.
Communication Overhead at Scale
Direct inter-agent communication and event-driven systems both accumulate overhead as agent count grows. With N agents all potentially communicating with each other, the potential number of communication channels grows as N*(N-1)/2. At ten agents, that's forty-five potential channels.
In practice, well-designed multi-agent systems have sparse communication graphs — most agents communicate with only a few others, and the orchestrator or event bus mediates all communication. Designing the communication topology explicitly — deciding which agents need to exchange information and through what mechanism — is as important as designing the agents themselves.
When you've found communication schemas that work reliably, save them to PromptABCD alongside the agent system prompts they're paired with. A great schema is as valuable as a great system prompt — it makes the whole pipeline reliable.
Versioning Communication Schemas
Agent communication schemas evolve as your system matures. New fields get added, field types change, optional fields become required. Without schema versioning, these changes break existing agent pairs that depend on the old format.
The simplest versioning approach adds a schema_version field to every inter-agent message:
message = {
,[object Object],: ,[object Object],,
,[object Object],: ,[object Object],,
,[object Object],: ,[object Object],,
,[object Object],: { ,[object Object],: [...] }
}When the receiving agent deserializes the message, it checks the schema version and applies the appropriate parser. This allows old and new message formats to coexist during a migration period — senders can emit the new format while receivers are updated in a rolling deployment.
The alternative to explicit versioning is compatibility rules: new fields are always optional, old fields are never removed, and type changes are never made (new fields replace old ones). Compatibility rules are simpler but create schemas that accumulate technical debt — fields that are deprecated but never removed, optional fields that all current senders happen to always include. Explicit versioning is more work up front and significantly cleaner over time. When your system has been running in production for six months and you need to introduce a schema change, having version numbers in your messages means you can plan and execute the migration without coordinating a simultaneous deploy across all agents.
Continue Reading
Save the prompts from this post
PromptABCD is a free prompt manager. Paste, organize, and reuse your best AI prompts — no more hunting through chat history.
