Handoffs Between Specialized Agents
Picture this: you're debugging a multi-agent pipeline where the final output is wrong, and you can't tell whether the error came from Agent A or Agent B, because the handoff between them swallowed all the context. Here's how to build handoffs that preserve what matters.
from anthropic import Anthropic
client = Anthropic()
def researcher(topic: str) -> str:
response = client.messages.create(
model="claude-sonnet-4-6",
max_tokens=2048,
system="You are a research analyst. Research the topic thoroughly.",
messages=[{"role": "user", "content": f"Research: {topic}"}]
)
return response.content[0].text # Full research output — unstructured
def writer(research_output: str, topic: str) -> str:
# Receives entire research output — must figure out what matters
response = client.messages.create(
model="claude-sonnet-4-6",
max_tokens=2048,
system="You are a writer. Write a clear article based on the research.",
messages=[{
"role": "user",
"content": f"Topic: {topic}\n\nResearch:\n{research_output}\n\nWrite the article."
}]
)
return response.content[0].text
topic = "AI inference cost trends 2025"
research = researcher(topic)
article = writer(research, topic) # Writer gets raw research dump — implicit handoffPicture this: you're debugging a multi-agent pipeline where the final output is wrong. The research agent produced good output. The writing agent produced good output given its input. But the final article has a subtle error — a claim that contradicts the research. You spend two hours tracing the problem and eventually find it: during the handoff from researcher to writer, a key qualifier ("in Q3 2024, not currently") was buried in a long research document and the writer missed it.
The handoff wasn't broken. It just didn't surface the most important information where the receiving agent would attend to it.
Agent handoff pattern failures are almost never caused by the data being absent. They're caused by the receiving agent not knowing which data matters most. The difference between a handoff that works and one that fails is how explicitly the sending agent communicates priority.
Before: The Implicit Handoff
The most common handoff implementation passes the sending agent's complete output to the receiving agent and expects the receiver to extract what's relevant:
[object Object], anthropic ,[object Object], Anthropic
client = Anthropic()
,[object Object], ,[object Object],(,[object Object],) -> ,[object Object],:
response = client.messages.create(
model=,[object Object],,
max_tokens=,[object Object],,
system=,[object Object],,
messages=[{,[object Object],: ,[object Object],, ,[object Object],: ,[object Object],}]
)
,[object Object], response.content[,[object Object],].text ,[object Object],
,[object Object], ,[object Object],(,[object Object],) -> ,[object Object],:
,[object Object],
response = client.messages.create(
model=,[object Object],,
max_tokens=,[object Object],,
system=,[object Object],,
messages=[{
,[object Object],: ,[object Object],,
,[object Object],: ,[object Object],
}]
)
,[object Object], response.content[,[object Object],].text
topic = ,[object Object],
research = researcher(topic)
article = writer(research, topic) ,[object Object],What this does: The writer receives the researcher's full output and must determine what to emphasize, what to include, and what to set aside. This works when research output is short and well-structured. It breaks when research output is long, when some findings are more reliable than others, or when there are specific caveats that must appear in the final article.
Why Implicit Handoffs Fail
The researcher and the writer have different expertise. The researcher knows which findings are solid and which are tentative. The researcher knows which sources are high-confidence and which are questionable. The researcher knows which data points are essential versus supplementary. But in an implicit handoff, none of this meta-knowledge transfers.
The writer has to guess which part of a 2,000-word research output is the critical insight and which is background context. Sometimes it guesses wrong. And because the writer produces confident-sounding prose, the wrong guess isn't flagged — it's just embedded in the article.
⚠️ Common mistake: Treating the handoff as just "passing the output" when it should be "passing the output plus what the receiver needs to know about the output." The sending agent has knowledge about its own output that the receiving agent doesn't have — confidence levels, important caveats, prioritization. That meta-knowledge is what makes handoffs work.
After: The Structured Handoff With Exit Interview
The fix is to add an "exit interview" step before handoff: before passing its output to the next agent, the sending agent produces a structured handoff package that explicitly marks what's most important, what's uncertain, and what caveats the receiving agent must respect.
[object Object], json
,[object Object], anthropic ,[object Object], Anthropic
client = Anthropic()
,[object Object], ,[object Object],(,[object Object],) -> ,[object Object],:
,[object Object],
raw_research = client.messages.create(
model=,[object Object],,
max_tokens=,[object Object],,
system=,[object Object],,
messages=[{,[object Object],: ,[object Object],, ,[object Object],: ,[object Object],}]
).content[,[object Object],].text
,[object Object],
handoff_response = client.messages.create(
model=,[object Object],,
max_tokens=,[object Object],,
system=,[object Object],,
messages=[{
,[object Object],: ,[object Object],,
,[object Object],: ,[object Object],
}]
).content[,[object Object],].text
,[object Object],:
handoff_package = json.loads(handoff_response)
handoff_package[,[object Object],] = raw_research
,[object Object], handoff_package
,[object Object], json.JSONDecodeError:
,[object Object], {,[object Object],: raw_research, ,[object Object],: handoff_response}
,[object Object], ,[object Object],(,[object Object],) -> ,[object Object],:
,[object Object],
structured_input = json.dumps({
,[object Object],: handoff_package.get(,[object Object],, []),
,[object Object],: handoff_package.get(,[object Object],, {}),
,[object Object],: handoff_package.get(,[object Object],, []),
,[object Object],: handoff_package.get(,[object Object],, []),
,[object Object],: handoff_package.get(,[object Object],, ,[object Object],)
}, indent=,[object Object],)
response = client.messages.create(
model=,[object Object],,
max_tokens=,[object Object],,
system=,[object Object],,
messages=[{
,[object Object],: ,[object Object],,
,[object Object],: ,[object Object],
}]
)
,[object Object], response.content[,[object Object],].text
topic = ,[object Object],
handoff_package = researcher_with_handoff(topic)
article = writer_with_handoff(handoff_package, topic)
,[object Object],(article)What this does: After completing research, the researcher runs a second LLM call — the exit interview — where it explicitly packages its knowledge about its own output: what's high-confidence, what must be caveated, what shouldn't be overstated. The writer receives this structured package alongside the raw research and can write with appropriate precision about each claim.
⚡ Pro tip: The exit interview is the most underused technique in multi-agent design. The sending agent knows things about its own output that no downstream agent can reliably infer. Explicit handoff packaging is usually a 200-400 token investment that produces significant improvements in handoff accuracy — especially for tasks where precision and caveat accuracy matter.
Breaking Down the Handoff Package Structure
A well-structured handoff package has five components:
Primary findings: The 3–5 most important things the receiving agent should use. Without this, receiving agents weight all research findings equally — which they're not.
Confidence notes: Which claims are high-confidence versus speculative. This prevents the writer from presenting a tentative finding with the same certainty as an established fact.
Required caveats: Specific qualifications that must appear in the output. "This data is from Q3 2024 and may not reflect current conditions" must appear if the research is time-sensitive.
Avoid overclaiming: Specific claims that must not be amplified. Research might say "some studies suggest" — the writer must not say "research proves."
Handoff timestamp and version: For auditing. When was the handoff created? What version of the source data does it reference? This matters in production systems where research data updates over time.
Variations for Different Contexts
Code handoffs: When a code-generation agent hands off to a code-review agent, the exit interview should include: known limitations in the implementation, assumptions made that weren't in the original spec, edge cases that weren't handled, and specific things the reviewer should check carefully.
Data analysis handoffs: When an analysis agent hands off to a reporting agent, the exit interview should include: statistical confidence intervals, sample size caveats, potential confounds, and claims that are descriptive vs. causal.
Customer interaction handoffs: When a triage agent hands off to a specialist agent, the exit interview should include: the customer's emotional state, key facts established so far, promises already made, and questions still unanswered.
Save and Reuse This Pattern
The agent handoff pattern with exit interview is one of the highest-return improvements you can make to any multi-agent pipeline with sequential dependencies. It costs one additional API call per handoff and dramatically reduces the probability that critical nuance gets lost between agents.
Handoff Templates for Common Domain Pairs
Different domain combinations require different handoff field sets. A handoff from a research agent to a writing agent needs different metadata than a handoff from a code generator to a code reviewer. Defining handoff templates by domain pair prevents agents from needing to invent handoff structure each time:
[object Object], dataclasses ,[object Object], dataclass
,[object Object], typing ,[object Object], ,[object Object],, ,[object Object],
,[object Object],
,[object Object], ,[object Object],:
,[object Object],
research_summary: ,[object Object],
key_facts: ,[object Object],
confidence_level: ,[object Object], ,[object Object],
source_quality: ,[object Object], ,[object Object],
content_gaps: ,[object Object], ,[object Object],
recommended_tone: ,[object Object], ,[object Object],
target_word_count: ,[object Object],
do_not_include: ,[object Object], ,[object Object],
,[object Object],
,[object Object], ,[object Object],:
,[object Object],
code: ,[object Object],
language: ,[object Object],
implementation_notes: ,[object Object], ,[object Object],
known_limitations: ,[object Object], ,[object Object],
test_cases_needed: ,[object Object], ,[object Object],
security_concerns: ,[object Object], ,[object Object],
performance_considerations: ,[object Object],
,[object Object], ,[object Object],(,[object Object],) -> ,[object Object],:
,[object Object],
,[object Object], json
,[object Object], dataclasses ,[object Object], asdict
,[object Object], json.dumps(asdict(handoff), indent=,[object Object],)What this does: Each handoff type is a typed Python dataclass with domain-specific fields. The sending agent populates the dataclass; format_handoff_for_agent converts it to structured JSON that the receiving agent receives as context. Type enforcement at the Python level catches missing required fields before they reach the receiving agent. When you add a new required field to a handoff template, the type system immediately surfaces all senders that need to be updated.
⚡ Pro tip: Include a do_not_include or avoid field in every handoff template. The sending agent knows which of its outputs are weakly supported, speculative, or outside scope — knowledge that the receiving agent doesn't have. Explicit exclusion fields are the fastest way to prevent the receiving agent from elaborating on content that the sender itself flagged as uncertain.
Detecting Handoff Failures
Handoff failures are subtler than agent failures. An agent that errors raises an exception. An agent that receives a bad handoff produces confident, well-formatted output that answers the wrong question. Detecting handoff failures requires validation at the receiving agent's entry point:
[object Object], pydantic ,[object Object], BaseModel, ValidationError
,[object Object], json
,[object Object], ,[object Object],(,[object Object],):
research_summary: ,[object Object],
key_facts: ,[object Object],
confidence_level: ,[object Object],
content_gaps: ,[object Object],
,[object Object], ,[object Object],(,[object Object],) -> ,[object Object],:
,[object Object],:
data = json.loads(raw_handoff)
validated = model_class(**data)
,[object Object], {,[object Object],: ,[object Object],, ,[object Object],: validated.,[object Object],()}
,[object Object], (json.JSONDecodeError, ValidationError) ,[object Object], e:
,[object Object], {
,[object Object],: ,[object Object],,
,[object Object],: ,[object Object],(e),
,[object Object],: raw_handoff[:,[object Object],]
}
,[object Object], ,[object Object],(,[object Object],) -> ,[object Object],:
result = receive_with_validation(handoff_json, ResearchHandoffPayload)
,[object Object], result[,[object Object],] == ,[object Object],:
,[object Object],
,[object Object],(,[object Object],)
,[object Object], ,[object Object],
payload = result[,[object Object],]
,[object Object],
,[object Object], ,[object Object],What this does: The receiving agent validates the handoff payload against a Pydantic schema before processing it. If the handoff is malformed — missing fields, wrong types, truncated JSON — the validator catches it and returns an error rather than attempting to process a partial handoff. The error is specific enough to diagnose: whether the handoff JSON was invalid, which fields were missing, or which fields had the wrong type.
⚡ Pro tip: Add handoff validation metrics to your monitoring. Track what percentage of handoffs pass validation on first attempt and what percentage fail. A handoff failure rate above two percent indicates that the sending agent's output format is drifting from the expected schema — either the sender's prompt needs refinement or the receiving agent's schema needs updating to match what senders reliably produce.
Handoff Monitoring in Production
Handoffs are the most common source of quality degradation in production multi-agent systems. This is true for two reasons: handoffs are changed more frequently than individual agent prompts (as capabilities evolve), and handoffs are tested less rigorously (because quality issues in handoffs manifest as downstream agent failures, not as handoff-layer failures, making the root cause less obvious).
Building a monitoring layer for handoffs pays dividends disproportionate to its implementation cost. Most multi-agent debugging time is spent tracing where a quality problem was introduced; handoff monitoring surfaces this in seconds rather than hours. Treat handoff monitoring as a required feature, not an optional improvement. A research agent that produces excellent findings can still produce poor downstream content if the handoff to the writer loses critical nuance. Monitoring handoff quality requires measuring output quality relative to handoff completeness:
Track three metrics for each handoff point in your pipeline: schema validation pass rate (what percentage of handoffs have all required fields), downstream quality score (how the receiving agent's output quality correlates with handoff completeness), and handoff latency (the time the exit interview adds to the pipeline). The first two metrics tell you whether handoffs are working; the third tells you what they cost.
When downstream quality drops — the writing agent starts producing less accurate or less well-sourced content — the first place to investigate is the handoff layer. Degraded handoffs are a more common cause of downstream quality degradation than degraded agent prompts, because handoffs are changed more frequently (as agent capabilities evolve) and tested less rigorously.
Store your exit interview prompt templates in PromptABCD. Different domain combinations need different handoff structures — a research-to-writing handoff needs different fields than a code-generation-to-review handoff. Maintaining a library of domain-specific handoff templates saves the iteration time of rebuilding them for each new pipeline.
Continue Reading
Save the prompts from this post
PromptABCD is a free prompt manager. Paste, organize, and reuse your best AI prompts — no more hunting through chat history.
