How to Evaluate a Multi-Agent System
Most guides tell you to evaluate a multi agent system by its final output. That's wrong — final-output metrics hide which agent is failing. Here's how to evaluate the system by its seams, where problems actually live.
def evaluate_pipeline(pipeline, test_cases):
results = []
for case in test_cases:
trace = pipeline.run_with_trace(case) # capture every handoff
results.append({
"final_score": score_output(trace.final, case.expected),
"per_handoff": [score_handoff(h) for h in trace.handoffs],
"case": case.id,
})
return resultsMost guides tell you to evaluate a multi agent system by measuring its final output — did the pipeline produce a good answer? That advice is not just incomplete, it's actively misleading, because a good final output can hide a broken agent that a lucky downstream agent happened to compensate for, and a bad output tells you nothing about which of five agents caused it. To evaluate a multi agent system properly, you measure the seams — the handoffs between agents — where problems actually originate and where a final-output metric is blind by construction.
This piece lays out what proper evaluation looks like, why output-only metrics fail, how to measure per-agent contribution, and the mistakes that make evaluation useless.
What does it mean to evaluate a multi agent system?
To evaluate a multi agent system is to measure not just whether it produces good results but where along the chain of agents quality is created and destroyed. A single-agent evaluation is one number: input, output, quality. A multi-agent evaluation is a decomposition: how good was the planner's plan, did the researcher's output actually serve that plan, did the synthesizer use what it was given. The unit of evaluation shifts from the system to the handoff.
This matters because multi-agent failures are compositional. The system can fail even when every agent is individually fine, because the handoffs lose information — agent A produces something correct that agent B misinterprets. And the system can succeed despite a broken agent, because a strong downstream agent papers over an upstream mistake. Neither case is visible in the final output alone. You have to look at the seams.
[object Object], ,[object Object],(,[object Object],):
results = []
,[object Object], ,[object Object], ,[object Object], test_cases:
trace = pipeline.run_with_trace(,[object Object],) ,[object Object],
results.append({
,[object Object],: score_output(trace.final, ,[object Object],.expected),
,[object Object],: [score_handoff(h) ,[object Object], h ,[object Object], trace.handoffs],
,[object Object],: ,[object Object],.,[object Object],,
})
,[object Object], resultsWhat this does: it runs each test case while capturing every handoff, then scores both the final output and each individual handoff — so a failure can be traced to the specific agent transition that caused it rather than attributed vaguely to "the system."
Why output-only evaluation fails
The core failure of output-only metrics is attribution blindness. When your final score drops, an output-only evaluation tells you the system got worse but not which agent regressed. In a five-agent pipeline, that's a debugging nightmare — you're bisecting blind. Per-handoff scoring tells you immediately: the planner's plans got vaguer, or the verifier started passing bad claims. You fix the actual regression instead of guessing.
The second failure is compensated-error masking. Suppose your researcher agent degrades and starts returning weaker findings, but your synthesizer is strong enough to still produce a decent report. Output-only evaluation shows no problem — until the day a slightly harder case exceeds the synthesizer's ability to compensate, and quality falls off a cliff with no warning. Per-agent evaluation would have flagged the researcher's decline weeks earlier, while it was still masked.
⚠️ Common mistake: treating a good average final score as proof the system is healthy. Averages hide the compensated errors and the high-variance failures that matter most. A system averaging well while quietly depending on one agent to rescue another's mistakes is fragile, and the average is exactly the metric that won't show it. Evaluate the distribution and the seams, not the mean.
The third failure is that output-only metrics can't distinguish a coordination problem from an agent problem. If the system fails, is it because an agent is bad or because two good agents don't hand off cleanly? These demand completely different fixes — retune an agent versus redesign an interface — and the final score can't tell them apart. Handoff-level evaluation separates "the agent produced garbage" from "the agent produced something fine that the next agent couldn't use."
How to measure per-agent contribution
The most useful technique is ablation: swap one agent for a known-good baseline and measure how much the final quality changes. If replacing your researcher with a perfect oracle barely improves the output, the researcher isn't your bottleneck — something downstream is. If it improves dramatically, the researcher is where to invest. Ablation turns "which agent should I improve?" from a guess into a measurement.
[object Object], ,[object Object],(,[object Object],):
baseline = mean(score(pipeline.run(c), c) ,[object Object], c ,[object Object], test_cases)
patched = pipeline.replace(agent_name, oracle)
improved = mean(score(patched.run(c), c) ,[object Object], c ,[object Object], test_cases)
,[object Object], improved - baseline ,[object Object],What this does: it replaces one agent with an oracle and measures the resulting quality jump, quantifying exactly how much that agent's imperfection costs the whole system — so you invest tuning effort where it actually moves the final result.
The second technique is handoff fidelity: for each handoff, measure whether the receiving agent used the information the sending agent provided. An agent that ignores its input is a broken seam even if both agents are individually good. You can measure this by checking whether the downstream output is sensitive to changes in the upstream output — if it isn't, the handoff is dead and one of the agents is redundant.
⚡ Pro tip: build a small set of "trap" test cases where you know the correct answer requires information that only one specific agent can provide. If the system still gets these right when you sabotage that agent's output, you've proven the agent isn't actually contributing — the rest of the pipeline is guessing correctly. Trap cases surface phantom agents that add cost without adding value, which are common and invisible to normal metrics.
How do you score a handoff you can't easily grade?
The method above assumes you can score each handoff, but scoring is where multi-agent evaluation gets genuinely hard. A planner's plan or a researcher's findings don't have a single right answer to check against, so teams reach for a model to grade them — an LLM-as-judge. That works, with caveats you have to respect or your whole evaluation becomes confidently wrong.
The first caveat is that a judge model shares blind spots with the agents it's grading. If your researcher and your judge are the same model family, the judge will rate the researcher's plausible-but-wrong output highly, because it makes the same mistakes and finds them plausible for the same reasons. Use a different model as judge where you can, or at minimum a different prompt lineage, so the judge's failure modes don't perfectly overlap the agent's. A judge that agrees with the agent's errors is worse than no judge, because it manufactures false confidence.
The second caveat is that judges drift toward superficial quality — length, fluency, confident tone — over substance. A judge scoring a research handoff will reward a long, well-written finding over a short, correct one unless you explicitly anchor it to the task. Give the judge a concrete rubric tied to what the downstream agent actually needs, not a vague "rate the quality" instruction.
[object Object], ,[object Object],(,[object Object],):
system = (,[object Object],
,[object Object],
,[object Object],
,[object Object],
,[object Object],)
,[object Object], llm(system=system, user=handoff)What this does: it anchors the judge to the downstream agent's actual need and forbids rewarding fluency over correctness, producing a grounded per-item score instead of a length-biased overall impression.
The most reliable check on your judge is to periodically grade a sample by hand and measure how well the judge agrees with you. When judge-human agreement drops, your automated evaluation has quietly stopped measuring what you care about, and every downstream decision based on it is suspect. The judge needs evaluating just like the agents do — an unevaluated judge is an unexamined assumption sitting under your entire quality process.
⚡ Pro tip: calibrate the judge against a small human-graded set before trusting it, and re-calibrate after any model change. A judge that agreed with human graders last quarter can drift with a model update, silently shifting your quality bar. Treating the judge as a fixed instrument rather than a component that needs its own evaluation is how teams end up optimizing confidently toward the wrong target for months.
Common mistakes
Teams evaluate the whole system as a black box and then wonder why they can't improve it. Without per-agent visibility, every improvement is a shot in the dark. Instrument the handoffs first; optimize second.
Teams also over-invest in the wrong agent because it's the one they understand best, not the one the ablation identifies as the bottleneck. The agent you find most interesting is rarely the one costing you the most quality. Let the ablation numbers, not your intuition, direct where you spend tuning effort.
And teams evaluate on too few cases, so high-variance failures — the ones that matter in production — never show up. Multi-agent systems have more failure modes than single agents because they have more moving parts and more seams, so they need larger, more varied evaluation sets to surface the tail failures that a small set misses.
⚡ Pro tip: log every production run's handoffs, not just failures, and sample them into your evaluation set continuously. Production traffic surfaces failure modes your hand-written test cases never imagined, and multi-agent systems fail in combinatorially more ways than single agents. A living evaluation set fed by real handoffs catches regressions your static tests structurally can't, and it's the single highest-value evaluation investment for a system in production.
Conclusion
To evaluate a multi agent system, measure the seams, not just the output. Score each handoff, ablate agents to find the real bottleneck, check handoff fidelity to catch dead seams, and evaluate on enough varied cases to surface the tail failures that averages hide. The final output tells you something is wrong; the seams tell you what.
The evaluation prompts and scoring rubrics you develop are as reusable as the pipeline itself. Store them in PromptABCD alongside your agent prompts so your evaluation method travels with the system it measures, and every new pipeline starts with a proven way to find its own weak seams instead of relying on a final score that hides them.
Continue Reading
Save the prompts from this post
PromptABCD is a free prompt manager. Paste, organize, and reuse your best AI prompts — no more hunting through chat history.
