Guardrails in Multi-Agent Systems
A single bad output from one agent can propagate through five others before anyone notices. Multi agent guardrails have to work at the seams, not just the edges. Here's how.
# A guardrail at the seam - validate output before it becomes input
def handoff(from_agent, to_agent, output):
check = validate_output(output, schema=to_agent.input_schema)
if not check.ok:
# contain it here, before it poisons the next agent
return quarantine(output, reason=check.reason)
return deliver(to_agent, output)In a single-agent system, one bad output is one bad output - the user sees it and moves on. In a multi-agent system, a single bad output from one agent can propagate through five others before anyone notices, each downstream agent treating the poisoned input as trustworthy and building on it. That amplification is why multi agent guardrails can't just sit at the edges of the system checking the final output - they have to work at the seams between agents, where a small error is still small and hasn't yet been laundered into an authoritative-looking conclusion.
Most teams put a guardrail on the user-facing input and the user-facing output and call it done. That's necessary but nowhere near sufficient for a team of agents, because the dangerous failures happen inside, between agents, where edge guardrails never look. Let me lay out what guardrails an agent team actually needs and where to place them.
What Are Multi Agent Guardrails?
Multi agent guardrails are the checks that constrain what agents can produce, consume, and do - placed not only at the system's edges but at every handoff between agents. They come in three kinds: input validation (is what this agent received well-formed and safe to act on?), output validation (is what this agent produced within acceptable bounds before it moves on?), and action constraints (is this agent allowed to take this action at all?).
The defining difference from single-agent guardrails is placement. Because errors amplify across handoffs, a guardrail at each seam catches a bad output while it's still contained to one agent, before the next agent builds on it. Waiting until the final output means catching the error after five agents have compounded it, when it's far harder to trace and far more expensive to have produced.
[object Object],
,[object Object], ,[object Object],(,[object Object],):
check = validate_output(output, schema=to_agent.input_schema)
,[object Object], ,[object Object], check.ok:
,[object Object],
,[object Object], quarantine(output, reason=check.reason)
,[object Object], deliver(to_agent, output)What this does: It validates an agent's output against what the receiving agent expects before the handoff completes, quarantining anything that fails instead of passing it on. A malformed or out-of-bounds output is stopped at the seam where it was produced, so it never becomes another agent's trusted input - which is the whole point of guardrails at the seams.
Why It Matters
Guardrails at the seams matter because of amplification, and amplification is what makes multi-agent failures qualitatively worse than single-agent ones. When agent one produces a subtly wrong number and agent two computes on it, agent three summarizes that, and agent four acts on the summary, the original small error becomes a confident, well-formatted, thoroughly-propagated wrong conclusion. Nobody can easily see it's wrong because four agents have dressed it up.
Three scenarios where seam guardrails earn their place:
A financial-analysis team had an extraction agent that occasionally misread a figure, and a downstream chain that computed projections from it. Without a guardrail checking the extracted figure against plausible ranges at the seam, one misread number produced a confident, entirely wrong financial projection. A range check at the extraction output would have caught it while it was still one bad number.
A content-moderation pipeline had agents that each refined a decision. A too-permissive early decision propagated through the chain and out to users, because the guardrail only checked the final action, by which point the reasoning looked sound. Checking each agent's decision at its own seam caught the permissive call early.
A customer-service system let a summarizer agent feed an action agent that could issue refunds. Without an action constraint at the seam, a hallucinated "customer is owed a refund" from the summarizer flowed straight into an actual refund. An action guardrail requiring corroboration before a financial action stopped it.
⚡ Pro tip: Place your first new guardrail at the seam just before any agent that takes a real-world action - sends an email, issues a refund, writes to a database. Action agents are where a propagated error stops being a wrong answer and becomes a wrong deed, so the handoff into them is the highest-value place to validate. Guard the actuators first.
What Guardrails Does an Agent Team Need?
Three layers, placed deliberately.
Input validation at each agent checks that what an agent received matches the schema and constraints it expects. This catches upstream errors and malformed handoffs before the agent acts on them, and it's cheap - a schema check is fast and stops a whole class of propagation.
Output validation at each agent checks that what the agent produced is within bounds before it moves on. This is the seam guardrail, and it's the one most systems lack. Range checks, schema checks, and sanity checks on each output contain errors at their source.
[object Object], ,[object Object],(,[object Object],):
,[object Object], ,[object Object], matches_schema(output, schema):
,[object Object], Check(ok=,[object Object],, reason=,[object Object],)
,[object Object], field, (lo, hi) ,[object Object], schema.ranges.items():
,[object Object], ,[object Object], lo <= output.get(field, lo) <= hi:
,[object Object], Check(ok=,[object Object],, reason=,[object Object],)
,[object Object], Check(ok=,[object Object],)What this does: It checks an agent's output against both structure (does it match the expected schema?) and plausibility (are the values in sane ranges?) before the output is allowed to proceed. The range check is what catches the subtly-wrong-number case that a pure schema check misses - a value can be perfectly well-formed and still absurd.
Action constraints gate what agents are permitted to do. An agent that can take consequential actions should be constrained to only the actions its role permits, and high-consequence actions should require corroboration or human approval. This is the guardrail that keeps a propagated error from becoming an irreversible action.
⚡ Pro tip: Make guardrails fail loud and visible, not silent. A guardrail that quietly drops a bad output can hide a systemic problem - if one agent's outputs are constantly being quarantined, you need to know, because that agent is broken. Log every guardrail trigger and alert on rising trigger rates, so the guardrail protects the system and also tells you which agent needs fixing.
How Do You Keep Guardrails From Slowing Everything Down?
A fair concern: a validation at every seam sounds like a lot of overhead. The resolution is to match each guardrail's cost to its risk. Cheap checks - schema validation, range checks - go everywhere, because they cost microseconds and catch the most common errors. Expensive checks - a model-based safety evaluation, a corroboration lookup - go only at the high-consequence seams, before action agents.
This tiering keeps guardrails affordable. You're not running a heavyweight safety model on every handoff; you're running a fast schema check on every handoff and reserving the expensive checks for the few seams where the consequence justifies them. The result is comprehensive coverage where errors amplify and heavy scrutiny where actions happen, without a validation tax on every internal step.
There's a second efficiency worth knowing: guardrails compound. A cheap schema check at every seam catches most malformed outputs so cheaply that your expensive checks run far less often, because fewer bad outputs survive to reach them. The cheap first layer does the bulk filtering; the expensive layer only ever sees the small fraction that passed structural checks but might still be substantively wrong. Layering cheap-then-expensive isn't just about placement, it's about letting the cheap layer reduce the load on the expensive one, which is what keeps comprehensive guardrails from becoming a cost center.
⚡ Pro tip: Order your guardrails cheapest-first at each seam, and stop at the first failure. A schema check that rejects a malformed output means you never pay for the expensive semantic check on that output. Fail-fast ordering makes multi agent guardrails cheaper the more layers you add, because each cheap layer shrinks the input to the expensive ones - the opposite of the "more checks, more cost" intuition that makes teams skimp on validation.
⚠️ Common mistake: Putting guardrails only at the system's input and output, leaving the agent-to-agent seams unguarded. This is the single most common guardrail gap in multi-agent systems, and it's exactly backwards for how these systems fail - the amplification happens between agents, so edge-only guardrails watch the two places where errors are least dangerous and ignore the many places where a small error becomes a big one. Guard the seams, not just the edges.
Conclusion
Multi agent guardrails have to account for amplification: a small error at one agent becomes a large error after propagating through several. So place validation at every seam, not just the edges - cheap schema and range checks everywhere, expensive safety and corroboration checks before action agents. Make guardrails fail loud so they double as a signal of which agent is misbehaving.
Keep your guardrail definitions and validation schemas versioned so you can reuse a safety configuration you trust. I store the validation schemas and action-constraint rules in PromptABCD, because the schemas encode what each agent seam should accept and the constraints encode what each agent is allowed to do, and reusing a proven set of guardrails means a new agent team inherits containment at every seam instead of shipping with only edge protection and discovering the amplification problem in production.
⚡ Pro tip: When a guardrail fires in production, treat it as a finding, not just a block. A quarantined output tells you an agent produced something out of bounds, which is a bug worth investigating at its source - not merely a thing to stop and forget. Feed guardrail triggers back into fixing the agent that caused them, and over time your seams get quieter because the agents stop producing the bad outputs, rather than because you stopped looking.
Continue Reading
Save the prompts from this post
PromptABCD is a free prompt manager. Paste, organize, and reuse your best AI prompts — no more hunting through chat history.
