How to Handle a Failing Agent in a Team
Most multi agent failure handling guides are wrong: they treat a failed agent like a crashed server. Agents fail differently. Here's a copy-paste guide to detecting and containing them.
import time
class AgentFailure(Exception):
def __init__(self, kind, detail):
self.kind = kind # "error" | "timeout" | "invalid" | "loop"
self.detail = detail
def run_agent_guarded(agent, task, validate, timeout=60, max_steps=8):
start = time.time()
steps = 0
for output in agent.stream(task):
steps += 1
if time.time() - start > timeout:
raise AgentFailure("timeout", f"exceeded {timeout}s")
if steps > max_steps:
raise AgentFailure("loop", f"exceeded {max_steps} steps")
ok, reason = validate(output)
if not ok:
raise AgentFailure("invalid", reason)
return outputMost multi agent failure handling guides are wrong in the same way: they treat a failing agent like a crashed microservice. Restart it, retry the call, move on. But agents don't fail like servers. A server that fails throws an error and stops. An agent that "fails" often keeps going confidently - producing plausible, wrong output, or looping, or quietly ignoring half its instructions. The scariest agent failures return HTTP 200.
So detection is the hard part, not recovery. This guide gives you code you can paste in today to catch the failures that don't announce themselves, then contain them before they poison the rest of the team.
Quick-Start (Copy This Right Now)
Here's a wrapper that catches the three failure classes agents actually exhibit: hard errors, timeouts, and - the important one - silent bad output that passes a validity check.
[object Object], time
,[object Object], ,[object Object],(,[object Object],):
,[object Object], ,[object Object],(,[object Object],):
,[object Object],.kind = kind ,[object Object],
,[object Object],.detail = detail
,[object Object], ,[object Object],(,[object Object],):
start = time.time()
steps = ,[object Object],
,[object Object], output ,[object Object], agent.stream(task):
steps += ,[object Object],
,[object Object], time.time() - start > timeout:
,[object Object], AgentFailure(,[object Object],, ,[object Object],)
,[object Object], steps > max_steps:
,[object Object], AgentFailure(,[object Object],, ,[object Object],)
ok, reason = validate(output)
,[object Object], ,[object Object], ok:
,[object Object], AgentFailure(,[object Object],, reason)
,[object Object], outputWhat this does: It runs an agent under three simultaneous guards - a wall-clock timeout, a step ceiling to catch loops, and a validator that inspects the final output. The validator is what separates this from naive error handling: it catches the confident-but-wrong outputs that never raise an exception on their own.
Understanding the Variables
The pieces you'll tune are validate, timeout, and max_steps, and they map to the three failure modes.
validate is the most important and the most neglected. It's a function that returns whether the output is actually usable, not just well-formed. For a JSON-emitting agent, "valid" means the schema matches AND the values are sane - a price of -$4,000,000 is well-formed and obviously broken. Good multi agent failure handling starts here, because this is the only guard that catches silent failures.
timeout bounds wall-clock time. Set it from real percentiles, not a guess - measure your p95 successful run and add headroom. Too tight and you kill slow-but-correct agents; too loose and hung agents waste your budget.
max_steps catches loops - an agent calling the same tool repeatedly, or two agents ping-ponging. Agents loop more than servers do because a confused model often "tries again" instead of erroring.
⚡ Pro tip: Write the validator before you write the agent. If you can't specify what a good output looks like precisely enough to check it in code, your agent doesn't have a clear enough job - and you'll never be able to tell success from failure at runtime.
Step-by-Step: Multi Agent Failure Handling
Here's the full flow, from detection to containment to recovery, as a coordinator would run it.
Step one, isolate the failing agent's effects. Before an agent's output is allowed to touch shared state, it must pass its guard. A failed agent's partial writes never reach the team.
[object Object], ,[object Object],(,[object Object],):
staged = {} ,[object Object],
,[object Object],:
result = run_agent_guarded(agent, task, validate)
store.commit(staged | result.writes) ,[object Object],
,[object Object], (,[object Object],, result)
,[object Object], AgentFailure ,[object Object], f:
,[object Object],
,[object Object], (,[object Object],, f)What this does: It stages an agent's writes in a private buffer and only commits them to shared state if the agent passes its guard. A failing agent leaves shared state exactly as it was, so one bad agent can't corrupt the team's working data - the single most important property in failure handling.
Step two, decide recovery based on failure kind, because the right response differs. A transient error deserves a retry. An "invalid output" failure usually deserves a retry with a corrective note. A loop deserves a hard stop and escalation - retrying a looping agent just loops again.
[object Object], ,[object Object],(,[object Object],):
,[object Object], kind ,[object Object], (,[object Object],, ,[object Object],) ,[object Object], attempt < ,[object Object],:
,[object Object], (,[object Object],, task)
,[object Object], kind == ,[object Object], ,[object Object], attempt < ,[object Object],:
note = ,[object Object],
,[object Object], (,[object Object],, task.with_note(note))
,[object Object], kind == ,[object Object],:
,[object Object], (,[object Object],, task) ,[object Object],
,[object Object], (,[object Object],, task)What this does: It routes recovery by failure kind. Transient failures retry, invalid outputs retry with a correction hint, and loops escalate immediately instead of retrying. This prevents the common anti-pattern of retrying a failure that will deterministically fail again.
Step three, escalate with context. When an agent can't recover, don't silently drop its task - hand it to a fallback (a different agent, a simpler model, or a human) with the failure detail so the fallback doesn't repeat the mistake.
There's a fourth failure mode the quick-start doesn't cover, and it's the one that ruins on-call nights: the stuck-but-not-crashed agent. It didn't time out (it's technically still producing tokens), it didn't loop (each step is different), and its output isn't invalid yet (it hasn't finished). It's just wandering - exploring, second-guessing, rewriting. A wall-clock timeout eventually catches it, but only after wasting the full budget. Better multi agent failure handling adds a progress check: is the agent getting measurably closer to done, or spinning?
[object Object], ,[object Object],(,[object Object],):
,[object Object],
,[object Object], ,[object Object],(history) < window:
,[object Object], ,[object Object],
recent = history[-window:]
,[object Object],
,[object Object], ,[object Object],(sim > ,[object Object], ,[object Object], sim ,[object Object], recent)What this does: It compares how similar each step's state is to the previous one. When the last few steps barely change - the agent keeps rephrasing the same partial answer - it flags a stall early, so you can intervene long before the timeout instead of paying for the full window of wandering.
Pro-Level Variations
For high-stakes pipelines, add a second validator agent - a critic - that independently checks outputs. Two different checks (code validator plus model critic) catch more silent failures than either alone, at the cost of extra tokens. Use it where a wrong answer is expensive.
For latency-sensitive systems, run a cheap fast agent and a slow reliable one in parallel, take the fast one if it validates, and fall back to the slow one only when the fast one fails its guard. This is hedged execution, and it bounds your worst-case latency.
For teams where one agent's failure should sometimes cancel the whole run - say, a safety check that, if it fails, means the entire output is unsafe to ship - wire failures into a shared cancellation signal. When the critical agent fails its guard, it flips a "stop" flag that every other agent reads before continuing, so the team abandons the run cleanly instead of wasting budget finishing work that will be thrown away. The distinction to design deliberately is which failures are local (contain and continue) versus fatal (stop everyone). Most teams treat every failure as local by default and then act surprised when a failed safety check didn't halt production - decide this per agent, in advance, and encode it rather than discovering it during an incident.
Getting that classification right is genuinely the heart of multi agent failure handling. A local failure that should have been fatal ships bad output; a fatal failure that should have been local turns a minor hiccup into a full outage. There's no universal rule - it depends on what each agent guarantees - so the useful practice is to annotate every agent with its failure blast radius when you add it, the same way you'd annotate a function with whether it can throw.
⚡ Pro tip: Log every failure with its kind and the input that caused it. After a week you'll have a failure taxonomy specific to your system, and you'll usually find that 80% of failures come from two or three input patterns you can handle explicitly - which beats generic retries every time.
Troubleshooting Common Issues
If retries seem to make things worse, you're probably retrying non-transient failures. Check that your recover logic distinguishes transient (retry) from deterministic (escalate). Retrying a deterministic failure burns tokens to fail identically.
If the whole team stalls when one agent fails, your agents are probably waiting on the failed one synchronously. Add per-dependency timeouts so a downstream agent proceeds with a "dependency unavailable" marker rather than hanging.
If your validator passes bad outputs, it's too shallow. A validator that only checks JSON shape will wave through a perfectly-formatted wrong answer. Strengthen it incrementally: every time a bad output reaches production, add the specific check that would have caught it. Over a few weeks the validator accretes real domain knowledge and becomes the most valuable piece of your failure handling - far more than the retry logic everyone obsesses over.
If failures cluster around one agent, the problem may be its job, not its reliability. An agent asked to do too much - research and analyze and write and format in one turn - fails more often and more ambiguously than three focused agents. Splitting an unreliable agent into narrower agents each with a crisp validator often does more for stability than any amount of retry tuning. Reliability is frequently a decomposition problem wearing a fault-tolerance costume.
⚡ Pro tip: Give each agent a "confidence to escalate" instruction - permission to stop and say "I'm not sure, hand this off" instead of guessing. Agents that can admit uncertainty fail loudly and early, which is exactly what you want. The dangerous agent is the one trained by its prompt to always produce a confident answer, because it converts "I don't know" into a plausible fabrication your validator then has to catch.
⚠️ Common mistake: Catching agent failures but not containing their side effects. If a failing agent already wrote to the shared store, tool, or database before you detected the failure, cleaning up the exception doesn't undo the damage. Stage all effects and commit only on success - detection without containment is a false sense of safety, because the corrupt write is already out there.
Your Turn
Take your riskiest agent - the one whose bad output would cost the most - and wrap it with run_agent_guarded plus a real validator today. Not a schema check; a check that a domain expert would agree means "this is actually correct." That one change catches more real failures than any retry logic.
⚡ Pro tip: Keep a running "failure museum" - a saved input for every distinct failure you've seen - and replay it against your guards after any change. It's the cheapest regression suite you'll ever build, and it stops you from re-shipping a failure you already fixed once, which is otherwise depressingly common as prompts drift.
Then version your validators and recovery prompts alongside your agent prompts. I keep the corrective-retry note and the critic prompt in PromptABCD, because the exact wording that gets an agent to fix a flagged field - without over-correcting everything else - takes iteration to land, and it's worth reusing verbatim once it works.
Continue Reading
Save the prompts from this post
PromptABCD is a free prompt manager. Paste, organize, and reuse your best AI prompts — no more hunting through chat history.
