PromptABCD
FeaturesLearnGuideBlogContext Blocks
Sign inGet started free
Sign inSign up
PromptABCD

A calm home for your best AI prompts. Save them once, find them in seconds, reuse them forever.

Product

  • Features
  • Chrome Extension
  • Free Courses
  • How it works
  • Use cases
  • Blog
  • Context Blocks
  • Export Anywhere
  • FAQ

Resources

  • User guide
  • Learn prompting
  • Sign in
  • Get started free

© 2026 PromptABCD. All rights reserved.

AboutPrivacy PolicyTerms and Conditions
Home/Blog/Multi-Agent Systems/How to Handle a Failing Agent in a Team
Multi-Agent Systems

How to Handle a Failing Agent in a Team

Most multi agent failure handling guides are wrong: they treat a failed agent like a crashed server. Agents fail differently. Here's a copy-paste guide to detecting and containing them.

September 23, 2026·9 min read
ShareShare
⚡Featured Prompt— copy and use right now
import time

class AgentFailure(Exception):
    def __init__(self, kind, detail):
        self.kind = kind      # "error" | "timeout" | "invalid" | "loop"
        self.detail = detail

def run_agent_guarded(agent, task, validate, timeout=60, max_steps=8):
    start = time.time()
    steps = 0
    for output in agent.stream(task):
        steps += 1
        if time.time() - start > timeout:
            raise AgentFailure("timeout", f"exceeded {timeout}s")
        if steps > max_steps:
            raise AgentFailure("loop", f"exceeded {max_steps} steps")
    ok, reason = validate(output)
    if not ok:
        raise AgentFailure("invalid", reason)
    return output

Most multi agent failure handling guides are wrong in the same way: they treat a failing agent like a crashed microservice. Restart it, retry the call, move on. But agents don't fail like servers. A server that fails throws an error and stops. An agent that "fails" often keeps going confidently - producing plausible, wrong output, or looping, or quietly ignoring half its instructions. The scariest agent failures return HTTP 200.

So detection is the hard part, not recovery. This guide gives you code you can paste in today to catch the failures that don't announce themselves, then contain them before they poison the rest of the team.

Quick-Start (Copy This Right Now)

Here's a wrapper that catches the three failure classes agents actually exhibit: hard errors, timeouts, and - the important one - silent bad output that passes a validity check.

python
[object Object], time

,[object Object], ,[object Object],(,[object Object],):
    ,[object Object], ,[object Object],(,[object Object],):
        ,[object Object],.kind = kind      ,[object Object],
        ,[object Object],.detail = detail

,[object Object], ,[object Object],(,[object Object],):
    start = time.time()
    steps = ,[object Object],
    ,[object Object], output ,[object Object], agent.stream(task):
        steps += ,[object Object],
        ,[object Object], time.time() - start > timeout:
            ,[object Object], AgentFailure(,[object Object],, ,[object Object],)
        ,[object Object], steps > max_steps:
            ,[object Object], AgentFailure(,[object Object],, ,[object Object],)
    ok, reason = validate(output)
    ,[object Object], ,[object Object], ok:
        ,[object Object], AgentFailure(,[object Object],, reason)
    ,[object Object], output

What this does: It runs an agent under three simultaneous guards - a wall-clock timeout, a step ceiling to catch loops, and a validator that inspects the final output. The validator is what separates this from naive error handling: it catches the confident-but-wrong outputs that never raise an exception on their own.

Understanding the Variables

The pieces you'll tune are validate, timeout, and max_steps, and they map to the three failure modes.

validate is the most important and the most neglected. It's a function that returns whether the output is actually usable, not just well-formed. For a JSON-emitting agent, "valid" means the schema matches AND the values are sane - a price of -$4,000,000 is well-formed and obviously broken. Good multi agent failure handling starts here, because this is the only guard that catches silent failures.

timeout bounds wall-clock time. Set it from real percentiles, not a guess - measure your p95 successful run and add headroom. Too tight and you kill slow-but-correct agents; too loose and hung agents waste your budget.

max_steps catches loops - an agent calling the same tool repeatedly, or two agents ping-ponging. Agents loop more than servers do because a confused model often "tries again" instead of erroring.

⚡ Pro tip: Write the validator before you write the agent. If you can't specify what a good output looks like precisely enough to check it in code, your agent doesn't have a clear enough job - and you'll never be able to tell success from failure at runtime.

Step-by-Step: Multi Agent Failure Handling

Here's the full flow, from detection to containment to recovery, as a coordinator would run it.

Step one, isolate the failing agent's effects. Before an agent's output is allowed to touch shared state, it must pass its guard. A failed agent's partial writes never reach the team.

python
[object Object], ,[object Object],(,[object Object],):
    staged = {}   ,[object Object],
    ,[object Object],:
        result = run_agent_guarded(agent, task, validate)
        store.commit(staged | result.writes)   ,[object Object],
        ,[object Object], (,[object Object],, result)
    ,[object Object], AgentFailure ,[object Object], f:
        ,[object Object],
        ,[object Object], (,[object Object],, f)

What this does: It stages an agent's writes in a private buffer and only commits them to shared state if the agent passes its guard. A failing agent leaves shared state exactly as it was, so one bad agent can't corrupt the team's working data - the single most important property in failure handling.

Step two, decide recovery based on failure kind, because the right response differs. A transient error deserves a retry. An "invalid output" failure usually deserves a retry with a corrective note. A loop deserves a hard stop and escalation - retrying a looping agent just loops again.

python
[object Object], ,[object Object],(,[object Object],):
    ,[object Object], kind ,[object Object], (,[object Object],, ,[object Object],) ,[object Object], attempt < ,[object Object],:
        ,[object Object], (,[object Object],, task)
    ,[object Object], kind == ,[object Object], ,[object Object], attempt < ,[object Object],:
        note = ,[object Object],
        ,[object Object], (,[object Object],, task.with_note(note))
    ,[object Object], kind == ,[object Object],:
        ,[object Object], (,[object Object],, task)   ,[object Object],
    ,[object Object], (,[object Object],, task)

What this does: It routes recovery by failure kind. Transient failures retry, invalid outputs retry with a correction hint, and loops escalate immediately instead of retrying. This prevents the common anti-pattern of retrying a failure that will deterministically fail again.

Step three, escalate with context. When an agent can't recover, don't silently drop its task - hand it to a fallback (a different agent, a simpler model, or a human) with the failure detail so the fallback doesn't repeat the mistake.

There's a fourth failure mode the quick-start doesn't cover, and it's the one that ruins on-call nights: the stuck-but-not-crashed agent. It didn't time out (it's technically still producing tokens), it didn't loop (each step is different), and its output isn't invalid yet (it hasn't finished). It's just wandering - exploring, second-guessing, rewriting. A wall-clock timeout eventually catches it, but only after wasting the full budget. Better multi agent failure handling adds a progress check: is the agent getting measurably closer to done, or spinning?

python
[object Object], ,[object Object],(,[object Object],):
    ,[object Object],
    ,[object Object], ,[object Object],(history) < window:
        ,[object Object], ,[object Object],
    recent = history[-window:]
    ,[object Object],
    ,[object Object], ,[object Object],(sim > ,[object Object], ,[object Object], sim ,[object Object], recent)

What this does: It compares how similar each step's state is to the previous one. When the last few steps barely change - the agent keeps rephrasing the same partial answer - it flags a stall early, so you can intervene long before the timeout instead of paying for the full window of wandering.

Pro-Level Variations

For high-stakes pipelines, add a second validator agent - a critic - that independently checks outputs. Two different checks (code validator plus model critic) catch more silent failures than either alone, at the cost of extra tokens. Use it where a wrong answer is expensive.

For latency-sensitive systems, run a cheap fast agent and a slow reliable one in parallel, take the fast one if it validates, and fall back to the slow one only when the fast one fails its guard. This is hedged execution, and it bounds your worst-case latency.

For teams where one agent's failure should sometimes cancel the whole run - say, a safety check that, if it fails, means the entire output is unsafe to ship - wire failures into a shared cancellation signal. When the critical agent fails its guard, it flips a "stop" flag that every other agent reads before continuing, so the team abandons the run cleanly instead of wasting budget finishing work that will be thrown away. The distinction to design deliberately is which failures are local (contain and continue) versus fatal (stop everyone). Most teams treat every failure as local by default and then act surprised when a failed safety check didn't halt production - decide this per agent, in advance, and encode it rather than discovering it during an incident.

Getting that classification right is genuinely the heart of multi agent failure handling. A local failure that should have been fatal ships bad output; a fatal failure that should have been local turns a minor hiccup into a full outage. There's no universal rule - it depends on what each agent guarantees - so the useful practice is to annotate every agent with its failure blast radius when you add it, the same way you'd annotate a function with whether it can throw.

⚡ Pro tip: Log every failure with its kind and the input that caused it. After a week you'll have a failure taxonomy specific to your system, and you'll usually find that 80% of failures come from two or three input patterns you can handle explicitly - which beats generic retries every time.

Troubleshooting Common Issues

If retries seem to make things worse, you're probably retrying non-transient failures. Check that your recover logic distinguishes transient (retry) from deterministic (escalate). Retrying a deterministic failure burns tokens to fail identically.

If the whole team stalls when one agent fails, your agents are probably waiting on the failed one synchronously. Add per-dependency timeouts so a downstream agent proceeds with a "dependency unavailable" marker rather than hanging.

If your validator passes bad outputs, it's too shallow. A validator that only checks JSON shape will wave through a perfectly-formatted wrong answer. Strengthen it incrementally: every time a bad output reaches production, add the specific check that would have caught it. Over a few weeks the validator accretes real domain knowledge and becomes the most valuable piece of your failure handling - far more than the retry logic everyone obsesses over.

If failures cluster around one agent, the problem may be its job, not its reliability. An agent asked to do too much - research and analyze and write and format in one turn - fails more often and more ambiguously than three focused agents. Splitting an unreliable agent into narrower agents each with a crisp validator often does more for stability than any amount of retry tuning. Reliability is frequently a decomposition problem wearing a fault-tolerance costume.

⚡ Pro tip: Give each agent a "confidence to escalate" instruction - permission to stop and say "I'm not sure, hand this off" instead of guessing. Agents that can admit uncertainty fail loudly and early, which is exactly what you want. The dangerous agent is the one trained by its prompt to always produce a confident answer, because it converts "I don't know" into a plausible fabrication your validator then has to catch.

⚠️ Common mistake: Catching agent failures but not containing their side effects. If a failing agent already wrote to the shared store, tool, or database before you detected the failure, cleaning up the exception doesn't undo the damage. Stage all effects and commit only on success - detection without containment is a false sense of safety, because the corrupt write is already out there.

Your Turn

Take your riskiest agent - the one whose bad output would cost the most - and wrap it with run_agent_guarded plus a real validator today. Not a schema check; a check that a domain expert would agree means "this is actually correct." That one change catches more real failures than any retry logic.

⚡ Pro tip: Keep a running "failure museum" - a saved input for every distinct failure you've seen - and replay it against your guards after any change. It's the cheapest regression suite you'll ever build, and it stops you from re-shipping a failure you already fixed once, which is otherwise depressingly common as prompts drift.

Then version your validators and recovery prompts alongside your agent prompts. I keep the corrective-retry note and the critic prompt in PromptABCD, because the exact wording that gets an agent to fix a flagged field - without over-correcting everything else - takes iteration to land, and it's worth reusing verbatim once it works.

multi-agent-systemsfailure-handlingreliabilityfault-toleranceai-agents

Continue Reading

A Reusable Prompt Kit for Agent Teams
Multi-Agent Systems

A Reusable Prompt Kit for Agent Teams

A team rebuilt their agent prompts from memory every project, and every project drifted a little worse. That failure is why a multi agent prompt kit matters. Here's the reusable set of role prompts every team should keep.

October 2, 2026·9 min read
Multi-Agent Systems on a Budget: An Interactive Guide
Multi-Agent Systems

Multi-Agent Systems on a Budget: An Interactive Guide

Most multi-agent tutorials assume you'll burn tokens freely. That's wrong for anyone shipping on real constraints. A cheap multi agent system can match an expensive one with the right moves. Here's how to build one.

October 2, 2026·9 min read
The Manager Agent Anti-Pattern: A Teardown
Multi-Agent Systems

The Manager Agent Anti-Pattern: A Teardown

Why does your orchestrator agent keep becoming a bottleneck that mangles every handoff? You've hit the manager agent anti pattern — one agent trying to coordinate everything. Here's why it fails and what replaces it.

October 2, 2026·9 min read

Save the prompts from this post

PromptABCD is a free prompt manager. Paste, organize, and reuse your best AI prompts — no more hunting through chat history.

Start free →
← PreviousDeadlocks in Multi-Agent SystemsNext →Retry and Fallback Across Agents
Share this post:
ShareShare