PromptABCD
FeaturesLearnGuideBlogContext Blocks
Sign inGet started free
Sign inSign up
PromptABCD

A calm home for your best AI prompts. Save them once, find them in seconds, reuse them forever.

Product

  • Features
  • Chrome Extension
  • Free Courses
  • How it works
  • Use cases
  • Blog
  • Context Blocks
  • Export Anywhere
  • FAQ

Resources

  • User guide
  • Learn prompting
  • Sign in
  • Get started free

© 2026 PromptABCD. All rights reserved.

AboutPrivacy PolicyTerms and Conditions
Home/Blog/Multi-Agent Systems/Timeouts and Circuit Breakers for Agent Teams
Multi-Agent Systems

Timeouts and Circuit Breakers for Agent Teams

Most reliability advice says 'add retries.' For agent teams that's backwards - a multi agent circuit breaker that stops trying is often what saves you. Here's the case study.

September 25, 2026·9 min read
ShareShare
⚡Featured Prompt— copy and use right now
# Attempt 1: gentler retries. Still fundamentally wrong under load.
def call_search(query):
    for attempt in range(3):
        try:
            return search_api(query, timeout=8)
        except SlowResponse:
            sleep(2 ** attempt)   # backoff, but still retrying into a fire
    raise SearchFailed()

Most reliability advice for agents says "add retries." For agent teams, that advice is often exactly backwards. The thing that saves a fleet under stress isn't trying harder - it's a multi agent circuit breaker that stops trying when a dependency is failing, so your agents don't pile onto a struggling service and turn a partial outage into a total one. I've watched retries turn a two-minute blip into a twenty-minute outage, and I've watched a circuit breaker turn the same blip into a barely-noticed dip.

Here's a case study where the counterintuitive move - failing fast and refusing to retry - was what restored stability. It's a pattern worth internalizing because the instinct to retry is strong and, for fleets, frequently wrong.

The Problem This Team Faced

A team ran a fleet of research agents, each of which called an external search API as part of its work. Dozens of agent instances, all depending on that one API. One afternoon the search API got slow - not down, just slow, responses climbing from 200 milliseconds to several seconds.

Here's where it cascaded. Each agent had a generous timeout and a retry policy: if a call was slow, wait, then retry. So as the API slowed, every agent started waiting longer and retrying. The retries added load to the already-struggling API, which made it slower, which triggered more waiting and more retries. Within minutes the search API went from slow to effectively unusable, and the entire research fleet was frozen - not because any agent had failed, but because they'd collectively hammered a wobbling dependency into the ground.

The symptom was a total fleet stall from a dependency that was only degraded, not down. The agents' own retry logic manufactured the outage.

The Wrong Approach

The team's first reaction was, predictably, to tune the retries - fewer attempts, longer backoff.

python
[object Object],
,[object Object], ,[object Object],(,[object Object],):
    ,[object Object], attempt ,[object Object], ,[object Object],(,[object Object],):
        ,[object Object],:
            ,[object Object], search_api(query, timeout=,[object Object],)
        ,[object Object], SlowResponse:
            sleep(,[object Object], ** attempt)   ,[object Object],
    ,[object Object], SearchFailed()

What this does: It reduces retry aggression with exponential backoff. It still misses the core problem - every agent independently decides to keep trying a dependency that's failing for everyone. Gentler retries slow the pile-on but don't stop it, because there's no shared awareness that the dependency is in trouble.

The flaw is that each agent makes its retry decision in isolation. Agent A doesn't know that agents B through Z are also retrying the same failing API. So even polite per-agent retries sum to an impolite fleet-wide assault. You can't fix a coordination problem with per-agent tuning; the agents need to share the knowledge that the dependency is down and collectively back off.

⚡ Pro tip: When many agents share a dependency, per-agent retry tuning can't save you, because the problem is the sum of everyone's retries. You need shared state that all agents consult - a circuit breaker - so that when the dependency is failing, the whole fleet backs off together instead of each agent discovering the failure independently and adding to it.

The Correct Pattern

The fix was a shared multi agent circuit breaker: one piece of state, visible to every agent, tracking the health of the shared dependency. When failures cross a threshold, the breaker "opens" and every agent immediately fails fast - skipping the call entirely - for a cooldown period. This gives the dependency room to recover instead of being pounded.

python
[object Object], ,[object Object],:
    ,[object Object], ,[object Object],(,[object Object],):
        ,[object Object],.failures = ,[object Object],
        ,[object Object],.state = ,[object Object],     ,[object Object],
        ,[object Object],.opened_at = ,[object Object],
        ,[object Object],.threshold, ,[object Object],.cooldown = threshold, cooldown

    ,[object Object], ,[object Object],(,[object Object],):
        ,[object Object], ,[object Object],.state == ,[object Object],:
            ,[object Object], time.time() - ,[object Object],.opened_at < ,[object Object],.cooldown:
                ,[object Object], CircuitOpen()          ,[object Object],
            ,[object Object],.state = ,[object Object],         ,[object Object],
        ,[object Object],:
            result = fn(*args)
            ,[object Object],._on_success()
            ,[object Object], result
        ,[object Object], Exception:
            ,[object Object],._on_failure()
            ,[object Object],

    ,[object Object], ,[object Object],(,[object Object],):
        ,[object Object],.failures += ,[object Object],
        ,[object Object], ,[object Object],.failures >= ,[object Object],.threshold:
            ,[object Object],.state, ,[object Object],.opened_at = ,[object Object],, time.time()

    ,[object Object], ,[object Object],(,[object Object],):
        ,[object Object],.failures = ,[object Object],
        ,[object Object],.state = ,[object Object],

What this does: It tracks failures against a shared dependency and, once they cross a threshold, flips to "open" - where every agent fails fast without calling the dependency at all. After a cooldown it goes "half-open" to test whether the dependency recovered, closing fully on success. The failing-fast is the point: it stops the fleet from hammering a service that needs breathing room.

The other half was fixing timeouts. The generous eight-second timeout meant a slow dependency tied up agents for eight seconds each. They cut it to a value just above the normal response time, so a slow dependency fails quickly instead of holding agents hostage.

python
[object Object],
NORMAL_P95 = ,[object Object],              ,[object Object],
SEARCH_TIMEOUT = NORMAL_P95 * ,[object Object],   ,[object Object],

What this does: It derives the timeout from measured normal latency rather than a generous guess. A call that takes three times the usual p95 is treated as failed quickly, so a degraded dependency stops consuming agent time almost immediately instead of after eight seconds per attempt.

Results and What Changed

The next time the search API degraded, the outcome was completely different. Failures crossed the threshold, the shared breaker opened, and the whole fleet stopped calling the API within seconds - falling back to a cached-results path. With the load removed, the API recovered in about a minute, the breaker went half-open, tested successfully, and closed. Total user impact: a brief period of slightly staler search results. No fleet stall.

The insight the team took away, and the reason this is worth a case study: for a shared dependency, not calling is a first-class strategy. Retrying is what you do when a failure is likely transient and isolated. When a failure is systemic and shared, the correct move is to stop, collectively, and let the thing recover. The circuit breaker encodes that "stop" as shared state so the whole fleet does it in unison.

There was a secondary lesson too, about the fallback. A circuit breaker that opens onto nothing - that just fails fast with an error - protects the dependency but not the user, who now gets errors instead of slow responses. The breaker is only half the pattern; the other half is a meaningful fallback for when it's open. For this team, cached search results were good enough - slightly stale but useful. The general principle is that every breaker needs a defined open-state behavior that degrades gracefully, whether that's a cache, a simpler computation, or an honest "this feature is briefly unavailable." Failing fast into a void is better than a cascade, but failing fast into a fallback is better still.

⚡ Pro tip: Pair every circuit breaker with a fallback that produces a usable-if-degraded result, not just a fast error. The breaker's job is to protect the dependency; the fallback's job is to protect the user. A breaker without a fallback trades a slow system for a broken one, which your users may not thank you for. Design the open-state behavior with as much care as the threshold.

⚡ Pro tip: Set timeouts from measured latency percentiles, never from a comfortable-feeling round number. A timeout should sit just above normal behavior so abnormal behavior fails fast. An eight-second timeout on a call that normally takes 400 milliseconds isn't safety margin - it's an invitation for a slow dependency to hold your agents hostage seven and a half seconds at a time.

How to Apply This to Your Situation

Find every external dependency your agent fleet shares - APIs, databases, tool servers - and wrap each in a shared circuit breaker. Shared is the key word: one breaker per dependency, consulted by all agents, so the fleet backs off in unison. Set each breaker's threshold from how many failures genuinely indicate trouble, and its cooldown from how long the dependency typically needs to recover.

Then audit your timeouts against real latency data. Most fleets that stall under load have timeouts set too generously, so degraded dependencies drain agent time slowly instead of failing fast. Tighten them to just above measured normal latency.

One detail that trips people up: the half-open state needs care. When the breaker tests recovery, let one request through, not the whole fleet. If you reopen the floodgates and send every waiting agent's request the instant the cooldown ends, you re-overwhelm a dependency that was just starting to breathe, and the breaker flaps open and closed. A single probe request decides whether to close; everyone else keeps failing fast until that probe succeeds.

python
[object Object], ,[object Object],(,[object Object],):
    ,[object Object],
    ,[object Object], ,[object Object],._probe_in_flight:
        ,[object Object], CircuitOpen()
    ,[object Object],._probe_in_flight = ,[object Object],

What this does: It ensures that when the breaker goes half-open, exactly one request tests the dependency while all others continue failing fast. This prevents the thundering-herd-on-recovery problem where a fleet's worth of queued requests all hit the barely-recovered dependency at once and knock it back down.

⚡ Pro tip: Guard the half-open probe so only one request tests recovery. The failure mode of a naive breaker isn't opening - it's the flapping that happens when recovery lets the whole fleet back in simultaneously. A single-probe half-open state is what makes the difference between a breaker that stabilizes a dependency and one that oscillates it.

⚠️ Common mistake: Giving each agent its own private circuit breaker instead of sharing one per dependency. Private breakers don't coordinate - each agent has to independently rack up failures before its own breaker opens, so the fleet still collectively hammers the dependency while each agent "learns" it's down. The breaker's power comes from being shared; a per-agent breaker throws away the coordination that makes it work.

Next Steps

Wrap your most-shared dependency in a single fleet-wide breaker, set its timeout from measured p95, and give it a fallback path for when it's open. Watch it during the next dependency hiccup and you'll see the fleet back off together instead of piling on.

Keep your breaker thresholds and fallback prompts versioned so you can reuse a reliability configuration you trust. I store the fallback prompts - the "breaker is open, use cached results" behavior - in PromptABCD, because a good degraded-mode response is something you tune once and want to reproduce everywhere, and reusing a proven multi agent circuit breaker configuration beats rediscovering the right thresholds during your next incident.

multi-agent-systemscircuit-breakertimeoutsreliabilityai-agents

Continue Reading

A Reusable Prompt Kit for Agent Teams
Multi-Agent Systems

A Reusable Prompt Kit for Agent Teams

A team rebuilt their agent prompts from memory every project, and every project drifted a little worse. That failure is why a multi agent prompt kit matters. Here's the reusable set of role prompts every team should keep.

October 2, 2026·9 min read
Multi-Agent Systems on a Budget: An Interactive Guide
Multi-Agent Systems

Multi-Agent Systems on a Budget: An Interactive Guide

Most multi-agent tutorials assume you'll burn tokens freely. That's wrong for anyone shipping on real constraints. A cheap multi agent system can match an expensive one with the right moves. Here's how to build one.

October 2, 2026·9 min read
The Manager Agent Anti-Pattern: A Teardown
Multi-Agent Systems

The Manager Agent Anti-Pattern: A Teardown

Why does your orchestrator agent keep becoming a bottleneck that mangles every handoff? You've hit the manager agent anti pattern — one agent trying to coordinate everything. Here's why it fails and what replaces it.

October 2, 2026·9 min read

Save the prompts from this post

PromptABCD is a free prompt manager. Paste, organize, and reuse your best AI prompts — no more hunting through chat history.

Start free →
← PreviousLoad Balancing Across Agent InstancesNext →Observability for Multi-Agent Systems
Share this post:
ShareShare