PromptABCD
FeaturesLearnGuideBlogContext Blocks
Sign inGet started free
Sign inSign up
PromptABCD

A calm home for your best AI prompts. Save them once, find them in seconds, reuse them forever.

Product

  • Features
  • Chrome Extension
  • Free Courses
  • How it works
  • Use cases
  • Blog
  • Context Blocks
  • Export Anywhere
  • FAQ

Resources

  • User guide
  • Learn prompting
  • Sign in
  • Get started free

© 2026 PromptABCD. All rights reserved.

AboutPrivacy PolicyTerms and Conditions
Home/Blog/Multi-Agent Systems/Deadlocks in Multi-Agent Systems
Multi-Agent Systems

Deadlocks in Multi-Agent Systems

Why did your agent team just freeze with no error? Multi agent deadlock is usually the answer. This teardown shows the exact prompt pattern that causes it and how to break the cycle.

September 23, 2026·9 min read
ShareShare
⚡Featured Prompt— copy and use right now
REVIEWER_PROMPT = """You review draft sections. If a section lacks
sources, ask the researcher for sources and WAIT for their reply
before continuing. Do not proceed without sources."""

RESEARCHER_PROMPT = """You find sources. Before researching, confirm
the reviewer has approved the section scope. WAIT for the reviewer's
approval before you begin searching."""

Why does your agent team sometimes just... stop? No error, no crash, no output. Every agent shows "waiting," the clock runs, and eventually something times out at the top level with a useless message. If you've built anything beyond two agents, you've probably hit this, and you've probably blamed the model. It usually isn't the model. It's a multi agent deadlock - a cycle where agents are each waiting on each other, and none can move.

Let's tear down a real one, because deadlock in agent systems looks different from the textbook version, and the fix is more about prompt design than lock ordering.

Before: The Weak Prompt

Here's the setup that produced a multi agent deadlock in a document-processing team. Two agents, a "reviewer" and a "researcher," were told to collaborate.

python
REVIEWER_PROMPT = ,[object Object],

RESEARCHER_PROMPT = ,[object Object],

What this does: It makes the reviewer wait for sources before approving, and the researcher wait for approval before finding sources. Read those two together and the trap is obvious in hindsight - each is waiting for an action only the other can take first.

Why It Fails

This is the classic circular-wait condition, dressed in natural language. The reviewer holds "approval" and wants "sources." The researcher holds "sources" and wants "approval." Neither will release what it has until it gets what it wants. In operating systems you'd draw this as a cycle in a wait-for graph. In an agent system, it manifests as two agents politely deferring to each other forever.

What makes agent deadlocks sneakier than thread deadlocks is that the "locks" are implicit and semantic. There's no mutex.acquire() you can grep for. The waiting is encoded in instructions like "wait for approval" and "confirm before proceeding." So the deadlock is invisible in your code and only emerges from the combination of two reasonable-sounding prompts.

⚡ Pro tip: Any time two agent prompts both contain a "wait for the other" clause, you have a latent deadlock. Search your prompts for words like "wait," "confirm first," "do not proceed until," and map who waits on whom. If you can draw a cycle, you have a bug that will trigger under the right inputs.

There's a second failure mode layered on top. Even without an explicit "wait," agents deadlock on shared resources. If Agent A holds a lock on document X and needs document Y, while Agent B holds Y and needs X, you get the same cycle - now over data instead of turns. I've seen this with agents that each grab an exclusive claim on a file before editing.

⚠️ Common mistake: Assuming a top-level timeout "handles" deadlocks. A timeout ends the symptom - the frozen run - but it wastes the entire time budget first, produces no partial result, and gives you no signal about why it hung. Timeouts are a safety net, not a fix. If your only deadlock defense is a 5-minute ceiling, every deadlock costs you 5 minutes and a blank output.

After: The Improved Prompt

The fix breaks the cycle by removing the mutual wait. You do that by establishing a strict ordering and giving each agent a way to make progress without the other. Here's the rewrite:

python
REVIEWER_PROMPT = ,[object Object],

RESEARCHER_PROMPT = ,[object Object],

What this does: It orders the interaction - scope approval always happens first and never blocks - so the researcher always has a clear starting signal. And it gives both agents an escape hatch ("proceed with a stated assumption," "review with what exists and note the gap") so neither can wait forever. The cycle is broken because the wait is now bounded and one-directional.

Breaking Down Each Element

Three specific changes did the work, and each maps to a known deadlock-prevention technique.

The two-pass split enforces a resource ordering. In OS terms, if every agent acquires resources in the same global order, you can't form a cycle. Scope-before-sources is exactly that: a fixed order that makes the circular wait impossible.

The "never wait indefinitely" clauses add timeouts at the agent level, not just the top level. Each agent has a bounded wait and a defined fallback. When the wait expires, the agent produces a partial, honest result ("reviewed without full sources, flagged the gap") instead of hanging. Now a stuck dependency degrades output quality slightly rather than freezing the whole run.

The "make a reasonable assumption and state it" clause removes the need for confirmation round-trips, which are where most agent deadlocks hide. Every "confirm before you proceed" is a potential wait edge. Replacing confirmation with "assume, declare, proceed" cuts those edges while keeping the assumption visible for later correction.

⚡ Pro tip: Give every blocking agent interaction a deadline and a defined fallback behavior. "Wait for X, but if X isn't here in N seconds, do Y and flag it" turns a hard freeze into graceful degradation. The flag matters as much as the fallback - it tells downstream agents the result is partial.

Variations for Different Contexts

For hierarchical teams (a manager agent directing workers), deadlocks are rarer because communication flows through one coordinator - but you can still deadlock if a worker waits on a sibling. Route all cross-worker requests through the manager so the manager can detect and break cycles centrally.

For peer swarms with no coordinator, add a "no circular wait" invariant enforced in code: before an agent blocks waiting on another, check the wait-for graph for a cycle and refuse the wait if one would form.

python
[object Object], ,[object Object],(,[object Object],):
    ,[object Object],
    seen, stack = ,[object Object],(), [target]
    ,[object Object], stack:
        node = stack.pop()
        ,[object Object], node == waiter:
            ,[object Object], ,[object Object],   ,[object Object],
        ,[object Object], node ,[object Object], seen:
            ,[object Object],
        seen.add(node)
        stack.extend(wait_graph.get(node, []))
    ,[object Object], ,[object Object],

What this does: It walks the wait-for graph before allowing a new wait edge. If adding "waiter is blocked on target" would let you reach the waiter again, that's a cycle, and the function returns False so the caller can take the fallback path instead of deadlocking.

For pipelines that mix human approval into the agent flow, the deadlock risk multiplies, because humans are the slowest and least predictable resource in the graph. An agent waiting on a human who's waiting on the agent's output is a real deadlock I've watched happen - the human wanted to see the draft before approving scope, the agent wanted approved scope before drafting. Treat human steps exactly like agent steps in your wait-for analysis: give them deadlines and fallbacks too. "If the human hasn't approved scope in an hour, proceed with the default scope and mark the run for later review" keeps the pipeline moving instead of parking it indefinitely on someone's lunch break.

A last variation worth naming: livelock, deadlock's quieter twin. Instead of freezing, two agents keep reacting to each other without converging - Agent A revises to satisfy B, which makes B revise to satisfy A, forever. No one is blocked; everyone is busy; nothing finishes. The fix is the same family of tools - a step ceiling and a "good enough, stop" threshold - so agents settle rather than chase each other around the ring.

How Do You Detect a Multi Agent Deadlock in Production?

Prevention is better, but you also need to catch a multi agent deadlock when one slips through, because a hung team looks identical to a slow team from the outside. The trick is a heartbeat plus a wait-for graph you can inspect at runtime.

Every agent emits a heartbeat with what it's doing and what it's blocked on. A monitor watches those heartbeats. When several agents report "blocked on X" for longer than any single operation should take, and their blocked-on targets form a cycle, you've found a deadlock - and you know exactly which agents are in it.

python
[object Object], ,[object Object],(,[object Object],):
    ,[object Object],
    ,[object Object], start ,[object Object], heartbeats:
        node, seen = start, ,[object Object],()
        ,[object Object], node ,[object Object], ,[object Object], ,[object Object],:
            ,[object Object], node ,[object Object], seen:
                ,[object Object], ,[object Object],(seen)   ,[object Object],
            seen.add(node)
            node = heartbeats.get(node, {}).get(,[object Object],)
    ,[object Object], ,[object Object],

What this does: It follows each agent's "blocked on" pointer through the heartbeat data. If following the chain returns to an agent already in the path, that set of agents is a live deadlock cycle - reported by name so your monitor can break it (by forcing a fallback on one member) instead of waiting for a top-level timeout.

The reason this beats a plain timeout is that it identifies the deadlock while there's still time budget left, and it names the participants. You can break the cycle deterministically - pick the lowest-priority agent in the cycle, force it down its fallback path, and the whole team unblocks - instead of losing the entire run and every downstream result.

⚡ Pro tip: Have your monitor break a detected cycle by victimizing one specific agent, not by killing the run. Choose the victim by a stable rule (lowest priority, or most-recently-started) so the choice is deterministic and reproducible in tests. Non-deterministic cycle-breaking is nearly impossible to debug.

Save and Reuse This

The reusable core here isn't any single prompt - it's the pattern: fixed resource ordering, bounded waits with declared fallbacks, and "assume and state" instead of "confirm and wait." That combination prevents the vast majority of multi agent deadlock situations I've encountered.

⚡ Pro tip: Write a tiny test that constructs your agents' wait-for graph from their prompts and asserts it's acyclic, and run it in CI. Deadlocks are combinatorial - they emerge from prompt pairs, not single prompts - so a human reviewing one prompt at a time will miss them. A machine checking the whole graph will not.

Keep the two-pass reviewer prompt and the "assume, declare, proceed" clause somewhere you can drop them into new agents. I save both in PromptABCD as building blocks, because the phrasing that reliably stops an agent from waiting forever is fiddly to get right, and once you have it, you want to reuse the exact wording rather than risk reintroducing a cycle.

multi-agent-systemsdeadlockcoordinationtimeoutsai-agents

Continue Reading

A Reusable Prompt Kit for Agent Teams
Multi-Agent Systems

A Reusable Prompt Kit for Agent Teams

A team rebuilt their agent prompts from memory every project, and every project drifted a little worse. That failure is why a multi agent prompt kit matters. Here's the reusable set of role prompts every team should keep.

October 2, 2026·9 min read
Multi-Agent Systems on a Budget: An Interactive Guide
Multi-Agent Systems

Multi-Agent Systems on a Budget: An Interactive Guide

Most multi-agent tutorials assume you'll burn tokens freely. That's wrong for anyone shipping on real constraints. A cheap multi agent system can match an expensive one with the right moves. Here's how to build one.

October 2, 2026·9 min read
The Manager Agent Anti-Pattern: A Teardown
Multi-Agent Systems

The Manager Agent Anti-Pattern: A Teardown

Why does your orchestrator agent keep becoming a bottleneck that mangles every handoff? You've hit the manager agent anti pattern — one agent trying to coordinate everything. Here's why it fails and what replaces it.

October 2, 2026·9 min read

Save the prompts from this post

PromptABCD is a free prompt manager. Paste, organize, and reuse your best AI prompts — no more hunting through chat history.

Start free →
← PreviousPreventing Agents From Duplicating WorkNext →How to Handle a Failing Agent in a Team
Share this post:
ShareShare