Preventing Agents From Duplicating Work
Multi agent duplicate work quietly triples your token bill and produces contradictory results. Here's a real case study on how one team fixed it with claims and idempotency keys.
# Attempt 1: router picks one specialist. Brittle. ROUTER_PROMPT = """Read the ticket and output exactly one of: BILLING, TECHNICAL, ACCOUNT. Pick the single best fit."""
Picture this: you're an engineering lead at a mid-size SaaS company, and you just shipped an agent team that answers support tickets. It works in the demo. Then the first real batch of 200 tickets comes in, and your token dashboard spikes to 4x what you projected. You dig in and find that three agents each independently researched the same knowledge-base article for the same ticket, then wrote three slightly different answers, and a fourth agent had to pick one. That's multi agent duplicate work, and it's one of the most expensive failure modes in agent systems precisely because nothing crashes.
I want to walk through a real version of this - anonymized, but the shape is exactly as it happened - because the fix is instructive and it's not "add a mutex."
The Problem the Support Team Faced
The team ran a fan-out pattern. A router agent received a ticket, spawned several specialist agents (billing, technical, account), and each specialist did its own retrieval and drafted a response. A synthesizer merged the drafts.
The intent was good: parallel specialists are faster than one generalist doing everything in sequence. But there was no coordination about who does what. Every specialist re-read the ticket, re-searched the knowledge base, and often pulled the same three articles. Multi agent duplicate work showed up as redundant retrieval, redundant model calls, and contradictory drafts that the synthesizer had to reconcile.
The measurable symptoms: token spend per ticket was roughly triple the estimate, latency was worse than a single agent (because the synthesizer waited on all specialists), and about one in eight tickets got an answer that contradicted itself because two specialists made different assumptions about the same missing fact.
The Wrong Approach
The first instinct - and the team tried this - was to make the router smarter. Give it a better prompt so it only spawns the relevant specialist. This helped a little and created a new problem.
[object Object],
ROUTER_PROMPT = ,[object Object],What this does: It forces the router to choose one category so only one specialist runs. It reduces duplication by eliminating parallelism - which throws away the speed benefit and fails on tickets that genuinely span two areas (a billing problem caused by a technical bug).
The deeper issue is that "prevent duplication by never running agents in parallel" isn't a fix, it's a retreat. The team wanted parallel specialists. They just needed those specialists to not step on each other. Tightening the router also concentrated risk: one bad routing decision now meant zero relevant specialists ran, and the ticket got a useless answer.
⚡ Pro tip: When a coordination bug tempts you to remove concurrency entirely, stop. The goal is coordinated concurrency, not serialization. Serializing is almost always leaving performance on the table to paper over a design gap.
The Correct Prompt (and the Coordination Layer)
The fix had two parts: a claim system so agents reserve work before doing it, and idempotency keys so repeated work is deduplicated instead of re-executed.
First, the claim board. Before an agent does a unit of work (like "retrieve articles about refund policy"), it claims that unit. If another agent already claimed it, the second agent waits for the result instead of redoing it.
[object Object], ,[object Object],:
,[object Object], ,[object Object],(,[object Object],):
,[object Object],._claims = {} ,[object Object],
,[object Object],._lock = threading.Lock()
,[object Object], ,[object Object],(,[object Object],):
,[object Object], ,[object Object],._lock:
,[object Object], work_key ,[object Object], ,[object Object],._claims:
,[object Object], (,[object Object],, ,[object Object],._claims[work_key][,[object Object],])
fut = Future()
,[object Object],._claims[work_key] = (agent_id, fut)
,[object Object], (,[object Object],, fut)
,[object Object], ,[object Object],(,[object Object],):
,[object Object], ,[object Object],._lock:
,[object Object],._claims[work_key][,[object Object],].set_result(result)What this does: The first agent to claim a work_key gets a "go" and a future to fill in. Any later agent asking for the same key gets "wait" and the same future, so it receives the result the first agent produces instead of computing it again. One retrieval, shared by everyone who needs it.
Second, the specialist prompt changed from "answer the ticket" to "answer only your slice, and declare your slice up front." Declaring the slice is what makes work_keys stable across agents.
SPECIALIST_PROMPT = ,[object Object],What this does: It forces each specialist to name the data it depends on as explicit keys before working. Those keys feed the claim board, so two specialists needing the same fact resolve to one claim - eliminating the duplicated retrieval and the contradictory assumptions that came from each guessing independently.
Results and What Changed
After the claim board and keyed idempotency went in, token spend per ticket dropped from roughly 3x baseline to about 1.2x - the remaining overhead being genuine coordination cost, which is fine. Latency improved because specialists stopped redundantly hitting the knowledge base, and the contradictory-answer rate fell from around 12% to under 2%, because when two specialists needed the same fact, they now got the same fact.
⚡ Pro tip: Track a "dedup ratio" - claims served from cache divided by total claims. When I onboard a new agent team, that single number tells me more about coordination health than any latency graph. A ratio near zero means agents aren't sharing work; a healthy fan-out system sees meaningful reuse.
The subtle win, and the thing I didn't expect, was correctness. Everyone frames multi agent duplicate work as a cost problem. It's also a consistency problem. When two agents independently derive the same fact, they can derive it differently, and now your system holds two contradictory beliefs. Deduplication forces a single source of truth for each fact, which quietly kills a whole category of contradiction bugs.
How Do You Measure Multi Agent Duplicate Work?
You can't fix what you can't see, and duplicate work is nearly invisible in normal logs because every duplicated action succeeds. Nothing errors. So the first move, before any claim board, is instrumentation that makes duplication visible.
The cheapest useful measurement is a content-hash tally. Hash the inputs of every expensive operation - each retrieval query, each tool call, each model prompt - and count how often the same hash appears within a single run. A hash that shows up three times is three agents doing identical work.
[object Object], collections ,[object Object], Counter
,[object Object], hashlib
,[object Object], ,[object Object],(,[object Object],):
counts = Counter()
,[object Object], e ,[object Object], events:
key = hashlib.sha256(e[,[object Object],].encode()).hexdigest()[:,[object Object],]
counts[key] += ,[object Object],
repeated = {k: n ,[object Object], k, n ,[object Object], counts.items() ,[object Object], n > ,[object Object],}
waste = ,[object Object],(n - ,[object Object], ,[object Object], n ,[object Object], repeated.values())
,[object Object], {,[object Object],: ,[object Object],(repeated), ,[object Object],: waste}What this does: It hashes the inputs of every operation in a run and reports how many operations were repeats. The wasted_ops number is the count of calls you could have avoided - a direct, dollar-convertible measure of multi agent duplicate work that you can watch trend down as you add coordination.
Run this against a day of real traffic before you build anything. Teams are routinely shocked - I've seen waste ratios north of 50%, meaning half of all expensive operations were redundant. That number also tells you whether the problem is worth solving; if waste is 3%, a claim board is over-engineering and you should spend your effort elsewhere.
⚡ Pro tip: Hash inputs, never outputs, when detecting duplicate work. Two agents doing the identical retrieval may get slightly different model-rephrased outputs, so output hashes miss the duplication. The inputs are what reveal that the same work was requested twice.
How to Apply This to Your Situation
You don't need this on day one. Apply it when you see the symptoms: token spend scaling faster than task count, latency that gets worse as you add agents, or outputs that contradict themselves.
Start by making work claimable. Identify the units of work agents repeat - retrieval, tool calls, sub-computations - and give each a stable key derived from its inputs. Then put a claim board in front so the first agent to need a unit computes it and the rest subscribe to the result.
⚠️ Common mistake: Deriving work_keys from agent identity or timestamp instead of from the work's inputs. If the key includes who's asking or when, two agents doing identical work get different keys and the dedup never triggers. The key must be a pure function of the input - "kb:refund-policy" not "billing_agent:kb:refund-policy:1699". This is the single most common reason a claim system quietly does nothing.
Next Steps
Instrument first. Add the dedup ratio and per-ticket token count before you change anything, so you can prove the fix worked. Then introduce claiming for your single most-duplicated unit of work - usually retrieval - and expand from there.
One caution as you roll this out: don't try to deduplicate everything. Some apparent duplication is legitimate - two agents checking the same fact for independent verification, for instance, where you actually want two opinions. Deduplicate the expensive, deterministic work (retrieval, tool calls, computations that always return the same answer for the same input) and leave the judgment work alone. The claim board should cover facts, not opinions. Collapsing two independent judgments into one because they share an input defeats the reason you built a team in the first place.
⚡ Pro tip: When you introduce claiming, roll it out behind a flag and compare the dedup report before and after on the same replayed traffic. Duplicate-work fixes are easy to think you shipped and easy to have silently no-op because of a key-derivation bug, so prove the reduction against real numbers rather than trusting that the mechanism fired.
Also plan for claims that fail. If the agent holding a claim errors out mid-work, everyone waiting on that claim inherits the failure. Give claims a timeout and a "reclaim" path so a dead claim gets released and re-attempted by another agent rather than blocking every subscriber forever - otherwise your dedup layer becomes a new single point of failure.
Keep the specialist and router prompts under version control so you can diff the "declare your slice" behavior across iterations. I store these coordination prompts in PromptABCD and reuse the claim-declaration pattern across projects, because getting an agent to name its dependencies up front is the hard-won part, and it transfers cleanly from one agent team to the next.
Continue Reading
Save the prompts from this post
PromptABCD is a free prompt manager. Paste, organize, and reuse your best AI prompts — no more hunting through chat history.
