Multi-Agent Systems for Software Development
Picture an agent that plans, another that codes, and a third that reviews before anything merges. Multi agent software development is real now — here's how the roles actually divide the work and where the seams break.
from dataclasses import dataclass
@dataclass
class Task:
ticket: str
spec: str = ""
diff: str = ""
review: str = ""
def planner(task: Task) -> Task:
task.spec = llm(
system="You turn tickets into precise specs. List inputs, "
"outputs, edge cases, and one acceptance test per case. "
"Do NOT write implementation code.",
user=task.ticket)
return task
def implementer(task: Task) -> Task:
task.diff = llm(
system="You implement exactly the spec. If the spec is "
"ambiguous, STOP and return a QUESTION, not a guess.",
user=task.spec)
return task
def reviewer(task: Task) -> Task:
task.review = llm(
system="You review a diff you did NOT write. Find spec "
"violations, missing edge cases, and untested paths. "
"Approve only if every acceptance test is covered.",
user=f"SPEC:\n{task.spec}\n\nDIFF:\n{task.diff}")
return taskPicture this: you're a staff engineer who just watched a single coding agent confidently ship a function that passed its own tests and broke in production. The tests were wrong — the same agent wrote them. That failure is the reason multi agent software development exists. When one model writes code, its tests, and its own review, it isn't checking anything. It's agreeing with itself three times.
Splitting those jobs across agents changes the dynamic. Not because any agent is smarter, but because a reviewer who didn't write the code has no ego invested in it. This piece is about how that division actually works, where it pays off, and where the seams tear.
What is multi agent software development?
Multi agent software development is an approach where distinct agents own distinct phases of building software — typically planning, implementation, and review — and coordinate through shared state rather than a single conversation. Each agent has its own instructions and often its own toolset. A planner turns a ticket into a spec, an implementer writes code against that spec, and a reviewer critiques the diff before it merges.
The key difference from a single powerful coding assistant is accountability. In a single-agent flow, the model that wrote a bug is the same model asked whether the bug exists. It will usually say no. In a multi-agent flow, the reviewer starts from the diff, not from the reasoning that produced it, so it evaluates what the code does rather than what the author intended.
Here's a minimal three-role setup using explicit handoffs:
[object Object], dataclasses ,[object Object], dataclass
,[object Object],
,[object Object], ,[object Object],:
ticket: ,[object Object],
spec: ,[object Object], = ,[object Object],
diff: ,[object Object], = ,[object Object],
review: ,[object Object], = ,[object Object],
,[object Object], ,[object Object],(,[object Object],) -> Task:
task.spec = llm(
system=,[object Object],
,[object Object],
,[object Object],,
user=task.ticket)
,[object Object], task
,[object Object], ,[object Object],(,[object Object],) -> Task:
task.diff = llm(
system=,[object Object],
,[object Object],,
user=task.spec)
,[object Object], task
,[object Object], ,[object Object],(,[object Object],) -> Task:
task.review = llm(
system=,[object Object],
,[object Object],
,[object Object],,
user=,[object Object],)
,[object Object], taskWhat this does: it forces three separate reasoning passes — spec, implementation, critique — where the reviewer only sees the spec and the diff, never the implementer's private justification, so it judges outcomes instead of intentions.
Why it matters for real teams
The value shows up most on the boring failures, not the impressive ones. A single agent rarely fails to write plausible code. It fails to notice that plausible code doesn't handle an empty list, or that the spec asked for idempotency and the implementation isn't. A reviewer with fresh eyes catches those because it's not carrying the implementer's assumptions.
Consider three concrete settings. A platform team at a logistics company uses a planner-implementer-reviewer loop for internal CRUD endpoints, where the spec is mechanical and the win is consistency. A game studio uses it for shader tweaks, where the reviewer's job is mostly catching performance regressions the implementer optimized past. A fintech uses it for data-validation rules, where a separate reviewer that starts from the compliance spec catches "the code works but violates the rule" cases the implementer glossed.
⚡ Pro tip: make the spec the contract, not the ticket. Most multi agent software development failures trace back to a vague spec, because every downstream agent inherits the ambiguity. Spend your best prompt on the planner and require it to emit acceptance tests — those tests become the objective thing the reviewer checks against.
The economic case is subtle. You're paying two-to-four times the tokens for work a single agent could attempt. That only pays off where a caught defect is expensive: production incidents, compliance violations, data corruption. For a throwaway script, it's overkill. The teams getting value are deliberately selective about which work goes through the full loop.
How the roles divide the work
The cleanest division assigns each agent a single question. The planner answers "what does done look like?" The implementer answers "how do I make it so?" The reviewer answers "did we actually get there?" When an agent starts answering someone else's question, the system degrades toward a single agent wearing three hats.
Tooling should follow roles. The planner usually needs read access to the codebase and issue tracker but should not write files. The implementer needs write access and a way to run tests. The reviewer needs read access to the diff and test results but, again, no write access. Handing every agent every tool is the most common way multi-agent setups collapse back into chaos.
AGENT_TOOLS = {
,[object Object],: [,[object Object],, ,[object Object],], ,[object Object],
,[object Object],: [,[object Object],, ,[object Object],, ,[object Object],],
,[object Object],: [,[object Object],, ,[object Object],], ,[object Object],
}
,[object Object], ,[object Object],(,[object Object],):
,[object Object], action.tool ,[object Object], ,[object Object], AGENT_TOOLS[agent_name]:
,[object Object], PermissionError(
,[object Object],)
,[object Object], TOOLS[action.tool](action.args)What this does: it enforces least-privilege per role at the dispatch layer, so a reviewer literally cannot rewrite the code it's supposed to critique — a guardrail that keeps the separation of duties from eroding under pressure.
⚡ Pro tip: give the implementer permission to refuse. The single most valuable behavior in the whole loop is an implementer that returns "the spec is ambiguous about X" instead of guessing. A guess propagates silently; a question surfaces the gap while it's cheap to fix. Reward the refusal in your prompt explicitly.
There's a coordination pattern worth naming that the framework tutorials skip: the review-repair loop needs a hard iteration cap. Left uncapped, reviewer and implementer can ping-pong — the reviewer requests a change, the implementer makes it and introduces a new issue, the reviewer flags that, and so on. Cap it at two or three rounds, then escalate to a human. The cap isn't a limitation; it's what keeps the loop from burning tokens on diminishing returns.
Where the seams break
The failure modes cluster in three places. First, spec drift: the implementer quietly reinterprets an ambiguous spec and the reviewer, lacking the original intent, approves the reinterpretation. Fix this by keeping the ticket alongside the spec so the reviewer can catch drift from the source.
Second, reviewer capture: over many iterations the reviewer starts pattern-matching the implementer's style and rubber-stamping it. A fresh reviewer context per task, rather than a long-running conversation, keeps the critique sharp.
Third, integration blindness: each agent reasons about its slice, and no agent owns how the slices fit. A function can satisfy its spec and still break a caller. This is why the loop needs an integration test run outside any single agent's judgment — an objective check that doesn't care what the agents concluded.
⚠️ Common mistake: measuring the system by how much code it produces. Multi agent software development should produce less code that ships cleaner, not more code faster. If your agent team is merging more lines per day than your humans did, the reviewer is probably rubber-stamping. Watch the defect-escape rate, not the throughput.
How do you debug a multi-agent dev loop?
The hardest part of multi agent software development isn't building it — it's knowing why it failed when it does. With one agent, you read one transcript. With three agents and a repair loop, a failure could be a bad spec, a spec-faithful implementation of a wrong spec, or a reviewer that missed the issue. You need to see the seams to debug them.
The practical answer is to log every handoff as a discrete record: input, agent, output, and the state that moved forward. When a bug ships, you replay the handoffs and find the exact stage where correct input produced wrong output. Frameworks with built-in checkpointing make this easier because each node's state is already captured — you can rewind to the planner's output and re-run just the implementer with a fixed spec, without re-running the whole pipeline.
[object Object], ,[object Object],(,[object Object],):
,[object Object], ,[object Object],(,[object Object],):
before = task.__dict__.copy()
out = fn(task)
AUDIT.append({,[object Object],: agent_name, ,[object Object],: before,
,[object Object],: out.__dict__.copy()})
,[object Object], out
,[object Object], wrapped
planner = logged(,[object Object],, planner)
implementer = logged(,[object Object],, implementer)
reviewer = logged(,[object Object],, reviewer)What this does: it wraps each agent so every handoff records the state before and after, giving you a replayable trail that pinpoints which stage turned good input into a bad result instead of guessing across three agents.
The early-warning signal for a degrading loop is rising repair rounds. If the average number of reviewer-implementer round-trips creeps up over a week, either your specs are getting vaguer or your reviewer is nitpicking. Both are fixable, but only if you're tracking the number. A loop that silently averages three rounds instead of one is quietly tripling your cost.
⚡ Pro tip: pin each agent to a specific model deliberately, not uniformly. The reviewer benefits from your strongest reasoning model because catching subtle spec violations is hard. The implementer can often run a cheaper, faster model because writing code against a precise spec is comparatively mechanical. Matching model strength to role difficulty cuts cost without cutting quality where it counts.
⚡ Pro tip: seed the reviewer with your last five real production incidents as examples of what slips through. A reviewer that has seen your actual failure modes catches them far better than one running on generic "review this code" instructions. Your incident history is the best reviewer-tuning data you have, and almost nobody uses it.
Common mistakes
Teams routinely over-decompose. Five agents for a task that needs two adds coordination overhead with no quality gain. Start with planner-implementer-reviewer and only add roles when a specific failure demands one.
They also let agents share a single context to "save tokens," which quietly destroys the separation that made the approach work. The isolation is the feature. A reviewer that can see the implementer's chain-of-thought will be swayed by it.
Finally, teams skip the human escalation path. Every multi-agent loop needs a defined exit to a person — after the iteration cap, on low reviewer confidence, on any change touching security-sensitive code. The agents handle the volume; the human handles the judgment calls the volume would otherwise bury.
Conclusion
Multi agent software development works when the roles enforce accountability a single agent can't give itself: a spec that acts as a contract, an implementer that asks instead of guessing, and a reviewer with no stake in the code. Keep the isolation strict, cap the repair loop, and route the hard calls to a human.
The prompts that make each role work — the spec-writing planner, the refuse-when-ambiguous implementer, the adversarial reviewer — are worth version-controlling like any other critical config. Store them in PromptABCD so your whole team runs the same reviewer standard instead of re-inventing it per project and drifting apart.
Continue Reading
Save the prompts from this post
PromptABCD is a free prompt manager. Paste, organize, and reuse your best AI prompts — no more hunting through chat history.
