Multi-Agent Systems for Research: A Case Study
A market analyst replaced a single research bot with a multi agent research system and cut a two-day report to four hours. Here's the exact architecture, the prompts, and the tradeoffs nobody warns you about.
You are a research assistant. Research {topic} thoroughly.
Search the web, read sources, and write a detailed report with
citations. Be objective and consider multiple viewpoints.
Flag anything uncertain.Only about one in five research automations that reach production use more than a single agent — yet the ones that do report the largest quality jumps of any agent category. That gap is the whole story here. Teams assume more agents means more complexity for marginal gain. Sometimes true. But for research, a multi agent research system solves a problem a single model structurally cannot: one context window can't hold a broad survey, deep verification, and synthesis without one of the three starving the others.
This is a walk-through of a real rebuild — names changed, numbers real — where splitting the work across agents turned an unreliable two-day slog into a four-hour pipeline an analyst actually trusts.
What problem does a multi agent research system solve?
Priya runs competitive intelligence at a mid-market fintech. Her job: weekly deep-dives on competitor pricing, feature launches, and hiring signals. Each report pulled from 30–40 sources and had to be defensible enough to put in front of the product VP.
Her first automation was a single large-context agent. Feed it a topic, let it search, get a report. It worked in demos and fell apart in practice. The failure wasn't hallucination in the obvious sense. It was subtler: the agent would find a strong claim early, latch onto it, and let it color the rest of the report. Contradicting sources got quietly dropped. By the time it wrote the synthesis, its context was so full of its own intermediate notes that fresh evidence got crowded out.
The honest diagnosis: a single agent conflates three jobs that want different temperaments. Searching broadly rewards curiosity and coverage. Verifying rewards suspicion. Synthesizing rewards restraint. You can't tune one prompt to be simultaneously curious, suspicious, and restrained — you get a mush that's mediocre at all three.
Why the single-agent approach kept failing
The first instinct was to fix the prompt. More instructions. "Consider contradicting evidence." "Cite every claim." "Don't over-index on the first source." Each addition helped one failure mode and worsened another, because the instructions competed for the model's attention inside one turn.
Here's the weak version that shipped first:
You are a research assistant. Research {topic} thoroughly.
Search the web, read sources, and write a detailed report with
citations. Be objective and consider multiple viewpoints.
Flag anything uncertain.
What this does: it asks one model to hold coverage, skepticism, and synthesis in a single pass — which is exactly why it drifts toward whatever it read first.
The measurable symptom: Priya spot-checked 20 reports and found that in 14 of them, a claim in the executive summary wasn't actually supported by the cited source when she opened it. The agent had summarized its own earlier paraphrase, not the source. That's a fatal trust problem for intelligence work.
⚠️ Common mistake: treating research quality as a prompting problem when it's a separation-of-duties problem. No amount of instruction stacking fixes a single agent grading its own homework.
What the correct multi agent research system looks like
The rebuild split the pipeline into four roles, each a separate agent with its own instructions, its own tools, and — critically — its own fresh context. A lead agent decomposes the question and spawns parallel gatherers, a verifier checks claims against raw sources, and a synthesizer writes only from verified material.
[object Object], langgraph.graph ,[object Object], StateGraph, END
,[object Object], typing ,[object Object], TypedDict, Annotated
,[object Object], operator
,[object Object], ,[object Object],(,[object Object],):
topic: ,[object Object],
subquestions: ,[object Object],[,[object Object],]
findings: Annotated[,[object Object],[,[object Object],], operator.add] ,[object Object],
verified: ,[object Object],[,[object Object],]
report: ,[object Object],
,[object Object], ,[object Object],(,[object Object],):
,[object Object],
subs = llm_plan(state[,[object Object],])
,[object Object], {,[object Object],: subs}
,[object Object], ,[object Object],(,[object Object],):
,[object Object],
,[object Object], {,[object Object],: [search_and_extract(q) ,[object Object], q ,[object Object], state[,[object Object],]]}
,[object Object], ,[object Object],(,[object Object],):
kept = []
,[object Object], f ,[object Object], state[,[object Object],]:
,[object Object], claim_matches_source(f[,[object Object],], f[,[object Object],]):
kept.append(f)
,[object Object], {,[object Object],: kept}
,[object Object], ,[object Object],(,[object Object],):
,[object Object], {,[object Object],: write_report(state[,[object Object],], state[,[object Object],])}
g = StateGraph(ResearchState)
,[object Object], name, fn ,[object Object], [(,[object Object],,planner),(,[object Object],,gather),(,[object Object],,verify),(,[object Object],,synthesize)]:
g.add_node(name, fn)
g.set_entry_point(,[object Object],)
g.add_edge(,[object Object],,,[object Object],); g.add_edge(,[object Object],,,[object Object],)
g.add_edge(,[object Object],,,[object Object],); g.add_edge(,[object Object],,END)
app = g.,[object Object],()What this does: it wires four specialized agents into a directed graph where each stage hands clean state to the next, so the synthesizer never sees unverified raw text and can't launder a bad claim into the summary.
The verifier is the piece that changed everything. It doesn't trust the gatherer's paraphrase — it re-opens the source URL and checks that the specific claim actually appears there. Claims that fail get dropped before synthesis, not after.
⚡ Pro tip: give the verifier a deliberately adversarial system prompt. Tell it its job is to find reasons a claim is unsupported, not to confirm it. A verifier prompted to "check" agrees too easily; one prompted to "disprove" catches the paraphrase drift.
Here's the verifier instruction that worked:
You are a fact-checker whose reputation depends on catching
unsupported claims. For each claim, open the cited source and
decide: does the source STATE this, or did the writer infer it?
Mark INFERRED or UNSUPPORTED claims for removal. When the source
only partially supports the claim, rewrite the claim to match
exactly what the source says. Assume every claim is wrong until
the source proves otherwise.
What this does: it flips the verifier from a rubber stamp into a skeptic, which is the only posture that reliably catches the "summarized my own paraphrase" failure.
Results and what actually changed
After two weeks of parallel running, the numbers were clear. Unsupported claims in the executive summary dropped from 14-in-20 to 1-in-20. Wall-clock time per report fell from roughly two working days to about four hours, most of that now spent by Priya reviewing rather than fixing. Token cost roughly tripled — four agents plus re-verification isn't free — but at report volume the analyst's time was worth far more than the extra spend.
The insight that isn't in the framework docs: the verifier catches more when gatherers are forbidden from summarizing. The gatherers were changed to return verbatim quotes plus URLs, never paraphrases. Paraphrase is where drift enters. If the gatherer never paraphrases, the verifier has a clean quote to check and the synthesizer has exact material to work from. Coverage went slightly down, precision went way up, and for intelligence work precision wins.
⚡ Pro tip: measure your research pipeline on "claims that survive verification," not "claims produced." A single agent looks more productive because nothing culls its output. The multi agent research system produces fewer claims and far more trustworthy ones — track the survival rate and you'll see the real quality difference.
How to apply this to your own research
Start by writing down the three temperaments your current bot is failing to hold at once. For most research tasks it's coverage, skepticism, and restraint. Assign each to a separate agent before you touch any framework.
Then enforce three rules. Gatherers return quotes and URLs, never summaries. The verifier re-opens sources and assumes claims are wrong. The synthesizer writes only from the verified pool and cannot search. That last rule prevents the synthesizer from "topping up" with unverified last-minute finds — the exact move that reintroduces drift.
⚡ Pro tip: parallelize the gatherers but keep the verifier serial. Fanning out search is where you save wall-clock time; verification is cheap per claim and benefits from consistent judgment, so running it as one agent over all findings keeps the skeptic's bar uniform.
Consider a legal researcher building case summaries, a pharma analyst tracking trial results, or a journalist verifying a tip. All three share Priya's shape: broad gathering, hard verification, careful synthesis. The roles transfer even when the domain doesn't.
⚠️ Common mistake: adding a fifth "editor" agent too early. The editor mostly rephrases, which reintroduces the paraphrase drift the verifier just removed. Add polish agents only after your survival rate is stable, and forbid them from altering any verified claim's substance.
What does the coordination tax actually cost?
The part the case study glosses is the bill. A multi agent research system isn't free, and being honest about the tax is how you decide where it belongs. Priya's four-agent pipeline cost roughly three times the tokens of the single agent, plus the verifier re-opening sources adds real latency — the four-hour figure includes verification time the single agent skipped entirely by never checking.
The tax breaks down usefully. The gatherers, run in parallel, add cost but not much wall-clock time. The verifier is the expensive stage in latency terms because it re-fetches sources, but it's the cheapest to reason about because each claim is a small independent check. The synthesizer is a single pass over a small verified pool, so it's cheap. Which means the obvious optimization — a smaller model for the synthesizer — barely helps, while a smaller model for the gatherers helps a lot, since they generate the bulk of the tokens.
Here's the migration pattern that worked without a risky big-bang switch. Priya ran the single agent and the multi-agent pipeline in parallel for two weeks, comparing claim survival on the same topics. Only once the multi-agent survival rate was demonstrably higher did she cut over. Running both is expensive for a fortnight and far cheaper than shipping a broken pipeline into a VP's decision-making.
⚡ Pro tip: watch verifier deletion rate as an early-warning signal on gatherer quality. If the verifier starts deleting a rising share of claims, your gatherers are drifting toward paraphrase or weak sources — fix the gatherers before the synthesizer starves. The deletion rate is a leading indicator that surfaces upstream rot before it reaches the report.
The strategic question isn't whether the multi agent research system is more expensive — it clearly is. It's whether trustworthy claims are worth three times the token cost for this particular research. For competitive intelligence feeding executive decisions, obviously yes. For a low-stakes internal summary nobody acts on, obviously no. The teams that get value are ruthless about that distinction and don't run the full pipeline on research that doesn't warrant it.
Next steps
Rebuild your worst research automation as three agents this week — gatherer, verifier, synthesizer — and measure claim survival before and after. That single metric tells you whether the coordination tax is buying you anything.
Once your role prompts earn their keep, stop rewriting them from memory each time. Save the gatherer, verifier, and synthesizer instructions as versioned, reusable prompts in PromptABCD so the whole team runs the same skeptic verifier instead of quietly drifting back to a single agent that grades its own work.
Continue Reading
Save the prompts from this post
PromptABCD is a free prompt manager. Paste, organize, and reuse your best AI prompts — no more hunting through chat history.
