Building a Writer-Editor-Fact-Checker Agent Team
One agent wrote a beautiful article with three invented statistics. That failure is why a writer editor agent team beats a single writer. Here's the case study, the roles, and the handoffs that make it work.
Write a compelling, data-driven thought leadership article on
{topic}. Use specific statistics and examples to support your
points. Sound authoritative and back up claims with evidence.The article was genuinely good. Clean structure, confident voice, three specific statistics that made the argument land — and all three were invented. Not misremembered. Invented. The single writer agent had produced numbers that didn't exist because specific numbers read as authoritative, and reading authoritative was its actual objective. That one incident is the entire case for a writer editor agent team: a writer optimizing for persuasive prose will fabricate, and only a separate checker with different incentives catches it.
This case study follows a content team's move from one writer to a three-agent team, the specific failure that forced it, and how the handoffs work.
What went wrong with a single writer agent?
Maya leads content at a B2B analytics company. Her team shipped thought-leadership pieces, and she'd automated first drafts with a strong single-agent prompt. The drafts were fluent. The problem surfaced when a prospect emailed to ask for the source of a stat in a published post. There was no source. The number was plausible, on-brand, and fabricated.
An audit found the pattern was systematic. Across thirty drafts, the writer had introduced an average of two unsourced quantitative claims each. Most were directionally reasonable, which made them dangerous — a wildly wrong number gets caught, a plausible fake gets published. The writer wasn't malfunctioning. It was doing exactly what "write a compelling data-driven article" rewards: producing compelling data.
Here's the single-agent prompt that caused it:
Write a compelling, data-driven thought leadership article on
{topic}. Use specific statistics and examples to support your
points. Sound authoritative and back up claims with evidence.
What this does: it explicitly rewards specific statistics and authority without providing any source of real data, so the model manufactures the statistics the instruction demands.
⚠️ Common mistake: reading "use specific statistics" as a quality instruction. To a model with no data source, it's a fabrication instruction. If you want real numbers, you must give the agent a way to get real numbers or a separate agent whose job is to remove fake ones. Asking nicely doesn't work.
The wrong approach: telling the writer to cite
Maya's first fix was the obvious one: add "cite every statistic with a real source" to the writer's prompt. The result was worse in a specific way. The writer now invented statistics and invented sources for them — plausible-looking citations to real publications that had never published those numbers. The instruction to cite didn't create honesty; it created better-camouflaged fabrication.
This is the trap that makes single-agent fixes fail. Every instruction you add is executed by the same model with the same objective. Telling a persuasion-optimized writer to cite just teaches it to persuade harder. You cannot fix a role conflict with more instructions to the conflicted role. You need a second role with a different objective.
The correct writer editor agent team
The rebuild split the work into three agents with deliberately opposed incentives. The writer writes and is told to tag every factual claim. The editor improves clarity and flow but cannot add facts. The fact-checker verifies or kills every tagged claim and has no stake in the prose reading well.
[object Object], ,[object Object],(,[object Object],):
system = (,[object Object],
,[object Object],
,[object Object],
,[object Object],)
,[object Object], llm(system=system, user=brief)
,[object Object], ,[object Object],(,[object Object],):
system = (,[object Object],
,[object Object],
,[object Object],
,[object Object],
,[object Object],)
,[object Object], llm(system=system, user=draft, tools=[,[object Object],])
,[object Object], ,[object Object],(,[object Object],):
system = (,[object Object],
,[object Object],
,[object Object],
,[object Object],)
,[object Object], llm(system=system, user=checked)
,[object Object], ,[object Object],(,[object Object],):
,[object Object], editor(fact_checker(writer(brief)))What this does: it orders the roles so the fact-checker runs before the editor, freezing verified claims so the editor's polishing can't reintroduce or distort them — the ordering is the safeguard.
The ordering matters more than the prompts. Run the editor before the checker and the editor smooths over the claim tags, making them harder to verify. Run the checker first and it works on raw, tagged claims, then hands a clean frozen set to the editor. The sequence writer → checker → editor is load-bearing.
⚡ Pro tip: give the fact-checker explicit permission to make the article worse. The instruction "you do not care whether the article reads well" is doing critical work. A checker that also wants good prose will keep a plausible unsourced claim to preserve flow. A checker indifferent to prose deletes it without hesitation. Opposed incentives are the point.
Results and what changed
After the switch, unsourced quantitative claims in published pieces dropped from an average of two per article to zero across the next forty pieces. The articles got noticeably shorter — the fact-checker deleted real content — and Maya considered that a feature, because the deleted content was the liability.
The unexpected result was reader trust. Pieces now carried real citations, and the prospect emails changed from "where's this from?" to "thanks for the source." The team's credibility improved not because the writing got better but because it stopped lying.
Token cost roughly doubled. For thought-leadership content where a single fabricated stat could cost a deal, that was trivial. Maya was explicit that she wouldn't run the full team on low-stakes internal copy — the writer editor agent team is for content where being wrong is expensive.
⚡ Pro tip: track "claims deleted per article" as a leading indicator. A rising deletion rate means your writer is fabricating more, often because briefs are pushing it toward claims it can't support. The metric surfaces a prompt problem upstream before it becomes a published problem downstream.
How to apply this to your situation
Identify where your content has fabrication risk — usually anywhere you ask for specific data, statistics, or named examples. Those are the pieces that need the full team. Marketing copy with no factual claims doesn't; a data-driven report does.
Then build the three roles with opposed incentives and get the order right: write with tags, check and freeze, then edit. Give the checker real search tools and explicit permission to prioritize truth over polish. Keep a human on the final read for anything high-stakes — the team removes fabrication, but a person still owns the judgment call on what to publish.
⚠️ Common mistake: letting the editor run with search tools "to fill gaps." An editor that can search will add facts, which reopens the fabrication door the checker just closed. The editor works only with frozen, verified material. If a gap exists, that's a note back to the writer, not a license for the editor to research.
When does the writer editor agent team stop paying off?
Every case study makes its approach sound universal. This one isn't, and knowing the boundary keeps you from over-applying it. The three-agent team earns its cost only where fabrication is both likely and expensive. Strip either condition and the math flips.
Fabrication is likely wherever you ask for specifics the model can't actually retrieve — statistics, dated events, named studies, precise figures. It's unlikely in opinion, framing, and argument, where there's nothing to fabricate. So a persuasive essay with no factual claims doesn't need the fact-checker at all; the writer and editor suffice. Running the full team there just burns tokens verifying claims that don't exist.
Fabrication is expensive where a wrong fact damages trust or triggers consequences — external thought leadership, anything a prospect or regulator reads, content with your name on it. It's cheap where mistakes are caught internally before harm, like first-draft brainstorms a human will heavily rewrite. Match the pipeline depth to that cost.
There's a scaling wrinkle Maya hit at volume. When the team ran across forty pieces a week, the fact-checker became the bottleneck — verification is inherently serial per claim and can't be parallelized as cleanly as drafting. The fix was to have the writer batch its claims into a structured list the checker could verify in one pass, rather than re-reading prose to find each tag. Structuring the handoff, not adding agents, broke the bottleneck.
⚡ Pro tip: give the fact-checker a confidence field, not just a verdict. When it marks a claim "verified — low confidence," that's a flag for human review rather than an automatic keep. The binary verified/deleted split loses the middle cases where a source half-supports a claim, and those middle cases are exactly where a human's judgment adds the most value.
⚡ Pro tip: freeze citations as structured data, not inline text. When the checker outputs claims and sources as a separate structured block, the editor physically cannot alter them while polishing prose — the facts live outside the text it's allowed to touch. Inline citations invite accidental editing; structured citations enforce the freeze the pipeline depends on.
Next steps
Rebuild your highest-stakes content type as a writer editor agent team this week and measure claims deleted per article before and after. That number tells you how much fabrication your single writer was quietly shipping.
One habit makes the difference between a team that improves and one that plateaus: after every published piece, spend two minutes noting which claims the checker deleted and why. Over a month that log becomes a map of where your writer fabricates most and which topics your team genuinely lacks sources for. Both are actionable — the first tunes the writer prompt, the second tells you where to commission original research. Most teams throw this signal away; the ones that keep it get compounding quality gains for almost no effort.
Once the three role prompts are working — especially the checker's opposed-incentive framing — save them as a reusable set in PromptABCD. The exact wording that makes a fact-checker indifferent to prose is hard-won, and you don't want to reconstruct it from memory every time you spin up a new content project.
Continue Reading
Save the prompts from this post
PromptABCD is a free prompt manager. Paste, organize, and reuse your best AI prompts — no more hunting through chat history.
