PromptABCD
FeaturesLearnGuideBlogContext Blocks
Sign inGet started free
Sign inSign up
PromptABCD

A calm home for your best AI prompts. Save them once, find them in seconds, reuse them forever.

Product

  • Features
  • Chrome Extension
  • Free Courses
  • How it works
  • Use cases
  • Blog
  • Context Blocks
  • Export Anywhere
  • FAQ

Resources

  • User guide
  • Learn prompting
  • Sign in
  • Get started free

© 2026 PromptABCD. All rights reserved.

AboutPrivacy PolicyTerms and Conditions
Home/Blog/Multi-Agent Systems/Building a Sales-Research-Outreach Agent Team
Multi-Agent Systems

Building a Sales-Research-Outreach Agent Team

One agent sent 200 personalized emails that all opened with the same fake-specific line. That failure shows why a multi agent sales team splits research from writing. Here's how to build one that sounds human.

October 1, 2026·9 min read
ShareShare
⚡Featured Prompt— copy and use right now
def researcher(prospect):
    system = ("Find 3 SPECIFIC, verifiable facts about this prospect "
              "or their company: a recent announcement, a role change, "
              "a public post, a product launch. Output each fact with "
              "its source URL. If you cannot find 3 real facts, output "
              "fewer — do NOT invent generic ones like 'doing "
              "interesting work'.")
    return llm(system=system, user=json.dumps(prospect),
               tools=["web_search"])

def qualifier(prospect, facts):
    system = ("Given these facts, decide if this prospect fits our "
              "ICP. Output {fit: bool, reason, best_hook}. If not a "
              "fit, say so — a rejected lead is better than a wasted "
              "email.")
    return json.loads(llm(system=system,
                          user=f"{prospect}\n{facts}"))

One sales team wired up a single agent to research prospects and write outreach, then sent two hundred emails in a week. Every one opened with a line that sounded personalized and was actually the same hollow template: "I noticed [company] is doing interesting work in [industry]." Prospects saw through it instantly. Reply rates were worse than a plain mass email, because fake personalization reads as more insulting than none. That failure is the case for a multi agent sales team — separating research from writing so the personalization is real, not performed.

This is a build guide: what a multi agent sales team is, why the split matters, how the roles divide, and where it goes wrong.

What is a multi agent sales team?

A multi agent sales team assigns distinct agents to the stages of outbound — researching a prospect, deciding whether they fit, writing the outreach, and sending — coordinating through shared prospect state rather than one agent doing everything in a single pass. Each stage produces a distinct artifact the next stage consumes.

The reason to split is specific to why the single agent failed. When one agent researches and writes in the same breath, it doesn't actually do research — it generates plausible-sounding personalization on the fly, because writing a smooth email is easier than digging up a real, specific hook. Separate the researcher and force it to output concrete facts, and the writer has something genuine to work with instead of a template to fill.

python
[object Object], ,[object Object],(,[object Object],):
    system = (,[object Object],
              ,[object Object],
              ,[object Object],
              ,[object Object],
              ,[object Object],
              ,[object Object],)
    ,[object Object], llm(system=system, user=json.dumps(prospect),
               tools=[,[object Object],])

,[object Object], ,[object Object],(,[object Object],):
    system = (,[object Object],
              ,[object Object],
              ,[object Object],
              ,[object Object],)
    ,[object Object], json.loads(llm(system=system,
                          user=,[object Object],))

What this does: it forces the researcher to produce specific, sourced facts or admit it found none, and adds a qualifier that can reject a prospect — so the writer only ever works from real hooks on real fits, not manufactured personalization.

Why it matters

The economics of outbound are brutal and specific: a bad email doesn't just fail, it burns the prospect and can hurt your domain reputation. Volume with poor personalization is negative-value work. A multi agent sales team is worth building precisely because it trades raw volume for genuine relevance, and in outbound, relevance is the only thing that moves reply rates.

Consider three settings. A B2B SaaS SDR team uses the researcher to find real trigger events — funding, hiring, product launches — so outreach lands when timing is right. A recruiting agency uses it to find specifics about a candidate's recent work so messages don't read as spray-and-pray. An agency doing partnership outreach uses the qualifier hard, because sending to poor-fit prospects wastes the one warm introduction they get.

⚡ Pro tip: make the researcher's "found nothing" output a first-class result, not a failure. The most valuable thing a research agent can tell you is that a prospect has no real hook — that lead should be dropped or routed to a nurture sequence, not force-fed a generic email. Teams that punish the researcher for finding nothing get invented facts; teams that reward honest "no hook here" get clean lists.

The qualifier is the piece most teams skip and shouldn't. An agent that can reject prospects keeps your list clean and your sending reputation intact. Sending to everyone the researcher touched is how the single-agent version generated its two hundred insulting emails. A qualifier that says "not a fit, skip" is doing high-value work by preventing bad sends.

How the roles divide the work

Give each agent one job and one output. The researcher produces sourced facts. The qualifier produces a fit decision and the best hook. The writer produces a draft from that hook. The sender handles delivery within guardrails. When the writer starts "researching" to fill a thin hook, you're back to the single-agent failure.

The writer's constraint is the key design choice: it may only use hooks the researcher surfaced. It cannot invent context. If the researcher found nothing specific, the writer either sends a deliberately short, honest message or the prospect gets skipped — but the writer never manufactures a fake specific.

python
[object Object], ,[object Object],(,[object Object],):
    system = (,[object Object],
              ,[object Object],
              ,[object Object],
              ,[object Object],
              ,[object Object],
              ,[object Object],)
    ,[object Object], llm(system=system, user=,[object Object],)

,[object Object], ,[object Object],(,[object Object],):
    ,[object Object], guardrails[,[object Object],] >= guardrails[,[object Object],]:
        ,[object Object], {,[object Object],: ,[object Object],, ,[object Object],: ,[object Object],}
    ,[object Object], ,[object Object], passes_spam_check(email):
        ,[object Object], {,[object Object],: ,[object Object],, ,[object Object],: ,[object Object],}
    ,[object Object], queue_for_human_approval(email)   ,[object Object],

What this does: it constrains the writer to the real hook and forbids fake personalization, then routes every email through caps, a spam check, and human approval before it sends — so nothing goes out unreviewed or over-volume.

⚠️ Common mistake: letting the sender agent send autonomously. Outbound email is a domain where a mistake at scale damages your domain reputation for months. The sender should queue for human approval or, at minimum, send within tight caps with a kill switch. Full autonomy on sending is the one place in this pipeline where the downside dwarfs the time saved.

⚡ Pro tip: separate the voice template from the hook. Store your outreach voice — tone, length, structure — as a stable reusable prompt, and let only the hook vary per prospect. This keeps your outreach consistent and on-brand while the personalization stays genuinely specific. Mixing voice and hook into one prompt makes both harder to tune and tends to produce that averaged, templated feel.

How do you measure a multi agent sales team?

The single-agent version of outbound has one metric — emails sent — and that metric actively lies, because it rewards exactly the spray-and-pray behavior that tanks results. A multi agent sales team gives you honest measurement points, and using them is what turns the system from a volume machine into a relevance machine.

Track four numbers, in order of importance. First, reply rate — the only metric that reflects whether the outreach actually landed. Second, the researcher's hook rate: what fraction of prospects yield three real, sourced facts. A falling hook rate means your list quality is degrading, and it explains reply-rate drops before they show up. Third, qualifier rejection rate: what share of researched prospects get filtered out. A rejection rate near zero means the qualifier is a rubber stamp and bad-fit prospects are getting emailed. Fourth, per-prospect cost, which rises when the researcher works hard on prospects the qualifier then rejects — a sign you should filter earlier.

The relationship between these tells a story a single number can't. If reply rate falls while hook rate holds, your writer or timing is the problem. If reply rate falls and hook rate fell first, your list is the problem. The multi-stage design gives you the diagnostic granularity to know which agent to fix, instead of guessing at one opaque pipeline.

⚡ Pro tip: attribute replies back to the specific hook the researcher found. Over a few hundred sends, patterns emerge — funding-round hooks reply at one rate, product-launch hooks at another, role-change hooks at a third. Feed that back into the researcher's priorities so it hunts hardest for the hook types that actually convert. Most teams never connect replies to hook types and leave this compounding advantage on the table.

There's an ordering optimization the metrics reveal: run the qualifier before the expensive deep research, not after. A cheap first-pass fit check on basic firmographics filters out obvious non-fits before you spend research tokens on them. Deep research is your costliest stage — spending it only on pre-qualified prospects can cut pipeline cost substantially while improving list quality, because the researcher's effort concentrates where it can pay off.

Common mistakes

The biggest is optimizing for volume. A multi agent sales team should send fewer, better emails — if it's sending more than your reps did, the qualifier is too loose and you're back to spray-and-pray with extra steps. Watch reply rate, not send count, and treat any week where sends rise but replies stay flat as a signal the qualifier has loosened and needs tightening before your list quality erodes further.

The second is trusting the researcher's facts without verification. Give the researcher real search tools and require source URLs, then spot-check. An unverified "fact" in a cold email that turns out wrong is worse than no personalization at all — it signals you didn't actually do the homework you implied.

The third is skipping human approval on send. The volume the pipeline enables is exactly why a human gate matters — automation multiplies both good and bad sends, and the bad ones compound in ways a single careless afternoon can't undo for months.

Conclusion

A multi agent sales team works because it forces real research before writing, lets a qualifier reject poor fits, and keeps a human on the send. The single agent fails because writing a smooth fake beats doing real research, every time, unless you separate the jobs.

One operational habit keeps the whole system honest over time: sample a handful of sent emails every week and read them as a prospect would. Metrics can look healthy while the writer slowly drifts back toward hollow personalization, because reply rates lag and small quality erosions hide in aggregates. A two-minute weekly read catches the drift a dashboard misses, and it keeps the researcher's real hooks from being wasted by a writer that started padding again.

The researcher, qualifier, writer, and sender prompts — especially the writer's "no fake personalization" constraint and your voice template — are worth versioning as a set. Store them in PromptABCD so your whole team runs the same qualified, genuinely-personalized outreach instead of each rep quietly reinventing a slightly worse template.

multi-agent-systemssales-agentsoutreach-automationlead-researchsales-teampersonalization

Continue Reading

A Reusable Prompt Kit for Agent Teams
Multi-Agent Systems

A Reusable Prompt Kit for Agent Teams

A team rebuilt their agent prompts from memory every project, and every project drifted a little worse. That failure is why a multi agent prompt kit matters. Here's the reusable set of role prompts every team should keep.

October 2, 2026·9 min read
Multi-Agent Systems on a Budget: An Interactive Guide
Multi-Agent Systems

Multi-Agent Systems on a Budget: An Interactive Guide

Most multi-agent tutorials assume you'll burn tokens freely. That's wrong for anyone shipping on real constraints. A cheap multi agent system can match an expensive one with the right moves. Here's how to build one.

October 2, 2026·9 min read
The Manager Agent Anti-Pattern: A Teardown
Multi-Agent Systems

The Manager Agent Anti-Pattern: A Teardown

Why does your orchestrator agent keep becoming a bottleneck that mangles every handoff? You've hit the manager agent anti pattern — one agent trying to coordinate everything. Here's why it fails and what replaces it.

October 2, 2026·9 min read

Save the prompts from this post

PromptABCD is a free prompt manager. Paste, organize, and reuse your best AI prompts — no more hunting through chat history.

Start free →
← PreviousMulti-Agent Simulations of Human Behavior: A Case StudyNext →Multi-Agent Systems for Due Diligence: A Teardown
Share this post:
ShareShare