Roles and Responsibilities in an Agent Crew
Most guides on agent roles responsibilities focus on what each agent should do. The more important question is what each agent should refuse to do — and most multi-agent systems never answer it.
from anthropic import Anthropic
client = Anthropic()
# Role with explicit refusals — the missing ingredient in most designs
RESEARCHER_SYSTEM = """You are a research specialist in a content production team.
YOUR JOB:
- Find specific facts, data points, and sources about the given topic
- Assess source reliability and note confidence levels
- Identify gaps where evidence is weak or missing
- Return structured findings with citations
NOT YOUR JOB (refuse these if asked):
- Do not write prose, articles, or formatted content — that's the writer's job
- Do not make editorial judgments about what the audience will find interesting
- Do not evaluate the quality of the brief you receive — just execute it
- Do not suggest expanding the scope of the research task
When you finish, output your findings in this format and nothing else:
FINDINGS: [your structured findings]
CONFIDENCE: [high/medium/low]
GAPS: [identified gaps in the evidence]"""
WRITER_SYSTEM = """You are a technical content writer in a content production team.
YOUR JOB:
- Transform research findings into clear, well-structured prose
- Match tone and complexity to the specified audience
- Ensure every claim is traceable to the provided research
- Produce complete drafts ready for review
NOT YOUR JOB (refuse these if asked):
- Do not conduct additional research — use only what the researcher provided
- Do not assess the accuracy of the research you receive — that's the reviewer's job
- Do not make decisions about what to include or exclude based on your own knowledge
- Do not change the scope or focus of the assignment
Produce a complete draft. Do not summarize or shorten the task."""
REVIEWER_SYSTEM = """You are a content quality reviewer in a content production team.
YOUR JOB:
- Check that every factual claim in the draft is supported by the provided research
- Identify logical gaps, unclear explanations, and structural issues
- Verify the draft matches the original brief
- Produce a specific, actionable review
NOT YOUR JOB (refuse these if asked):
- Do not rewrite content — flag issues for the writer to fix
- Do not conduct additional research to verify claims not in the provided sources
- Do not give overall quality scores or grades — give specific, actionable feedback only
- Do not approve drafts that contain unsupported claims, regardless of how minor they seem"""
def run_crew(topic: str, audience: str) -> dict:
research = client.messages.create(
model="claude-sonnet-4-6",
max_tokens=1024,
system=RESEARCHER_SYSTEM,
messages=[{"role": "user", "content": f"Research topic: {topic}"}]
).content[0].text
draft = client.messages.create(
model="claude-sonnet-4-6",
max_tokens=2048,
system=WRITER_SYSTEM,
messages=[{"role": "user", "content": f"Audience: {audience}\nResearch:\n{research}\nWrite the article."}]
).content[0].text
review = client.messages.create(
model="claude-sonnet-4-6",
max_tokens=1024,
system=REVIEWER_SYSTEM,
messages=[{"role": "user", "content": f"Brief: {topic} for {audience}\nResearch:\n{research}\nDraft:\n{draft}\nReview the draft."}]
).content[0].text
return {"research": research, "draft": draft, "review": review}Most guides on multi-agent design spend all their time on what each agent should do. The research agent researches. The writing agent writes. The review agent reviews. This seems complete until your research agent starts offering editorial opinions, your writing agent starts conducting additional research to fill gaps it notices, and your review agent starts rewriting content rather than just flagging issues.
What actually defines a role isn't just its responsibilities. It's its constraints — specifically, what the agent explicitly refuses to do.
This is the most underappreciated principle in agent roles responsibilities design, and it's counterintuitive: the most useful thing you can put in an agent's system prompt is often a list of things it won't do.
What Are Agent Roles in a Multi-Agent System?
An agent role is a configuration that defines an agent's identity, scope of action, and behavioral constraints within a larger system. Roles serve two functions: they shape how the model attends to inputs (a legal reviewer attends to different cues than a content writer), and they prevent role drift — agents gradually expanding beyond their intended scope as they try to be helpful.
In a well-designed agent crew, each agent has a specific capability profile, a defined set of inputs it can receive, a defined set of outputs it should produce, and explicit boundaries on what it doesn't do. Those last boundaries are what most implementations skip.
[object Object], anthropic ,[object Object], Anthropic
client = Anthropic()
,[object Object],
RESEARCHER_SYSTEM = ,[object Object],
WRITER_SYSTEM = ,[object Object],
REVIEWER_SYSTEM = ,[object Object],
,[object Object], ,[object Object],(,[object Object],) -> ,[object Object],:
research = client.messages.create(
model=,[object Object],,
max_tokens=,[object Object],,
system=RESEARCHER_SYSTEM,
messages=[{,[object Object],: ,[object Object],, ,[object Object],: ,[object Object],}]
).content[,[object Object],].text
draft = client.messages.create(
model=,[object Object],,
max_tokens=,[object Object],,
system=WRITER_SYSTEM,
messages=[{,[object Object],: ,[object Object],, ,[object Object],: ,[object Object],}]
).content[,[object Object],].text
review = client.messages.create(
model=,[object Object],,
max_tokens=,[object Object],,
system=REVIEWER_SYSTEM,
messages=[{,[object Object],: ,[object Object],, ,[object Object],: ,[object Object],}]
).content[,[object Object],].text
,[object Object], {,[object Object],: research, ,[object Object],: draft, ,[object Object],: review}What this does: Each agent's system prompt explicitly lists what it refuses to do alongside what it does. The researcher won't write prose. The writer won't conduct research. The reviewer won't rewrite. These constraints prevent role drift and keep each agent's behavior within its intended scope, even when the input or task context might seem to invite expansion.
Why Role Definitions Matter Beyond Job Descriptions
When you give an agent a job description without constraints, it optimizes for being helpful within its own context. A research agent that encounters a gap in the available evidence will try to fill the gap — by drawing on its training data, by making educated guesses, or by generating plausible-sounding content. This is helpful in the sense that it produces output rather than an error. It's harmful in the sense that it produces output that looks like research but isn't.
The same dynamic affects every role. Writers who notice factual gaps add claims not in the research. Reviewers who spot a fixable error fix it rather than flagging it. Agents optimize for task completion, and task completion often means overstepping the role boundary.
Explicit refusals address this by creating negative constraints that the model treats as strongly as positive instructions. "Do not add information from your training data" is a specific instruction that shapes behavior more reliably than hoping the agent naturally stays within scope.
⚡ Pro tip: Test your role boundaries explicitly by giving each agent inputs that would tempt it to overstep. Give the writer a research document with an obvious gap and see if it adds unsupported content. Give the reviewer a draft with a fixable error and see if it flags or rewrites. This testing reveals role drift before it reaches production.
The Core Components of a Well-Defined Agent Role
Every effective agent role definition has five elements:
Role title and context: Who the agent is within the team structure. "You are a research specialist in a content production team" sets both the identity and the organizational context. The organizational context matters — "in a content production team" signals that other agents handle adjacent responsibilities.
Primary responsibilities: What the agent does, stated specifically. "Find specific facts, data points, and sources" is better than "do research." Specificity shapes which features of the input the model attends to.
Output format: What the agent produces and how it's structured. Explicit output format requirements prevent agents from inventing formats that downstream agents can't parse. "Return your findings in this format and nothing else" is often the most important sentence in the output section.
Explicit refusals: What the agent does not do. This is the element most system prompts omit and the element that most reliably prevents role drift. Each refusal should name the specific tempting behavior and the agent it belongs to. "Do not write prose — that's the writer's job" is more effective than "stay in your lane."
Escalation instructions: What the agent does when it encounters something outside its scope. "If you're uncertain whether a task falls within your responsibilities, flag it rather than attempting it" prevents silent scope expansion.
⚡ Pro tip: Write your refusals in the same specific tone as your responsibilities. "Do not make editorial judgments about what the audience will find interesting" is more effective than "stay focused on research." The more specific the refusal, the more reliably the model follows it.
How Role Definitions Work Across Different Agent Types
Research agents: The most important constraint is data sourcing. Research agents should refuse to use training data as a primary source, refuse to fill gaps with plausible-sounding guesses, and refuse to evaluate or editorialize on what they find. Their job is to find, not to judge.
Execution agents: Code generators, form fillers, data processors. Refusals matter most around scope — don't expand the task, don't make design decisions that weren't specified, don't optimize for anything other than the stated requirements. Execution agents should refuse to make judgment calls that require human authorization.
Review agents: The critical refusal for reviewers is rewriting. A review agent that fixes errors rather than flagging them eliminates the writer's ability to understand and correct their own work. Another important refusal: review agents should refuse to approve content that contains any unsupported claim, even if flagging it will trigger another revision cycle.
Coordination agents: Supervisors and orchestrators should refuse to do domain work. If your supervisor agent is writing content, analyzing data, or doing the task that workers are supposed to do, its coordination quality degrades. Supervisors should refuse to produce domain output — their job is to direct and evaluate, not to execute.
Common Mistakes in Role Design
⚠️ Common mistake: Writing role descriptions that are too abstract to actually constrain behavior. "Be a professional researcher" is a description, not a constraint. "Return only findings from the provided sources — do not add claims from your training data" is a constraint. Abstract role descriptions produce general helpfulness; specific constraints produce role-appropriate behavior.
Three additional failure patterns:
Roles that overlap: When two agents' responsibilities share territory, both agents try to do the overlapping work, producing duplicated or conflicting output. Map your agent responsibilities explicitly and identify any overlaps before implementation.
Roles without output format: An agent that knows what to do but not what to produce outputs content in whatever format seems natural for the task. This format varies with input and is often incompatible with downstream agents' expectations. Always specify output format explicitly.
No escalation path: Agents without escalation instructions handle out-of-scope requests by either attempting them poorly or refusing without context. Both behaviors break the pipeline. Specify exactly what an agent should do when it encounters something outside its defined role.
Building Your Agent Crew
Agent roles responsibilities design is upstream of everything else in multi-agent system development. Poorly defined roles produce agents that drift, overlap, and produce inconsistent output regardless of how well the orchestration logic is built. Well-defined roles — with explicit responsibilities, output formats, refusals, and escalation instructions — produce agents whose behavior is predictable and whose problems are diagnosable.
Start by listing every type of work your system needs to perform. Group related work into roles. For each role, draft the responsibilities and the refusals simultaneously. Test for role drift explicitly. Iterate.
Role Validation Through Adversarial Testing
The most reliable test for a well-defined agent role is adversarial: deliberately give the agent inputs designed to trigger the behaviors the role definition is supposed to prevent.
For a research agent defined to "report only findings from provided sources, never from training knowledge," adversarial test inputs include questions where the provided sources are incomplete or silent on the topic. A well-defined researcher refuses to fill gaps from training knowledge and explicitly flags what it couldn't find. A poorly-defined researcher fills the gaps confidently, producing output that looks research-backed but isn't.
For a writer agent defined to "produce output only from the provided research brief," adversarial inputs include briefs with obvious gaps or interesting-looking tangents. A well-defined writer sticks to the brief and notes the gaps. A poorly-defined writer embellishes, producing content that sounds good but exceeds what the research supports.
⚡ Pro tip: Create an "adversarial input set" for each agent role in your system — five to ten inputs specifically designed to trigger role violations. Run this set against the agent every time you update the system prompt. If an adversarial input that previously caused a role violation now passes, that's a genuine prompt improvement. If a previously passing input now fails, the prompt change introduced a regression.
Role definitions that survive adversarial testing are the most reusable. They've been tested against their own failure modes, not just their happy path.
When your role definitions are working well, save them to PromptABCD. Well-defined agent roles are your most reusable assets in multi-agent development — they transfer between projects with minimal modification.
Continue Reading
Save the prompts from this post
PromptABCD is a free prompt manager. Paste, organize, and reuse your best AI prompts — no more hunting through chat history.
