Agent Specialization Through Prompting: A Teardown
Most guides say specialize agents by giving each a job title in its prompt. That's shallow and it shows. Real agent specialization prompts change what an agent notices and refuses. Here's the weak version and the rebuild.
You are an expert security engineer with 15 years of experience. Review this code and find security issues. You are highly skilled at identifying vulnerabilities. Be thorough and professional.
Most guides on building specialized agents say the same shallow thing: give each agent a job title in its system prompt. "You are a security expert." "You are a financial analyst." Done, supposedly. But a title is not a specialization — it's a costume. Agent specialization prompts that actually work change what an agent notices, what it refuses to do, and what it prioritizes when goals conflict. This teardown takes apart the title-only approach and rebuilds it into specialization that survives contact with a real task.
Before: the weak specialization prompt
Here's the specialization approach ninety percent of tutorials teach:
You are an expert security engineer with 15 years of experience.
Review this code and find security issues. You are highly skilled
at identifying vulnerabilities. Be thorough and professional.
What this does: it assigns a role and some flattering backstory, then hopes the model behaves like a specialist — relying on a title and adjectives to produce expertise it hasn't actually been given any structure to apply.
It reads like specialization. Run it against a real codebase next to a genuinely specialized prompt and the gap is obvious.
Why it fails
A job title changes the model's tone, not its behavior. "You are a security expert" makes the output sound more security-flavored — more jargon, more confidence — without changing what the agent actually examines or how it decides. The costume is convincing; the expertise underneath is the same generalist reasoning it always had. You get security-flavored generalism, not security specialization.
The specific failure is that title-only prompts don't change attention. A real security specialist doesn't scan code the way a generalist does — they look for specific things in a specific order, they know which patterns are dangerous, and they refuse to sign off on certain constructs regardless of context. A title gives the agent none of that structure. It still looks at everything with even, generalist attention, just describing what it finds in a more expert voice.
⚠️ Common mistake: believing a longer, more impressive backstory makes a better specialist. "Fifteen years of experience at a top firm" adds tokens and zero capability. The model doesn't have fifteen years of experience; telling it that it does just biases it toward overconfidence. Specialization comes from structure — what to notice, what to prioritize, what to refuse — not from an increasingly elaborate fictional resume.
The deepest failure is that title-only agents don't refuse. Real specialists have hard lines — a security specialist won't approve code that logs credentials, no matter how the rest looks. A title-only agent has no such lines, so under pressure it rationalizes, because nothing in its prompt tells it what it must never accept. Specialization without refusal boundaries is just a preference, and preferences bend.
After: the improved agent specialization prompts
Real specialization is built from four components: what the agent attends to, the order it examines things, what it prioritizes when goals conflict, and what it refuses absolutely. Watch how much more structure this has than a job title.
FOCUS: Examine this code ONLY for security. Attend, in order, to:
(1) untrusted input reaching a sink (injection, SSRF, deserialize),
(2) authn/authz on every state-changing path,
(3) secrets or credentials in code, logs, or errors,
(4) cryptographic misuse.
PRIORITY: When security conflicts with performance or readability,
security wins. Say so explicitly rather than balancing them.
REFUSE: Never approve code that logs credentials, disables cert
verification, or builds SQL by string concatenation — regardless
of surrounding context or author justification. These are blocking,
not advisory.
METHOD: For each finding, name the exact line, the attack it
enables, and the minimal fix. Rank by exploitability, not by how
easy the fix is.
What this does: it specifies attention order, a conflict-resolution priority, hard refusals, and a method — giving the agent the actual structure of a security specialist's judgment instead of a title and hoping expertise emerges from it.
The difference is that this prompt changes behavior, not just tone. The attention order makes the agent examine the dangerous things first and thoroughly. The priority rule tells it how to resolve the conflicts that generalists fudge. The refusals give it hard lines it won't rationalize past. The method standardizes its output. None of that comes from "you are an expert."
⚡ Pro tip: build specialization from refusals first, then attention, then priority. The refusals are what most sharply distinguish a specialist from a generalist, because they're the lines the specialist holds that the generalist negotiates. Start by writing down what this agent must never accept, and you'll find the rest of the specialization follows from defending those lines. Most specialization prompts skip refusals entirely, which is exactly why they're shallow.
Breaking down each element
The attention order is doing the heaviest lifting. By telling the agent what to look at first and in what sequence, you replicate the specialist's trained instinct for where danger lives. A generalist distributes attention evenly; a specialist front-loads it on the high-risk patterns, and the explicit order encodes that.
The priority rule handles the conflicts that reveal whether specialization is real. Every non-trivial task has moments where security and performance, or thoroughness and speed, pull against each other. A generalist balances them mushily. A specialist has a clear hierarchy — security wins, and here's why — and stating that hierarchy explicitly is what makes the agent decide like a specialist under conflict.
The refusals are the backbone. They're the difference between an agent that flags a credential-logging line as a concern and one that blocks the merge over it. Advisory specialists get overridden; specialists with hard refusals hold the line, and that holding is often the entire value of having a specialist in the loop.
⚡ Pro tip: give each specialist a short list of its own domain's most common false positives to suppress. A security specialist that flags every string operation as potential injection becomes noise. Telling it explicitly which safe-looking patterns are actually fine — parameterized queries, escaped output — sharpens it from a paranoid generalist into a calibrated specialist. Real expertise is as much about what you correctly ignore as what you catch, and generalists-in-costume flag everything.
Variations for different contexts
A financial-analysis specialist gets an attention order over the statements that hide problems, a priority rule favoring conservatism over optimism, and refusals around unsupported projections. A medical-information specialist gets attention on contraindications, a priority favoring safety over completeness, and hard refusals around anything that could be read as diagnosis. A legal-review specialist gets attention on liability clauses, a priority favoring risk-flagging over reassurance, and refusals around giving definitive legal conclusions.
The four-component structure is identical across all of them — attention, order, priority, refusal. Only the domain content changes. That structure is the reusable template for building any real specialist, and it's what "you are an expert" fundamentally lacks.
⚠️ Common mistake: copying one specialist's structure to a new domain without rethinking the refusals. The attention order and priority often transfer loosely, but the refusals are deeply domain-specific — a security specialist's hard lines mean nothing to a financial one. Teams that clone a specialist and swap the domain word produce a costume again, because they kept the structure but not the domain-specific judgment that made it real. Rewrite the refusals for every new specialist from scratch.
How do you test whether specialization actually took?
Writing structured agent specialization prompts is only half the job — you have to verify the specialization actually changed behavior rather than just tone, because a well-written prompt can still produce costume specialization if the model doesn't internalize the structure. The test is a differential one: run the same input through your specialist and through a plain generalist prompt, then compare not the wording but the decisions.
Real specialization shows up as different decisions, not different vocabulary. If your security specialist and a generalist both flag the same issues in the same order and only the specialist uses more security jargon, the specialization didn't take — you have a generalist in a costume. If the specialist catches issues the generalist missed, examines the dangerous patterns first, and holds a refusal the generalist negotiated away, the structure worked. The decision delta, not the tone delta, is the measure.
The sharpest test is the refusal test. Construct an input that a generalist would rationalize approving but a true specialist must refuse — code that logs credentials wrapped in an otherwise clean, well-justified diff. A generalist balances the good against the bad and often approves. A properly specialized agent blocks on the refusal regardless of the surrounding quality. If your specialist approves that input, its refusals are decorative, and you need to strengthen them until they hold under pressure.
There's a calibration test too. Feed the specialist inputs that look dangerous but are actually safe — the domain's common false positives. A miscalibrated specialist flags them all, revealing it's a paranoid generalist rather than a real expert. A calibrated one correctly passes them. Running both the refusal test and the false-positive test together tells you whether your agent specialization prompts produced genuine judgment or just a more confident voice.
⚡ Pro tip: keep a small regression suite of decision tests for each specialist and re-run it whenever you change the prompt or the model. Specialization is fragile — a prompt tweak that seems harmless can quietly turn a specialist back into a costume, and a model update can shift how it weighs your refusals. A handful of decision tests, checking that the specialist still catches what it should and refuses what it must, catches that regression before it ships. Tone is easy to eyeball; behavior needs a test.
Save and reuse this
The four-component structure here — attention, order, priority, refusal — is the reusable pattern for building any specialist that's more than a costume. The domain content changes; the structure that turns a generalist into a specialist stays constant.
Save your specialist prompts in PromptABCD as a structured library, organized by the four components rather than as opaque blobs of text. When you build a new specialist, you'll start from the proven structure and fill in the domain, and when you improve how you write refusals, every specialist can inherit the improvement — which is impossible if each one lives as a copy-pasted job title in a different config file.
Continue Reading
Save the prompts from this post
PromptABCD is a free prompt manager. Paste, organize, and reuse your best AI prompts — no more hunting through chat history.
