Multi-Agent Systems for Trading and Finance: A Teardown
Picture a quant who wired one agent to research, decide, and execute trades — and watched it talk itself into a position. A multi agent trading system separates those jobs. Here's the weak design and the rebuild.
You are an expert trading agent. Analyze {ticker} using recent
price action, news, and fundamentals. Decide whether to BUY,
SELL, or HOLD, choose a position size, and place the order.
Be decisive and act on your best judgment.Picture this: you're a quant who wired a single agent to research a stock, decide whether to buy, and place the order. In backtests it looked sharp. Live, it did something no backtest showed — it found a bullish thesis, then interpreted every subsequent data point as confirming that thesis, and sized the position as if it were certain. It didn't just make a bad trade. It talked itself into one and then congratulated itself. That confirmation spiral is the core reason a multi agent trading system exists, and it's what this teardown is about.
We'll take apart the naive single-agent design, explain precisely why it fails in markets specifically, and rebuild it as separated agents with an adversarial risk check. To be clear up front: this is about architecture, not trading advice, and nothing here is a recommendation to trade.
Before: the weak single-agent design
Here's the design that looks reasonable and fails in production:
You are an expert trading agent. Analyze {ticker} using recent
price action, news, and fundamentals. Decide whether to BUY,
SELL, or HOLD, choose a position size, and place the order.
Be decisive and act on your best judgment.
What this does: it hands one agent research, decision, sizing, and execution with an instruction to be decisive — which in a market means committing hard to whatever narrative it constructs first.
It reads like competence. It's a loaded gun pointed at your account, for reasons specific to how markets and language models interact.
Why it fails
Markets punish exactly the failure mode language models are most prone to: narrative lock-in. A model builds a story and then weights evidence to fit it. In most tasks this is a mild bias. In trading it's ruinous, because the whole game is updating against your prior when new evidence arrives, and a single agent that generated the thesis is the worst possible judge of evidence against it.
The "be decisive" instruction makes it worse. Decisiveness is a virtue when the analysis is sound and a vice when it isn't, and the agent can't tell the difference from inside its own reasoning. So it sizes a shaky thesis and a strong one identically — with confidence. There's no independent check on conviction because the agent that researched is the agent that sized.
⚠️ Common mistake: assuming a good backtest validates the architecture. Backtests reward the single agent's decisiveness because you're testing on data where the outcome is fixed. Live, the same decisiveness applied to genuine uncertainty is what blows up accounts. A design that backtests beautifully can be structurally unsafe, and the teardown here is about structure, not returns.
The deepest problem is that execution is welded to decision. An agent that decides and executes in one breath has no gap where a risk limit can intervene. By the time anything could say "that position is too large," the order is placed. The missing seam is the whole safety mechanism.
After: the improved multi agent trading system
The rebuild separates four roles and, crucially, makes one of them adversarial. A researcher builds the thesis. A skeptic argues the opposite case. A risk manager enforces hard limits and can veto. An executor places only what survives, and only within pre-set bounds.
[object Object], ,[object Object],(,[object Object],):
system = (,[object Object],
,[object Object],
,[object Object],)
,[object Object], llm(system=system, user=,[object Object],)
,[object Object], ,[object Object],(,[object Object],):
system = (,[object Object],
,[object Object],
,[object Object],
,[object Object],
,[object Object],)
,[object Object], json.loads(llm(system=system,
user=,[object Object],))What this does: it forces the researcher to hold both sides of the trade with explicit probabilities and gives a separate risk manager a hard veto keyed to portfolio limits — so no single agent's conviction can size a position on its own.
The researcher is required to build the bear case, not just the bull case. That single requirement breaks the confirmation spiral, because the agent can't pretend the counter-evidence doesn't exist when its own instruction demands it articulate the counter-thesis. Forcing both sides is cheap and changes the reasoning quality dramatically.
[object Object], ,[object Object],(,[object Object],):
,[object Object],
,[object Object], order[,[object Object],] > limits[,[object Object],]:
,[object Object], {,[object Object],: ,[object Object],, ,[object Object],: ,[object Object],}
,[object Object], ,[object Object], order[,[object Object],]:
,[object Object], {,[object Object],: ,[object Object],, ,[object Object],: ,[object Object],}
,[object Object], place(order) ,[object Object],What this does: it makes execution a dumb gate that refuses anything without explicit risk approval and anything over the hard size limit, so the decision-to-execution seam a single agent lacks is enforced in code, not prompt.
⚡ Pro tip: put the hard limits in code, not in a prompt. A risk instruction in a system prompt is a suggestion the model can rationalize past when a thesis feels compelling. A size check in the executor is a wall. Anything that protects capital belongs in deterministic code where no amount of persuasive reasoning can override it.
Breaking down each element
The researcher's power is being forbidden to recommend. By stripping its authority to decide, you turn it into an evidence generator rather than an advocate, which is what you actually want from analysis. Advocacy is where narrative lock-in lives.
The skeptic and the two-sided thesis are the antidote to confirmation bias. An independent agent whose only job is the counter-case has no stake in the trade going through, so it surfaces the risks the researcher would minimize. Opposed incentives, again, are the mechanism.
The risk manager's power is the veto plus indifference to upside. A risk agent that also cared about returns would approve marginal trades. One that cares only about surviving losses holds the line. And the executor's power is being deliberately unintelligent — a gate, not a decider, so there's nothing to persuade.
⚡ Pro tip: log every vetoed trade and review the vetoes weekly, not just the executed trades. Your risk manager's rejections are a record of the bad positions your researcher wanted to take. A rising veto rate on a particular sector or thesis type tells you the researcher is developing a blind spot — an early-warning signal you'd never see if you only reviewed trades that happened.
Variations for different contexts
A long-term value strategy weights the researcher heavily and sets loose size limits, since positions are held through volatility. A high-frequency-adjacent setup can't afford a language model in the execution path at all — there the multi-agent design lives entirely in pre-trade research and risk, with execution fully deterministic. A portfolio-management use case adds an allocator agent that the risk manager still gets to veto, keeping the "one agent can't size alone" rule intact.
Across all of them, two things never change: the researcher argues both sides, and the risk manager holds a hard, code-enforced veto. Those are the load-bearing walls.
⚠️ Common mistake: giving the risk manager a language-model veto without a code backstop. If the risk manager is "just another agent," a cleverly framed proposal can talk it into approval, and you're back to one persuadable system. The code-level limits must exist independently of any agent's judgment, so that even a compromised risk agent can't authorize a catastrophic size.
How do you test a multi agent trading system safely?
The teardown fixed the architecture, but a safe architecture deployed carelessly still loses money. Testing a multi agent trading system demands more caution than a normal software rollout, because the failure mode isn't a crash — it's a plausible-looking bad decision that costs capital before anyone notices.
The staged path that limits damage has three gates. First, paper trading: run the full pipeline against live market data but route the executor to a simulated account, so the agents make real decisions with zero capital at risk. Watch not just the returns but the reasoning — read the researcher's bull and bear cases and the risk manager's vetoes. You're checking whether the agents reason soundly, not whether they got lucky on a sample.
Second, shadow mode with tiny size: once paper trading looks sound over enough varied conditions, let the executor place real orders at a fraction of your intended size — small enough that a total loss is an acceptable tuition payment. This surfaces the gap between simulated fills and real ones, and the psychological reality of watching an agent trade your actual money.
Third, capped live with a kill switch: scale size gradually, always behind a hard daily loss limit enforced in code that halts all trading when hit. The kill switch is non-negotiable. It's the one control that turns a runaway failure from catastrophic into merely bad.
⚡ Pro tip: test the agents specifically on regime changes, not just calm markets. A multi agent trading system that reasons well in a steady trend can fall apart when volatility spikes, because the researcher's confidence and the risk manager's limits interact differently under stress. Deliberately replay historical crisis periods and watch whether the risk manager's vetoes increase appropriately — if they don't, your limits are miscalibrated for the conditions that matter most.
The measurement that matters during testing isn't raw return — it's whether the risk manager's vetoes correlate with trades that would have lost money. A risk manager vetoing trades that would have been fine is too tight; one approving trades that lose is too loose. Tracking veto quality, not just profit, tells you whether the safety layer is actually working or just occasionally guessing right.
⚠️ Common mistake: scaling size based on a good week. A short winning streak in a favorable regime tells you almost nothing about the system's behavior when conditions turn. Scale on demonstrated risk discipline across varied conditions, not on recent returns — the returns that convince you to size up are exactly the ones that precede the drawdown that hurts.
Save and reuse this
The researcher, skeptic, risk-manager, and executor prompts here are a safe skeleton for a multi agent trading system — a research and risk architecture, not a money-printing machine, and worth treating with the seriousness the domain demands.
Save this prompt set in PromptABCD as a versioned kit, and pair it with a written record of your hard limits so the code backstops and the prompts never drift apart. In a domain where a prompt regression can cost real money, having one reviewed source of truth for your research and risk agents is worth far more than the few minutes it takes to store them properly.
Continue Reading
Save the prompts from this post
PromptABCD is a free prompt manager. Paste, organize, and reuse your best AI prompts — no more hunting through chat history.
