Multi-Agent Simulations of Human Behavior: A Case Study
Most behavior-simulation guides oversell it — agents don't predict what people will do. A multi agent behavior simulation is a hypothesis generator, not a crystal ball. Here's a real study of what it got right and wrong.
def customer_agent(profile, proposed_pricing):
system = f"""You are a customer with this profile:
Usage: {profile['usage']}. Price sensitivity: {profile['sensitivity']}.
Tenure: {profile['tenure']}. What you value: {profile['values']}.
You just received this pricing change: {proposed_pricing}.
React honestly as THIS person. Output JSON:
{{"reaction": "...", "likely_action": "stay|upgrade|downgrade|churn",
"why": "...", "what_would_change_my_mind": "..."}}"""
return json.loads(llm(system=system, user="React now."))Most guides on simulating people with agents oversell the whole idea. They imply that if you spin up enough agent "personas," you can predict how a real crowd will behave — how customers react to a price change, how a community responds to a policy. That's the wrong promise. A multi agent behavior simulation doesn't predict what people will do. It generates hypotheses about what might happen and surfaces failure modes you hadn't considered. Treat it as a crystal ball and it will mislead you. Treat it as a stress-test and it earns its keep.
This case study follows a product team that used agent-based simulation before a pricing change, what it got right, what it got embarrassingly wrong, and how to use the tool honestly.
What was the team actually trying to answer?
Devin's team at a subscription software company was about to restructure pricing — moving from flat to usage-based. The finance model said revenue would rise. The nervous question was behavioral: how would existing customers react? Would power users churn? Would light users upgrade or leave? A survey would take weeks and anchor on hypotheticals. Devin wanted a faster way to pressure-test the plan.
The team built a multi agent behavior simulation: a few hundred agents, each seeded with a customer profile — usage level, price sensitivity, tenure, stated priorities drawn from real segments — reacting to the proposed pricing. The goal wasn't a revenue number. It was to find reactions the team hadn't anticipated.
[object Object], ,[object Object],(,[object Object],):
system = ,[object Object],
,[object Object], json.loads(llm(system=system, user=,[object Object],))What this does: it seeds each agent with a concrete customer profile and asks for a specific action plus the reasoning behind it, producing a spread of reactions the team can inspect for surprises rather than a single averaged prediction.
The wrong way the team first used it
Devin's first instinct was to treat the aggregate as a forecast. The simulation showed twelve percent churn, so the team started planning around twelve percent churn. That was the mistake, and it took a reality check to see it.
The number was meaningless as a prediction. The agents didn't share real customers' inertia, switching costs, or the simple fact that most people ignore pricing emails entirely. The simulation's churn rate was an artifact of how articulate and rational the agents were — real humans are neither as articulate nor as prompt to act. Any team that had budgeted around that twelve percent would have been planning around fiction.
⚠️ Common mistake: reading the aggregate output of a behavior simulation as a probability. The agents are too rational, too responsive, and too verbal to match real population dynamics. The numbers look precise and mean almost nothing. The value was never in the aggregate — it was in the individual reactions, and the team nearly missed it by staring at the total.
The correct use: mining reactions, not predictions
The reframe was to ignore the churn percentage entirely and read the individual why and what_would_change_my_mind fields. That's where the real value hid. Three specific reactions changed the launch plan.
First, several high-usage agents flagged that the new pricing punished exactly the loyal power users the company most wanted to keep — a fairness perception the finance model was blind to. Second, several agents raised a billing-predictability concern: usage-based meant unpredictable bills, and predictability mattered more to them than absolute cost. Third, a cluster of agents pointed out that the change would be read as a stealth increase regardless of actual math, a messaging landmine.
None of these were predictions. All three were hypotheses the team could then check against real customers cheaply, because now they knew what to ask. The simulation didn't forecast behavior — it told the team which questions were worth a real survey.
⚡ Pro tip: prompt agents for what would change their mind, not just what they'd do. The "what_would_change_my_mind" field is where actionable insight lives. It surfaces the levers — a grandfather clause, a bill cap, clearer messaging — that turn a churning customer into a retained one. A simulation that only asks "what will you do" gives you a number; one that asks "what would move you" gives you a plan.
Results and what changed
The team didn't launch the original plan. Based on the mined reactions, they added a grandfather clause for long-tenured power users, a monthly bill cap to preserve predictability, and reframed the announcement around value rather than usage. Then they ran a real, targeted survey on exactly the three concerns the simulation surfaced — a survey they could write precisely because the simulation had pointed at the questions.
The honest outcome: the simulation's headline number was wrong, and its individual reactions were genuinely useful. The team got value by using it as a hypothesis generator and a survey-design tool, not as a forecast. That distinction is the whole lesson.
⚡ Pro tip: seed agents from real segments, not imagined ones. The simulation's reactions are only as good as the profiles behind them. Pull usage levels, tenures, and stated values from your actual customer data so the agents reason from real distributions. Agents built from made-up personas generate plausible-sounding reactions that don't map to anyone real, and you can't tell the difference from the output.
How to apply this to your situation
Use a multi agent behavior simulation anywhere you're about to make a decision that affects many people and you want to find objections before they find you — a policy change, a feature deprecation, a communication rollout. Seed agents from real segments, ask for reactions and mind-changers, and read the individuals, never the aggregate.
Then treat every interesting reaction as a hypothesis to validate cheaply with real people, not as a conclusion. The simulation's job is to make your subsequent research sharper, not to replace it. The teams that get burned are the ones who stop at the simulation; the teams that get value use it to aim what comes next.
⚠️ Common mistake: scaling up the agent count to make the numbers more "reliable." More agents give you a more precise wrong number. The aggregate doesn't become predictive with scale — it becomes more confidently non-predictive. Spend the effort on better profiles and richer reaction prompts, not on running ten thousand agents to get a decimal place on a figure that was never real.
What can a multi agent behavior simulation actually be trusted for?
After the pricing project, Devin's team drew a boundary they now apply to every simulation, and it's the most useful takeaway from the whole exercise. A multi agent behavior simulation is trustworthy for divergent thinking and untrustworthy for convergent estimation. It's good at generating a wide set of possible reactions; it's bad at telling you how likely any of them is.
That boundary maps to a simple rule for when to reach for it. Use it when your risk is missing an objection — a reaction, a fairness concern, a messaging trap you didn't anticipate. Don't use it when your risk is mis-estimating a magnitude — how many will churn, what percentage will upgrade. For the first kind of question, agents are a cheap idea machine. For the second, you need real data, because the agents' rationality systematically distorts magnitudes.
There's a subtler validity issue worth understanding. Agent reactions cluster around what's articulable, and much real human behavior isn't articulable at all. People stay with a product out of habit they'd never state, or churn over an accumulated irritation no single email would capture. The simulation can only reason about reasons it can express, so it's blind to the inarticulate inertia that drives much real behavior. Knowing this blind spot tells you exactly where to distrust the output.
⚡ Pro tip: run the simulation twice with deliberately different framings of the same change and compare which reactions are stable across both. Reactions that appear regardless of framing are stable signals worth investigating; reactions that flip with the wording are artifacts of how you prompted, not real concerns. This cross-framing check separates genuine hypotheses from prompt-sensitivity noise, and almost no one does it. The stable-across-framings reactions are the ones worth spending real survey budget to confirm.
The validity also degrades with how far the scenario sits from ordinary experience. Agents reason reasonably about familiar situations — a price change, a new feature — because those resemble things in their training. They reason poorly about genuinely novel situations with no precedent, confidently generating reactions that have no grounding. The further your scenario is from the everyday, the more you should treat the output as pure brainstorm and the less as informed hypothesis. When in doubt about which side of that line you're on, assume the scenario is novel enough to demand real validation before you act on anything the agents produced.
Next steps
Take a decision you're nervous about and run fifty profile-seeded agents through it this week. Ignore the totals completely and read every "what would change my mind" — then write your real survey from what you find.
As you build agent profiles and reaction prompts that surface useful hypotheses, save them in PromptABCD. A multi agent behavior simulation is a reusable instrument once tuned, and keeping your profile templates and reaction schemas versioned means your next launch stress-test starts from a proven setup instead of a blank page.
Continue Reading
Save the prompts from this post
PromptABCD is a free prompt manager. Paste, organize, and reuse your best AI prompts — no more hunting through chat history.
