Voting and Consensus in Agent Teams
When multiple agents disagree, how does your system decide what's true? Multi agent consensus is more nuanced than majority vote. Here's the full decision framework.
import asyncio
import json
from anthropic import Anthropic
client = Anthropic()
SPECIALIST_CONFIGS = [
{"name": "analyst_1", "focus": "First-principles reasoning from known facts"},
{"name": "analyst_2", "focus": "Pattern matching from historical precedent"},
{"name": "analyst_3", "focus": "Risk assessment and failure mode analysis"},
{"name": "analyst_4", "focus": "Quantitative modeling and estimation"},
{"name": "analyst_5", "focus": "Stakeholder impact and practical implementation"},
]
async def run_specialist(config: dict, question: str) -> dict:
"""Run one specialist and collect their response with confidence."""
system = f"""You are an analytical agent focused on: {config['focus']}.
Analyze the question from your specific analytical perspective.
Return JSON:
{{
"answer": "Your conclusion",
"confidence": 0.0-1.0,
"reasoning": "Key reasoning steps",
"uncertainty_factors": ["factor1", "factor2"]
}}"""
response = await asyncio.to_thread(
client.messages.create,
model="claude-opus-4-5",
max_tokens=1024,
system=system,
messages=[{"role": "user", "content": question}]
)
result = json.loads(response.content[0].text)
result["specialist"] = config["name"]
return result
async def collect_specialist_responses(question: str) -> list[dict]:
tasks = [run_specialist(cfg, question) for cfg in SPECIALIST_CONFIGS]
return await asyncio.gather(*tasks)
def simple_majority_vote(responses: list[dict]) -> dict:
"""Count the most common answer."""
from collections import Counter
answers = [r["answer"] for r in responses]
vote_counts = Counter(answers)
winner = vote_counts.most_common(1)[0][0]
agreement_rate = vote_counts[winner] / len(responses)
return {
"method": "majority_vote",
"conclusion": winner,
"agreement_rate": agreement_rate,
"vote_breakdown": dict(vote_counts)
}
def confidence_weighted_vote(responses: list[dict]) -> dict:
"""Weight each answer by stated confidence."""
from collections import defaultdict
weighted_scores = defaultdict(float)
for r in responses:
weighted_scores[r["answer"]] += r["confidence"]
winner = max(weighted_scores, key=weighted_scores.get)
total_weight = sum(weighted_scores.values())
return {
"method": "confidence_weighted",
"conclusion": winner,
"confidence_score": weighted_scores[winner] / total_weight,
"weighted_breakdown": dict(weighted_scores)
}
def synthesize_responses(responses: list[dict], question: str) -> dict:
"""Use a synthesis agent to evaluate reasoning quality and produce a conclusion."""
panel_text = "\n\n".join([
f"Specialist: {r['specialist']}\nFocus: {next(c['focus'] for c in SPECIALIST_CONFIGS if c['name'] == r['specialist'])}\nAnswer: {r['answer']}\nConfidence: {r['confidence']}\nReasoning: {r['reasoning']}"
for r in responses
])
response = client.messages.create(
model="claude-opus-4-5",
max_tokens=1024,
system="""You are a synthesis analyst. Evaluate the quality of reasoning from each specialist — not just their conclusion — and produce the most defensible answer.
Return JSON:
{
"conclusion": "Best-supported conclusion",
"confidence": 0.0-1.0,
"strongest_reasoning": "Which specialist's reasoning was most sound and why",
"key_uncertainties": ["remaining uncertainty 1", "uncertainty 2"],
"dissenting_points": "Where minority views had valid points"
}""",
messages=[{"role": "user", "content": f"Question: {question}\n\nSpecialist panel:\n\n{panel_text}"}]
)
result = json.loads(response.content[0].text)
result["method"] = "expert_synthesis"
return result
async def run_consensus_pipeline(question: str, method: str = "auto") -> dict:
responses = await collect_specialist_responses(question)
avg_confidence = sum(r["confidence"] for r in responses) / len(responses)
# Check for natural consensus
unique_answers = set(r["answer"] for r in responses)
if len(unique_answers) == 1:
return {
"method": "unanimous",
"conclusion": responses[0]["answer"],
"confidence": avg_confidence,
"specialist_count": len(responses)
}
if method == "auto":
# High disagreement → synthesis; moderate → confidence-weighted; low → majority
if len(unique_answers) > len(responses) * 0.6:
method = "synthesis"
elif avg_confidence > 0.75:
method = "confidence_weighted"
else:
method = "majority_vote"
if method == "majority_vote":
return {**simple_majority_vote(responses), "specialist_responses": responses}
elif method == "confidence_weighted":
return {**confidence_weighted_vote(responses), "specialist_responses": responses}
else:
return {**synthesize_responses(responses, question), "specialist_responses": responses}Multi agent consensus solves a specific problem: when you run multiple agents on the same task and they disagree, your system needs a principled way to combine or adjudicate their outputs. The naive answer — take the majority — works adequately for simple classification tasks but fails for complex reasoning tasks where the most common answer is not the most correct one.
Understanding when to use simple voting, confidence-weighted voting, and expert synthesis is the difference between a consensus mechanism that improves accuracy and one that launders low-quality answers into false confidence.
Three Consensus Approaches
The choice between consensus methods depends on your task type, the independence of your agents, and what accuracy means for your use case.
Simple majority voting works when: responses are categorical (classification labels, yes/no decisions), agents are genuinely independent, and the task has a single correct answer. It fails when answers are continuous, agents share training data biases, or the minority view is systematically more likely to be correct (contrarian consensus problems).
Confidence-weighted voting extends majority voting by weighting each agent's vote by its stated confidence. Agents that express high confidence have more influence over the final answer. This works when agent confidence is well-calibrated — when high-confidence answers are actually more likely to be correct. It fails when agents are systematically overconfident (common with LLMs on tasks outside their core training distribution).
Expert synthesis runs a dedicated synthesis agent that evaluates the reasoning quality of each participating agent's response and produces a conclusion informed by all of them. This is the most expensive approach and the most effective for high-stakes tasks where the quality of reasoning matters as much as the answer.
Building the Interactive Consensus Pipeline
[object Object], asyncio
,[object Object], json
,[object Object], anthropic ,[object Object], Anthropic
client = Anthropic()
SPECIALIST_CONFIGS = [
{,[object Object],: ,[object Object],, ,[object Object],: ,[object Object],},
{,[object Object],: ,[object Object],, ,[object Object],: ,[object Object],},
{,[object Object],: ,[object Object],, ,[object Object],: ,[object Object],},
{,[object Object],: ,[object Object],, ,[object Object],: ,[object Object],},
{,[object Object],: ,[object Object],, ,[object Object],: ,[object Object],},
]
,[object Object], ,[object Object], ,[object Object],(,[object Object],) -> ,[object Object],:
,[object Object],
system = ,[object Object],
response = ,[object Object], asyncio.to_thread(
client.messages.create,
model=,[object Object],,
max_tokens=,[object Object],,
system=system,
messages=[{,[object Object],: ,[object Object],, ,[object Object],: question}]
)
result = json.loads(response.content[,[object Object],].text)
result[,[object Object],] = config[,[object Object],]
,[object Object], result
,[object Object], ,[object Object], ,[object Object],(,[object Object],) -> ,[object Object],[,[object Object],]:
tasks = [run_specialist(cfg, question) ,[object Object], cfg ,[object Object], SPECIALIST_CONFIGS]
,[object Object], ,[object Object], asyncio.gather(*tasks)
,[object Object], ,[object Object],(,[object Object],) -> ,[object Object],:
,[object Object],
,[object Object], collections ,[object Object], Counter
answers = [r[,[object Object],] ,[object Object], r ,[object Object], responses]
vote_counts = Counter(answers)
winner = vote_counts.most_common(,[object Object],)[,[object Object],][,[object Object],]
agreement_rate = vote_counts[winner] / ,[object Object],(responses)
,[object Object], {
,[object Object],: ,[object Object],,
,[object Object],: winner,
,[object Object],: agreement_rate,
,[object Object],: ,[object Object],(vote_counts)
}
,[object Object], ,[object Object],(,[object Object],) -> ,[object Object],:
,[object Object],
,[object Object], collections ,[object Object], defaultdict
weighted_scores = defaultdict(,[object Object],)
,[object Object], r ,[object Object], responses:
weighted_scores[r[,[object Object],]] += r[,[object Object],]
winner = ,[object Object],(weighted_scores, key=weighted_scores.get)
total_weight = ,[object Object],(weighted_scores.values())
,[object Object], {
,[object Object],: ,[object Object],,
,[object Object],: winner,
,[object Object],: weighted_scores[winner] / total_weight,
,[object Object],: ,[object Object],(weighted_scores)
}
,[object Object], ,[object Object],(,[object Object],) -> ,[object Object],:
,[object Object],
panel_text = ,[object Object],.join([
,[object Object],
,[object Object], r ,[object Object], responses
])
response = client.messages.create(
model=,[object Object],,
max_tokens=,[object Object],,
system=,[object Object],,
messages=[{,[object Object],: ,[object Object],, ,[object Object],: ,[object Object],}]
)
result = json.loads(response.content[,[object Object],].text)
result[,[object Object],] = ,[object Object],
,[object Object], result
,[object Object], ,[object Object], ,[object Object],(,[object Object],) -> ,[object Object],:
responses = ,[object Object], collect_specialist_responses(question)
avg_confidence = ,[object Object],(r[,[object Object],] ,[object Object], r ,[object Object], responses) / ,[object Object],(responses)
,[object Object],
unique_answers = ,[object Object],(r[,[object Object],] ,[object Object], r ,[object Object], responses)
,[object Object], ,[object Object],(unique_answers) == ,[object Object],:
,[object Object], {
,[object Object],: ,[object Object],,
,[object Object],: responses[,[object Object],][,[object Object],],
,[object Object],: avg_confidence,
,[object Object],: ,[object Object],(responses)
}
,[object Object], method == ,[object Object],:
,[object Object],
,[object Object], ,[object Object],(unique_answers) > ,[object Object],(responses) * ,[object Object],:
method = ,[object Object],
,[object Object], avg_confidence > ,[object Object],:
method = ,[object Object],
,[object Object],:
method = ,[object Object],
,[object Object], method == ,[object Object],:
,[object Object], {**simple_majority_vote(responses), ,[object Object],: responses}
,[object Object], method == ,[object Object],:
,[object Object], {**confidence_weighted_vote(responses), ,[object Object],: responses}
,[object Object],:
,[object Object], {**synthesize_responses(responses, question), ,[object Object],: responses}What this does: Five specialist agents run in parallel, each analyzing the question through a different analytical lens. The pipeline checks for unanimous agreement first — the cheapest outcome. If disagreement exists, the method selection is automatic: high disagreement levels trigger expert synthesis because simple voting on widely divergent answers produces noise; moderate disagreement with high confidence triggers weighted voting; low confidence triggers simple majority. The automatic routing avoids paying for synthesis on easy questions while ensuring it's used when the quality of reasoning actually matters.
⚡ Pro tip: Diversify your specialists' information sources, not just their system prompt framing. Two agents with different prompts but identical model weights and no tools have correlated failures — they make similar errors on similar inputs. For genuine diversity, give some specialists web search access while others reason from training knowledge only. Different information sources produce more independent disagreement than different prompt phrasings.
Choosing the Right Method for Your Task
Use majority voting when: The task is classification or categorical selection; you have three or more agents; and you've validated that your agents produce genuinely independent outputs (not correlated by shared model biases).
Use confidence-weighted voting when: Your agents' confidence scores have been calibrated against ground truth (i.e., when an agent says 0.9, it's right about 90% of the time). Without calibration, confidence weighting can amplify systematic overconfidence rather than improve accuracy.
Use expert synthesis when: Answers are continuous or complex; the task rewards reasoning quality over answer frequency; or the cost of a wrong answer justifies an additional LLM call. Synthesis is roughly 20-30% more expensive than the specialist panel alone but produces meaningfully better results on open-ended reasoning tasks.
⚡ Pro tip: Run a calibration study before committing to confidence-weighted voting. Collect 100 questions with known answers, run your specialist panel, and plot confidence scores against accuracy. If the correlation is weak (agents claiming 0.9 confidence are right only 70% of the time), confidence weighting will make your consensus worse, not better. Use simple majority or synthesis instead.
When to Avoid Consensus Mechanisms
Multi agent consensus adds cost and latency. It's not always justified.
Consensus adds no value when agents are not independent. If all five specialists use the same model without tool access on the same input, their outputs will be correlated. Voting correlated outputs doesn't improve accuracy — it just makes the same error more confidently.
Consensus is wrong for divergence tasks. Asking five agents to brainstorm and then voting on the most common suggestion produces the most average idea, not the best one. For creative or generative tasks, collect all responses and present them as options rather than collapsing them into a single consensus answer.
⚠️ Common mistake: Building consensus pipelines without measuring whether they improve accuracy over a single agent. The assumption that "more agents = more accurate" is wrong in many domains. Measure first. Add agents second.
Building Your Panel Prompt Library
A well-tuned specialist panel for a specific domain — legal analysis, financial modeling, technical review — is a reusable infrastructure asset. The specialist system prompts that produce reliable, genuinely independent analysis are the hard-won result of calibration work that shouldn't be redone for each project.
Measuring Consensus Quality
Consensus mechanisms are only useful if they improve accuracy over a single agent. Measuring this requires a ground truth dataset: questions with verifiable correct answers that you can compare against both single-agent and consensus outputs.
The measurement methodology is straightforward:
[object Object], ,[object Object],(,[object Object],) -> ,[object Object],:
single_correct = ,[object Object],
consensus_correct = ,[object Object],
,[object Object], question, truth ,[object Object], ,[object Object],(test_questions, ground_truth):
single_answer = single_agent_fn(question)
consensus_answer = consensus_fn(question)
single_correct += ,[object Object],(single_answer.lower() == truth.lower())
consensus_correct += ,[object Object],(consensus_answer.lower() == truth.lower())
n = ,[object Object],(test_questions)
,[object Object], {
,[object Object],: single_correct / n,
,[object Object],: consensus_correct / n,
,[object Object],: (consensus_correct - single_correct) / n,
,[object Object],: ,[object Object],(SPECIALIST_CONFIGS) + ,[object Object], ,[object Object],
}What this does: The measurement function runs both a single agent and the consensus pipeline on the same test questions and compares accuracy against known correct answers. The cost_multiplier field captures the cost ratio — how many times more expensive is the consensus approach. If the consensus approach is 20% more accurate at 5x the cost, you have concrete data to decide whether the tradeoff is worth it for your specific use case.
⚡ Pro tip: Run the accuracy measurement on a stratified sample: easy questions, medium questions, and hard questions (based on single-agent confidence or question type). Consensus mechanisms typically add the most value on hard questions and the least on easy ones. If your workload is 80% easy questions, the consensus overhead may not be justified even if consensus improves accuracy significantly on the 20% that are hard. Stratified measurement reveals whether selective consensus (only for hard inputs) is a better approach than universal consensus.
When Confidence Scores Mislead
LLMs are systematically overconfident on certain input types. Questions that touch on topics the model has seen frequently in training — even when asking about specific edge cases in that domain — often receive high confidence scores regardless of the actual correctness. Questions about less common domains receive more calibrated (lower) confidence scores.
This means confidence-weighted voting can amplify high-confidence errors rather than correcting them. If your specialists consistently express high confidence on a category of inputs where they're systematically wrong, weighting by confidence makes the consensus worse than simple majority voting on that category.
The practical detection method: plot confidence scores against accuracy by question category. If there are categories where high-confidence answers are wrong more than 20% of the time, exclude those categories from confidence-weighted voting and use simple majority or synthesis instead.
Designing Independent Specialists
The quality of multi agent consensus depends critically on the independence of the specialists. Two agents with the same model, same tools, and only slightly different system prompts will produce highly correlated outputs — they'll make the same errors on the same inputs. Voting correlated outputs doesn't improve accuracy; it just makes the same error more confidently.
Achieving genuine independence requires at least one of: different information access (one agent can search the web, another reasons from training data only), different analytical frames (one focuses on quantitative evidence, another on qualitative signals), or genuinely different models with different training distributions.
For most production systems, different analytical frames is the most practical approach. A financial specialist focused on "what do the numbers say" and a strategic specialist focused on "what does the competitive context suggest" will have genuinely independent failure modes even if they share the same underlying model. Frame diversity produces more genuine disagreement than prompt wording variation alone.
Ensemble Size and Diminishing Returns
Multi agent consensus improves accuracy up to a point, then plateaus. The accuracy improvement from adding a fourth specialist is smaller than the improvement from adding a third, which is smaller than the improvement from adding a second. At some ensemble size — typically five to seven specialists for most tasks — additional agents contribute noise rather than signal.
The plateau point depends on specialist independence. If your specialists are genuinely independent (different information sources, different analytical frames), the plateau comes later. If your specialists are highly correlated (same model, similar prompts, no tool access), two or three specialists captures almost all the accuracy benefit that consensus provides.
Finding the right ensemble size requires empirical testing: run your benchmark questions with 1, 3, 5, and 7 specialists and plot accuracy against count. The inflection point — where additional specialists produce minimal accuracy gain — is your target size. Building beyond that point adds cost without adding quality. Run this calibration once when you first build the consensus system, and again after significant changes to the specialist prompts or the task distribution.
Store your specialist configurations and synthesis agent prompts in PromptABCD. When the next project needs a consensus mechanism on a similar problem type, your tuned agent definitions are the starting point, not the end product.
Continue Reading
Save the prompts from this post
PromptABCD is a free prompt manager. Paste, organize, and reuse your best AI prompts — no more hunting through chat history.
