Building a Multi-Agent System With AutoGen
This AutoGen tutorial covers the setup mistake that trips up 80% of first-time users, plus working code for a real multi-agent pipeline using ConversableAgent and GroupChat.
import autogen
config_list = [{"model": "claude-sonnet-4-6", "api_type": "anthropic", "api_key": "YOUR_KEY"}]
llm_config = {"config_list": config_list, "temperature": 0.1}
# The critical fix: set human_input_mode to "NEVER" for autonomous pipelines
researcher = autogen.ConversableAgent(
name="Researcher",
system_message="You are a research specialist. Find key facts and data about the topic. Be specific and cite your sources where possible.",
llm_config=llm_config,
human_input_mode="NEVER", # Do NOT leave this as ALWAYS
max_consecutive_auto_reply=3
)
writer = autogen.ConversableAgent(
name="Writer",
system_message="You are a technical writer. Transform research findings into clear, well-structured prose. Ask for clarification if research is ambiguous.",
llm_config=llm_config,
human_input_mode="NEVER",
max_consecutive_auto_reply=3
)
# Initiate the conversation
result = researcher.initiate_chat(
writer,
message="Research this topic and then write a 300-word summary: quantum error correction breakthroughs in 2025",
max_turns=4
)
print(result.summary)AutoGen ships with a default configuration that trips up most first-time users: every agent has human_input_mode set to "ALWAYS" by default. This means every agent will pause and wait for a human to type something before proceeding — the exact opposite of what you want in an automated multi-agent pipeline. You run your first autogen tutorial example, watch the terminal wait indefinitely, assume the library is broken, and move on.
It isn't broken. It's just optimized for interactive research demos, not autonomous production pipelines. One configuration change fixes it — and understanding why it exists helps you use it intentionally when you actually want human-in-the-loop behavior.
Quick-Start: AutoGen Multi-Agent Pipeline in 30 Lines
[object Object], autogen
config_list = [{,[object Object],: ,[object Object],, ,[object Object],: ,[object Object],, ,[object Object],: ,[object Object],}]
llm_config = {,[object Object],: config_list, ,[object Object],: ,[object Object],}
,[object Object],
researcher = autogen.ConversableAgent(
name=,[object Object],,
system_message=,[object Object],,
llm_config=llm_config,
human_input_mode=,[object Object],, ,[object Object],
max_consecutive_auto_reply=,[object Object],
)
writer = autogen.ConversableAgent(
name=,[object Object],,
system_message=,[object Object],,
llm_config=llm_config,
human_input_mode=,[object Object],,
max_consecutive_auto_reply=,[object Object],
)
,[object Object],
result = researcher.initiate_chat(
writer,
message=,[object Object],,
max_turns=,[object Object],
)
,[object Object],(result.summary)What this does: Creates two ConversableAgents that communicate directly with each other. The researcher goes first, produces findings, then the writer converts them to prose. human_input_mode="NEVER" ensures the pipeline runs autonomously. max_consecutive_auto_reply prevents infinite loops by capping how many times each agent can respond without external input.
Understanding the AutoGen Variables
AutoGen's architecture is built around ConversableAgent — a base class that handles conversation history, termination conditions, and tool calling. Before wiring agents together, you need to understand four key parameters:
human_input_mode: Controls when agents pause for human input. Options: "ALWAYS" (pauses before every response — good for interactive demos, bad for automation), "NEVER" (fully autonomous — good for production pipelines), "TERMINATE" (pauses only when the conversation would otherwise end — useful for review workflows). Set this intentionally; the default will surprise you.
max_consecutive_auto_reply: The maximum number of replies an agent can generate before requiring human input (with ALWAYS) or terminating (with TERMINATE). This is your primary loop-prevention mechanism. Set it to 3–5 for most workflows; the exact number depends on how many turns your task needs.
is_termination_msg: A function that evaluates each message and returns True if the conversation should end. By default, conversations end when an agent says "TERMINATE". You can override this with a custom function — for example, ending when a JSON response is detected or when a specific keyword appears.
system_message: The agent's role definition. Unlike raw API calls where you set system prompts directly, AutoGen injects system messages into the conversation history. This has implications for token usage in long conversations — system context is repeated in conversation history, which AutoGen manages automatically but which you should account for in cost estimates.
⚡ Pro tip: AutoGen agents accumulate conversation history across turns, which affects token usage and reasoning in long conversations. For pipelines where agents should start each task fresh, create new agent instances per task rather than reusing the same instance. The reset() method clears history, but creating a fresh instance is more predictable.
Step-by-Step: Building a Three-Agent GroupChat Pipeline
For more complex coordination — three or more agents where the routing isn't simple ping-pong between two — AutoGen's GroupChat is the right tool:
[object Object], autogen
,[object Object], json
config_list = [{,[object Object],: ,[object Object],, ,[object Object],: ,[object Object],, ,[object Object],: ,[object Object],}]
llm_config = {,[object Object],: config_list, ,[object Object],: ,[object Object],}
,[object Object],
analyst = autogen.ConversableAgent(
name=,[object Object],,
system_message=,[object Object],,
llm_config=llm_config,
human_input_mode=,[object Object],
)
strategist = autogen.ConversableAgent(
name=,[object Object],,
system_message=,[object Object],,
llm_config=llm_config,
human_input_mode=,[object Object],
)
summarizer = autogen.ConversableAgent(
name=,[object Object],,
system_message=,[object Object],,
llm_config=llm_config,
human_input_mode=,[object Object],
)
,[object Object],
group_chat = autogen.GroupChat(
agents=[analyst, strategist, summarizer],
messages=[],
max_round=,[object Object],,
speaker_selection_method=,[object Object], ,[object Object],
)
manager = autogen.GroupChatManager(
groupchat=group_chat,
llm_config=llm_config
)
,[object Object],
result = analyst.initiate_chat(
manager,
message=,[object Object],
)
,[object Object],
final_messages = group_chat.messages
,[object Object], msg ,[object Object], final_messages:
,[object Object], msg.get(,[object Object],) == ,[object Object],:
,[object Object],(,[object Object],)
,[object Object],What this does: A three-agent GroupChat where the analyst processes data, the strategist builds recommendations, and the summarizer produces the executive output. speaker_selection_method="round_robin" ensures predictable ordering. Using "auto" instead would let an LLM decide who speaks next — more flexible but less deterministic.
⚡ Pro tip: Use speaker_selection_method="round_robin" while building your system, and switch to "auto" only after you've validated that the round-robin order works correctly for your task. Debugging why "auto" selected an unexpected agent is significantly harder than debugging a round-robin that selected the wrong number of turns.
Pro-Level Variations
Tool-calling agents: AutoGen supports agents that can call Python functions. Register tools using the register_for_execution and register_for_llm decorators. A research agent that can call a web search function, or an execution agent that can run Python code, adds significant capability without complex prompt engineering.
Nested chat patterns: AutoGen allows one agent to initiate a sub-conversation with another group of agents and return the summary to the original conversation. This enables hierarchical coordination — a supervisor agent that runs sub-chats for complex subtasks and reports back to the main group.
Custom termination conditions: Instead of ending on "TERMINATE", implement a termination function that checks for specific output quality: has the response reached a minimum length? Does it contain required fields? Is it valid JSON? Custom termination gives you much more control over when conversations end.
State preservation across conversations: For agents that need to remember context from previous sessions, implement a context injection pattern: at the start of each new conversation, inject a summary of relevant prior context into the first message. AutoGen doesn't provide built-in cross-session memory — you manage this externally and inject it at conversation start.
Troubleshooting Common AutoGen Issues
Conversation loops that don't terminate: If agents keep responding to each other beyond max_turns, check your termination signal. Agents must output exactly "TERMINATE" (case-sensitive, on its own in the message) for the default termination to work. If your agents are instructed to "end with TERMINATE" but also produce other content on the same line, termination may not fire. Use a custom is_termination_msg function for more reliable detection.
GroupChatManager selecting wrong speaker: When speaker_selection_method="auto", the manager uses an LLM to select the next speaker based on conversation context. If it's selecting unexpectedly, check your agent system prompts — the manager reads them to understand each agent's role and infers who should speak next. Vague role descriptions produce unpredictable selection.
Token costs higher than expected: AutoGen agents maintain full conversation history, which grows with each turn. For long group chats, the later agents' turns include the full history as context, which can make token costs grow quadratically. Set a max_round limit appropriate for your task and consider whether all agents need the full history or just their relevant portion.
⚠️ Common mistake: Not setting max_consecutive_auto_reply on agents in a GroupChat. Without this limit, an agent that's determined to be helpful can respond multiple times in a row before yielding to others, throwing off the conversation dynamics and burning tokens on repeated variations of the same response.
Your Turn
Start with the two-agent example from the Quick-Start and modify the system prompts for your specific use case. Run it on 10 representative inputs. Check whether the researcher's findings are actually useful to the writer, and whether max_turns=4 is the right number for your task — too few and the task isn't complete; too many and agents start padding.
Configuring AutoGen for Production
Moving an AutoGen system from notebook prototype to production requires three configuration changes that documentation often skips.
LLM cache management. By default, AutoGen caches LLM responses to identical inputs. In development this speeds up iteration; in production it can cause stale responses. For systems where inputs change frequently, disable the cache or set a short TTL:
[object Object], autogen
,[object Object],
llm_config = {
,[object Object],: [{,[object Object],: ,[object Object],, ,[object Object],: ,[object Object],}],
,[object Object],: ,[object Object], ,[object Object],
}What this does: Setting cache_seed to None disables response caching. Each LLM call hits the API fresh. For production systems processing unique user inputs, this ensures responses are always generated from the current input rather than retrieved from a cache that may no longer be relevant.
⚡ Pro tip: For AutoGen pipelines that process user-specific content — where each run should produce fresh output based on the current input — always set cache_seed=None. The default caching behavior is appropriate for development but produces confusing results in production when the same question from different users returns identical cached responses from a prior run.
Termination conditions. Without explicit termination, AutoGen conversations can run indefinitely. Define clear termination conditions on every agent:
[object Object], ,[object Object],(,[object Object],) -> ,[object Object],:
content = msg.get(,[object Object],, ,[object Object],)
,[object Object], ,[object Object], ,[object Object], content ,[object Object], ,[object Object], ,[object Object], content
agent = autogen.ConversableAgent(
name=,[object Object],,
system_message=,[object Object],,
is_termination_msg=is_termination_msg,
max_consecutive_auto_reply=,[object Object],,
llm_config=llm_config,
human_input_mode=,[object Object],
)What this does: The termination function checks each message for a completion signal. The agent's system prompt instructs it to include TASK_COMPLETE when done. max_consecutive_auto_reply=5 acts as a safety ceiling — the conversation ends after five consecutive auto-replies even if the termination signal hasn't appeared. Two termination mechanisms are better than one.
Structured output extraction. AutoGen conversations produce conversation history, not clean structured output. For systems that need to parse the final agent output programmatically, extract the last relevant message explicitly:
[object Object], ,[object Object],(,[object Object],) -> ,[object Object],:
,[object Object],
messages = chat_result.chat_history
,[object Object], msg ,[object Object], ,[object Object],(messages):
,[object Object], msg.get(,[object Object],) == ,[object Object], ,[object Object], ,[object Object], ,[object Object], ,[object Object], msg.get(,[object Object],, ,[object Object],):
,[object Object], msg[,[object Object],]
,[object Object], ,[object Object],What this does: Iterates the conversation history in reverse and returns the last substantive assistant message, skipping the termination signal itself. This gives you the final output as a clean string that downstream code can process without parsing the full conversation history.
When your agent system prompts are working well, save them to PromptABCD. AutoGen pipelines are only as good as their individual agent prompts, and well-tuned prompts are worth preserving across project iterations.
What Makes AutoGen Different
The core architectural difference between AutoGen and other frameworks is the conversation model. In AutoGen, agents are fundamentally conversational entities — they speak, receive replies, and respond again. The GroupChat manager decides who speaks next. This model is natural for tasks that benefit from back-and-forth deliberation between agents but less natural for tasks that are sequential transformations with no deliberation needed.
When to choose AutoGen over a simpler approach: your task genuinely involves agents needing to negotiate, challenge each other's reasoning, or iterate based on each other's feedback before reaching a conclusion. When agents have nothing to say to each other until they're done — one produces output, the next processes it — the conversation model adds overhead without benefit. A sequential API call chain is simpler.
The autogen tutorial pattern above covers the most common production use case: a GroupChat where each agent has a specialized role, a strict maximum turn limit, and human_input_mode="NEVER". That configuration handles the vast majority of AutoGen production deployments.
Continue Reading
Save the prompts from this post
PromptABCD is a free prompt manager. Paste, organize, and reuse your best AI prompts — no more hunting through chat history.
