PromptABCD
FeaturesLearnGuideBlogContext Blocks
Sign inGet started free
Sign inSign up
PromptABCD

A calm home for your best AI prompts. Save them once, find them in seconds, reuse them forever.

Product

  • Features
  • Chrome Extension
  • Free Courses
  • How it works
  • Use cases
  • Blog
  • Context Blocks
  • Export Anywhere
  • FAQ

Resources

  • User guide
  • Learn prompting
  • Sign in
  • Get started free

© 2026 PromptABCD. All rights reserved.

AboutPrivacy PolicyTerms and Conditions
Home/Blog/Multi-Agent Systems/Building a Multi-Agent System With MCP
Multi-Agent Systems

Building a Multi-Agent System With MCP

MCP passed 97 million monthly SDK downloads by early 2026, and most teams still wire it wrong for multi-agent work. This mcp multi agent teardown shows the fix.

September 25, 2026·9 min read
ShareShare
⚡Featured Prompt— copy and use right now
# Each agent pins a stateful session to each server - the 2025 default
class Agent:
    def __init__(self, server_urls):
        self.sessions = [open_stateful_session(u) for u in server_urls]
        # session holds server-side state tied to THIS agent instance

    def call_tool(self, name, args):
        return self.sessions[route(name)].invoke(name, args)

By early 2026, the Model Context Protocol had passed 97 million monthly SDK downloads and become the default way agents connect to tools - and yet most teams building an mcp multi agent system wire it in a way that quietly caps how far they can scale. They treat every MCP server as a permanent, stateful attachment to one agent, and then wonder why the system won't spread across machines cleanly. The July 2026 spec revision changed the rules here, and if you learned MCP in 2025 your mental model is now slightly wrong.

Let me tear down a typical setup, show exactly where the coupling hides, and rebuild it against the current stateless-core design. The fix is less code than you'd expect and it unlocks the horizontal scaling people assume MCP can't do.

Before: The Tightly-Coupled Setup

Here's a common mcp multi agent arrangement. Each agent owns a long-lived, stateful connection to its MCP servers, and the servers hold session state per client.

python
[object Object],
,[object Object], ,[object Object],:
    ,[object Object], ,[object Object],(,[object Object],):
        ,[object Object],.sessions = [open_stateful_session(u) ,[object Object], u ,[object Object], server_urls]
        ,[object Object],

    ,[object Object], ,[object Object],(,[object Object],):
        ,[object Object], ,[object Object],.sessions[route(name)].invoke(name, args)

What this does: It gives each agent a set of persistent, stateful sessions to its MCP servers. Every tool call rides an open session that the server keeps alive with per-client state - which works fine for one agent on one machine and becomes the exact thing that blocks you from running many agent instances behind a load balancer.

Why It Fails

The problem is that stateful sessions pin an agent instance to a server instance. When you want to scale - run ten copies of an agent across five machines - each copy needs its own sticky session, and the server has to hold state for all of them. Now you need sticky-session routing at the gateway, a shared session store so a reconnecting agent finds its state, and careful handling when a server instance dies and takes live sessions with it. You've imported all the operational pain of stateful services into what should be a stateless tool layer.

There's a second, subtler failure. In a multi-agent setup, several agents often need the same MCP server (a shared database tool, a shared search tool). With per-agent stateful sessions, you can't freely pool or share those server connections, because each carries agent-specific state. So you either over-provision servers or build a connection-sharing layer the protocol was fighting you on.

⚠️ Common mistake: Treating an MCP server like a stateful microservice that each agent must hold open. That model made sense for a single desktop client in 2025, but for a fleet of agents it forces sticky routing and shared session stores you don't need. The whole point of the 2026 revision was to let servers run behind a plain round-robin load balancer - if your design still needs stickiness, you're fighting the current protocol instead of using it.

After: The Stateless-Core Rebuild

The July 2026 MCP spec made the core stateless. Requests carry the state they need in the payload, so any server instance can handle any request. Multi-round-trip interactions - the ones that used to require holding a stream open, like an elicitation asking the user to confirm - now work by the server returning a result that echoes an opaque state token, which the client sends back on the follow-up. No sticky session required.

python
[object Object],
,[object Object], ,[object Object],:
    ,[object Object], ,[object Object],(,[object Object],):
        ,[object Object],.pool = server_pool   ,[object Object],

    ,[object Object], ,[object Object],(,[object Object],):
        payload = {,[object Object],: name, ,[object Object],: args}
        ,[object Object], resume_state:                 ,[object Object],
            payload[,[object Object],] = resume_state
        ,[object Object], ,[object Object],.pool.post(name, payload)   ,[object Object],

What this does: It sends each tool call to a pool of interchangeable server instances instead of a pinned session. When an interaction needs a follow-up (confirm a deletion, supply more input), the agent just echoes the opaque state token the server returned - so the retry can land on any server instance, because everything needed is in the payload. Stickiness disappears.

Breaking Down Each Element

Three changes carry the rebuild, and each maps to a scaling win.

The pool replaces pinned sessions. Because the core is stateless, agents talk to a load-balanced pool of server instances rather than holding one open. Any instance can serve any request, so you scale servers by adding instances behind the balancer - ordinary horizontal scaling, no session affinity.

The echoed state token replaces the held-open stream. Interactions that need multiple round-trips used to keep a connection alive; now the server hands back an opaque token summarizing where the interaction is, and the client returns it on the next call. This is the single change that made statelessness possible for interactive tools, and it's the part most 2025-era code doesn't do.

Method-based routing replaces per-agent stickiness. The 2026 transport lets gateways route on a method header and lets clients cache the server's tool list for as long as the server permits. So a shared tool server can be freely pooled across many agents - the exact multi-agent sharing that stateful sessions blocked.

python
[object Object],
,[object Object], ,[object Object],(,[object Object],):
    cached = ttl_cache.get(,[object Object],)
    ,[object Object], cached:
        ,[object Object], cached
    tools, ttl = pool.list_tools()   ,[object Object],
    ttl_cache.,[object Object],(,[object Object],, tools, ttl)
    ,[object Object], tools

What this does: It caches the server's advertised tool list for the duration the server allows, so agents don't re-discover tools on every call. Across a fleet of agents hitting shared servers, this cuts a large volume of redundant discovery traffic - and because the cache respects the server's stated lifetime, it stays correct when tools change.

⚡ Pro tip: If you built on MCP in 2025, audit your servers for anything that assumes a sticky session - in-memory per-client state, an open stream awaiting a reply, a session ID that must survive reconnects. Each of those is now a scaling ceiling you can remove by moving that state into the request payload. The migration is mechanical once you find the stateful bits.

Variations for Different Contexts

For local single-agent setups (a coding assistant on your laptop talking to local tools over stdio), the stateful model is still fine - you have one client and one server, and stickiness costs nothing. Don't rearchitect a local tool for statelessness it doesn't need. The stateless core matters when you scale horizontally, not when you run one of everything.

For large fleets where many agents share many servers, lean fully into the stateless design: pool every shared server, cache tool lists per TTL, and route on method headers. This is where the 2026 revision pays off most - you get the tool-integration benefits of MCP with the operational simplicity of stateless HTTP services, which is a combination that genuinely wasn't available a year earlier.

For agent-to-agent coordination specifically, watch the A2A direction maturing alongside MCP in 2026 - MCP standardizes agent-to-tool, and the emerging agent-to-agent layer sits on top for cross-agent orchestration. Keeping those two concerns separate in your architecture (tools via MCP, coordination via a dedicated layer) ages better than cramming inter-agent messaging into tool calls.

⚡ Pro tip: Keep a hard line between "an agent using a tool" (MCP's job) and "an agent talking to another agent" (coordination's job). Teams that route agent-to-agent messages through MCP tool calls end up with a tangled layer that's hard to observe and secure. MCP is the tool interface; put your inter-agent messaging in the event bus or queue where it belongs.

How Do You Migrate a Stateful MCP Setup Without Downtime?

If you already run a stateful mcp multi agent system, you can't flip to stateless overnight - you have live agents depending on the old behavior. The migration that works is running both models side by side and moving traffic gradually, the same strangler pattern you'd use for any protocol upgrade.

Start by making your servers able to handle stateless requests without removing statefulness. A server that accepts an echoed state token in the payload but also still honors an open session can serve both kinds of client. Then migrate clients one at a time to the stateless calling convention, watching for correctness regressions on the multi-round-trip interactions, which are where subtle bugs hide.

python
[object Object],
,[object Object], ,[object Object],(,[object Object],):
    ,[object Object], ,[object Object], ,[object Object], request:          ,[object Object],
        state = decode(request[,[object Object],])
    ,[object Object], request.session_id ,[object Object], sessions:   ,[object Object],
        state = sessions[request.session_id]
    ,[object Object],:
        state = fresh_state()
    ,[object Object], process(request, state)

What this does: It lets one server serve both stateless and stateful clients during the transition by checking for an echoed state token first and falling back to a legacy session. You migrate clients on your own schedule instead of coordinating a risky big-bang cutover, and you can roll a client back instantly if it misbehaves.

The interactions to watch most carefully are the ones that ask the user something mid-call - confirmations, follow-up inputs. Those are where the old held-open-stream model and the new echoed-token model differ most, so test them first and hardest. Simple one-shot tool calls migrate almost trivially by comparison; the round-trips are the real work.

⚡ Pro tip: Migrate your simplest one-shot tools to stateless first to build confidence, then tackle the interactive multi-round-trip tools last. Doing it in that order means the tricky migrations happen when your team already understands the new model from the easy ones, rather than learning the pattern on the hardest case.

Save and Reuse This

The reusable core is the stateless pattern: pool interchangeable server instances, carry interaction state in the payload via echoed tokens, and cache tool discovery per the server's TTL. That combination is what lets an mcp multi agent system scale horizontally instead of hitting a sticky-session wall.

Keep your MCP client configuration and the tool-discovery-caching logic versioned so you can reproduce a known-good stateless setup. I store these patterns in PromptABCD alongside the agent prompts that use the tools, because getting the stateless client wiring right took real study of the 2026 spec, and reusing the exact configuration beats re-deriving it - protocols evolve, and you want to carry forward the version you've verified works under scale.

⚡ Pro tip: Pin and record the exact MCP spec version your system targets, and re-audit when it changes. The gap between the 2025 stateful model and the 2026 stateless core is large enough that code written against one behaves subtly wrong against the other. Treat the protocol version like a dependency version - explicit, recorded, and reviewed on upgrade - so a spec revision is a deliberate migration rather than a surprise.

multi-agent-systemsmcpmodel-context-protocolarchitectureai-agents

Continue Reading

A Reusable Prompt Kit for Agent Teams
Multi-Agent Systems

A Reusable Prompt Kit for Agent Teams

A team rebuilt their agent prompts from memory every project, and every project drifted a little worse. That failure is why a multi agent prompt kit matters. Here's the reusable set of role prompts every team should keep.

October 2, 2026·9 min read
Multi-Agent Systems on a Budget: An Interactive Guide
Multi-Agent Systems

Multi-Agent Systems on a Budget: An Interactive Guide

Most multi-agent tutorials assume you'll burn tokens freely. That's wrong for anyone shipping on real constraints. A cheap multi agent system can match an expensive one with the right moves. Here's how to build one.

October 2, 2026·9 min read
The Manager Agent Anti-Pattern: A Teardown
Multi-Agent Systems

The Manager Agent Anti-Pattern: A Teardown

Why does your orchestrator agent keep becoming a bottleneck that mangles every handoff? You've hit the manager agent anti pattern — one agent trying to coordinate everything. Here's why it fails and what replaces it.

October 2, 2026·9 min read

Save the prompts from this post

PromptABCD is a free prompt manager. Paste, organize, and reuse your best AI prompts — no more hunting through chat history.

Start free →
← PreviousPub/Sub Patterns for AI AgentsNext →Agent Registries and Discovery
Share this post:
ShareShare