PromptABCD
FeaturesLearnGuideBlogContext Blocks
Sign inGet started free
Sign inSign up
PromptABCD

A calm home for your best AI prompts. Save them once, find them in seconds, reuse them forever.

Product

  • Features
  • Chrome Extension
  • Free Courses
  • How it works
  • Use cases
  • Blog
  • Context Blocks
  • Export Anywhere
  • FAQ

Resources

  • User guide
  • Learn prompting
  • Sign in
  • Get started free

© 2026 PromptABCD. All rights reserved.

AboutPrivacy PolicyTerms and Conditions
Home/Blog/Multi-Agent Systems/Pub/Sub Patterns for AI Agents
Multi-Agent Systems

Pub/Sub Patterns for AI Agents

A team's agent added a new step and silently broke three downstream agents nobody remembered depended on it. Pub sub ai agents patterns prevent exactly this. Here's the case study.

September 24, 2026·8 min read
ShareShare
⚡Featured Prompt— copy and use right now
# The wiki approach - a doc nobody updates
# scorer expects: {enriched_score: float, region: str}
# router expects: {score: float 0-100, priority: str}

Here's the failure that made a team take pub/sub seriously. An engineer added a data-cleaning step to an agent pipeline - a small, sensible improvement. It shipped. And within an hour, three downstream agents started producing garbage, because they'd been quietly depending on the uncleaned data format in ways nobody had documented. The change was correct in isolation and catastrophic in context, because the dependencies between agents were invisible. That invisibility is exactly what pub sub ai agents patterns exist to fix.

Publish-subscribe isn't just a messaging trick. Applied to agents, it makes dependencies explicit and turns "adding a step secretly breaks things" into "adding a step is safe by default." Here's how that team got there, and the specific patterns that did the work.

The Problem This Team Faced

The team ran a data-enrichment pipeline: agents that ingested records, cleaned them, enriched them with external data, scored them, and routed them. The agents were wired together with direct calls in a fixed order. Each agent called the next by name.

The hidden problem was that agents made undocumented assumptions about each other's output. The scoring agent assumed the enrichment agent's exact field names. The routing agent assumed the scorer's exact score range. None of this was written down; it was baked into prompts and parsing code. So any change to one agent's output could break any agent downstream, and there was no way to know which without running everything.

When they inserted the cleaning agent, it changed a field format the scorer depended on. The scorer's parsing silently produced nulls, the router treated nulls as "low priority," and high-priority records got dropped. It took hours to trace, because nothing errored - the dependency was invisible and the failure was silent.

The Wrong Approach

Their first instinct was documentation: write down every agent's input and output contract in a wiki so people would know what depended on what.

python
[object Object],
,[object Object],
,[object Object],

What this does: It records the implicit contracts in prose so engineers can check dependencies before changing an agent. It also rots immediately - documentation drifts from code the moment someone changes an output without updating the wiki, which is always. A contract that isn't enforced is a contract that's already wrong.

Documentation describes dependencies; it doesn't enforce them, and it doesn't make them visible at runtime. The team needed the dependencies to be structural - part of how the system worked - not a description sitting in a doc that goes stale. Prose contracts fail for the same reason comments fail: nothing checks them.

⚡ Pro tip: If a dependency matters, encode it in the system, not in a doc. Documentation is where dependencies go to become lies. The pub/sub approach that follows works precisely because the dependencies become executable subscriptions rather than described intentions.

The Correct Pattern

The fix was pub/sub with typed topics and explicit subscriptions. Agents publish to named topics with a declared schema, and consumers subscribe to topics. The subscription is the dependency - visible, enumerable, and checkable.

python
[object Object], ,[object Object],:
    ,[object Object], ,[object Object],(,[object Object],):
        ,[object Object],._topics = {}      ,[object Object],
        ,[object Object],._subs = defaultdict(,[object Object],)

    ,[object Object], ,[object Object],(,[object Object],):
        ,[object Object],._topics[topic] = schema

    ,[object Object], ,[object Object],(,[object Object],):
        ,[object Object],._subs[topic].append(agent)

    ,[object Object], ,[object Object],(,[object Object],):
        validate(payload, ,[object Object],._topics[topic])   ,[object Object],
        ,[object Object], agent ,[object Object], ,[object Object],._subs[topic]:
            agent.deliver(topic, payload)

What this does: It makes every publish validate against the topic's declared schema, so an agent can't silently change its output shape - a schema violation errors at publish time. And because subscriptions are registered explicitly, you can list exactly which agents depend on any topic, turning invisible dependencies into a queryable graph.

The critical addition was schema validation at publish. Now when the cleaning agent changed a field, the publish either conformed to the declared schema or failed loudly at the source - not silently three agents downstream.

python
[object Object], ,[object Object],(,[object Object],):
    ,[object Object], [agent.name ,[object Object], agent ,[object Object], bus._subs[topic]]
,[object Object],

What this does: It answers "what breaks if I change this topic?" by listing every subscriber. Before touching any agent's output, an engineer runs this and sees the exact set of downstream agents affected - the blast radius that was previously invisible.

Results and What Changed

After the migration, the class of failure that started this story became impossible. Changing an agent's output either matched the declared schema or failed at publish, immediately and at the source. Adding a new agent meant declaring a subscription - a visible, reviewable act - rather than a silent assumption. And who_depends_on gave engineers the blast radius of any change before they made it.

The unexpected benefit was onboarding speed. New engineers could understand the system by reading the topic declarations and subscriptions - the entire dependency structure was enumerable in one place. Previously, understanding the pipeline meant tracing direct calls through five agents' worth of code and prompts. The pub/sub topic map became the documentation that the wiki was always supposed to be, except it couldn't go stale because it was the system.

⚡ Pro tip: Version your topic schemas, and when you need a breaking change, publish to a new topic version rather than mutating the old one. Consumers migrate on their own schedule, and you never get the "one change breaks everyone at once" failure. This is the single practice that lets a pub/sub agent system evolve safely over time.

When Does Pub/Sub Beat Direct Calls for Agents?

Pub/sub isn't automatically better - it trades directness for flexibility, and that trade is worth it only under specific conditions. Knowing when pub sub ai agents patterns pay off keeps you from adding a bus to a system that didn't need one.

The clearest signal is fan-out: one event that multiple agents should react to independently. A record gets scored, and now the router wants it, the audit-logger wants it, and the alerting agent wants it. With direct calls, the scorer must know about all three and call each - and adding a fourth consumer means editing the scorer. With pub/sub, the scorer publishes once and the three (or four, or ten) consumers self-select. Fan-out is where pub/sub earns its keep, because it's exactly the case direct calls handle worst.

python
[object Object],
bus.publish(,[object Object],, {,[object Object],: ref, ,[object Object],: ,[object Object],})
,[object Object],

What this does: It shows a single publish triggering multiple independent consumers. The scorer's code is identical whether one agent or ten react to the score - adding a consumer never touches the producer, which is the property that makes fan-out-heavy systems dramatically easier to grow.

The second signal is volatility - if your set of agents changes often, pub/sub's decoupling means each change is local. The third is independent evolution - when different teams own different agents, schemas-and-topics give them a contract to evolve against without coordinating every change through a meeting.

Conversely, for a fixed, linear, three-agent pipeline that never changes, direct calls are simpler and you should use them. The indirection of a bus is a cost you only recoup when you actually exercise the flexibility it buys. Adding pub/sub to a system with no fan-out and no churn is architecture cosplay - it looks sophisticated and buys you nothing but a harder-to-trace call graph.

⚡ Pro tip: Count the fan-out in your system - how many events have more than one interested consumer. If the answer is "almost none," you probably don't need pub/sub yet. If several events fan out to three or more consumers, and that number is growing, pub/sub will pay for itself quickly. Let the fan-out count, not fashion, drive the decision.

How to Apply This to Your Situation

Start by finding your invisible dependencies. For each agent, ask what it assumes about its inputs' exact shape, and write those assumptions as schemas. You'll likely uncover assumptions nobody knew existed - that discovery alone is worth the exercise.

Then introduce a typed bus for your riskiest handoff - the one where a change would cause the quietest, most expensive failure. Declare the topic schema, validate on publish, and register subscriptions explicitly. Expand from there.

The mindset shift with pub sub ai agents is that dependencies should be declared and enforced, not assumed and documented. Once dependencies live in subscriptions and schemas rather than in tribal knowledge, the whole system gets safer to change - which is what lets an agent team grow past the size where one person can hold it all in their head.

⚠️ Common mistake: Adopting pub/sub topic names but skipping schema validation, so you get the decoupling but not the safety. Topics without enforced schemas still let an agent silently change its output shape - you've renamed the problem, not solved it. The schema enforcement at publish is the part that actually prevents the silent-breakage failure; the topic naming alone is cosmetic.

Next Steps

Map one agent's input assumptions into a schema, stand up a typed topic for its most dangerous handoff, and enforce validation on publish. Add who_depends_on so you can see blast radius before every change.

Keep your topic schemas and the agent subscription declarations versioned so the dependency structure travels with the code. I store the topic definitions and subscription patterns in PromptABCD, because the schemas encode hard-won knowledge about what each agent actually requires, and reusing a proven set of topic contracts saves you from rediscovering the same silent dependencies in the next pub/sub agent system you build.

⚡ Pro tip: Generate your dependency diagram automatically from the live subscription registry, and put it in front of engineers during code review. When someone changes an agent's output, the diff to the dependency graph shows up right there - the reviewer sees exactly which subscribers are affected without anyone having to remember. Auto-generated from the running system, this diagram can't drift the way a hand-drawn one always does.

multi-agent-systemspub-subdecouplingevent-drivenai-agents

Continue Reading

A Reusable Prompt Kit for Agent Teams
Multi-Agent Systems

A Reusable Prompt Kit for Agent Teams

A team rebuilt their agent prompts from memory every project, and every project drifted a little worse. That failure is why a multi agent prompt kit matters. Here's the reusable set of role prompts every team should keep.

October 2, 2026·9 min read
Multi-Agent Systems on a Budget: An Interactive Guide
Multi-Agent Systems

Multi-Agent Systems on a Budget: An Interactive Guide

Most multi-agent tutorials assume you'll burn tokens freely. That's wrong for anyone shipping on real constraints. A cheap multi agent system can match an expensive one with the right moves. Here's how to build one.

October 2, 2026·9 min read
The Manager Agent Anti-Pattern: A Teardown
Multi-Agent Systems

The Manager Agent Anti-Pattern: A Teardown

Why does your orchestrator agent keep becoming a bottleneck that mangles every handoff? You've hit the manager agent anti pattern — one agent trying to coordinate everything. Here's why it fails and what replaces it.

October 2, 2026·9 min read

Save the prompts from this post

PromptABCD is a free prompt manager. Paste, organize, and reuse your best AI prompts — no more hunting through chat history.

Start free →
← PreviousUsing a Message Queue for Agent CoordinationNext →Building a Multi-Agent System With MCP
Share this post:
ShareShare