Shared Memory in Multi-Agent Systems
Multi-agent shared memory is where most agent architectures get lazy — and where the most important bugs hide. Here's how to design memory that agents can actually rely on.
import json
import time
from typing import Optional, Any
from anthropic import Anthropic
client = Anthropic()
class SharedMemory:
"""In-memory shared state for multi-agent coordination."""
def __init__(self, task_id: str, original_request: str):
self._store = {
"task_id": task_id,
"original_request": original_request,
"created_at": time.time(),
"agent_outputs": {},
"status": "pending",
"metadata": {}
}
self._write_log = []
def write(self, agent_name: str, key: str, value: Any) -> None:
"""Write a value with audit logging."""
full_key = f"{agent_name}.{key}"
self._store[full_key] = value
self._write_log.append({
"agent": agent_name,
"key": full_key,
"written_at": time.time(),
"value_type": type(value).__name__
})
def read(self, key: str, default: Any = None) -> Any:
"""Read a value by key."""
return self._store.get(key, default)
def get_original_request(self) -> str:
return self._store["original_request"]
def get_all_agent_results(self) -> dict:
"""Retrieve all results written by agents."""
return {k: v for k, v in self._store.items()
if "." in k and not k.startswith("_")}Most tutorials on multi-agent systems get the memory question backwards. They treat shared memory as a convenience feature — a place to store things so agents don't have to re-derive them. That's not what multi agent shared memory is.
Shared memory is the coordination primitive. It's how agents that have never directly communicated still work toward a coherent goal. It's how a reviewer agent knows what a researcher agent found two steps ago. It's how a supervisor agent knows whether workers have completed their subtasks. Without well-designed shared memory, you don't have a multi-agent system — you have multiple independent agents that happen to run in sequence.
Getting shared memory right is one of the hardest parts of building reliable multi-agent systems, and it's the one that gets the least attention.
What is Multi-Agent Shared Memory?
Shared memory in a multi-agent system is any persistent data store that multiple agents can read from and write to during the course of a task. "Persistent" here means persistent across agent calls — not necessarily across system restarts, though production systems need that too.
There are four levels of shared memory, each appropriate for different use cases:
In-memory dictionary: The simplest option. A Python dict that lives for the duration of one pipeline run. Zero setup, fast, not durable. Good for development and for pipelines that run synchronously and complete in a single process.
Shared database (relational or document): SQL or NoSQL storage that persists across process restarts, supports concurrent reads and writes from multiple agents, and enables audit logging. The right choice for production systems where tasks span multiple processes or need durability.
Vector store: Specialized memory for semantic retrieval. Agents store results as embeddings; other agents retrieve them by semantic similarity rather than exact key lookup. Necessary when agents need to find relevant prior work without knowing its exact identifier.
Hybrid memory: A combination of fast in-memory state for current task coordination and persistent storage for long-term agent knowledge. Most sophisticated production systems use this approach.
[object Object], json
,[object Object], time
,[object Object], typing ,[object Object], ,[object Object],, ,[object Object],
,[object Object], anthropic ,[object Object], Anthropic
client = Anthropic()
,[object Object], ,[object Object],:
,[object Object],
,[object Object], ,[object Object],(,[object Object],):
,[object Object],._store = {
,[object Object],: task_id,
,[object Object],: original_request,
,[object Object],: time.time(),
,[object Object],: {},
,[object Object],: ,[object Object],,
,[object Object],: {}
}
,[object Object],._write_log = []
,[object Object], ,[object Object],(,[object Object],) -> ,[object Object],:
,[object Object],
full_key = ,[object Object],
,[object Object],._store[full_key] = value
,[object Object],._write_log.append({
,[object Object],: agent_name,
,[object Object],: full_key,
,[object Object],: time.time(),
,[object Object],: ,[object Object],(value).__name__
})
,[object Object], ,[object Object],(,[object Object],) -> ,[object Object],:
,[object Object],
,[object Object], ,[object Object],._store.get(key, default)
,[object Object], ,[object Object],(,[object Object],) -> ,[object Object],:
,[object Object], ,[object Object],._store[,[object Object],]
,[object Object], ,[object Object],(,[object Object],) -> ,[object Object],:
,[object Object],
,[object Object], {k: v ,[object Object], k, v ,[object Object], ,[object Object],._store.items()
,[object Object], ,[object Object], ,[object Object], k ,[object Object], ,[object Object], k.startswith(,[object Object],)}What this does: The SharedMemory class wraps a Python dictionary with audit logging — every write is tracked with the agent name, key, timestamp, and value type. This makes debugging straightforward: when an agent reads a wrong value, the write log shows which agent wrote it and when.
Why Shared Memory Design Matters
The most common failure in multi-agent shared memory is write collision: two agents writing to the same key and one overwriting the other's work.
This happens more often than you'd expect. Parallel agents running simultaneously might both try to update a task status field, or both try to log their findings to a "results" key. Without write coordination, the last writer wins and the first writer's output is silently lost.
The fix isn't complicated: namespace your keys by agent name. Instead of memory.write("results", output), use memory.write("research_agent", "results", output) — which stores the value at key research_agent.results. Every agent owns its own namespace in shared memory, and only the orchestrator reads across namespaces to synthesize final output.
⚡ Pro tip: Design your shared memory schema before you write your first agent. List every key that will be read and written, which agents will read and write each key, and what the expected data type is. This schema is as important as your agent system prompts — it defines the contract between agents. Changes to the schema mid-project require updating multiple agents simultaneously.
The Core Components of Shared Memory Architecture
Key namespacing: Each agent writes to its own namespace (agent_name.key_name) to prevent collisions. The orchestrator has read access to all namespaces. Agents have write access to their own namespace and read access to any namespace specified in their instructions.
Write-through vs write-back: Write-through means every agent write is immediately persisted to the backing store. Write-back means agents write to fast in-memory state first and sync to persistent storage periodically or at task completion. Write-through is safer (no data loss if the process crashes mid-task) but slower. Write-back is faster but requires careful synchronization logic.
Read consistency: In a parallel multi-agent system, Agent B might read from shared memory while Agent A is mid-write. Define your consistency requirements upfront: is eventual consistency acceptable (Agent B might read slightly stale data) or does your system require strong consistency (Agent B always reads the latest data, even if that means blocking)?
Memory TTL: Shared memory grows over time. Define time-to-live (TTL) policies for different types of data. Task-specific intermediate results might expire after 24 hours. Agent learned preferences might persist indefinitely. Without TTL policies, memory stores grow without bound and become expensive to query.
[object Object], ,[object Object],(,[object Object],) -> ,[object Object],:
,[object Object],
original_request = memory.get_original_request()
response = client.messages.create(
model=,[object Object],,
max_tokens=,[object Object],,
system=,[object Object],,
messages=[{
,[object Object],: ,[object Object],,
,[object Object],: ,[object Object],
}]
)
result = json.loads(response.content[,[object Object],].text) ,[object Object], response.content[,[object Object],].text.strip().startswith(,[object Object],) ,[object Object], {,[object Object],: response.content[,[object Object],].text}
,[object Object],
memory.write(,[object Object],, ,[object Object],, result)
memory.write(,[object Object],, ,[object Object],, topic)
memory.write(,[object Object],, ,[object Object],, time.time())
,[object Object], json.dumps(result)
,[object Object], ,[object Object],(,[object Object],) -> ,[object Object],:
,[object Object],
research_findings = memory.read(,[object Object],, {})
original_request = memory.get_original_request()
,[object Object], ,[object Object], research_findings:
,[object Object], json.dumps({,[object Object],: ,[object Object],})
response = client.messages.create(
model=,[object Object],,
max_tokens=,[object Object],,
system=,[object Object],,
messages=[{
,[object Object],: ,[object Object],,
,[object Object],: ,[object Object],
}]
)
result = json.loads(response.content[,[object Object],].text) ,[object Object], response.content[,[object Object],].text.strip().startswith(,[object Object],) ,[object Object], {,[object Object],: response.content[,[object Object],].text}
memory.write(,[object Object],, ,[object Object],, result)
memory.write(,[object Object],, ,[object Object],, audience)
memory.write(,[object Object],, ,[object Object],, time.time())
,[object Object], json.dumps(result)What this does: Each agent reads from the shared memory using namespaced keys and writes to its own namespace. The writer agent reads researcher.findings directly from memory rather than from the researcher's return value — the two agents are decoupled at the data level. You can change the researcher's implementation without changing how the writer reads its output, as long as both respect the same memory schema.
How Multi-Agent Shared Memory Works in Practice
Three patterns appear most often in production systems:
Task workspace: Each task gets its own memory instance, scoped to the duration of that task. Agents write to the task's workspace, and the orchestrator reads the final state to compose the response. Task workspaces are isolated — no cross-contamination between concurrent tasks.
Persistent agent memory: Each agent maintains a long-term memory store of what it has learned — useful facts, user preferences, successful patterns. This memory persists across tasks and is consulted at the start of each new task. Implementation requires a vector store for semantic retrieval or a key-value store for exact lookups.
Shared knowledge base: All agents can read from a shared knowledge store populated with domain information — product documentation, policy documents, customer profiles. This is write-once, read-many — unlike task-specific memory, the knowledge base is not modified during task execution.
⚡ Pro tip: Start with task workspace memory only. Don't implement persistent agent memory or shared knowledge bases until you have specific evidence that your system would benefit from them. Task workspace memory solves 80% of multi-agent coordination problems, is easy to reason about, and requires no complex storage infrastructure. Add persistence when you find yourself re-deriving the same information across tasks.
Common Mistakes with Multi-Agent Shared Memory
⚠️ Common mistake: Treating shared memory as an append-only log that grows indefinitely. Agents that write to memory without reading what's already there, or without defined retention policies, create memory stores that become expensive to query and contain contradictory information. Define who can write to each namespace, what they're expected to write, and when old values can be discarded.
Three other patterns that cause consistent problems:
Reading memory before it's written: If Agent B runs before Agent A writes its results, Agent B reads an empty or stale value. Design your pipeline so agents only read from namespaces that are guaranteed to be populated by the time they run. Make dependencies explicit in your orchestration logic.
Memory schema drift: The memory schema that works for your first five agents becomes inconsistent as you add agents that write data in slightly different formats. One agent writes {"findings": [...]}, another writes {"research_findings": [...]}. Downstream agents that read memory need to handle both formats or break. Enforce schema consistency with validation at write time.
No memory visibility during debugging: When your multi-agent system produces a wrong answer, you need to inspect shared memory at each step to find where the error was introduced. Add a memory snapshot capability — a way to read the full state of shared memory at any point in the pipeline — before you need it.
Getting Started with Shared Memory
Multi agent shared memory done right is a force multiplier: agents can build on each other's work without tight coupling, and the orchestrator has a complete audit trail of what every agent produced.
The minimal viable approach is a task-scoped dictionary with namespaced keys, write logging, and explicit read-before-write dependency tracking. That's four hours of implementation work that prevents months of debugging.
Cache Invalidation in Agent Memory
When shared memory values become stale — when an agent writes an update that supersedes a prior value — downstream agents need to either read the updated value or know that their cached version is outdated. Cache invalidation in multi-agent systems has three practical patterns:
TTL-based invalidation: Each memory entry has a time-to-live. After the TTL expires, the value is considered stale and agents must re-fetch from the source or flag the absence. This is simple to implement and works for values that naturally expire (stock prices, weather data, recent events).
Write-invalidate: When an agent writes a new value for a key, it broadcasts an invalidation signal to all agents that might have cached the old value. Agents that receive the signal discard their cached version. This requires agents to subscribe to invalidation events for the keys they cache.
Read-through with freshness checks: Agents always write to and read from the canonical shared memory store, never from a local cache. This is the simplest pattern — no invalidation complexity — but incurs a read overhead on every access. For most multi-agent systems running in a single process with in-memory shared state, read-through is the correct default.
⚡ Pro tip: Tag each shared memory write with the writing agent's name and the task ID.
⚡ Pro tip: Implement a shared memory health check that runs before your pipeline starts. Verify that all required keys exist with the correct types, that namespace prefixes are consistent, and that no keys contain stale values from a prior failed run. A five-second pre-flight check on shared memory state prevents the class of bugs where a prior run's residue corrupts the current run's results.
When debugging a wrong value in shared memory, you need to know not just what the value is, but which agent wrote it and in the context of which task. Without write attribution, shared memory bugs can take hours to trace. With it, they take minutes.
PromptABCD is a useful complement here — store your agents' system prompts and memory schema definitions together, so the prompt and the expected memory contract stay synchronized as you iterate.
Continue Reading
Save the prompts from this post
PromptABCD is a free prompt manager. Paste, organize, and reuse your best AI prompts — no more hunting through chat history.
