Sequential vs Parallel Agent Execution
Choosing between sequential vs parallel agents sounds like a performance question, but it's really a dependency question. Get it wrong and you'll either slow down tasks that could parallelize, or corrupt state between agents running simultaneously.
import concurrent.futures
import time
from anthropic import Anthropic
client = Anthropic()
def call_agent(system_prompt: str, user_message: str) -> str:
response = client.messages.create(
model="claude-sonnet-4-6",
max_tokens=1024,
system=system_prompt,
messages=[{"role": "user", "content": user_message}]
)
return response.content[0].text
# Sequential execution: each agent waits for the previous to finish
def run_sequential(topic: str) -> dict:
start = time.time()
research = call_agent("You are a researcher. Find key facts.", f"Research: {topic}")
analysis = call_agent("You are an analyst. Analyze these findings.", research) # Depends on research
summary = call_agent("You are a summarizer. Write a concise summary.", analysis) # Depends on analysis
return {"output": summary, "time": time.time() - start}
# Parallel execution: independent tasks run simultaneously
def run_parallel(topics: list[str]) -> dict:
start = time.time()
research_prompt = "You are a researcher. Find key facts and return a structured summary."
with concurrent.futures.ThreadPoolExecutor(max_workers=5) as executor:
futures = {executor.submit(call_agent, research_prompt, f"Research: {t}"): t for t in topics}
results = {}
for future in concurrent.futures.as_completed(futures):
topic = futures[future]
results[topic] = future.result()
return {"outputs": results, "time": time.time() - start}
# Sequential: ~9 seconds for 3 steps (3 × ~3 seconds each)
seq_result = run_sequential("machine learning infrastructure costs 2025")
print(f"Sequential time: {seq_result['time']:.1f}s")
# Parallel: ~3 seconds for 3 topics (all run simultaneously)
par_result = run_parallel(["Anthropic", "OpenAI", "Google DeepMind"])
print(f"Parallel time: {par_result['time']:.1f}s")A data engineering team built a market research system where three agents analyzed different competitor categories simultaneously. Performance was excellent — three parallel calls instead of three sequential. Then they discovered that one of the three agents was also reading from a shared context dictionary to look up recent market trends, and a second agent was writing to that same dictionary as it ran. The third agent sometimes read stale data, sometimes read partially-written data, and occasionally produced output that contradicted the first agent's findings.
The bug was subtle: the agents were parallel in execution but had a hidden sequential dependency through shared state. Parallel agents that share state are one of the most common sources of non-deterministic behavior in multi-agent systems — and the failure doesn't produce an error, it produces wrong output.
Understanding when to use sequential vs parallel agents — and what makes parallelism safe vs unsafe — is a fundamental multi-agent engineering skill.
Quick-Start: Sequential vs Parallel Side by Side
[object Object], concurrent.futures
,[object Object], time
,[object Object], anthropic ,[object Object], Anthropic
client = Anthropic()
,[object Object], ,[object Object],(,[object Object],) -> ,[object Object],:
response = client.messages.create(
model=,[object Object],,
max_tokens=,[object Object],,
system=system_prompt,
messages=[{,[object Object],: ,[object Object],, ,[object Object],: user_message}]
)
,[object Object], response.content[,[object Object],].text
,[object Object],
,[object Object], ,[object Object],(,[object Object],) -> ,[object Object],:
start = time.time()
research = call_agent(,[object Object],, ,[object Object],)
analysis = call_agent(,[object Object],, research) ,[object Object],
summary = call_agent(,[object Object],, analysis) ,[object Object],
,[object Object], {,[object Object],: summary, ,[object Object],: time.time() - start}
,[object Object],
,[object Object], ,[object Object],(,[object Object],) -> ,[object Object],:
start = time.time()
research_prompt = ,[object Object],
,[object Object], concurrent.futures.ThreadPoolExecutor(max_workers=,[object Object],) ,[object Object], executor:
futures = {executor.submit(call_agent, research_prompt, ,[object Object],): t ,[object Object], t ,[object Object], topics}
results = {}
,[object Object], future ,[object Object], concurrent.futures.as_completed(futures):
topic = futures[future]
results[topic] = future.result()
,[object Object], {,[object Object],: results, ,[object Object],: time.time() - start}
,[object Object],
seq_result = run_sequential(,[object Object],)
,[object Object],(,[object Object],)
,[object Object],
par_result = run_parallel([,[object Object],, ,[object Object],, ,[object Object],])
,[object Object],(,[object Object],)What this does: Sequential runs agents one after another — necessary when each agent's output feeds the next. Parallel runs the same agent against different inputs simultaneously — appropriate when each task is independent. The timing difference is significant: sequential takes roughly 3N seconds for N agents, parallel takes roughly the time of the slowest single agent regardless of N.
Understanding the Variables
Two questions determine whether parallel execution is safe:
Are the tasks independent? Independent means: Agent A's output is not needed to start Agent B's task, and Agent A's execution does not affect Agent B's input data. If both answers are "no," parallelism is safe. If either answer is "yes," you have a dependency and must sequence those agents.
Do agents share mutable state? If parallel agents write to the same dictionary, database row, or file without locking, you have a race condition. The fix: either make shared state read-only during parallel execution, give each agent its own write namespace, or use thread-safe data structures.
⚡ Pro tip: A quick way to determine if two tasks can parallelize: ask whether the output of one needs to be in the prompt of the other. If yes, they must be sequential. If the only connection is that both contribute to a final aggregation step, they can parallelize. Draw this dependency graph on paper before writing code — it's faster than discovering race conditions in testing.
Step-by-Step: Building a Hybrid Sequential-Parallel Pipeline
Most real-world systems need both sequential and parallel execution in different parts of the pipeline. Here's a practical pattern:
[object Object], concurrent.futures
,[object Object], anthropic ,[object Object], Anthropic
,[object Object], json
client = Anthropic()
,[object Object], ,[object Object],(,[object Object],) -> ,[object Object],:
,[object Object],
response = client.messages.create(
model=,[object Object],,
max_tokens=,[object Object],,
system=,[object Object],,
messages=[{
,[object Object],: ,[object Object],,
,[object Object],: ,[object Object],
}]
)
,[object Object],:
,[object Object], json.loads(response.content[,[object Object],].text)
,[object Object], json.JSONDecodeError:
,[object Object], {,[object Object],: company, ,[object Object],: [response.content[,[object Object],].text], ,[object Object],: ,[object Object],, ,[object Object],: ,[object Object],}
,[object Object], ,[object Object],(,[object Object],) -> ,[object Object],:
,[object Object],
research_summary = json.dumps(all_research, indent=,[object Object],)
response = client.messages.create(
model=,[object Object],,
max_tokens=,[object Object],,
system=,[object Object],,
messages=[{
,[object Object],: ,[object Object],,
,[object Object],: ,[object Object],
}]
)
,[object Object], response.content[,[object Object],].text
,[object Object], ,[object Object],(,[object Object],) -> ,[object Object],:
,[object Object],
response = client.messages.create(
model=,[object Object],,
max_tokens=,[object Object],,
system=,[object Object],,
messages=[{
,[object Object],: ,[object Object],,
,[object Object],: ,[object Object],
}]
)
,[object Object], response.content[,[object Object],].text
,[object Object], ,[object Object],(,[object Object],) -> ,[object Object],:
,[object Object],
,[object Object], concurrent.futures.ThreadPoolExecutor(max_workers=,[object Object],) ,[object Object], executor:
research_futures = [executor.submit(research_competitor, c) ,[object Object], c ,[object Object], companies]
all_research = [f.result() ,[object Object], f ,[object Object], concurrent.futures.as_completed(research_futures)]
,[object Object],
synthesis = synthesize_research(all_research)
,[object Object],
report = write_report(synthesis, report_request)
,[object Object], {
,[object Object],: all_research,
,[object Object],: synthesis,
,[object Object],: report
}
result = run_hybrid_pipeline(
companies=[,[object Object],, ,[object Object],, ,[object Object],, ,[object Object],, ,[object Object],],
report_request=,[object Object],
)
,[object Object],(result[,[object Object],])What this does: Phase 1 runs five competitor research agents simultaneously — they're independent and safe to parallelize. Phase 2 (synthesis) and Phase 3 (report writing) are sequential because each depends on the previous step's output. The hybrid model gets the speed benefits of parallelism where safe and the correctness benefits of sequential execution where dependencies exist.
⚡ Pro tip: The most important thing to do before implementing parallel agents is to make all shared state read-only during the parallel phase. In the example above, companies and report_request are read-only during the parallel research phase. The all_research list is built by aggregating results only after all parallel tasks complete — not during. This eliminates race conditions at the architecture level.
Pro-Level Variations
Timeout handling for parallel agents: Some parallel agents might take significantly longer than others. Use concurrent.futures.wait() with a timeout parameter to collect results from faster agents and handle stragglers explicitly, rather than waiting indefinitely for the slowest one.
Partial results on failure: If one of five parallel research agents fails, should the pipeline abort or continue with four results? Implement explicit failure tolerance: as_completed() with try-except around each future.result() lets you collect successful results and log failures without aborting the entire parallel phase.
Dynamic parallelism: For use cases where the number of parallel tasks isn't known upfront — crawling a variable-size list of documents, processing an inbox of unknown size — use a semaphore to cap concurrent API calls at a sensible limit rather than spawning unbounded threads.
Context contamination risk: A non-obvious risk in parallel agent systems is context contamination through shared LLM state. Each parallel call is stateless at the API level — but if you're using a framework that maintains conversation history (like AutoGen's ConversableAgent), parallel instances of the same agent might share history in unexpected ways. Test your parallel setup with a framework that you're sure is creating independent instances for each parallel task.
Troubleshooting Common Issues
Non-deterministic output between runs: If your parallel pipeline produces different results on identical inputs, you likely have a shared mutable state issue. Add logging to track write operations on any shared dictionaries, databases, or files during the parallel phase. The first write that happens in a non-deterministic order will show up in the log.
Parallel tasks not actually running simultaneously: Python's Global Interpreter Lock (GIL) doesn't block I/O-bound operations, so thread-based parallelism works well for API calls. But if you're inadvertently doing CPU-bound work between API calls, threads won't help — use ProcessPoolExecutor instead for CPU-bound parallel work.
Cost scaling unexpectedly: Parallel agents consume API rate limits simultaneously. Five parallel agents making calls at the same time might hit rate limits that sequential execution would avoid. Implement exponential backoff retry logic on each parallel task, or add a rate limiter that caps concurrent API calls.
⚠️ Common mistake: Running validation or quality checks in parallel with the tasks they're supposed to check. If your reviewer agent runs in parallel with your writer agent because "they can both work at the same time," the reviewer is reviewing something the writer hasn't finished yet. Validation must be sequential — always run after the task being validated.
Your Turn
Audit your current multi-agent pipeline and map it as a dependency graph: draw each agent as a node, draw an edge from A to B if B's input depends on A's output. Any agents with no incoming edges from other agents can run in parallel. Any with incoming edges must sequence after their dependencies.
Implement the parallel phase first, validate that it handles failures and race conditions correctly, then wire in the sequential synthesis step. The quality and performance improvement from this hybrid pattern is usually the most impactful single change you can make to an existing multi-agent system.
Safe Parallelism: The Isolation Checklist
Before running any two agents in parallel, verify all four isolation properties:
No shared writes. Neither agent writes to a field that the other agent reads or writes. Write conflicts in parallel agents produce race conditions: the last agent to write wins, and the order of writes is non-deterministic.
Read-only access to shared data. If both agents read from the same source (a shared context dict, a database, a cached document), verify that neither agent modifies that source during execution. Read-only shared access is safe; read-write shared access is not.
Independent inputs. Each agent receives its inputs before the parallel phase starts. Inputs aren't computed by one parallel agent and consumed by another — that's a sequential dependency, not a parallel one.
Aggregation happens after completion. Both agents complete fully before their outputs are combined. No partial aggregation during the parallel phase.
[object Object], asyncio
,[object Object], ,[object Object], ,[object Object],(,[object Object],) -> ,[object Object],:
,[object Object],
,[object Object],
tasks = [
asyncio.to_thread(agent, ,[object Object],(inp)) ,[object Object],
,[object Object], agent, inp ,[object Object], ,[object Object],(agents, inputs)
]
results = ,[object Object], asyncio.gather(*tasks, return_exceptions=,[object Object],)
,[object Object],
successful = [r ,[object Object], r ,[object Object], results ,[object Object], ,[object Object], ,[object Object],(r, Exception)]
failed = [r ,[object Object], r ,[object Object], results ,[object Object], ,[object Object],(r, Exception)]
,[object Object], failed:
,[object Object],(,[object Object],)
,[object Object], successfulWhat this does: Each agent receives dict(inp) — a copy of its input — rather than a reference to a shared object. Modifying the copy doesn't affect other agents' inputs. asyncio.gather runs all tasks concurrently and waits for all to complete before returning results. Aggregation of successful results happens after the gather returns — never during the parallel phase.
⚡ Pro tip: Log the start and end timestamps for each parallel agent with a shared task ID. When debugging non-deterministic output from a parallel system, the timestamps reveal whether agents overlapped as intended, whether one agent took significantly longer than others, and whether the aggregation happened before all agents completed. Timestamp logging costs almost nothing and is the first thing you'll want when a parallel execution produces unexpected results.
Choosing Between Sequential and Parallel
The choice between sequential vs parallel agents is ultimately a dependency question, not a performance question. If Agent B needs Agent A's output as input, they must be sequential. If Agent B and Agent A both receive the same input and produce independent outputs for a downstream aggregator, they can parallelize.
Draw the dependency graph explicitly before writing code. Sequential dependencies appear as directed edges between agents. Agents with no edges between them and a common output sink are candidates for parallelism. This graph makes the safe parallelism boundaries visible without requiring you to reason about shared state in code.
Debugging Non-Deterministic Parallel Output
When a parallel agent system produces different output on the same input across runs, the cause is almost always shared mutable state — state that one agent writes to while another agent reads from it. The output depends on which agent runs first, which is non-deterministic.
The fix is always isolation, not synchronization. Adding locks and semaphores to shared state in a parallel agent system is the wrong solution: it serializes the agents (defeating the purpose of parallelism) and introduces deadlock risks. The right solution is eliminating shared mutable state. Each agent should receive immutable inputs and write only to its own output. The orchestrator aggregates after all agents complete.
Save your orchestration patterns and agent prompts to PromptABCD to keep them accessible as you iterate on the pipeline structure.
Continue Reading
Save the prompts from this post
PromptABCD is a free prompt manager. Paste, organize, and reuse your best AI prompts — no more hunting through chat history.
