Tracing a Request Across Many Agents
In one system, the average request touched 11 agents before producing an answer. Multi agent distributed tracing is the only way to follow that path. Here's how to build it.
import time, uuid
class TraceContext:
def __init__(self, trace_id=None, parent_span=None):
self.trace_id = trace_id or uuid.uuid4().hex
self.parent_span = parent_span
def child(self, name):
return Span(self.trace_id, name, parent=self.parent_span)
class Span:
def __init__(self, trace_id, name, parent=None):
self.trace_id, self.name = trace_id, name
self.span_id = uuid.uuid4().hex
self.parent = parent
self.start = time.time()
self.attrs = {}
def set(self, key, value):
self.attrs[key] = value
def end(self):
emit_span(self.trace_id, self.span_id, self.name,
self.parent, time.time() - self.start, self.attrs)In one agent system I profiled, the average user request touched eleven different agents before producing an answer - a router, several specialists, a couple of tool-wrapping agents, a critic, and a synthesizer. Eleven. When one of those requests produced a bad result, the team's debugging process was to guess which agent was at fault and read its logs in isolation. Multi agent distributed tracing exists precisely because that guess-and-check approach doesn't scale past two or three agents, and almost every real system has more.
Tracing follows a single request as it moves through many agents, stitching together each agent's work into one coherent story you can read start to finish. It's borrowed directly from microservices - where a request hops across many services - and it adapts cleanly to agents, because an agent team is a distributed system whether you designed it as one or not. This guide builds tracing you can add to an existing team without rewriting it.
Quick-Start (Copy This Right Now)
Here's the core: a trace context that flows with a request and a span that each agent opens to record its slice of the work.
[object Object], time, uuid
,[object Object], ,[object Object],:
,[object Object], ,[object Object],(,[object Object],):
,[object Object],.trace_id = trace_id ,[object Object], uuid.uuid4().,[object Object],
,[object Object],.parent_span = parent_span
,[object Object], ,[object Object],(,[object Object],):
,[object Object], Span(,[object Object],.trace_id, name, parent=,[object Object],.parent_span)
,[object Object], ,[object Object],:
,[object Object], ,[object Object],(,[object Object],):
,[object Object],.trace_id, ,[object Object],.name = trace_id, name
,[object Object],.span_id = uuid.uuid4().,[object Object],
,[object Object],.parent = parent
,[object Object],.start = time.time()
,[object Object],.attrs = {}
,[object Object], ,[object Object],(,[object Object],):
,[object Object],.attrs[key] = value
,[object Object], ,[object Object],(,[object Object],):
emit_span(,[object Object],.trace_id, ,[object Object],.span_id, ,[object Object],.name,
,[object Object],.parent, time.time() - ,[object Object],.start, ,[object Object],.attrs)What this does: It creates a trace ID that identifies one request and lets each agent open a span - a timed record of its work - that references its parent. When every agent opens a span as a child of the caller's span, the emitted spans reconstruct the full tree of who called whom, how long each took, and what each recorded.
Understanding the Variables
Three ideas make multi agent distributed tracing work: trace context, spans, and propagation.
Trace context is the small bundle of identifiers that travels with a request - at minimum a trace ID (identifying the whole request) and the current span ID (identifying the agent currently working). Every agent that touches the request reads this context and adds to it. The trace ID is what lets you later gather every span belonging to one request out of millions.
Spans are the individual units of work. Each agent, and often each significant operation within an agent (a tool call, a model call), opens a span, does its work, and closes it. A span records timing, the agent's name, and attributes you attach - the decision made, tokens used, confidence. Spans nest: a synthesizer's span is the parent of the specialist spans it invoked.
Propagation is how the trace context moves from one agent to the next. This is the part that breaks most often. When agent A calls agent B, A must pass the trace context to B so B's spans attach to the same trace. Miss this and B's work shows up as a separate, orphaned trace - you've lost the thread exactly where the request crossed an agent boundary.
⚡ Pro tip: Propagation is where traces break, so make it automatic, not manual. If passing the trace context depends on every agent-to-agent call remembering to include it, someone will forget, and that request's trace will fragment. Wrap your inter-agent call mechanism so the context rides along by default - the same way HTTP tracing libraries inject headers automatically rather than trusting each call site.
Step-by-Step: Building Multi Agent Distributed Tracing
Let's instrument a real multi-agent flow so one request produces one readable trace.
Step one, start a trace at the entry point - wherever a request enters your system - and open a root span:
[object Object], ,[object Object],(,[object Object],):
ctx = TraceContext() ,[object Object],
root = ctx.child(,[object Object],)
root.,[object Object],(,[object Object],, ,[object Object],(user_input))
,[object Object],:
,[object Object], orchestrator.run(user_input, ctx) ,[object Object],
,[object Object],:
root.end()What this does: It mints a fresh trace for each incoming request and opens the root span everything else nests under. Passing ctx into the orchestrator is what starts propagation - every downstream agent will receive and extend this same context, so their spans join this trace.
Step two, have each agent open a child span and propagate the context to any agent it calls:
[object Object], ,[object Object],(,[object Object],):
span = ctx.child(agent.name)
child_ctx = TraceContext(ctx.trace_id, parent_span=span.span_id)
,[object Object],:
span.,[object Object],(,[object Object],, task.,[object Object],)
result = agent.execute(task, child_ctx) ,[object Object],
span.,[object Object],(,[object Object],, result.confidence)
,[object Object], result
,[object Object],:
span.end()What this does: It opens a span for the agent's work and creates a child context whose parent is this span, then passes that child context into the agent's execution. Any agent this one calls will attach beneath it in the trace tree, so the nesting mirrors the actual call structure.
Step three, build the view that assembles spans into a trace. Given a trace ID, gather all its spans and order them into a tree by parent references. This is what turns raw span records into the readable timeline that makes an eleven-agent request comprehensible at a glance.
Pro-Level Variations
For systems using asynchronous coordination - queues, event buses - propagation gets trickier because the caller and callee are decoupled. Carry the trace context inside the message or event payload so a consumer picks it up when it processes the message, even though the producer is long gone. Tracing across async boundaries is entirely possible; you just move the context from the call stack into the message body.
For high-volume systems, sample traces rather than capturing every one - trace a percentage of requests fully, plus every request that errors. Full tracing on every request at scale is expensive, and a representative sample plus all failures gives you nearly all the debugging value at a fraction of the cost.
⚡ Pro tip: Always trace the failures, sample the successes. A uniform 5% sample might miss the exact failed request you need to debug. Trace 100% of errors and flagged requests, and sample the rest - so you never lose the trace of the thing that actually went wrong, which is the whole reason you built tracing.
What Should Each Span Actually Record?
A trace tree that shows timing and nesting is useful, but the real debugging power comes from the attributes you attach to each span. An empty span tells you an agent ran and how long it took; a well-populated span tells you what it decided and why. The difference between the two is what separates a trace you can debug from a trace you can only admire.
For agent spans specifically, capture the task type the agent handled, the model it used, the tokens it consumed, its confidence in the result, and the key decision it made - which tool it picked, which branch it took. These are the attributes you'll filter and sort by when hunting a problem across thousands of traces.
[object Object], ,[object Object],(,[object Object],):
span.,[object Object],(,[object Object],, task.,[object Object],)
span.,[object Object],(,[object Object],, result.model)
span.,[object Object],(,[object Object],, result.total_tokens)
span.,[object Object],(,[object Object],, result.confidence)
span.,[object Object],(,[object Object],, result.key_decision) ,[object Object],
,[object Object], result.error:
span.,[object Object],(,[object Object],, result.error) ,[object Object],What this does: It attaches the attributes that make a span queryable - task type, model, cost, confidence, and the decision the agent made - plus any error. With these on every span, you can ask questions like "show me every trace where a low-confidence agent chose the search tool," which is the kind of query that finds a systemic bug rather than a single instance.
The attribute that pays off most unexpectedly is token count per span. Because it rides the same trace, you get a per-request cost breakdown for free - you can see exactly which agent in an eleven-agent request burned the tokens. Multi agent distributed tracing thus doubles as cost attribution, turning "our bill went up" into "this specific agent's token use tripled last Tuesday," which is a fixable statement rather than a mystery.
⚡ Pro tip: Put token count on every span and you get cost attribution for free alongside your debugging. The same trace that tells you where a request went wrong tells you where a request got expensive, because both questions are answered by breaking the request down per agent. One instrumentation effort, two high-value payoffs.
Troubleshooting Common Issues
If traces come out fragmented - one request appearing as several disconnected traces - your context propagation is breaking at some agent boundary. Find the boundary where the trace ID changes or a new trace starts unexpectedly; that's the call that isn't passing context. This is the most common tracing bug and it's always a propagation gap.
If spans have no timing or the tree looks flat, agents probably aren't setting parent references correctly, so everything attaches to the root instead of nesting. Check that each agent creates a child context pointing at its own span before calling downstream agents.
⚠️ Common mistake: Adding tracing but never building the view that assembles spans into a readable trace. Emitting spans into a log is necessary but not sufficient - if the only way to see a trace is to grep for a trace ID and mentally reconstruct the tree, nobody will do it under incident pressure. The payoff of multi agent distributed tracing comes entirely from the assembled view; without the timeline that shows the eleven agents in order with their timings and attributes, you've paid the instrumentation cost and left the benefit on the table.
Your Turn
Add a trace context to your entry point, propagate it through one agent-to-agent call, and confirm both agents' spans share a trace ID. Then build the simplest possible trace view - trace ID in, ordered spans out. That view is what converts tracing from a data-collection exercise into a debugging superpower.
Keep your span attribute conventions and instrumentation wrapper versioned so every agent traces consistently. I store the tracing conventions - which attributes every span should carry - in PromptABCD, because consistent span attributes across a whole team are what make traces comparable and searchable, and defining that convention once and reusing it beats letting each agent record a different, incompatible set of attributes that fragments your ability to query across the fleet.
⚡ Pro tip: Adopt an existing tracing standard's data model rather than inventing your own span format. The concepts here - traces, spans, parent references, attributes - match established distributed-tracing conventions, so shaping your spans the same way means you can later feed them into mature tracing tools instead of building your own viewer forever. Borrowing the proven data model costs nothing now and saves you from a homegrown format you'll outgrow.
Continue Reading
Save the prompts from this post
PromptABCD is a free prompt manager. Paste, organize, and reuse your best AI prompts — no more hunting through chat history.
