Multi-Agent Systems for Incident Response: A Case Study
When production breaks at 3am, who investigates while everyone's asleep? A multi agent incident response system triages, gathers context, and drafts a timeline before a human even joins. Here's a real deployment.
def on_alert(alert):
# fire specialists in parallel the instant an alert lands
context = parallel(
triage_agent(alert), # severity + blast radius
change_agent(alert), # recent deploys / config changes
log_agent(alert), # relevant error patterns
dependency_agent(alert), # upstream/downstream health
)
return coordinator(alert, context) # drafts the human briefingWhen production breaks at 3am, who investigates while the on-call engineer is still rubbing sleep out of their eyes? For most teams the answer is nobody — the first fifteen minutes of an incident are lost to a groggy human gathering context that was fully available the whole time. That gap is where a multi agent incident response system earns its place. It can't fix the outage, but it can do all the context-gathering a human would spend precious minutes on, so that when the engineer joins, the investigation is already half done. This case study follows a real deployment, what it automated safely, and the hard line it never crossed.
What problem was the team actually solving?
Sana runs platform reliability at a payments company where downtime costs real money per minute. Her on-call engineers were competent but human — woken at 3am, they'd spend the first ten to fifteen minutes doing the same mechanical work every time: which service is alerting, what changed recently, what do the logs say, is this affecting customers. Only then could real diagnosis begin. That mechanical opening was identical across incidents and entirely automatable.
The insight that shaped the design: the slow part of early incident response isn't thinking, it's gathering. A human diagnostician is valuable for judgment, wasteful for log-scraping. So the system would gather; the human would judge. Sana built a multi agent incident response system where specialist agents assemble context in parallel the moment an alert fires, and a coordinator drafts a briefing for the human who's just waking up.
[object Object], ,[object Object],(,[object Object],):
,[object Object],
context = parallel(
triage_agent(alert), ,[object Object],
change_agent(alert), ,[object Object],
log_agent(alert), ,[object Object],
dependency_agent(alert), ,[object Object],
)
,[object Object], coordinator(alert, context) ,[object Object],What this does: it launches four context-gathering specialists the moment an alert fires and hands their combined output to a coordinator, so the mechanical first fifteen minutes happen automatically before the engineer is even at their laptop.
The wrong way the team first tried it
Sana's first version overreached. She gave the agents the ability to act — restart services, roll back deploys, scale resources — reasoning that automating remediation would cut downtime further. It cut downtime in the demo and caused a worse incident in week two.
An agent misread a symptom, "remediated" by rolling back a deploy that wasn't the cause, and turned a single degraded service into a cascading failure because the rollback broke a dependency the agent hadn't reasoned about. The lesson was sharp: the agents were good at gathering context and bad at the judgment that remediation requires, because remediation demands understanding consequences across a system no agent fully models.
⚠️ Common mistake: giving incident-response agents the power to remediate autonomously. Gathering context is safe — reading logs changes nothing. Taking action is dangerous, because a wrong action during an incident can convert a small problem into a large one, and the agent can't fully reason about the blast radius of its own fix. Keep agents on the investigation side of the line and humans on the action side.
The correct design: gather, brief, and stop
The rebuild drew a hard boundary. Agents gather and analyze; agents propose; humans decide and act. The coordinator produces a briefing that includes suggested actions, but every action requires a human to execute it. The agents make the human faster, not absent.
[object Object], ,[object Object],(,[object Object],):
system = (,[object Object],
,[object Object],
,[object Object],
,[object Object],
,[object Object],
,[object Object],)
,[object Object], llm(system=system, user=,[object Object],)What this does: it turns gathered context into a decision-ready briefing with ranked, risk-labeled suggestions the human can act on immediately — while keeping every actual action behind a human decision, the boundary that prevents the agents from making incidents worse.
The change agent turned out to be the highest-value specialist. Most incidents follow a change, and an agent that instantly correlates the alert with recent deploys and config changes answers the single most useful question — "what did we just change?" — before the human asks it. That correlation alone routinely cut diagnosis time dramatically.
⚡ Pro tip: give the change agent a tight time window and rank changes by proximity to the alert. An incident's cause is usually the most recent change to the affected service, so an agent that surfaces "this service was deployed four minutes before the alert" points straight at the likely culprit. Ranking changes by recency and relevance, rather than dumping a change log, is what makes this specialist genuinely useful rather than another wall of data.
Results and what changed
Time-to-first-meaningful-action dropped substantially, because the engineer now joined an incident that was already triaged, contextualized, and paired with ranked suggestions instead of a blank alert. The engineers reported the biggest win wasn't speed — it was starting from a briefing instead of a cold start, which meant less panic and clearer thinking at 3am.
The system never took an action on its own after the rebuild, and that constraint was what made the team trust it. A tool that only informs can't make things worse, so engineers leaned on it fully. Sana was explicit that the earlier action-taking version, for all its demo appeal, had been on track to erode the trust the informing version earned.
⚡ Pro tip: have the coordinator explicitly state its confidence and what it's uncertain about. A briefing that says "most likely the cache deploy, but I could not confirm the error spike is causally linked" is more useful than false confidence, because it tells the engineer exactly where to apply human judgment. The agents' honesty about their own uncertainty is what lets a human trust the parts they're confident about.
How to apply this to your situation
Start by timing your own incidents' opening minutes and listing the mechanical gathering steps your engineers repeat every time. Those steps — triage, change correlation, log scraping, dependency checks — are your specialist agents. Automate the gathering, draft a briefing, and stop there.
Draw the action boundary in code, not just policy. The agents should have read access to your observability stack and no write access to anything that changes production state. That constraint is what makes the system safe to trust during the exact moments when trust matters most.
⚠️ Common mistake: letting the briefing grow into an exhaustive report. A groggy engineer at 3am needs the three things that matter, not everything the agents found. A briefing longer than a screen defeats its purpose — the human ends up scrolling instead of acting. Force the coordinator to be ruthlessly concise, and keep the full gathered context one click away for the engineer who wants to dig deeper.
What does a multi agent incident response system cost to run?
The obvious cost of a multi agent incident response system is tokens, and that's the one that matters least. Firing four specialists on every alert costs real money at scale, but the number that actually determines whether the system helps or harms is the false-alarm cost — what happens when the agents gather context for an alert that turns out to be noise.
Sana's team learned this quickly. Their alerting was noisy, and the incident system dutifully spun up a full investigation for every flapping alert, flooding the on-call channel with briefings for non-incidents. The context-gathering was correct; the problem was that it fired indiscriminately. The fix was a triage gate that decides whether an alert warrants full investigation before the expensive specialists run, so the system's effort scales with actual severity rather than raw alert volume.
[object Object], ,[object Object],(,[object Object],):
system = (,[object Object],
,[object Object],
,[object Object],
,[object Object],
,[object Object],
,[object Object],)
,[object Object], json.loads(llm(system=system, user=json.dumps(alert)))What this does: it puts a cheap triage decision in front of the expensive specialists, suppressing investigation of known noise while erring toward investigating anything genuinely uncertain — so cost scales with real severity, not alert volume.
The asymmetry in that prompt is deliberate and worth internalizing: a false investigation wastes some tokens and a moment of attention, while a missed incident can cost far more. So the triage gate is tuned to over-investigate rather than under-investigate, accepting some waste to avoid the expensive failure. This is the same risk-asymmetry reasoning that governs the whole system — the cost of the agents being wrong is not symmetric, and the design should reflect which direction of error hurts more.
The real return isn't measured in tokens saved but in engineer well-being and retention, which rarely make it into a cost model. On-call is a leading cause of burnout, and a system that removes the 3am cold-start panic makes on-call meaningfully less brutal. That's a cost-benefit line that doesn't show up in a token budget but shows up clearly in whether your senior engineers stay.
⚡ Pro tip: track how often engineers act on the coordinator's suggestions versus overriding them, per suggestion type. Where they consistently override, the agents are wrong about that scenario and you should stop suggesting it. Where they consistently act, you've found a pattern reliable enough to consider elevating. This override-rate signal tells you exactly where the system is trusted and where it's noise, which is the only honest measure of whether it's working.
Next steps
Instrument your next incident to see how much of the first fifteen minutes is mechanical gathering versus real diagnosis. That measurement tells you exactly how much a multi agent incident response system could give back to your on-call engineers.
As you tune the specialist and coordinator prompts to your stack, save them as a versioned set in PromptABCD. Incident response works best when the process is identical every time, and keeping your gathering and briefing prompts in one place means every on-call engineer gets the same fast, consistent start instead of a system that drifts as different people tweak it.
Continue Reading
Save the prompts from this post
PromptABCD is a free prompt manager. Paste, organize, and reuse your best AI prompts — no more hunting through chat history.
