PromptABCD
FeaturesLearnGuideBlogContext Blocks
Sign inGet started free
Sign inSign up
PromptABCD

A calm home for your best AI prompts. Save them once, find them in seconds, reuse them forever.

Product

  • Features
  • Chrome Extension
  • Free Courses
  • How it works
  • Use cases
  • Blog
  • Context Blocks
  • Export Anywhere
  • FAQ

Resources

  • User guide
  • Learn prompting
  • Sign in
  • Get started free

© 2026 PromptABCD. All rights reserved.

AboutPrivacy PolicyTerms and Conditions
Home/Blog/CLI AI Agents/Streaming Output in a Terminal AI Agent
CLI AI Agents

Streaming Output in a Terminal AI Agent

A terminal agent that goes silent feels broken even when it works perfectly. Here's how cli agent streaming output makes the same task feel twice as fast — without changing a single tool.

September 16, 2026·9 min read
ShareShare
⚡Featured Prompt— copy and use right now
resp = client.messages.create(
    model="claude-sonnet-5",
    max_tokens=1024,
    tools=TOOLS,
    messages=messages,
)
print(resp.content[0].text)   # everything appears at once, at the very end

Here's a finding from usability research that reshapes how you build terminal tools: people perceive a wait with visible progress as roughly half as long as an identical wait staring at a frozen cursor. Same forty seconds, half the felt pain — purely because something was moving on screen. That single fact is the entire argument for cli agent streaming output, and it's why the first thing seasoned builders add after the core loop is streaming, not more tools.

This is the story of a data team whose agent worked perfectly and got abandoned anyway, and how adding cli agent streaming output brought it back from the dead — without changing a single tool.

The Problem the Data Team Faced

A three-person analytics team at a subscription company built a terminal agent to answer questions about their warehouse: "which cohorts churned hardest last quarter," "why did signups dip on the 14th." The agent was accurate. It pulled the right tables, ran the right aggregations, and summarized clearly. On paper it was a success.

In practice, people stopped using it. The complaint was always the same: "it feels broken." A real question took twenty to forty seconds — a model call, a couple of SQL tool runs, another model call to summarize. For all of that time, the terminal showed nothing. No output, no spinner, no sign of life. Users assumed it had hung, hit Ctrl-C, and went back to writing SQL by hand. The tool wasn't slow by machine standards. It was slow by human standards, because silence reads as failure.

The Wrong Approach

The first instinct was to make it faster. The team optimized the SQL, switched to a quicker model for simple questions, and shaved maybe eight seconds off. It didn't help. People still bailed.

python
resp = client.messages.create(
    model=,[object Object],,
    max_tokens=,[object Object],,
    tools=TOOLS,
    messages=messages,
)
,[object Object],(resp.content[,[object Object],].text)   ,[object Object],

What this does: Blocks until the entire response is generated, then prints it in one dump. The user sees nothing for the whole generation and then a wall of text arrives. Faster generation doesn't fix the core problem, which is the silence, not the speed.

The mistake was treating a perception problem as a performance problem. Eight seconds faster is still fifteen seconds of dead air, and dead air is what killed the tool. You can't optimize your way out of a feeling.

⚠️ Common mistake: Chasing raw speed when the real issue is perceived responsiveness. A tool that streams output over thirty seconds feels faster than one that goes silent for twenty. Fix the feedback loop before you touch the model.

The Right Approach: Stream Everything

The rewrite changed nothing about the tools or the SQL. It only changed when the user saw output — from "all at the end" to "as it happens." Python's SDK makes this a small change.

python
[object Object], client.messages.stream(
    model=,[object Object],,
    max_tokens=,[object Object],,
    tools=TOOLS,
    messages=messages,
) ,[object Object], stream:
    ,[object Object], text ,[object Object], stream.text_stream:
        ,[object Object],(text, end=,[object Object],, flush=,[object Object],)
    final = stream.get_final_message()

What this does: Prints the model's text token-by-token as it's generated using text_stream, while still collecting the complete final message so the loop can extract tool calls afterward. The flush=True is essential — without it, Python buffers the output and you're back to a delayed dump.

The moment reasoning appeared live on screen, the abandonment stopped. Users watched the agent think — "Let me check the churn table... that shows a spike in the enterprise tier..." — and a task that used to feel like a gamble now felt like a conversation. Nothing was actually faster. It just stopped feeling broken.

⚡ Pro tip: Always pass flush=True (Python) or write directly to process.stdout (Node). Terminal output is line-buffered by default, which means your carefully streamed tokens can sit in a buffer until a newline — recreating the exact lag you were trying to remove.

Streaming Around Tool Calls

The subtle part is streaming an agent, not just a single reply. An agent alternates between talking and acting: it narrates, calls a tool, waits for the result, narrates again. Naive streaming shows the narration but leaves another silent gap during the tool run. The fix is to narrate the tool itself.

python
[object Object], ,[object Object],(,[object Object],):
    ,[object Object], client.messages.stream(
        model=,[object Object],, max_tokens=,[object Object],,
        tools=TOOLS, messages=messages,
    ) ,[object Object], stream:
        ,[object Object], event ,[object Object], stream:
            ,[object Object], event.,[object Object], == ,[object Object], ,[object Object], \
               event.content_block.,[object Object], == ,[object Object],:
                ,[object Object],(,[object Object],, flush=,[object Object],)
            ,[object Object], event.,[object Object], == ,[object Object],:
                ,[object Object],(event.text, end=,[object Object],, flush=,[object Object],)
        ,[object Object], stream.get_final_message()

What this does: Iterates the full event stream instead of only text, so when a tool-use block starts, it prints a live "→ running query_warehouse..." line. Now even the tool-execution phase shows motion, closing the last silent gap in the loop.

This is the detail generic streaming tutorials skip. They stream one model reply and call it done. An agent has multiple silent phases, and each one needs its own signal or the "did it freeze?" feeling creeps back in during tool runs. Narrate the whole cycle, not just the talking parts.

⚡ Pro tip: You can even stream the tool's arguments as they form. Tool inputs arrive as input_json_delta events, so you can print "query_warehouse(table=churn, quarter=..." building up character by character. It's oddly reassuring to watch the agent construct its query in real time — and it exposes a bad tool call a beat sooner.

Results and What Changed

Within two weeks of shipping the streaming version, daily active use of the agent tripled back to where the demo had suggested it would be. The team measured one number that told the whole story: time-to-first-visible-output dropped from around eighteen seconds to under one. Total task time barely moved. Perceived responsiveness moved enormously.

The qualitative feedback shifted too. "It feels broken" became "it feels fast," about a tool that was, by the stopwatch, no faster at all. One analyst described watching it reason through a tricky cohort question as "like pairing with someone who thinks out loud." That's the effect cli agent streaming output buys you — presence, not speed.

⚡ Pro tip: Print a subtle separator between streamed reasoning and tool output — a dim line or a blank line. Without visual breaks, the model's narration and a command's raw output smear into one block that's miserable to read back later. Structure the stream, don't just dump it.

How to Apply This to Your Situation

The pattern transfers to any terminal agent, and the persona barely matters.

A DevOps engineer running an incident-triage agent wants to watch it check pods, tail logs, and correlate timestamps live — because during an outage, a silent tool is one they'll kill and distrust exactly when they need it most.

A bioinformatician running long analysis steps needs the agent to narrate "loading the variant file... 40,000 rows... filtering by quality" so a genuinely slow step is visibly working rather than ambiguously hung.

A customer support lead using an agent to draft replies wants the draft to appear as it's written, so they can start reading — and start editing in their head — before it's finished, shaving real minutes off each ticket.

Every one of them benefits from the same move: show motion, narrate the silent phases, flush your output. The tools don't change. The feeling does.

How Do You Stream Markdown Without It Looking Broken?

Here's a wrinkle that trips people up the moment their agent's output contains formatting: streamed markdown looks wrong mid-flight. A code fence that hasn't closed yet, a half-written bullet, a table with one row — rendering those partial structures live produces flickering garbage. The naive fix is to render markdown on every token, and it looks terrible.

python
buffer = ,[object Object],
,[object Object], client.messages.stream(...) ,[object Object], stream:
    ,[object Object], text ,[object Object], stream.text_stream:
        buffer += text
        ,[object Object], ,[object Object], ,[object Object], text:                 ,[object Object],
            ,[object Object],(render_ready_lines(buffer), end=,[object Object],, flush=,[object Object],)
            buffer = keep_incomplete_tail(buffer)

What this does: Accumulates streamed text and flushes only when a line completes, so you never try to render a half-open code fence or a partial table. The user still sees near-real-time output, just at line granularity instead of token granularity.

The honest tradeoff: for plain prose, stream every token — it's smoother and there's nothing to break. For structured markdown, buffer to line boundaries or skip rich rendering entirely and show raw text, which many terminal agents do because raw streamed markdown is perfectly readable and never looks broken. Match the strategy to the content, and don't let a fancy renderer reintroduce the exact jank you added streaming to remove.

⚡ Pro tip: If you add ANSI color to streamed output, reset the color at the end of every flush with \033[0m. A stream interrupted mid-escape-sequence — by a Ctrl-C or an error — can leave the user's terminal stuck in a color, which looks like your tool corrupted their shell. Always close what you open.

When Not to Stream

Honesty about tradeoffs: streaming isn't free and isn't always right. If your agent's output is being piped into another program rather than read by a human — agent "..." | jq — streaming partial tokens is worse than useless, because the consumer wants one clean, complete payload. Detect whether stdout is a TTY and stream only when a human is watching.

python
[object Object], sys
INTERACTIVE = sys.stdout.isatty()
,[object Object],

What this does: Checks whether output is going to a terminal or being piped elsewhere. When piped, you switch off streaming and emit a single complete result, so downstream tools get well-formed input instead of a dribble of tokens.

This TTY check is the mark of a tool built by someone who's actually shipped one. Streaming is a human affordance; machines want atomic output. Serving both correctly is a small isatty() branch, and skipping it is why some agents are maddening to script around.

Next Steps

Add streaming to your loop today: swap messages.create for messages.stream, print with flush=True, narrate tool starts, and guard it behind a TTY check. It's an afternoon of work and it's the highest-impact change you can make to how a terminal agent feels.

The narration lines you write — "→ running...", "checking the logs..." — are small prompts and templates in their own right, and getting their tone right (informative, not chatty) takes iteration. Keeping the phrasings that work in a prompt library like PromptABCD means your next agent inherits a streaming experience you've already tuned, instead of rediscovering from scratch what good live narration sounds like.

cli agentsstreamingterminalai agentsdeveloper toolsux

Continue Reading

Managing Reusable Prompts for Terminal Workflows
CLI AI Agents

Managing Reusable Prompts for Terminal Workflows

Retyping your best prompt from memory loses its refinements every time. Managing cli agent reusable prompts as named, parameterized, versioned assets keeps the prompt quality you earned — and lets you share it.

September 19, 2026·9 min read
Distributing System Prompts With Your CLI Tool
CLI AI Agents

Distributing System Prompts With Your CLI Tool

Hardcoding your agent's system prompt as a string is the wrong place for it. Treating cli agent system prompt distribution as content — versioned, overridable, updatable — is how prompts evolve independently of code.

September 19, 2026·9 min read
Building a Plugin System for Your CLI Agent
CLI AI Agents

Building a Plugin System for Your CLI Agent

How do you let people add tools to your agent without forking it? A cli agent plugin system lets users extend the agent with their own tools. Here's how to rebuild a hardcoded tool list into a real plugin system.

September 19, 2026·8 min read

Save the prompts from this post

PromptABCD is a free prompt manager. Paste, organize, and reuse your best AI prompts — no more hunting through chat history.

Start free →
← PreviousAdding Autocomplete to Your CLI AgentNext →Handling Interactive Prompts in a CLI Agent
Share this post:
ShareShare