Handling Interactive Prompts in a CLI Agent
When a command your agent runs stops to ask a question, the agent hangs forever waiting on nobody. This teardown rebuilds a naive command runner to handle cli agent interactive prompts safely.
import subprocess
def run(cmd):
r = subprocess.run(cmd, shell=True, capture_output=True,
text=True, timeout=30)
return r.stdout or r.stderrPicture this: you're a backend engineer, you've built a terminal agent that can run your project's setup commands, and you ask it to initialize a new service. It runs npm init, which stops and waits for someone to type a package name. But there's no someone — the agent is a subprocess, nothing is connected to that prompt, and your agent sits there forever, frozen on a question nobody will ever answer. You Ctrl-C, curse a little, and realize handling cli agent interactive prompts is the problem nobody warned you about.
This teardown takes the naive way of running commands and rebuilds it into something that survives contact with tools that expect a human on the keyboard.
Before: The Naive Command Runner
Here's the version almost everyone writes first. It runs whatever command the model asked for and captures the output.
[object Object], subprocess
,[object Object], ,[object Object],(,[object Object],):
r = subprocess.run(cmd, shell=,[object Object],, capture_output=,[object Object],,
text=,[object Object],, timeout=,[object Object],)
,[object Object], r.stdout ,[object Object], r.stderrWhat this does: Runs a command and returns its output, with a timeout as the only guardrail. It works fine for ls and git log — and hangs the instant a command decides to ask a question.
Why It Fails
The failure is invisible until it isn't. Plenty of common commands pause for input under conditions you don't control: npm init wants a package name, git rebase opens an editor, apt install asks "Do you want to continue? [Y/n]", psql prompts for a password, an rm -i asks for per-file confirmation. When the process pauses for stdin and there's no stdin, it waits indefinitely.
Your timeout=30 eventually fires, which feels like a fix but isn't. You've turned a hang into a failure with no useful information — the agent gets "timed out" and has no idea a prompt was waiting. Worse, some tools detect they're not attached to a terminal and change behavior silently, so you get subtly different output than a human would, and the model reasons from the wrong data.
The root problem is that cli agent interactive prompts assume a two-way channel your subprocess doesn't have. You captured output but never provided a way to answer.
⚠️ Common mistake: Relying on a timeout to catch interactive hangs. A timeout tells you something went wrong, not that a prompt was waiting, and it wastes the full timeout duration every time. Prevent the prompt or answer it — don't wait for it to expire.
After: Prevent, Then Answer
The rebuilt runner does two things in order. First, it tries to stop commands from prompting at all, because the best interactive prompt is the one that never happens. Second, for the cases where a prompt is unavoidable, it drives a real pseudo-terminal and answers programmatically.
[object Object], os, subprocess
,[object Object],
NONINTERACTIVE_ENV = {
**os.environ,
,[object Object],: ,[object Object],,
,[object Object],: ,[object Object],, ,[object Object],
,[object Object],: ,[object Object],,
,[object Object],: ,[object Object],, ,[object Object],
}
,[object Object], ,[object Object],(,[object Object],):
r = subprocess.run(cmd, shell=,[object Object],, capture_output=,[object Object],, text=,[object Object],,
timeout=,[object Object],, stdin=subprocess.DEVNULL,
env=NONINTERACTIVE_ENV)
,[object Object], r.stdout ,[object Object], r.stderrWhat this does: Sets environment variables that many tools respect to skip prompts, and — critically — points stdin at DEVNULL so a prompting command gets an immediate EOF and errors out fast instead of hanging. A clean error the model can read beats an infinite wait every time.
The stdin=DEVNULL line is the quiet hero. With it, a command that tries to read input gets end-of-file instantly, fails in a second, and returns an error the agent can actually reason about ("this wants confirmation") rather than freezing. Non-interactive flags handle the rest: apt -y, npm init -y, pip install without prompts.
For the genuinely unavoidable cases — a tool with no non-interactive mode — you drive it through a pseudo-terminal and pattern-match the prompts.
[object Object], pexpect
,[object Object], ,[object Object],(,[object Object],):
child = pexpect.spawn(cmd, timeout=,[object Object],, encoding=,[object Object],)
transcript = []
,[object Object], ,[object Object],:
i = child.expect([,[object Object],, ,[object Object],, pexpect.EOF])
transcript.append(child.before)
,[object Object], i == ,[object Object],: ,[object Object],
,[object Object],
child.sendline(answers.get(i, ,[object Object],)) ,[object Object],
,[object Object], ,[object Object],.join(transcript)What this does: Uses pexpect to spawn the command inside a pseudo-terminal, watch for known prompt patterns, and send predefined answers. The command believes it's talking to a real terminal, so it behaves normally, and your code supplies the responses a human otherwise would.
Breaking Down Each Element
The pseudo-terminal matters because a plain pipe isn't a terminal. Many programs call isatty() and behave differently — or refuse to prompt — when they detect a pipe. pexpect allocates a real PTY, so the command runs exactly as it would for a person, which is both the power and the danger of the approach.
The expect pattern list is a small state machine. You enumerate the prompts you're willing to handle, pair each with an answer, and treat anything unexpected as a stop condition rather than guessing. That last part is a safety rule, not a nicety: an unrecognized prompt is a command asking something you didn't anticipate, and blindly sending "y" to an unknown question is how agents cause damage.
The layered order — suppress first, answer second — reflects a real preference. Answering prompts programmatically is fragile: prompt text changes between tool versions, locales shift the wording, and your patterns rot. Suppressing prompts with flags and environment variables is durable. Reach for pexpect only when a tool leaves you no other option.
⚡ Pro tip: Keep a small registry mapping commands to their non-interactive form — apt install becomes apt-get -y install, npm init becomes npm init -y. Rewrite the model's requested command through this registry before running it, so you fix prompts at dispatch time instead of scattering flags through your prompt.
⚡ Pro tip: When you can't suppress and can't confidently answer, escalate to the human instead of guessing. Pause the agent, surface the exact prompt text, and let the user type the answer. A three-second question to a person beats a wrong automated "yes" to an irreversible action.
Surfacing Prompts to the User
Sometimes the right answer genuinely isn't yours to give — a destructive confirmation, a credential, a choice with real consequences. The mature pattern hands those specific prompts back to the operator while handling the mechanical ones automatically.
DANGEROUS = (,[object Object],, ,[object Object],, ,[object Object],, ,[object Object],)
,[object Object], ,[object Object],(,[object Object],):
,[object Object], ,[object Object],(d ,[object Object], prompt_text.lower() ,[object Object], d ,[object Object], DANGEROUS):
,[object Object], ,[object Object],(,[object Object],) ,[object Object],
,[object Object], ,[object Object], ,[object Object],What this does: Routes benign prompts to an automatic default but bounces anything that looks consequential to the actual user. The agent stays autonomous for the boring cases and correctly humble for the risky ones.
This split is the whole philosophy. Handling interactive prompts well isn't about answering everything automatically — it's about knowing which questions you're allowed to answer and which ones belong to a human. An agent that suppresses noise and escalates real decisions is one people trust with real work.
How Do You Know a Command Is Stuck on a Prompt?
Prevention and PTY handling cover the prompts you anticipate. The remaining danger is the prompt you didn't — a tool you've never run before that pauses on some question you never mapped. For those, you need to detect a hang while it's happening, not after a timeout expires. The signal is a live process that has produced no output for a while.
[object Object], subprocess, time, select
,[object Object], ,[object Object],(,[object Object],):
p = subprocess.Popen(cmd, shell=,[object Object],, stdout=subprocess.PIPE,
stderr=subprocess.STDOUT, stdin=subprocess.DEVNULL, text=,[object Object],)
last = time.time()
out = []
,[object Object], p.poll() ,[object Object], ,[object Object],:
r, _, _ = select.select([p.stdout], [], [], ,[object Object],)
,[object Object], r:
out.append(p.stdout.readline()); last = time.time()
,[object Object], time.time() - last > idle_limit:
p.terminate()
,[object Object], ,[object Object],.join(out) + ,[object Object],
,[object Object], ,[object Object],.join(out)What this does: Watches the process for output and flags it as stalled if it goes idle for several seconds while still running — the classic signature of a command blocked on a prompt. Instead of a blank timeout, the model gets a specific message telling it the command wanted input, which it can act on.
This idle-detection layer is what makes handling cli agent interactive prompts dependable against the unknown. You suppress what you can predict, and you detect-and-report what you can't, so a novel prompt degrades into a clear signal rather than a silent freeze.
⚡ Pro tip: Tune idle_limit to the command class. A git status that's idle for eight seconds is stuck; a docker build legitimately goes quiet for minutes mid-layer. Pass a per-command idle budget rather than one global number, or you'll kill slow-but-healthy commands and frustrate everyone.
⚡ Pro tip: When you detect a stall, capture the command's last line of output in the message you hand back — that partial text is almost always the prompt itself. "Overwrite existing file? [y/N]" in the stall report tells the model exactly what was being asked, so it can either answer sensibly or explain to the user what the tool wanted.
Variations for Different Contexts
A platform engineer wrapping infrastructure tools sets TF_INPUT=0 and AWS_PAGER="" so Terraform and the AWS CLI never block on prompts or pagers, then handles the rare apply-confirmation through an explicit escalation.
A data scientist running database migrations pipes credentials through environment variables (PGPASSWORD) rather than answering a psql password prompt, keeping secrets out of the agent's reasoning and out of any transcript.
A release engineer automating a deploy checklist uses pexpect for one legacy tool that has no non-interactive mode, with a tightly scoped pattern list and a hard rule that any unmatched prompt aborts the run and pages a human.
The constant across all three: prevent prompts where you can, drive them deliberately where you must, and escalate the ones that carry real risk.
Save and Reuse This
The durable asset here isn't the pexpect code — it's the mapping of commands to their non-interactive forms and the list of prompts your agent knows how to handle safely. That knowledge accumulates painfully, one hung process at a time, and it's exactly the kind of thing that gets lost between projects.
Storing that registry and the escalation rules alongside your prompts in a library like PromptABCD means the next terminal agent you build already knows that apt needs -y and that a [y/N] prompt goes to a human. You solve the interactive-prompt problem once and carry the solution forward, instead of rediscovering every hang from scratch.
Continue Reading
Save the prompts from this post
PromptABCD is a free prompt manager. Paste, organize, and reuse your best AI prompts — no more hunting through chat history.
