Building a CLI Agent With Node.js
A case study in building a CLI agent nodejs teams actually install: named tasks over raw shell, confirmation gates on writes, and streaming output. See the wrong way first, then the fix.
import Anthropic from "@anthropic-ai/sdk";
import { exec } from "node:child_process";
const client = new Anthropic();
async function run(goal) {
const msg = await client.messages.create({
model: "claude-sonnet-5",
max_tokens: 1024,
messages: [{ role: "user", content: goal }],
});
// BAD: blindly execute whatever text comes back
exec(msg.content[0].text, (err, stdout) => console.log(stdout));
}Picture this: you're a platform engineer at a mid-size fintech, and every new hire spends their first week drowning in your internal tooling. Which script deploys to staging? Where do the feature flags live? You've written a wiki nobody reads. So you decide to wrap that tribal knowledge in a terminal tool — a CLI agent nodejs project that answers "how do I deploy the payments service" by actually reading your repo and running the safe commands. This is the story of building exactly that, the wrong way first, then the right way.
The keyword here is nodejs on purpose: the JavaScript ecosystem's streaming primitives and package tooling make it a strong fit for a CLI agent nodejs teams can install with a single npm i -g.
The Problem the Platform Engineer Faced
The onboarding wiki was stale the day it shipped. New engineers asked the same five questions in Slack every week, and the answers lived in scripts nobody could find. The engineer wanted a tool where someone could type devhelp "run the payments integration tests" and get either the exact command or, better, have the agent run it and report back.
The constraints were real: it had to install cleanly on macOS and Linux, stream output so long-running commands didn't look frozen, and never — ever — run a destructive command without asking first. A crashed onboarding tool that rm -rf'd someone's node_modules would end the project on day one.
The Wrong Approach
The first version glued a model call to child_process.exec and shipped. It worked in the demo and fell apart in real use.
[object Object], ,[object Object], ,[object Object], ,[object Object],;
,[object Object], { exec } ,[object Object], ,[object Object],;
,[object Object], client = ,[object Object], ,[object Object],();
,[object Object], ,[object Object], ,[object Object],(,[object Object],) {
,[object Object], msg = ,[object Object], client.,[object Object],.,[object Object],({
,[object Object],: ,[object Object],,
,[object Object],: ,[object Object],,
,[object Object],: [{ ,[object Object],: ,[object Object],, ,[object Object],: goal }],
});
,[object Object],
,[object Object],(msg.,[object Object],[,[object Object],].,[object Object],, ,[object Object], ,[object Object],.,[object Object],(stdout));
}What this does: Asks the model for a command and runs its raw text output directly. This is the anti-pattern — there's no tool schema, no allowlist, and the model's prose ("You can run npm test...") gets shoved into a shell verbatim.
Three things broke immediately. The model often replied with explanation around the command, so the shell got garbage. There was no confirmation, so the tool once ran a git reset --hard because a user's question was ambiguous. And output arrived all at once after a 40-second wait, so users assumed it had hung and killed it.
⚠️ Common mistake: Treating the model's free-text reply as a command. Always use structured tool calls so you receive a clean, parsed { cmd: "npm test" } object — never scrape a command out of a paragraph.
The Correct Approach
The rewrite used tool calls, an allowlist with a confirmation gate for writes, and streaming. Here's the core of the working CLI agent nodejs version.
[object Object], ,[object Object], ,[object Object], ,[object Object],;
,[object Object], { execFile } ,[object Object], ,[object Object],;
,[object Object], { promisify } ,[object Object], ,[object Object],;
,[object Object], pexec = ,[object Object],(execFile);
,[object Object], ,[object Object], = ,[object Object], ,[object Object],([,[object Object],, ,[object Object],, ,[object Object],, ,[object Object],, ,[object Object],]);
,[object Object], ,[object Object], = ,[object Object], ,[object Object],([,[object Object],, ,[object Object],]);
,[object Object], tools = [{
,[object Object],: ,[object Object],,
,[object Object],: ,[object Object],,
,[object Object],: {
,[object Object],: ,[object Object],,
,[object Object],: { ,[object Object],: { ,[object Object],: ,[object Object], }, ,[object Object],: { ,[object Object],: ,[object Object], } },
,[object Object],: [,[object Object],],
},
}];What this does: Defines a single run_task tool keyed to named tasks, not raw commands. The model can only ask for tasks you've defined, which collapses the entire "arbitrary shell" risk into a small, auditable map.
[object Object], readline ,[object Object], ,[object Object],;
,[object Object], ,[object Object], ,[object Object],(,[object Object],) {
,[object Object], rl = readline.,[object Object],({ ,[object Object],: process.,[object Object],, ,[object Object],: process.,[object Object], });
,[object Object], ans = ,[object Object], rl.,[object Object],(,[object Object],);
rl.,[object Object],();
,[object Object], ans.,[object Object],().,[object Object],() === ,[object Object],;
}
,[object Object], ,[object Object], ,[object Object],(,[object Object],) {
,[object Object], (,[object Object],.,[object Object],(task)) {
,[object Object], (!(,[object Object], ,[object Object],(,[object Object],)))
,[object Object], ,[object Object],;
}
,[object Object], (!,[object Object],.,[object Object],(task) && !,[object Object],.,[object Object],(task))
,[object Object], ,[object Object],;
,[object Object], { stdout, stderr } = ,[object Object], ,[object Object],(,[object Object],(task), args.,[object Object],(,[object Object],).,[object Object],(,[object Object],));
,[object Object], (stdout || stderr).,[object Object],(,[object Object],, ,[object Object],);
}What this does: Runs known tasks, and for anything that writes, blocks on a y/N confirmation before touching the system. execFile (not exec) avoids shell-injection because arguments are passed as an array, not interpolated into a string.
⚡ Pro tip: Use execFile/spawn with argument arrays instead of exec with a command string. It's not just cleaner — it structurally prevents shell injection, because there's no shell parsing the arguments. This single choice removes an entire class of bugs.
Results and What Changed
After the rewrite, onboarding Slack questions dropped noticeably — the team estimated it saved each new hire two to three hours in their first week, and saved the senior engineers from answering the same questions on repeat. The confirmation gate caught two would-be mistakes in the first month, both write tasks a user triggered by accident.
The streaming fix mattered more than expected. Long tasks now print output as it arrives:
[object Object], stream = client.,[object Object],.,[object Object],({
,[object Object],: ,[object Object],,
,[object Object],: ,[object Object],,
tools,
messages,
});
stream.,[object Object],(,[object Object],, ,[object Object], process.,[object Object],.,[object Object],(t));
,[object Object], final = ,[object Object], stream.,[object Object],();What this does: Streams the model's reasoning token-by-token to the terminal while accumulating the full message for tool dispatch. Users see progress immediately, so nobody kills the process thinking it froze.
⚡ Pro tip: Stream the model's narration even when the real work is a tool call. A line like "Checking the deploy script..." printed before a slow command turns dead air into a sense of progress, and it costs nothing extra.
Handling Errors So the Agent Recovers
The version that survived contact with real users had one property the demo lacked: when a tool failed, the agent adapted instead of crashing. The trick is to treat errors as information for the model, not exceptions that kill the process. When run_task fails, you return the error text as a normal tool result and let the model decide what to do next.
[object Object], ,[object Object], ,[object Object],(,[object Object],) {
,[object Object], {
,[object Object], ,[object Object], ,[object Object],(input);
} ,[object Object], (err) {
,[object Object],
,[object Object], ,[object Object],;
}
}What this does: Converts a thrown exception into a tool result the model reads on its next turn. A failed npm test becomes context the agent can reason about — "the build is broken, here's why" — instead of a stack trace that ends the session.
This changes the whole feel of the tool. A user asks to deploy, the pre-check fails because tests are red, and instead of dying the agent reports why and offers to run the tests. That recovery loop is what people mean when they say an agent feels "smart." It isn't smarter — it's just handed its own failures as input.
⚡ Pro tip: Include a short hint in the error string about what the model can do next ("ask the user" or "try the read-only version"). Models act on error messages far better when the message suggests a path forward rather than just stating what broke.
Making the CLI Agent Auditable
For anything touching a shared environment, you want a record of what actually ran — not what the model said it would do, but what your code executed. A structured log per dispatch turns "the agent did something weird" into a debuggable timeline.
[object Object], { appendFile } ,[object Object], ,[object Object],;
,[object Object], ,[object Object], ,[object Object],(,[object Object],) {
,[object Object], line = ,[object Object],.,[object Object],({ ,[object Object],: ,[object Object], ,[object Object],().,[object Object],(), ...entry }) + ,[object Object],;
,[object Object], ,[object Object],(,[object Object],, line);
}
,[object Object],What this does: Appends a timestamped JSON line for every task dispatch, including rejected ones. When someone asks "why did staging restart at 2 a.m.," the answer is one grep away instead of a guessing game.
The fintech team wired this from day one, and it paid off the first time a user swore the agent "deleted my branch." The audit log showed the agent had asked for confirmation, the user had typed y, and the task was a normal checkout — not a delete at all. Logs end arguments.
⚡ Pro tip: Log the rejected requests too, not just the executed ones. The commands your allowlist blocks are the single best signal for what capability to add next — users are telling you what they wish the tool could do, one denied request at a time.
Cross-Platform Gotchas That Break Node CLIs
The requirement to install cleanly on macOS and Linux hides a few traps that only surface on a colleague's machine. The first is the shebang. Your executable needs #!/usr/bin/env node as its literal first line, and the file has to stay executable — a .gitattributes that normalizes line endings to CRLF will silently break that shebang on Unix, producing a baffling "bad interpreter" error nobody can reproduce on their own laptop.
[object Object],
,[object Object],[object Object], ,[object Object], ,[object Object],[object Object], ,[object Object], ,[object Object],[object Object],
,[object Object],[object Object], ,[object Object], ,[object Object],[object Object], ,[object Object], ,[object Object],
,[object Object],What this does: Registers the global command and pins a minimum Node version. The engines field means someone on Node 16 gets a clear warning at install time instead of a cryptic syntax error when your code uses a newer feature.
Paths are the other quiet killer. Hardcoding / as a separator or assuming $HOME exists breaks the moment someone runs your tool on Windows. Use node:path and os.homedir() rather than string-building paths by hand. None of this is hard, but each one produces a bug report that reads "works on mine" — the most expensive kind to chase down.
How to Apply This to Your Situation
The pattern generalizes well beyond fintech onboarding. Three variations from teams I've compared notes with:
A QA lead at a games studio built a CLI agent nodejs tool that maps English test descriptions to named Playwright suites, with a confirmation before anything touches the shared test environment.
A site reliability engineer wrapped runbooks: "restart the cache tier" maps to a named task with a mandatory confirm and an audit log entry, so incident actions are both fast and traceable.
A technical writer — not even an engineer — uses one to run the docs build and preview locally, because "the deploy preview command" was the one thing she could never remember.
The shared recipe: named tasks over raw commands, confirmation gates on writes, streaming for feedback. Start there and adapt the task map to your stack.
Next Steps
Wrap the whole thing as a global npm binary with a bin entry in package.json, add a --dry-run flag that prints the task it would run, and log every dispatch to a local file for audit. Each of those is an afternoon of work on top of the core loop.
From there, the highest-value addition is usually a small set of named tasks that cover your team's real questions rather than a big generic toolset. The fintech team's tool did four things well — check deploy status, run integration tests, tail service logs, and roll back — and that focus is exactly why people trusted it. A CLI agent nodejs project that does four things reliably beats one that does forty things unpredictably, because trust is the whole product. Add tasks in response to the rejected-request log, not in anticipation of what someone might want, and the toolset stays tight and the confidence stays high.
The trickiest part to maintain over time isn't the code — it's the prompts. Your tool descriptions and the task-mapping instructions in your system prompt need tuning as the model and your tooling evolve. Storing those prompts in a reusable library like PromptABCD means when you build the next internal agent, you start from the tool schemas and system prompts you already refined here, instead of rewriting them from memory. The code is the easy part to rebuild; the tuned prompts are what you don't want to lose.
Continue Reading
Save the prompts from this post
PromptABCD is a free prompt manager. Paste, organize, and reuse your best AI prompts — no more hunting through chat history.
