Parsing Natural Language Into CLI Commands
Turn natural language to CLI actions the safe way — map requests to typed intents instead of generating shell strings. Copy-paste patterns for a terminal front end you can trust.
import anthropic, json
client = anthropic.Anthropic()
INTENTS = {
"list_deploys": {"limit": "int", "status": "str"},
"show_logs": {"service": "str", "since": "str"},
"disk_usage": {"path": "str"},
}
SYSTEM = f"""Map the user's request to exactly one intent.
Valid intents and their fields: {json.dumps(INTENTS)}
Reply ONLY with JSON: {{"intent": "...", "args": {{...}}}}.
If nothing fits, use {{"intent": "none"}}."""
def parse(nl_request):
resp = client.messages.create(
model="claude-sonnet-5", max_tokens=300, system=SYSTEM,
messages=[{"role": "user", "content": nl_request}],
)
return json.loads(resp.content[0].text)Most natural-language-to-CLI guides are wrong about the core step. They tell you to prompt a model with "convert this English request into a shell command" and run whatever string comes back. That approach feels obvious and it's a security incident waiting to happen. The reliable way to handle natural language to CLI translation isn't to generate a command string at all — it's to map the request onto a small, structured set of intents you already trust. This guide shows you that pattern, copy-paste ready.
By the end you'll have a terminal tool that turns "show me the last five failed deploys" into a safe, parameterized action — no raw shell generation anywhere in the pipeline.
Quick-Start (Copy This Right Now)
Here's the whole idea in one function: constrain the model to a fixed menu of intents plus typed arguments.
[object Object], anthropic, json
client = anthropic.Anthropic()
INTENTS = {
,[object Object],: {,[object Object],: ,[object Object],, ,[object Object],: ,[object Object],},
,[object Object],: {,[object Object],: ,[object Object],, ,[object Object],: ,[object Object],},
,[object Object],: {,[object Object],: ,[object Object],},
}
SYSTEM = ,[object Object],
,[object Object], ,[object Object],(,[object Object],):
resp = client.messages.create(
model=,[object Object],, max_tokens=,[object Object],, system=SYSTEM,
messages=[{,[object Object],: ,[object Object],, ,[object Object],: nl_request}],
)
,[object Object], json.loads(resp.content[,[object Object],].text)What this does: Forces the model to pick from a fixed set of intents and return typed arguments as JSON, instead of emitting an executable command. The model becomes a classifier, not a code generator — a much smaller and safer job.
Now you dispatch that structured result through code you control, where every intent maps to a hand-written, parameterized command.
Understanding the Variables
The INTENTS map is the entire contract. Each key is an action your tool supports; each value describes the arguments and their types. This is the menu the model orders from, and it can't order off-menu.
SYSTEM pins the model to that menu and to a strict output format. The instruction to reply only with JSON matters — any prose wrapping the JSON will break your parser, so you say it plainly and you validate afterward.
The return value is a plain dict: {"intent": "show_logs", "args": {"service": "api", "since": "1h"}}. Structured, typed, and trivial to validate before anything runs. This is the heart of turning natural language to CLI actions without ever trusting a generated string.
⚡ Pro tip: Add an "explanation" field to the required JSON so the model returns why it chose that intent. You can print it back to the user ("Interpreting as: show api logs from the last hour") and let them catch a misread before you execute.
How Do You Turn Natural Language to CLI Commands Reliably?
Validate the model's output against your schema, then build the real command yourself. Never let the model's text near a shell.
[object Object], ,[object Object],(,[object Object],):
intent = parsed.get(,[object Object],)
,[object Object], intent ,[object Object], ,[object Object], INTENTS:
,[object Object], ValueError(,[object Object],)
args = parsed.get(,[object Object],, {})
,[object Object], field, typ ,[object Object], INTENTS[intent].items():
,[object Object], field ,[object Object], args ,[object Object], typ == ,[object Object],:
args[field] = ,[object Object],(args[field]) ,[object Object],
,[object Object], intent == ,[object Object],:
,[object Object], [,[object Object],, ,[object Object],, ,[object Object],, ,[object Object],(args.get(,[object Object],, ,[object Object],)),
,[object Object],, args.get(,[object Object],, ,[object Object],)]
,[object Object], intent == ,[object Object],:
,[object Object], [,[object Object],, args[,[object Object],], ,[object Object],, args.get(,[object Object],, ,[object Object],)]What this does: Turns the validated intent into a real argument array using templates you wrote. The model chose the intent; your code owns every character that reaches the system. Type coercion here also catches a model that returns "5" where you need an integer.
The step-by-step for any request: parse to intent, validate against schema, build the command from your template, optionally confirm with the user, then execute with subprocess.run(cmd_array). Five stages, and the model only touches the first one.
⚠️ Common mistake: Asking the model to output the final shell command and running it. Even with a careful prompt, a request like "clean up old logs" can produce rm -rf /var/log/*. Constrain to intents and the model literally cannot express a destructive action you didn't define.
Pro-Level Variations
Once the intent classifier works, teams extend it in ways that fit their domain.
A network engineer maps intents to read-only show commands on switches — "is port 24 up on the core switch" becomes a structured query, and configuration-changing verbs simply aren't in the intent map.
A data analyst at a retail company points intents at a SQL query builder: "revenue by region last quarter" maps to a parameterized query template, so the model picks the shape but never writes raw SQL that could touch the wrong table.
A support engineer wires intents to a ticketing CLI — "escalate the Johnson ticket to tier two" becomes {"intent": "escalate", "args": {...}}, validated against real ticket IDs before anything changes.
The pattern scales because adding a capability means adding one intent and one template, not loosening the security model.
⚡ Pro tip: When a request maps to {"intent": "none"}, don't fail silently — echo back the menu of things the tool can do. Users learn your vocabulary fast when the tool tells them its verbs on a miss.
Confirming Before You Execute
Structured intents make confirmation trivial and worth doing. Because you know the exact intent and arguments before running anything, you can show the user precisely what's about to happen in plain terms — not a cryptic command, but a sentence they can approve or reject.
[object Object], ,[object Object],(,[object Object],):
i, a = parsed[,[object Object],], parsed.get(,[object Object],, {})
,[object Object], i == ,[object Object],:
,[object Object], ,[object Object],
,[object Object], i == ,[object Object],:
,[object Object], ,[object Object],
,[object Object], ,[object Object],
,[object Object], ,[object Object],(,[object Object],):
parsed = parse(nl_request)
,[object Object],(,[object Object],)
,[object Object], ,[object Object], auto ,[object Object], ,[object Object],(,[object Object],).lower() != ,[object Object],:
,[object Object],
execute(build_command(parsed))What this does: Renders the parsed intent as a readable sentence and asks for approval before executing. The user confirms an interpretation, not a command string — which is exactly the review step that catches a misread before it does damage.
This is where the intent approach pays a second dividend. With raw command generation, a confirmation prompt shows the user a shell string they may not understand well enough to judge. With intents, you show them "Roll back the payments service to the previous version" — a decision anyone can make correctly.
⚡ Pro tip: Offer an --auto flag that skips confirmation for read-only intents but never for write intents. Frequent, safe queries shouldn't nag; irreversible actions always should. Tie the gate to the intent's category, not to a global setting.
Handling Ambiguity and Multiple Intents
Real requests are messy. "Fix the api service" could mean restart it, roll it back, or check its logs. When the model isn't confident, the worst outcome is a confident wrong guess. Build a clarification path for exactly this.
SYSTEM_MULTI = SYSTEM + ,[object Object],
,[object Object], ,[object Object],(,[object Object],):
parsed = parse(nl_request)
,[object Object], parsed[,[object Object],] == ,[object Object],:
,[object Object],(parsed[,[object Object],])
,[object Object], n, opt ,[object Object], ,[object Object],(parsed[,[object Object],], ,[object Object],):
,[object Object],(,[object Object],)
,[object Object], ,[object Object],
execute(build_command(parsed))What this does: Lets the model signal uncertainty explicitly instead of forcing a pick. An ambiguous "fix the api" comes back as a question with options, so the user disambiguates in one keystroke rather than discovering the wrong action after it ran.
Adding two or three worked examples to your system prompt — a sample request paired with its correct intent JSON — sharpens classification more than any amount of prose instruction. Models pattern-match from examples far better than they follow abstract rules, and a few-shot intent classifier is noticeably more reliable than a zero-shot one.
⚡ Pro tip: Log every "clarify" response and every misclassification you catch. Those are your training signal — the requests that map cleanly to new few-shot examples, which you fold back into the system prompt to close the gap for next time.
Troubleshooting Common Issues
If the model returns JSON wrapped in markdown code fences, strip those fences before parsing, or instruct the model more firmly to emit bare JSON. If it invents an intent, your schema check catches it — return the error to the model and let it retry once.
If arguments come back as the wrong type, coerce and validate in build_command rather than trusting the model. And if the same request keeps mapping to the wrong intent, your intent descriptions are ambiguous — sharpen them, because the fix is almost always in the schema, not the prompt.
Keeping the Intent Map Maintainable
An intent map grows, and a sprawling one gets its own problems. Two intents start overlapping, the model can't tell them apart, and classification accuracy quietly drops. Treat the intent map like the API it actually is.
Version it. When you rename or remove an intent, keep the old name mapped to a "this moved" response for a release or two, so users and scripts that learned the old vocabulary get a helpful redirect instead of a silent failure. Test it: keep a small file of example requests paired with the intent each should produce, and run your classifier against it whenever you change the schema or the system prompt. That regression suite catches the case where sharpening one intent's description accidentally breaks another.
And prune it. An intent nobody has triggered in months is a candidate for removal — every extra option is one more thing the model has to disambiguate on every request, which slightly dulls its accuracy on the intents that matter. A tight map of a dozen well-separated intents classifies more reliably than a bloated map of fifty fuzzy ones. Turning natural language to CLI actions stays dependable only when the menu stays clean.
How Do You Know the Classifier Is Accurate Enough?
You measure it, and the way you measure it should reflect an asymmetry most people miss. Not every wrong answer costs the same. A classifier that says "clarify" when it should have been confident is mildly annoying. A classifier that confidently fires roll_back when the user meant show_logs is a production incident. So the metric that matters isn't raw accuracy — it's how often the tool takes a wrong action without asking first.
Build a small labeled set: thirty to fifty real requests, each tagged with the intent it should produce. Run the classifier over them after every prompt change and count three buckets separately — correct, wrongly-clarified (a false hesitation), and wrongly-executed (a confident miss). The last bucket is the one you drive toward zero, even if it means tolerating a few more clarifications. Tuning the system prompt to prefer "clarify" on any doubt trades a little friction for a lot of safety, and that's almost always the right trade for a tool that touches real systems.
Track accuracy per intent, not just overall. An aggregate of "94% correct" hides the fact that one dangerous intent is your worst performer. When you find the weak one, the fix is usually a sharper description and one or two more few-shot examples aimed squarely at the confusion. This per-intent view is the single most useful habit I've picked up building these — it turns a vague sense that "the model gets it wrong sometimes" into a specific, fixable list.
Your Turn
Take three commands you run constantly, write them as intents with typed arguments, and wire up the parse-validate-build-execute chain. You'll have a working natural-language front end to your own tooling in an afternoon.
The intent schemas and the system prompt that drives them are the assets worth keeping — they encode exactly how your team talks about its tools. Storing them in a prompt library like PromptABCD means the next tool you build reuses the same vocabulary and the same carefully-worded classifier prompt, instead of starting the mapping from zero.
Continue Reading
Save the prompts from this post
PromptABCD is a free prompt manager. Paste, organize, and reuse your best AI prompts — no more hunting through chat history.
