PromptABCD
FeaturesLearnGuideBlogContext Blocks
Sign inGet started free
Sign inSign up
PromptABCD

A calm home for your best AI prompts. Save them once, find them in seconds, reuse them forever.

Product

  • Features
  • Chrome Extension
  • Free Courses
  • How it works
  • Use cases
  • Blog
  • Context Blocks
  • Export Anywhere
  • FAQ

Resources

  • User guide
  • Learn prompting
  • Sign in
  • Get started free

© 2026 PromptABCD. All rights reserved.

AboutPrivacy PolicyTerms and Conditions
Home/Blog/CLI AI Agents/Testing a CLI AI Agent
CLI AI Agents

Testing a CLI AI Agent

"You can't test AI, it's non-deterministic" is wrong. When you test a CLI AI agent you're testing the harness, not the model — here's the two-layer strategy that stops shipping broken agents.

September 17, 2026·9 min read
ShareShare
⚡Featured Prompt— copy and use right now
def test_agent_lists_files():
    result = agent("list the python files")
    assert result == "Found 3 files: a.py, b.py, c.py"   # brittle, fails constantly

Most advice about testing AI tools starts with a surrender: "you can't really test it, the model is non-deterministic." That's wrong, and believing it is why so many agents ship untested and break in embarrassing ways. When you test a CLI AI agent, you're mostly not testing the model at all — you're testing the deterministic harness around it: the loop, the tool dispatch, the parsing, the guardrails. All of that is perfectly testable with ordinary unit tests. This is the story of a team that learned to separate the two and stopped shipping broken agents.

The reframe is the whole insight. "Non-deterministic" describes the model. It does not describe your code, and your code is where almost all the bugs live.

The Problem the Team Faced

A four-person developer-tools startup shipped a coding agent that kept breaking in ways that felt random. A release would fix one thing and silently break tool dispatch. An edge case in argument parsing would slip through and corrupt a file. Every incident felt unpredictable, so the team treated the whole agent as untestable and relied on manual spot-checks before each release.

The spot-checks missed things constantly, because a human clicking through a few flows can't cover the combinatorial space of tool calls, error paths, and malformed inputs. Confidence before each release was low, and the fear of touching the loop grew until refactoring felt too dangerous to attempt. They had convinced themselves the model's randomness made testing pointless, when the actual bugs were in boring, deterministic code.

The Wrong Approach

Their first stab at testing tried to assert on the model's output directly, and it failed exactly the way everyone predicts.

python
[object Object], ,[object Object],():
    result = agent(,[object Object],)
    ,[object Object], result == ,[object Object],   ,[object Object],

What this does: Runs the full agent against a live model and asserts on its exact words. This is the test everyone writes first and abandons first — the model phrases things differently every run, so the assertion fails on wording even when the behavior is correct.

This approach conflates two questions that need separating: "does my harness work" (deterministic, testable) and "does the model behave well" (probabilistic, needs a different tool). By trying to pin the model's exact output, they got flaky tests that failed for the wrong reasons, which taught them the false lesson that testing was hopeless. They deleted the tests and went back to manual checks.

⚠️ Common mistake: Writing tests that assert on the model's exact output. Model responses vary by design, so these tests fail on harmless wording changes and train you to ignore or delete them. Test the deterministic behavior of your harness; evaluate the model's quality separately with a different method.

The Right Approach: Two Layers

The fix was to split testing into two layers that use completely different techniques. Layer one unit-tests the harness with a fake model, deterministically. Layer two evaluates real model behavior with pass rates, not assertions.

The first layer replaces the model with a scripted stand-in, so the loop's logic becomes fully testable.

python
[object Object], ,[object Object],:
    ,[object Object], ,[object Object],(,[object Object],):
        ,[object Object],.script = ,[object Object],(script)   ,[object Object],
        ,[object Object],.calls = ,[object Object],
    ,[object Object], ,[object Object],(,[object Object],):
        ,[object Object],.calls += ,[object Object],
        ,[object Object], ,[object Object],.script.pop(,[object Object],)

,[object Object], ,[object Object],():
    fake = FakeModel([
        tool_use(,[object Object],, {,[object Object],: ,[object Object],}),   ,[object Object],
        text(,[object Object],),                      ,[object Object],
    ])
    result = run_loop(,[object Object],, client=fake)
    ,[object Object], fake.calls == ,[object Object],                ,[object Object],
    ,[object Object], ,[object Object], ,[object Object], result              ,[object Object],

What this does: Feeds the loop a scripted sequence of model responses instead of a live model, making the harness behavior deterministic and assertable. You test that the loop dispatches the tool, appends the result, and stops correctly — real logic, real bugs caught, zero flakiness and zero API cost.

The second layer handles the model's actual behavior, and here you accept probability instead of fighting it. You run a set of realistic tasks against the real model and measure how often it does the right thing.

python
[object Object], ,[object Object],(,[object Object],):
    passed = ,[object Object],(,[object Object], ,[object Object], c ,[object Object], cases ,[object Object], check(agent(c[,[object Object],]), c[,[object Object],]))
    rate = passed / ,[object Object],(cases)
    ,[object Object], rate >= threshold, ,[object Object],

What this does: Runs the agent against many real cases and asserts on the pass rate, not any single output. A model that gets 92% of cases right passes; a regression that drops it to 70% fails the build. You're measuring behavior statistically, which is the honest way to test something probabilistic.

Results and What Changed

Once the team could test a CLI AI agent this way, the "random" breakage stopped being random. The harness unit tests — dozens of them, running in under a second with no API calls — caught the tool-dispatch and parsing bugs that had been slipping through, immediately and on every commit. Refactoring the loop stopped being scary because a green test suite proved the behavior held.

The eval layer caught a different class of problem: a prompt change that improved one task while quietly degrading three others. Before, that would have shipped and generated support tickets. Now the pass rate dropped in CI and the change got fixed first. The team's release confidence went from "cross our fingers" to "the suite is green," and their incident rate fell accordingly.

⚡ Pro tip: Run harness unit tests on every commit and evals on a schedule or before releases. Unit tests are instant and free, so gate every push on them. Evals cost real API calls and take minutes, so run them nightly and pre-release rather than on every tiny change — you get fast feedback where it's cheap and thorough feedback where it counts.

⚡ Pro tip: Test your tool executors in complete isolation, with no model involved at all. The function that runs a shell command or reads a file is ordinary code — feed it a malformed path, an oversized output, a blocked command, and assert it handles each. Most "agent" bugs are really tool-executor bugs wearing a costume.

How Do You Build a Good Eval Set to Test a CLI AI Agent?

The eval layer is only as good as its cases, and building those cases well is a skill of its own. A weak eval set — three happy-path tasks the agent obviously handles — tells you nothing and passes forever. A good one is deliberately stocked with the situations that actually break agents.

Pull real cases from three sources. First, past failures: every bug report and every incident becomes a permanent eval case, so a problem you fixed can never silently return. Second, edge cases: empty results, enormous outputs, ambiguous requests, malformed inputs — the situations that expose brittle handling. Third, representative real usage: the actual questions users ask, so your pass rate reflects genuine experience rather than a curated demo.

python
CASES = [
    {,[object Object],: ,[object Object],,           ,[object Object],
     ,[object Object],: {,[object Object],: ,[object Object],}},
    {,[object Object],: ,[object Object],,
     ,[object Object],: {,[object Object],: ,[object Object],, ,[object Object],: ,[object Object],}},
    {,[object Object],: ,[object Object],,          ,[object Object],
     ,[object Object],: {,[object Object],: ,[object Object],}},
]

What this does: Defines eval cases that check behaviors — asks for confirmation, handles a missing file, picks the right tool — rather than exact wording. Each encodes a property the agent must hold, so the check survives the model phrasing its response however it likes.

When you test a CLI AI agent this way, the eval set becomes a living specification of what "working" means, and it grows every time reality surprises you. A set that starts at ten cases and reaches a hundred over a year is a hundred ways the agent is now guaranteed to behave, which is worth far more than any single passing run.

⚡ Pro tip: Check behavioral properties, not string equality. Assert that the agent called the right tool, asked for confirmation, or produced valid output — not that it said an exact sentence. Behavior-based checks are stable across model updates and rephrasings, which is the entire reason the eval layer is worth maintaining.

How to Apply This to Your Situation

The two-layer strategy fits any agent, and the split is always the same.

A fintech developer unit-tests that the confirmation gate blocks every destructive tool deterministically, then evals whether the model correctly chooses safe actions across a suite of ambiguous requests — separating "does the guard work" from "does the model behave."

A DevOps engineer mocks the model to test that the loop retries a failed tool the right number of times, then runs a nightly eval checking the agent picks the correct diagnostic commands for a set of simulated incidents.

A QA engineer builds a golden set of tricky inputs — empty results, huge outputs, malformed tool arguments — as deterministic harness tests, and maintains a separate behavioral eval that tracks the agent's task success rate release over release.

The pattern holds regardless of domain: unit-test the deterministic machine, eval the probabilistic model, and never mix the two.

⚡ Pro tip: Pin a fixed seed or temperature where your SDK allows it when running evals, to cut noise in the pass rate. You can't make a model fully deterministic, but reducing sampling variance makes a dropped pass rate more likely to signal a real regression than random chance — which means you chase fewer phantom failures and trust a red build when you see one.

Next Steps

Start with the cheap, high-value layer: write a FakeModel, and unit-test your loop's dispatch, stop condition, and error handling. That alone catches most of the bugs that make agents feel unreliable, and it runs in milliseconds. Add an eval set of ten to twenty real tasks next, and track the pass rate as a release gate.

The eval cases — the tasks paired with what a good response looks like — are a genuine asset that grows more valuable with every regression it catches. Keeping those cases and the checking prompts in a library like PromptABCD means your test suite for the next agent starts from real, battle-tested scenarios instead of a blank file, and the behaviors you learned to guard against stay guarded.

cli agentstestingevalsai agentsqualityci

Continue Reading

Managing Reusable Prompts for Terminal Workflows
CLI AI Agents

Managing Reusable Prompts for Terminal Workflows

Retyping your best prompt from memory loses its refinements every time. Managing cli agent reusable prompts as named, parameterized, versioned assets keeps the prompt quality you earned — and lets you share it.

September 19, 2026·9 min read
Distributing System Prompts With Your CLI Tool
CLI AI Agents

Distributing System Prompts With Your CLI Tool

Hardcoding your agent's system prompt as a string is the wrong place for it. Treating cli agent system prompt distribution as content — versioned, overridable, updatable — is how prompts evolve independently of code.

September 19, 2026·9 min read
Building a Plugin System for Your CLI Agent
CLI AI Agents

Building a Plugin System for Your CLI Agent

How do you let people add tools to your agent without forking it? A cli agent plugin system lets users extend the agent with their own tools. Here's how to rebuild a hardcoded tool list into a real plugin system.

September 19, 2026·8 min read

Save the prompts from this post

PromptABCD is a free prompt manager. Paste, organize, and reuse your best AI prompts — no more hunting through chat history.

Start free →
← PreviousCLI Agent Config: Environment Variables and SecretsNext →Packaging a CLI Agent for Distribution
Share this post:
ShareShare