Giving a CLI Agent Web Search
An agent confidently recommended a flag removed months earlier. Giving a cli agent web search fixes stale, overconfident answers — here's native vs custom, with citations and domain filtering.
SYSTEM = f"""You are a frontend assistant.
Here are the latest docs as of last Tuesday:
{paste_of_changelog}
{paste_of_api_reference}
""" # stale the moment the next release shipsA frontend team shipped an agent that confidently told a developer to install a package flag that had been removed eight months earlier. The agent wasn't broken. It was answering from a model whose knowledge had a cutoff, with total confidence and zero awareness that the world had moved on. The developer burned an hour debugging a flag that no longer existed. That failure is the whole case for giving a cli agent web search — not as a bonus feature, but as the fix for an agent that doesn't know what it doesn't know.
This is the story of how that team added cli agent web search, the choice they faced between two very different approaches, and what it actually took to make the agent's answers trustworthy again.
The Problem the Frontend Team Faced
Their agent lived in the terminal and answered questions about their stack — build config, dependency choices, framework APIs. It was genuinely useful for stable knowledge. But their ecosystem moved fast: libraries shipped breaking changes monthly, best practices shifted, and APIs got deprecated. The agent's training cutoff meant it was quietly, confidently wrong about anything recent.
The insidious part was the confidence. A model doesn't hedge just because information is stale — it states outdated facts in the same authoritative tone as timeless ones. Developers couldn't tell which answers were current and which were historical fiction. After the removed-flag incident, trust cratered. An agent you have to double-check on the web is an agent that isn't saving you anything.
The Wrong Approach
The team's first fix was to stuff recent docs into the system prompt — paste in the changelog, the latest API reference, the migration guide. It helped for exactly as long as it took those docs to go stale, which wasn't long.
SYSTEM = ,[object Object], ,[object Object],What this does: Injects a snapshot of documentation into the prompt so the model has recent context. The flaw is structural — it's a manual snapshot that someone has to keep refreshing, it bloats every request, and it's outdated the instant a new version lands.
This approach trades one staleness problem for a maintenance treadmill. Someone has to notice a release, find the new docs, and update the prompt — and the moment they miss a beat, the agent is confidently wrong again. Worse, the giant doc paste inflated token costs on every single call, whether or not the question needed recent information.
⚠️ Common mistake: Pasting "current" documentation into a static system prompt as your freshness strategy. It's stale on arrival, expensive on every request, and it fails silently — nobody notices the docs went out of date until the agent gives another wrong answer.
The Right Approach: Let the Agent Search on Demand
The fix was to give the agent the ability to fetch current information itself, only when a question actually needed it. Anthropic's API offers this as a server-side tool you add to the tools array — the model decides when to search, the search runs on Anthropic's side, and results come back with citations.
resp = client.messages.create(
model=,[object Object],,
max_tokens=,[object Object],,
messages=messages,
tools=[{
,[object Object],: ,[object Object],,
,[object Object],: ,[object Object],,
,[object Object],: ,[object Object],,
}],
)What this does: Adds the native web search tool to the request. Claude decides whether a question needs fresh information, runs up to five searches server-side, and returns an answer grounded in current results with source citations attached. You write no search-plumbing code at all.
The change was immediate. "What's the current recommended way to configure the router" now triggered a search and returned an answer citing this month's docs, not last year's. The removed flag never came back as a suggestion, because the agent could see it was gone.
⚡ Pro tip: The model, not your code, decides when to search — so shape that decision in your system prompt. A line like "search the web for any question about library versions, APIs, or recent releases; answer from your own knowledge for stable concepts" keeps it from searching for basics while ensuring it checks anything time-sensitive.
Choosing Between Native and Custom Search
The native tool isn't the only option, and the team weighed both. The alternative is a custom search tool: you define a normal tool, and when the model calls it, your code hits a search API (Brave, Tavily, or similar) and returns the results yourself.
TOOLS = [{
,[object Object],: ,[object Object],,
,[object Object],: ,[object Object],,
,[object Object],: {,[object Object],: ,[object Object],,
,[object Object],: {,[object Object],: {,[object Object],: ,[object Object],}},
,[object Object],: [,[object Object],]},
}]
,[object Object], ,[object Object],(,[object Object],):
hits = brave_client.search(query, count=,[object Object],) ,[object Object],
,[object Object], format_results(hits)[:,[object Object],]What this does: Defines search as a tool you execute, giving you full control over the search backend, result formatting, caching, and cost. The model requests a query; your code runs it however you like and returns the results into the loop.
The tradeoff is real and worth stating plainly. The native tool is effortless — no search backend, citations handled, maintained by Anthropic — but it's billed per search (on the order of ten dollars per thousand searches on top of tokens) and you don't control the underlying engine. A custom tool means you own the plumbing and the bill, can cache aggressively, and can point at an internal search index — but you're now maintaining search infrastructure. For most teams starting out, native wins on simplicity; teams with heavy volume or a private corpus eventually want custom.
⚡ Pro tip: Cap max_uses deliberately. Set it to 1–2 for quick fact-checks and 5+ only for research-style questions. An agent left uncapped can fire a dozen searches on one question, and at per-search pricing that turns a cheap query into a surprising line item.
Results and What Changed
After switching to on-demand search, the team measured the thing that mattered: how often the agent gave a stale-but-confident answer. It went from a recurring weekly complaint to nearly zero. Developers stopped reflexively double-checking the agent's version-specific claims, which is the exact behavior that had made the tool pointless.
Cost went down, not up, despite paying per search — because dropping the giant doc paste from every system prompt saved more tokens than occasional searches cost. The agent now sent a lean prompt and searched only the ~15% of questions that actually needed current data, instead of carrying stale docs on 100% of them.
⚡ Pro tip: Surface citations in the terminal. The native tool returns sources; print them under the answer as "sources: [1] docs.example.com...". In a CLI, a visible citation lets the developer click through and verify in one step — which is what rebuilds trust after a wrong-answer incident.
Keeping Search Trustworthy With Domain Filtering
One more refinement made the difference between "searches the web" and "searches the right web." The team constrained searches to sources they trusted, using the tool's domain filtering so the agent cited official docs rather than random blog posts.
tools=[{
,[object Object],: ,[object Object],,
,[object Object],: ,[object Object],,
,[object Object],: ,[object Object],,
,[object Object],: [,[object Object],, ,[object Object],, ,[object Object],],
}]What this does: Restricts web search to an allowlist of trusted domains, so the agent grounds its answers in authoritative sources instead of whatever ranks well that week. You use allowed_domains or blocked_domains, not both.
For an agent giving technical advice, this mattered enormously. An unfiltered search can surface an outdated Stack Overflow answer or a content-farm tutorial and cite it with the same confidence as the official docs. Pinning the agent to sources the team already trusted turned web search from a possible new source of wrong answers into a reliable one. Newer dated versions of the tool add finer dynamic filtering, but even the basic allowlist covered the team's needs.
When Is Giving a CLI Agent Web Search the Wrong Call?
Honesty about the downside: cli agent web search isn't free, and it isn't always appropriate. Search adds latency — a searched answer is seconds slower than one from the model's own knowledge — and per-search billing means a chatty agent gets expensive. For a tool answering mostly stable questions, defaulting to search on everything is both slower and pricier than it needs to be.
There's also a data-sensitivity angle. A search query leaves your environment, so an agent working on confidential material shouldn't blindly search text that might contain secrets or internal names. For those cases, a custom search tool pointed at an internal index keeps queries in-house, which is one more reason teams graduate from the native tool to their own.
⚡ Pro tip: Cache identical searches within a session. Agents often re-ask the same question across turns of a task, and caching the result for a few minutes cuts both latency and cost with zero downside. A tiny in-memory dict keyed on the query string is usually all it takes.
How to Apply This to Your Situation
The pattern fits any agent whose domain changes faster than a model's training cutoff.
A DevOps engineer gives an incident agent web search scoped to their cloud provider's status page and docs, so "is there a known outage in us-east-1" pulls real-time information mid-incident.
A financial analyst wraps a search tool around a licensed market-data feed rather than the open web, keeping the agent's answers grounded in data the firm is allowed to use and pays for.
A technical writer points the agent's search at the official docs of the tools they document, so drafts cite current APIs and never resurrect a deprecated method from the model's memory.
The recipe is consistent: add search as a tool, let the model decide when to use it, constrain it to trusted sources, cap the uses, and show the citations.
Next Steps
Start with the native tool — it's one entry in your tools array and it solves the staleness problem today. Add a system-prompt rule about when to search, cap max_uses, allowlist your trusted domains, and print citations. If volume or a private corpus later pushes you toward a custom search backend, you'll switch with the loop already in place.
The system-prompt guidance that governs when and where the agent searches is the part you'll tune most, and it's easy to lose between projects. Keeping those search-behavior instructions and domain allowlists in a prompt library like PromptABCD means your next agent inherits a search policy you've already gotten right, instead of relearning the balance between fresh answers and needless searching from scratch.
Continue Reading
Save the prompts from this post
PromptABCD is a free prompt manager. Paste, organize, and reuse your best AI prompts — no more hunting through chat history.
