PromptABCD
FeaturesLearnGuideBlogContext Blocks
Sign inGet started free
Sign inSign up
PromptABCD

A calm home for your best AI prompts. Save them once, find them in seconds, reuse them forever.

Product

  • Features
  • Chrome Extension
  • Free Courses
  • How it works
  • Use cases
  • Blog
  • Context Blocks
  • Export Anywhere
  • FAQ

Resources

  • User guide
  • Learn prompting
  • Sign in
  • Get started free

© 2026 PromptABCD. All rights reserved.

AboutPrivacy PolicyTerms and Conditions
Home/Blog/CLI AI Agents/Rate Limiting and Retries in a CLI Agent
CLI AI Agents

Rate Limiting and Retries in a CLI Agent

By the time a rate-limit error reaches your code, the SDK already retried it. Good cli agent retry logic tells a 429 from a 529, honors retry-after, and avoids the double-retry deadlock.

September 17, 2026·9 min read
ShareShare
⚡Featured Prompt— copy and use right now
import anthropic, time, random

client = anthropic.Anthropic(max_retries=0)   # we own the retry loop

def call(**kwargs):
    delay = 1.0
    for attempt in range(5):
        try:
            return client.messages.create(**kwargs)
        except anthropic.RateLimitError as e:          # 429: your account's limit
            wait = float(e.response.headers.get("retry-after", delay))
            time.sleep(wait)                            # honor the server exactly
        except anthropic.APIStatusError as e:          # 529/503/500: server-side
            if e.status_code in (529, 503, 500):
                time.sleep(delay + random.uniform(0, delay))  # backoff + jitter
                delay = min(delay * 2, 60)
            else:
                raise
    raise RuntimeError("Exhausted retries")

Here's something that surprises almost everyone building their first resilient agent: by the time a rate-limit error reaches your code, the SDK has usually already retried it three times and failed. The official Anthropic SDKs auto-retry transient failures twice by default with exponential backoff, so a 429 that surfaces to you has burned real time and, in the case of rate limits, real money — you're billed for failed 429 requests. That single fact reshapes how you should write cli agent retry logic: the goal isn't to retry harder, it's to retry correctly, and to know when retrying is the wrong move entirely.

This guide gives you retry logic that tells the two failure types apart, honors the server's instructions, and avoids the deadlock trap that turns a soft limit into a hard outage.

Quick-Start (Copy This Right Now)

The core insight is that a 429 and a 529 mean opposite things and need opposite handling. Here's a wrapper that treats them differently.

python
[object Object], anthropic, time, random

client = anthropic.Anthropic(max_retries=,[object Object],)   ,[object Object],

,[object Object], ,[object Object],(,[object Object],):
    delay = ,[object Object],
    ,[object Object], attempt ,[object Object], ,[object Object],(,[object Object],):
        ,[object Object],:
            ,[object Object], client.messages.create(**kwargs)
        ,[object Object], anthropic.RateLimitError ,[object Object], e:          ,[object Object],
            wait = ,[object Object],(e.response.headers.get(,[object Object],, delay))
            time.sleep(wait)                            ,[object Object],
        ,[object Object], anthropic.APIStatusError ,[object Object], e:          ,[object Object],
            ,[object Object], e.status_code ,[object Object], (,[object Object],, ,[object Object],, ,[object Object],):
                time.sleep(delay + random.uniform(,[object Object],, delay))  ,[object Object],
                delay = ,[object Object],(delay * ,[object Object],, ,[object Object],)
            ,[object Object],:
                ,[object Object],
    ,[object Object], RuntimeError(,[object Object],)

What this does: Sets the client's own max_retries to zero so it doesn't retry underneath you, then handles the two error classes distinctly — a 429 waits exactly as long as the retry-after header says, while a 529 uses exponential backoff with jitter. One retry layer, two correct behaviors.

Wire this around every model call and your agent stops falling over on the transient failures that are a normal part of production traffic.

Understanding the Variables

The max_retries=0 on the client is the most important line, and it's the one people miss. The SDK retries by default. If you also wrap calls in your own retry loop without disabling that, you get two retry layers fighting each other — nested bursts that exhaust your request budget faster and can deadlock a rate-limited account. Own the loop or configure the SDK's, never both.

The retry-after header is the server telling you exactly how long to wait on a 429. It's not a suggestion to improve upon — waiting less just wastes a request and extends your limit window, and waiting on your own guessed schedule instead ignores the one piece of ground truth you were handed. Read it and obey it.

The jitter — that random.uniform(0, delay) — matters specifically for 529s. When Anthropic's infrastructure is overloaded, every client hitting it is retrying at once. If they all back off on the same fixed schedule, they retry in synchronized waves that keep the service saturated. Randomizing each client's delay spreads the load and helps the system recover.

⚡ Pro tip: Never retry a spend or credit-cap error, even though it also arrives as a 429-family failure. Those don't clear on their own — no amount of waiting refills an empty account. Read the error body, and if the problem is billing rather than rate, fail fast with a clear message instead of sleeping in a doomed loop.

Why Does cli agent retry logic Need to Distinguish 429 From 529?

Because the correct fix for each is the opposite of the other, and treating them the same guarantees you handle one of them wrong. A 429 is your account crossing a rate limit — requests, input tokens, or output tokens per minute. It's your responsibility, it comes with a retry-after, and if you keep hitting it, retrying won't help; you have a capacity problem that needs a higher tier, prompt caching, or lower concurrency.

A 529 is Anthropic's infrastructure being overloaded across all users. It has nothing to do with your account, there's no reliable retry-after, and those requests don't count against your billing. The right response is backoff with jitter and, if it persists, failing over to the same model on another provider like Bedrock or Vertex — because more retries against an overloaded endpoint won't help, but a different endpoint might.

python
[object Object], ,[object Object],(,[object Object],):
    ,[object Object], ,[object Object],(err, anthropic.RateLimitError):
        ,[object Object], ,[object Object],   ,[object Object],
    ,[object Object], ,[object Object],(err, ,[object Object],, ,[object Object],) == ,[object Object],:
        ,[object Object], ,[object Object], ,[object Object],
    ,[object Object], ,[object Object],

What this does: Labels a failure so your agent can respond with the right strategy — throttle back on a 429, ride out or fail over on a 529. Logging this classification also tells you, over time, whether your pain is self-inflicted (429s) or upstream (529s), which are two very different problems to fix.

⚠️ Common mistake: Treating every 4xx and 5xx from the API the same way — one retry loop, one backoff, one bucket. A 429 and a 529 have different headers, different meanings, and opposite fixes. Lumping them together means you'll retry a billing error forever, ignore a retry-after, or hammer an overloaded server when you should fail over. Branch on the error type.

Step-by-Step: Building the Full Retry Path

Start by disabling SDK retries and owning one loop. Then, for each call: attempt it; on a 429, read retry-after and sleep exactly that long; on a 529 or 5xx, sleep a jittered exponential backoff and try again; on anything else, raise immediately because it's not transient. Cap total attempts at four or five so a genuinely stuck request fails cleanly instead of looping.

The order of those branches is deliberate — check the specific rate-limit type before the generic status-code branch, so a 429 gets its header-honoring path rather than falling into blind backoff. And cap your maximum delay (60 seconds is reasonable) so an absurd retry-after value can't park your agent for ten minutes.

⚡ Pro tip: Enable prompt caching on your long system prompt and any stable context. Cached tokens don't count toward your input-tokens-per-minute limit on most models, which can lift your effective throughput several times over. For a rate-limited agent, caching is often a bigger win than any retry tuning — it stops you hitting the 429 in the first place.

⚡ Pro tip: Emit a metric for every retry — the error type, the attempt number, and the final outcome. Most 529 storms are short, but you only know that if you can graph them, and a rising 429 rate is an early warning that you need to tier up before it becomes an outage.

Should You Fail Over to Another Provider?

For a 529, the honest answer is often yes. When Anthropic's infrastructure is saturated, retrying the same endpoint harder does nothing — the capacity isn't there for anyone. The same Claude model is available through AWS Bedrock and Google Vertex, which draw on separate capacity pools, so a failover can succeed exactly when the direct API can't.

python
[object Object], ,[object Object],(,[object Object],):
    ,[object Object],:
        ,[object Object], call(**kwargs)                      ,[object Object],
    ,[object Object], anthropic.APIStatusError ,[object Object], e:
        ,[object Object], e.status_code == ,[object Object],:
            ,[object Object], call_bedrock(**kwargs)          ,[object Object],
        ,[object Object],

What this does: Falls over to the same model on a secondary provider when the primary returns an overloaded error, using a capacity pool that isn't affected by the primary's saturation. For a 429 you would not do this — a rate limit follows your account, not the endpoint — which is one more reason to branch on error type.

The tradeoff is added complexity and a second set of credentials to manage, so failover earns its place in production agents that can't tolerate a brief outage and skips the ones where "try again in a minute" is a fine answer.

⚡ Pro tip: Only fail over on 529, never on 429. A rate limit travels with your account, so a different provider that bills the same account may share the quota — and even when it doesn't, you're papering over a capacity problem you should fix at the source. Failover is for upstream overload, not your own limits.

Pro-Level Variations

Teams adapt retry logic to their risk and scale.

A backend developer running an interactive agent keeps retries short and few — an interactive user would rather see "the service is busy, try again" after ten seconds than wait through a minute of silent backoff, so they cap attempts low and surface failures fast.

A data engineer running a large batch job over thousands of items does the opposite: long backoff, provider failover to Bedrock on sustained 529s, and a per-workspace concurrency limit so the batch can't starve the company's interactive traffic into a self-inflicted 429 storm.

A platform team serving many internal users wraps every agent call in a shared client with tuned max_retries, centralized metrics, and a circuit breaker that stops sending requests entirely for a cooldown after repeated failures — protecting both their users and the upstream API from a thundering herd.

Each is the same core logic with the knobs turned to match latency tolerance and volume.

Troubleshooting Common Issues

If your agent seems to hang for exactly two minutes on failures, you've likely stacked your retry loop on top of the SDK's and the delays are compounding — disable one layer. If you keep hitting 429s no matter how you tune backoff, retries aren't your fix; you're over your tier's throughput, so cache, reduce concurrency, or upgrade. If a billing error loops forever, you're retrying a spend-cap 429 that will never clear — inspect the error body and bail on non-transient causes.

And if 529s dominate during peak hours, that's upstream saturation you can't tune away — add provider failover so the same model on Bedrock or Vertex catches the overflow.

Your Turn

Set your client's max_retries to zero, write one retry loop that branches on 429 versus 529, honor retry-after, add jitter to backoff, and cap attempts. Then add prompt caching to stop hitting limits in the first place. You'll have an agent that survives the transient failures that are a normal, expected part of running against any large API.

The retry policy — the branch logic, the caps, the failover rules — is the same across every agent you build, and the system-prompt notes about how the agent should behave when throttled belong with it. Keeping that policy and those prompts in a library like PromptABCD means your next agent inherits battle-tested cli agent retry logic instead of rediscovering the 429-versus-529 distinction during its first production incident.

cli agentsrate limitingretriesai agentsreliabilityapi

Continue Reading

Managing Reusable Prompts for Terminal Workflows
CLI AI Agents

Managing Reusable Prompts for Terminal Workflows

Retyping your best prompt from memory loses its refinements every time. Managing cli agent reusable prompts as named, parameterized, versioned assets keeps the prompt quality you earned — and lets you share it.

September 19, 2026·9 min read
Distributing System Prompts With Your CLI Tool
CLI AI Agents

Distributing System Prompts With Your CLI Tool

Hardcoding your agent's system prompt as a string is the wrong place for it. Treating cli agent system prompt distribution as content — versioned, overridable, updatable — is how prompts evolve independently of code.

September 19, 2026·9 min read
Building a Plugin System for Your CLI Agent
CLI AI Agents

Building a Plugin System for Your CLI Agent

How do you let people add tools to your agent without forking it? A cli agent plugin system lets users extend the agent with their own tools. Here's how to rebuild a hardcoded tool list into a real plugin system.

September 19, 2026·8 min read

Save the prompts from this post

PromptABCD is a free prompt manager. Paste, organize, and reuse your best AI prompts — no more hunting through chat history.

Start free →
← PreviousPackaging a CLI Agent for DistributionNext →Building a Read-Only Mode for Safe Exploration
Share this post:
ShareShare