Rate Limiting and Retries in a CLI Agent
By the time a rate-limit error reaches your code, the SDK already retried it. Good cli agent retry logic tells a 429 from a 529, honors retry-after, and avoids the double-retry deadlock.
import anthropic, time, random
client = anthropic.Anthropic(max_retries=0) # we own the retry loop
def call(**kwargs):
delay = 1.0
for attempt in range(5):
try:
return client.messages.create(**kwargs)
except anthropic.RateLimitError as e: # 429: your account's limit
wait = float(e.response.headers.get("retry-after", delay))
time.sleep(wait) # honor the server exactly
except anthropic.APIStatusError as e: # 529/503/500: server-side
if e.status_code in (529, 503, 500):
time.sleep(delay + random.uniform(0, delay)) # backoff + jitter
delay = min(delay * 2, 60)
else:
raise
raise RuntimeError("Exhausted retries")Here's something that surprises almost everyone building their first resilient agent: by the time a rate-limit error reaches your code, the SDK has usually already retried it three times and failed. The official Anthropic SDKs auto-retry transient failures twice by default with exponential backoff, so a 429 that surfaces to you has burned real time and, in the case of rate limits, real money — you're billed for failed 429 requests. That single fact reshapes how you should write cli agent retry logic: the goal isn't to retry harder, it's to retry correctly, and to know when retrying is the wrong move entirely.
This guide gives you retry logic that tells the two failure types apart, honors the server's instructions, and avoids the deadlock trap that turns a soft limit into a hard outage.
Quick-Start (Copy This Right Now)
The core insight is that a 429 and a 529 mean opposite things and need opposite handling. Here's a wrapper that treats them differently.
[object Object], anthropic, time, random
client = anthropic.Anthropic(max_retries=,[object Object],) ,[object Object],
,[object Object], ,[object Object],(,[object Object],):
delay = ,[object Object],
,[object Object], attempt ,[object Object], ,[object Object],(,[object Object],):
,[object Object],:
,[object Object], client.messages.create(**kwargs)
,[object Object], anthropic.RateLimitError ,[object Object], e: ,[object Object],
wait = ,[object Object],(e.response.headers.get(,[object Object],, delay))
time.sleep(wait) ,[object Object],
,[object Object], anthropic.APIStatusError ,[object Object], e: ,[object Object],
,[object Object], e.status_code ,[object Object], (,[object Object],, ,[object Object],, ,[object Object],):
time.sleep(delay + random.uniform(,[object Object],, delay)) ,[object Object],
delay = ,[object Object],(delay * ,[object Object],, ,[object Object],)
,[object Object],:
,[object Object],
,[object Object], RuntimeError(,[object Object],)What this does: Sets the client's own max_retries to zero so it doesn't retry underneath you, then handles the two error classes distinctly — a 429 waits exactly as long as the retry-after header says, while a 529 uses exponential backoff with jitter. One retry layer, two correct behaviors.
Wire this around every model call and your agent stops falling over on the transient failures that are a normal part of production traffic.
Understanding the Variables
The max_retries=0 on the client is the most important line, and it's the one people miss. The SDK retries by default. If you also wrap calls in your own retry loop without disabling that, you get two retry layers fighting each other — nested bursts that exhaust your request budget faster and can deadlock a rate-limited account. Own the loop or configure the SDK's, never both.
The retry-after header is the server telling you exactly how long to wait on a 429. It's not a suggestion to improve upon — waiting less just wastes a request and extends your limit window, and waiting on your own guessed schedule instead ignores the one piece of ground truth you were handed. Read it and obey it.
The jitter — that random.uniform(0, delay) — matters specifically for 529s. When Anthropic's infrastructure is overloaded, every client hitting it is retrying at once. If they all back off on the same fixed schedule, they retry in synchronized waves that keep the service saturated. Randomizing each client's delay spreads the load and helps the system recover.
⚡ Pro tip: Never retry a spend or credit-cap error, even though it also arrives as a 429-family failure. Those don't clear on their own — no amount of waiting refills an empty account. Read the error body, and if the problem is billing rather than rate, fail fast with a clear message instead of sleeping in a doomed loop.
Why Does cli agent retry logic Need to Distinguish 429 From 529?
Because the correct fix for each is the opposite of the other, and treating them the same guarantees you handle one of them wrong. A 429 is your account crossing a rate limit — requests, input tokens, or output tokens per minute. It's your responsibility, it comes with a retry-after, and if you keep hitting it, retrying won't help; you have a capacity problem that needs a higher tier, prompt caching, or lower concurrency.
A 529 is Anthropic's infrastructure being overloaded across all users. It has nothing to do with your account, there's no reliable retry-after, and those requests don't count against your billing. The right response is backoff with jitter and, if it persists, failing over to the same model on another provider like Bedrock or Vertex — because more retries against an overloaded endpoint won't help, but a different endpoint might.
[object Object], ,[object Object],(,[object Object],):
,[object Object], ,[object Object],(err, anthropic.RateLimitError):
,[object Object], ,[object Object], ,[object Object],
,[object Object], ,[object Object],(err, ,[object Object],, ,[object Object],) == ,[object Object],:
,[object Object], ,[object Object], ,[object Object],
,[object Object], ,[object Object],What this does: Labels a failure so your agent can respond with the right strategy — throttle back on a 429, ride out or fail over on a 529. Logging this classification also tells you, over time, whether your pain is self-inflicted (429s) or upstream (529s), which are two very different problems to fix.
⚠️ Common mistake: Treating every 4xx and 5xx from the API the same way — one retry loop, one backoff, one bucket. A 429 and a 529 have different headers, different meanings, and opposite fixes. Lumping them together means you'll retry a billing error forever, ignore a retry-after, or hammer an overloaded server when you should fail over. Branch on the error type.
Step-by-Step: Building the Full Retry Path
Start by disabling SDK retries and owning one loop. Then, for each call: attempt it; on a 429, read retry-after and sleep exactly that long; on a 529 or 5xx, sleep a jittered exponential backoff and try again; on anything else, raise immediately because it's not transient. Cap total attempts at four or five so a genuinely stuck request fails cleanly instead of looping.
The order of those branches is deliberate — check the specific rate-limit type before the generic status-code branch, so a 429 gets its header-honoring path rather than falling into blind backoff. And cap your maximum delay (60 seconds is reasonable) so an absurd retry-after value can't park your agent for ten minutes.
⚡ Pro tip: Enable prompt caching on your long system prompt and any stable context. Cached tokens don't count toward your input-tokens-per-minute limit on most models, which can lift your effective throughput several times over. For a rate-limited agent, caching is often a bigger win than any retry tuning — it stops you hitting the 429 in the first place.
⚡ Pro tip: Emit a metric for every retry — the error type, the attempt number, and the final outcome. Most 529 storms are short, but you only know that if you can graph them, and a rising 429 rate is an early warning that you need to tier up before it becomes an outage.
Should You Fail Over to Another Provider?
For a 529, the honest answer is often yes. When Anthropic's infrastructure is saturated, retrying the same endpoint harder does nothing — the capacity isn't there for anyone. The same Claude model is available through AWS Bedrock and Google Vertex, which draw on separate capacity pools, so a failover can succeed exactly when the direct API can't.
[object Object], ,[object Object],(,[object Object],):
,[object Object],:
,[object Object], call(**kwargs) ,[object Object],
,[object Object], anthropic.APIStatusError ,[object Object], e:
,[object Object], e.status_code == ,[object Object],:
,[object Object], call_bedrock(**kwargs) ,[object Object],
,[object Object],What this does: Falls over to the same model on a secondary provider when the primary returns an overloaded error, using a capacity pool that isn't affected by the primary's saturation. For a 429 you would not do this — a rate limit follows your account, not the endpoint — which is one more reason to branch on error type.
The tradeoff is added complexity and a second set of credentials to manage, so failover earns its place in production agents that can't tolerate a brief outage and skips the ones where "try again in a minute" is a fine answer.
⚡ Pro tip: Only fail over on 529, never on 429. A rate limit travels with your account, so a different provider that bills the same account may share the quota — and even when it doesn't, you're papering over a capacity problem you should fix at the source. Failover is for upstream overload, not your own limits.
Pro-Level Variations
Teams adapt retry logic to their risk and scale.
A backend developer running an interactive agent keeps retries short and few — an interactive user would rather see "the service is busy, try again" after ten seconds than wait through a minute of silent backoff, so they cap attempts low and surface failures fast.
A data engineer running a large batch job over thousands of items does the opposite: long backoff, provider failover to Bedrock on sustained 529s, and a per-workspace concurrency limit so the batch can't starve the company's interactive traffic into a self-inflicted 429 storm.
A platform team serving many internal users wraps every agent call in a shared client with tuned max_retries, centralized metrics, and a circuit breaker that stops sending requests entirely for a cooldown after repeated failures — protecting both their users and the upstream API from a thundering herd.
Each is the same core logic with the knobs turned to match latency tolerance and volume.
Troubleshooting Common Issues
If your agent seems to hang for exactly two minutes on failures, you've likely stacked your retry loop on top of the SDK's and the delays are compounding — disable one layer. If you keep hitting 429s no matter how you tune backoff, retries aren't your fix; you're over your tier's throughput, so cache, reduce concurrency, or upgrade. If a billing error loops forever, you're retrying a spend-cap 429 that will never clear — inspect the error body and bail on non-transient causes.
And if 529s dominate during peak hours, that's upstream saturation you can't tune away — add provider failover so the same model on Bedrock or Vertex catches the overflow.
Your Turn
Set your client's max_retries to zero, write one retry loop that branches on 429 versus 529, honor retry-after, add jitter to backoff, and cap attempts. Then add prompt caching to stop hitting limits in the first place. You'll have an agent that survives the transient failures that are a normal, expected part of running against any large API.
The retry policy — the branch logic, the caps, the failover rules — is the same across every agent you build, and the system-prompt notes about how the agent should behave when throttled belong with it. Keeping that policy and those prompts in a library like PromptABCD means your next agent inherits battle-tested cli agent retry logic instead of rediscovering the 429-versus-529 distinction during its first production incident.
Continue Reading
Save the prompts from this post
PromptABCD is a free prompt manager. Paste, organize, and reuse your best AI prompts — no more hunting through chat history.
