PromptABCD
FeaturesLearnGuideBlogContext Blocks
Sign inGet started free
Sign inSign up
PromptABCD

A calm home for your best AI prompts. Save them once, find them in seconds, reuse them forever.

Product

  • Features
  • Chrome Extension
  • Free Courses
  • How it works
  • Use cases
  • Blog
  • Context Blocks
  • Export Anywhere
  • FAQ

Resources

  • User guide
  • Learn prompting
  • Sign in
  • Get started free

© 2026 PromptABCD. All rights reserved.

AboutPrivacy PolicyTerms and Conditions
Home/Blog/Multi-Agent Systems/Rate Limiting a Fleet of Agents
Multi-Agent Systems

Rate Limiting a Fleet of Agents

Per-agent rate limits feel safe and are quietly useless - a fleet of agents under individual limits can still blow a shared quota. Multi agent rate limiting has to be fleet-wide.

September 28, 2026·9 min read
ShareShare
⚡Featured Prompt— copy and use right now
import time

class SharedRateLimiter:
    def __init__(self, store, limit_per_sec):
        self.store = store          # shared across all agents
        self.limit = limit_per_sec

    def allow(self):
        now = int(time.time())
        key = f"ratelimit:{now}"
        count = self.store.incr(key)   # atomic across the whole fleet
        if count == 1:
            self.store.expire(key, 2)
        return count <= self.limit

Here's a piece of conventional wisdom that's quietly wrong: "give each agent a rate limit and you're protected." Per-agent rate limits feel safe and are nearly useless for a fleet, because twenty agents each politely under their individual limit can still collectively blow a shared quota to pieces. The API you're calling doesn't care that each agent stayed under its own limit - it sees the sum, and the sum is what gets you throttled or hit with a surprise bill. Multi agent rate limiting has to be fleet-wide and shared, not per-agent, and that requires coordination that per-agent limits specifically avoid.

This guide builds fleet-wide rate limiting that actually holds a shared quota, covers the coordination problem that makes it harder than single-process rate limiting, and shows how to distribute a shared budget fairly across agents. It's the difference between a limit that looks protective and one that is.

Quick-Start (Copy This Right Now)

Here's a shared rate limiter that every agent consults before making a call, backed by a store all agents can see:

python
[object Object], time

,[object Object], ,[object Object],:
    ,[object Object], ,[object Object],(,[object Object],):
        ,[object Object],.store = store          ,[object Object],
        ,[object Object],.limit = limit_per_sec

    ,[object Object], ,[object Object],(,[object Object],):
        now = ,[object Object],(time.time())
        key = ,[object Object],
        count = ,[object Object],.store.incr(key)   ,[object Object],
        ,[object Object], count == ,[object Object],:
            ,[object Object],.store.expire(key, ,[object Object],)
        ,[object Object], count <= ,[object Object],.limit

What this does: It keeps the request count for the current second in a shared store that every agent increments atomically, so the limit applies to the fleet's total rather than to any single agent. Any agent about to exceed the shared limit gets told no, regardless of how little that particular agent has done - which is exactly what per-agent limits can't do.

Understanding the Variables

Three things make multi agent rate limiting different from the single-process kind: shared state, atomicity, and fairness.

Shared state is the core requirement. The limit is on the fleet's combined rate, so the counter has to live somewhere every agent can see and update - a shared store, not each agent's local memory. The instant you have per-agent local counters, you've lost, because no agent can see the fleet's total, and the total is the only number that matters to the quota you're protecting.

Atomicity is what keeps the shared counter correct under concurrency. When twenty agents check and increment the counter simultaneously, a non-atomic "read then write" lets them all read the same value and all proceed, blowing the limit. The increment has to be atomic - one operation that reads and writes together - so concurrent agents see each other's increments. This is the subtle bug that makes naive shared limiters leak.

Fairness is how you divide a shared budget so one greedy agent doesn't starve the others. A shared limit protects the quota, but without fairness, a single high-volume agent can consume the entire fleet budget and leave nothing for the rest. Fair distribution ensures every agent gets a workable share of the common limit.

⚡ Pro tip: The moment your rate limit counter lives in an agent's local memory instead of a shared store, your fleet limit is fiction. Each agent enforcing its own local view enforces nothing on the total, which is the only thing the downstream quota measures. If you can't point to the one shared place the count lives, you don't have fleet-wide rate limiting - you have a collection of independent limits that sum to something you're not controlling.

Step-by-Step: Building Multi Agent Rate Limiting

Let's build a fleet limiter that holds a shared quota and distributes it fairly.

Step one, make the counter shared and atomic. Use a store with atomic increment so concurrent agents can't race past the limit.

python
[object Object], ,[object Object],(,[object Object],):
    now = ,[object Object],(time.time() / window)
    key = ,[object Object],
    count = store.incr(key)         ,[object Object],
    store.expire(key, window * ,[object Object],)
    ,[object Object], count <= limit

What this does: It buckets requests by time window and atomically increments a shared per-window counter, allowing the request only if the fleet's total for that window is under the limit. Because the increment is atomic and the key is shared, every agent's request counts against the same total, and concurrent requests can't both slip through on a stale read.

Step two, add fairness so no single agent monopolizes the shared budget. Give each agent a share of the total and let it borrow unused capacity from others only when available.

python
[object Object], ,[object Object],(,[object Object],):
    fair_share = fleet_limit // n_agents
    agent_key = ,[object Object],
    agent_count = store.incr(agent_key)
    store.expire(agent_key, ,[object Object],)
    ,[object Object], agent_count <= fair_share:
        ,[object Object], ,[object Object],                       ,[object Object],
    ,[object Object], allow_request(store, fleet_limit)  ,[object Object],

What this does: It guarantees each agent its fair share of the fleet limit, and lets an agent exceed its share only if the fleet as a whole still has slack. A greedy agent can use spare capacity when the fleet is quiet but can't starve others when the fleet is busy, because past its fair share it competes for the shared remainder rather than claiming it outright.

Step three, handle rejection gracefully. When an agent is rate-limited, it should back off with jitter and retry, not spin. Return a retry-after hint so the agent waits the right amount rather than hammering the limiter.

Pro-Level Variations

For token-based limits (like model APIs that limit tokens per minute, not just requests), count tokens instead of requests - estimate a call's token cost before making it and reserve that many from the shared budget. Rate limiting agents on model APIs almost always means limiting tokens, since that's the real quota, and a request-count limit misses it entirely because requests vary enormously in token cost.

For multiple downstream quotas (several APIs, each with its own limit), run a separate shared limiter per quota, since each downstream service throttles independently. One fleet-wide limiter per protected resource keeps each quota safe without one limit's slack masking another's exhaustion.

⚡ Pro tip: Use a token-bucket algorithm rather than a fixed window when burst tolerance matters. A fixed window resets abruptly and allows a double-rate burst at the boundary - the last requests of one window plus the first of the next. A token bucket refills smoothly and caps bursts naturally, which is usually what you actually want when protecting a downstream service that dislikes spikes. The smoother shaping is worth the slightly more complex implementation.

Should Agents Wait for the Limiter or Avoid Hitting It?

Here's a contrarian point about multi agent rate limiting that most implementations miss: the best rate limiter is the one your agents rarely trigger, because they're shaping their own demand to fit the budget. Treating the limiter purely as a gate agents slam into and bounce off is reactive and wasteful - every rejection is work an agent has to retry, and a fleet that's constantly bouncing off the limit is a fleet doing a lot of redundant back-and-forth.

The proactive alternative is to let agents see the remaining budget and pace themselves against it. If an agent can check how much fleet capacity is left this window, it can decide to batch its calls, defer non-urgent work, or slow its own cadence - avoiding rejection rather than absorbing it. This turns the limiter from a wall into a shared signal the fleet coordinates around, which is far more efficient than every agent probing the wall independently.

python
[object Object], ,[object Object],(,[object Object],):
    now = ,[object Object],(time.time())
    used = ,[object Object],(store.get(,[object Object],) ,[object Object], ,[object Object],)
    ,[object Object], ,[object Object],(,[object Object],, limit - used)   ,[object Object],

,[object Object], ,[object Object],(,[object Object],):
    slack = remaining_budget(store, limit) / limit
    ,[object Object],
    ,[object Object], slack < ,[object Object], ,[object Object], urgency == ,[object Object],

What this does: It exposes the remaining fleet budget so agents can pace themselves, and lets low-urgency work voluntarily defer when capacity is tight - reserving the scarce remaining budget for urgent calls. Instead of every agent racing to consume the budget and bouncing off the limit, the fleet self-regulates, with important work getting priority access to a shared resource that low-priority work steps back from.

This proactive shaping also solves a problem pure limiting can't: priority. A blind rate limiter is first-come-first-served, so a flood of trivial calls can consume the budget and block an urgent one. Letting agents consult remaining budget and defer by urgency means the fleet naturally prioritizes, keeping headroom for what matters instead of letting whatever arrives first win.

⚡ Pro tip: Expose the remaining budget as a shared signal and let low-priority agents defer when it's scarce. A limiter agents can see lets the fleet shape its own demand and reserve capacity for urgent work; a limiter agents can only hit forces every request to compete equally and wastes effort on rejections and retries. The visible-budget version does more with the same quota because the fleet stops fighting the limit and starts cooperating with it.

Troubleshooting Common Issues

If you're still getting throttled by a downstream API despite a fleet limiter, your atomicity is probably broken - agents are racing on a non-atomic check-and-increment and slipping past the limit under concurrency. Confirm the increment is a single atomic operation, not a separate read and write, because the gap between them is exactly where concurrent agents leak through.

If some agents are starved while others are busy, your fairness distribution needs work - a purely first-come limiter lets fast agents consume everything. Add per-agent fair shares so every agent gets guaranteed capacity regardless of how aggressive its peers are.

⚠️ Common mistake: Implementing per-agent rate limits and believing they protect a shared downstream quota. They don't, and the failure is silent until you hit the quota - each agent stays under its own limit, everyone feels compliant, and the fleet still collectively exceeds the shared limit because nobody was counting the total. The only limit that protects a shared resource is one that counts the shared total, which per-agent limits fundamentally cannot do.

Your Turn

Stand up a shared store, move your rate counter into it with an atomic increment, and have every agent consult it before calling the protected resource. Then add fair shares so no agent can starve the others, and confirm the fleet total holds under a burst of concurrent requests.

Keep your rate-limit configurations and fairness policies versioned so you can reuse limits you've tuned. I store the per-quota limits and fair-share settings in PromptABCD, because the right limit and distribution for each downstream resource is tuned from real quota data, and reusing a proven multi agent rate limiting configuration means a new fleet protects its shared quotas correctly from day one instead of discovering per-agent limits don't add up the way everyone assumed.

⚡ Pro tip: Set your fleet limit slightly below the downstream quota, not exactly at it, to leave a safety margin for clock skew and counting lag. A shared counter is never perfectly synchronized across a distributed fleet, so a limit set exactly at the quota will occasionally tip over it during a burst. Aiming a little under the real ceiling absorbs that imprecision and keeps you on the safe side of throttling, at the cost of a few percent of theoretical capacity you'd rather not risk anyway.

multi-agent-systemsrate-limitingscalabilityreliabilityai-agents

Continue Reading

A Reusable Prompt Kit for Agent Teams
Multi-Agent Systems

A Reusable Prompt Kit for Agent Teams

A team rebuilt their agent prompts from memory every project, and every project drifted a little worse. That failure is why a multi agent prompt kit matters. Here's the reusable set of role prompts every team should keep.

October 2, 2026·9 min read
Multi-Agent Systems on a Budget: An Interactive Guide
Multi-Agent Systems

Multi-Agent Systems on a Budget: An Interactive Guide

Most multi-agent tutorials assume you'll burn tokens freely. That's wrong for anyone shipping on real constraints. A cheap multi agent system can match an expensive one with the right moves. Here's how to build one.

October 2, 2026·9 min read
The Manager Agent Anti-Pattern: A Teardown
Multi-Agent Systems

The Manager Agent Anti-Pattern: A Teardown

Why does your orchestrator agent keep becoming a bottleneck that mangles every handoff? You've hit the manager agent anti pattern — one agent trying to coordinate everything. Here's why it fails and what replaces it.

October 2, 2026·9 min read

Save the prompts from this post

PromptABCD is a free prompt manager. Paste, organize, and reuse your best AI prompts — no more hunting through chat history.

Start free →
← PreviousHandling Cascading Failures Between AgentsNext →Versioning Prompts Across Agents
Share this post:
ShareShare