API Rate Limiter Strategies for High-Volume SaaS Integrations
Five algorithms solve every rate-limiting choice, but enforcement location matters just as much.
Most engineering teams can recite their provider's rate limit from memory. The integration still falls over in production, because rate limiting isn't one decision, it's several layers that all have to hold at the same time, and most teams only build one of them. That last one catches people off guard, because concurrency is a different ceiling than request rate. Latency creeps up somewhere downstream, requests pile up waiting on it, and the system runs out of threads, database connections, or memory long before it ever gets near its requests-per-second cap.
Providers don't make this easier by being consistent with each other. Some send a 200 OK that looks like success but hides an error in the response body. Data quietly goes missing unless something is reading the body every time. And some use a status code that has nothing to do with rate limits on paper but gets used for one anyway, which can trick client code into refreshing an auth token when the real issue is traffic volume, not credentials.
A team building a Zendesk sync ran into a version of this that's almost too on-the-nose to be a cautionary tale: they built fast, multi-update agent workflows, shipped them, and only then discovered Zendesk caps updates at 30 per ticket per short window. The system hadn't been built to expect a limit that specific, applied at that granularity, until it was already live and failing, because the limit was never secret, but the architecture assumed one number would cover every case, and one number never does.
The five canonical algorithms and what each one protects
Five algorithms cover almost every rate-limiting decision a platform team will make, and each one is built to protect against a different shape of traffic. It either blocks real users who didn't do anything wrong, or hands an attacker a loophole that doubles the limit you thought you'd set. These five, Fixed Window Counter, Sliding Window Log, Sliding Window Counter, Token Bucket, and Leaky Bucket, are the actual menu every platform team chooses from when the system goes into production.
Fixed Window Counter protects against the simplest case: one counter, one increment per request, cheap to run and easy to reason about. The edge of the window itself is the catch. That makes it fine for internal limits or login throttling, where nobody is losing sleep over a user sneaking in a few extra requests, but risky anywhere precision actually matters.
Sliding Window Log protects accuracy itself. It keeps a timestamp for every single request, so there's no boundary trick to exploit and no approximation involved. The tradeoff is memory: storing a timestamp per request scales with how much traffic a client sends, and a single noisy client can make that expensive fast. It earns its keep on strict audit APIs or high-value APIs where getting the count exactly right is worth the memory cost.
Sliding Window Counter is one of the five canonical algorithms platform teams pick from. It blends the counts from the current and previous windows, which gets you accuracy close to the log approach while keeping memory flat, just two keys per client no matter how much traffic comes in. Cloudflare runs this approach across its global network, and it's the version most public APIs default to in practice.
Token Bucket protects burst tolerance. Capacity fills up to a set cap, each request spends a token, and once the bucket is empty, requests get turned away until it refills. That makes it the natural fit for developer-facing APIs that need to let a client fire off a quick burst of calls while still holding them to a steady average over time. Stripe tells its own API consumers to build a client-side token-bucket limiter for exactly this reason.
Leaky Bucket protects the thing downstream of the client: a database, a queue, a third-party service that can't handle spikes no matter how short. Requests drain out at one constant rate, and anything over that gets either dropped or held until there's room. Shopify uses a leaky-bucket model to pace traffic hitting its API.
None of these five count requests by request alone once the API gets complex enough. GraphQL is the clearest case: a single query can ask for a deep, tangled tree of data that costs the server far more than a hundred simple ones combined. Limiting by request count alone misses that entirely, so modern GraphQL rate limiting measures query depth, complexity, and the actual cost of resolving it, not just how many times someone hit the endpoint.
Where enforcement lives
Choosing an algorithm answers one question. Deciding where it runs answers a different one, and most teams only ever answer the first. There are three places enforcement can live, each with a different protection scope.
At the edge, an API gateway or CDN enforces limits broadly and cheaply, catching volumetric abuse before it ever reaches a backend service, and doing it with minimal added delay. That's useful protection, but it's also blind in a specific way: an edge gateway typically has no idea which tenant, which customer, or which subscription tier a request belongs to. It just sees traffic.
One level in, at the service mesh, enforcement shifts to protecting services from each other. Mesh-level limits exist to stop that kind of internal blast radius.
The application layer is one of three enforcement locations, each with a different protection scope. A limit that only exists in the application layer has no defense against a flood of traffic large enough to eat up compute before any application code even runs. Both gaps are real, and the only way to close them is to run enforcement at more than one layer at once.
Making that work across a fleet of servers requires a shared place to keep count. Redis is the common answer: every worker checks in with the same counter before deciding whether to send a request, so the system behaves like one limiter instead of workers each guessing independently and collectively blowing past the limit. The implementation detail that matters here: Redis-based limiters hold up better using Lua scripts run through EVAL rather than MULTI/EXEC transactions, because Lua runs atomically on the server in a single round trip, sidestepping the kind of check-then-act race condition that MULTI/EXEC retry loops exist to patch over.
Concurrency limits deserve a separate mention here, because they live at the application layer too, sometimes scoped per endpoint or per tenant, and they guard something request-rate limits never touch: the actual threads and database connections a system has to spend to serve each request.
Client-side queuing and backoff as the other half of a working rate-limit system
Server-side limits only work if the client on the other end knows how to behave when it gets told no. Teams often don't build that behavior until errors start showing up in logs, even though the client needs to back off correctly for server-side limits to protect anything. Get the backoff wrong and server-side limiting doesn't protect anything: it just creates a thundering herd, where every client that got rejected retries at roughly the same moment and slams the same limit right back into place. AWS tested this directly and found that Full Jitter exponential backoff, adding randomness to each retry delay, cut the total number of calls by more than half compared to backoff with no jitter at all, when many clients were competing for the same resource. Plain exponential backoff without jitter lost badly by comparison.
Queuing lets a well-built client put requests into a queue and release them at a controlled pace, instead of firing them the moment they're ready, which stops sudden spikes before they ever reach the provider's limiter. That matters most for the heavy jobs, things like pulling in a customer's full history of invoices or contacts in one go, where an unthrottled client could otherwise try to push years of data through in a few seconds. But a queue isn't just a buffer, it has to understand what's actually urgent. A payment status update needs to clear before a backlog of historical invoice imports does, and a sync job a customer is actively watching shouldn't wait behind a nightly reporting job that nobody's checking in real time. Building the queue around FIFO order alone ignores all of that business context.
Polling makes all of this worse than it needs to be. Checking an API over and over for updates burns through quota even when nothing has changed, which is quota spent for zero information gained. Webhooks and other event-driven approaches flip that: the provider notifies the integration only when something actually changes, so quota gets spent on real updates instead of repeated no-op checks.
Handling provider inconsistency also falls on the client, and it has to be handled case by case, because there's no universal rule that covers it. A 429 that comes with a Retry-After header should be parsed and respected exactly as given. A 429 without one needs the fallback of jittered exponential backoff rather than a guess at timing. A 200 OK response that still contains an error in its body needs to be parsed on every single call, because skipping that check is how data quietly goes missing without ever throwing a visible error. And a 403 that's actually standing in for a rate-limit warning needs to be told apart from a genuine authentication failure before the client reflexively tries to refresh a token that was never the problem.
Batching compounds all of these savings. Grouping related objects into fewer, larger calls cuts that volume down substantially. Less quota gets spent per unit of actual work done.
Multi-tenant fairness and per-tenant enforcement
Everything so far assumes one client talking to one provider. Multi-tenant SaaS platforms break that assumption immediately, because a single aggressive customer, running an unusually large sync or an unusually chatty integration, can burn through a shared provider quota and leave every other customer on the platform degraded as a result. Stopping that requires tracking rate-limit state per tenant, in addition to watching the platform's total usage.
When one platform is making calls to a provider on behalf of many different customers, per-account tracking isn't optional. One customer's overly aggressive sync schedule has to be contained so it doesn't quietly eat the quota that every other customer is also relying on. Tenant identity becomes a key dimension of that shared state at this layer, attached directly to the data rather than tacked on in a log line. Some teams go a layer deeper still, adding a per-user or per-token sub-limit inside a single tenant, so one noisy automation script inside a customer's account doesn't end up throttling that customer's entire organization.
Tiering is where all this gets translated into something a customer can actually understand. The fix isn't fewer tiers for the sake of simplicity, it's tiers that map to real categories of use case rather than arbitrary slices of a number line.
The provider limits that platforms actually have to plan around vary enough to make per-tenant tracking a necessity rather than a nice-to-have. QuickBooks Online allows a few hundred concurrent requests per minute with no fixed daily cap mentioned. Xero caps usage at 60 calls per minute and 5,000 calls per day per organization, on a rolling basis. Zoho Books varies its concurrent call allowance by subscription plan, with both per-minute and per-day ceilings that scale up as a customer pays for a higher tier. Notion's API holds to an average of 3 requests per second, the tightest ceiling among the major SaaS platforms, allowing short bursts but triggering 429s the moment traffic sustains above that average. Zendesk Suite Pro allows 100 requests per minute (more with its High Volume add-on), but also layers in a tight cap on how many times a single ticket can be updated in a short window, a constraint that directly hits fast multi-update agent workflows. Salesforce enforces its limit on a 24-hour rolling basis per organization; HubSpot combines per-app burst limits with a shared daily limit at the account level; NetSuite caps concurrent requests per account. No two of these behave the same way, so a platform serving all of them at once needs tenant-level accounting instead of one global number.
Observability as the feedback loop that keeps the layered system calibrated
A layered rate-limit system built exactly right on launch day will still be wrong a few months later, because traffic patterns shift, new tenants onboard, and providers change their own limits without much warning.
Measurement needs to happen at every layer discussed so far, not just one. Each of those signals maps directly back to a layer built earlier: queue depth reveals whether client-side queuing is keeping up, tenant-level quota tracking reveals whether per-tenant fairness is holding, and rate-limit response logging reveals whether the algorithm chosen for a given endpoint still fits the traffic it's actually seeing.
None of the algorithms, enforcement points, or client behaviors described earlier are static choices made once and left alone. They're settings that need revisiting as the system they protect keeps changing shape, and the only way to know when to revisit them is to be watching the numbers that tell you they've drifted.



