Part 2 · 2 chapters · ~12 min

Rate Limiter

Token buckets, sliding windows and fixed windows, atomic checks in Redis from many gateway instances, failure modes, and limits that clients can understand.

5

Brief, questions and numbers

the brief
  1. Limit each API client to its plan rate across all gateway instances, with fair behaviour for bursts.
questionanswer we assume
limit by?API key, plus IP for unauthenticated routes
precision?within a few percent is fine
latency budget?under 2 ms per check
rules?per plan and per endpoint group
when the limiter is down?fail open, except login and OTP
code
traffic       = 50,000 req/s across 20 gateway instances
clients       = 200,000 keys
state per key = ~50 bytes  → 10 MB in Redis
checks        = 50,000 Redis script calls/s: a small Redis cluster
DISTRIBUTED RATE LIMITER
a token bucket per client in Redis, checked atomically by every gateway instance
clientAPI key k1gateway × Ncheck before routingRedisbucket:k1 = tokens, tsLua scriptrefill + take, atomicserviceif allowed
swipe the figure sideways, or tap expand for full screen
1/5
token bucket
Each client has a bucket of capacity C refilled at R tokens per second; each request takes one. Bursts up to C are allowed, sustained rate is R.
capacity C, refill rate Rbursts allowed, average capped
6

v1, the break, and v2

v1. Each gateway instance keeps in-memory counters per key with a fixed one-minute window.

The break. With 20 instances, each sees a fraction of a client's traffic, so the effective limit is 20× the plan. Fixed windows also allow double bursts at window edges.

v2. A token bucket (or sliding window log/counter) stored in Redis and updated by an atomic Lua script, keys sharded across a Redis cluster, 429 with Retry-After headers, and a local fallback when Redis is unreachable.

the sentence
v2 buys accurate limits across any number of instances, and pays with a Redis round trip per request and a new dependency on the request path.
code
-- token bucket in Redis (Lua), called with: key, capacity, refill_per_sec, now_ms
local b = redis.call('HMGET', KEYS[1], 'tokens', 'ts')
local cap, rate, now = tonumber(ARGV[1]), tonumber(ARGV[2]), tonumber(ARGV[3])
local tokens = tonumber(b[1]) or cap
local ts = tonumber(b[2]) or now
tokens = math.min(cap, tokens + (now - ts) / 1000 * rate)
local allowed = tokens >= 1
if allowed then tokens = tokens - 1 end
redis.call('HSET', KEYS[1], 'tokens', tokens, 'ts', now)
redis.call('PEXPIRE', KEYS[1], math.ceil(cap / rate * 1000) * 2)
return { allowed and 1 or 0, math.floor(tokens) }