Part 2 · 2 chapters · ~12 min
Rate Limiter
Token buckets, sliding windows and fixed windows, atomic checks in Redis from many gateway instances, failure modes, and limits that clients can understand.
5
Brief, questions and numbers
the brief
- Limit each API client to its plan rate across all gateway instances, with fair behaviour for bursts.
| question | answer we assume |
|---|---|
| limit by? | API key, plus IP for unauthenticated routes |
| precision? | within a few percent is fine |
| latency budget? | under 2 ms per check |
| rules? | per plan and per endpoint group |
| when the limiter is down? | fail open, except login and OTP |
code
traffic = 50,000 req/s across 20 gateway instances clients = 200,000 keys state per key = ~50 bytes → 10 MB in Redis checks = 50,000 Redis script calls/s: a small Redis cluster
DISTRIBUTED RATE LIMITER
a token bucket per client in Redis, checked atomically by every gateway instance
swipe the figure sideways, or tap expand for full screen
1/5
token bucket
Each client has a bucket of capacity C refilled at R tokens per second; each request takes one. Bursts up to C are allowed, sustained rate is R.
capacity C, refill rate Rbursts allowed, average capped
6
v1, the break, and v2
v1. Each gateway instance keeps in-memory counters per key with a fixed one-minute window.
The break. With 20 instances, each sees a fraction of a client's traffic, so the effective limit is 20× the plan. Fixed windows also allow double bursts at window edges.
v2. A token bucket (or sliding window log/counter) stored in Redis and updated by an atomic Lua script, keys sharded across a Redis cluster, 429 with Retry-After headers, and a local fallback when Redis is unreachable.
the sentence
v2 buys accurate limits across any number of instances, and pays with a Redis round trip per request and a new dependency on the request path.
code
-- token bucket in Redis (Lua), called with: key, capacity, refill_per_sec, now_ms
local b = redis.call('HMGET', KEYS[1], 'tokens', 'ts')
local cap, rate, now = tonumber(ARGV[1]), tonumber(ARGV[2]), tonumber(ARGV[3])
local tokens = tonumber(b[1]) or cap
local ts = tonumber(b[2]) or now
tokens = math.min(cap, tokens + (now - ts) / 1000 * rate)
local allowed = tokens >= 1
if allowed then tokens = tokens - 1 end
redis.call('HSET', KEYS[1], 'tokens', tokens, 'ts', now)
redis.call('PEXPIRE', KEYS[1], math.ceil(cap / rate * 1000) * 2)
return { allowed and 1 or 0, math.floor(tokens) }