Design document
Moving rate limiting to the edge
Replace per-service limiters with one token bucket per client in the gateway, backed by Redis, so a client's limit holds however far we scale.
- Problem: per-service limits multiply with services and replicas. One client sent 9× its contract without being limited.
- Proposal: enforce limits once, in the gateway, with an atomic Redis script. Services keep only a coarse safety limit.
- Ask: review the fail-open policy and the rollout by 17 October.
Context & problem
Since ADR-001, each service has run its own token bucket. That shipped quickly, but a limit enforced in nine places isn't a limit. During the September incident, one client's traffic spread across three services with three replicas each, and every replica admitted it.
Goals & non-goals
Goals
- One limit per client, accurate to within 5%
- Rejected requests never reach a service
- Under 2 ms added at p99
Non-goals
- DDoS protection: that stays with the CDN
- Per-endpoint quotas (a later phase)
Proposed design
Compare the two versions below. Today the limiters sit inside each service. In the proposal, the gateway checks one bucket per client in Redis before routing to any service.
The request path
The gateway makes one Redis round trip per request. The script refills and takes tokens atomically, so concurrent requests can never spend the same token twice.
A request reaches the gateway, keyed by its API key.
The gateway runs the bucket script in Redis. Refill and take happen in one atomic step.
Admitted requests are routed on as before.
Once a burst empties the bucket, the gateway answers 429 itself, with a Retry-After the client can trust.
The limiter script
-- KEYS[1] = bucket key (one per API key)
-- ARGV = capacity, rate (tokens/s), now (ms, from Redis TIME), cost
local b, r, now, cost = tonumber(ARGV[1]), tonumber(ARGV[2]), tonumber(ARGV[3]), tonumber(ARGV[4])
local s = redis.call('HMGET', KEYS[1], 'tokens', 'ts')
local tokens = tonumber(s[1]) or b
local ts = tonumber(s[2]) or now
tokens = math.min(b, tokens + math.max(0, now - ts) / 1000 * r)
local ok = tokens >= cost
if ok then tokens = tokens - cost end
redis.call('HSET', KEYS[1], 'tokens', tokens, 'ts', now)
redis.call('PEXPIRE', KEYS[1], math.ceil(b / r * 1000) + 1000)
return { ok and 1 or 0, tostring(tokens) }
Inputs. The gateway passes capacity and rate from the client's plan, plus the current time from Redis itself, so gateway clocks never matter.
Read the bucket. A client seen for the first time starts with a full bucket.
Refill, then take. This is the same arithmetic as the in-process version. Running it as a script makes the read-modify-write atomic.
Write back and expire. An idle key expires once it would be full anyway, so memory tracks active clients only.
Answer. Admitted or not, plus the tokens left, which the gateway turns into Retry-After.
Alternatives considered
The gateway option wins under the weights we agreed in the kickoff. Move the weights to test how sensitive the choice is: favouring time to ship above accuracy brings back per-service limits.
| Criterion | Weight | Per service | Gateway + Redis | CDN rules |
|---|---|---|---|---|
| Accuracy of the limit | 5 | 1 | 5 | 3 |
| Latency added | 3 | 5 | 3 | 4 |
| Operational cost | 3 | 4 | 3 | 4 |
| Sees API keys and plans | 4 | 5 | 5 | 1 |
| Time to ship | 2 | 5 | 3 | 4 |
Canary results
We ran the gateway limiter for the noisiest 5% of clients for a week. Everyone else's tail latency stopped tracking those clients' bursts.
| Day | Per-service limits | Edge limiter |
|---|---|---|
| Mon | 380 | 371 |
| Tue | 412 | 214 |
| Wed | 395 | 188 |
| Thu | 441 | 181 |
| Fri | 404 | 179 |
Risks & mitigations
| Risk | Likelihood | Impact | Mitigation |
|---|---|---|---|
| Redis unavailable | Low | High | Fail open after 5 ms; services' coarse safety limits still apply |
| Hot key for one huge client | Medium | Medium | Shard that client's bucket by region; reconcile every second |
| Clock skew between gateways | Low | Low | None needed: the script reads time from Redis |
Rollout plan
- Canary: noisiest 5% of clients
Done. The results are above.
- All clients, shadow mode
Decide but don't reject; compare against per-service decisions.
- Enforce at the edge
Exit criterion: under 0.1% disagreement in shadow mode.
- Retire per-service limiters
Keep a coarse safety limit at 10× the plan rate.
Decision
We will enforce client rate limits at the gateway with the Redis bucket script, failing open, and retire the per-service limiters.
Recorded as ADR-002, which supersedes ADR-001.