Design document

Moving rate limiting to the edge

Replace per-service limiters with one token bucket per client in the gateway, backed by Redis, so a client's limit holds however far we scale.

  • Problem: per-service limits multiply with services and replicas. One client sent 9× its contract without being limited.
  • Proposal: enforce limits once, in the gateway, with an atomic Redis script. Services keep only a coarse safety limit.
  • Ask: review the fail-open policy and the rollout by 17 October.

Context & problem

Since ADR-001, each service has run its own token bucket. That shipped quickly, but a limit enforced in nine places isn't a limit. During the September incident, one client's traffic spread across three services with three replicas each, and every replica admitted it.

9×Peak traffic vs contract, never limited
3 × 3Services × replicas the limit was split across
412 msp99 for everyone else during the incident

Goals & non-goals

Goals

  • One limit per client, accurate to within 5%
  • Rejected requests never reach a service
  • Under 2 ms added at p99

Non-goals

  • DDoS protection: that stays with the CDN
  • Per-endpoint quotas (a later phase)

Proposed design

Compare the two versions below. Today the limiters sit inside each service. In the proposal, the gateway checks one bucket per client in Redis before routing to any service.

Client Gateway Orders Billing limiter limiter
Each service limits on its own, so a client's real limit is the per-service limit × replicas × services.
Client Gateway Redisbucket per client Orders Billing take tokens
One bucket per client, checked once at the gateway: the limit holds however many services and replicas there are.

The request path

The gateway makes one Redis round trip per request. The script refills and takes tokens atomically, so concurrent requests can never spend the same token twice.

Client Gateway Redis Orders GET /orders EVALSHA bucket refill + take, atomically allowed · 7.4 left GET /orders 200 OK a burst empties the bucket GET /orders EVALSHA bucket denied · retry 0.4 s 429 · Retry-After: 1
The service only ever sees admitted requests; Redis answers in well under a millisecond in-region.

A request reaches the gateway, keyed by its API key.

The gateway runs the bucket script in Redis. Refill and take happen in one atomic step.

Admitted requests are routed on as before.

Once a burst empties the bucket, the gateway answers 429 itself, with a Retry-After the client can trust.

The limiter script

-- KEYS[1] = bucket key (one per API key)
-- ARGV    = capacity, rate (tokens/s), now (ms, from Redis TIME), cost
local b, r, now, cost = tonumber(ARGV[1]), tonumber(ARGV[2]), tonumber(ARGV[3]), tonumber(ARGV[4])

local s = redis.call('HMGET', KEYS[1], 'tokens', 'ts')
local tokens = tonumber(s[1]) or b
local ts = tonumber(s[2]) or now

tokens = math.min(b, tokens + math.max(0, now - ts) / 1000 * r)
local ok = tokens >= cost
if ok then tokens = tokens - cost end

redis.call('HSET', KEYS[1], 'tokens', tokens, 'ts', now)
redis.call('PEXPIRE', KEYS[1], math.ceil(b / r * 1000) + 1000)
return { ok and 1 or 0, tostring(tokens) }

Inputs. The gateway passes capacity and rate from the client's plan, plus the current time from Redis itself, so gateway clocks never matter.

Read the bucket. A client seen for the first time starts with a full bucket.

Refill, then take. This is the same arithmetic as the in-process version. Running it as a script makes the read-modify-write atomic.

Write back and expire. An idle key expires once it would be full anyway, so memory tracks active clients only.

Answer. Admitted or not, plus the tokens left, which the gateway turns into Retry-After.

Alternatives considered

The gateway option wins under the weights we agreed in the kickoff. Move the weights to test how sensitive the choice is: favouring time to ship above accuracy brings back per-service limits.

Scores 1–5, higher is better. Weights 0–5 say how much each criterion matters.
CriterionWeightPer serviceGateway + RedisCDN rules
Accuracy of the limit5153
Latency added3534
Operational cost3434
Sees API keys and plans4551
Time to ship2534

Canary results

We ran the gateway limiter for the noisiest 5% of clients for a week. Everyone else's tail latency stopped tracking those clients' bursts.

DayPer-service limitsEdge limiter
Mon380371
Tue412214
Wed395188
Thu441181
Fri404179
The canary started on Tuesday. p99 for the other 95% of clients fell by more than half and stayed flat through the noisy clients' bursts.

Risks & mitigations

RiskLikelihoodImpactMitigation
Redis unavailableLowHighFail open after 5 ms; services' coarse safety limits still apply
Hot key for one huge clientMediumMediumShard that client's bucket by region; reconcile every second
Clock skew between gatewaysLowLowNone needed: the script reads time from Redis

Rollout plan

  1. Canary: noisiest 5% of clients

    Done. The results are above.

  2. All clients, shadow mode

    Decide but don't reject; compare against per-service decisions.

  3. Enforce at the edge

    Exit criterion: under 0.1% disagreement in shadow mode.

  4. Retire per-service limiters

    Keep a coarse safety limit at 10× the plan rate.

Decision

We will enforce client rate limits at the gateway with the Redis bucket script, failing open, and retire the per-service limiters.

Recorded as ADR-002, which supersedes ADR-001.