Architecture decision record
ADR-002: Rate limit at the edge gateway
Enforce one per-client token bucket in the gateway, with state in Redis, so limits hold however many services and replicas run behind it.
We will enforce client rate limits once, in the gateway, using a token bucket updated atomically in Redis, because per-service limits no longer bound what a client can send.
Services keep only a coarse safety limit against runaway internal callers.
Context
Under the per-service limits of ADR-001, one client sent 9× its contracted rate during the September incident without hitting a single limit: their traffic was spread across 3 services with 3 replicas each. We now have a gateway that runs plugins and a managed Redis cluster with p99 under 2 ms in-region. The token bucket explainer covers the algorithm, and the design document covers the request path, the script and the canary results.
Options considered
Keep per-service limits
- No migration
- The limit still multiplies as we scale
Gateway + Redis token bucket
- One accurate limit per client
- Rejections never reach services
- Redis on the request path: +1–2 ms, needs a fail-open policy
CDN provider limiting
- Absorbs floods before our network
- Can't see API keys or weight endpoints
Consequences
What gets easier
- Limits are contractual numbers that hold
- One place to tune, observe and explain a 429
What gets harder
- Redis becomes part of every request; if it's down we fail open and lean on service safety limits