Architecture decision record

ADR-002: Rate limit at the edge gateway

Enforce one per-client token bucket in the gateway, with state in Redis, so limits hold however many services and replicas run behind it.

We will enforce client rate limits once, in the gateway, using a token bucket updated atomically in Redis, because per-service limits no longer bound what a client can send.

Services keep only a coarse safety limit against runaway internal callers.

Context

Under the per-service limits of ADR-001, one client sent 9× its contracted rate during the September incident without hitting a single limit: their traffic was spread across 3 services with 3 replicas each. We now have a gateway that runs plugins and a managed Redis cluster with p99 under 2 ms in-region. The token bucket explainer covers the algorithm, and the design document covers the request path, the script and the canary results.

Options considered

Keep per-service limits

  • No migration
  • The limit still multiplies as we scale

Gateway + Redis token bucket

  • One accurate limit per client
  • Rejections never reach services
  • Redis on the request path: +1–2 ms, needs a fail-open policy

CDN provider limiting

  • Absorbs floods before our network
  • Can't see API keys or weight endpoints

Consequences

What gets easier

  • Limits are contractual numbers that hold
  • One place to tune, observe and explain a 429

What gets harder

  • Redis becomes part of every request; if it's down we fail open and lean on service safety limits