Architecture decision record
ADR-001: Rate limit inside each service
Each service enforces its own per-client token bucket in process, so we can ship limits before launch without new infrastructure.
We will embed a per-client token bucket in every service, keyed by API key, because launch is three weeks away and this needs no shared infrastructure.
Context
The public API launches on 2 December. Two design partners have scripted clients that retry aggressively, and nothing today stops one of them saturating a service. We have no gateway able to run custom logic, and no shared low-latency store in production.
Options considered
In-process limiter per service
- Ships with a library and a config file
- No new runtime dependency
- Limits multiply with replicas and services
Central limiter service
- One accurate limit per client
- New service, new on-call, before launch
Consequences
What gets easier
- Each team tunes limits for its own endpoints
What gets harder
- A client's real limit is per-service limit × replicas × services, so it grows as we scale
- Limits drift between services