Architecture decision record

ADR-001: Rate limit inside each service

Each service enforces its own per-client token bucket in process, so we can ship limits before launch without new infrastructure.

We will embed a per-client token bucket in every service, keyed by API key, because launch is three weeks away and this needs no shared infrastructure.

Context

The public API launches on 2 December. Two design partners have scripted clients that retry aggressively, and nothing today stops one of them saturating a service. We have no gateway able to run custom logic, and no shared low-latency store in production.

Options considered

In-process limiter per service

  • Ships with a library and a config file
  • No new runtime dependency
  • Limits multiply with replicas and services

Central limiter service

  • One accurate limit per client
  • New service, new on-call, before launch

Consequences

What gets easier

  • Each team tunes limits for its own endpoints

What gets harder

  • A client's real limit is per-service limit × replicas × services, so it grows as we scale
  • Limits drift between services