Skip to content

Reference architecture

Edge rate-limiting and DDoS mitigation architecture

Protect backend APIs from scraping, credential stuffing and volumetric DDoS attacks with token-bucket rate limiting at the edge.

Design constraints

What the design has to hold to

Targets for the scenario this reference is sized for. A real engagement starts by replacing them with your own numbers.

  • 01Under 2 ms of rate-limit overhead per API request
  • 02Distributed token-bucket and sliding-window algorithms
  • 03Tiered limits by subscription plan and API key
  • 04Automated IP blocking for volumetric brute-force bots

Component topology

System components and technologies

Subsystems with separate responsibilities, clear contracts between them and storage that scales on its own. The stack named for each is typical, not mandatory.

Stack topology

Edge rate-limiting and DDoS mitigation architecture

Illustrative reference architecture

  1. 01

    Edge security proxy

    Anycast DDoS absorption, TLS termination and WAF rules

    Cloudflare Enterprise / AWS WAF

  2. 02

    Rate-limiting gateway

    Checks token-bucket allowances and rejects excess traffic with HTTP 429

    Envoy Proxy / Traefik

  3. 03

    Distributed counter store

    In-memory sliding-window counters

    Redis Cluster / Upstash

  4. 04

    Security analytics (SIEM)

    Bot detection and IP reputation scoring

    Datadog Security / CloudWatch

Subsystem 01

Edge security proxy

Anycast DDoS absorption, TLS termination and WAF rules

Typical stack

Cloudflare Enterprise / AWS WAF

Subsystem 02

Rate-limiting gateway

Checks token-bucket allowances and rejects excess traffic with HTTP 429

Typical stack

Envoy Proxy / Traefik

Subsystem 03

Distributed counter store

In-memory sliding-window counters

Typical stack

Redis Cluster / Upstash

Subsystem 04

Security analytics (SIEM)

Bot detection and IP reputation scoring

Typical stack

Datadog Security / CloudWatch

Data lifecycle

End-to-end data flow

  1. A client sends a request with Authorization: Bearer <key> to the API gateway.

  2. Cloudflare WAF checks IP reputation and passes clean requests to the Envoy gateway.

  3. Envoy runs a Lua script and Redis pipeline that evaluates the sliding window in a single round trip.

  4. Over the tier limit, the gateway returns HTTP 429 Too Many Requests with a Retry-After header.

  5. Allowed requests pass to the backend service, and Redis increments the counter with an automatic TTL.

Reliability and resilience

Failure modes and how each is contained

Failure mode 01

Redis latency spike stalling the gateway

Mitigation

Fail open (allow requests) when Redis takes longer than a 5 ms timeout, so the limiter can't cause an outage.

Failure mode 02

Botnets rotating residential IPs

Mitigation

Fingerprinting heuristics (JA4 TLS fingerprints, canvas hashes) throttle bots even as their IPs change.

Failure mode 03

Thundering herd on counter reset

Mitigation

Sliding-window logs with jitter, so clients don't all retry on the minute boundary.

Questions

What teams ask about this design

A token bucket allows short bursts while holding an average rate. A sliding window prevents the spike that fixed one-minute windows allow at their boundary.

Every 429 carries a Retry-After header, and every response carries the widely used X-RateLimit-Limit, X-RateLimit-Remaining and X-RateLimit-Reset headers, or the IETF draft RateLimit headers if your clients support them.

Planning a system like this?

Send us your requirements, expected load and budget. We'll reply within one business day with an honest read on the design, and on whether we're the right team to build it.