Skip to content

Free calculator · Engineering

Plan Redis capacity for API rate limits and caching

Enter your peak traffic, cacheable share, hit ratio and rate-limit policy. The planner shows the load left on your origin, the Redis throughput and shards you need, cache and limiter memory, and how one client is treated under fixed window, sliding window or token bucket limits.

How this is calculated

The planner works from your peak requests per second. The cache answers the cacheable share at your hit ratio, and the rest reaches the origin. Rate limiting adds Redis commands for every request, cache reads add one for every cacheable request, and memory comes from key and value sizes with Redis's per-key overhead, fragmentation, copies and headroom.

Step by step

  1. Origin load: peak requests per second × (1 − cacheable share × hit ratio). Compared with the origin capacity you enter, when you enter one.
  2. Redis commands per second: peak requests per second × commands per limiter check for the chosen algorithm, plus peak requests per second × cacheable share for cache reads.
  3. Shards: commands per second ÷ (commands per shard × (1 − headroom)), rounded up.
  4. Cache memory: hot objects × (key + value + per-key overhead) × (1 + fragmentation) × copies ÷ (1 − headroom).
  5. Rate-limit memory: clients × bytes per client. A fixed window or token bucket keeps one key per client, a sliding window counter two, and a sliding window log one key plus up to N entries.
  6. Token bucket: a client can send B + r × T requests in T seconds, and a client sending d requests a second (faster than r) empties the bucket after B ÷ (d − r) seconds.
  7. Fixed window: the worst burst across a window boundary is 2N, because a client can spend the whole limit at the end of one window and again at the start of the next.

Default assumptions

Assumptions marked adjustable can be changed in the calculator; the others are fixed parts of the model.

Default assumptions and their sources
AssumptionDefaultSources
Peak requests per secondadjustable2,000 a secondNone
Share of requests that are cacheableadjustable70%
Cache hit ratioadjustable85%None
Origin capacityadjustable1,000 a secondNone
Clients with a live limiter keyadjustable50,000 clientsNone
Window limit (N requests)adjustable100 requestsNone
Window length (W)adjustable60 secondsNone
Token bucket capacity (B)adjustable20 requests
Token refill rate (r)adjustable5 a second
One client's burst rate (d)adjustable10 a secondNone
Horizon for allowed requests (T)adjustable60 secondsNone
Bytes per limiter keyadjustable100 bytes
Bytes per sliding-log entryadjustable64 bytes
Redis commands per check: fixed windowadjustable2 commands
Redis commands per check: sliding window counteradjustable3 commands
Redis commands per check: sliding window logadjustable4 commands
Redis commands per check: token bucketadjustable3 commands
Sustained commands per second per Redis shardadjustable50,000 a second
Headroom kept freeadjustable25%
Hot objects in the cacheadjustable1,000,000 objectsNone
Average cache key sizeadjustable50 bytesNone
Average cached value sizeadjustable1,000 bytesNone
Redis overhead per cached keyadjustable60 bytes
Memory fragmentation over data sizeadjustable20%
Copies of the cache (primary plus replicas)adjustable2 copiesNone
Live keys per client: fixed window1 key
Live keys per client: sliding window counter2 keys
Live keys per client: token bucket1 key
Worst burst across a fixed-window boundary2× the limit

What this doesn’t model

  • Cache writes on misses, TTL expiry, eviction and key churn aren't counted in Redis commands or memory.
  • Rate-limit memory is one copy, before fragmentation, replicas and headroom; apply the same factors as the cache if the limiter shares its cluster.
  • Redis throughput varies with CPU, network, virtualisation, connection count, pipelining and the cost of Lua scripts. The shard count is a starting point for a load test, not a result of one.
  • Network bandwidth between clients, Redis and the origin isn't modelled.
  • Traffic is treated as a steady peak; real traffic arrives in bursts, and hit ratios fall after deploys and cache flushes.
  • The sliding window counter is an approximation; it can let slightly more or fewer requests through than the limit.

Sources

  1. Redis, Build 5 rate limiters with Redis: fixed window, sliding window, token bucket, and leaky bucket (26 Feb 2026). Accessed . Lua scripts per algorithm: fixed window INCR, EXPIRE and PTTL on one key; sliding counter two GETs and an INCR across two keys; sliding log ZREMRANGEBYSCORE, ZCARD, ZADD and EXPIRE on a sorted set; token bucket HGETALL, HSET and EXPIRE on one hash. Fixed windows allow up to twice the limit across a boundary.
  2. Redis, MEMORY USAGE. Accessed . Reports the bytes a key and its value take in RAM, including administrative overhead; an empty string key measures 56 bytes on Redis 7.2 (64-bit, jemalloc).
  3. Redis, Memory optimization. Accessed . Small sorted sets (128 entries or fewer by default) use a compact encoding; provision memory for peak usage; fragmentation ratio is RSS divided by memory in use.
  4. Redis, Redis benchmark. Accessed . Example runs: 180,180 SET/s on one key and 72,144 SET/s over 100,000 random keys, without pipelining; results depend on CPU, network, virtualisation and connection count.
  5. IETF, RFC 6585: Additional HTTP status codes (section 4, 429 Too Many Requests) (Apr 2012). Accessed . 429 means the user sent too many requests in a given amount of time; the response may include Retry-After.
  6. IETF HTTPAPI working group, RateLimit header fields for HTTP (draft-ietf-httpapi-ratelimit-headers-11) (23 May 2026). Accessed . An active Internet-Draft, not yet an RFC, defining the RateLimit-Policy and RateLimit fields.
  7. IETF, RFC 9111: HTTP caching (Jun 2022). Accessed . Which responses a cache may store and when a stored response is fresh enough to reuse.
  8. Stripe, Scaling your API with rate limiters (30 Mar 2017). Accessed . Stripe's request rate limiter uses the token bucket algorithm, implemented with Redis.
  9. Cloudflare, How we built rate limiting capable of scaling to millions of domains (7 Jun 2017). Accessed . Sliding window counter: previous count × (time left in the window ÷ window) + current count. Across 400 million requests, 0.003% were wrongly allowed or limited.

Last reviewed by the QuantmHill engineering team. Found an error?

Link to or cite this tool

Writing about this topic? Link to the calculator or cite it. Its method, defaults and sources are all on this page, so readers can check the numbers.

Embed this calculator

You can put this calculator on your own site for free. Paste the code below where it should appear. It loads the same calculator in a frame, with a link back to this page for the full method and sources.

The credit line links to this page with the anchor text “QuantmHill”. You may edit it, add rel="nofollow" or remove it — the calculator works the same either way. Add ?theme=light or ?theme=dark to the iframe address to fix its colour scheme; otherwise it follows the visitor's system setting.

Add this once per page, after the iframe, if you want the frame to grow and shrink with the calculator instead of using the fixed height above. It accepts messages from quantmhill.com only and resizes only the frame that sent them.

Frequently asked questions

Five things: the requests per second that still reach your origin after caching, the Redis commands per second that rate limiting and cache reads add, the number of Redis shards for that throughput, the memory to provision for the cache, and the memory the rate limiter's keys take. It also shows how much one client can send under the policy you enter.

It depends on how strict you need to be. A fixed window is simplest but lets a client send up to twice the limit across a window boundary. A sliding window counter smooths that out with two keys per client; Cloudflare found its approximation wrong for 0.003% of 400 million requests. A sliding window log is exact but stores every request. A token bucket allows a burst up to the bucket size and then the refill rate; Stripe uses one for its request rate limiter.

Each limiter check runs the commands in Redis's reference scripts: two for a fixed window, three for a sliding window counter or token bucket, and four for a sliding window log. Each cacheable request adds one cache read. You can change every count under Adjust assumptions.

The default of 50,000 commands a second sits below the 72,144 SET commands a second shown on Redis's benchmark page for one instance over 100,000 random keys without pipelining. Your figure depends on the CPU, network, virtualisation, number of connections and Lua scripts, so benchmark your own instance type and enter that.

It is a sizing estimate. Redis reports 56 bytes for an empty string key on a 64-bit build, and real keys add their name and value length, so the defaults are close to a typical small key. Run MEMORY USAGE on a few of your own keys and read mem_fragmentation_ratio from INFO memory, then enter those figures.

HTTP status 429 Too Many Requests, defined in RFC 6585, optionally with a Retry-After header saying how long to wait. The IETF's RateLimit header fields draft, still an Internet-Draft, describes headers that tell clients their remaining quota.

Want an engineer to check your numbers?

Send us your inputs and the decision you're weighing. We'll reply within one business day with an honest read on whether we can help.