How to Implement Distributed Rate Limiting with Redis: Algorithms and Code Examples

Implement distributed rate limiting with Redis by using atomic Lua scripts to manage token buckets or sliding-window counters, ensuring consistent enforcement across multiple application nodes with minimal latency.

Distributed systems running across multiple nodes or data centers require centralized rate limiting to prevent abuse and ensure fair resource allocation. Using Redis as a fast, in-memory store allows you to enforce limits consistently while maintaining sub-millisecond latency for high-throughput applications. The liquidslr/system-design-notes repository provides detailed architectural guidance in 04. Rate Limiter/Readme.md covering deterministic algorithms and distributed considerations.

Choose Between Token Bucket and Sliding-Window Algorithms

The repository identifies two primary algorithms for Redis-backed rate limiting, each suited to different traffic patterns.

Token Bucket allows short bursts while maintaining a steady-state rate, making it ideal for APIs that experience traffic spikes. This algorithm stores the current token count and last refill timestamp in a Redis hash, calculating elapsed time to add tokens back to the bucket.

Sliding-Window Counter offers precise per-window limits by tracking individual request timestamps in a Redis sorted set. This approach provides better accuracy for strict rate limits but consumes more memory than token bucket implementations.

Store Counters in Redis with Expiring Keys

Regardless of the algorithm chosen, each client or API key requires a dedicated Redis key that automatically expires to prevent memory bloat.

  • Use keys formatted as rate:{clientId} or rl:{clientId} to namespace rate limit data
  • Set TTL equal to the rate limiting window so stale counters are reclaimed automatically
  • Store token bucket state as Redis hashes with tokens and ts fields
  • Store sliding-window data as Redis sorted sets where the score equals the request timestamp

Ensure Atomic Operations with Lua Scripts

In distributed environments, multiple workers may simultaneously attempt to decrement the same counter, creating race conditions. The repository emphasizes wrapping logic in Redis Lua scripts to guarantee atomic read-modify-write operations.

The Lua script executes entirely within the Redis server, eliminating network round-trips between the check and update operations. This approach prevents overshooting limits when concurrent requests hit different application instances.

Token Bucket Lua Script

For token bucket implementations, the script calculates elapsed time since the last request, refills tokens accordingly, decrements if available, and updates the timestamp:

const tokenBucketLua = `
local key = KEYS[1]
local capacity = tonumber(ARGV[1])
local refill = tonumber(ARGV[2])
local now = tonumber(ARGV[3])
local ttl = tonumber(ARGV[4])

local bucket = redis.call('HMGET', key, 'tokens', 'ts')
local tokens = tonumber(bucket[1]) or capacity
local last = tonumber(bucket[2]) or now

local elapsed = now - last
tokens = math.min(capacity, tokens + (elapsed * refill))
if tokens < 1 then
  return 0
end
tokens = tokens - 1
redis.call('HMSET', key, 'tokens', tokens, 'ts', now)
redis.call('EXPIRE', key, ttl)
return 1
`;

Sliding-Window Lua Script

For sliding-window counters, the script removes entries outside the current time window, adds the new request, and checks the cardinality:

local key = KEYS[1]
local now = tonumber(ARGV[1])
local period = tonumber(ARGV[2])
local limit = tonumber(ARGV[3])

redis.call('ZADD', key, now, now)
redis.call('ZREMRANGEBYSCORE', key, 0, now - period)
local count = redis.call('ZCARD', key)
redis.call('EXPIRE', key, period + 1)
if count > limit then
  return 0
else
  return 1
end

Load this script once using SCRIPT LOAD and invoke it with EVALSHA to minimize bandwidth. The script works atomically across all workers regardless of the programming language used.

Implement the Middleware Layer

According to the architecture described in 04. Rate Limiter/images/architecture.png, the typical request flow places Redis between the API gateway and your service.

Client requests first hit an API gateway or middleware layer that queries Redis before forwarding to the backend service. The middleware executes the appropriate Lua script using EVAL or EVALSHA (for cached scripts), returning an immediate allow/reject decision.

Node.js Implementation Example

Below is a complete Node.js implementation using the token bucket algorithm with Redis:

const redis = require('redis');
const client = redis.createClient({ url: 'redis://localhost:6379' });
await client.connect();

const tokenBucketLua = `
local key = KEYS[1]
local capacity = tonumber(ARGV[1])
local refill = tonumber(ARGV[2])
local now = tonumber(ARGV[3])
local ttl = tonumber(ARGV[4])

local bucket = redis.call('HMGET', key, 'tokens', 'ts')
local tokens = tonumber(bucket[1]) or capacity
local last = tonumber(bucket[2]) or now

local elapsed = now - last
tokens = math.min(capacity, tokens + (elapsed * refill))
if tokens < 1 then
  return 0
end
tokens = tokens - 1
redis.call('HMSET', key, 'tokens', tokens, 'ts', now)
redis.call('EXPIRE', key, ttl)
return 1
`;

async function allowRequest(clientId, capacity = 10, refill = 1, windowSec = 60) {
  const now = Math.floor(Date.now() / 1000);
  const result = await client.eval(tokenBucketLua, {
    keys: [`rate:${clientId}`],
    arguments: [capacity, refill, now, windowSec],
  });
  return result === 1;
}

The function returns true if the request is allowed or false if the client has exceeded their quota. The script updates the token count and timestamp in a single Redis call, guaranteeing correctness even when many instances call the same key concurrently.

Python Implementation Example

For sliding-window counters, use Redis pipelines in Python to batch operations:

import time
import redis

r = redis.Redis(host='localhost', port=6379, db=0)

def allow_request(client_id: str, limit: int = 100, period: int = 60) -> bool:
    """Return True if the request is allowed, False otherwise."""
    now = int(time.time())
    key = f"rl:{client_id}"
    pipeline = r.pipeline()
    pipeline.zadd(key, {now: now})
    pipeline.zremrangebyscore(key, 0, now - period)
    pipeline.zcard(key)
    pipeline.expire(key, period + 1)
    _, _, count, _ = pipeline.execute()
    return count <= limit

The sorted-set approach records each request timestamp, removes entries older than the defined period, and checks the cardinality to decide if the limit is exceeded.

Handle Multi-Data-Center Synchronization

When deploying across multiple data centers, consistency becomes critical. The "Distributed Environments" section in 04. Rate Limiter/Readme.md recommends several strategies:

  • Single global Redis cluster: Route all rate limit checks through a centralized cluster with replication
  • Write replication: If using multiple Redis instances, replicate counter writes between clusters
  • Sorted-set structures: Use Redis sorted sets with precise timestamps to maintain consistency across regions

Return Proper Client Feedback

When rejecting requests, always return HTTP 429 (Too Many Requests) with a Retry-After header indicating when the client should retry. This allows API consumers to implement exponential backoff strategies and improves the user experience during rate limit episodes.

Summary

  • Token bucket algorithms work best for burst-friendly limits using Redis hashes, while sliding-window counters provide stricter accuracy using sorted sets
  • Store rate limit data in expiring Redis keys (e.g., rate:{clientId}) to ensure automatic cleanup of stale counters
  • Wrap all read-modify-write operations in Lua scripts to guarantee atomicity across distributed application instances
  • Query Redis from an API gateway or middleware layer before forwarding requests to backend services
  • Return HTTP 429 status codes with Retry-After headers when limits are exceeded

Frequently Asked Questions

What is the difference between token bucket and sliding-window rate limiting?

Token bucket allows short bursts of traffic up to a maximum capacity while maintaining an average refill rate, making it suitable for APIs that experience occasional traffic spikes. Sliding-window counters track every request timestamp within the current window, providing stricter per-window limits but requiring more memory to store individual event records.

Why use Lua scripts for Redis rate limiting instead of standard GET/SET operations?

Standard GET/SET operations create race conditions in distributed environments where multiple application instances might simultaneously read the same counter value before either has incremented it. Lua scripts execute atomically on the Redis server, ensuring that the check-and-decrement logic runs as a single indivisible operation regardless of how many clients are connected.

How do I scale rate limiting across multiple data centers?

Deploy a single globally-available Redis cluster or use Redis replication to synchronize counter state between regional clusters. For eventually consistent scenarios, use CRDT-based Redis Enterprise features, or implement the sliding-window algorithm which naturally handles clock skew through timestamp-based expiration of old entries.

What Redis data structure should I use for sliding-window rate limiting?

Use Redis sorted sets (ZADD, ZREMRANGEBYSCORE, ZCARD) where the score represents the request timestamp. This structure allows efficient removal of entries outside the current time window and accurate counting of requests within the window, though it consumes more memory than the hash-based token bucket approach.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →