How Bella OpenAPI Implements Distributed Rate Limiting Using Redis and Lua Scripts
Bella OpenAPI enforces cluster-wide rate limits by executing atomic Lua scripts inside Redis, guaranteeing consistent throttling across multiple server nodes without external coordination services.
Bella OpenAPI is an open-source AI gateway that manages high-volume API traffic through sophisticated rate limiting mechanisms. The system implements distributed rate limiting using Redis and Lua scripts to maintain a single source of truth across the entire cluster. By moving the limiting logic from application code into Redis atomic operations, Bella eliminates race conditions while supporting multiple algorithms including per-minute, per-second, and concurrent request limits.
Architecture Overview
Java Layer: LimiterManager
The Java side handles key construction and parameter preparation. The LimiterManager class located in api/server/src/main/java/com/ke/bella/openapi/protocol/limiter/LimiterManager.java determines the appropriate Redis keys and arguments for each request. It delegates execution to LuaScriptExecutor, which manages script caching and EVALSHA optimization.
Redis Layer: Atomic Lua Execution
All rate limiting decisions happen inside Redis through purpose-built Lua scripts stored in api/server/src/main/resources/lua/limiter/. Because Redis executes Lua scripts atomically, the check-and-increment operations are inherently race-condition free, even under heavy distributed load.
The Four Rate Limiting Algorithms
Bella implements four distinct limiting strategies, each optimized for specific traffic patterns:
RPM Limiter: Sliding Window with Sorted Sets
The requests-per-minute limiter tracks API usage using Redis sorted sets. It stores request timestamps in bella-openapi-limiter-rpm:{entity}:{ak} and maintains a counter in bella-openapi-limiter-rpm-count:{entity}:{ak}. The rpm.lua script uses ZADD, ZCOUNT, and ZREMRANGEBYSCORE to maintain a rolling 60-second window, removing expired entries and counting current usage without scanning the entire set.
QPS Limiter: Segmented Hash Fields
For queries-per-second throttling, Bella splits each second into configurable segments (e.g., 100ms intervals). The qps.lua script stores counts in hash fields within bella-openapi-limiter-qps:{ak}. By aggregating recent segments, it achieves fine-grained per-second control while minimizing memory overhead compared to tracking individual timestamps.
Concurrent Request Limiter: Bucket Counters
To enforce maximum simultaneous operations, the concurrent limiter uses time-bucketed counters stored in bella-openapi-limiter-concurrent:{entity}:{ak}. The concurrent.lua script atomically increments active request counts at the start of processing and decrements them upon completion, automatically cleaning up empty buckets to prevent memory leaks.
Channel-Level RPM: Segmented Sliding Window
The channel limiter applies a segmented sliding window algorithm across the entire cluster for model-specific throttling. The channel_rpm.lua script divides the 60-second window into slices (e.g., 10 seconds each) stored as hash fields. It weights each slice's contribution based on temporal overlap with the current window, providing accurate rate limiting without storing every individual request.
Execution Flow
The request lifecycle follows four distinct steps:
-
Key Preparation:
LimiterManagerbuilds Redis keys using the API key, entity (model), and operation type, then prepares parameters including timestamps and request IDs. -
Script Execution: The manager calls
executor.execute(scriptPath, ScriptType.limiter, keys, params).LuaScriptExecutorloads the script from the classpath and optimizes execution usingEVALSHA(falling back toEVALon first run). -
Atomic Decision: The Lua script runs atomically inside Redis, calculating current usage against limits and conditionally incrementing counters.
-
Result Processing: The script returns a tuple
[allowed, current, remaining, message]. Java interprets this result, throwing aRateLimitExceptionif the request exceeds limits.
Code Implementation
Java Integration Example
Controllers interact with the limiter through LimiterManager:
LimiterManager limiter = applicationContext.getBean(LimiterManager.class);
String akCode = processData.getAkCode();
String entity = processData.getModel();
String requestId = processData.getRequestId();
// Record RPM count
limiter.incrementRequestCountPerMinute(
akCode, entity, requestId, DateTimeUtils.getCurrentSeconds()
);
// Increment concurrent counter at request start
limiter.incrementConcurrentCount(akCode, entity);
// Process AI provider call...
// Decrement after completion
limiter.decrementConcurrentCount(akCode, entity);
The getCurrentRequests method supports routing decisions:
long currentLoad = limiter.getCurrentRequests(entityCode);
Lua Script: Channel RPM Algorithm
The channel_rpm.lua script demonstrates the segmented window calculation:
local key = KEYS[1] -- hash key: "bella-channel-rpm:123"
local limit = tonumber(ARGV[1]) -- max requests: 100
local window = tonumber(ARGV[2]) -- 60 seconds
local seg = tonumber(ARGV[3]) -- slice size: 10 seconds
local cost = tonumber(ARGV[4]) -- request cost: 1
local now = tonumber(ARGV[5]) -- current timestamp
local curSeg = math.floor(now / seg)
local segCnt = math.ceil(window / seg)
-- Clean expired segments
for _, field in ipairs(redis.call('HKEYS', key)) do
if tonumber(field) < curSeg - segCnt then
redis.call('HDEL', key, field)
end
end
-- Calculate weighted usage
local total = 0
local winStart = now - window
for i = 0, segCnt-1 do
local id = curSeg - i
local v = tonumber(redis.call('HGET', key, id) or 0)
if v > 0 then
local segStart = id * seg
local segEnd = segStart + seg
local weight = (math.min(segEnd, now) - math.max(segStart, winStart)) / seg
total = total + v * weight
end
end
-- Enforce limit
if cost > 0 and total + cost > limit then
return {0, math.floor(total), math.floor(limit-total), "Rate limit exceeded"}
end
-- Record consumption
if cost > 0 then
redis.call('HINCRBY', key, curSeg, cost)
end
redis.call('EXPIRE', key, window*2)
return {1, math.floor(total+cost), math.floor(limit-total-cost), "OK"}
Summary
- Bella OpenAPI implements distributed rate limiting using Redis and Lua scripts to maintain atomic, cluster-wide throttling without external coordination.
- The
LimiterManagerJava class orchestrates four algorithms: RPM (sorted sets), QPS (segmented hashes), concurrent requests (bucket counters), and channel-level RPM (weighted sliding windows). - All limiting logic executes atomically inside Redis via scripts stored in
api/server/src/main/resources/lua/limiter/, guaranteeing consistency under high concurrency. - The architecture separates concerns: Java handles key construction and business logic, while Lua scripts handle the mathematical limiting calculations and state management.
Frequently Asked Questions
Why does Bella OpenAPI use Lua scripts instead of Redis transactions?
Lua scripts execute atomically as a single operation, eliminating race conditions between the read-check and write-increment steps. Redis transactions (MULTI/EXEC) can fail if watched keys change between commands, requiring complex retry logic. Bella's approach guarantees that rate limit calculations happen in isolation, ensuring accurate counts even when hundreds of nodes simultaneously access the same Redis keys.
How does the segmented sliding window algorithm improve accuracy over simple counters?
The segmented sliding window used in channel_rpm.lua divides time into slices (e.g., 10-second segments within a 60-second window) and weights each slice based on its temporal overlap with the current window. This provides smoother limiting than fixed windows (which allow traffic spikes at window boundaries) while consuming significantly less memory than tracking every individual request timestamp in a sorted set.
What happens if the Redis connection fails during a rate limit check?
If Redis becomes unavailable, the LuaScriptExecutor cannot execute the limiting scripts. According to the Bella OpenAPI source code, the Java layer would fail to obtain a limit check result, typically resulting in either a fallback to allow traffic (depending on configuration) or a service degradation. The concurrent request limiter specifically requires careful cleanup handling to prevent "ghost" counts if a node crashes between increment and decrement operations.
Can the rate limiting algorithms handle variable request costs?
Yes. The Lua scripts accept a cost parameter (shown as ARGV[4] in the channel RPM example) allowing different operations to consume different amounts of the rate limit budget. For instance, a large model inference might cost 10 units while a simple embedding request costs 1. The scripts atomically verify that current_usage + cost <= limit before incrementing counters.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →