Token Bucket Algorithm for Rate Limiting in Claude-Code-Telegram
The bot implements a classic token bucket algorithm in src/security/rate_limiter.py to enforce rate limits, using lazy timestamp-based refilling to balance burst tolerance with steady-state throttling.
The Claude-Code-Telegram bot protects itself from API overload and abuse by applying a token bucket algorithm to incoming Telegram updates. This lightweight, in-memory implementation lives in the security module and integrates directly into the request pipeline through dedicated middleware, ensuring fair resource allocation without database dependencies.
Core Token Bucket Implementation
The heart of the system is the RateLimiter class defined in src/security/rate_limiter.py. This class maintains the bucket state and implements the refill logic using a lazy evaluation pattern based on timestamps rather than background timers.
Bucket State Management
Each RateLimiter instance tracks three critical pieces of state:
capacity: The maximum number of tokens the bucket can hold, determining burst tolerance_tokens: The current available token count (float, allowing fractional tokens)_last_refill: The UTC timestamp of the last refill operation
# src/security/rate_limiter.py
class RateLimiter:
def __init__(self, capacity: int, refill_rate: float):
self.capacity = capacity # max tokens
self.refill_rate = refill_rate # tokens per second
self._tokens = capacity
self._last_refill = datetime.now(tz=UTC)
Lazy Refill Mechanism
Instead of using a background coroutine, the implementation uses lazy refilling. The _refill() method calculates elapsed time since the last check and adds proportional tokens only when allow_request() is called. This approach eliminates the need for scheduled tasks or locks while maintaining mathematical precision.
def _refill(self) -> None:
now = datetime.now(tz=UTC)
elapsed = (now - self._last_refill).total_seconds()
added = elapsed * self.refill_rate
if added > 0:
self._tokens = min(self.capacity, self._tokens + added)
self._last_refill = now
Request Processing and Consumption
The allow_request() method implements the consumption logic. It first triggers the lazy refill, then checks if at least one token is available. If sufficient tokens exist, it decrements the count and permits the request; otherwise, it returns False, triggering rate limit enforcement.
def allow_request(self) -> bool:
self._refill()
if self._tokens >= 1:
self._tokens -= 1
return True
return False
Middleware Integration
The rate limiter integrates into the bot's request pipeline through RateLimitMiddleware in src/bot/middleware/rate_limit.py. This middleware wraps every incoming Telegram update and applies the rate limit check before reaching business logic handlers.
# src/bot/middleware/rate_limit.py
class RateLimitMiddleware:
def __init__(self, limiter: RateLimiter):
self.limiter = limiter
async def __call__(self, update: Update, next_handler):
if not self.limiter.allow_request():
await update.message.reply_text(
"⚠️ You're sending messages too fast. Please slow down."
)
return # drop the update
return await next_handler(update)
Configuration and Initialization
The system pulls configuration from environment variables and src/config/features.py, allowing operators to tune limits without code changes. The src/bot/orchestrator.py file wires everything together during bot initialization.
# src/bot/orchestrator.py (excerpt)
from src.security.rate_limiter import RateLimiter
from src.bot.middleware.rate_limit import RateLimitMiddleware
def build_bot():
limiter = RateLimiter(
capacity=int(os.getenv("RATE_LIMIT_CAPACITY", "20")),
refill_rate=float(os.getenv("RATE_LIMIT_RATE", "1.0")),
)
dispatcher = Dispatcher(...)
dispatcher.middleware.append(RateLimitMiddleware(limiter))
# … register other handlers …
return dispatcher
Default values provide immediate protection (20 tokens capacity, 1 token per second refill), while environment variables enable deployment-specific tuning for different traffic patterns.
Algorithm Mechanics
The token bucket algorithm operates through three distinct phases:
- Initialization: The bucket starts at full capacity, allowing immediate burst traffic up to the limit
- Consumption: Each request removes exactly one token; if the bucket empties, subsequent requests fail
- Refilling: Tokens regenerate based on elapsed time and the configured rate, never exceeding capacity
This design guarantees burst tolerance (users can send up to capacity requests instantly) while enforcing steady-state throttling (sustained traffic cannot exceed refill_rate requests per second). The in-memory implementation avoids database latency, making each check computationally cheap with O(1) complexity.
Summary
- The
RateLimiterclass insrc/security/rate_limiter.pyimplements a lazy token bucket using timestamps to track refill state - Lazy refilling calculates token additions on-demand rather than using background threads, reducing complexity and resource overhead
- The middleware in
src/bot/middleware/rate_limit.pyintercepts all updates and rejects requests when the bucket is empty - Configuration through environment variables in
src/bot/orchestrator.pyallows runtime tuning ofcapacityandrefill_rate - The algorithm provides burst tolerance up to the capacity limit while enforcing a maximum sustained request rate
Frequently Asked Questions
How does the token bucket algorithm handle burst traffic?
The token bucket algorithm allows bursts up to the configured capacity because the bucket initializes full and refills continuously. A user can consume all available tokens immediately (sending up to 20 messages instantly with default settings), but must then wait for the refill rate to regenerate tokens before making additional requests.
Why does the implementation use lazy refilling instead of a background timer?
The RateLimiter uses lazy refilling (calculating elapsed time only during allow_request() calls) rather than asynchronous background tasks to eliminate the need for locks, scheduled coroutines, or thread management. This approach reduces resource overhead and eliminates race conditions while maintaining mathematical accuracy in token calculations.
What happens when a user exceeds the rate limit?
When the bucket empties, the allow_request() method returns False, triggering the middleware in src/bot/middleware/rate_limit.py to send a warning message ("⚠️ You're sending messages too fast...") and abort processing before the request reaches the bot's handlers. The update is dropped without consuming additional resources.
Can rate limits be adjusted without restarting the bot?
The current implementation initializes limits from environment variables in src/bot/orchestrator.py at startup. While changing these values requires a restart, the modular design allows extending the bot with a runtime configuration command (such as /set_rate_limit) that could instantiate a new RateLimiter with updated parameters and hot-swap the middleware reference.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →