# Token Bucket Algorithm for Rate Limiting in Claude-Code-Telegram

> Learn how the token bucket algorithm, implemented in Python, effectively rate limits requests to the Claude Code Telegram bot balancing bursts and steady traffic.

- Repository: [Richard A/claude-code-telegram](https://github.com/richardatct/claude-code-telegram)
- Tags: deep-dive
- Published: 2026-02-19

---

**The bot implements a classic token bucket algorithm in [`src/security/rate_limiter.py`](https://github.com/RichardAtCT/claude-code-telegram/blob/main/src/security/rate_limiter.py) to enforce rate limits, using lazy timestamp-based refilling to balance burst tolerance with steady-state throttling.**

The Claude-Code-Telegram bot protects itself from API overload and abuse by applying a token bucket algorithm to incoming Telegram updates. This lightweight, in-memory implementation lives in the security module and integrates directly into the request pipeline through dedicated middleware, ensuring fair resource allocation without database dependencies.

## Core Token Bucket Implementation

The heart of the system is the `RateLimiter` class defined in [`src/security/rate_limiter.py`](https://github.com/RichardAtCT/claude-code-telegram/blob/main/src/security/rate_limiter.py). This class maintains the bucket state and implements the refill logic using a lazy evaluation pattern based on timestamps rather than background timers.

### Bucket State Management

Each `RateLimiter` instance tracks three critical pieces of state:

- **`capacity`**: The maximum number of tokens the bucket can hold, determining burst tolerance
- **`_tokens`**: The current available token count (float, allowing fractional tokens)
- **`_last_refill`**: The UTC timestamp of the last refill operation

```python

# src/security/rate_limiter.py

class RateLimiter:
    def __init__(self, capacity: int, refill_rate: float):
        self.capacity = capacity                 # max tokens

        self.refill_rate = refill_rate           # tokens per second

        self._tokens = capacity
        self._last_refill = datetime.now(tz=UTC)

```

### Lazy Refill Mechanism

Instead of using a background coroutine, the implementation uses **lazy refilling**. The `_refill()` method calculates elapsed time since the last check and adds proportional tokens only when `allow_request()` is called. This approach eliminates the need for scheduled tasks or locks while maintaining mathematical precision.

```python
    def _refill(self) -> None:
        now = datetime.now(tz=UTC)
        elapsed = (now - self._last_refill).total_seconds()
        added = elapsed * self.refill_rate
        if added > 0:
            self._tokens = min(self.capacity, self._tokens + added)
            self._last_refill = now

```

## Request Processing and Consumption

The `allow_request()` method implements the consumption logic. It first triggers the lazy refill, then checks if at least one token is available. If sufficient tokens exist, it decrements the count and permits the request; otherwise, it returns `False`, triggering rate limit enforcement.

```python
    def allow_request(self) -> bool:
        self._refill()
        if self._tokens >= 1:
            self._tokens -= 1
            return True
        return False

```

## Middleware Integration

The rate limiter integrates into the bot's request pipeline through `RateLimitMiddleware` in [`src/bot/middleware/rate_limit.py`](https://github.com/RichardAtCT/claude-code-telegram/blob/main/src/bot/middleware/rate_limit.py). This middleware wraps every incoming Telegram update and applies the rate limit check before reaching business logic handlers.

```python

# src/bot/middleware/rate_limit.py

class RateLimitMiddleware:
    def __init__(self, limiter: RateLimiter):
        self.limiter = limiter

    async def __call__(self, update: Update, next_handler):
        if not self.limiter.allow_request():
            await update.message.reply_text(
                "⚠️ You're sending messages too fast. Please slow down."
            )
            return  # drop the update

        return await next_handler(update)

```

## Configuration and Initialization

The system pulls configuration from environment variables and [`src/config/features.py`](https://github.com/RichardAtCT/claude-code-telegram/blob/main/src/config/features.py), allowing operators to tune limits without code changes. The [`src/bot/orchestrator.py`](https://github.com/RichardAtCT/claude-code-telegram/blob/main/src/bot/orchestrator.py) file wires everything together during bot initialization.

```python

# src/bot/orchestrator.py (excerpt)

from src.security.rate_limiter import RateLimiter
from src.bot.middleware.rate_limit import RateLimitMiddleware

def build_bot():
    limiter = RateLimiter(
        capacity=int(os.getenv("RATE_LIMIT_CAPACITY", "20")),
        refill_rate=float(os.getenv("RATE_LIMIT_RATE", "1.0")),
    )
    dispatcher = Dispatcher(...)
    dispatcher.middleware.append(RateLimitMiddleware(limiter))
    # … register other handlers …

    return dispatcher

```

Default values provide immediate protection (20 tokens capacity, 1 token per second refill), while environment variables enable deployment-specific tuning for different traffic patterns.

## Algorithm Mechanics

The token bucket algorithm operates through three distinct phases:

1. **Initialization**: The bucket starts at full capacity, allowing immediate burst traffic up to the limit
2. **Consumption**: Each request removes exactly one token; if the bucket empties, subsequent requests fail
3. **Refilling**: Tokens regenerate based on elapsed time and the configured rate, never exceeding capacity

This design guarantees **burst tolerance** (users can send up to `capacity` requests instantly) while enforcing **steady-state throttling** (sustained traffic cannot exceed `refill_rate` requests per second). The in-memory implementation avoids database latency, making each check computationally cheap with O(1) complexity.

## Summary

- The `RateLimiter` class in [`src/security/rate_limiter.py`](https://github.com/RichardAtCT/claude-code-telegram/blob/main/src/security/rate_limiter.py) implements a lazy token bucket using timestamps to track refill state
- **Lazy refilling** calculates token additions on-demand rather than using background threads, reducing complexity and resource overhead
- The middleware in [`src/bot/middleware/rate_limit.py`](https://github.com/RichardAtCT/claude-code-telegram/blob/main/src/bot/middleware/rate_limit.py) intercepts all updates and rejects requests when the bucket is empty
- Configuration through environment variables in [`src/bot/orchestrator.py`](https://github.com/RichardAtCT/claude-code-telegram/blob/main/src/bot/orchestrator.py) allows runtime tuning of `capacity` and `refill_rate`
- The algorithm provides burst tolerance up to the capacity limit while enforcing a maximum sustained request rate

## Frequently Asked Questions

### How does the token bucket algorithm handle burst traffic?

The token bucket algorithm allows bursts up to the configured `capacity` because the bucket initializes full and refills continuously. A user can consume all available tokens immediately (sending up to 20 messages instantly with default settings), but must then wait for the refill rate to regenerate tokens before making additional requests.

### Why does the implementation use lazy refilling instead of a background timer?

The `RateLimiter` uses **lazy refilling** (calculating elapsed time only during `allow_request()` calls) rather than asynchronous background tasks to eliminate the need for locks, scheduled coroutines, or thread management. This approach reduces resource overhead and eliminates race conditions while maintaining mathematical accuracy in token calculations.

### What happens when a user exceeds the rate limit?

When the bucket empties, the `allow_request()` method returns `False`, triggering the middleware in [`src/bot/middleware/rate_limit.py`](https://github.com/RichardAtCT/claude-code-telegram/blob/main/src/bot/middleware/rate_limit.py) to send a warning message ("⚠️ You're sending messages too fast...") and abort processing before the request reaches the bot's handlers. The update is dropped without consuming additional resources.

### Can rate limits be adjusted without restarting the bot?

The current implementation initializes limits from environment variables in [`src/bot/orchestrator.py`](https://github.com/RichardAtCT/claude-code-telegram/blob/main/src/bot/orchestrator.py) at startup. While changing these values requires a restart, the modular design allows extending the bot with a runtime configuration command (such as `/set_rate_limit`) that could instantiate a new `RateLimiter` with updated parameters and hot-swap the middleware reference.