# How the Notification Service Handles Message Delivery with Per-Chat Rate Limiting

> Discover how the NotificationService handles message delivery with per-chat rate limiting using asyncio sleep and last send timestamps to enforce a 1.1-second interval between messages.

- Repository: [Richard A/claude-code-telegram](https://github.com/richardatct/claude-code-telegram)
- Tags: internals
- Published: 2026-02-20

---

**The NotificationService enforces a strict 1.1-second interval between messages to the same Telegram chat by tracking last send timestamps in a dictionary and using `asyncio.sleep` to delay delivery when necessary.**

In the `RichardAtCT/claude-code-telegram` repository, the **NotificationService** bridges Claude-generated agent responses to Telegram chats while respecting Telegram's per-chat rate limits. This component works alongside the **EventBus** and **AgentResponseEvent** types to ensure no single chat receives more than one message per second, allowing higher global throughput without violating API constraints.

## Core Architecture of the Notification Service

The service operates as an asynchronous mediator between event production and Telegram API consumption. It decouples fast response generation from slower, rate-limited message delivery through internal queueing mechanisms.

### Event Subscription and Queueing

When `register()` is called, the service subscribes to `AgentResponseEvent` on the central `EventBus`. According to the source code in [`src/events/bus.py`](https://github.com/RichardAtCT/claude-code-telegram/blob/main/src/events/bus.py), every time the Claude integration finishes a response, an `AgentResponseEvent` is published and routed to the notification service.

The `handle_response()` coroutine validates the event type and pushes it onto an internal `asyncio.Queue` named `_send_queue`. This queue decouples event production from the slower sending path, preventing backpressure when Telegram's API throttles delivery.

### The Processing Loop

The `start()` method spawns a private background task `_process_send_queue()` that runs continuously until explicitly stopped. As implemented in [`src/notifications/service.py`](https://github.com/RichardAtCT/claude-code-telegram/blob/main/src/notifications/service.py), the loop pulls events from the queue with a 1-second timeout, ensuring responsive shutdown while maintaining high throughput for active message streams.

## Per-Chat Rate Limiting Implementation

The service implements strict per-chat throttling to comply with Telegram's approximate 1 message per second limit per chat. This prevents API bans while allowing concurrent messaging to multiple chats.

### Tracking Last Send Timestamps

The service maintains a private dictionary `_last_send_per_chat` that maps `chat_id` integers to timestamp floats representing when the last message was sent to that specific chat. This state persists for the lifetime of the service instance, enabling accurate rate calculations across multiple message batches.

### Calculating Wait Times

Inside `_rate_limited_send()` in [`src/notifications/service.py`](https://github.com/RichardAtCT/claude-code-telegram/blob/main/src/notifications/service.py), the service calculates the required delay using the constant `SEND_INTERVAL_SECONDS` set to **1.1 seconds**. The algorithm computes:

```python
wait_time = SEND_INTERVAL_SECONDS - (now - last_send)

```

Where `now` is retrieved from `asyncio.get_event_loop().time()`. If `wait_time` is positive, the coroutine executes `await asyncio.sleep(wait_time)`, guaranteeing at least a 1.1-second gap between consecutive sends to the same `chat_id`. This conservative 1.1-second buffer accounts for network latency and ensures compliance even under high load.

### Handling Multi-Chunk Messages

Telegram caps single messages at 4096 characters. The `_split_message()` function breaks longer texts at paragraph, newline, or space boundaries as defined in [`src/notifications/service.py`](https://github.com/RichardAtCT/claude-code-telegram/blob/main/src/notifications/service.py) lines 34-58.

After each chunk transmits, the timestamp for that chat updates to the current loop time. If multiple chunks exist, the service invokes another `await asyncio.sleep(SEND_INTERVAL_SECONDS)` between chunks, ensuring the per-chat rate limit applies even when sending lengthy Claude responses as sequential messages.

## Message Delivery Workflow

The delivery pipeline resolves recipients, segments content, and handles errors without breaking the processing loop.

### Chat Resolution

For each `AgentResponseEvent`, the `_resolve_chat_ids()` method determines target recipients. As shown in [`src/notifications/service.py`](https://github.com/RichardAtCT/claude-code-telegram/blob/main/src/notifications/service.py) lines 86-90, it returns either the specific `chat_id` embedded in the event or falls back to the service-wide `default_chat_ids` configured during initialization. This supports both targeted replies and broadcast scenarios.

### Message Splitting Logic

Long messages undergo intelligent fragmentation in `_split_message()`. The algorithm prioritizes splitting at paragraph breaks, then newlines, then spaces, ensuring readable chunks that respect Telegram's 4096-character hard limit. Each chunk passes through the same rate-limiting pipeline as standalone messages.

### Error Handling

The actual delivery uses `bot.send_message()` with optional HTML parsing based on `event.parse_mode`. Errors from the Telegram API are caught and logged without breaking the processing loop, ensuring that one failed delivery does not block subsequent messages to other chats or future retries to the same chat.

## Lifecycle Management

The service provides clean startup and shutdown semantics. The `start()` method initializes the background processor, while `stop()` cancels the background task and ensures the queue drains properly. As implemented in [`src/notifications/service.py`](https://github.com/RichardAtCT/claude-code-telegram/blob/main/src/notifications/service.py) lines 53-64, this prevents message loss during application shutdown and guarantees all pending sends complete before the event loop closes.

## Implementation Example

```python
from telegram import Bot
from src.events.bus import EventBus
from src.notifications.service import NotificationService

# Initialize core objects

bot = Bot(token="YOUR_TELEGRAM_BOT_TOKEN")
bus = EventBus()

# Create the notification service with default chats

notif = NotificationService(
    event_bus=bus, 
    bot=bot, 
    default_chat_ids=[123456789]
)

# Register to listen for AgentResponseEvent

notif.register()

# Start the background processor

await notif.start()

# Publish events from anywhere in your codebase

from src.events.types import AgentResponseEvent
await bus.publish(
    AgentResponseEvent(
        chat_id=123456789,
        text="Your Claude answer here",
        parse_mode="HTML",
        originating_event_id="some-event-id",
    )
)

# Shutdown gracefully

await notif.stop()

```

## Summary

- The **NotificationService** uses an internal `asyncio.Queue` to decouple event production from rate-limited delivery.
- **Per-chat rate limiting** is enforced through a timestamp dictionary (`_last_send_per_chat`) and a fixed 1.1-second interval between messages to the same chat.
- **Message splitting** handles Telegram's 4096-character limit while maintaining the same throttling between chunks.
- The service supports both targeted chat IDs and default broadcast lists through `_resolve_chat_ids()`.
- **Error isolation** ensures Telegram API failures do not crash the processing loop or block other chats.

## Frequently Asked Questions

### How does the notification service prevent hitting Telegram's rate limits?

The service tracks the last send time for each chat ID in `_last_send_per_chat` and calculates a dynamic wait time before each send. If the time elapsed since the last message is less than 1.1 seconds, it sleeps for the remaining duration, ensuring no chat receives messages faster than the configured interval.

### Can the service handle multiple chats simultaneously?

Yes. Because rate limiting tracks timestamps per individual `chat_id`, the service can deliver messages to many chats concurrently. While one chat waits for its 1.1-second interval, the processing loop continues servicing other chats, maximizing global throughput while respecting per-chat constraints.

### What happens when a Claude response exceeds 4096 characters?

The `_split_message()` method breaks long text into smaller chunks at logical boundaries (paragraphs, newlines, then spaces). Each chunk passes through the normal `_rate_limited_send()` pipeline, with the 1.1-second delay enforced between consecutive chunks to the same chat, ensuring compliance even with lengthy outputs.

### How does the service handle Telegram API errors?

Errors during `bot.send_message()` are caught and logged without breaking the main processing loop in `_process_send_queue()`. This design ensures that a network failure or temporary API outage for one chat does not block message delivery to other chats or prevent retry attempts for subsequent events.