# How LiteLLM Handles Rate Limiting and Cooldown Periods: A Complete Technical Guide

> Explore how LiteLLM manages rate limiting and cooldown periods with TPM/RPM limits, priority reservations, and automatic backend rotation. Protect your inference workloads effectively.

- Repository: [Berri AI/litellm](https://github.com/BerriAI/litellm)
- Tags: how-to-guide
- Published: 2026-03-26

---

**LiteLLM protects inference workloads through three integrated mechanisms: per-deployment TPM/RPM limits enforced via pre-call checks, saturation-aware priority reservations for global capacity management, and automatic cooldown periods that remove unhealthy backends from rotation.**

LiteLLM, the popular open-source LLM gateway by BerriAI, implements sophisticated **rate limiting and cooldown period** management to prevent individual model deployments from overwhelming upstream providers while maintaining high availability. The architecture combines model-level token and request throttling with dynamic, priority-aware global limits and automated failure isolation. This article examines the implementation details found in the `litellm` source code, including specific file paths and configuration patterns used in production environments.

## Model-Level TPM and RPM Enforcement

When `enforce_model_rate_limits=True` is configured, LiteLLM instantiates the **`ModelRateLimitingCheck`** class from [`litellm/router_utils/pre_call_checks/model_rate_limit_check.py`](https://github.com/BerriAI/litellm/blob/main/litellm/router_utils/pre_call_checks/model_rate_limit_check.py) to intercept requests before they reach the provider.

### Deployment-Specific Limit Checks

The pre-call check extracts TPM (tokens-per-minute) and RPM (requests-per-minute) limits from the deployment configuration—searching top-level keys, `litellm_params`, or `model_info` dictionaries (lines 52-78). For each request, LiteLLM constructs cache keys scoped to the model ID, deployment name, and current minute using the format `model_id:deployment_name:tpm:HH-MM` (lines 81-89).

**TPM enforcement** operates via local cache reads only. If the stored token count meets or exceeds the limit, the system immediately raises a `RateLimitError` (HTTP 429) without incrementing counters (lines 112-122). **RPM enforcement** uses `DualCache.increment_cache` for atomic increments; if the new count exceeds the configured limit, the request is rejected (lines 132-140).

Successful requests update counters asynchronously via `async_log_success_event`, which adds actual token usage to the TPM counter (lines 46-84).

## Dynamic Priority-Aware Rate Limiting

The **`DynamicRateLimiter`** in [`litellm/proxy/hooks/dynamic_rate_limiter_v3.py`](https://github.com/BerriAI/litellm/blob/main/litellm/proxy/hooks/dynamic_rate_limiter_v3.py) provides global capacity management with saturation awareness and priority reservations. This enterprise feature operates as a proxy hook (`async_pre_call_hook`) using a three-phase validation strategy.

### Three-Phase Atomic Validation

**Phase 1: Read-Only Saturation Check**  
The `_check_model_saturation` method queries Redis-backed counters to calculate a saturation ratio (0 = empty, 1 = at capacity) without mutating state (lines 30-96).

**Phase 2: Priority Evaluation**  
When saturation exceeds the configured `saturation_threshold` (default requires Enterprise configuration), the system enforces priority-based limits. High-priority callers receive reserved capacity percentages defined in `priority_reservation` mappings (lines 40-45).

**Phase 3: Atomic Increment**  
Using Lua-scripted operations, the limiter increments model-wide counters always, and priority counters only if the request passes all limits. This separation prevents "phantom usage" where blocked requests would still consume capacity (lines 20-31, 63-84).

## Automatic Cooldown Periods for Failing Deployments

LiteLLM uses the **`CooldownCache`** mechanism in [`litellm/router_utils/cooldown_cache.py`](https://github.com/BerriAI/litellm/blob/main/litellm/router_utils/cooldown_cache.py) to temporarily remove unhealthy deployments from the routing pool after repeated failures.

### CooldownCache Implementation

When a deployment fails more than `allowed_fails` times (default 1), the router calls `add_deployment_to_cooldown` with a key formatted as `deployment:<model_id>:cooldown` (line 9). The cached value contains:

- **Masked exception messages** processed by `SensitiveDataMasker` to prevent secret leakage (lines 36-41)
- HTTP status codes
- Unix timestamps
- Configurable TTL (default 60 seconds via `DEFAULT_COOLDOWN_TIME_SECONDS`)

During routing, `_filter_cooldown_deployments` excludes any deployment with an active cooldown entry, ensuring traffic routes only to healthy backends.

### Configuration Parameters

| Setting | Description | Default |
|---------|-------------|---------|
| `allowed_fails` | Number of consecutive failures before cooldown | `1` |
| `cooldown_time` | Duration in seconds to withhold traffic | `60` |
| `enforce_model_rate_limits` | Enable TPM/RPM pre-call checks | `False` |

## Practical Configuration Examples

### Enabling Model-Level Rate Limits

```python
from litellm import Router

router = Router(
    model_list=[
        {
            "model_name": "openai-gpt-4",
            "litellm_params": {"model": "gpt-4"},
            "tpm": 300_000,            # 300k tokens per minute

            "rpm": 60,                 # 60 requests per minute

        },
    ],
    enforce_model_rate_limits=True,   # Activate pre-call checks

    allowed_fails=2,                  # Cooldown after 2 failures

    cooldown_time=120,                # 2-minute cooldown period

)

```

### Configuring Priority-Aware Dynamic Limits

```yaml

# config.yaml

litellm_settings:
  dynamic_rate_limiter: true
  
  priority_reservation:
    high: 0.7      # Reserve 70% for high priority when saturated

    medium: 0.2
    low: 0.1
  
  saturation_threshold: 0.8
  saturation_check_cache_ttl: 5

```

### Automatic Cooldown Configuration

```python
router = Router(
    model_list=[{
        "model_name": "azure-gpt-35",
        "litellm_params": {
            "model": "azure/deployment-1",
            "api_key": "<key>",
            "api_base": "https://example.openai.azure.com/",
        },
    }],
    allowed_fails=1,          # Immediate cooldown on first failure

    cooldown_time=90,         # 90-second exclusion from rotation

)

```

On failure, the system stores a `CooldownCacheValue` with masked error details and a 90-second TTL, automatically reintegrating the deployment once healthy.

## Summary

- **Per-deployment protection**: The `ModelRateLimitingCheck` class enforces TPM/RPM limits using minute-scoped cache keys before requests leave the proxy, preventing individual providers from receiving excessive traffic.
- **Global capacity management**: The dynamic rate limiter v3 implements three-phase atomic checks with priority reservations, ensuring high-priority workloads receive guaranteed capacity during saturation.
- **Failure isolation**: The `CooldownCache` automatically removes failing deployments using masked error tracking and configurable TTLs, preventing cascade failures across the model pool.
- **Flexible configuration**: All mechanisms support granular tuning via Router constructor parameters or `litellm_settings` in YAML configuration files.

## Frequently Asked Questions

### How does LiteLLM track rate limits for individual model deployments?

LiteLLM tracks rate limits using scoped cache keys in the format `model_id:deployment_name:tpm:HH-MM` or `model_id:deployment_name:rpm:HH-MM`. TPM counters are read locally before incrementing, while RPM counters use `DualCache.increment_cache` for atomic updates. The system checks these counters in the `pre_call_check` method of `ModelRateLimitingCheck` before forwarding requests to providers.

### What happens when a deployment exceeds its configured rate limits?

When limits are exceeded, LiteLLM raises a `RateLimitError` with HTTP status 429 immediately during the pre-call phase. For TPM violations, this occurs after reading the local cache (lines 112-122 in [`model_rate_limit_check.py`](https://github.com/BerriAI/litellm/blob/main/model_rate_limit_check.py)). For RPM violations, the error triggers after the atomic increment reveals the count exceeds the threshold (lines 132-140). The request never reaches the upstream provider.

### How does LiteLLM prioritize requests when the system approaches capacity?

The dynamic rate limiter v3 calculates a saturation ratio by querying current usage counters. When this ratio exceeds `saturation_threshold` (e.g., 0.8), the system enforces priority reservations configured in `priority_reservation`. High-priority requests proceed if capacity exists within their reserved percentage, while lower-priority requests may be rate-limited. The three-phase atomic increment ensures accurate capacity tracking without phantom usage.

### Can I customize how long a deployment stays in cooldown after failures?

Yes, configure the `cooldown_time` parameter in the Router constructor (in seconds) or set the `LITELLM_COOLDOWN_TIME` environment variable. The default is 60 seconds (`DEFAULT_COOLDOWN_TIME_SECONDS`). You can also adjust `allowed_fails` to control how many consecutive errors trigger cooldown. The `CooldownCache` stores entries with your specified TTL and automatically expires them, reintegrating the deployment into the routing pool.